In-vehicle call method and device, storage medium and electronic terminal

Through the time domain convolutional network model and adaptive filtering algorithm, the problems of incomplete echo cancellation and howling in the in-car call system are solved, the call quality and driving safety are improved, and it adapts to complex acoustic environments.

CN120748424AActive Publication Date: 2025-10-03AUTOMOBILE RES INST OF TSINGHUA UNIV IN SUZHOU XIANGCHENG
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511172456.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-10-03
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

The existing in-car communication system has problems such as incomplete echo cancellation and howling during two-way calls, which affects the call quality and even threatens driving safety. In addition, the existing neural network algorithm lacks accuracy in phase estimation and waveform reconstruction.

Method used

A time-domain convolutional network model combined with an adaptive filtering algorithm is used to extract user voice from the in-car voice. The residual network and depthwise separable convolution are used for feature extraction and masking value output, and the adaptive filtering algorithm is combined to perform howling suppression and voice enhancement.

Benefits of technology

It effectively eliminates echoes, suppresses howling, improves call quality, reduces driving safety risks, optimizes voice waveform accuracy and phase estimation, and adapts to the complex acoustic environment inside the car.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748424A_ABST
    Figure CN120748424A_ABST
Patent Text Reader

Abstract

The invention discloses an in-vehicle call method and device, a storage medium and an electronic terminal, and the method comprises the steps: obtaining a first voice signal of a sound source seat, and the first voice signal comprises a user voice signal and an environment voice signal; a time domain convolutional network model is created, the time domain convolutional network model comprises an encoder, a residual network and a decoder, and the time domain convolutional network model is used for extracting the user voice signal from the first voice signal to obtain a second voice signal; and performing howling suppression and voice enhancement on the second voice signal through an adaptive filtering algorithm to obtain a third voice signal, and sending the third voice signal to a loudspeaker of a non-sound-source seat. According to the invention, feature extraction and processing are optimized, the voice waveform precision and the phase estimation effect are improved, and the method adapts to a complex acoustic environment in a vehicle; the in-vehicle communication quality is greatly improved, a driver can communicate with passengers in the back row conveniently, distracting operation is not needed, and driving potential safety hazards are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle-mounted voice processing, and in particular to an in-vehicle communication method, device, storage medium and electronic terminal. Background Art

[0002] With the development of the automotive industry, in-car communication systems are becoming increasingly important for improving the in-car communication experience and ensuring driving safety. They primarily create a high-quality communication environment for passengers, allowing drivers to conveniently communicate with rear-seat passengers, avoiding the potential safety hazards associated with turning or looking away. However, current in-car communication systems face numerous issues. Some systems lack support for simultaneous two-way communication, or are prone to incomplete echo cancellation and howling during two-way communication, resulting in poor call quality and even threatening driving safety.

[0003] In-car communication systems require the integration of microphones, speakers, voice processing, and onboard audio signal processing. Within the voice front-end, echo cancellation plays a crucial role in impacting call quality. Existing neural network-based echo cancellation algorithms often employ a time-frequency masking strategy, which has significant drawbacks. These include the need for a short-time Fourier transform (SFT), which negatively impacts echo cancellation. Furthermore, neural networks based on time-frequency training strategies often have phase estimation flaws when reconstructing the voice signal phase during echo cancellation, resulting in inaccurate reconstructed waveforms.

[0004] Therefore, how to accurately extract effective user voice from a mixed signal containing user voice and environmental voice, and further optimize call quality, has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The purpose of the present invention is to address the above-mentioned problems and provide a method, device, storage medium and electronic terminal for in-car communication.

[0006] The technical solution of the present invention is: a method for in-car communication, comprising the following steps: obtaining a first voice signal of a sound source seat, wherein the first voice signal includes a user voice signal and an environmental voice signal; creating a time domain convolutional network model, wherein the time domain convolutional network model includes an encoder, a residual network and a decoder, wherein the time domain convolutional network model is used to extract the user voice signal from the first voice signal to obtain a second voice signal; wherein the residual network includes normalization, N convolutional blocks 、 ,..., And activation function, the N convolution blocks are connected by residual paths, the convolution blocks The input is the convolution block The residual path output of , where i and N are both positive integers, ; The encoder is used to convert the first speech signal to obtain a feature space representation of the first speech signal; the residual network is used to extract features from the feature space representation and output a masking value of the user speech signal; the decoder is used to convert the masking value to obtain a second speech signal; howling suppression and speech enhancement are performed on the second speech signal through an adaptive filtering algorithm to obtain a third speech signal, and the third speech signal is sent to the speaker of the non-sound source seat.

[0007] As an improvement to the embodiment of the present invention, the convolution block The convolution is a depth-separable convolution; the expression of the depth-separable convolution is: , ,in, For the convolution block Input, and The sizes are and a convolution kernel of 1, is the convolution operation, It is a splicing operation.

[0008] As an improvement to the embodiment of the present invention, the incremental expansion factor is introduced into the residual network. , the convolution block Expansion factor .

[0009] As an improvement to an embodiment of the present invention, the “converting the first speech signal to obtain a feature space representation of the first speech signal” specifically includes: converting the discrete waveform of the first speech signal by the encoder to obtain a feature space representation of the first speech signal; the discrete waveform expression of the first speech signal is: , where M is the number of source signals of the first speech signal, is the source signal of the first speech signal, wherein j and M are both positive integers, .

[0010] As an improvement of the embodiment of the present invention, the “extracting features from the feature space representation and outputting the masking value of the user voice signal” specifically includes: extracting features from the feature space representation through the time domain convolutional network model and outputting M of the source signals The masking value of , ,in, is a constraint condition, H() is an optional activation function, is the basis function of the encoder.

[0011] As an improvement to the embodiment of the present invention, the step of “converting the masking value to obtain the second speech signal” specifically includes: reconstructing the masking value by a decoder to obtain the second speech signal, ,in, is the basis function of the decoder.

[0012] As an improvement to the embodiment of the present invention, the adaptive filtering algorithm is normalized least squares adaptive filtering.

[0013] To achieve one of the above-mentioned purposes of the invention, an embodiment of the present invention provides an in-car communication device, comprising the following modules: a signal acquisition module for acquiring a first voice signal of a sound source seat, wherein the first voice signal includes a user voice signal and an environmental voice signal; a model creation module for creating a time domain convolutional network model, wherein the time domain convolutional network model includes an encoder, a residual network and a decoder, wherein the time domain convolutional network model is used to extract the user voice signal from the first voice signal to obtain a second voice signal; the residual network includes normalization, N convolutional blocks 、 ,..., And activation function, the N convolution blocks are connected by residual paths, the convolution blocks The input is the convolution block The encoder is used to convert the first speech signal to obtain a feature space representation of the first speech signal; the residual network is used to extract features from the feature space representation and output a masking value of the user speech signal; the decoder is used to convert the masking value to obtain a second speech signal; a signal optimization module is used to perform howling suppression and speech enhancement on the second speech signal through an adaptive filtering algorithm to obtain a third speech signal, and send the third speech signal to the speaker of the non-sound source seat.

[0014] To achieve one of the above-mentioned objects of the invention, an embodiment of the present invention provides a storage medium storing program instructions, which, when executed, implement the in-vehicle call method as described in any one of the above items.

[0015] To achieve one of the above-mentioned objects of the invention, an embodiment of the present invention provides an electronic terminal, including a processor and a memory, wherein the memory stores program instructions, and the processor executes the program instructions to implement the in-vehicle call method as described in any one of the above items.

[0016] The in-car call method, device, storage medium and electronic terminal provided by the embodiments of the present invention have the following advantages: the present invention uses a time-domain convolutional network model to extract user voice from the in-car voice, thereby eliminating echo; suppresses howling and enhances voice through an adaptive filtering algorithm, thereby solving the problem of in-car call quality; facilitates communication between the driver and the rear passengers without distraction, reducing driving safety hazards; compared with traditional algorithms, it does not require short-time Fourier transform, optimizes feature extraction and processing, improves voice waveform accuracy and phase estimation effect, and adapts to the complex acoustic environment in the car. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 1 is a flow chart of the in-car communication method of the present invention; Figure 2 It is a schematic diagram of the structure of the time domain convolutional network model of the present invention; Figure 3 Schematic diagram of an application scenario of the in-car communication method of the present invention; Figure 4 It is a structural schematic diagram of the in-car communication device of the present invention; Figure 5 It is a structural schematic diagram of the electronic terminal of the present invention. DETAILED DESCRIPTION

[0018] The present invention will be described in detail below with reference to the specific embodiments shown in the accompanying drawings. However, these embodiments do not limit the present invention, and any structural, methodological, or functional changes made by those skilled in the art based on these embodiments are all within the scope of protection of the present invention.

[0019] If the present invention involves directions (for example, up, down, left, right, front, back, outside, inside, etc.) when describing, the directions involved need to be defined.

[0020] The scope of the embodiments herein includes the entire scope of the claims, and all available equivalents of the claims. Herein, the terms "first", "second", etc. are only used to distinguish one element from another, without requiring or implying any actual relationship or order between these elements. In fact, the first element can also be called the second element, and vice versa. Moreover, the terms "comprise", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that the structure, device or equipment including a series of elements includes not only those elements, but also includes other elements not clearly listed, or also includes elements inherent to such structure, device or equipment. In the absence of more restrictions, the elements limited by the statement "comprising one..." do not exclude the presence of other identical elements in the structure, device or equipment including the elements. Each embodiment is described in a progressive manner herein, and each embodiment focuses on the differences from other embodiments, and the same similar parts between the embodiments can be referred to each other.

[0021] The terms "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like used herein to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are intended only to facilitate the description of this document and simplify the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention. In the description herein, unless otherwise specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, they can be mechanical or electrical connections, or they can be internal connections between two elements, they can be directly connected, or they can be indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to the specific circumstances.

[0022] The first embodiment of the present invention provides a method for in-car communication. Figure 1 As shown, the following steps are included: Step 101: Acquire a first voice signal at a sound source seat, where the first voice signal includes a user voice signal and an environmental voice signal; In the car environment, such as Figure 2 As shown, the first microphone 201, the second microphone 202, the third microphone 203, and the fourth microphone 204 are provided in the vehicle to collect the voice signal in the vehicle; and the first speaker 301, the second speaker 302, the third speaker 303, and the fourth speaker 304 are provided in the vehicle to play the received voice signal. The voice signal collected by each microphone is subjected to Fourier transform to obtain the frequency domain expression of the voice signal collected by each microphone. ,in, , corresponding to microphones 201-204, The time domain speech signal collected by each microphone is used to extract the speech domain amplitude signal Each frequency domain speech signal is segmented according to a preset frequency domain sub-band, which may be: 80-250Hz, 250-500Hz, 500-1000Hz, 1000-2000Hz, 2000-4000Hz. For each sub-band, according to the formula Calculate the signal energy, where For the The microphone collects the signal at The energy brought by the child, The microphone signal is The frequency domain amplitude of the subband. Then calculate the weighted energy of the corresponding signal of each microphone ,in, For the Subband weighting coefficients. Compare the weighted energy of the 4 microphones The in-vehicle voice signal collected by the microphone with the highest energy value is the first voice signal. The seat corresponding to the microphone with the highest energy value is the sound source seat, and the seats corresponding to the other microphones are non-sound source seats. The first voice signal includes the user's voice and environmental voice signals. The environmental voice signals include environmental interference sounds such as speaker echo signals from non-sound source seats, in-vehicle device noise, external noise, and howling signals.

[0023] Step 102: Create a time-domain convolutional network model, which includes an encoder, a residual network, and a decoder. The time-domain convolutional network model is used to extract the user voice signal from the first voice signal to obtain a second voice signal. Here, the time-domain convolutional network model is a deep learning architecture designed specifically for processing time series signals. It effectively captures long-term and short-term dependencies in the signal through methods such as causal convolution, dilated convolution, and residual connection, and is particularly suitable for speech separation tasks.

[0024] The residual network includes normalization, N convolution blocks 、 ,..., and activation function, preferably, the residual network introduces an incremental expansion factor , the convolution block Expansion factor , the N convolution blocks are connected by residual paths, and the convolution blocks The input is the convolution block The residual path output of , where i and N are both positive integers, In this way, the residual network can exponentially expand its receptive field without increasing the number of parameters or computational complexity, enabling it to simultaneously capture both microscopic details and macroscopic structure in speech signals. The sparse connectivity of the dilated convolution reduces redundant computation between parameters, improving model training efficiency while mitigating the risk of overfitting, enabling the model to maintain good generalization capabilities even with limited training data. By connecting residual paths, the gradient diffusion problem caused by dilated convolution is effectively alleviated, ensuring smooth information transfer within the deep network and further optimizing the model's ability to represent speech signals.

[0025] In the present invention, the convolution block The convolution is a depth-separable convolution; the expression of the depth-separable convolution is: , ,in, For the convolution block Input, and The sizes are and a convolution kernel of 1, is the convolution operation, Here, depthwise separable convolution can significantly reduce the number of parameters and computational effort, lowering the computational overhead of the in-car communication system when processing voice signals. Furthermore, depthwise separable convolution finely extracts intra-channel features and flexibly integrates cross-channel features, adapting to the diverse characteristics of voice signals in the complex acoustic environment inside the car. This improves the model's ability to distinguish between user voice and ambient speech, enabling more efficient extraction of user voice.

[0026] During the specific implementation, a model architecture consisting of an encoder, a residual network and a decoder is constructed. After the encoder receives the first speech signal, it extracts features from the time domain signal through a one-dimensional convolution layer and converts it into a feature space representation; the residual network includes a normalization operation, N convolution blocks and an activation function, wherein the N convolution blocks are connected by a residual path, and the input of each convolution block is the residual path output of the previous convolution block. The residual network performs deep feature mining on the feature space representation and outputs a masking value for identifying the user's speech features; here, preferably, the activation function used in the present invention is PReLu. Compared with the ReLU activation function, PReLU introduces a learnable parameter and can adaptively adjust the slope of the negative half-axis, so that the network can better retain negative feature information when processing complex speech signals containing user speech and environmental noise, avoid information loss due to feature truncation, and effectively improve the model's ability to capture different speech signal features.

[0027] The decoder converts the masked values ​​back to the time domain through a deconvolution operation, generating a second voice signal containing only the user's voice. This time-domain convolutional network model effectively separates the user's voice from ambient noise, improving in-car call quality. To optimize model performance, mixed voice data from different in-car environments is collected for training. The network parameters are adjusted using a mean squared error loss function and a backpropagation algorithm to adapt to complex acoustic environments. The trained model can be deployed on an in-vehicle processing chip or in the cloud, enabling real-time processing of the first voice signal, providing the foundation for subsequent voice enhancement and howling suppression.

[0028] Step 103: Perform howling suppression and speech enhancement on the second speech signal through an adaptive filtering algorithm to obtain a third speech signal, and send the third speech signal to a speaker at a non-sound source seat. Preferably, the adaptive filtering algorithm is normalized least squares adaptive filtering.

[0029] In practice, a Fourier transform is performed on the current frame of the second speech signal to obtain a frequency domain signal. The ratio of the average power of the frame spectrum to the power and average power of each frequency point is calculated. The frequency points with a ratio greater than a preset threshold are marked as howling frequencies. A corresponding number of notch filter transfer functions are preset based on the number of howling frequencies. After being connected in series, howling suppression is performed on the frequency domain speech signal. Parameters such as the control bandwidth α and attenuation depth β in the transfer function can be assigned according to actual working conditions or prior information. Afterwards, the noise signal is estimated through the silent segment and the initial filter coefficients are given. The current frame speech signal is used as the target signal. Combined with the estimated noise signal, the normalized least squares adaptive filtering algorithm is used to further remove residual noise and obtain the third speech signal. Finally, the third speech signal is sent to the speaker in the non-sound source seat.

[0030] In this embodiment, “converting the first speech signal to obtain a feature space representation of the first speech signal” specifically includes: converting the discrete waveform of the first speech signal by the encoder to obtain a feature space representation of the first speech signal; the discrete waveform expression of the first speech signal is: , where M is the number of source signals of the first speech signal, is the source signal of the first speech signal, wherein j and M are both positive integers, .

[0031] In practice, the encoder can perform a sliding convolution operation on the discrete waveform through one-dimensional convolution to generate a feature map. The feature map is then normalized through normalization, and an activation function is used to introduce a nonlinear transformation to obtain a feature space representation. This feature space representation preserves the time-frequency and spatial characteristics of the first speech signal, providing structured input for the residual network and effectively supporting the extraction and separation of the user's speech signal.

[0032] In this embodiment, the “extracting features from the feature space representation and outputting the masking value of the user voice signal” specifically includes: extracting features from the feature space representation through the time domain convolutional network model and outputting M source signals The masking value of , ,in, is a constraint condition, H() is an optional activation function, is the basis function of the encoder.

[0033] In practice, through the residual network Convolution blocks perform layer-by-layer feature extraction to generate intermediate feature representations . Then you can use the mapping function Calculate each source signal The masking value of , where the mapping process satisfies the constraints To ensure energy conservation. The basis function here is As the feature conversion basis of the encoder, it is used to map the mask value back to the original signal space, and the mask value of the final output is Represents the first The probability distribution of the source signal is multiplied element by element with the feature space representation to separate the feature components corresponding to the user voice signal and realize the extraction of the user voice.

[0034] In this embodiment, the “converting the masking value to obtain the second speech signal” specifically includes: reconstructing the masking value by a decoder to obtain the second speech signal, ,in, is the basis function of the decoder.

[0035] In practice, the decoder first receives the M source signal mask values ​​output by the residual network , through the basis function Mapping it back to the time domain signal space can be done by linear combination operation Implementation, where Represents the basis function of the decoder, which defines the mapping relationship from feature space to time domain. To ensure the accuracy of the reconstructed signal, the basis function With encoder basis function Duality conditions must be met , to achieve lossless reconstruction of the speech signal. It is understandable that this mapping process converts the masking value representing the user's speech characteristics into a time-domain waveform, generating a second speech signal containing only the user's speech, thereby completing the task of extracting the user's speech from the first speech signal.

[0036] The second embodiment of the present invention provides an in-car communication device, such as Figure 4 As shown, it includes the following modules: A signal acquisition module 401 is configured to acquire a first voice signal at a sound source seat, wherein the first voice signal includes a user voice signal and an environmental voice signal; The model creation module 402 is used to create a time domain convolutional network model, which includes an encoder, a residual network and a decoder. The time domain convolutional network model is used to extract the user voice signal from the first voice signal to obtain a second voice signal; the residual network includes normalization, N convolution blocks 、 ,..., And activation function, the N convolution blocks are connected by residual paths, the convolution blocks The input is the convolution block The encoder is used to convert the first speech signal to obtain a feature space representation of the first speech signal; the residual network is used to extract features from the feature space representation and output a mask value of the user speech signal; the decoder is used to convert the mask value to obtain a second speech signal; The signal optimization module 403 is configured to perform howling suppression and speech enhancement on the second speech signal through an adaptive filtering algorithm to obtain a third speech signal, and send the third speech signal to a speaker at a non-sound source seat.

[0037] A third embodiment of the present invention provides a storage medium storing program instructions, which, when executed, implements the in-vehicle call method as described in any one of the above items.

[0038] A fourth embodiment of the present invention provides an electronic terminal, such as Figure 5 As shown, it includes a processor and a memory, the memory stores program instructions, and the processor runs the program instructions to implement the in-car call method as described in any one of the above items.

[0039] The present invention may be an apparatus, a method and / or a computer program product. The computer program product may include a readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0040] A storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. Storage media may include, for example, but are not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof.

[0041] It should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each implementation method can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

[0042] The series of detailed descriptions listed above are only specific descriptions of feasible implementation methods of the present invention. They are not intended to limit the scope of protection of the present invention. Any equivalent implementation methods or changes that do not deviate from the technical spirit of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for in-car communication, characterized in that: The following steps are involved: Acquire a first voice signal at a sound source seat, where the first voice signal includes a user voice signal and an environmental voice signal; Creating a time-domain convolutional network model, the time-domain convolutional network model including an encoder, a residual network, and a decoder, the time-domain convolutional network model being used to extract the user voice signal from the first voice signal to obtain a second voice signal; The residual network includes normalization, N convolution blocks 、 ,..., And activation function, the N convolution blocks are connected by residual paths, the convolution blocks The input is the convolution block The residual path output of , where i and N are both positive integers, ; The encoder is used to convert the first speech signal to obtain a feature space representation of the first speech signal; the residual network is used to extract features from the feature space representation and output a mask value of the user speech signal; the decoder is used to convert the mask value to obtain a second speech signal; Howling suppression and speech enhancement are performed on the second speech signal through an adaptive filtering algorithm to obtain a third speech signal, and the third speech signal is sent to a speaker at a seat other than the sound source seat.

2. The in-car communication method according to claim 1, characterized in that: The convolution block The convolution is a depth-separable convolution; the expression of the depth-separable convolution is: , ,in, For the convolution block Input, and The sizes are and a convolution kernel of 1, is the convolution operation, It is a splicing operation.

3. The in-car communication method according to claim 1, characterized in that: The residual network introduces an incremental dilation factor , the convolution block Expansion factor .

4. The in-car communication method according to claim 1, characterized in that: The “converting the first speech signal to obtain a feature space representation of the first speech signal” specifically includes: The encoder converts the discrete waveform of the first speech signal into a feature space representation of the first speech signal; the discrete waveform expression of the first speech signal is: , where M is the number of source signals of the first speech signal, is the source signal of the first speech signal, wherein j and M are both positive integers, .

5. The in-car communication method according to claim 4, characterized in that: The “extracting features from the feature space representation and outputting a masking value of the user voice signal” specifically includes: The feature space representation is extracted by the time domain convolutional network model and M source signals are output. The masking value of , ,in, is a constraint condition, H() is an optional activation function, is the basis function of the encoder.

6. The in-car communication method according to claim 1, characterized in that: The “converting the masking value to obtain the second speech signal” specifically includes: reconstructing the masking value by a decoder to obtain the second speech signal, ,in, is the basis function of the decoder.

7. The in-car communication method according to claim 1, characterized in that: The adaptive filtering algorithm is normalized least squares adaptive filtering.

8. An in-car communication device, characterized in that: Includes the following modules: A signal acquisition module, configured to acquire a first voice signal from a sound source seat, wherein the first voice signal includes a user voice signal and an environmental voice signal; A model creation module is used to create a time domain convolutional network model, wherein the time domain convolutional network model includes an encoder, a residual network and a decoder, and the time domain convolutional network model is used to extract the user voice signal from the first voice signal to obtain a second voice signal; the residual network includes normalization, N convolution blocks 、 ,..., And activation function, the N convolution blocks are connected by residual paths, the convolution blocks The input is the convolution block The encoder is used to convert the first speech signal to obtain a feature space representation of the first speech signal; the residual network is used to extract features from the feature space representation and output a mask value of the user speech signal; the decoder is used to convert the mask value to obtain a second speech signal; The signal optimization module is used to perform howling suppression and voice enhancement on the second voice signal through an adaptive filtering algorithm to obtain a third voice signal, and send the third voice signal to a speaker at a non-sound source seat.

9. A storage medium storing program instructions, characterized in that: When the program instructions are executed, the in-vehicle call method according to any one of claims 1 to 7 is implemented.

10. An electronic terminal, characterized in that: The method comprises a processor and a memory, wherein the memory stores program instructions, and the processor executes the program instructions to implement the in-vehicle call method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-feature fusion echo cancellation method and system based on self-attention transformation network

    CN113870874A

  • Speech enhancement method based on convolutional self-attention coding structure

    CN115700882A

  • Voice wake-up method and device for vehicle cabin and vehicle

    CN117789709A

  • Vehicle-mounted echo cancellation method and device, equipment and storage medium

    CN119028307A

  • In-vehicle call method and device, readable storage medium, electronic equipment and vehicle

    CN119152867A