Voice processing method and device, vehicle and equipment

By collecting audio signals inside the car and using a sound extraction model to identify and optimize the singing signal, the problem of noise interference in the car environment affecting singing is solved, achieving a high-quality singing experience inside the car.

CN120913604APending Publication Date: 2025-11-07CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511155829.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Singing in a car can disrupt the singer's rhythm due to ambient noise and interference, thus reducing the overall singing experience.

Method used

By collecting audio signals in the vehicle, using a sound extraction model to determine the singer's voice signal, and judging whether it is a singing signal based on spectral characteristics and voiceprint characteristics, the audio playback device is controlled to play the singing signal for voice enhancement and optimization.

Benefits of technology

It accurately identifies singing signals, improving the purity and personalization of singing in the car, enhancing the sound quality and immersion, and meeting the individual needs of different singers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913604A_ABST
    Figure CN120913604A_ABST
Patent Text Reader

Abstract

The invention relates to a voice processing method and device, a vehicle and equipment, and relates to the technical field of vehicles. The method comprises the following steps: in response to a singing mode of an audio playing device in a vehicle, collecting an audio signal in the vehicle; inputting the audio signal into a sound extraction model, and determining a voice signal of a singer; determining whether the voice signal is a singing signal; and if the voice signal is a singing signal, controlling the audio playing device to play the singing signal. Therefore, the audio playing experience which is purer and meets the singing scene requirement can be provided for the user, and the singing experience of the user singing in the vehicle is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicles, in particular to the technical field of intelligent vehicle cockpits, and specifically to a voice processing method and device, a vehicle and equipment. BACKGROUND

[0002] In the past two years, the entertainment demand of the vehicle cockpit has shown a diversified expansion trend. Among them, the in-vehicle seat singing without microphone is gradually gaining the favor of more and more users as a new form of entertainment. In addition, the introduction of personalized voice enhancement technology will undoubtedly inject new vitality into the vehicle entertainment experience and significantly improve its interactivity and immersion. This technology provides clearer and purer voice effects, and combines personalized sound effect processing, so that users can enjoy a pleasant and relaxing singing time in the car, achieving excellent live effects in all scenarios, multiple sound zones, high fidelity, low latency, anti-howling and noise reduction.

[0003] However, in the current singing scenario, whether it is solo singing or group singing, when there are disturbers talking or noise interference around, the rhythm of the singer will often be affected by the surrounding environment, thereby reducing the overall singing experience. Therefore, it is necessary to explore effective ways to improve the singing experience of the singer. SUMMARY

[0004] The present application provides a voice processing method, device, vehicle and equipment to at least solve the technical problem that the rhythm of the singer is often affected by the surrounding environment in related technologies, thereby reducing the overall singing experience. The technical solution of the present application is as follows:

[0005] According to the first aspect of the present application, a voice processing method is provided, comprising: collecting audio signals in the vehicle in response to the audio playback device in the vehicle being in a singing mode; inputting the audio signals to a sound extraction model to determine the voice signals of the singer; determining whether the voice signals are singing signals; and if the voice signals are singing signals, controlling the audio playback device to play the singing signals.

[0006] According to the above technical means, the present application can accurately determine the voice signals of the singer by collecting the audio signals in the vehicle and using the sound extraction model in response to the audio playback device in the vehicle being in a singing mode, and then determine whether it is a singing signal and control the audio playback device to play the signal after confirmation, avoiding the situation that the singing content cannot be accurately recognized or non-singing sound is mistaken for singing signal due to the interference of complex environmental sound in the vehicle, which can provide users with a more pure, accurate and current singing scene demand audio playback experience, and improve the singing experience of users in the car.

[0007] In a possible implementation, the determining whether the voice signal is a singing signal includes: performing feature extraction on the voice signal to obtain a spectrum feature of the voice signal; and determining whether the voice signal is a singing signal based on the spectrum feature.

[0008] According to the technical means, the voice signal is extracted based on the spectrum feature, and it is determined whether the voice signal is a singing signal, so that the accuracy and reliability of singing signal identification are improved, and the singing content is accurately played by the subsequent audio playing device.

[0009] In a possible implementation, the determining whether the voice signal is a singing signal based on the spectrum feature includes: if a similarity between the spectrum feature and a first spectrum feature is greater than or equal to a preset similarity threshold, determining that the voice signal is a singing signal; and the first spectrum feature is a spectrum feature of a song to be played by the audio playing device.

[0010] According to the technical means, the spectrum feature of the voice signal is compared with the spectrum feature (the first spectrum feature) of the song to be played by the audio playing device, and it is determined that the voice signal is a singing signal when the similarity is greater than or equal to the preset similarity threshold, so that the singing signal is not inaccurately identified due to factors such as in-vehicle environmental noise interference and different singing styles, the speaking content or noise of the singer is not misjudged as an effective singing signal, the singing behavior that matches the currently played song can be accurately identified, the user can be provided with a more pure, accurate and current singing scene demand audio playing experience, and the singing experience of the user in the vehicle is improved.

[0011] In a possible implementation, the voice signal of the singer is determined by inputting the audio signal into a sound extraction model, including: determining a voiceprint feature of the singer; inputting the voiceprint feature and the audio signal into the sound extraction model to obtain the voice signal output by the sound extraction model; and the sound extraction model is used to represent a mapping relationship between the voiceprint feature, the audio signal and the voice information.

[0012] According to the technical means, the voiceprint feature of the singer is determined first, and then the voiceprint feature and the audio signal are input into the sound extraction model having the mapping relationship between the voiceprint feature, the audio signal and the voice information to determine the voice signal of the singer, so that in a complex in-vehicle audio environment, the pure voice signal of the singer can be accurately extracted from the mixed audio by the sound extraction model, the problem that the extracted voice signal is inaccurate and contains a large amount of interference information is avoided, and the accurate analysis of the singing content, the personalized processing (such as adjusting the sound effect parameters according to different singers) and the high-quality singing experience are provided.

[0013] In a possible implementation, the method for determining the voiceprint feature of the singer comprises: collecting a voice signal sample of the singer by an audio collection device connected to the audio playing device; and performing feature extraction on the voice signal sample to obtain the voiceprint feature.

[0014] According to the technical means described above, the voice signal sample of the singer can be collected by the audio collection device connected to the audio playing device, and the voiceprint feature can be obtained by performing feature extraction thereon, thereby providing strong support for subsequent accurate identification and extraction of the voice signal of the specific singer based on the voiceprint feature, and further improving the identification and processing accuracy of the entire in-vehicle audio processing system for the voice of the singer.

[0015] In a possible implementation, the method for controlling the audio playing device to play the singing signal comprises: performing voice enhancement and voice optimization on the singing signal to obtain a target singing signal; and controlling the audio playing device to play the target singing signal.

[0016] According to the technical means described above, the voice enhancement and voice optimization can be performed on the singing signal to obtain the target singing signal, and the audio playing device can be controlled to play the target signal, which can significantly improve the sound quality of the singing signal, enhance the clarity, purity and fullness of the sound, and create a more high-quality and immersive in-vehicle singing atmosphere for the user, greatly improving the user's experience of singing in the vehicle.

[0017] In a possible implementation, the vehicle comprises a plurality of audio playing devices, and the method for controlling the audio playing device to play the target singing signal comprises: if there are a plurality of singers, determining a playing strategy of the plurality of audio playing devices based on audio preferences of the plurality of singers; the playing strategy is used to indicate the singing signal of the singer who needs to be highlighted; and based on the playing strategy, the plurality of audio playing devices are controlled to play the target singing signal.

[0018] According to the technical means described above, when there are a plurality of singers in the vehicle, the playing strategy of the plurality of audio playing devices can be determined based on the audio preferences of the plurality of singers, and the target singing signal can be played according to the strategy, which can avoid the situation that a single playing mode cannot meet the individualized needs of different singers for highlighting their own voices in the multi-singer singing scenario, resulting in poor experience of some singers, even mutual interference and covering of the voices, affecting the overall singing entertainment atmosphere, and can realize flexible and individualized in-vehicle audio playing, meet the expectations of different singers for highlighting their own voices, enhance the interest and participation of in-vehicle music interaction, and improve the experience of each passenger in the vehicle.

[0019] According to a second aspect provided in the present application, a voice processing apparatus is provided, comprising: an acquisition unit, a determination unit and a control unit; the acquisition unit is configured to acquire an audio signal in a vehicle in response to an audio playback device in the vehicle being in a singing mode; the determination unit is configured to input the audio signal to a sound extraction model to determine a voice signal of a singer; the determination unit is further configured to determine whether the voice signal is a singing signal; and the control unit is configured to control the audio playback device to play the singing signal if the voice signal is the singing signal.

[0020] In a possible implementation, the determination unit is specifically configured to: perform feature extraction on the voice signal to obtain a spectral feature of the voice signal; and determine whether the voice signal is the singing signal based on the spectral feature.

[0021] In a possible implementation, the determination unit is specifically configured to: determine that the voice signal is the singing signal if a similarity between the spectral feature and a first spectral feature is greater than or equal to a preset similarity threshold; and the first spectral feature is a spectral feature of a song to be played by the audio playback device.

[0022] In a possible implementation, the determination unit is specifically configured to: determine that the voice signal is the singing signal if a similarity between the spectral feature and a first spectral feature is greater than or equal to a preset similarity threshold; and the first spectral feature is a spectral feature of a song to be played by the audio playback device.

[0023] In a possible implementation, the determination unit is specifically configured to: determine a voiceprint feature of the singer, comprising: acquiring a voice signal sample of the singer through an audio acquisition device connected to the audio playback device; and performing feature extraction on the voice signal sample to obtain the voiceprint feature.

[0024] In a possible implementation, the control unit is specifically configured to: perform voice enhancement and voice optimization on the singing signal to obtain a target singing signal; and control the audio playback device to play the target singing signal.

[0025] In a possible implementation, the control unit is specifically configured to: if there are multiple singers, determine a playing strategy of multiple audio playback devices based on audio preferences of the multiple singers; the playing strategy is used to indicate a singing signal of a singer to be highlighted; and control the multiple audio playback devices to play the target singing signal based on the playing strategy.

[0026] According to a third aspect provided in the present application, a vehicle is provided, comprising: an audio playback device; and the audio playback device is configured to play a singing signal output by the voice processing apparatus in the second aspect.

[0027] According to a fourth aspect provided in the present application, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method of the first aspect and any possible implementation thereof.

[0028] According to a fifth aspect provided in the present application, a computer-readable storage medium is provided, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method of the first aspect and any possible implementation thereof.

[0029] According to a sixth aspect provided in the present application, a computer program product is provided, the computer program product comprising computer instructions, when the computer instructions are run on an electronic device, the electronic device performs the method of the first aspect and any possible implementation thereof.

[0030] It should be noted that the technical effects brought by any implementation of the second aspect to the fifth aspect can be referred to the technical effects brought by the corresponding implementation of the first aspect, which will not be repeated here.

[0031] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0032] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present application and serve to explain the principles of the present application, and do not constitute an undue limitation on the present application.

[0033] Figure 1 is a structural schematic diagram of a voice processing system according to an exemplary embodiment;

[0034] Figure 2 is a flowchart of a voice processing method according to an exemplary embodiment;

[0035] Figure 3 is a schematic diagram of a sound extraction model according to an exemplary embodiment;

[0036] Figure 4 is a schematic diagram of another sound extraction model according to an exemplary embodiment;

[0037] Figure 5 is a schematic diagram of a sound extraction process according to an exemplary embodiment;

[0038] Figure 6 is a framework diagram of a voice optimization model according to an exemplary embodiment;

[0039] Figure 7 is a schematic diagram of a voice processing flow according to an example embodiment;

[0040] Figure 8 is a block diagram of a voice processing apparatus according to an example embodiment;

[0041] Figure 9 is a block diagram of an electronic device according to an example embodiment. DETAILED DESCRIPTION

[0042] In order to make the skilled in the art better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings.

[0043] It should be noted that the terms "first", "second" and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily mean a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all the embodiments consistent with the present application. Rather, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0044] In the embodiments of the present application, the words "exemplary", "such as", or "for example" are used to mean serving as an example, instance, or illustration. Any embodiment or design described as "exemplary", "such as", or "for example" in the embodiments of the present application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of "exemplary", "such as", or "for example" is intended to present relevant concepts in a concrete manner.

[0045] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all.

[0046] The voice processing method provided in the embodiments of the present application can be applied in a vehicle. The vehicle can also be referred to as a vehicle, a mobile carrier, an electric vehicle (EV), a hybrid electric vehicle (HEV), a plug-in hybrid electric vehicle (PHEV), a fuel cell vehicle (FCV), an autonomous vehicle, an intelligent and connected vehicle (ICV), a driverless vehicle, and the like.

[0047] In the embodiments of the present application, the vehicle can be a car, a sport utility vehicle (SUV), a truck, an electric vehicle, a motorcycle, a tricycle, a special vehicle (such as an ambulance, a fire truck, a police car, and the like), a driverless taxi, an intelligent and connected bus, an autonomous logistics vehicle, an electric truck, and the like. In addition, the method is also applicable to various special vehicles, such as agricultural vehicles, mining vehicles, forestry vehicles, airport vehicles, port vehicles, and the like. The present application does not make specific limitations in this regard.

[0048] As shown in FIG. 1, the voice processing system in the vehicle includes a voice processing apparatus 101, an audio acquisition device 102, and an audio playback device 103. Figure 1 Figure 1 Optionally, a communication connection can be established between the voice processing apparatus 101 and the audio acquisition device 102, and a communication connection can be established between the voice processing apparatus 101 and the audio playback device 103.

[0049] Optionally, a communication connection can be established between the voice processing apparatus 101 and the audio acquisition device 102, and a communication connection can be established between the voice processing apparatus 101 and the audio playback device 103. Figure 1

[0050] In actual applications, the voice processing apparatus 101 can be in communication connection with one or more audio acquisition devices 102. Similarly, the voice processing apparatus 101 can be in communication connection with one or more audio playback devices 103. The audio acquisition device 102 can be in communication connection with the audio playback device 103.

[0051] For ease of understanding, the present application takes the voice processing apparatus 101 in communication connection with one audio acquisition device 102 and the voice processing apparatus 101 in communication connection with one audio playback device 103 as an example for illustration.

[0052] Optionally, a communication connection can be established between the voice processing apparatus 101 and the audio acquisition device 102, and a communication connection can be established between the voice processing apparatus 101 and the audio playback device 103. Figure 1 ​​The voice processing apparatus 101 and the audio acquisition device 102 in the voice processing system 100 can be functional modules integrated in the same device, or can be devices independently arranged from each other. The present application does not limit this.

[0053] It is easy to understand that when the voice processing apparatus 101, the audio acquisition device 102 and the audio playback device 103 are functional modules integrated in the same device, the communication mode between the voice processing apparatus 101, the audio acquisition device 102 and the audio playback device 103 is the communication between internal modules of the device. In this case, the communication flow between the two is the same as the communication flow when the voice processing apparatus 101, the audio acquisition device 102 and the audio playback device 103 are independently arranged from each other.

[0054] For ease of understanding, the present application mainly takes the voice processing apparatus 101, the audio acquisition device 102 and the audio playback device 103 as an example for description.

[0055] In one possible implementation manner, Figure 1 The audio acquisition device 102 in the voice processing system 100 can be arranged at a suitable position in the vehicle, such as above the center console, near the seat headrest, etc., to comprehensively collect the sound in the vehicle. The audio acquisition device 102 can collect the sound in the vehicle and convert the sound into an analog electric signal, i.e., an audio signal, in response to the audio playback device in the vehicle being in a singing mode. The audio acquisition device 102 can send the audio signal to the voice processing apparatus 101. The voice processing apparatus 101 can input the audio signal to a sound extraction model, determine the voice signal of the singer, and then determine whether the voice signal is a singing signal. If the voice signal is a singing signal, the voice processing apparatus 101 controls the audio playback device 103 to play the singing signal.

[0056] The audio signal can include the singing sound, the speaking sound, the environmental noise and other possible sound information of the people in the vehicle.

[0057] The audio playback device 103 can be arranged at various positions in the vehicle interior to facilitate uniform sound propagation. For example, the audio playback device 103 can be arranged at the center console, the roof and the rear central armrest.

[0058] Optionally, Figure 1 The voice processing apparatus 101 in the voice processing system 100 can be arranged in a terminal, a server or other types of electronic devices. Figure 1 The device form of the voice processing apparatus 101 shown in the voice processing system 100 is only one example and does not limit the voice processing apparatus 101.

[0059] When the voice processing device 101 is deployed in a terminal, the terminal can be a device providing voice and / or data connectivity to a user, a handheld device with wireless connectivity, or other processing devices connected to a wireless modem. The terminal can communicate with one or more core networks via a radio access network (RAN). The terminal can be a mobile terminal, such as a computer with a mobile terminal, or a mobile device built into the voice processing system that exchanges voice and / or data with the radio access network, such as a mobile phone, tablet computer, laptop computer, netbook, or personal digital assistant (PDA). This application does not impose any limitations on this.

[0060] When the voice processing device 101 is deployed on a server, the server can be a single server or a server cluster consisting of multiple servers. In some embodiments, the server cluster can also be a distributed cluster. This application does not impose any limitations on this.

[0061] It should be noted that the structures illustrated in the embodiments of this application do not constitute a limitation on the voice processing system. It may include more or fewer components than illustrated, or combine some components, split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0062] For ease of understanding, the speech processing method provided in this application will be described in detail below with reference to the accompanying drawings.

[0063] Figure 2 This is a flowchart illustrating a speech processing method according to an exemplary embodiment, such as... Figure 2 As shown, this speech processing method can be applied to Figure 1 The voice processing device 101 in the middle includes the following steps: S201-S204.

[0064] S201. In response to the audio playback device in the vehicle being in singing mode, the audio signal inside the vehicle is collected.

[0065] Among them, the singing mode is a specific working mode designed to meet the singing and entertainment needs of people in the vehicle.

[0066] In one possible implementation, the vehicle could activate the karaoke mode based on user selection. For example, the user could tap the "Karaoke" icon on the vehicle's central control screen or activate the function via voice command (such as "Hello XX, activate karaoke mode"). Upon recognizing the activation command, the vehicle could switch the audio playback device to karaoke mode.

[0067] In addition, in the singing mode, the vehicle can create a suitable singing environment. On the one hand, the vehicle can intelligently adjust the light system in the vehicle, switch the originally bright lighting to soft and rhythmic colored light, and the light color will change with the melody and rhythm of the music, enhancing the interest and atmosphere of singing. On the other hand, the vehicle can automatically close the windows and start the sound insulation and noise reduction function, effectively blocking the noise interference from the outside world, while accurately controlling the echo in the vehicle to avoid noise and howling, so that the singer can focus on their own singing and fully immerse themselves in the singing entertainment.

[0068] In one possible implementation, in response to the audio playback device in the vehicle being in the singing mode, the voice processing device can collect audio signals in the vehicle in all directions through an audio collection device (such as a microphone), including sounds such as conversations, audio playback, and environmental noise. Then, the audio collection device can convert the collected sound into an analog digital signal, i.e., an audio signal. The audio collection device can reduce noise in the audio signal as much as possible through echo cancellation algorithms, noise reduction algorithms, etc.

[0069] For example, the audio signal of an n-person chorus (including noise and speech) satisfies the following first formula:

[0070] The first formula.

[0071] Wherein, n can be used to represent the number of participants in the chorus. k can be used to represent the index variable of the summation. x(k) can be used to represent the signal of the kth component in the audio signal.

[0072] S202, input the audio signal to the sound extraction model to determine the voice signal of the singer.

[0073] Wherein, the singer can be one or more. When the audio playback device in the vehicle is in the singing mode, the singer of the song can be selected. For example, when the user selects a song to be sung on the vehicle screen, the vehicle screen can automatically pop up a member selection interface, and at this time the user can select the singer from the current vehicle passengers through voice instructions (such as "select the second row of two passengers") or touch screen operation (check the passenger's head portrait). The vehicle screen can also include preset selections of fixed combinations (such as pre-saving "family chorus group"), and also allows temporary dynamic adjustment, fully taking into account the flexibility and ritual sense in the multi-person travel scenario.

[0074] In one possible implementation, the voice processing device can determine the voiceprint features of the singer, and input the voiceprint features and the audio signal to the sound extraction model to obtain the voice signal output by the sound extraction model.

[0075] The voice extraction model is used to represent the mapping relationship between the voiceprint feature, the audio signal and the speech information. The voice extraction model can include a speaker encoder (Speaker Encoder) model constructed based on a long short-term memory (Long Short-Term Memory, LSTM) framework. The speaker encoder can extract the voiceprint feature to obtain a deep learning model of the speaker feature, and the fixed-length target vector (d-vector) output by the deep learning model can represent the target speech feature. The d-vector is a fixed-length vector obtained by averaging the time sequence feature sequence output by the last layer of the model in the time dimension when the speaker encoder extracts the feature from the target speech. The d-vector can be used for voiceprint embedding.

[0076] Specifically, the voice processing device can collect the voice signal sample of the singer through the audio acquisition device connected with the audio playing device, and extract the voice signal sample to obtain the voiceprint feature.

[0077] Illustratively, the voice processing device can prompt the collection of the voice sample of the singer after the singer determination is completed. For example, the voice processing device can display "Please say a sentence for voice sample collection to bring you a better singing experience" through the vehicle-mounted screen. The voice processing device can collect the voice signal sample through the audio acquisition device.

[0078] Further, the voice processing device can input the audio signal into the encoder to extract the feature of the audio signal and obtain the spectral feature of the audio signal. The voice processing device can input the voiceprint feature (d-vector) and the spectral feature of the audio signal into the voice extraction model.

[0079] The mask model of the voice extraction model can adopt a modular design strategy: taking a convolutional neural network (Convolutional Neural Network, CNN) as a basic architecture unit, combining multiple CNN modules and LSTM modules into a stack block in a series manner, and then flexibly adjusting the number of CNN modules and the hierarchical structure of the stack block according to the model performance requirements and algorithm optimization targets.

[0080] In the signal processing stage, the output of the mask model is mapped to the interval [0, 1] through a Sigmoid activation function, generating a time-frequency mask matrix w corresponding to each speaker. By performing element-wise point multiplication operation between the mask and the spectral features output by the encoder, a weighted masked spectrogram can be obtained, realizing the separation of the frequency domain information of the target speaker's voice signal. Finally, the decoder structure designed in advance is used to perform time-frequency inverse transformation on the masked spectrogram, and the time-domain voice waveform of the target singer is reconstructed, completing the process of accurately extracting the voice signal from the audio signal.

[0081] S203, determining whether the voice signal is a singing signal.

[0082] In a possible implementation manner, the voice processing apparatus can perform feature extraction on the voice signal to obtain a spectral feature of the voice signal. If the similarity between the spectral feature and the first spectral feature is greater than or equal to a preset similarity threshold, the voice processing apparatus can determine that the voice signal is a singing signal.

[0083] The first spectral feature is the spectral feature of a song to be played by the audio playback device. The spectral feature can reflect information such as energy distribution, time-varying characteristics, and harmonic structure of sound at different frequency components.

[0084] Optionally, the preset similarity threshold can be set according to actual needs. For example, the preset similarity threshold can be 80%, or 70%. The present application does not make specific limitations on this.

[0085] In another possible implementation manner, since the spectral feature during singing usually has stronger periodicity and stability (such as the pitch of consecutive notes being maintained), and the spectral feature during speaking fluctuates greatly (such as the intonation fluctuation), the voice processing apparatus can determine whether the spectral feature has periodicity and stability, thereby determining whether the voice signal is a singing signal.

[0086] S204, if the voice signal is a singing signal, controlling the audio playback device to play the singing signal.

[0087] In a possible implementation manner, if the voice signal is a singing signal, the voice processing apparatus can perform voice enhancement and voice optimization on the singing signal to obtain a target singing signal, and then the voice processing apparatus can control the audio playback device to play the target singing signal.

[0088] Specifically, the voice enhancement can include: inputting the voice signal into a voice filter (Voice Filter) network framework for decoding processing, generating a time-frequency mask and reconstructing the voice signal, and then using the enhanced voice signal as the input of the next round to carry out multi-round iteration and update, gradually suppressing the residual noise and interference components by using a closed-loop optimization mechanism, so that the voice signal is voice enhanced.

[0089] The voice optimization can include: inputting the voice signal after voice enhancement into a pre-trained sound effect enhancement network. The sound effect enhancement network can model the pitch curve of the off-key voice by using a multi-layer convolution and a Fast Fourier Transform (FFT) block architecture, and combining a Hidden Markov Model (HMM) smoothing technique.

[0090] Specifically, a standard Musical Instrument Digital Interface (MIDI) note sequence is used as a reference template to learn from off-key characteristics and generate an adapted pitch correction curve. Meanwhile, the uncompressed linear scale vocal spectrum envelope and note embedding features are fused, and after linear projection, the target frequency band is dynamically enhanced. Finally, the personalized voice with full frequency response and optimized pitch is output, and immersive sound effects consistent with the user's acoustic characteristics are superimposed, realizing voice optimization from pitch correction to sound quality beautification.

[0091] In one possible implementation, a vehicle can include multiple audio playback devices. If there are multiple singers, the voice processing device can determine a playback strategy for the multiple audio playback devices based on the audio preferences of the multiple singers, that is, based on the positions of the singers, adjust the output power and frequency band allocation of each audio playback device through the playback strategy, to highlight the singing signal of the target singer, so that the target singer can hear his own voice more.

[0092] The playback strategy can be used to indicate the singing signal of the singer to be highlighted. The target singer is any one or more of the multiple singers.

[0093] In one possible implementation, the voice processing device can score the singing signal. The voice processing device can input the singing segment with a score greater than a preset score threshold to a sound effect enhancement network to optimize the sound effect enhancement model.

[0094] Based on the technical scheme, the vehicle audio playing device is in the singing mode, the in-vehicle audio signal is collected, the singing voice signal is accurately determined by using the sound extraction model, and it is determined whether it is a singing signal, and the audio playing device is controlled to play the signal after confirmation, so as to avoid the situation that the singing content cannot be accurately recognized or non-singing sound is mistaken for a singing signal due to the interference of complex in-vehicle environment sound, and to provide a more pure, accurate and current singing scene demand audio playing experience for the user, and to improve the singing experience of the user in the vehicle.

[0095] In some embodiments, as Figure 3 shown, Figure 3 is a schematic diagram of a sound extraction model provided by an embodiment of the present application. Figure 3 The sound extraction model in the

[0096] In a possible implementation manner, the sound extraction model in the related art can include an encoder, an extraction network, a decoder, and an auxiliary network.

[0097] The audio signal is input into the encoder to obtain the encoded features. The speech signal sample is input into the encoder alone, and the encoded target embedding vector is provided as auxiliary information in the auxiliary network to the extraction network.

[0098] The encoded features are input into the extraction network, and the extraction network can extract the singer's voice features with the assistance of the auxiliary information.

[0099] The voice features are input into the decoder to reconstruct / separate the singer's voice signal.

[0100] In some embodiments, in combination with Figure 3 as Figure 4 shown, Figure 4 is a sound extraction model provided by an embodiment of the present application.

[0101] Among them, Figure 4 The sound extraction model in the Figure 3 compared with the sound extraction model in the introduced the singer's speaking state input into the encoder and the frequency spectrum features of the singer's voice to assist the sound extraction model to determine whether the singer's voice signal is a singing signal, so as to extract the singer's singing signal.

[0102] Further, the adult voice of the singing signal can be recovered for playing, and the singing signal can be scored to take the singing voice segment with a score greater than a preset score threshold as the reference singing signal of the singer.

[0103] Figure 4 In some embodiments, in combination with Figure 5 as Figure 5is a schematic diagram of a sound extraction process provided by an embodiment of the present application.

[0104] In a possible implementation manner, the speech processing process can include a model pre-training step and a sound extraction step.

[0105] The model pre-training step includes:

[0106] The encoder LSTM is trained alone. After training, the encoder LSTM is used to generate a D-vector.

[0107] The sound extraction step includes:

[0108] The singer's voice signal sample is input into the encoder LSTM to generate a D-vector (voiceprint feature) corresponding to the voice signal sample.

[0109] The mixed audio signal containing multi-chorus, noise, and speech is subjected to short-time Fourier transform to extract the spectral features of the mixed audio.

[0110] The clean audio signal is subjected to short-time Fourier transform to obtain the spectral features of the clean audio.

[0111] The spectral features and the D-vector are input into a trainable speech filter composed of a CNN, an LSTM, and a fully connected layer.

[0112] The extraction network can input a mask based on the spectral features and the D-vector.

[0113] Based on the mask and the spectral features of the mixed audio, the masked spectral features can be obtained, and the spectral features of the clean audio and the masked spectral features are used to calculate the difference through a loss function, and then the singer's voice signal is extracted through a decoder to complete the singer's sound extraction.

[0114] The singer's voice signal is subjected to speech enhancement, and the enhanced voice signal is obtained after multiple iterations.

[0115] In some embodiments, as shown in Figure 6 , a framework diagram of a speech optimization model is provided. Figure 6

[0116] In a possible implementation manner, the speech optimization model can include an encoder, an RNN module, a rectified linear unit (ReLU), an activation function, a predicted pitch module, and an output module (Output vpcal). The RNN module can include multiple RNN units to determine the sequence information and time dependency in the voice signal.

[0117] ​The singing signal and the score signal are input to the voice optimization model, and the voice optimization model can process the singing signal and the score signal, and finally output the voice-optimized singing signal through the output module.

[0118] In some embodiments, as shown in Figure 7 Figure 7 is a schematic diagram of a voice processing process.

[0119] In a possible implementation manner, the voice processing apparatus can input the audio signal and the voice signal sample to a sound extraction model after sound feedback suppression and noise reduction, to extract a singing signal. The voice processing apparatus can input the singing signal to a voice optimization model for voice optimization, and obtain a target singing signal based on a accompaniment signal, and output the target singing signal.

[0120] The voice processing apparatus can store the singing signal extracted by the sound extraction model, and perform local audio scoring and local audio beautification and sound correction.

[0121] The above mainly describes the solutions provided by the embodiments of the present application from the perspective of methods. In order to implement the above functions, the voice processing apparatus or the electronic device comprises a hardware structure and / or a software module corresponding to each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed in the present application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0122] The embodiments of the present application can divide the function modules of the voice processing apparatus or the electronic device according to the above method, for example, the voice processing apparatus or the electronic device can comprise function modules corresponding to each function division, or two or more functions can be integrated in one processing module. The integrated module can be realized in the form of hardware or software function module. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical function division. When actually implemented, there can be another division manner.

[0123] Figure 8 is a block diagram of a voice processing apparatus according to an example embodiment. Refer to Figure 8 ​The voice processing apparatus comprises: an acquisition unit 801, a determination unit 802, and a control unit 803; the acquisition unit 801 is configured to acquire an audio signal in the vehicle in response to an audio playback device in the vehicle being in a singing mode; the determination unit 802 is configured to input the audio signal to a sound extraction model to determine a voice signal of a singer; the determination unit 802 is further configured to determine whether the voice signal is a singing signal; and the control unit 803 is configured to control the audio playback device to play the singing signal if the voice signal is the singing signal.

[0124] Based on the above technical solution, the present application can acquire an audio signal in the vehicle in response to an audio playback device in the vehicle being in a singing mode, accurately determine a voice signal of a singer by using a sound extraction model, and then determine whether the voice signal is a singing signal and control the audio playback device to play the signal after confirmation, thereby avoiding the situation that singing content cannot be accurately recognized or non-singing sound is mistakenly played as a singing signal due to interference of complex environmental sound in the vehicle, providing a user with a more pure, accurate, and current singing scene demand-compliant audio playback experience, and improving the singing experience of the user in the vehicle.

[0125] In a possible manner, the determination unit 802 is specifically configured to: perform feature extraction on the voice signal to obtain a spectral feature of the voice signal; and determine whether the voice signal is a singing signal based on the spectral feature.

[0126] In a possible manner, the determination unit 802 is specifically configured to: if a similarity between the spectral feature and a first spectral feature is greater than or equal to a preset similarity threshold, determine that the voice signal is a singing signal; and the first spectral feature is a spectral feature of a song to be played by the audio playback device.

[0127] In a possible manner, the determination unit 802 is specifically configured to: if a similarity between the spectral feature and a first spectral feature is greater than or equal to a preset similarity threshold, determine that the voice signal is a singing signal; and the first spectral feature is a spectral feature of a song to be played by the audio playback device.

[0128] In a possible manner, the determination unit 802 is specifically configured to: determine a voiceprint feature of the singer, including: acquiring a voice signal sample of the singer by using an audio acquisition device connected to the audio playback device; and performing feature extraction on the voice signal sample to obtain the voiceprint feature.

[0129] In a possible manner, the control unit 803 is specifically configured to: perform voice enhancement and voice optimization on the singing signal to obtain a target singing signal; and control the audio playback device to play the target singing signal.

[0130] In a possible implementation, the control unit 803 is specifically configured to: if there are multiple singers, determine a playing strategy of the multiple audio playing devices based on audio preferences of the multiple singers; the playing strategy is used to indicate a singing signal of a singer that needs to be highlighted; and control the multiple audio playing devices to play the target singing signal based on the playing strategy.

[0131] As to the apparatus in the above embodiments, the specific manners in which the various modules perform operations have been described in detail in the embodiments of the method, and thus will not be described in detail here.

[0132] Figure 9 is a block diagram of an electronic device according to an example embodiment. As shown in Figure 9 The electronic device includes, but is not limited to, a processor 901 and a memory 902.

[0133] The memory 902 is configured to store executable instructions of the processor 901. It can be understood that the processor 901 is configured to execute the instructions to implement the voice processing method in the above embodiments.

[0134] It should be noted that those skilled in the art can understand, Figure 9 the structure of the electronic device shown in the above embodiments does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than Figure 9 those shown in the above embodiments, or combine certain components, or arrange different components.

[0135] The processor 901 is the control center of the electronic device, connects all parts of the electronic device through various interfaces and lines, and performs various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 902 and calling data stored in the memory 902, thereby overall monitoring the electronic device. The processor 901 can include one or more processing units. Optionally, the processor 901 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 901.

[0136] The memory 902 can be used to store software programs and various data. The memory 902 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs (such as a determination unit, a processing unit, etc.) required by at least one functional module, etc. In addition, the memory 902 can include a high-speed random access memory, and can also include a non-volatile memory, for example, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device.

[0137] In an example embodiment, a computer-readable storage medium, for example, the memory 902 including instructions, is also provided, which can be executed by the processor 901 of the electronic device to implement the method in the above-described embodiments.

[0138] In actual implementation, Figure 8 The functions of the acquisition unit 801, the determination unit 802, and the control unit 803 in the above-described embodiments can be implemented by the processor 901 calling the computer program stored in the memory 902. Figure 9 The specific execution process can refer to the description of the method part in the above-described embodiments, and will not be described here.

[0139] Alternatively, the computer-readable storage medium can be a non-transitory computer-readable storage medium, for example, a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc. In an example embodiment, the present application also provides a computer program product including one or more instructions, which can be executed by the processor 901 of the electronic device to complete the method in the above-described embodiments.

[0140] It should be noted that the instructions in the above-described computer-readable storage medium or the one or more instructions in the computer program product are executed by the processor of the electronic device to implement each process of the above-described method embodiments, and can achieve the same technical effects as the above-described method. To avoid repetition, it will not be described here.

[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-described division of functional modules is taken as an example for illustration. In actual application, the above-described functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0142] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the described apparatus embodiments are merely schematic. The division of the modules or units is merely logical function division. There can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another apparatus, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, apparatuses or units, and can be in electrical, mechanical or other forms.

[0143] The units described as separated components can or can not be physically separated, and the components displayed as units can be one physical unit or multiple physical units, i.e., can be located in one place, or can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments.

[0144] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware or software function unit.

[0145] If the integrated unit is implemented in the form of software function unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solutions of the embodiments of the present application essentially or substantially, or part of the classification of the technical solutions or the whole classification, can be embodied in the form of a software product. The software product is stored in a storage medium, and includes several instructions for causing an apparatus (which can be a single chip machine, a chip, etc.) or a processor to execute all or part of the steps of the methods of the embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, and various other media that can store program codes.

[0146] The embodiments of the present application provide a computer program product containing instructions, which, when executed on a computer, cause the computer to perform the voice processing method in the method embodiments.

[0147] The embodiments of the present application also provide a computer readable storage medium, which stores instructions. When the instructions are executed on a computer, the computer performs the voice processing method in the method flow shown in the method embodiments.

[0148] The computer readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a register, a hard disk, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing, or any other medium from which a computer can read instructions. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Of course, the storage medium can be a part of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). In some embodiments, the computer readable storage medium can be any tangible medium that can contain or store program code for use by or in connection with an instruction execution system, apparatus, or device.

[0149] The speech processing apparatus, the computer readable storage medium, and the computer program product in the embodiments of the present application can be applied to the method described above, and the technical effects that can be achieved thereby can be referred to the method embodiments described above, which will not be described herein again.

[0150] The above merely describes the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any change or replacement within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A voice processing method, characterized by, The voice processing method comprises: in response to an audio playing device in a vehicle being in a singing mode, collecting an audio signal in the vehicle; inputting the audio signal into a sound extraction model to determine a voice signal of a singer; determining whether the voice signal is a singing signal; if the voice signal is the singing signal, controlling the audio playing device to play the singing signal.

2. The voice processing method of claim 1, wherein, The determination of whether the voice signal is the singing signal comprises: performing feature extraction on the voice signal to obtain a spectral feature of the voice signal; based on the spectral feature, determining whether the voice signal is the singing signal.

3. The voice processing method of claim 2, wherein, The determination of whether the voice signal is the singing signal based on the spectral feature comprises: if a similarity between the spectral feature and a first spectral feature is greater than or equal to a preset similarity threshold, determining that the voice signal is the singing signal; the first spectral feature is a spectral feature of a song to be played by the audio playing device.

4. The voice processing method of claim 1, wherein, The inputting of the audio signal into the sound extraction model to determine the voice signal of the singer comprises: determining a voiceprint feature of the singer; inputting the voiceprint feature and the audio signal into the sound extraction model to obtain the voice signal output by the sound extraction model; wherein the sound extraction model is used to represent a mapping relationship between the voiceprint feature, the audio signal and the voice information.

5. The voice processing method of claim 4, wherein, The determination of the voiceprint feature of the singer comprises: collecting a voice signal sample of the singer through an audio collection device connected with the audio playing device; performing feature extraction on the voice signal sample to obtain the voiceprint feature.

6. The voice processing method of claim 1, wherein, The control of the audio playing device to play the singing signal comprises: performing voice enhancement and voice optimization on the singing signal to obtain a target singing signal; controlling the audio playing device to play the target singing signal.

7. The voice processing method of claim 6, wherein, The vehicle comprises a plurality of audio playing devices, and the control of the audio playing device to play the target singing signal comprises: if there are a plurality of singers, determining a playing strategy of a plurality of audio playing devices based on audio preferences of a plurality of singers; the playing strategy is used to indicate a singing signal of a singer to be highlighted; based on the playing strategy, controlling a plurality of audio playing devices to play the target singing signal.

8. A speech processing device, characterized by The voice processing device comprises a collection unit, a determination unit and a control unit; the collection unit is configured to collect an audio signal in a vehicle in response to an audio playing device in the vehicle being in a singing mode; the determination unit is configured to input the audio signal into a sound extraction model to determine a voice signal of a singer; the determination unit is further configured to determine whether the voice signal is a singing signal; the control unit is configured to control the audio playing device to play the singing signal if the voice signal is the singing signal.

9. A vehicle characterized by comprising: The vehicle comprises an audio playing device; the audio playing device is configured to play a singing signal output by the device of claim 8.

10. An electronic device, comprising: comprises: a processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method of any one of claims 1-7.