In-vehicle call voice processing method, system, electronic device and storage medium

By differentiating the in-vehicle call mode and encryption level and selecting the appropriate playback device and audio processing technology, the problems of noise interference and privacy leakage in the in-vehicle communication system are solved, and the call quality and driving experience are improved.

CN119183101BActive Publication Date: 2025-10-03CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411420957.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-12
Publication Date
2025-10-03
Estimated Expiration
2044-10-12

AI Technical Summary

Technical Problem

Existing in-vehicle communication systems have problems such as noise interference, privacy leakage and call interruption when making calls in the car, which affects call quality and driving experience.

Method used

By distinguishing the in-car call mode into private call mode and non-private call mode, selecting the appropriate playback device to play the call audio, and controlling the content of the second playback device according to the encryption level, combining sound source localization, voice emotion recognition and aliasing reconstruction technology to optimize audio processing.

Benefits of technology

It achieves privacy protection and improves call quality in the car, ensuring that the call content is only played near the caller, reducing interference to other passengers, and improving the driving experience and call safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119183101B_ABST
    Figure CN119183101B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of voice processing technology, and discloses a method, system, electronic device and storage medium for in-vehicle call voice processing. To address the problems of poor privacy and low call quality in existing in-vehicle communication systems, the present invention obtains the current in-vehicle call mode, selects a target playback device for call audio from the playback devices in the car according to the current in-vehicle call mode, and controls the target playback device to play the call audio; if the current in-vehicle call mode is a privacy call mode, obtains the current call encryption level; according to the current call encryption level, controls a second playback device to play content matching the encryption level, thereby realizing privacy call protection, effectively removing interference noise, protecting the privacy of in-vehicle calls, and improving the driving experience and enhancing the privacy and security of communications through aliasing reconstruction of call audio and music.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice processing technology, and in particular to a method, system, electronic device and storage medium for in-vehicle call voice processing. Background Art

[0002] With the development of automobile technology and the improvement of people's living standards, more and more families own private cars, and the in-vehicle communication system is one of the important components of modern cars. However, while the existing technology provides convenience, it also has some defects.

[0003] 1. When making calls in the car, the sound quality may be degraded due to interference from noise inside the car, such as engine roar and wind noise, affecting the call quality.

[0004] 2. Other passengers in the car will hear the conversation, which may leak the privacy of both parties;

[0005] 3. During a call, the in-car navigation, music, etc. will be interrupted, affecting the riding experience of other passengers in the car.

[0006] Therefore, it is necessary to provide a new in-vehicle call voice processing method to achieve noise reduction during in-vehicle calls, protect privacy in the car, improve call quality, and enhance the driving experience. Summary of the Invention

[0007] In order to solve the above technical problems, the present invention provides a method, system, electronic device and storage medium for in-vehicle call voice processing, which can achieve noise reduction effect during in-vehicle calls, protect privacy in the car, improve call quality, and enhance the driving experience.

[0008] In a first aspect, the present invention provides a method for processing in-vehicle call voice, comprising:

[0009] When the in-vehicle call is connected, obtaining the current in-vehicle call mode; the in-vehicle call mode includes a non-privacy call mode and a privacy call mode;

[0010] According to the current in-vehicle call mode, a target playback device for the call audio is selected from the playback devices in the vehicle, and the target playback device is controlled to play the call audio; wherein the playback devices include a first playback device and a second playback device; and the distance between the first playback device and the person in the vehicle who is talking on the phone is smaller than the distance between the second playback device and the person in the vehicle who is talking on the phone;

[0011] If the current in-car call mode is private call mode, obtain the current call encryption level;

[0012] According to the current call encryption level, controlling the second playback device to play content matching the encryption level;

[0013] According to the current in-car call mode, a target playback device for the call audio is selected from the playback devices in the car, in the following manner:

[0014] If the current in-vehicle call mode is a private call mode, playing the call audio using a first playback device;

[0015] When the current in-vehicle call mode is a non-privacy call mode, the call audio is played using the first playback device and / or the second playback device.

[0016] In a second aspect, the present invention provides a vehicle call voice processing system, which adopts the vehicle call voice processing method as described in any one of the above items, specifically comprising:

[0017] Audio receiving device: used to receive mixed audio;

[0018] Audio processing module: used to obtain the current in-vehicle call mode when a car call is connected; the in-vehicle call mode includes a non-private call mode and a private call mode; based on the current in-vehicle call mode, select a target playback device for the call audio from the playback devices in the car, and control the target playback device to play the call audio; if the current in-vehicle call mode is the private call mode, obtain the current call encryption level; based on the current call encryption level, control the second playback device to play content that matches the encryption level;

[0019] Playback device: used for playing audio; the playback device includes a first playback device and a second playback device; the distance between the first playback device and the person talking in the car is smaller than the distance between the second playback device and the person talking in the car.

[0020] In a third aspect, the present invention provides an electronic device, comprising:

[0021] processor and memory;

[0022] The processor is used to execute the steps of the in-vehicle call voice processing method as described above by calling the program or instructions stored in the memory.

[0023] In a fourth aspect, the present invention provides a storage medium, characterized in that the storage medium stores a computer executable program, which is called by a processor to execute the steps of any of the above-mentioned in-vehicle call voice processing methods.

[0024] The embodiments of the present invention have the following technical effects:

[0025] First, private call mode: The present invention selects a target playback device for the call audio from the playback devices in the car based on the current in-car call mode. In private mode, the system only plays the call audio content on the playback device closer to the caller, preventing the voices of other people in the car from being heard by the caller, thereby enhancing the privacy of the communication;

[0026] Second, the system intelligently matches encryption levels to playback modes: Based on the encryption level, the system intelligently adjusts the audio output of other playback devices, playing entertainment audio with level 1 encryption and distracting audio with level 2 encryption, further enhancing call privacy and security. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 This is a flow chart of a method for processing in-vehicle call voice provided by an embodiment of the present invention;

[0029] Figure 2 1 is a schematic diagram of an in-vehicle call voice processing system provided by an embodiment of the present invention;

[0030] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.

[0032] The specific embodiments of the present invention are described below with reference to the accompanying drawings.

[0033] The present invention addresses the problems of poor privacy and low call quality in existing in-vehicle communication systems. The present invention obtains the current in-vehicle call mode, selects a target playback device for the call audio from the playback devices in the car according to the current in-vehicle call mode, and controls the target playback device to play the call audio; if the current in-vehicle call mode is a privacy call mode, obtains the current call encryption level; according to the current call encryption level, controls the second playback device to play content matching the encryption level, thereby realizing privacy call protection, effectively removing interference noise, protecting the privacy of in-vehicle calls, and improving the driving experience and enhancing the privacy and security of communications through the aliasing reconstruction of call audio and music.

[0034] Example 1

[0035] The present invention proposes a method for processing vehicle-mounted call voice. Figure 1 This is a flow chart of a method for processing in-vehicle call voice provided by an embodiment of the present invention. Figure 1 , including steps S1 to S4:

[0036] Step S1, when the in-vehicle call is connected, obtaining the current in-vehicle call mode; the in-vehicle call mode includes a non-privacy call mode and a privacy call mode.

[0037] In-car call modes can be divided into encrypted and non-encrypted modes to meet different call needs and help improve the driving experience.

[0038] Non-private call mode: The call content can be heard by all passengers in the car. It is suitable for general calls that do not require confidentiality. When selecting this mode, the system may play the call audio to all or a specified playback device in the car.

[0039] Private Call Mode: When a call needs to be kept confidential, or when there are other passengers in the car and the caller does not want others to overhear the call, you can select Private Call Mode. In this mode, the system limits the playback range of the call audio, typically only to one or a few devices closest to the caller, to reduce the possibility of overhearing the call.

[0040] Step S2, according to the current in-vehicle call mode, select a target playback device for the call audio from the playback devices in the car, and control the target playback device to play the call audio; wherein, the playback device includes a first playback device and a second playback device; the distance between the first playback device and the person talking in the car is smaller than the distance between the second playback device and the person talking in the car.

[0041] Furthermore, according to the current in-car call mode, a target playback device for the call audio is selected from the playback devices in the car, as follows:

[0042] If the current in-car call mode is a private call mode, using the first playback device to play the call audio can reduce the possibility of the call content being heard by other passengers in the car;

[0043] When the current in-vehicle call mode is a non-private call mode, the call audio is played using the first playback device and / or the second playback device, and whether all passengers can hear the call content is selected as needed.

[0044] By locating the sound source, such as using array microphones and image sensors inside the car, the position of the voice call recipient in the car is determined, and then the distance between each playback device in the car and the caller is determined. According to the needs of the caller, the playback devices are divided into the first playback device and the second playback device. The system can dynamically adjust the playback device's capabilities and adjust the playback strategy in real time according to changes during the call (such as passenger movement, door opening and closing, etc.).

[0045] The playback devices are divided into a first playback device and a second playback device. The first playback device is usually located closer to the caller, while the second playback device is located farther away. This helps to limit the call audio to the vicinity of the caller in private call mode, thereby reducing interference to other passengers and the risk of privacy leakage.

[0046] Furthermore, before controlling the target playback device to play the call audio, the method further includes adjusting the volume of the playback device according to the caller's mood, call keywords, and response level, specifically:

[0047] Acquire the emotions of the caller through the voice emotion recognition model;

[0048] Use neural network models to model conversation keywords and response levels;

[0049] Adjust the volume of the playback device according to the caller's mood, call keywords, and response level.

[0050] Furthermore, before controlling the target playback device to play the call audio, the method further includes: performing call audio extraction processing on the transmitted mixed audio, and performing aliasing reconstruction processing on the call audio after the call audio extraction processing.

[0051] Before playing the call audio, the mixed audio undergoes call audio extraction processing, including noise reduction, voice separation, keyword recognition, and voiceprint feature extraction, aiming to accurately extract the caller's voice from the mixed audio. In non-privacy mode, aliasing reconstruction is performed to merge the extracted call audio with the original in-car audio (such as music or navigation sounds), allowing the call content and the original in-car audio to be played simultaneously.

[0052] Aliasing reconstruction processing ensures the harmonious unity of call audio and the ambient audio in the car, improving the clarity and privacy of calls while avoiding unnecessary interference to other passengers in the car, thereby optimizing the overall driving experience.

[0053] Specifically, the mixed audio extraction process includes steps S21-S25:

[0054] Step S21, performing noise reduction processing on the mixed audio;

[0055] Step S22, separating the human voice audio from the mixed audio after the noise reduction processing;

[0056] Step S23, performing keyword recognition on the human voice audio, and retaining sentence audio containing voice call keywords;

[0057] Step S24, performing voiceprint feature recognition on the sentence audio containing the voice call keyword, and outputting the voiceprint feature of the caller;

[0058] Step S25: extracting the call audio according to the voiceprint feature.

[0059] Through the call audio extraction processing of steps S21 to S25, the voice of the target caller can be accurately extracted from the complex in-vehicle audio environment, which not only improves the call quality and ensures the clarity and intelligibility of the voice, but also enhances the security and privacy of the call through the matching of voiceprint features, making the call experience smoother and more focused.

[0060] If the audio is mixed with broadcast sounds, car horns and other people's chatting sounds, then step S21 is used to obtain audio with reduced background noise, step S22 is used to obtain audio containing only human voices, step S23 is used to obtain sentence audio containing call keywords, such as audio containing "hello", "hello", "where are you", etc., step S24 is used to obtain the voiceprint features of the callers, and step S25 is used to obtain all the audios of the callers in the audio.

[0061] In the present invention, the call audio is not directly extracted from the mixed audio. This is because not every sentence spoken by the caller contains the call keyword. In order to avoid the system missing part of the call audio, the present invention reduces the background noise through noise reduction processing, and then separates the human voice part, and then identifies the keyword to locate the call audio, and obtains the voiceprint features of the caller through voiceprint recognition technology. In this way, the audio is subsequently extracted according to the voiceprint features, and finally the complete and clear voice of the specific caller is extracted.

[0062] The method of separating the human voice audio from the mixed audio after the noise reduction process includes:

[0063] Input the noise-reduced mixed audio into the trained separation model and output the human voice audio;

[0064] The separation model is used to extract human voice audio from a variety of mixed audios. All existing methods for achieving sound separation in the prior art can be used in the present invention.

[0065] The aliasing reconstruction process is performed on the call audio after the call audio extraction process, specifically:

[0066] The call audio extracted and processed by the call audio is mixed and reconstructed with the audio being played in the car before the car call is turned on, and the mixed and reconstructed audio is then played through the target playback device.

[0067] Aliasing reconstruction is an advanced audio processing technology that involves intelligently blending processed call audio with existing in-car audio content (such as music, navigation instructions, or other media audio). This process ensures a natural transition and smoothness in call audio while minimizing disruption to other passengers in the car. This technology ensures a consistent and enjoyable in-car audio experience even during calls, with a smooth transition between call content and background audio, improving overall listening comfort and call quality. Furthermore, aliasing reconstruction facilitates shared audio playback in non-privacy call mode, or restricts call audio playback to a specific area in privacy mode, ensuring the privacy of calls and the listening enjoyment of passengers.

[0068] Furthermore, the sound categories contained in the mixed audio can be distinguished and processed differently.

[0069] (1) If the picked-up audio is pure music (or other audio that does not contain human voice), it is not difficult for the user to recognize the voice of the caller after reconstructing the music and the call. The frequency range and sound intensity of the human voice can be enhanced by sound enhancement technology (digital signal processing technology) to make the human voice clearer than the audio. Or noise suppression technology can analyze the characteristics of background noise and adopt filtering, compensation and other technologies to eliminate echo interference and improve the clarity of the human voice. Or audio coding technology can compress and decode audio data to reduce the size of audio files and improve the playback speed and sound quality. Or microphone array technology can realize the directional collection and positioning of human voice through an array composed of multiple microphones, thereby improving the clarity of human voice.

[0070] (2) If the picked-up audio contains human voices, then based on the above-mentioned converted text, the caller's emotions, speaking speed, keywords in the call content, and the user's response level are identified according to the call text information to make corresponding decisions, and determine whether it is necessary to stop playing the in-car audio, reduce the volume of the in-car audio, or enhance the noise reduction function of the in-car audio to make the human voice clearer, etc., so that the user can better obtain call information and improve call quality;

[0071] Identifying the emotions of the caller is done through voice emotion recognition technology. This involves collecting a certain amount of voice data, including data in different emotional states, such as happiness, sadness, and anger. Feature extraction is performed on the collected voice data, converting the voice signal into a set of numerical features, such as fundamental frequency, energy, and harmonics. The extracted features are trained using machine learning algorithms or deep learning models to establish a sentiment classification model. New voice data is then sentimentally classified into different emotion types. Different emotion types can also be prioritized or ranked, and the system adaptively adjusts the in-car audio based on the priority or ranking of different emotions.

[0072] Keywords during calls are also collected by collecting the target user's daily call content. Using databases and data training, the target user's attention and interest in different call contents are judged, and the corresponding keywords in the sentences are extracted and saved. Keywords in new conversations are compared and classified into different priorities and importance levels to adaptively adjust the in-car audio.

[0073] The user's responsiveness refers to the degree to which the user responds to the content of the call, including the number of replies, reply duration, speaking frequency, speaking duration, etc. It is used to measure whether the user is interested in the call or wants to answer, and to determine whether the in-car audio needs to be adjusted to improve call quality.

[0074] In addition to the aforementioned processing of call audio, the present invention can also be used to perform the aforementioned processing on the audio inside the vehicle. During a call, the system automatically picks up the audio inside the vehicle, performs the aforementioned call audio extraction processing on the audio inside the vehicle, and obtains the in-vehicle call audio. Only the in-vehicle call audio is transmitted to the other party, thereby protecting the privacy of the vehicle.

[0075] The microphone picks up sounds such as in-car audio and calls and then identifies their digital signals. The digital audio signals of music are usually transmitted from music files (MP3, WAV, etc.) or network streaming media, and usually have a higher sampling rate, bit depth and number of channels to ensure the details and quality of the music; the digital audio signals during phone calls are generated by the telephone communication system, and usually have a lower sampling rate and bit depth to save bandwidth and storage space, and ensure the fluency and intelligibility of the voice; the difference in the above digital signals is used to distinguish between call and audio signals when picking up sound.

[0076] Step S3: If the current in-vehicle call mode is the private call mode, obtain the current call encryption level;

[0077] Step S4: According to the current call encryption level, control the second playback device to play content matching the encryption level.

[0078] When the current call encryption level is level one, controlling the second playback device to play entertainment audio;

[0079] When the current call encryption level is level 2, controlling the second playback device to play interference audio;

[0080] The interference audio is entertainment audio whose voiceprint features are consistent with the voiceprint features of the caller.

[0081] The caller selects the encryption level for their call based on their needs. Encryption level control is a privacy protection mechanism that determines the content played on a secondary playback device (located farther away from the caller) based on the call's sensitivity. Level 1 encryption plays standard entertainment audio, suitable for calls where the content is not private. Level 2 encryption plays interference audio that matches the caller's voiceprint. This interference audio effectively obscures the content of the call and is suitable for calls requiring higher privacy protection. This design ensures call privacy while providing different levels of privacy to suit different call needs.

[0082] The audio playback device reconstructs the call audio and the pure audio information in the car to obtain a new mixed audio signal, while highlighting the human voice signal of the call partner, and uses this audio signal for the in-car call audio.

[0083] Furthermore, the new mixed audio signal can be played on the first playback device, while the playback devices in other areas continue to play normally, which ensures that the call recipient can receive clear call voice audio, and also ensures that the audio-visual experience of passengers who are not the call recipients is not interrupted by the call or causes discomfort due to the call.

[0084] When the call privacy mode is turned on, the microphone picks up the audio of non-calling parties in the car, extracts audio features, and uses them for active noise reduction or voice masking in the car area. By playing, reducing the audio, or voice masking the audio on the playback device in the target area, non-calling parties in the car cannot obtain the accurate voice and audio information of the calling party. At the same time, since voiceprint recognition and tracking are turned on, the in-car call microphone will not pick up other audio signals of the non-calling party in the car. This ensures that the content of the call cannot be heard by other passengers in the car, and the conversations or entertainment audio of other passengers in the car will not be obtained by the calling party outside the car, ensuring the privacy of the call.

[0085] Audio mixing technology is used to mix multiple different audio signals into a new audio signal in a certain proportion. The volume, balance and frequency characteristic parameters of the signals are adjusted according to the characteristics of the audio signals and call signals to ensure coordination between the various signals. In the later stage, equalizers, compressors, reverberators, etc. can be added as needed to enhance the expressiveness of the audio signal, or post-processing such as noise reduction, denoising, gain control, etc. can be performed to improve audio quality and effects.

[0086] In non-private call mode, the present invention reconstructs the audio of the call content by mixing it with the audio played in the car, thereby achieving simultaneous playback of call audio and music. This can make the call sound more naturally integrated into the car environment, enhance the driving experience, and at the same time ensure the clarity of the call content and the enjoyment of the music.

[0087] Furthermore, emotion recognition processing can be added to the in-vehicle call processing method of the present invention. Through audio features, voice emotion recognition can be performed on the current object, and different in-vehicle control parameters can be preset according to different emotional states. For example, the call volume, entertainment audio playback volume, air conditioning, windows, fragrance and other in-vehicle environmental parameters can be adjusted, etc.; when the corresponding emotional changes are matched, the matching in-vehicle control parameter scene is called to enhance the in-vehicle passenger experience.

[0088] The in-vehicle call processing method provided by the present invention can effectively reduce non-human voice audio noise interference in the car when making a voice call through the in-vehicle equipment, and even reduce the human voice audio interference of non-calling parties, while protecting the privacy in the car and improving the call quality. In addition, during the voice call, other audio in the car will not be interrupted, thereby improving the user experience and comfort, and will not incur additional hardware costs or modification costs.

[0089] Example 2

[0090] The present invention also provides a vehicle-mounted call voice processing system, which executes the vehicle-mounted call voice processing method provided in embodiment 1. Figure 2 Schematic diagram of a vehicle call voice processing system provided by an embodiment of the present invention, specifically comprising:

[0091] Audio receiving device: used to receive mixed audio;

[0092] Audio processing module: used to obtain the current in-vehicle call mode when a car call is connected; the in-vehicle call mode includes a non-private call mode and a private call mode; based on the current in-vehicle call mode, select a target playback device for the call audio from the playback devices in the car, and control the target playback device to play the call audio; if the current in-vehicle call mode is the private call mode, obtain the current call encryption level; based on the current call encryption level, control the second playback device to play content that matches the encryption level;

[0093] Playback device: used for playing audio; the playback device includes a first playback device and a second playback device; the distance between the first playback device and the person talking in the car is smaller than the distance between the second playback device and the person talking in the car.

[0094] Each part in this application can be embedded in existing vehicle-mounted equipment without the need for major modifications, thereby reducing the cost of equipment upgrades.

[0095] Example 3

[0096] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 3 As shown, the electronic device 500 includes one or more processors 501 and a memory 502 .

[0097] The processor 501 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 500 to perform desired functions.

[0098] The memory 502 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 501 may execute the program instructions to implement the in-vehicle call voice processing method of any embodiment of the present application described above and / or other desired functions. The computer-readable storage medium may also store various contents such as initial external parameters and thresholds.

[0099] In one example, electronic device 500 may further include an input device 503 and an output device 504, which are interconnected via a bus system and / or other connection mechanisms (not shown). Input device 503 may include, for example, a keyboard, a mouse, etc. Output device 504 may output various information to the outside, including warning information, braking force, etc. Output device 504 may include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto.

[0100] Of course, to simplify, Figure 3Only some of the components related to the present application in the electronic device 500 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 500 may further include any other appropriate components according to specific application scenarios.

[0101] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute a vehicle call voice processing method provided by any embodiment of the present application.

[0102] The computer program product may be written in any combination of one or more programming languages ​​to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0103] In addition, an embodiment of the present application may also be a computer-readable storage medium on which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor executes a method for in-vehicle call voice processing provided by any embodiment of the present application.

[0104] The computer-readable storage medium may be any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0105] It should be noted that the terms used in the present invention are only for describing specific embodiments and are not intended to limit the scope of this application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular and may also include the plural. The terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method or device comprising a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of further restrictions, the elements defined by the sentence "comprise a..." do not exclude the presence of other identical elements in the process, method or device comprising the elements.

[0106] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the present invention. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the internal parts of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing in-vehicle call voice, characterized in that: include: When a car call is connected, get the current car call mode; The in-car call mode includes a non-privacy call mode and a privacy call mode; According to the current in-vehicle call mode, a target playback device for the call audio is selected from the playback devices in the vehicle, and the target playback device is controlled to play the call audio; wherein the playback devices include a first playback device and a second playback device; and the distance between the first playback device and the person in the vehicle who is talking on the phone is smaller than the distance between the second playback device and the person in the vehicle who is talking on the phone; If the current in-car call mode is private call mode, obtain the current call encryption level; According to the current call encryption level, controlling the second playback device to play content matching the encryption level; According to the current in-car call mode, a target playback device for the call audio is selected from the playback devices in the car, in the following manner: If the current in-vehicle call mode is a private call mode, playing the call audio using a first playback device; When the current in-vehicle call mode is a non-privacy call mode, the call audio is played using the first playback device and / or the second playback device.

2. The vehicle-mounted call voice processing method according to claim 1, characterized in that: Before controlling the target playback device to play the call audio, the method further includes adjusting the volume of the playback device according to the caller's emotions, call keywords, and response level, specifically: Acquire the emotions of the caller through the voice emotion recognition model; Obtain call keywords and response levels through a neural network model; Adjust the volume of the playback device according to the caller's mood, call keywords, and response level.

3. The method for processing in-vehicle call voice according to claim 1, characterized in that: Before controlling the target playback device to play the call audio, the method further includes: performing call audio extraction processing on the transmitted mixed audio, and performing aliasing reconstruction processing on the call audio after the call audio extraction processing in a non-privacy call mode.

4. The method for processing in-vehicle call voice according to claim 3, wherein: The method for performing call audio extraction processing on the transmitted mixed audio is as follows: Perform noise reduction on the incoming mixed audio; Separate the human voice audio from the mixed audio after noise reduction processing; Perform keyword recognition on the human voice audio and retain sentence audio containing voice call keywords; Perform voiceprint feature recognition on the sentence audio containing the voice call keywords, and output the voiceprint features of the caller; Extract call audio based on the voiceprint features.

5. The method for processing in-vehicle call voice according to claim 4, characterized in that: The method for separating the human voice audio from the mixed audio after noise reduction processing is as follows: Input the noise-reduced mixed audio into the trained separation model and output the human voice audio; The separation model is used to extract human voice audio from a variety of mixed audios.

6. The method for processing in-vehicle call voice according to claim 3, characterized in that: The method for performing aliasing reconstruction processing on the call audio after the call audio extraction processing is as follows: The call audio extracted and processed by the call audio is mixed and reconstructed with the audio being played in the car before the car call is turned on, and the mixed and reconstructed audio is then played through the target playback device.

7. The method for processing in-vehicle call voice according to claim 1, characterized in that: The method of controlling the second playback device to play content matching the encryption level according to the current call encryption level is as follows: When the current call encryption level is level one, controlling the second playback device to play entertainment audio; When the current call encryption level is level 2, controlling the second playback device to play interference audio; The interference audio is entertainment audio whose voiceprint features are consistent with the voiceprint features of the caller.

8. A vehicle-mounted call voice processing system, characterized in that: The in-vehicle call voice processing method according to any one of claims 1 to 7 specifically comprises: Audio receiving device: used to receive the incoming mixed audio; Audio processing module: used to obtain the current in-vehicle call mode when a car call is connected; the in-vehicle call mode includes a non-private call mode and a private call mode; based on the current in-vehicle call mode, select a target playback device for the call audio from the playback devices in the car, and control the target playback device to play the call audio; if the current in-vehicle call mode is the private call mode, obtain the current call encryption level; based on the current call encryption level, control the second playback device to play content that matches the encryption level; Playback device: used for playing audio; the playback device includes a first playback device and a second playback device; the distance between the first playback device and the person talking in the car is smaller than the distance between the second playback device and the person talking in the car.

9. An electronic device, characterized in that: The electronic device comprises: processor and memory; The processor is used to execute the steps of the in-vehicle call voice processing method as described in any one of claims 1 to 7 by calling the program or instructions stored in the memory.

10. A storage medium, characterized in that: The storage medium stores a computer executable program, which is called by a processor to execute the steps of the in-vehicle call voice processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Vehicle-mounted telephone privacy protection method and system and computer readable storage medium

    CN114979994A

  • Vehicle-mounted voice communication method, communication device, electronic equipment and storage medium

    CN117041917A