Audio signal processing method and apparatus, vehicle, and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG GEELY HLDG GRP CO LTD
- Filing Date
- 2026-06-26
- Publication Date
- 2026-08-04
AI Technical Summary
[0004]然而,相关技术中在降低多媒体音乐对导航语音的遮蔽时会存在播放的多媒体音乐的连续性较差
[0036]本申请实施例提供的音频信号处理方法、装置、车辆及存储介质,包括:对车载音频播报系统输出的音乐音频信号进行分解得到多路子音频信号,在检测到车载音频播报系统输出导航音频信号的情况下,根据各子音频信号的声源类别对多路子音频信号中目标声源类别的子音频信号进行衰减处理,生成衰减后子音频信号,并根据衰减后子音频信号、多路子音频信号中未衰减的子音频信号以及导航音频信号生成混合后音频信号,将混合后音频信号发送给车载音频播报系统输出,其中,目标声源类别表示对导航音频信号形成声学干扰的声源类别;上述方法在检测到导航音频信号输出,也即多媒体音乐和导航语音同步播放时,可以基于各路子音频信号对应的声源类别,仅对多路子音频信号中造成声学干扰的目标声源类别的子音频信号进行衰减,其余不产生干扰的子音频信号保持原有音量不变,这样能够精准抑制多媒体音乐中干扰导航语音的声源分量,最大程度保留音乐中无干扰声部的音量、律动与听感,也即最大程度保留多媒体音乐的完整性和连续性,使得多媒体音乐和导航语音同步播放时,降低多媒体音乐对导航语音的遮蔽,避免同步播放过程中音乐整体音量骤降而破坏多媒体音乐的连续性,提高了同步播放过程中多媒体音乐的连续性,能够保障驾驶员清晰接收导航提示信息,兼顾语音播报清晰度与车内音乐聆听体验,提升行车安全性与车载音频使用舒适度;同时,上述方法无需通过多个独立的扬声器同步播放不同声音,就能够解决同步播放会破坏多媒体音乐连续性的问题,降低了解决技术问题所需的硬件成本,并且能够提高该方案的广泛适应性。
Smart Images

Figure CN122511284A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle technology, and in particular to an audio signal processing method, apparatus, vehicle, and storage medium. Background Technology
[0002] With the development of automotive intelligent technology, in-vehicle audio broadcasting systems have become an essential feature of vehicles. In practical applications, users often use in-vehicle audio broadcasting systems to simultaneously play navigation voice prompts and multimedia music to enhance driving enjoyment. When navigation voice prompts and multimedia music are played simultaneously, minimizing the obstruction of navigation voice prompts by multimedia music becomes particularly important to avoid interfering with the navigation voice prompts.
[0003] When navigation voice and multimedia music are played simultaneously, the main technology used is to globally attenuate the audio signal of the multimedia music in order to reduce the obscuring of the navigation voice by the multimedia music.
[0004] However, in related technologies, when reducing the obscuring of navigation voice by multimedia music, the continuity of the played multimedia music is poor. Summary of the Invention
[0005] Therefore, it is necessary to provide an audio signal processing method, apparatus, vehicle, and storage medium to address the aforementioned technical problems.
[0006] In a first aspect, embodiments of this application provide an audio signal processing method, the method comprising:
[0007] The music audio signal output by the vehicle audio broadcasting system is decomposed to obtain multiple sub-audio signals; among them, different sub-audio signals correspond to different sound source categories;
[0008] When the navigation audio signal output by the vehicle audio broadcasting system is detected, the sub-audio signal of the target sound source category in the multi-channel sub-audio signal is attenuated according to the sound source category of each sub-audio signal to generate the attenuated sub-audio signal; the target sound source category indicates the sound source category that causes acoustic interference to the navigation audio signal;
[0009] A mixed audio signal is generated based on the attenuated sub-audio signal, the unattenuated sub-audio signal from the multiple sub-audio signals, and the navigation audio signal, and then sent to the vehicle audio broadcasting system for output.
[0010] In one embodiment, based on the sound source category of each sub-audio signal, the sub-audio signal of the target sound source category in the multi-channel sub-audio signals is attenuated to generate attenuated sub-audio signals, including:
[0011] Based on the sound source category of each sub-audio signal, the sub-audio signals that cause acoustic interference to the navigation audio signal are selected from the multiple sub-audio signals to obtain the sub-audio signals of the target sound source category;
[0012] The sub-audio signals of the target sound source category are attenuated to obtain attenuated sub-audio signals.
[0013] In one embodiment, the sub-audio signal of the target sound source category includes a human voice sub-audio signal and / or a musical instrument sub-audio signal; based on the sound source category of each sub-audio signal, sub-audio signals that cause acoustic interference to the navigation audio signal are filtered from the multiple sub-audio signals to obtain the sub-audio signal of the target sound source category, including:
[0014] For any given sub-audio signal, if the source type of the sub-audio signal is human voice, the sub-audio signal is identified as a human voice sub-audio signal; and / or,
[0015] When the sound source category of the sub-audio signal is an instrument type, the sub-audio signal is identified as a candidate sub-audio signal, and the instrument sub-audio signal is determined based on the candidate sub-audio signal.
[0016] In one embodiment, determining the instrument sub-audio signal based on the candidate sub-audio signals includes:
[0017] Audio components whose frequency bands overlap with those of the navigation audio signal are selected from the candidate sub-audio signals and identified as instrument sub-audio signals.
[0018] In one embodiment, the sub-audio signal of the target sound source category is attenuated to obtain the attenuated sub-audio signal, including:
[0019] Obtain the first attenuation trigger signal of the sub-audio signal of the target sound source category;
[0020] When the first attenuation trigger signal is valid, the sub-audio signal of the target sound source category is attenuated to obtain the attenuated sub-audio signal.
[0021] In one embodiment, determining the instrument sub-audio signal based on the candidate sub-audio signals includes:
[0022] Energy detection is performed on the audio components of the candidate sub-audio signals whose sound source category is a solo instrument to obtain the energy value of the audio components of the solo instrument.
[0023] If the energy value is greater than a preset threshold, the audio components of the candidate sub-audio signals whose sound source category is not a solo instrument are determined as instrument sub-audio signals.
[0024] In one embodiment, the sub-audio signal of the target sound source category is attenuated to obtain the attenuated sub-audio signal, including:
[0025] Acquire the second attenuation trigger signal of the sub-audio signal of the target sound source category;
[0026] When the second attenuation trigger signal is valid, the main frequency band signal in the sub-audio signal of the target sound source category is subjected to gain attenuation processing to generate the attenuated sub-audio signal.
[0027] In one embodiment, the dominant frequency band signal in the sub-audio signal of the target sound source category is subjected to gain attenuation processing to generate an attenuated sub-audio signal, including:
[0028] Dynamic gain attenuation processing is performed on target frequency band signals whose amplitude is greater than a preset value in the main frequency band signal to generate attenuated sub-audio signals.
[0029] Secondly, embodiments of this application provide an audio signal processing apparatus, the apparatus comprising:
[0030] The signal decomposition module is used to decompose the music audio signal output by the vehicle audio broadcasting system into multiple sub-audio signals; among them, different sub-audio signals correspond to different sound source categories;
[0031] The attenuation processing module is used to attenuate the sub-audio signals of the target sound source category among the multiple sub-audio signals according to the sound source category of each sub-audio signal when the navigation audio signal output by the vehicle audio broadcasting system is detected, and generate attenuated sub-audio signals; the target sound source category indicates the sound source category that causes acoustic interference to the navigation audio signal;
[0032] The signal output module is used to generate a mixed audio signal based on the attenuated sub-audio signal, the unattenuated sub-audio signal from the multiple sub-audio signals, and the navigation audio signal, and then send the mixed audio signal to the vehicle audio broadcasting system for output.
[0033] Thirdly, embodiments of this application also provide a vehicle, which includes a memory, a processor, and an in-vehicle audio broadcasting system. The memory stores a computer program, and the processor executes the computer program to implement the steps of the method in any of the embodiments of the first aspect described above.
[0034] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the method in any of the embodiments of the first aspect described above.
[0035] Fifthly, embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the method in any of the embodiments of the first aspect described above.
[0036] The audio signal processing method, apparatus, vehicle, and storage medium provided in this application include: decomposing a music audio signal output by an in-vehicle audio broadcasting system to obtain multiple sub-audio signals; when a navigation audio signal is detected output by the in-vehicle audio broadcasting system, attenuating the sub-audio signals of the target sound source category among the multiple sub-audio signals according to the sound source category of each sub-audio signal to generate attenuated sub-audio signals; generating a mixed audio signal based on the attenuated sub-audio signals, the unattenuated sub-audio signals among the multiple sub-audio signals, and the navigation audio signal; and sending the mixed audio signal to the in-vehicle audio broadcasting system for output. The target sound source category represents the sound source category that causes acoustic interference to the navigation audio signal. When the above method detects the output of the navigation audio signal, i.e., when multimedia music and navigation voice are played synchronously, it can, based on the sound source category corresponding to each sub-audio signal, only process the sub-audio signals of the target sound source category that cause acoustic interference among the multiple sub-audio signals. By attenuating the audio signal while maintaining the original volume of the remaining non-interfering sub-audio signals, the system can precisely suppress the sound source components in multimedia music that interfere with navigation voice, preserving the volume, rhythm, and listening experience of interference-free parts of the music to the greatest extent possible. This means preserving the integrity and continuity of the multimedia music to the greatest extent possible. When multimedia music and navigation voice are played synchronously, the system reduces the obscuring of navigation voice by the multimedia music, preventing a sudden drop in the overall volume of the music during synchronous playback and thus improving its continuity. This ensures that the driver can clearly receive navigation prompts, balancing voice clarity with the in-car music listening experience, enhancing driving safety and the comfort of in-car audio use. Furthermore, this method solves the problem of synchronous playback disrupting the continuity of multimedia music without requiring multiple independent speakers to play different sounds simultaneously, reducing the hardware costs required to solve the technical problem and increasing the applicability of the solution. Attached Figure Description
[0037] Figure 1 This is a flowchart illustrating the audio signal processing method in one embodiment of this application.
[0038] Figure 2 This is a flowchart illustrating an audio signal processing method in another embodiment of this application;
[0039] Figure 3 This is a flowchart illustrating an audio signal processing method in another embodiment of this application;
[0040] Figure 4This is a flowchart illustrating an audio signal processing method in another embodiment of this application;
[0041] Figure 5 This is a flowchart illustrating an audio signal processing method in another embodiment of this application;
[0042] Figure 6 This is a flowchart illustrating an audio signal processing method in another embodiment of this application;
[0043] Figure 7 This is a structural block diagram of an audio signal processing device in one embodiment of this application;
[0044] Figure 8 This is an internal structural diagram of a vehicle in one embodiment of this application. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0046] In the field of in-vehicle audio playback, users often use in-vehicle audio broadcast systems to simultaneously play navigation voice and multimedia music to enhance driving enjoyment. When navigation voice and multimedia music play simultaneously, the mixed audio signal exceeds full amplitude, and the multimedia music masks the navigation voice. Therefore, reducing the masking effect of multimedia music on navigation voice becomes crucial. Related technologies primarily rely on the intermittent nature of navigation voice broadcasts, prioritizing navigation voice. When simultaneous playback of navigation voice and multimedia music is detected, global attenuation processing is applied to the multimedia music audio signal to reduce the amplitude of the mixed audio signal and decrease the masking effect of multimedia music on navigation voice. However, in reducing the masking effect of multimedia music on navigation voice, related technologies sometimes result in fluctuating volume of the multimedia music, significantly diminishing its quality and disrupting its continuity. Therefore, this application provides an audio signal processing method that can reduce the masking effect of multimedia music on navigation voice and improve the continuity of simultaneously played multimedia music.
[0047] The audio signal processing method provided in this application is applicable to any type of vehicle, which includes a processor. The vehicle can be a truck, train, bus, coach, car, etc., and this application does not limit the type. The following embodiments of this application will specifically describe the specific process of the audio signal processing method for an in-vehicle audio broadcasting system, using the processor as the execution subject. The following embodiments will also specifically describe the specific process of the audio signal processing method, using the cloud as the execution subject.
[0048] like Figure 1 The diagram shown is a flowchart of an audio signal processing method provided in an embodiment of this application. The method may include the following steps:
[0049] Step S100: Decompose the music audio signal output by the vehicle audio broadcasting system to obtain multiple sub-audio signals; wherein, different sub-audio signals correspond to different sound source categories.
[0050] In practical applications, the processor can use audio decomposition to decompose the music audio signal output by the vehicle audio broadcasting system to obtain multiple sub-audio signals.
[0051] Among them, the above-mentioned audio decomposition methods can be, but are not limited to, time-frequency texture decomposition, matrix decomposition, and subband transformation decomposition.
[0052] In this embodiment, the processor can acquire a pre-trained audio separation model, then input the music audio signal output from the vehicle audio broadcasting system into the audio separation model, decompose the music audio signal, and output multiple sub-audio signals. Optionally, the above-mentioned audio separation model can be implemented by at least one of a convolutional recurrent neural network model, a Transformer model, or a Conformer model.
[0053] Optionally, different sub-audio signals correspond to different sound source categories, which can be vocal categories and instrument categories (solo categories and non-solo categories). Furthermore, the aforementioned multiple sub-audio signals can include audio signals corresponding to vocal tracks and instrument tracks (such as drum, bass, piano, guitar, etc.) in music.
[0054] It should be noted that decomposing a music audio signal can be understood as the process of separating the music into tracks or the music's sound source.
[0055] Step S200: Upon detecting that the vehicle audio broadcast system outputs a navigation audio signal, the sub-audio signals of the target sound source category among the multiple sub-audio signals are attenuated according to the sound source category of each sub-audio signal, generating attenuated sub-audio signals. Here, the target sound source category represents the sound source category that causes acoustic interference to the navigation audio signal.
[0056] It should be noted that the above-mentioned method of detecting the navigation audio signal output by the vehicle audio broadcasting system can be either the processor receiving the navigation audio signal sent by the vehicle audio broadcasting system, or the processor receiving the broadcast reminder message of the navigation audio signal sent by the vehicle audio broadcasting system.
[0057] Specifically, when the processor detects that the in-vehicle audio broadcasting system outputs navigation audio signals, it can use an attenuation algorithm to attenuate the sub-audio signals of the target sound source category in the multiple sub-audio signals according to the sound source category of each sub-audio signal, and generate attenuated sub-audio signals.
[0058] The sub-audio signals of the aforementioned target sound source category may include at least one sub-audio signal from a set of multiple sub-audio signals. Furthermore, the aforementioned attenuation algorithm may be parametric equalization, graphic equalization, high-pass / low-pass filtering, amplitude limiting, etc.
[0059] Step S300: Generate a mixed audio signal based on the attenuated sub-audio signal, the unattenuated sub-audio signal from the multiple sub-audio signals, and the navigation audio signal, and send the mixed audio signal to the vehicle audio broadcasting system for output.
[0060] In practical applications, the processor can employ an audio signal mixing method to mix or remix the attenuated sub-audio signals, the unattenuated sub-audio signals from multiple sub-audio signals, and the navigation audio signal to generate a mixed audio signal, which is then sent to the vehicle audio broadcasting system for output. Optionally, the aforementioned audio signal mixing method can be, but is not limited to, time-domain addition or frequency-domain addition.
[0061] The technical solution in this application embodiment decomposes the music audio signal output by the vehicle audio broadcasting system into multiple sub-audio signals. When a navigation audio signal is detected output by the vehicle audio broadcasting system, the sub-audio signals of the target sound source category in the multiple sub-audio signals are attenuated according to the sound source category of each sub-audio signal to generate attenuated sub-audio signals. A mixed audio signal is then generated based on the attenuated sub-audio signals, the unattenuated sub-audio signals in the multiple sub-audio signals, and the navigation audio signal. The mixed audio signal is then sent to the vehicle audio broadcasting system for output. Here, the target sound source category represents the sound source category that causes acoustic interference to the navigation audio signal. When the above method detects the output of the navigation audio signal, that is, when multimedia music and navigation voice are played synchronously, it can attenuate only the sub-audio signals of the target sound source category that cause acoustic interference in the multiple sub-audio signals based on the sound source category corresponding to each sub-audio signal, while keeping the original volume of the other sub-audio signals that do not cause interference unchanged. This method precisely suppresses sound source components in multimedia music that interfere with navigation voice, preserving the volume, rhythm, and auditory quality of interference-free parts of the music to the greatest extent possible. In other words, it maximizes the preservation of the integrity and continuity of the multimedia music, avoiding a drop in overall music volume that would disrupt its continuity. When multimedia music and navigation voice are played synchronously, it reduces the obscuring of navigation voice by the multimedia music, preventing a sudden drop in overall music volume that would disrupt the continuity of the multimedia music. This improves the continuity of multimedia music during synchronous playback, ensuring the driver clearly receives navigation prompts while balancing voice clarity with the in-car music listening experience, thus enhancing driving safety and the comfort of using in-car audio. Furthermore, this method solves the problem of synchronous playback disrupting multimedia music continuity without requiring multiple independent speakers to play different sounds simultaneously, reducing the hardware costs required to solve the technical problem, improving the economic adaptability of the solution, and enhancing its broad applicability.
[0062] The following describes the process of attenuating the sub-audio signals of the target sound source category from multiple sub-audio signals according to the sound source category of each sub-audio signal, and generating attenuated sub-audio signals. In one embodiment, as... Figure 2 As shown, the process of attenuating the sub-audio signal of the target sound source category in the multi-channel sub-audio signals according to the sound source category of each sub-audio signal in step S200 above can be implemented in the following way:
[0063] Step S210: Based on the sound source category of each sub-audio signal, filter out the sub-audio signals that cause acoustic interference to the navigation audio signal from the multiple sub-audio signals to obtain the sub-audio signal of the target sound source category.
[0064] Specifically, the processor can acquire a pre-trained algorithm model, and then input the sound source category of each sub-audio signal and the multiple sub-audio signals into the algorithm model. After filtering the sub-audio signals that cause acoustic interference to the navigation audio signal from the multiple sub-audio signals, the sub-audio signal of the target sound source category is obtained.
[0065] Simultaneously, the processor can perform matching, comparison, and / or analysis processing on multiple sub-audio signals according to the sound source category of each sub-audio signal, so as to filter out sub-audio signals that cause acoustic interference to the navigation audio signal from the multiple sub-audio signals and obtain the sub-audio signal of the target sound source category.
[0066] Step S220: Attenuate the sub-audio signal of the target sound source category to obtain the attenuated sub-audio signal.
[0067] Furthermore, the processor can employ an attenuation algorithm to attenuate the sub-audio signals of the target sound source category, generating attenuated sub-audio signals.
[0068] The technical solution in this application embodiment filters out sub-audio signals that cause acoustic interference to the navigation audio signal from multiple sub-audio signals according to the sound source category of each sub-audio signal, obtains the sub-audio signal of the target sound source category, and attenuates the sub-audio signal of the target sound source category to obtain the attenuated sub-audio signal. The above method only attenuates the sub-audio signal of the target sound source category that causes acoustic interference from the multiple sub-audio signals, while the other sub-audio signals that do not cause interference maintain their original volume. This can accurately suppress the sound source components that interfere with the navigation voice in the multimedia music, and preserve the volume, rhythm and listening experience of the non-interfering sound parts in the music to the greatest extent. That is, it preserves the integrity and continuity of the multimedia music to the greatest extent. When the multimedia music and navigation voice are played synchronously, the obscuring of the navigation voice by the multimedia music is reduced, and the overall volume of the music is prevented from dropping suddenly during synchronous playback, which would destroy the continuity of the multimedia music. This improves the continuity of the multimedia music during synchronous playback.
[0069] The process of filtering sub-audio signals that cause acoustic interference to the navigation audio signal from multiple sub-audio signals based on the sound source category of each sub-audio signal, to obtain the sub-audio signal of the target sound source category, will be described below. In one embodiment, the sub-audio signal of the target sound source category includes human voice sub-audio signals and / or instrument sub-audio signals; such as Figure 3 As shown, the process of filtering out sub-audio signals that cause acoustic interference to the navigation audio signal from multiple sub-audio signals according to the sound source category of each sub-audio signal in step S210 above can be achieved in the following way:
[0070] Step S211: For any sub-audio signal, if the sound source category of the sub-audio signal is human voice type, the sub-audio signal is determined to be a human voice sub-audio signal.
[0071] It should be noted that, because the human ear has a high degree of recognition for similar human voice signals, compared to instrumental sounds, the superposition of musical human voices and navigation voices results in the human voices in the music having a masking effect on the navigation voices. Therefore, in this embodiment, human voices are the highest priority target for suppression.
[0072] Specifically, for any given sub-audio signal, the processor can determine that the sub-audio signal is a human voice sub-audio signal if it is determined that the sound source category of the sub-audio signal is human voice type.
[0073] Step S212: If the sound source category of the sub-audio signal is an instrument type, the sub-audio signal is determined as a candidate sub-audio signal, and the instrument sub-audio signal is determined based on the candidate sub-audio signal.
[0074] Simultaneously, at least a portion of the instrument-type sub-audio signals in the multi-channel sub-audio signals can be attenuated. Specifically, when the processor determines that the sound source category of a sub-audio signal is an instrument type, it can identify the sub-audio signal as a candidate sub-audio signal and determine the instrument sub-audio signal based on the candidate sub-audio signal.
[0075] In practical applications, candidate sub-audio signals can be compared, matched, and analyzed, and instrument sub-audio signals can be obtained from multiple sub-audio signals based on the processing results.
[0076] When the sound source category of the above-mentioned sub-audio signal is a musical instrument type, the sub-audio signal can be the audio signal corresponding to non-solo instrument 1, non-solo instrument 2, ..., non-solo instrument m, solo instrument 1, solo instrument 2, ..., or solo instrument n.
[0077] In this application embodiment, two methods can be used to obtain the instrument sub-audio signal, and one of these methods will be described below. In one embodiment, the process of determining the instrument sub-audio signal based on the candidate sub-audio signals in step S212 above may include: filtering audio components whose frequency bands overlap with the frequency bands of the navigation audio signal from the candidate sub-audio signals, and determining them as the instrument sub-audio signals.
[0078] Specifically, the processor can invoke a signal filtering tool and then input both the candidate sub-audio signal and the navigation audio signal into the signal filtering tool to filter out audio components whose frequency bands overlap with those of the navigation audio signal from the candidate sub-audio signal, as instrument sub-audio signals.
[0079] Additionally, the processor can acquire a pre-trained signal filtering model, and then input both the candidate sub-audio signals and the navigation audio signals into the signal filtering model. This allows for the selection of audio components whose frequency bands overlap with the navigation audio signal from the candidate sub-audio signals, resulting in the output of the instrument sub-audio signal. Optionally, the aforementioned signal filtering model can be implemented using at least one of the following: a convolutional neural network model, a recurrent recurrent neural network model, a fully connected neural network model, a residual neural network model, or a long short-term memory neural network model.
[0080] Optionally, the audio components in the candidate sub-audio signals whose frequency bands overlap with the navigation audio signal's frequency band may include audio components of instruments such as the erhu, drums, bass, piano, and guitar. Specifically, the audio components in the candidate sub-audio signals whose frequency bands overlap with the navigation audio signal's frequency band may include audio components whose frequency bands overlap with the navigation audio signal's frequency band, and may also include audio components whose frequency bands overlap with the navigation audio signal's frequency band, as well as audio components in a preset small frequency band before and after the overlapping frequency band.
[0081] The following describes another process for acquiring sub-audio signals from musical instruments. In one embodiment, as... Figure 4 As shown, the process of determining the instrument sub-audio signal based on the candidate sub-audio signal in step S212 above can be implemented in the following way:
[0082] Step S2121: Perform energy detection on the audio component of the candidate sub-audio signal whose sound source category is a solo instrument to obtain the energy value of the audio component of the solo instrument.
[0083] The processor can use energy detection to detect the energy of candidate sub-audio signals and obtain the energy value of the audio components of the solo instrument.
[0084] Optionally, the energy detection method described above can be a sliding window energy detection method, an instantaneous amplitude calculation method, a short-time average amplitude calculation method, a mean square value calculation method, a short-time average energy calculation method, a specified sub-band energy calculation method, a fixed threshold detection method, etc.
[0085] Step S2122: If the energy value is greater than a preset threshold, the audio components of the candidate sub-audio signals whose sound source category is non-solo instrument are determined as instrument sub-audio signals.
[0086] In practical applications, the processor can determine whether the energy value of the audio component of a solo instrument is greater than a preset threshold. If the energy value of the audio component of a solo instrument is greater than the preset threshold, in order to ensure that the rhythm and melody of the multimedia music are not interrupted, the audio component of the candidate sub-audio signal whose sound source category is not a solo instrument can be identified as the instrument sub-audio signal.
[0087] The aforementioned solo instruments can include drums, bass, and some melodic instruments, while non-solo instruments can be any instrument other than a solo instrument. Optionally, the aforementioned preset threshold can be user-defined or determined based on historical experience values.
[0088] The technical solution in this application embodiment, for any sub-audio signal, if the sound source category of the sub-audio signal is human voice, the sub-audio signal is determined to be a human voice sub-audio signal; if the sound source category of the sub-audio signal is musical instrument, the sub-audio signal is determined to be a candidate sub-audio signal, and the musical instrument sub-audio signal is determined based on the candidate sub-audio signal. In the process of synchronous playback of multimedia music and navigation voice, the above method can filter sub-audio signals that cause acoustic interference to the navigation audio signal from multiple sub-audio signals according to the sound source category of the sub-audio signal, so as to avoid false attenuation and damage to the continuity of music in synchronous playback audio, thereby improving the continuity of synchronously played multimedia music.
[0089] The following describes how the attenuation process is triggered in one scenario. In one embodiment, as... Figure 5 As shown, the attenuation processing of the sub-audio signal of the target sound source category in step S220 above may include:
[0090] Step S221: Obtain the first attenuation trigger signal of the sub-audio signal of the target sound source category.
[0091] In practical applications, after the processor filters out the human voice sub-audio signal and the instrument sub-audio signal from the multiple sub-audio signals (that is, it selects the audio components whose frequency bands overlap with the frequency bands of the navigation audio signal from the candidate sub-audio signals to determine the instrument sub-audio signal), it can generate the first attenuation trigger signal corresponding to the sub-audio signal of the target sound source category.
[0092] Step S222: When the first attenuation trigger signal is a valid signal, the sub-audio signal of the target sound source category is attenuated to obtain the attenuated sub-audio signal.
[0093] Specifically, the processor can determine whether the first attenuation trigger signal is a valid signal. If the first attenuation trigger signal is determined to be a valid signal, a gain attenuation algorithm is used to attenuate the sub-audio signal of the target sound source category to obtain the attenuated sub-audio signal.
[0094] Optionally, the above gain attenuation algorithm can be a linear fixed gain attenuation method, a filter-type gain attenuation method, a frequency domain gain attenuation method, a track-by-track gain attenuation method, etc.
[0095] Alternatively, the processor can acquire a pre-trained gain attenuation model, then input the sub-audio signal of the target sound source category into the gain attenuation model, which performs gain attenuation processing on the sub-audio signal of the target sound source category and outputs the attenuated sub-audio signal.
[0096] Optionally, the above-mentioned gain attenuation model may include, but is not limited to, deep learning neural network models, embedded logic processing network models, passive attenuation network models, and / or active attenuation network models.
[0097] The following describes how the attenuation process is triggered in another scenario. In one embodiment, as... Figure 6 As shown, the process of attenuating the sub-audio signal of the target sound source category to obtain the attenuated sub-audio signal in step S220 above can be achieved in the following way:
[0098] Step S223: Obtain the second attenuation trigger signal of the sub-audio signal of the target sound source category.
[0099] In practical applications, after the processor selects the sub-audio signal of the target sound source category from multiple sub-audio signals, it can generate a second attenuation trigger signal corresponding to the sub-audio signal of the target sound source category.
[0100] It should be noted that if the target sound source category is multiple sound source categories, the second attenuation trigger signal corresponding to the sub-audio signal of the target sound source category can include multiple sub-trigger signals, and each sub-trigger signal can control the attenuation processing of the sub-audio signal of one sound source category.
[0101] Step S224: When the second attenuation trigger signal is a valid signal, the gain attenuation processing is performed on the main frequency band signal in the sub-audio signal of the target sound source category to generate the attenuated sub-audio signal.
[0102] Specifically, the processor can determine whether the second attenuation trigger signal is a valid signal. If the second attenuation trigger signal is determined to be a valid signal, the processor can perform gain attenuation processing on the main frequency band signal in the sub-audio signal of the target sound source category based on the sidechain equalizer to generate the attenuated sub-audio signal.
[0103] One method for performing gain attenuation processing on the main frequency band signal in the sub-audio signal of the target sound source category can be to use a gain attenuation algorithm to perform gain attenuation processing on the main frequency band signal that causes acoustic interference to the navigation audio signal in the sub-audio signal of the target sound source category, thereby generating an attenuated sub-audio signal.
[0104] Optionally, the above-mentioned gain attenuation algorithm can be a linear fixed gain attenuation method, a filter-type gain attenuation method, a frequency domain gain attenuation method, a track-by-track gain attenuation method, etc. In the embodiments of this application, the above-mentioned main frequency band signal can be a mid-frequency band signal.
[0105] Meanwhile, another way to perform gain attenuation processing on the main frequency band signal in the sub-audio signal of the target sound source category is to obtain a pre-trained gain attenuation model, and then input both the sub-audio signal of the target sound source category and the navigation audio signal into the gain attenuation model. After performing gain attenuation processing on the main frequency band signal in the sub-audio signal of the target sound source category, the gain attenuation model outputs the attenuated sub-audio signal.
[0106] Optionally, the above-mentioned gain attenuation model may include, but is not limited to, deep learning neural network models, embedded logic processing network models, passive attenuation network models, and / or active attenuation network models.
[0107] In one embodiment, the process of performing gain attenuation processing on the main frequency band signal in the sub-audio signal of the target sound source category in step S224 above may include: performing dynamic gain attenuation processing on the target frequency band signal whose signal amplitude is greater than a preset amplitude in the main frequency band signal to generate the attenuated sub-audio signal.
[0108] In practical applications, a multi-segment dynamic range control method can be used to dynamically attenuate the main frequency band signal in the sub-audio signal of the target sound source category. Specifically, the processor can dynamically attenuate the target frequency band signal whose signal amplitude is greater than a preset value in real time, generating an attenuated sub-audio signal.
[0109] The target frequency band signal whose amplitude is greater than the preset amplitude can be a discontinuous multi-segment audio signal in the main frequency band signal, or a continuous multi-segment audio signal in the main frequency band signal. This application embodiment does not limit this.
[0110] The technical solution in this application embodiment obtains a second attenuation trigger signal for the sub-audio signal of the target sound source category. When the second attenuation trigger signal is a valid signal, the gain attenuation processing is performed on the main frequency band signal in the sub-audio signal of the target sound source category to generate the attenuated sub-audio signal. In the process of synchronous playback of multimedia music and navigation voice, the above method only performs audio signal attenuation processing when the attenuation processing is triggered, which can improve the effectiveness of attenuation processing. At the same time, the above method only attenuates the core frequency band signal, i.e., the main frequency band signal, in the music audio signal that causes acoustic interference to the navigation audio signal. This can preserve the volume, rhythm, and listening experience of the interference-free parts of the music to a great extent, that is, preserve the integrity and continuity of the multimedia music to the greatest extent. When multimedia music and navigation voice are played synchronously, the obscuring of navigation voice by multimedia music is reduced, and the overall volume of the music is prevented from dropping suddenly during synchronous playback, thus disrupting the continuity of the multimedia music and improving the continuity of multimedia music during synchronous playback.
[0111] In one embodiment, this application also provides an audio signal processing method applied to a processor in a vehicle, the method comprising the following steps:
[0112] (1) Decompose the music audio signal output by the vehicle audio broadcasting system to obtain multiple sub-audio signals; among which, different sub-audio signals correspond to different sound source categories;
[0113] (2) When the navigation audio signal is detected to be output by the vehicle audio broadcasting system, for any sub-audio signal, if the sound source category of the sub-audio signal is human voice, the sub-audio signal is determined to be a human voice sub-audio signal; and / or, if the sound source category of the sub-audio signal is musical instrument, the sub-audio signal is determined to be a candidate sub-audio signal, and the energy of the audio component with the sound source category of a solo instrument in the candidate sub-audio signal is detected to obtain the energy value of the audio component of the solo instrument. If the energy value is greater than a preset threshold, the audio component with the sound source category of a non-solo instrument in the candidate sub-audio signal is determined to be a musical instrument sub-audio signal; the target sound source category indicates the sound source category that causes acoustic interference to the navigation audio signal;
[0114] (3) Obtain the first attenuation trigger signal of the human voice sub-audio signal and / or the instrument sub-audio signal, and when the first attenuation trigger signal is a valid signal, perform attenuation processing on the sub-audio signal of the target sound source category to obtain the attenuated sub-audio signal;
[0115] (4) Obtain the second attenuation trigger signal of the human voice sub-audio signal and / or the instrument sub-audio signal, and when the second attenuation trigger signal is a valid signal, perform gain attenuation processing on the main frequency band signal that forms acoustic interference with the navigation audio signal in the human voice sub-audio signal and / or the instrument sub-audio signal according to the dynamic gain attenuation condition that the signal amplitude is greater than the preset amplitude, and generate the attenuated sub-audio signal.
[0116] (5) Generate a mixed audio signal based on the attenuated sub-audio signal, the unattenuated sub-audio signal in the multi-channel sub-audio signal and the navigation audio signal, and send the mixed audio signal to the vehicle audio broadcasting system for output.
[0117] For details of the execution process of (1) to (5) above, please refer to the description of the above embodiments. The implementation principle and technical effect are similar, and will not be repeated here.
[0118] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0119] Based on the same inventive concept, this application also provides an audio signal processing apparatus for implementing the audio signal processing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more audio signal processing apparatus embodiments provided below can be found in the limitations of the audio signal processing method described above, and will not be repeated here.
[0120] In one embodiment, Figure 7 This is a schematic diagram of an audio signal processing device in one embodiment of this application. The audio signal processing device provided in this embodiment can be applied to a processor. Figure 7 As shown, the audio signal processing apparatus of this application embodiment may include: a signal decomposition module 11, an attenuation processing module 12, and a signal output module 13, wherein:
[0121] The signal decomposition module 11 is used to decompose the music audio signal output by the vehicle audio broadcasting system to obtain multiple sub-audio signals; wherein, different sub-audio signals correspond to different sound source categories;
[0122] The attenuation processing module 12 is used to attenuate the sub-audio signals of the target sound source category among the multiple sub-audio signals according to the sound source category of each sub-audio signal when the navigation audio signal output by the vehicle audio broadcasting system is detected, and generate attenuated sub-audio signals; the target sound source category indicates the sound source category that causes acoustic interference to the navigation audio signal.
[0123] The signal output module 13 is used to generate a mixed audio signal based on the attenuated sub-audio signal, the unattenuated sub-audio signal in the multi-channel sub-audio signal and the navigation audio signal, and send the mixed audio signal to the vehicle audio broadcasting system for output.
[0124] The audio signal processing apparatus provided in this application embodiment can be used to execute the technical solutions in the above-described audio signal processing method embodiments of this application. Its implementation principle and technical effects are similar, and will not be repeated here.
[0125] In one embodiment, the attenuation processing module 12 includes: a filtering unit and an attenuation processing unit, wherein:
[0126] The filtering unit is used to filter out sub-audio signals that cause acoustic interference to the navigation audio signal from multiple sub-audio signals according to the sound source category of each sub-audio signal, so as to obtain the sub-audio signal of the target sound source category.
[0127] The attenuation processing unit is used to attenuate the sub-audio signals of the target sound source category to obtain attenuated sub-audio signals.
[0128] The audio signal processing apparatus provided in this application embodiment can be used to execute the technical solutions in the above-described audio signal processing method embodiments of this application. Its implementation principle and technical effects are similar, and will not be repeated here.
[0129] In one embodiment, the sub-audio signals of the target sound source category include human voice sub-audio signals and / or instrument sub-audio signals; the filtering unit includes: a first determining sub-unit and a second determining sub-unit, wherein:
[0130] The first determining subunit is configured to, for any given sub-audio signal, determine the sub-audio signal as a human voice sub-audio signal if the source type of the sub-audio signal is human voice; and / or,
[0131] The second determining subunit is used to determine the sub-audio signal as a candidate sub-audio signal when the sound source category of the sub-audio signal is an instrument type, and to determine the instrument sub-audio signal based on the candidate sub-audio signal.
[0132] The audio signal processing apparatus provided in this application embodiment can be used to execute the technical solutions in the above-described audio signal processing method embodiments of this application. Its implementation principle and technical effects are similar, and will not be repeated here.
[0133] In one embodiment, the second determining subunit includes: a filtering subunit, wherein:
[0134] The filtering subunit is used to filter out audio components whose frequency bands overlap with those of the navigation audio signal from the candidate sub-audio signals, and identify them as instrument sub-audio signals.
[0135] The audio signal processing apparatus provided in this application embodiment can be used to execute the technical solutions in the above-described audio signal processing method embodiments of this application. Its implementation principle and technical effects are similar, and will not be repeated here.
[0136] In one embodiment, the attenuation processing unit is specifically used for:
[0137] Obtain the first attenuation trigger signal of the sub-audio signal of the target sound source category;
[0138] When the first attenuation trigger signal is valid, the sub-audio signal of the target sound source category is attenuated to obtain the attenuated sub-audio signal.
[0139] The audio signal processing apparatus provided in this application embodiment can be used to execute the technical solutions in the above-described audio signal processing method embodiments of this application. Its implementation principle and technical effects are similar, and will not be repeated here.
[0140] In one embodiment, the second determining subunit includes: an energy detection subunit and a signal extraction subunit, wherein:
[0141] The energy detection subunit is used to detect the energy of the audio components of the candidate sub-audio signals whose sound source category is a solo instrument, and to obtain the energy value of the audio components of the solo instrument.
[0142] The signal extraction subunit is used to identify the audio components of the candidate sub-audio signals whose sound source category is not a solo instrument as instrument sub-audio signals when the energy value is greater than a preset threshold.
[0143] The audio signal processing apparatus provided in this application embodiment can be used to execute the technical solutions in the above-described audio signal processing method embodiments of this application. Its implementation principle and technical effects are similar, and will not be repeated here.
[0144] In one embodiment, the attenuation processing unit includes: an acquisition subunit and an attenuation processing subunit, wherein:
[0145] The acquisition subunit is used to acquire the second attenuation trigger signal of the sub-audio signal of the target sound source category;
[0146] The attenuation processing subunit is used to perform gain attenuation processing on the main frequency band signal in the sub-audio signal of the target sound source category when the second attenuation trigger signal is an effective signal, so as to generate the attenuated sub-audio signal.
[0147] The audio signal processing apparatus provided in this application embodiment can be used to execute the technical solutions in the above-described audio signal processing method embodiments of this application. Its implementation principle and technical effects are similar, and will not be repeated here.
[0148] For specific limitations regarding the audio signal processing device, please refer to the limitations on the audio signal processing method above, which will not be repeated here. Each module in the aforementioned audio signal processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the vehicle's processor, or stored in software in the vehicle's memory, so that the processor can call and execute the corresponding operations of each module.
[0149] In one embodiment, a vehicle is provided, the internal structure diagram of which can be as follows: Figure 8 As shown, the vehicle includes an in-vehicle audio broadcasting system, a processor connected via a system bus, memory, and a network interface. The vehicle's processor provides processing power. The vehicle's memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The vehicle's database stores music and navigation audio signals output by the in-vehicle audio broadcasting system. The vehicle's network interface is used for communication with external endpoints via a network connection. When the computer program is executed by the processor, it implements an audio signal processing method.
[0150] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the vehicle to which the present application is applied. A specific vehicle may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0151] In one embodiment, a vehicle is also provided, including a memory, a processor, and an in-vehicle audio broadcasting system, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0152] The music audio signal output by the vehicle audio broadcasting system is decomposed to obtain multiple sub-audio signals; among them, different sub-audio signals correspond to different sound source categories;
[0153] When the navigation audio signal output by the vehicle audio broadcasting system is detected, the sub-audio signal of the target sound source category in the multi-channel sub-audio signal is attenuated according to the sound source category of each sub-audio signal to generate the attenuated sub-audio signal; the target sound source category indicates the sound source category that causes acoustic interference to the navigation audio signal;
[0154] A mixed audio signal is generated based on the attenuated sub-audio signal, the unattenuated sub-audio signal from the multiple sub-audio signals, and the navigation audio signal, and then sent to the vehicle audio broadcasting system for output.
[0155] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0156] The music audio signal output by the vehicle audio broadcasting system is decomposed to obtain multiple sub-audio signals; among them, different sub-audio signals correspond to different sound source categories;
[0157] When the navigation audio signal output by the vehicle audio broadcasting system is detected, the sub-audio signal of the target sound source category in the multi-channel sub-audio signal is attenuated according to the sound source category of each sub-audio signal to generate the attenuated sub-audio signal; the target sound source category indicates the sound source category that causes acoustic interference to the navigation audio signal;
[0158] A mixed audio signal is generated based on the attenuated sub-audio signal, the unattenuated sub-audio signal from the multiple sub-audio signals, and the navigation audio signal, and then sent to the vehicle audio broadcasting system for output.
[0159] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:
[0160] The music audio signal output by the vehicle audio broadcasting system is decomposed to obtain multiple sub-audio signals; among them, different sub-audio signals correspond to different sound source categories;
[0161] When the navigation audio signal output by the vehicle audio broadcasting system is detected, the sub-audio signal of the target sound source category in the multi-channel sub-audio signal is attenuated according to the sound source category of each sub-audio signal to generate the attenuated sub-audio signal; the target sound source category indicates the sound source category that causes acoustic interference to the navigation audio signal;
[0162] A mixed audio signal is generated based on the attenuated sub-audio signal, the unattenuated sub-audio signal from the multiple sub-audio signals, and the navigation audio signal, and then sent to the vehicle audio broadcasting system for output.
[0163] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0164] The technical features in the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0165] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method of audio signal processing, characterized by, The method includes: The music audio signal output by the vehicle audio broadcasting system is decomposed to obtain multiple sub-audio signals; among them, different sub-audio signals correspond to different sound source categories; When the vehicle audio broadcasting system detects that it outputs a navigation audio signal, the sub-audio signal of the target sound source category in the multi-channel sub-audio signal is attenuated according to the sound source category of each sub-audio signal to generate an attenuated sub-audio signal; the target sound source category indicates the sound source category that causes acoustic interference to the navigation audio signal. A mixed audio signal is generated based on the attenuated sub-audio signal, the unattenuated sub-audio signal from the multi-channel sub-audio signal, and the navigation audio signal, and the mixed audio signal is sent to the vehicle audio broadcasting system for output.
2. The method of claim 1, wherein, The step of attenuating the sub-audio signals of the target sound source category among the multiple sub-audio signals according to the sound source category of each sub-audio signal to generate attenuated sub-audio signals includes: Based on the sound source category of each of the sub-audio signals, sub-audio signals that cause acoustic interference to the navigation audio signal are filtered from the multiple sub-audio signals to obtain the sub-audio signal of the target sound source category; The sub-audio signal of the target sound source category is attenuated to obtain the attenuated sub-audio signal.
3. The method of claim 2, wherein, The sub-audio signals of the target sound source category include human voice sub-audio signals and / or instrument sub-audio signals; the step of filtering sub-audio signals that cause acoustic interference to the navigation audio signal from the multiple sub-audio signals according to the sound source category of each sub-audio signal to obtain the sub-audio signals of the target sound source category includes: For any sub-audio signal, if the sound source category of the sub-audio signal is human voice type, the sub-audio signal is identified as the human voice sub-audio signal; and / or, If the sound source category of the sub-audio signal is a musical instrument, the sub-audio signal is identified as a candidate sub-audio signal, and the musical instrument sub-audio signal is determined based on the candidate sub-audio signal.
4. The method of claim 3, wherein, Determining the instrument sub-audio signal based on the candidate sub-audio signals includes: Audio components whose frequency bands overlap with those of the navigation audio signal are selected from the candidate sub-audio signals and identified as the instrument sub-audio signals.
5. The method of claim 4, wherein, The attenuation process of the sub-audio signal of the target sound source category to obtain the attenuated sub-audio signal includes: Obtain the first attenuation trigger signal of the sub-audio signal of the target sound source category; When the first attenuation trigger signal is a valid signal, the sub-audio signal of the target sound source category is attenuated to obtain the attenuated sub-audio signal.
6. The method of claim 3, wherein, Determining the instrument sub-audio signal based on the candidate sub-audio signals includes: Energy detection is performed on the audio components of the candidate sub-audio signals whose sound source category is a solo instrument to obtain the energy value of the audio components of the solo instrument. If the energy value is greater than a preset threshold, the audio component of the candidate sub-audio signal whose sound source category is a non-solo instrument is determined as the instrument sub-audio signal.
7. The method of claim 6, wherein, The attenuation process of the sub-audio signal of the target sound source category to obtain the attenuated sub-audio signal includes: Obtain the second attenuation trigger signal of the sub-audio signal of the target sound source category; When the second attenuation trigger signal is a valid signal, the main frequency band signal in the sub-audio signal of the target sound source category is subjected to gain attenuation processing to generate the attenuated sub-audio signal.
8. The method of claim 7, wherein, The step of performing gain attenuation processing on the dominant frequency band signal in the sub-audio signal of the target sound source category to generate the attenuated sub-audio signal includes: Dynamic gain attenuation processing is performed on the target frequency band signal whose signal amplitude is greater than a preset value in the main frequency band signal to generate the attenuated sub-audio signal.
9. An audio signal processing apparatus, characterized by comprising: The device includes: The signal decomposition module is used to decompose the music audio signal output by the vehicle audio broadcasting system into multiple sub-audio signals; among them, different sub-audio signals correspond to different sound source categories; The attenuation processing module is used to attenuate the sub-audio signals of the target sound source category among the multiple sub-audio signals according to the sound source category of each sub-audio signal when the navigation audio signal is detected to be output by the vehicle audio broadcasting system, thereby generating attenuated sub-audio signals; the target sound source category represents the sound source category that causes acoustic interference to the navigation audio signal; The signal output module is used to generate a mixed audio signal based on the attenuated sub-audio signal, the unattenuated sub-audio signal among the multiple sub-audio signals, and the navigation audio signal, and send the mixed audio signal to the vehicle audio broadcasting system for output.
10. A vehicle comprising a memory, a processor and an on-board audio announcement system, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-8.