Audio processing method, storage medium, program product, electronic device, and vehicle

By using sound pickup devices and separation model technology, the problem of poor user experience in karaoke systems has been solved, enabling users to enjoy karaoke entertainment at any time and with any sound source, thus improving user experience and driving pleasure.

WO2026153188A1PCT designated stage Publication Date: 2026-07-23BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BYD CO LTD
Filing Date
2026-01-06
Publication Date
2026-07-23

Smart Images

  • Figure CN2026070895_23072026_PF_FP_ABST
    Figure CN2026070895_23072026_PF_FP_ABST
Patent Text Reader

Abstract

An audio processing method, a storage medium, a program product, an electronic device, and a vehicle. The method comprises: when a music audio is played back, detecting an audio signal of a user, and outputting an accompaniment audio in the music audio and the audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing methods, storage media, software products, electronic devices, and vehicles

[0001] This application claims priority to Chinese Patent Application No. 202510061444.4, filed on January 14, 2025, entitled "Audio Processing Method, Storage Medium, Program Product, Electronic Device and Vehicle", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to, but is not limited to, the field of audio processing technology, specifically to an audio processing method, storage medium, program product, electronic device, and vehicle. Background Technology

[0003] In modern entertainment systems, karaoke offers a novel form of interactive entertainment. Users sing through a microphone or a built-in microphone in the space, and the singing is simply mixed into the audio output of a dedicated application or hardware device, resulting in a poor user experience. Technical solutions

[0004] This application provides an audio processing method, storage medium, program product, electronic device, and vehicle to address the problem of poor user experience in karaoke. The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.

[0005] To achieve the above objectives, according to a first aspect of this application, an audio processing method is provided, comprising:

[0006] When playing music audio, the system detects the user's audio signal and outputs the accompaniment audio from the music audio along with the audio signal.

[0007] Optionally, when playing music audio, detecting the user's audio signal and outputting the accompaniment audio in the music audio along with the audio signal includes:

[0008] Identify the user's audio signal; and

[0009] If the audio signal is a singing signal, then the accompaniment audio in the music audio and the audio signal are output.

[0010] Optionally, identifying the user's audio signal includes:

[0011] Based on the audio signal and the source human voice audio corresponding to the music audio, determine whether the audio signal is a singing signal.

[0012] Optionally, determining whether the audio signal is a singing signal based on the audio signal and the source human voice audio corresponding to the music audio includes:

[0013] Based on the similarity between the audio features of the audio signal and the source human voice audio corresponding to the music audio, it is determined whether the audio signal is a singing signal.

[0014] Optionally, before outputting the accompaniment audio in the music audio and the audio signal, the method includes:

[0015] The music audio is input into a separation model for processing to obtain accompaniment audio and / or source vocal audio.

[0016] Optionally, the step of inputting the music audio into a separation model for processing to obtain accompaniment audio includes:

[0017] The music audio is input into the separation model for processing to obtain the source human voice audio:

[0018] The accompaniment audio is obtained based on the difference between the music audio and the source vocal audio.

[0019] Optionally, the step of inputting the music audio into a separation model for processing to obtain accompaniment audio and / or source vocal audio includes:

[0020] Extract the music audio of the first preset frame length, input it into the separation model for processing, and obtain the accompaniment audio and / or the source vocal audio of the first preset frame length.

[0021] Optionally, the method further includes:

[0022] When outputting the accompaniment audio of the first preset frame length, the music audio of the second preset frame length is extracted and input into the separation model for processing to obtain the accompaniment audio and / or the source vocal audio of the second preset frame length.

[0023] Optionally, the extraction of music audio of the second preset frame length includes:

[0024] After extracting the music audio of the first preset frame length, the second preset frame length is determined in order to extract the music audio of the second preset frame length.

[0025] Optionally, the time taken for the separation model to process the music audio of the second preset frame length is less than the playback time of the accompaniment audio of the first preset frame length.

[0026] Optionally, the separation model is trained in the following manner:

[0027] Obtain sample audio data, wherein the sample audio data includes corresponding sample accompaniment audio and / or sample source vocal audio; and

[0028] Based on the sample audio data, the trained separation model is obtained through the separation model to be trained.

[0029] Optionally, the method further includes:

[0030] Output the accompaniment audio, the original vocal audio, and the audio signal from the music audio.

[0031] Optionally, the method further includes:

[0032] An operation to adjust the volume of the source human voice was detected, and the volume of the output source human voice audio was adjusted.

[0033] Optionally, the method further includes:

[0034] In response to user control commands, determine the audio mode; and

[0035] Output music audio based on the audio pattern.

[0036] Optionally, the audio mode includes a first audio mode, and the step of outputting music audio according to the audio mode includes:

[0037] In response to being in the first audio mode, when playing music audio, the user's audio signal is detected, and the accompaniment audio in the music audio and the audio signal are output.

[0038] Optionally, the method further includes:

[0039] During the process of outputting the accompaniment audio in the music audio and the audio signal, if no user audio signal is detected, the accompaniment audio in the music audio or the music audio is output.

[0040] Optionally, the step of outputting either the accompaniment audio or the music audio in response to the absence of a user's audio signal includes:

[0041] Based on the duration during which no user's audio signal is detected, output either the accompaniment audio in the music audio or the music audio itself.

[0042] Optionally, the step of outputting the accompaniment audio or the music audio based on the duration for which no user's audio signal is detected includes:

[0043] In response to the duration being less than a first preset duration, the accompaniment audio in the music audio is output.

[0044] Optionally, the step of outputting the accompaniment audio or the music audio based on the duration for which no user's audio signal is detected includes:

[0045] In response to the duration being greater than or equal to a first preset duration, the music audio is output.

[0046] Optionally, the step of outputting the music audio in response to the duration being greater than or equal to a first preset duration includes:

[0047] In response to the absence of the user's audio signal within the first preset duration, and the presence of source vocal audio corresponding to the accompaniment audio in the music audio, the music audio is output.

[0048] Optionally, the source vocal audio corresponding to the accompaniment audio in the music audio includes:

[0049] The energy of the music audio is calculated. If the energy of the music audio is less than a preset energy threshold, then it is determined that there is a source human voice audio corresponding to the accompaniment audio in the music audio.

[0050] Optionally, the audio mode includes a second audio mode, and the step of outputting music audio according to the audio mode includes:

[0051] In response to being in the second audio mode, the accompaniment audio in the music audio is output.

[0052] Optionally, the method further includes:

[0053] During the process of outputting the accompaniment audio in the music audio, in response to detecting the user's audio signal, the accompaniment audio in the music audio and the audio signal are output.

[0054] Optionally, the method further includes:

[0055] In response to detecting the user's audio signal, the accompaniment audio in the music audio stored in the first buffer and the audio signal are output.

[0056] Optionally, before the accompaniment audio in the music audio stored in the first buffer and the audio signal are output, the method further includes:

[0057] Store the accompaniment audio from the music audio into the first buffer area.

[0058] Optionally, storing the accompaniment audio from the music audio into the first buffer includes:

[0059] Store the accompaniment audio from the music audio with a preset frame length into the first buffer area.

[0060] Optionally, the method further includes:

[0061] Upon detecting a progress bar adjustment, output the accompaniment audio from the music audio at the corresponding moment after the operation, indicating the progress bar's position.

[0062] Optionally, after the output operation, the position of the progress bar corresponds to the time before the accompaniment audio in the music audio. The method further includes:

[0063] Upon detecting the progress bar adjustment operation, determine whether the accompaniment audio from the music audio corresponding to the position of the progress bar after the operation exists in the first buffer:

[0064] In response to the presence of accompaniment audio in the music audio corresponding to the position of the progress bar after the operation in the first buffer, the music audio after the position is extracted based on the preset frame length corresponding to the initial position; and

[0065] In response to the absence of the accompaniment audio in the music audio corresponding to the position of the progress bar after the operation in the first buffer, the music audio is extracted again according to the time sequence and based on the preset frame length corresponding to the current time.

[0066] Optionally, the method further includes:

[0067] The music audio is processed to obtain music audio in a preset format.

[0068] Optionally, processing the music audio to obtain music audio in a preset format includes:

[0069] The music audio is converted into music audio with a preset number of channels.

[0070] Optionally, processing the music audio to obtain music audio in a preset format includes:

[0071] The music audio is converted into music audio with a preset sampling rate.

[0072] Optionally, the method further includes:

[0073] The music audio in the preset format is stored in the second buffer.

[0074] Optionally, the method further includes:

[0075] If the user is detected to have disabled or enabled the singing function, a synthesized audio is output based on a preset fade-in / fade-out strategy. The synthesized audio is a combination of at least two audio components: the music audio, the accompaniment audio in the music audio, and the audio signal.

[0076] Optionally, the output of synthesized audio based on a preset fade-in / fade-out strategy includes...

[0077] Based on the preset fade-in / fade-out strategy, the second preset duration is determined;

[0078] Within the second preset duration, each frame of music audio with a preset frame length and its corresponding target audio with a preset frame length are obtained, wherein the target audio is the accompaniment audio in the music audio or an audio synthesized based on the accompaniment audio in the music audio and the audio signal; and

[0079] The music audio and the corresponding target audio are weighted and summed to output the synthesized audio.

[0080] Optionally, the step of weighted summing of the music audio and the corresponding target audio to output the synthesized audio includes:

[0081] In response to the user activating the singing function, a weighted sum is performed based on the first weight and the music audio, and the second weight and the corresponding target audio, to output the synthesized audio.

[0082] The first weight is negatively correlated with the frame count of the music audio with the preset frame length in the second preset duration, and the second weight is positively correlated with the frame count of the music audio with the preset frame length in the second preset duration.

[0083] Optionally, the method further includes:

[0084] In response to the user disabling the singing function, a weighted sum is performed based on the third weight and the music audio, and the fourth weight and the corresponding target audio, to output the synthesized audio.

[0085] The third weight is positively correlated with the frame count of the music audio with the preset frame length in the second preset duration, and the fourth weight is negatively correlated with the frame count of the music audio with the preset frame length in the second preset duration.

[0086] Optionally, the music audio is obtained through at least one of music applications, video applications, storage devices, and user terminals.

[0087] According to a second aspect of this application, embodiments of this application also provide a computer-readable storage medium storing instructions that, when executed by a computer, cause the computer to implement any of the audio processing methods described in the embodiments of this application.

[0088] According to a third aspect of this application, embodiments of this application also provide a computer program product storing instructions that, when executed by a computer, cause the computer to implement any of the audio processing methods described in the embodiments of this application.

[0089] According to a fourth aspect of this application, embodiments of this application also provide an electronic device, comprising:

[0090] Memory, on which computer instructions are stored;

[0091] A processor is configured to execute the computer instructions in the memory to implement any of the audio processing methods provided in the embodiments of this application.

[0092] According to a fifth aspect of this application, embodiments of this application also provide a multimedia system, the multimedia system comprising: a microphone, a processor, and a power amplifier, wherein,

[0093] The pickup device is used to acquire the user's audio signal;

[0094] The processor is configured to: detect the user's audio signal when playing music audio, and output the accompaniment audio in the music audio and the audio signal;

[0095] The power amplifier is used to play the accompaniment audio and the audio signal.

[0096] According to a sixth aspect of this application, embodiments of this application also provide a vehicle including the aforementioned electronic device or the aforementioned multimedia system.

[0097] Some embodiments of this specification include at least the following beneficial effects: by detecting the user's audio signal and combining the accompaniment audio in the music audio with the user's audio signal to output a mixed sound, the reliance on third-party karaoke applications is broken, so as to realize the karaoke experience. Users can sing anytime and anywhere through various devices, thereby providing a more convenient entertainment experience and improving the user's experience.

[0098] Other features and advantages of this application will be described in detail in the following detailed description section. Attached Figure Description

[0099] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0100] To gain a more complete understanding of this application and its beneficial effects, the following description will be provided in conjunction with the accompanying drawings, wherein the same reference numerals in the following description denote the same parts.

[0101] Figure 1 is a schematic diagram of the structure of a multimedia system according to some embodiments of this specification;

[0102] Figure 2 is an exemplary flowchart of an audio processing method according to some embodiments of this specification;

[0103] Figure 3 is an exemplary schematic diagram of a separation model according to some embodiments of this specification;

[0104] Figure 4 is an exemplary flowchart illustrating the determination of an audio mode according to some embodiments of this specification;

[0105] Figure 5 is an exemplary schematic diagram of preprocessing according to some embodiments of this specification;

[0106] Figure 6 is an exemplary schematic diagram of a cache according to some embodiments of this specification;

[0107] Figure 7 is an exemplary schematic diagram of an interactive interface according to some embodiments of this specification;

[0108] Figure 8 is an exemplary schematic diagram of human voice detection according to some embodiments of this specification;

[0109] Figure 9 is an exemplary schematic diagram illustrating the detection of whether a user is singing, according to some embodiments of this specification;

[0110] Figure 10 is an exemplary schematic diagram of exit time control according to some embodiments of this specification;

[0111] Figure 11 is an exemplary schematic diagram of a signal link according to some embodiments of this specification;

[0112] Figure 12 is a schematic diagram of the structure of an electronic device according to some embodiments of this specification.

[0113] Implementation methods of this application

[0114] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the protection scope of this application.

[0115] To facilitate understanding of the implementation schemes provided in this application, the relevant application background of the audio processing method provided in this application will be explained first.

[0116] In modern entertainment systems, karaoke offers an interactive form of entertainment, utilizing a music system to play background music while users sing using a microphone or a microphone built into the space. However, in certain scenarios (such as while driving), holding a microphone is neither safe nor convenient. Another related technology uses microphone-free karaoke, employing a multi-microphone array within the space to capture the user's voice more naturally and mix it into the audio output of the sound system in real time. However, this method cannot overcome the dependence of traditional karaoke applications on music libraries. It still relies on players or applications fixed to specific locations to play songs for entertainment purposes. These players or applications are difficult to update in terms of song content, and users need to go through multiple steps to find and play their favorite songs. This complex process not only affects the user experience but also fails to meet the user's need for timeliness in song selection, personal preference in music selection, and real-time singing.

[0117] In view of this, some embodiments of this specification provide an audio processing method that separates any music into accompaniment audio and source vocal audio. The accompaniment audio can be used for karaoke by the user, and the source vocal audio can be output with the volume freely adjusted by the user, breaking the constraints of traditional reliance on third-party karaoke applications. This allows users to use any sound source to complete the karaoke experience at any time, enhancing the user's driving pleasure.

[0118] The subject implementing the technical solution of the embodiments of this application can be an electronic device.

[0119] In some embodiments, the electronic device may be deployed on or connected to a mobile device via a wired or wireless means (e.g., a control device), and of course, the electronic device may also be the mobile device itself. The mobile device can have any appearance, such as a vehicle or other means of transportation. When the electronic device is connected to a mobile device, it can be a terminal device, such as a smartphone, tablet, laptop, desktop computer, etc., but is not limited thereto.

[0120] In some embodiments, the electronic device may also be a server, which may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0121] The following explanation uses a mobile device, specifically a smart vehicle (or simply vehicle), as an example.

[0122] The audio processing methods provided in some embodiments of this specification can be applied to various scenarios, such as in-vehicle karaoke systems, home entertainment, mobile devices, and other fields, to provide users with a richer and more personalized audio experience.

[0123] To at least partially address the aforementioned problems, this application provides an audio processing method, a storage medium, a program product, an electronic device, and a vehicle. Exemplary embodiments according to this application will now be described in more detail with reference to the accompanying drawings.

[0124] Figure 1 is a schematic diagram of the structure of a multimedia system according to some embodiments of this specification.

[0125] As shown in Figure 1, the multimedia system 100 may include a sound pickup device 110 for acquiring sound signals within a space. The space can refer to any physical environment or area. For example, it can be an enclosed or semi-enclosed room, conference room, classroom, recording studio, or other indoor space; or an outdoor environment such as a square, park, or street; or certain specific small areas, such as a performance area on a stage or the interior of a vehicle.

[0126] The microphone 110 is used to acquire the user's audio signal to pick up the sound emitted by the user. In some embodiments, the microphone 110 can be fixedly arranged in one or more locations in a space. The microphone 110 differs from existing wired and wireless handheld microphones in that it can achieve the sound pickup function without requiring the user to hold the microphone, freeing the user's hands and improving the singing experience. For example, the microphone 110 can be placed on the interior of a vehicle, which can improve driver safety and make the vehicle interior more aesthetically pleasing and tidy.

[0127] The processor 120 executes program instructions based on data, information, and / or processing results to perform one or more functions described in this application. For example, if the multimedia system is in karaoke mode, the processor (such as an AUDIO DSP) synthesizes and optimizes the user's audio signal obtained from the pickup device and the accompanying audio in the music audio according to the audio optimization strategy of karaoke mode to obtain the target audio; if the multimedia system is not in karaoke mode, i.e., in normal mode, the processor (such as an AUDIO DSP) optimizes the original music audio according to the audio optimization strategy of normal mode.

[0128] Amplifier device 130 is used to amplify and play audio. In some embodiments, amplifier device 130 may include a speaker.

[0129] In some embodiments, the multimedia system has multiple audio modes, including a karaoke mode and a normal mode. The karaoke mode is designed for singing, allowing users to sing along with the music. The normal mode is suitable for regular music playback and other non-singing uses, such as listening to music, watching movies, or playing games.

[0130] The karaoke mode includes an intelligent karaoke mode (i.e., the first audio mode) and an instant karaoke mode (i.e., the second audio mode). The sound pickup device is used to acquire the user's audio signal in the space. If the multimedia system is in karaoke mode, the processor 120 can optimize the accompanying audio and the user's audio signal according to the audio optimization strategy of karaoke mode to obtain the target audio. The power amplifier device amplifies the target audio and plays it. By setting up a built-in microphone in the space, the multimedia system can switch between karaoke mode and other working modes, which enables the multimedia system to have multiple working modes. It realizes the coupling of karaoke function and existing audio processing function, reduces system complexity, increases the functional diversity of the multimedia system, and allows users to sing in the car without holding a microphone. This solves the problem of system complexity and low convenience caused by the independent settings of the karaoke system and the existing multimedia system.

[0131] In some embodiments, the sound pickup device 110 may be multiple microphones arranged in space. After the multiple microphones pick up the sound, they send the ambient audio to the processor for fusion processing, making the sound signal clearer and more complete, improving the sound pickup device's ability to acquire sound signals, and thus improving the effect of subsequent sound playback.

[0132] For example, the sound pickup device 110 can be installed on the vehicle interior such as the center console, A / B pillars, front seats, rear seats, and roof to improve the sound pickup device 110's ability to acquire sound signals, thereby improving the effect of subsequent sound playback.

[0133] In some embodiments, the microphone 110 is provided on the vehicle's center console, A / B pillars, front seats, rear seats, and roof lining only as an example. In other embodiments, the microphone 110 may also be provided in other places inside the vehicle or in other locations, which will not be described in detail here.

[0134] It is important to note that the multimedia system 100 is provided for illustrative purposes only and is not intended to limit the scope of this specification. Various changes and modifications can be made by those skilled in the art based on the description herein. For example, the multimedia system 100 may also include storage devices, information sources, etc. Furthermore, the multimedia system 100 may be implemented on other devices to achieve similar or different functions. However, these changes and modifications will not depart from the scope of this specification.

[0135] For ease of explanation, the following description uses the scenario of multimedia system 100 being used in a vehicle as an example.

[0136] Figure 2 is an exemplary flowchart of an audio processing method according to some embodiments of this specification. In some embodiments, process 200 may be executed on a processor basis. As shown in Figure 2, process 200 includes the following steps.

[0137] Step 210: When playing music audio, the user's audio signal is detected, and the accompaniment audio and audio signal in the music audio are output.

[0138] In some embodiments, when playing music audio, the user's audio signal is detected, and the accompaniment audio and the audio signal from the music audio are output. Alternatively, when the multimedia system is turned on, it may detect the user's audio signal and output the accompaniment audio and the audio signal from the music audio. The processor can capture and process the user's audio signal (e.g., the user's singing signal), adjust parameters such as the volume, rhythm, or pitch of the accompaniment audio to ensure harmony between the two, and mix the processed accompaniment audio with the user's audio signal before outputting them to the power amplifier device.

[0139] In some embodiments, the input music audio can be processed to obtain accompaniment audio and source vocal audio, wherein the accompaniment audio and source vocal audio can be used in karaoke mode.

[0140] In some embodiments, the user's audio signal and the accompaniment audio can be synthesized to obtain the target audio, such as by performing echo cancellation, noise reduction, reverberation processing, etc.

[0141] In some embodiments, the karaoke mode can be controlled and related settings can be configured through the interactive interface. For example, options such as setting an on / off switch, selecting different karaoke modes, and adjusting the volume of the musician's voice can be implemented.

[0142] In some embodiments, the processed target audio is output via a power amplifier device.

[0143] In some embodiments, there can be multiple methods for acquiring music audio. For example, multimedia applications such as music or video can be installed in the in-vehicle multimedia system, and the system can acquire music audio from these applications. Users can also acquire music audio from storage devices (such as USB flash drives, hard drives, internal storage, etc.) and then read it through the in-vehicle multimedia system. Another example is that users can connect to the in-vehicle multimedia system via Bluetooth, wireless, or other means using their own terminals to transmit music audio played on their terminals to the system, enabling the system to acquire the music audio.

[0144] In-vehicle multimedia systems refer to multimedia systems installed inside vehicles.

[0145] In some embodiments of this specification, existing in-vehicle multimedia systems rely on third-party karaoke applications, preventing users from utilizing a wider range of music sources for karaoke and from entering karaoke mode at any time. By outputting the accompaniment audio and the audio signal from the music when the user's audio signal is detected, the system breaks free from the constraints of traditional reliance on third-party karaoke software apps. This allows users to control and complete the karaoke experience using any audio source within the in-vehicle multimedia system, providing a smart karaoke experience where users can enter karaoke mode at any time while playing any music.

[0146] In some embodiments, when a user's audio signal is detected, the accompaniment audio and the audio signal in the music audio are output, including:

[0147] Recognize the user's audio signal;

[0148] If the audio signal is a singing signal, then the accompaniment audio and the audio signal in the music audio will be output.

[0149] In some embodiments, whether a user is singing can be detected in a variety of ways. One possible way is to directly detect whether the user's audio signal exists in the ambient audio when the audio mode is smart karaoke mode. If it exists, it can be assumed that the user is singing, and then audio output is performed based on karaoke mode.

[0150] Figure 8 is an exemplary schematic diagram of human voice detection according to some embodiments of this specification.

[0151] For example, as shown in Figure 8, the ambient audio collected by the microphone can be processed for echo cancellation and noise reduction, the music audio played by the speaker can be removed and non-human voice noise can be suppressed, and the processed ambient audio can be passed through Voice Activity Detection (VAD) to determine whether the user is singing at the current moment.

[0152] In some embodiments of this specification, when a user's singing signal is detected, the system can dynamically obtain the accompaniment audio, making the accompaniment audio more compatible with the user's singing signal, thereby providing a smoother, more harmonious, and personalized karaoke experience.

[0153] In some embodiments, the method further includes:

[0154] Determine whether the audio signal is a singing signal based on the source human voice audio corresponding to the audio signal and the music audio.

[0155] In some embodiments, when the audio mode is karaoke mode, the processor can identify audio segments containing human voices from ambient audio using voice detection (VAD) technology and extract key features of the audio segments, such as spectral characteristics, fundamental frequency (F0), zero-crossing rate, and rhythmic pattern. Simultaneously, it analyzes the source human voice audio in the music audio and extracts key features of the source human voice audio. It then uses a machine learning model or rule-based algorithm to compare the differences between two key features, such as calculating similarity or classification judgment, to determine whether the audio signal is a singing signal.

[0156] In some embodiments of this specification, by comparing the audio signal with the source human voice audio, it is possible to effectively determine whether the audio signal is a singing signal and optimize the karaoke experience.

[0157] In some embodiments, determining whether an audio signal is a singing signal based on the source human voice audio corresponding to the audio signal and the music audio includes:

[0158] Based on the similarity between the audio features of the audio signal and the source human voice audio corresponding to the music audio, it is determined whether the audio signal is a singing signal.

[0159] The source vocal audio can be the singer's audio in a music audio file, or it can be the singer's audio obtained through other means (such as from music applications, video applications, user terminals, etc.).

[0160] Figure 9 is an exemplary schematic diagram illustrating the detection of whether a user is singing, according to some embodiments of this specification.

[0161] In some embodiments, as shown in Figure 9, another possible implementation involves performing similarity detection between the user's audio signal and the source human voice audio. After human voice detection, the user's audio signal is extracted from the ambient audio. The user's audio signal is then subjected to a short-time Fourier transform (time-frequency analysis) to obtain the first audio feature of the user's audio signal (e.g., time spectrum, Mel spectrum, or MFCC spectrum). Similarly, the second audio feature of the separated source human voice audio (features after time-frequency analysis) is obtained. The first and second audio features are then subjected to similarity detection. If the similarity between the two is greater than a set similarity threshold, it indicates that the user's audio signal is similar to the lyrics, melody, etc. of the source human voice audio, suggesting that the user is singing rather than speaking randomly, thus completing the similarity detection based on time-frequency changes.

[0162] In some embodiments of this specification, the similarity detection method can significantly reduce latency and improve user experience compared to methods such as speech recognition.

[0163] Figure 3 is an exemplary schematic diagram of a separation model according to some embodiments of this specification.

[0164] In some embodiments, as shown in FIG3, before outputting the accompaniment audio and audio signal in the music audio, the method includes:

[0165] The music audio is input into the separation model for processing to obtain the accompaniment audio and / or the original vocal audio.

[0166] A separation model is a mathematical or computational model used to separate different audio components from a mixed audio signal.

[0167] In some embodiments, the separation model is a machine learning model. For example, the separation model may include any one or a combination of Convolutional Neural Networks (CNN) models, Neural Networks (NN) models, or other custom model structures.

[0168] In some embodiments, the input to the separation model may include music audio, and the output may include accompanying audio and / or source vocal audio.

[0169] Accompaniment audio refers to all parts of a musical audio file other than the vocals, including but not limited to instrumental performances, background music, and sound effects. Accompaniment audio forms the basic melody and rhythm of a song or audio file.

[0170] Source vocal audio refers to the vocal portion of an audio file, i.e., the singer's voice. This can be a solo, a chorus, or any other form of vocal performance.

[0171] In some embodiments, music audio in a preset format can be component-processed using a separation model. Further details of this embodiment can be found in the following description.

[0172] In some embodiments, the music audio is processed by the input separation model to obtain the accompaniment audio, including:

[0173] The music audio is input into the separation model for processing, resulting in the source human voice audio:

[0174] The accompaniment audio is obtained based on the difference between the music audio and the source vocal audio.

[0175] In some embodiments, for the input and output of the separation model, one possible approach is that the input is music audio, and the output consists of two audio signals: the source vocal audio and the accompaniment audio. Another possible approach is that the input is music audio, and the output is either the source vocal audio or the accompaniment audio. Subtracting the source vocal audio or the accompaniment audio from the music audio yields either the accompaniment audio or the source vocal audio. The trained separation model can satisfy the performance requirement of separating the source vocal audio from the accompaniment audio in the audio.

[0176] In some embodiments of this specification, the trained separation model intelligently separates the music audio from any music source to obtain the accompaniment audio, breaking the reliance on third-party karaoke applications. This allows users to enjoy a karaoke experience at any time based on any music source, improving the user's driving experience.

[0177] In some embodiments of this specification, the accuracy and efficiency of audio separation are improved by using a trained separation model, making it simpler and more reliable to extract pure accompaniment audio from complex music audio.

[0178] In some embodiments, the music audio is processed by an input separation model to obtain accompaniment audio and / or source vocal audio, including:

[0179] Extract the music audio of the first preset frame length, input it into the separation model for processing, and obtain the accompaniment audio and / or the source vocal audio of the first preset frame length.

[0180] The first preset frame length refers to the duration of one frame of music audio. The first preset frame length can be expressed in units such as the number of samples or milliseconds.

[0181] In some embodiments, the music audio can be divided into multiple frames, each frame having a length of a first preset frame length.

[0182] In some embodiments, the method includes:

[0183] When outputting the accompaniment audio of the first preset frame length, the music audio of the second preset frame length is extracted, input into the separation model for processing, and the accompaniment audio and / or source vocal audio of the second preset frame length are obtained.

[0184] The second preset frame length refers to the duration of one frame of music audio. The second preset frame length can be expressed in units such as sample count or milliseconds. The second preset frame length can be the same as or different from the first preset frame length.

[0185] In some embodiments, the frame length used to extract the music audio is redefined after each frame of music audio is acquired. For example, one or more thresholds can be set, and when a feature of the music audio (such as transient intensity) exceeds the corresponding threshold, a shorter frame length is used; otherwise, a longer frame length is used.

[0186] In some embodiments, extracting music audio of a second preset frame length includes:

[0187] After extracting the music audio of the first preset frame length, a second preset frame length is determined in order to extract the music audio of the second preset frame length.

[0188] In some embodiments, the time taken for the separation model to process the music audio of the second preset frame length is less than the playback time of the accompaniment audio of the first preset frame length.

[0189] In some embodiments, a multi-threaded architecture or a multi-process architecture can be used to assign different tasks to different threads or processes. For example, one thread is responsible for obtaining a frame of music audio, another thread processes the frame of music audio into a separation model to obtain the corresponding accompaniment audio and / or the source vocal audio, and a third thread is responsible for outputting the processed accompaniment audio and the user's audio signal to the playback device.

[0190] In some embodiments, when the accompaniment audio of the current frame begins to play, the processor can asynchronously process the music audio of the next frame in the background. If the processing time of a music audio frame is short, the processor continues to process the music audio of the frame after that.

[0191] By ensuring that the time taken for the separation model to process the music audio of the second preset frame length is less than the playback time of the accompaniment audio of the first preset frame length, it helps to ensure smooth audio output and avoid problems such as insufficient buffering or playback interruption.

[0192] In some embodiments, the separation model is trained in the following manner:

[0193] Obtain sample audio data, which includes the corresponding sample accompaniment audio and / or sample source human voice audio;

[0194] Based on the sample audio data, the trained separation model is obtained through the separation model to be trained.

[0195] In some embodiments, the separation model can be trained using a large number of labeled separation model training samples through various feasible methods. For example, parameters can be updated using gradient descent. An exemplary training process includes: inputting multiple labeled separation model training samples into an initial separation model; constructing a loss function using the labels and the results of the initial separation model; and iteratively updating the parameters of the initial separation model based on the loss function using gradient descent or other methods. The model training is complete when preset conditions are met, resulting in a trained separation model. These preset conditions may include loss function convergence, the number of iterations reaching a threshold, etc.

[0196] In some embodiments, the training samples for the separation model include at least sample audio data.

[0197] In some embodiments, the tags may include sample background audio and sample source human voice audio corresponding to the sample audio data. Tags can be obtained through a processor or manual annotation.

[0198] In some embodiments, a music source database can be collected and created. The music source database can consist of files such as original music, vocals, and accompaniment. The music in the music source database covers most commonly used instruments and genres.

[0199] In some embodiments, the separation model can be trained based on a music source database.

[0200] In some embodiments, source vocal audio and accompaniment audio are collected to establish a music source library containing known accompaniment and vocal data. The separation model is trained based on this music source library. The separation model can include open-source music source separation algorithms such as MdxNet, Demucs, and BandSplitRNN.

[0201] In some embodiments, the separation model can be trained frame by frame. For example, if the frame length is 2.72s and the sampling rate is 48kHz, then the sample audio of each frame of dual channels has 2×130560 sampling points.

[0202] It should be noted that the input and output of the separation model during training are audio frames with a length of 2.72 seconds. In practical applications, if the input audio is shorter than 2.72 seconds, for example, only 0.17 seconds, the separation model can output a separation result of 0.17 seconds at the expense of some performance. Based on the above separation model, any audio can be separated to obtain its source vocal audio and accompaniment audio.

[0203] First, a music source database for separating vocals and accompaniment was collected and constructed. This database can be used to train a separation model capable of separating vocals and accompaniment for any music. Second, based on this separation model, any music audio is separated into accompaniment audio and source vocal audio. The accompaniment audio can be used for karaoke, and the source vocal audio can be output with the user's volume adjusted according to the accompaniment.

[0204] In some embodiments of this specification, the trained separation model can infer and extract the source vocal audio and accompaniment audio from any music source. The separation model has good separation performance and variable frame length processing capability, enabling vocal and accompaniment separation under different input frame lengths.

[0205] Figure 4 is an exemplary flowchart illustrating the determination of an audio mode according to some embodiments of this specification. In some embodiments, process 400 may be processor-based. As shown in Figure 4, process 400 includes the following steps.

[0206] Step 410: In response to the user's control command, determine the audio mode.

[0207] Control commands are operational commands used to control audio modes.

[0208] Step 420: Output music audio according to the audio mode.

[0209] Figure 7 is an exemplary schematic diagram of an interactive interface shown according to some embodiments of this specification.

[0210] In some embodiments, as shown in Figure 7, users in the vehicle can control the karaoke mode via a vehicle-mounted tablet or voice commands. For example, taking vehicle-mounted tablet control as an example, the interactive interface of the vehicle-mounted tablet has a karaoke mode switch button (as shown in Figure 7, virtual button 701, etc.). In the karaoke mode of the interactive interface, there are corresponding switch buttons for a first audio mode (as shown in Figure 7, smart karaoke switch button 703) and a second audio mode (as shown in Figure 7, i.e. karaoke switch button 702).

[0211] In some embodiments, the method further includes:

[0212] Outputs the accompaniment audio, source vocal audio, and audio signal from the music audio.

[0213] In some embodiments, the processor can synthesize the user's audio signal, accompaniment audio, and source vocal audio to obtain the target audio, such as by performing echo cancellation, noise reduction, reverberation processing, etc.

[0214] In some embodiments, the method further includes:

[0215] An operation to adjust the volume of the source human voice was detected, and the volume of the output source human voice audio was adjusted accordingly.

[0216] The original vocals refer to the singer's voice separated from the original audio file, that is, the original vocal part in the song or audio data.

[0217] In some embodiments, to enable users to adjust the output audio, the interactive interface typically uses a series of user-friendly controls (such as buttons, sliders, input boxes, etc.). For example, the first object for adjusting the source voice volume (as shown in Figure 7, slider 705), which is the source voice separated from the music audio by the separation model, can be manipulated to adjust the volume of the output source voice, allowing a portion of the source voice to be heard during karaoke mode. Specifically, this is achieved by adding source voice audio with a certain volume to the target audio. The maximum adjustable volume of this source voice is the source voice volume, and the minimum is 0. This source voice volume adjustment function can be used in both instant karaoke mode and intelligent karaoke mode. Another example is a second object (as shown in Figure 7, numeric input box 704, etc.), which can be set to exit time, automatically exiting karaoke mode after no user singing is detected within the exit time.

[0218] In some embodiments, the processor can obtain user control commands through an interactive interface to determine the audio mode.

[0219] In some embodiments, when not in karaoke mode, the processor can process the music audio using the audio optimization strategy corresponding to the normal mode to obtain and play the processed music audio.

[0220] In some embodiments, when in karaoke mode, the processor can optimize the accompanying audio and the user's audio signal by processing them through the audio optimization strategy corresponding to karaoke mode to obtain the target audio and play it.

[0221] Audio optimization strategies are a set of pre-defined rules or methods for processing audio.

[0222] When not in karaoke mode (i.e., normal mode), the processor processes audio based on music (such as regular music, radio, etc.) using the audio optimization strategies corresponding to normal mode to ensure the best listening experience. Normal mode primarily focuses on improving the quality of regular audio playback.

[0223] Audio optimization strategies in normal mode include, but are not limited to: equalizer adjustment, dynamic range compression, noise reduction, volume normalization, and audio format conversion.

[0224] The audio optimization strategies for karaoke mode include, but are not limited to: vocal enhancement, reverb and echo control, pitch correction, background music optimization, and multi-channel processing.

[0225] In some embodiments, the audio mode includes a first audio mode, and the method includes:

[0226] In the first audio mode, when playing music audio, the steps include detecting the user's audio signal and outputting the accompaniment audio and audio signal from the music audio.

[0227] The first audio mode is the smart karaoke mode.

[0228] Figure 10 is an exemplary schematic diagram of exit time control according to some embodiments of this specification.

[0229] In some embodiments, as shown in Figure 10, when a user switches to the smart karaoke mode, if no singing is detected, the audio being played or about to be played is processed based on the normal mode, and the amplifier plays the audio processed based on the normal mode. When a user starts singing, the audio being played or about to be played is processed based on the karaoke mode, and the amplifier plays the audio processed based on the karaoke mode, that is, the target audio is played based on the mixture of the accompaniment audio and the user's audio. When the user stops singing, there is accompaniment audio output. If there is source human voice audio but the user has stopped singing for more than a set exit time, the audio processed in the normal mode will be output.

[0230] In some embodiments of this specification, the intelligent karaoke mode helps to intelligently identify the user's singing activities and automatically switch the corresponding audio without manual intervention, thus simplifying the operation process.

[0231] In some embodiments, the method further includes:

[0232] If the user's audio signal is not detected during the output of the accompaniment audio and the audio signal in the music audio, then the accompaniment audio or the music audio will be output.

[0233] If no user's audio signal is detected during the process of controlling the output music audio and the accompaniment audio, the music audio will be output.

[0234] In some embodiments, when the audio mode is Smart Karaoke mode, the processor monitors the output of the pickup device in real time through voice detection (VAD) technology to determine whether the user's singing voice is present; when a valid voice signal is not identified, it indicates that the user may not have started singing, has paused temporarily, etc., and outputs the accompaniment audio or the original music audio to the power amplifier device.

[0235] In some embodiments of this specification, audio music is output by determining that there is currently no valid user audio signal; ensuring a smooth and natural transition from an audio signal with a user to an audio signal without a user, and avoiding sudden volume changes or audio interruptions.

[0236] In some embodiments, if no user audio signal is detected, the accompaniment audio or music audio in the music audio is output, including:

[0237] Based on the duration for which no user's audio signal is detected, output the accompaniment audio or music audio from the music audio.

[0238] In some embodiments, based on the duration for which no user's audio signal is detected, the accompaniment audio or music audio in the music audio is output, including:

[0239] If the duration is less than the first preset duration, output the accompaniment audio in the music audio.

[0240] In some embodiments, based on the duration for which no user's audio signal is detected, the accompaniment audio or music audio in the music audio is output, including:

[0241] If the duration is greater than or equal to the first preset duration, output music audio.

[0242] In some embodiments, when the audio mode is Smart Karaoke mode, the processor can monitor the user's audio signal in real time through human voice detection (VAD) technology: if the user's audio signal is not detected, the processor starts timing and records the duration of the state. When the duration reaches a first preset duration, the processor automatically switches to outputting only music audio to the amplifier device.

[0243] In some embodiments of this specification, a more flexible and natural singing environment is provided by setting a first preset duration, allowing for uninterrupted musical accompaniment whether during short breaks or preparation phases.

[0244] In some embodiments, when the duration is greater than or equal to a first preset duration, music audio is output, including:

[0245] If no user's audio signal is detected within the first preset duration, and the source human voice audio corresponds to the accompaniment audio in the music audio, the music audio is output.

[0246] In some embodiments, when the smart karaoke mode is enabled, the user-inputted exit time is obtained through the interactive interface as the first preset duration. The system also detects in real time whether the user is singing in the vehicle environment. When the user stops singing, the system also detects whether the original human voice audio exists. If it exists, a countdown starts; if the original human voice audio does not exist, the countdown stops. If the countdown does not reach the first preset duration but the system detects that the user is singing in the vehicle environment, the countdown is reset to zero. If the countdown reaches the first preset duration, the system exits the karaoke mode and switches to normal mode, at which point the original music audio is played.

[0247] In some embodiments of this specification, by detecting the user's audio signal and the source human voice audio, it is ensured that the background music can still play smoothly even when the user is not singing temporarily, avoiding sudden silence or interruption and providing a more consistent listening experience.

[0248] In some embodiments, the music audio includes source vocal audio corresponding to the accompaniment audio, including:

[0249] Calculate the energy of the music audio. If the energy of the music audio is less than a preset energy threshold, then it is determined that there is a source human voice audio corresponding to the accompaniment audio in the music audio.

[0250] In some embodiments, the presence of musical vocals can be determined based on the energy of the source vocal audio. One possible approach is to set a preset energy threshold. If the energy of the current frame of the source vocal audio is less than the preset energy threshold, it indicates that there are no musical vocals in the source vocal audio of that frame; otherwise, musical vocals are present.

[0251] In some embodiments of this specification, by setting a reasonable energy threshold, it is possible to more accurately distinguish between background music (such as instrumental music) and audio segments containing source human voices, which helps to avoid misjudgment.

[0252] In some embodiments, the audio mode includes a second audio mode, and the method further includes:

[0253] In the second audio mode, the accompaniment audio from the music audio is output.

[0254] The second audio mode is the instant karaoke mode.

[0255] In some embodiments, the method further includes:

[0256] If the user's audio signal is detected during the output of the accompaniment audio in the music audio, the accompaniment audio and the audio signal in the music audio will be output.

[0257] In some embodiments, after detecting the user's operation of turning on the karaoke mode switch, the system defaults to entering the instant karaoke mode. At this time, the audio being played or about to be played is separated and processed, and the audio source of the power amplifier device is switched from the audio processed based on the normal mode to the audio processed in the karaoke mode, that is, the accompaniment audio or the target audio based on the accompaniment audio and the user's audio, until the user turns off the karaoke mode and plays the audio processed in the normal mode.

[0258] In some embodiments of this manual, users only need to select the instant karaoke mode, and the system will automatically adjust to output accompaniment audio without any additional settings or manual parameter adjustments, which greatly simplifies the usage steps; when the user is ready to start singing, the system can quickly provide the most suitable accompaniment environment and reduce preparation time.

[0259] In some embodiments, when in the second audio mode, if the user's audio signal is detected, the system controls the output of the accompaniment audio in the music audio and the audio signal synthesized together.

[0260] In some embodiments of this specification, by detecting the user's audio signal, it is ensured that the user can immediately receive optimized accompaniment support and corresponding auditory feedback when it is detected that they have started singing, thereby enhancing the user's sense of participation and singing enjoyment.

[0261] The two karaoke modes mentioned above can provide users with richer karaoke effects. Users can only use one mode at a time. The instant karaoke mode provides users with the experience of singing as soon as they turn on the switch, while the smart karaoke mode provides users with the experience of singing anytime and automatically turning off when not singing.

[0262] Figure 5 is an exemplary schematic diagram of preprocessing according to some embodiments of this specification.

[0263] In some embodiments, as shown in FIG5, the method further includes:

[0264] The music audio is processed to obtain music audio in a preset format.

[0265] A preset format refers to a predefined format for music audio. Preset formats may include a preset number of channels and a preset sampling rate.

[0266] In some embodiments, the music audio is processed to obtain music audio in a preset format, including:

[0267] Convert music audio to music audio with a preset number of channels.

[0268] In some embodiments, music audio can be converted to a preset number of channels using channel number conversion.

[0269] Channel conversion refers to changing the number of channels in an audio file or audio stream. For example, converting mono audio to stereo, or 5.1 surround sound to stereo.

[0270] In some embodiments, channel number conversion includes:

[0271] If the music audio has a single channel, then the music audio will be used as the left and right channels of the music audio with the preset number of channels.

[0272] If the number of channels in the music audio is greater than two channels, then the music audio is downmixed to obtain the music audio with the preset number of channels.

[0273] In some embodiments, the in-vehicle multimedia system can acquire music in various ways, and the types of music audio can also be varied (e.g., music applications, video applications, storage devices, user terminals, etc.). To ensure that all audio is processed under the same algorithm logic, preprocessing of the music audio is required. Preprocessing mainly includes channel number conversion and sampling rate conversion. In practical applications, if the music audio acquired by the in-vehicle multimedia system is dual-channel, no preprocessing is performed; if the music audio is single-channel, the single-channel music audio is copied to obtain dual-channel music audio; if the number of channels of the music audio is greater than 2, such as 5.1 or 7.1 surround sound audio, or 7.1.4 Dolby Atmos audio, the music audio is downmixed to obtain dual-channel music audio.

[0274] In some embodiments of this specification, through reasonable mono-to-stereo and multi-channel downmixing processing, not only can the spatial sense and auditory experience of the audio be significantly improved, but the compatibility and consistency of the audio content on various playback devices can also be ensured.

[0275] In some embodiments, the downmixing process includes:

[0276] The preprocessed music audio is obtained by weighted summation of multiple audio channels.

[0277] In some embodiments, downmixing can be implemented in several ways. For example, one possible approach is to extract the left and right channels from the music audio as the preprocessed music audio, which will result in the loss of some channel audio information. Another approach is to weighted summation of the audio from multiple channels in the music audio to obtain the preprocessed audio for one channel. For example, in 5.1 surround sound, the distribution is L, C, R, LR, RR, and Sub, so the preprocessed left channel audio would be...

[0278] 0.707*L+0.5*C+0.25*LR+Sub can be used to obtain the audio from the right channel, similar to the above, thus converting multi-channel audio into dual-channel audio for easier subsequent audio processing.

[0279] In some embodiments of this specification, the spatial sense of the original audio is preserved as much as possible by merging information from multiple channels. For example, in a 5.1 channel system, the left and right channels can be directly mapped to the left and right channels of stereo, and the information from the center channel and surround channels is merged in to maintain a certain sense of direction and depth.

[0280] In some embodiments, the music audio is processed to obtain music audio in a preset format, including:

[0281] Convert music audio to music audio with a preset sampling rate.

[0282] In some embodiments, music audio can be converted to music audio with a preset sampling rate through sample rate conversion.

[0283] Sampling rate conversion refers to changing the number of times audio data is sampled per second, that is, changing the temporal resolution of the audio signal. Common sampling rates include 44.1kHz, 48kHz (digital video standard), and 96kHz.

[0284] In some embodiments, the sampling rate conversion includes upsampling and downsampling.

[0285] Upsampling increases the sampling rate. Downsampling decreases the sampling rate.

[0286] In some embodiments, if the sampling rate of the acquired music audio is not the set sampling rate, the sampling rate of the music audio is converted by upsampling or downsampling to make the sampling rate of the music audio the set sampling rate in karaoke mode. For example, if the set sampling rate is 48kHz, and the acquired music audio has a sampling rate of 96kHz, the music audio is passed through an anti-aliasing filter and then the sampling rate is converted to downsample the 96kHz audio to 48kHz. If the acquired music signal is 44.1kHz, the music audio is first upsampled and then passed through an anti-mirror filter to upsample the 44.1kHz audio to a 48kHz sampling rate signal. After channel number conversion and sampling rate conversion, the music audio output is a dual-channel audio at the set sampling rate. With preprocessing, any audio format can be used for karaoke.

[0287] In some embodiments of this specification, channel number conversion and sampling rate conversion are used to better adapt the audio to the requirements of separation processing and improve the efficiency of separation processing.

[0288] In some embodiments, the method further includes:

[0289] Store the music audio in the preset format into the second buffer.

[0290] The second buffer is the audio buffer area.

[0291] In some embodiments, the processor may acquire complete music audio data from a network stream, local file, or other source as the music audio to be played, which includes vocal audio and accompaniment audio. The processor may convert the music audio to be played into a preset format and store it in a second buffer to ensure smooth audio playback even under network fluctuations or read latency.

[0292] Figure 6 is an exemplary schematic diagram of a cache according to some embodiments of this specification.

[0293] In some embodiments, as shown in FIG6, the method further includes:

[0294] When the user's audio signal is detected, the accompaniment audio and the audio signal stored in the first buffer are output.

[0295] The first buffer is a buffer area for the separated results of music audio.

[0296] In some embodiments, before outputting the accompaniment audio and audio signal from the music audio stored in the first buffer, the method further includes:

[0297] Store the accompaniment audio from the music audio file into the first buffer.

[0298] In some embodiments, the processor may use audio processing techniques (such as sound source separation algorithms or pre-recorded accompaniment versions) to extract the accompaniment audio from the music audio in the second buffer and store the separated accompaniment audio in the first buffer, providing flexibility for subsequent operations.

[0299] In some embodiments, the processor may use a separation model to extract the accompaniment audio and / or the source vocal audio from the music audio in the second cache, and store the separated accompaniment audio and / or source vocal audio into the first cache.

[0300] In some embodiments of this specification, by storing the music audio and accompaniment audio in separate buffers, the system can more efficiently manage and adjust the audio output during playback. For example, in karaoke mode, when the user's audio signal is detected, the system can quickly read the accompaniment audio from the first buffer and mix it with the user's voice for output, ensuring real-time performance and interactivity.

[0301] In some embodiments, the processor can extract music audio segments from the second buffer according to a preset frame length (e.g., from a few milliseconds to several hundred milliseconds), i.e., frame-by-frame processing.

[0302] In some embodiments, the processor can extract each frame of music audio, use audio processing algorithms (such as sound source separation technology, spectrum analysis or pre-trained machine learning models) to identify and separate the accompaniment audio, and store the separated accompaniment audio in a first buffer.

[0303] In some embodiments of this specification, through frame-segmentation and caching strategies, the system can efficiently manage and process audio data, reducing memory usage and computational burden, while ensuring the continuity and stability of audio playback.

[0304] In some embodiments, storing the accompaniment audio from the music audio into a first buffer includes:

[0305] Store the accompaniment audio from the music audio with a preset frame length into the first buffer.

[0306] In some embodiments, the preset frame length is a fixed frame length, such as 23ms, or the preset frame length is a dynamically changing frame length.

[0307] In some embodiments, for each frame of music audio obtained, the relationship between a preset frame length and a frame length threshold is determined, and the preset frame length is increased or maintained, including:

[0308] If the preset frame length is less than the frame length threshold, then increase and update the preset frame length;

[0309] If the preset frame length is greater than or equal to the frame length threshold, then the preset frame length is maintained.

[0310] In some embodiments, the processor may set an initial preset frame length for extracting the first frame of music audio from the second buffer; whenever the system acquires a frame of music audio with a preset frame length, it evaluates the relationship between the current preset frame length and a predefined frame length threshold; if the current preset frame length is less than the frame length threshold, the preset frame length is appropriately increased based on the preset value; if the current preset frame length has reached or exceeded the frame length threshold, the preset frame length remains unchanged.

[0311] In some embodiments of this specification, by dynamically adjusting the frame length, the frame length can be flexibly adjusted to adapt to different audio characteristics and processing needs while ensuring the real-time performance of audio processing. For example, when the audio content is complex or changes rapidly, a shorter frame length can help capture more details; while in relatively stable parts, a longer frame length can reduce the number of processing steps and improve efficiency.

[0312] For example, the audio data being played or about to be played can be separated into source vocal audio and accompaniment audio in the in-vehicle multimedia system. Since the separation process requires a certain processing time, the audio stream is pre-buried when it arrives at the processor. One possible implementation is that the music stream enters an audio buffer area, which typically stores the audio currently being played and the audio that will be played in the future. When a user activates karaoke mode (instant karaoke mode, or smart karaoke mode where the user has started singing), to reduce algorithm processing latency and user waiting time and improve user experience, a variable frame length processing method can be adopted for vocal and accompaniment separation. For example, the separation model first processes the audio with an initial frame length of 0.17s after the current moment, outputs 0.17s of source vocal audio and accompaniment audio, and stores them in the separation result buffer area. When playing audio for this time period, the source vocal audio and / or accompaniment audio for this time period are extracted from the separation result buffer area and synthesized with the user's audio signal through karaoke mode to output the target audio for this time period. In the playback sequence, the audio processed in normal mode is output based on a fade-in / fade-out switching strategy. When the target audio of 0.17s is ready to be played and output, audio with a length of 4 times the initial frame length is taken from the audio buffer area, that is, a new audio input with a frame length of 0.68s is processed by the separation model. The separation result is placed in the separation result buffer area. The source human voice audio and / or accompaniment audio of this time period is extracted from the separation result buffer area and synthesized with the user's audio signal based on the karaoke mode to output the target audio, and so on. While ensuring real-time processing, the frame length of the signal processed by the separation model is gradually increased to 2.72s, and then the frame length of 2.72s is kept unchanged. After processing a certain music audio in the audio buffer area that has not been separated, the separated source human voice audio and accompaniment audio are stored in the separation result buffer area.

[0313] In some embodiments, the method further includes:

[0314] Upon detecting a progress bar adjustment, output the accompaniment audio from the music audio at the corresponding moment after the operation, indicating the progress bar's position.

[0315] In some embodiments, when a user drags the progress bar, the processor can obtain the music audio corresponding to the position of the progress bar after the operation and check whether the accompaniment audio corresponding to that moment exists in the first buffer. If it exists, it is used directly. If it does not exist, the processor can obtain the music audio of the corresponding frame length based on the preset frame length corresponding to the initial position or the preset frame length corresponding to the current moment. For the newly obtained music audio frame, the accompaniment audio is extracted through separation processing and stored in the first buffer.

[0316] The preset frame length corresponding to the initial position refers to the frame length extracted from the initial playback position of the music audio.

[0317] The preset frame length corresponding to the current moment can be any frame length, such as the preset frame length corresponding to the initial position, or the preset frame length after dynamic adjustment.

[0318] In some embodiments of this specification, in application scenarios that require frequent switching of playback positions (such as video viewing, karaoke singing, etc.), waiting time and stuttering can be effectively reduced, providing users with a smoother and more natural user experience and ensuring the continuity and real-time performance of subsequent playback.

[0319] In some embodiments, the method further includes outputting the progress bar position before the accompaniment audio in the music audio at the time corresponding to the output operation:

[0320] Upon detecting a progress bar adjustment operation, determine if the accompaniment audio from the music audio corresponding to the position of the progress bar after the operation exists in the first buffer:

[0321] If the first buffer contains the accompaniment audio in the music audio corresponding to the position of the progress bar after the operation, extract the music audio after the position based on the preset frame length corresponding to the initial position.

[0322] If the accompaniment audio in the music audio corresponding to the position of the progress bar after the operation is not found in the first buffer, the music audio will continue to be extracted according to the time sequence and based on the preset frame length corresponding to the current time.

[0323] In some embodiments, the processor can check the first buffer to confirm whether there is accompaniment audio at the time pointed to by the progress bar after dragging. If not, the processor can locate the audio data after the time corresponding to the position of the progress bar in the second buffer according to the position of the progress bar after the operation, and extract the corresponding frame of music audio according to the preset frame length corresponding to the initial position. For the newly extracted frame of music audio, the accompaniment audio part is separated by the separation model and stored in the first buffer.

[0324] In some embodiments of this specification, the system responds quickly when the user drags the progress bar, and can rapidly provide accurate accompaniment audio even when the corresponding data is missing in the first buffer; this not only improves flexibility and response speed, but also ensures continuity and stability during playback.

[0325] In some embodiments, the processor can check the first cache to confirm whether there is accompaniment audio at the time indicated by the progress bar after dragging. If there is, the processor can continue to extract music audio from the second cache in chronological order. For example, the processor can extract music audio based on the preset frame length corresponding to the current time, i.e., the frame length after dynamic adjustment. For the newly extracted music audio of the second frame length, the processor can use a separation model to separate the accompaniment audio and store it in the first cache.

[0326] In some embodiments of this specification, by continuing to extract and cache the accompaniment audio based on the dynamic frame length at the current moment, it is possible to provide a more personalized and stable audio playback service while maintaining efficient processing. This is suitable for application scenarios that require real-time interaction, such as karaoke singing.

[0327] In some embodiments, if a user drags the music playback progress bar during music playback, one possible approach is to continue processing the 2.72s frame length audio according to the normal timing if the audio at the dragged progress bar position has been separated; if the audio at the dragged progress bar position has not been separated, a variable frame length approach similar to the one described above is used to ensure the user's real-time experience.

[0328] In some embodiments, as shown in FIG6, the method further includes:

[0329] If the user is detected to have disabled or enabled the singing function, the synthesized audio is output based on a preset fade-in / fade-out strategy. The synthesized audio is a combination of at least two audio components: music audio, accompaniment audio from the music audio, and audio signals.

[0330] The fade-in / fade-out switching strategy is used to achieve a smooth transition between two audio signals, with the aim of allowing users to switch between karaoke audio output and normal sound effect processed audio output more smoothly in terms of listening experience.

[0331] Enabling the singing function activates the karaoke mode.

[0332] In some embodiments, when the user enables the singing function, the processor gradually reduces the volume of the music audio while gradually increasing the volume of the accompaniment audio and the user's audio signal that is about to be added; when the user disables the singing function, the processor gradually reduces the volume of the accompaniment audio and the user's audio signal, and restores the normal volume of the music audio, so that the music audio regains its dominant position.

[0333] In some embodiments of this specification, a fade-in / fade-out strategy is used to achieve smooth switching between different modes, reducing the discomfort caused by sudden volume changes or audio interruptions.

[0334] In some embodiments, synthesized audio is output based on a preset fade-in / fade-out strategy, including...

[0335] Based on the preset fade-in and fade-out strategy, determine the second preset duration;

[0336] Within the second preset duration, each frame of music audio with a preset frame length and its corresponding target audio with a preset frame length are obtained. The target audio is the accompaniment audio in the music audio or the audio synthesized based on the accompaniment audio in the music audio and the audio signal.

[0337] The music audio and the corresponding target audio are weighted and summed to output the synthesized audio.

[0338] The second preset duration is the time period used for a smooth transition.

[0339] In some embodiments, the processor sets an appropriate second preset duration (e.g., a few seconds) according to a preset fade-in / fade-out strategy; within the second preset duration, it acquires music audio of a preset frame length and its corresponding accompaniment audio frame by frame; if the singing function is enabled and the user's audio signal is detected, the accompaniment audio and the user's audio signal are combined to form the target audio; if the singing function is enabled but the user's audio signal is not detected, the first audio is only the accompaniment audio.

[0340] In some embodiments, the processor can perform a weighted summation of the original music audio and the corresponding first audio (accompaniment audio or target audio), with the weights depending on whether the current phase is fade-in (when the user enables the singing function) or fade-out (when the user disables the singing function), and gradually adjust the relative contributions between the music audio and the first audio to achieve a smooth transition.

[0341] In some embodiments of this specification, frame-by-frame processing and weighted summation are used to achieve a smooth transition between music audio, accompaniment audio, and user vocals, providing a more natural and fluid audio playback experience.

[0342] In some embodiments, a weighted sum is performed between the music audio and the corresponding target audio to output a synthesized audio, including:

[0343] When the user enables the singing function, a weighted sum is performed based on the first weight and the music audio, and the second weight and the corresponding target audio, to output the synthesized audio.

[0344] The first weight is negatively correlated with the frame count of the music audio with the preset frame length in the second preset duration, and the second weight is positively correlated with the frame count of the music audio with the preset frame length in the second preset duration.

[0345] In some embodiments, the method further includes:

[0346] When the user disables the singing function, a weighted sum is calculated based on the third weight and the music audio, and the fourth weight and the corresponding target audio, to output the synthesized audio.

[0347] The third weight is positively correlated with the frame count of the music audio with the preset frame length in the second preset duration, and the fourth weight is negatively correlated with the frame count of the music audio with the preset frame length in the second preset duration.

[0348] In some embodiments, the specific implementation of the fade-in / fade-out switching strategy is as follows: Obtain the audio x output based on normal mode and the audio y output based on karaoke mode at the current moment. Both have the same frame length and the same output timing in the audio link signal. The fade-in / fade-out coefficient is a1, a1 = (n-1) / N, where n is the frame count of the audio at the current moment, n = 1, 2, ..., N, where N is the total number of frames corresponding to the music audio to be processed. If karaoke mode is detected, the frame count of the currently output audio is 1, then the output synthesized audio is o = (1-a1)*x + a1*y, so that the audio output based on normal mode is switched to the karaoke output within N frames, i.e., the audio output in karaoke mode. For example, if the duration of a frame of music audio to be processed is 1.3ms and N is 1024, then the switch from normal mode output to karaoke mode output is completed within 1.3ms × 1024. If the operation of turning off the karaoke mode is detected, the fade-in and fade-out coefficient is a2, a2 ​​= (N-n+1) / N, where n is the frame count of the music audio at the current moment, so that the switch from the output of karaoke mode to the output of normal mode is completed within N frames.

[0349] In some embodiments, the normal mode refers to the default sound effect adjustment strategy of the in-vehicle multimedia system when playing music, while the karaoke mode refers to the sound effect adjustment strategy related to karaoke when playing music. Compared with the normal mode, the karaoke mode adds processing methods such as reverb, gain, and EQ. In some embodiments, the audio output to the amplifier can be switched between the normal mode and the karaoke mode. Specifically, when the karaoke mode is off, the output of the normal mode is played. When the karaoke mode is on, during the fade-in / fade-out switching strategy, the accompaniment audio and the original vocal audio are processed based on the normal mode. This makes the user feel that the original vocal audio gradually disappears, leaving only the accompaniment audio, based on the sound quality of the music audio processed in the normal mode. If the user starts singing, the user's audio signal and the accompaniment audio are combined and processed based on the normal mode to obtain the target audio. This results in a sound quality where the user does not perceive any change in the audio reverb between the karaoke mode and other modes.

[0350] In some embodiments, the karaoke mode includes, but is not limited to, algorithms and effects such as echo cancellation, noise reduction, reverberation, gain adjustment, and EQ adjustment. Echo cancellation is used to eliminate the music echo signal generated by the speakers playing music in the in-vehicle multimedia system. At the same time, noise reduction processing is used to reduce the noise inside and outside the vehicle that interferes with the user's audio signal. Finally, a clean user audio signal with echo and noise removed is obtained. The user's audio signal is then mixed with the accompanying audio, the source vocal audio with adjustable volume, and subjected to reverberation control, gain adjustment, EQ adjustment, etc., and finally the target audio is output to the power amplifier device for playback.

[0351] It is important to note that the above description of the process is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art can make various modifications and changes to the process under the guidance of this specification. However, these modifications and changes remain within the scope of this specification.

[0352] Figure 11 is an exemplary schematic diagram of a signal link according to some embodiments of this specification.

[0353] As shown in Figure 11, source music audio can be acquired from different signal sources 902. Users can control different karaoke modes through the interactive interface and voice control in the vehicle's multimedia system. When a user sings, their voice is captured by the microphone, converted from digital to analog, and then processed by the algorithm. The algorithm processing includes separation model, echo cancellation, howling suppression, noise reduction algorithm, and reverberation algorithm. Finally, the karaoke sound is played through a power amplifier and digital-to-analog conversion in a multi-channel speaker system.

[0354] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0355] Figure 12 is a schematic diagram of an electronic device according to some embodiments of this specification. As shown in Figure 12, the electronic device 1200 includes a processor 1201 with one or more processing cores, a memory 1202 with one or more computer-readable storage media, and computer instructions stored in the memory 1202 and executable on the processor. The processor 1201 and the memory 1202 are electrically connected.

[0356] The processor 1201 is the control center of the electronic device 1200. It connects various parts of the electronic device 1200 via various interfaces and lines. By running or loading software programs and / or units stored in the memory 1202, and by calling data stored in the memory 1202, it executes various functions and processes data of the electronic device 1200, thereby providing overall monitoring of the electronic device 1200. The processor 1201 can be a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a Network Processor (NP), etc., and can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0357] In this embodiment of the application, the processor 1201 in the electronic device 1200 loads the computer instructions corresponding to the processes of one or more applications into the memory 1202 according to the methods or steps of the above embodiments, and the processor 1201 runs the applications stored in the memory 1202 to execute the audio processing method.

[0358] Optionally, as shown in FIG12, the electronic device 1200 further includes: a touch display screen 1203, a radio frequency circuit 1204, an audio circuit 1205, an input unit 1206, and a power supply 1207. The processor 1201 is electrically connected to the touch display screen 1203, the radio frequency circuit 1204, the audio circuit 1205, the input unit 1206, and the power supply 1207. Those skilled in the art will understand that the structure shown in FIG12 does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0359] The touch display screen 1203 can be used to display a graphical user interface (GUI) and receive operation commands generated by the user interacting with the GUI. The touch display screen 1203 may include a display panel and a touch panel. The display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the vehicle. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Optionally, the display panel can be configured using a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar technologies. The touch panel can be used to collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel), generate corresponding operation commands, and execute the corresponding program according to the operation commands. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch location and the signal generated by the touch operation, transmitting the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 1201. It can also receive and execute commands from the processor 1201. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits the information to the processor 1201 to determine the type of touch event. Subsequently, the processor 1201 provides corresponding visual output on the display panel based on the type of touch event. In this embodiment, the touch panel and the display panel can be integrated into the touch display screen 1203 to achieve input and output functions. However, in some embodiments, the touch panel and the touch display screen 1203 can be implemented as two independent components to achieve input and output functions. That is, the touch display screen 1203 can also be used as part of the input unit 1206 to achieve input functions.

[0360] The radio frequency circuit 1204 can be used to transmit and receive radio frequency signals to establish wireless communication with network devices or other vehicles, and to transmit and receive signals with network devices or other vehicles.

[0361] Audio circuit 1205 can be used to provide an audio interface between the user and the vehicle via a speaker and a microphone. Audio circuit 1205 can convert received audio data into electrical signals and transmit them to the speaker, where the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuit 1205, converted back into audio data, and processed by processor 1201 before being transmitted via radio frequency circuit 1204 to, for example, another vehicle, or output to memory 1202 for further processing. Audio circuit 1205 may also include an earphone jack to provide communication between peripheral headphones and the vehicle.

[0362] The input unit 1206 can be used to receive input numbers, characters, or user characteristic information (such as fingerprints, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.

[0363] Power supply 1207 is used to supply power to various components of electronic device 1200. Optionally, power supply 1207 can be logically connected to processor 1201 through a power management device, thereby enabling functions such as charging, discharging, and power consumption management through the power management device. Power supply 1207 may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0364] Although not shown in Figure 12, the electronic device 1200 may also include a camera, sensor, wireless fidelity, Bluetooth, etc., which will not be described in detail here.

[0365] In another exemplary embodiment, a computer-readable storage medium including computer instructions is also provided, which, when executed by a processor, implement the steps of the audio processing method described above. For example, the computer-readable storage medium may be the memory 1202 including computer instructions, which may be executed by the processor 1201 of the electronic device 1200 to implement or execute the methods, steps, and logic diagrams disclosed in the embodiments of this application.

[0366] Alternatively, when executed by a computer, the instructions implement or execute the methods, steps, and logic diagrams disclosed in the embodiments of this application.

[0367] This application also provides a vehicle equipped with the electronic equipment or multimedia system provided in any of the above embodiments, wherein the electronic equipment is used to execute the audio processing method provided in any of the above embodiments. The vehicle may be a gasoline-powered vehicle, a plug-in hybrid electric vehicle, or a new energy vehicle, etc., and this specification does not specifically limit it.

[0368] The above-described embodiments are only used to illustrate the technical solutions of multimedia systems applied in vehicles, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art will understand that multimedia systems can also be used in home entertainment, commercial entertainment venues, online platforms, etc., without causing the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0369] In one embodiment, the vehicle can be configured for fully or partially autonomous driving. For example, the vehicle can control itself while in autonomous driving mode, and can determine the current state of the vehicle and its surrounding environment through human intervention, determine the possible behaviors of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the probability of that other vehicle performing a possible behavior, and control the vehicle based on the determined information. When the vehicle is in autonomous driving mode, it can be configured to operate without human interaction.

[0370] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0371] The embodiments, implementation methods, and related technical features of this application can be combined and substituted for each other without conflict.

[0372] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although the descriptions of each embodiment in this application have different focuses, and the parts not described in detail in a certain embodiment can be referred to the relevant embodiments of other embodiments, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of this application without departing from the content of the technical solution of this application shall still fall within the scope of the technical solution of this application.

Claims

1. An audio processing method, wherein, include: When playing music audio, the system detects the user's audio signal and outputs the accompaniment audio from the music audio along with the audio signal.

2. The method according to claim 1, wherein, When playing music audio, the system detects the user's audio signal and outputs the accompaniment audio from the music audio along with the audio signal, including: Identify the user's audio signal; and If the audio signal is a singing signal, then the accompaniment audio in the music audio and the audio signal are output.

3. The method according to claim 2, wherein, The process of identifying the user's audio signal includes: Based on the audio signal and the source human voice audio corresponding to the music audio, determine whether the audio signal is a singing signal.

4. The method according to claim 3, wherein, The step of determining whether the audio signal is a singing signal based on the audio signal and the source human voice audio corresponding to the music audio includes: Based on the similarity between the audio features of the audio signal and the source human voice audio corresponding to the music audio, it is determined whether the audio signal is a singing signal.

5. The method according to any one of claims 1 to 4, wherein, Before outputting the accompaniment audio in the music audio and the audio signal, the method includes: The music audio is input into a separation model for processing to obtain accompaniment audio and / or source vocal audio.

6. The method according to claim 5, wherein, The process of inputting the music audio into a separation model for processing to obtain accompaniment audio includes: The music audio is input into the separation model for processing to obtain the source human voice audio: The accompaniment audio is obtained based on the difference between the music audio and the source vocal audio.

7. The method according to claim 5, wherein, The step of inputting the music audio into a separation model for processing to obtain accompaniment audio and / or source vocal audio includes: Extract the music audio of the first preset frame length, input it into the separation model for processing, and obtain the accompaniment audio and / or the source vocal audio of the first preset frame length.

8. The method according to claim 7, wherein, The method further includes: When outputting the accompaniment audio of the first preset frame length, the music audio of the second preset frame length is extracted and input into the separation model for processing to obtain the accompaniment audio and / or the source vocal audio of the second preset frame length.

9. The method according to claim 8, wherein, The extraction of music audio of the second preset frame length includes: After extracting the music audio of the first preset frame length, the second preset frame length is determined in order to extract the music audio of the second preset frame length.

10. The method according to claim 8, wherein, The time taken for the separation model to process the music audio of the second preset frame length is less than the playback time of the accompaniment audio of the first preset frame length.

11. The method according to any one of claims 5 to 10, wherein, The separation model was trained in the following way: Obtain sample audio data, wherein the sample audio data includes corresponding sample accompaniment audio and / or sample source vocal audio; and Based on the sample audio data, the trained separation model is obtained through the separation model to be trained.

12. The method according to any one of claims 1 to 11, wherein, The method further includes: Output the accompaniment audio, the original vocal audio, and the audio signal from the music audio.

13. The method according to any one of claims 1 to 12, wherein, The method further includes: An operation to adjust the volume of the source human voice was detected, and the volume of the output source human voice audio was adjusted.

14. The method according to any one of claims 1 to 13, wherein, The method further includes: In response to user control commands, determine the audio mode; and Output music audio based on the audio pattern.

15. The method according to claim 14, wherein, The audio mode includes a first audio mode, and the step of outputting music audio according to the audio mode includes: In response to being in the first audio mode, when playing music audio, the user's audio signal is detected, and the accompaniment audio in the music audio and the audio signal are output.

16. The method according to claim 15, wherein, The method further includes: During the process of outputting the accompaniment audio in the music audio and the audio signal, if no user audio signal is detected, the accompaniment audio in the music audio or the music audio is output.

17. The method according to claim 16, wherein, The response of not detecting a user's audio signal to output either the accompaniment audio in the music audio or the music audio itself includes: Based on the duration during which no user's audio signal is detected, output either the accompaniment audio in the music audio or the music audio itself.

18. The method according to claim 17, wherein, The step of outputting either the accompaniment audio or the music audio based on the duration for which no user's audio signal is detected includes: In response to the duration being less than a first preset duration, the accompaniment audio in the music audio is output.

19. The method of claim 17, wherein, The step of outputting either the accompaniment audio or the music audio based on the duration for which no user's audio signal is detected includes: In response to the duration being greater than or equal to a first preset duration, the music audio is output.

20. The method according to claim 19, wherein, The response to the duration being greater than or equal to a first preset duration, outputting the music audio, includes: In response to the absence of the user's audio signal within the first preset duration, and the presence of source vocal audio corresponding to the accompaniment audio in the music audio, the music audio is output.

21. The method according to claim 20, wherein, The source vocal audio corresponding to the accompaniment audio in the music audio includes: The energy of the music audio is calculated. If the energy of the music audio is less than a preset energy threshold, then it is determined that there is a source human voice audio corresponding to the accompaniment audio in the music audio.

22. The method according to claim 14, wherein, The audio mode includes a second audio mode, and the step of outputting music audio according to the audio mode includes: In response to being in the second audio mode, the accompaniment audio in the music audio is output.

23. The method according to claim 22, wherein, The method further includes: During the process of outputting the accompaniment audio in the music audio, in response to detecting the user's audio signal, the accompaniment audio in the music audio and the audio signal are output.

24. The method according to any one of claims 1 to 23, wherein, The method further includes: In response to detecting the user's audio signal, the accompaniment audio in the music audio stored in the first buffer and the audio signal are output.

25. The method according to claim 24, wherein, Before the accompaniment audio in the music audio stored in the first buffer and the audio signal are output, the method further includes: Store the accompaniment audio from the music audio into the first buffer area.

26. The method of claim 25, wherein, The step of storing the accompaniment audio from the music audio into the first buffer includes: Store the accompaniment audio from the music audio with a preset frame length into the first buffer area.

27. The method according to claim 24, wherein, The method further includes: Upon detecting a progress bar adjustment, output the accompaniment audio from the music audio at the corresponding moment after the operation, indicating the progress bar's position.

28. The method according to claim 27, wherein, The method further includes: the position of the progress bar after the output operation corresponds to the time before the accompaniment audio in the music audio. Upon detecting the progress bar adjustment operation, determine whether the accompaniment audio from the music audio corresponding to the position of the progress bar after the operation exists in the first buffer: In response to the presence of accompaniment audio in the music audio corresponding to the position of the progress bar after the operation in the first buffer, the music audio after the position is extracted based on the preset frame length corresponding to the initial position; and In response to the absence of the accompaniment audio in the music audio corresponding to the position of the progress bar after the operation in the first buffer, the music audio is extracted again according to the time sequence and based on the preset frame length corresponding to the current time.

29. The method according to any one of claims 1 to 28, wherein, The method further includes: The music audio is processed to obtain music audio in a preset format.

30. The method according to claim 29, wherein, The process of processing the music audio to obtain music audio in a preset format includes: The music audio is converted into music audio with a preset number of channels.

31. The method according to claim 29, wherein, The process of processing the music audio to obtain music audio in a preset format includes: The music audio is converted into music audio with a preset sampling rate.

32. The method according to claim 29, wherein, The method further includes: The music audio in the preset format is stored in the second buffer.

33. The method according to any one of claims 1 to 32, wherein, The method further includes: If the user is detected to have disabled or enabled the singing function, a synthesized audio is output based on a preset fade-in / fade-out strategy. The synthesized audio is a combination of at least two audio components: the music audio, the accompaniment audio in the music audio, and the audio signal.

34. The method according to claim 33, wherein, The synthesized audio output based on the preset fade-in / fade-out strategy includes... Based on the preset fade-in / fade-out strategy, the second preset duration is determined; Within the second preset duration, each frame of music audio with a preset frame length and its corresponding target audio with a preset frame length are obtained. The target audio is the accompaniment audio in the music audio or an audio synthesized based on the accompaniment audio in the music audio and the audio signal. and The music audio and the corresponding target audio are weighted and summed to output the synthesized audio.

35. The method according to claim 34, wherein, The step of weighted summing of the music audio and the corresponding target audio to output synthesized audio includes: In response to the user activating the singing function, a weighted sum is performed based on the first weight and the music audio, and the second weight and the corresponding target audio, to output the synthesized audio. The first weight is negatively correlated with the frame count of the music audio with the preset frame length in the second preset duration, and the second weight is positively correlated with the frame count of the music audio with the preset frame length in the second preset duration.

36. The method according to claim 34, wherein, The method further includes: In response to the user disabling the singing function, a weighted sum is performed based on the third weight and the music audio, and the fourth weight and the corresponding target audio, to output the synthesized audio. The third weight is positively correlated with the frame count of the music audio with the preset frame length in the second preset duration, and the fourth weight is negatively correlated with the frame count of the music audio with the preset frame length in the second preset duration.

37. The method according to any one of claims 1 to 36, wherein, The music audio is obtained through at least one of the following methods: music application, video application, storage device, and user terminal.

38. A computer-readable storage medium, wherein, The computer-readable storage medium stores instructions that, when executed by a computer, cause the computer to perform the audio processing method according to any one of claims 1 to 37.

39. A computer program product, wherein, The computer program product stores instructions that, when executed by a computer, cause the computer to perform the audio processing method according to any one of claims 1 to 37.

40. An electronic device, wherein, include: Memory, on which computer instructions are stored; A processor for executing the computer instructions in the memory to implement the audio processing method according to any one of claims 1 to 37.

41. A multimedia system, wherein, The multimedia system includes: a microphone, a processor, and a power amplifier, wherein... The microphone is used to acquire the user's audio signal; The processor is configured to: detect the user's audio signal when playing music audio, and output the accompaniment audio in the music audio and the audio signal; The power amplifier is used to play the accompaniment audio and the audio signal.

42. A vehicle, wherein, This includes the electronic device of claim 40, or the multimedia system of claim 41.