Audio processing method, system and equipment
By acquiring user performance data and processing it to generate audio data, the problem of the lack of interactive music experience in existing audio systems has been solved. It enables the mixing of user-initiated performance with the original audio, thereby enhancing in-car music interaction and immersive experience.
Patent Information
- Application Number
- CN202511138700.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-10-28
AI Technical Summary
Existing audio systems lack an interactive music experience; users can only passively listen to audio and cannot interact with the system in real time.
By acquiring user performance data, the system generates audio data corresponding to the instrument tracks played by the user. The original audio data is then split into tracks, tracks to be muted are deleted, and the audio data corresponding to each track is synthesized and played through the speakers. Combined with seat vibration and sound field rendering, the system achieves a blending of the user's active performance with the original audio.
It enhances the user's interactive music experience, enabling users to actively play music and interact with the original audio, thus strengthening the immersive and creative experience within the car.
Smart Images

Figure CN120853531A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to an audio processing method, system and device. Background Technology
[0002] Currently, most audio systems focus on audio playback as their core function, employing fixed speaker layouts and preset sound effect parameters, primarily concentrating on sound quality optimization, noise reduction, and sound field compensation.
[0003] Existing audio systems typically only play audio, which is merely used as a background playback medium, resulting in a poor interactive music experience for users. Summary of the Invention
[0004] This application provides an audio processing method, system, and device to enhance the user's interactive music experience.
[0005] According to a first aspect of the embodiments of this application, an audio processing method is provided, comprising:
[0006] Obtain user performance data;
[0007] Based on the user's performance data, audio data corresponding to the track of the instrument played by the user is generated;
[0008] The original audio data is divided into tracks to obtain the audio data corresponding to each original track.
[0009] From the audio data corresponding to each original audio track, delete the audio data of the track to be muted to obtain the audio data of the target audio track.
[0010] Based on the audio data corresponding to the track of the instrument played by the user and the audio data of the target track, the audio data corresponding to each track is obtained.
[0011] According to a second aspect of the embodiments of this application, an audio processing system is provided, including an input component, a processing component, and an output component;
[0012] The input component is used to collect initial audio data; the initial audio data includes user performance data and original music audio data;
[0013] The processing component is configured to execute the audio processing method as described in the first aspect through a processor, convert the initial audio data collected by the input component into audio data corresponding to each audio track, and synthesize target audio data corresponding to each speaker based on the audio data corresponding to each audio track.
[0014] The output component is used to output the target audio data corresponding to each speaker to each speaker.
[0015] Optionally, the input component is specifically used to collect user performance data through an external musical instrument connected to the target vehicle, and / or to collect user performance data generated by user operation by displaying a virtual musical instrument interface on the vehicle's in-vehicle display screen; and to collect original audio data from the in-vehicle player.
[0016] Optionally, the output component is further configured to extract audio data of percussion instrument tracks from the target audio data corresponding to each speaker; determine vibration parameters based on the audio data of the percussion instrument tracks; and control the seat vibration module of the target vehicle to vibrate according to the vibration parameters while each speaker plays its corresponding target audio data.
[0017] Optionally, the processing component is specifically used to determine the mapping relationship between each audio track and each speaker; and to synthesize target audio data corresponding to each speaker based on the audio data corresponding to each audio track and the mapping relationship.
[0018] The output component is also used to render a sound field heatmap on the central control screen of the target vehicle; wherein the sound field heatmap includes the sound pressure level of each audio track and the mapping relationship.
[0019] According to a third aspect of the embodiments of this application, an electronic device is provided, including a memory and a processor;
[0020] The memory is connected to the processor and is used to store programs;
[0021] The processor is used to implement the audio processing method as described in the first aspect by running a program in the memory.
[0022] In this application, based on user performance data, audio data corresponding to the track of the user's instrument is generated. The original audio data is divided into tracks to obtain the audio data corresponding to each original track. From the audio data corresponding to each original track, the audio data of the track to be muted is deleted to obtain the audio data of the target track. Based on the audio data corresponding to the track of the user's instrument and the audio data of the target track, the audio data corresponding to each track is obtained. This allows the instrument position in the original song to be "made room" for the user during real-time performance. The audio data of the track to be muted in the original song audio data is replaced with the audio data corresponding to the track of the user's instrument. This achieves "controllable replacement" or muting of the original accompaniment content. It also enables the mixing of user performance data and original song audio data before playback, allowing the user to actively perform and interact with the original song audio data, thus enhancing the user's interactive music experience. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0024] Figure 1 This is a flowchart illustrating an audio processing method provided in an embodiment of this application.
[0025] Figure 2 This is a flowchart illustrating step 102 provided in an embodiment of this application.
[0026] Figure 3 This is a flowchart illustrating an audio processing method provided in an embodiment of this application.
[0027] Figure 4 This is a flowchart illustrating step 302 provided in an embodiment of this application.
[0028] Figure 5 This is a flowchart illustrating an audio processing method provided in an embodiment of this application.
[0029] Figure 6 This is a schematic diagram of the structure of an audio processing system provided in an embodiment of this application.
[0030] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0032] Exemplary Implementation Environment
[0033] The audio processing method according to the embodiments of this application can be executed by electronic devices such as terminal devices or servers. The terminal device can be an in-vehicle device, a user device, a mobile device, a computing device, a wearable device, etc., and the server can be a standalone physical server, a server cluster composed of multiple physical servers, or a cloud server capable of cloud computing. This method can be implemented by a processor calling computer-readable program instructions stored in memory. The following embodiments use the execution of the audio processing method by an in-vehicle device as an example for explanation, but do not limit the scope of the application.
[0034] Exemplary methods
[0035] Please see Figure 1 In one exemplary embodiment, an audio processing method is provided. For example... Figure 1 As shown, the audio processing method mainly includes the following steps:
[0036] Step 101: Obtain user performance data.
[0037] In an exemplary embodiment, the user performance data may include MIDI (Musical Instrument Digital Interface) signals played by the user, or pre-recorded audio data of the user performance, or audio signals of the user performance acquired in real time. This application does not limit this; the following embodiments use the example of user performance data including MIDI signals played by the user for explanation.
[0038] In some embodiments, step 101 may be implemented in ways including, but not limited to, the following:
[0039] Method 1
[0040] Obtain user performance data from the musical instruments connected to the target vehicle.
[0041] In an exemplary embodiment, an external MIDI instrument (such as a piano keyboard, drum pad, etc.) can be accessed through the USB (Universal Serial Bus) interface or Bluetooth interface of the target vehicle's in-vehicle device to obtain the MIDI signal played by the user output by the external MIDI instrument.
[0042] In the exemplary embodiment, the onboard equipment of the target vehicle is compatible with GM (General MIDI) and MPE (MIDI Polyphonic Expression) protocols, uniformly parsing the MIDI signals played by the user and input through the hardware. MPE allows each note to independently control multiple parameters (such as pitch and velocity).
[0043] In an exemplary embodiment, an external MIDI instrument can be accessed through the vehicle's in-vehicle equipment, or multiple external MIDI instruments can be accessed through the vehicle's in-vehicle equipment, enabling multiple users and multiple instruments to perform collaboratively within the vehicle.
[0044] Method 2
[0045] Display a virtual musical instrument interface on the in-vehicle display screen of the target vehicle, acquire user operations received by the virtual musical instrument interface, and generate user performance data based on the user operations.
[0046] In an exemplary embodiment, the virtual instrument interface provides a variety of virtual instruments for the user to choose from. After selecting a virtual instrument, the user performs user operations on it. These operations can include various types, such as clicking, swiping, dragging, and multi-touch. Based on the user operations, user performance data is generated, which may include: determining the user's touch position on the in-vehicle display screen based on the user operations; determining the position of the touch point on the virtual instrument based on the touch position, and then determining the note number corresponding to the touch; determining the velocity value based on the user's touch pressure on the in-vehicle display screen; and generating a MIDI signal for the user's performance based on the note number and the velocity value. The MIDI signal for the user's performance is independent of the type of user operation, but is related to the user's touch position on the in-vehicle display screen and the user's touch pressure on the in-vehicle display screen.
[0047] In the exemplary embodiment, a virtual instrument can be selected for performance on the virtual instrument interface, or multiple virtual instruments can be selected for performance on the virtual instrument interface, enabling multiple users and multiple instruments to perform collaboratively in the vehicle.
[0048] Method 3
[0049] Acquire user performance data from the musical instruments connected to the target vehicle; and display a virtual musical instrument interface on the vehicle's in-vehicle display screen, acquire user operations received by the virtual musical instrument interface, and generate user performance data based on the user operations.
[0050] In an exemplary embodiment, an external MIDI instrument can be connected through the vehicle's in-vehicle device, and a virtual instrument can be selected for performance in the virtual instrument interface; alternatively, an external MIDI instrument can be connected through the vehicle's in-vehicle device, and multiple virtual instruments can be selected for performance in the virtual instrument interface; or multiple external MIDI instruments can be connected through the vehicle's in-vehicle device, and multiple virtual instruments can be selected for performance in the virtual instrument interface; thus enabling collaborative performance by multiple users and multiple instruments within the vehicle.
[0051] With the continuous evolution of intelligent cockpits and in-vehicle entertainment systems, users' needs for in-vehicle audio experiences are shifting from "passive listening" to "active participation." In traditional in-vehicle audio systems, audio is mainly played passively, with audio serving only as a background playback medium. There is a lack of real-time music interaction mechanisms in the in-vehicle environment, which cannot support users to actively play music or engage in multimodal music interaction with the in-vehicle audio system, thus limiting the expansion boundaries of the in-vehicle entertainment experience.
[0052] In this application, an external musical instrument is connected through the vehicle's in-vehicle equipment, and / or a virtual musical instrument interface is displayed on the vehicle's in-vehicle display screen, allowing users to play the virtual musical instrument. This enables real-time acquisition of the user's MIDI signals, creating an intuitive and engaging performance scenario and constructing a creative and engaging "in-vehicle music stage." This greatly expands the entertainment potential of the smart cockpit and enhances the in-vehicle interactive music experience.
[0053] Step 102: Based on the user's performance data, generate audio data corresponding to the track of the instrument played by the user.
[0054] In some embodiments, such as Figure 2 As shown, step 102 includes:
[0055] Step 201: Extract the note number and dynamic value from the user's performance data.
[0056] In the exemplary embodiment, the user performance data includes the MIDI signal played by the user. Before step 401, the MIDI signal played by the user can be uniformly parsed into an internal data structure according to the timestamp, containing note number, velocity value, channel information, and control commands. This facilitates the subsequent extraction of the target timbre and the synthesis of audio data corresponding to the track of the user's instrument. Here, channel information refers to the track corresponding to each instrument; control commands refer to control instructions for the instrument, such as turning on a note, adjusting a button on the instrument, etc.
[0057] Step 202: Extract the target timbre from the preset timbre set based on the note number and velocity value.
[0058] In an exemplary embodiment, step 202 may include: extracting a target timbre from a local preset timbre set or a preset timbre set sampled with high precision by a third party, based on the note number and velocity value. For example, multiple timbres are found from the preset timbre set based on the note number and velocity value, and these timbres have priorities among them; the target timbre is extracted according to the priority order.
[0059] Step 203: Based on the user's performance data and the target timbre, synthesize the audio data corresponding to the track of the instrument played by the user.
[0060] In an exemplary embodiment, step 203 may include: synthesizing user performance data and target timbre into audio data corresponding to the track of the user's instrument using timbre mapping real-time synthesis technology, wherein the audio data corresponding to the track of the user's instrument may be a PCM (Pulse Code Modulation) waveform, thereby generating a high-fidelity "user performance stream".
[0061] It extracts note numbers and velocity values from user performance data, extracts target timbres from a preset timbre set based on note numbers and velocity values, and synthesizes audio data corresponding to the user's instrument track based on user performance data and target timbres. It can select target timbres that match the user's performance data according to the note numbers and velocity values in the user performance data, and synthesize audio data corresponding to the user's instrument track in real time, achieving high-fidelity audio synthesis of the user's instrument performance.
[0062] Step 103: Divide the original audio data into tracks to obtain the audio data corresponding to each original track.
[0063] In an exemplary embodiment, step 103 may include: using AI (Artificial Intelligence) to separate the original audio data into tracks, such as drums, bass, guitar, and vocals, to obtain the audio data corresponding to each original track.
[0064] In an exemplary embodiment, the original audio data may include the original audio data obtained from the in-vehicle player.
[0065] In some embodiments, in addition to acquiring user performance data and original audio data, auxiliary control signals can also be acquired through in-vehicle equipment to assist in audio processing.
[0066] In an exemplary embodiment, the auxiliary control signals may include user commands such as switching instruments, muting instrument tracks, recalling sound field presets, or adjusting performance parameters. Switching instruments refers to changing the virtual instrument A initially selected by the user in the virtual instrument interface to a newly selected virtual instrument B; muting instrument tracks refers to the user selecting a specific instrument track from the original audio data of the muted track; recalling sound field presets refers to recalling the preset mapping relationships between the stored tracks and speakers; adjusting performance parameters is a conventional user command and will not be described in detail here.
[0067] In the exemplary embodiment, the auxiliary control signal can be obtained by the user operating the in-vehicle display screen, or by the in-vehicle device performing voice recognition of the user's voice, or by the in-vehicle device performing gesture capture of the user's gesture. This application does not limit this.
[0068] Step 104: Delete the audio data of the track to be muted from the audio data corresponding to each original track to obtain the audio data of the target track.
[0069] In an exemplary embodiment, a mute control can be provided in the user interface of the vehicle display screen, and the mute interface control permission can be opened, allowing the user to mute any original audio track of the original music data. The original audio track selected by the user to be mute is the audio track to be mute.
[0070] Step 105: Based on the audio data corresponding to the track of the instrument played by the user and the audio data of the target track, obtain the audio data corresponding to each track.
[0071] Based on user performance data, the system generates audio data corresponding to the tracks of the user's instrument. The original audio data is then divided into tracks, resulting in audio data for each track. From the audio data of each original track, the audio data of the track to be muted is removed, yielding the audio data of the target track. Based on the audio data of the user's instrument and the target track, the system obtains the audio data for each track. This allows the system to "make room" for the instrument in the original music during real-time performance, replacing the audio data of the track to be muted with the audio data of the user's instrument. This enables "controllable replacement" or muting of the original accompaniment content. Furthermore, the system allows for the mixing of user performance data and original music audio data before playback, enabling active user interaction with the original music audio data and enhancing the user's interactive music experience.
[0072] In some embodiments, such as Figure 3 As shown, the audio processing method also includes:
[0073] Step 301: Determine the mapping relationship between each audio track and each speaker.
[0074] In the exemplary embodiment, each speaker can be multiple speakers arranged in a certain space. For example, each speaker can be a speaker in the target vehicle, including headrest speakers, door panel speakers, center speaker, etc., arranged in different positions in the target vehicle; each speaker can also be a speaker in other spaces, and this application does not limit this. In the following embodiments, each speaker is used as an example of each speaker in the target vehicle for explanation and illustration.
[0075] In some embodiments, step 301 includes: adjusting the preset mapping relationship between each audio track and each speaker based on the sound field positioning offset command, so as to obtain the mapping relationship between each audio track and each speaker.
[0076] In an exemplary embodiment, the preset mapping relationship between each audio track and each speaker may include a preset channel mapping matrix, for example, positioning the drum sound to the rear surround speakers and placing the main melody to the center speaker.
[0077] In the exemplary embodiment, the preset mapping relationship between each audio track and each speaker can be a one-to-one correspondence between audio tracks and speakers, or each audio track can correspond to multiple speakers, or a one-to-one correspondence between some audio tracks and speakers, with each audio track in some audio tracks corresponding to multiple speakers. This application does not limit this.
[0078] In an exemplary embodiment, the sound field positioning offset command may include a user dragging a certain audio track from an initial speaker position to a new speaker position on the in-vehicle display screen; the sound field positioning offset command may also include a sound field positioning offset command given by the user via voice.
[0079] Based on the sound field positioning offset command, the preset mapping relationship between each track and each speaker is adjusted to obtain the mapping relationship between each track and each speaker. Although the position of the speaker is fixed, the sound field positioning offset command allows users to flexibly adjust the mapping relationship between each track and each speaker. The sound image position is dynamically allocated and managed according to different instrument types, improving the flexibility of speaker sound effect feedback and enhancing the user's immersive music experience.
[0080] In some other embodiments, step 301 includes: determining a preset mapping relationship between each audio track and each speaker as a mapping relationship between each audio track and each speaker.
[0081] Step 302: Based on the audio data and mapping relationship corresponding to each audio track, synthesize the target audio data corresponding to each speaker.
[0082] In scenarios involving multi-instrument collaborative output, the audio from different instruments is typically assigned to different audio tracks. This application can determine the mapping relationship between each audio track and each speaker, and adopt different speaker allocation methods for the audio data corresponding to different audio tracks. Furthermore, based on the audio data corresponding to each audio track and the mapping relationship, it synthesizes the target audio data corresponding to each speaker, thereby realizing the allocation of audio data to different speakers according to the audio track, and the synthesis of the allocated audio data in each speaker. It can adjust the speaker allocation method corresponding to the audio data according to the audio track, and improve the flexibility of speaker sound effect feedback.
[0083] In some embodiments, such as Figure 4 As shown, step 302 includes:
[0084] Perform the following operations for any one speaker:
[0085] Step 401: Based on the mapping relationship, determine the audio adjustment parameters for each audio track played on any speaker.
[0086] In an exemplary embodiment, audio adjustment parameters may include gain and delay.
[0087] In an exemplary embodiment, step 401 may include: based on the mapping relationship between each audio track and each speaker, by looking up a table, the gain and delay of each audio track when played on any speaker can be found. The table may pre-store the correspondence between audio tracks, speakers, gain, and delay.
[0088] Step 402: Based on the audio adjustment parameters of each audio track played on any speaker, adjust the audio data corresponding to each audio track to obtain the adjusted audio data corresponding to each audio track.
[0089] Step 403: Synthesize the adjusted audio data corresponding to each audio track to obtain the target audio data corresponding to any speaker.
[0090] In an exemplary embodiment, step 403 may include: synthesizing the adjusted audio data corresponding to each audio track in a multitrack mixer to obtain target audio data corresponding to any one speaker. The target audio data corresponding to any one speaker is multi-channel audio data with sound field effects, i.e., multi-channel PCM.
[0091] Based on the mapping relationship, the audio data corresponding to each track is adjusted, and the adjusted audio data of each track is synthesized to obtain the target audio data for any speaker. This allows for corresponding adjustments to the target audio data for each speaker during the mapping adjustment process, improving the overall audio presentation effect of each speaker playing its respective target audio data and enhancing the user's immersive music experience. By generating audio data for each track based on user performance data and original song audio data, immersive multi-channel mixing output of the user's performance track and the accompaniment track can be achieved.
[0092] In some other embodiments, step 302 includes performing the following operations for any one speaker: determining the audio data corresponding to the audio track played on any one speaker based on the mapping relationship; and synthesizing the audio data corresponding to the audio track played on any one speaker to obtain the target audio data corresponding to any one speaker.
[0093] In some embodiments, the audio processing method further includes: sending the target audio data corresponding to each speaker to the audio signal processing unit for sound effect processing to obtain the processed audio data corresponding to each speaker; and outputting the processed audio data corresponding to each speaker to each speaker after digital-to-analog conversion and power amplification.
[0094] In an exemplary embodiment, the audio signal processing unit's sound effect processing process specifically includes: adapting the sound effects of the target audio data corresponding to each speaker according to the interior space layout of the target vehicle, adjusting sound effect parameters such as gain and delay to make the sound effects more compatible with the specific vehicle.
[0095] In some embodiments, such as Figure 5 As shown, the audio processing method also includes:
[0096] Step 501: Extract the audio data of the percussion instrument track from the target audio data corresponding to each speaker.
[0097] In an exemplary embodiment, step 501 may include: sending the target audio data corresponding to each speaker to the audio signal processing unit for sound effect processing to obtain the processed audio data corresponding to each speaker; and extracting the audio data of the percussion instrument track from the target audio data corresponding to each speaker.
[0098] In an exemplary embodiment, percussion instruments may include drums, gongs, etc.
[0099] Step 502: Determine vibration parameters based on the audio data of the percussion instrument track.
[0100] In an exemplary embodiment, vibration parameters may include amplitude, etc.
[0101] Step 503: While each speaker is playing its corresponding target audio data, the seat vibration module of the target vehicle is controlled to vibrate according to the vibration parameters.
[0102] In an exemplary embodiment, the seat vibration module may be a vibration exciter embedded in the seat. For example, the vibration exciter can be understood as a low-frequency speaker that does not emit sound (the frequency of the audio signal is too low for the human ear to hear), but only produces a vibration effect.
[0103] The audio data of the percussion instrument track is extracted from the target audio data corresponding to each speaker. Based on the audio data of the percussion instrument track, the vibration parameters are determined. While each speaker plays its corresponding target audio data, the seat vibration module of the target vehicle is controlled to vibrate according to the vibration parameters. This allows the user to feel the corresponding vibration effect according to the audio, providing tactile feedback that matches the striking force and enhancing the user's immersive experience.
[0104] In some embodiments, the audio processing method further includes: rendering a sound field heatmap on the central control screen of the target vehicle; wherein the sound field heatmap includes the sound pressure level of each audio track and the mapping relationship.
[0105] It can intuitively display the sound pressure level and mapping relationship of each audio track, helping users to observe and adjust the sound field arrangement at any time.
[0106] In summary, this application generates audio data corresponding to the track of the user's instrument based on user performance data. The original audio data is divided into tracks to obtain the audio data corresponding to each original track. Audio data of the track to be muted is deleted from the audio data of each original track to obtain the audio data of the target track. Based on the audio data corresponding to the user's instrument track and the audio data of the target track, the audio data corresponding to each track is obtained. This allows the instrument position in the original song to be "made room" for the user during real-time performance. The audio data of the track to be muted in the original song is replaced with the audio data corresponding to the track of the user's instrument, achieving "controllable replacement" or muting of the original accompaniment content. It also enables the mixing of user performance data and original song audio data before playback, allowing for active user interaction with the original song audio data and enhancing the user's interactive music experience.
[0107] Exemplary System
[0108] Accordingly, embodiments of this application also provide an audio processing system, such as... Figure 6 As shown, the audio processing system includes an input component 601, a processing component 602, and an output component 603;
[0109] Input component 601 is used to acquire initial audio data; the initial audio data includes user performance data and original music audio data.
[0110] The processing component 602 is used to execute any of the audio processing methods provided in the above embodiments of this application through the processor, convert the initial audio data collected by the input component 601 into audio data corresponding to each audio track, and synthesize the target audio data corresponding to each speaker based on the audio data corresponding to each audio track.
[0111] Output component 603 is used to output the target audio data corresponding to each speaker to each speaker.
[0112] In some embodiments, the input component 601 is specifically used to collect user performance data via an external musical instrument connected to the target vehicle, and / or to collect user performance data generated by user operation by displaying a virtual musical instrument interface on the vehicle's in-vehicle display screen; and to collect original audio data from the in-vehicle player.
[0113] In an exemplary embodiment, an external MIDI instrument (such as a piano keyboard, drum pad, etc.) can be connected via the USB or Bluetooth interface of the vehicle's in-vehicle equipment to obtain the MIDI signal played by the user output from the external MIDI instrument. The audio processing system is compatible with GM and MPE protocols and uniformly parses the MIDI signal played by the user input from the hardware.
[0114] In an exemplary embodiment, the virtual instrument interface provides a variety of virtual instruments for the user to choose from. After selecting a virtual instrument, the user performs user operations on it. These operations can include various types, such as clicking, swiping, dragging, and multi-touch. Based on the user operations, user performance data is generated, which may include: determining the user's touch position on the in-vehicle display screen based on the user operations; determining the position of the touch point on the virtual instrument based on the touch position, and then determining the note number corresponding to the touch; determining the velocity value based on the user's touch pressure on the in-vehicle display screen; and generating a MIDI signal for the user's performance based on the note number and the velocity value. The MIDI signal for the user's performance is independent of the type of user operation, but is related to the user's touch position on the in-vehicle display screen and the user's touch pressure on the in-vehicle display screen.
[0115] Input component 601 collects user performance data through an external musical instrument connected to the target vehicle, and / or collects user performance data generated by user operation by displaying a virtual musical instrument interface on the vehicle's in-vehicle display screen. This enables real-time acquisition of MIDI signals from the user's performance, creating an intuitive and engaging performance scenario and constructing a creative and engaging "in-vehicle music stage." This greatly expands the entertainment potential of the smart cockpit and enhances the in-vehicle interactive music experience.
[0116] In some embodiments, the input component 601 is also used to acquire auxiliary control signals. These auxiliary control signals are then used to assist in audio processing.
[0117] In an exemplary embodiment, the auxiliary control signals may include user commands such as switching instruments, muting instrument tracks, recalling sound field presets, or adjusting performance parameters. Switching instruments refers to changing the virtual instrument A initially selected by the user in the virtual instrument interface to a newly selected virtual instrument B; muting instrument tracks refers to the user selecting a specific instrument track from the original audio data of the muted track; recalling sound field presets refers to recalling the preset mapping relationships between the stored tracks and speakers; adjusting performance parameters is a conventional user command and will not be described in detail here.
[0118] In the exemplary embodiment, the auxiliary control signal can be obtained by the user operating the in-vehicle display screen, or by the in-vehicle device performing voice recognition of the user's voice, or by the in-vehicle device performing gesture capture of the user's gesture. This application does not limit this.
[0119] In some embodiments, the processing component 602 is specifically configured to: acquire user performance data; generate audio data corresponding to the track of the user's instrument based on the user performance data; divide the original audio data into tracks to obtain audio data corresponding to each original track; delete the audio data of the track to be muted from the audio data corresponding to each original track to obtain the audio data of the target track; and obtain the audio data corresponding to each track based on the audio data corresponding to the track of the user's instrument and the audio data of the target track.
[0120] In the exemplary embodiment, AI track splitting can be used to split the original audio data in the car player into tracks, breaking down the drum, bass, guitar, vocal, and other audio tracks to obtain the audio data corresponding to each original track.
[0121] In an exemplary embodiment, a mute control can be provided in the user interface of the vehicle display screen, and the mute interface control permission can be opened, allowing the user to mute any original audio track of the original music data. The original audio track selected by the user to be mute is the audio track to be mute.
[0122] It can "give up" the positions of instruments in the original song when users play in real time, and replace the audio data of the track to be muted in the original song audio data with the audio data corresponding to the track of the instrument played by the user. This enables "controllable replacement" or muting of the original accompaniment content. It can also mix the user's performance data and the original song audio data and then play it out, allowing users to actively play and interact with the original song audio data, thus enhancing the user's interactive music experience.
[0123] In some embodiments, the processing component 602 is specifically used to determine the mapping relationship between each audio track and each speaker; and to synthesize the target audio data corresponding to each speaker based on the audio data corresponding to each audio track and the mapping relationship.
[0124] In scenarios involving multi-instrument collaborative output, the audio from different instruments is typically assigned to different audio tracks. This application can determine the mapping relationship between each audio track and each speaker, and adopt different speaker allocation methods for the audio data corresponding to different audio tracks. Furthermore, based on the audio data corresponding to each audio track and the mapping relationship, it synthesizes the target audio data corresponding to each speaker, thereby realizing the allocation of audio data to different speakers according to the audio track, and the synthesis of the allocated audio data in each speaker. It can adjust the speaker allocation method corresponding to the audio data according to the audio track, and improve the flexibility of speaker sound effect feedback.
[0125] In some embodiments, the processing component 602 is specifically used to: extract note numbers and velocity values from user performance data; extract a target timbre from a preset timbre set based on the note numbers and velocity values; and synthesize audio data corresponding to the track of the user's instrument based on the user performance data and the target timbre.
[0126] In an exemplary embodiment, the MIDI signal played by the user can be uniformly parsed into an internal data structure according to the timestamp, including note number, velocity value, channel information and control commands, which facilitates the subsequent extraction of target timbre and synthesis of audio data corresponding to the track of the user's instrument.
[0127] In an exemplary embodiment, a target timbre can be extracted from a local set of preset timbres or a set of preset timbres sampled with high precision from a third party, based on the note number and velocity value. For example, multiple timbres can be found from the preset timbre set based on the note number and velocity value, and these timbres have priorities; the target timbre is extracted according to the priority order.
[0128] In an exemplary embodiment, timbre mapping real-time synthesis technology can be used to synthesize user performance data and target timbre into audio data corresponding to the track of the user's instrument.
[0129] It can select a target timbre that matches the user's performance data based on the note number and velocity value in the user's performance data, synthesize the audio data corresponding to the track of the user's instrument, and synthesize high-fidelity audio of the user's instrument in real time.
[0130] In some embodiments, the processing component 602 is specifically used to: adjust the preset mapping relationship between each audio track and each speaker based on the sound field positioning offset instruction, so as to obtain the mapping relationship between each audio track and each speaker.
[0131] In an exemplary embodiment, the preset mapping relationship between each audio track and each speaker may include a preset channel mapping matrix, for example, positioning the drum sound to the rear surround speakers and placing the main melody to the center speaker.
[0132] In the exemplary embodiment, the preset mapping relationship between each audio track and each speaker can be a one-to-one correspondence between audio tracks and speakers, or each audio track can correspond to multiple speakers, or a one-to-one correspondence between some audio tracks and speakers, with each audio track in some audio tracks corresponding to multiple speakers. This application does not limit this.
[0133] In an exemplary embodiment, the sound field positioning offset command may include a user dragging a certain audio track from an initial speaker position to a new speaker position on the in-vehicle display screen; the sound field positioning offset command may also include a sound field positioning offset command given by the user via voice.
[0134] Although the speaker positions are fixed, the sound field positioning offset command allows users to flexibly adjust the mapping relationship between each track and each speaker, dynamically allocate and manage the sound image position according to different instrument types, improve the flexibility of speaker sound effect feedback, and enhance the user's immersive music experience.
[0135] In some embodiments, the processing component 602 is specifically configured to perform the following operations for any one speaker: based on the mapping relationship, determine the audio adjustment parameters of each audio track played on any one speaker; based on the audio adjustment parameters of each audio track played on any one speaker, adjust the audio data corresponding to each audio track to obtain the adjusted audio data corresponding to each audio track; and synthesize the adjusted audio data corresponding to each audio track to obtain the target audio data corresponding to any one speaker.
[0136] In an exemplary embodiment, audio adjustment parameters may include gain and delay.
[0137] In the exemplary embodiment, the adjusted audio data corresponding to each audio track can be synthesized in a multitrack mixer to obtain the target audio data corresponding to any one speaker. The target audio data corresponding to any one speaker is multi-channel audio data with sound field effects, i.e., multi-channel PCM.
[0138] When adjusting the mapping relationship, it can also adjust the target audio data corresponding to each speaker, improving the overall audio presentation effect brought by each speaker playing its corresponding target audio data, and enhancing the user's immersive music experience. Based on the user's performance data and the original song's audio data, it generates audio data corresponding to each track, enabling immersive mixed output of the user's performance track and the accompaniment track in multiple channels.
[0139] In some embodiments, the output component 603 is further configured to send the target audio data corresponding to each speaker to the audio signal processing unit for sound effect processing to obtain the processed audio data corresponding to each speaker; and output the processed audio data corresponding to each speaker to each speaker after digital-to-analog conversion and power amplification.
[0140] In an exemplary embodiment, the audio signal processing unit's sound effect processing process specifically includes: adapting the sound effects of the target audio data corresponding to each speaker according to the interior space layout of the target vehicle, adjusting sound effect parameters such as gain and delay to make the sound effects more compatible with the specific vehicle.
[0141] In an exemplary embodiment, each speaker can be a speaker in the target vehicle, including headrest speakers, door panel speakers, center speaker, etc., arranged in different positions in the target vehicle.
[0142] In some embodiments, the output component 603 is further configured to extract audio data of percussion instrument tracks from the target audio data corresponding to each speaker; determine vibration parameters based on the audio data of the percussion instrument tracks; and control the seat vibration module of the target vehicle to vibrate according to the vibration parameters while each speaker plays its corresponding target audio data.
[0143] In an exemplary embodiment, the target audio data corresponding to each speaker can be sent to the audio signal processing unit for sound effect processing to obtain the processed audio data corresponding to each speaker; and the audio data of the percussion instrument track can be extracted from the target audio data corresponding to each speaker.
[0144] In an exemplary embodiment, percussion instruments may include drums, gongs, etc.
[0145] In an exemplary embodiment, vibration parameters may include amplitude, etc.
[0146] In an exemplary embodiment, the seat vibration module may be a vibration exciter embedded in the seat. For example, the vibration exciter can be understood as a low-frequency speaker that does not emit sound (the frequency of the audio signal is too low for the human ear to hear), but only produces a vibration effect.
[0147] It enables users to feel corresponding vibration effects based on audio, providing tactile feedback that matches the force of the impact, thus enhancing the user's immersive experience.
[0148] In some embodiments, the processing component 602 is specifically used to determine the mapping relationship between each audio track and each speaker; and to synthesize the target audio data corresponding to each speaker based on the audio data corresponding to each audio track and the mapping relationship.
[0149] The output component 603 is also used to render a sound field heatmap on the central control screen of the target vehicle; wherein the sound field heatmap includes the sound pressure level of each audio track and the mapping relationship.
[0150] It can intuitively display the sound pressure level and mapping relationship of each audio track, helping users to observe and adjust the sound field arrangement at any time.
[0151] The audio processing system provided in this embodiment belongs to the same concept as the audio processing method provided in the above embodiments of this application. It can execute the audio processing method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects for executing the audio processing method. Technical details not described in detail in this embodiment can be found in the specific processing content of the audio processing method provided in the above embodiments of this application, and will not be repeated here.
[0152] It should be understood that the above audio processing system can be implemented by a processor calling software. For example, the audio processing system includes a processor connected to memory, which stores instructions. The processor calls the instructions stored in memory to implement any of the above methods or to implement the functions of each component of the audio processing system. The processor can be a general-purpose processor, such as a CPU or microprocessor, and the memory can be internal or external to the audio processing system. Alternatively, the components in the audio processing system can be implemented as hardware circuits. By designing the hardware circuits, some or all of the component functions can be implemented. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above components are implemented by designing the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a PLD, such as an FPGA, which can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files to implement the functions of some or all of the above components. All components of the above system can be implemented entirely by a processor calling software, entirely by hardware circuits, or partially by a processor calling software with the remaining parts implemented by hardware circuits.
[0153] In this application embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, microprocessor, GPU, or DSP. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented as an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the processor loading instructions to implement the functions of some or all of the above components. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, or DPU.
[0154] As can be seen, each component in the above system can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor types.
[0155] Furthermore, the components in the above systems can be integrated in whole or in part, or they can be implemented independently. In one implementation, these components are integrated together as a System-on-Chip (SoC). This SoC may include at least one processor for implementing any of the above methods or for implementing the functions of the components in the system. The at least one processor can be of different types, such as CPU and FPGA, CPU and AI processor, CPU and GPU, etc.
[0156] Exemplary electronic devices
[0157] One embodiment of this application discloses an electronic device, see [link to relevant documentation] Figure 7 As shown, the device includes:
[0158] Memory 200 and processor 210;
[0159] The memory 200 is connected to the processor 210 and is used to store programs;
[0160] The processor 210 is configured to implement the audio processing method disclosed in any of the above embodiments by running the program stored in the memory 200.
[0161] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 220, an input device 230, and an output device 240.
[0162] The processor 210, memory 200, communication interface 220, input device 230, and output device 240 are interconnected via a bus. Among them:
[0163] A bus can include a pathway for transmitting information between various components of a computer system.
[0164] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0165] Processor 210 may include a main processor, as well as a baseband chip, modem, etc.
[0166] The memory 200 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 200 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.
[0167] Input device 230 may include a means of receiving data and information input by the user, such as an external MIDI instrument, a vehicle display screen, a camera, a voice input device, etc.
[0168] Output device 240 may include devices that allow information to be output to a user, such as an in-vehicle display, a seat vibration module, a speaker, etc.
[0169] The communication interface 220 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0170] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement the various steps of any of the audio processing methods provided in the above embodiments of this application.
[0171] In an exemplary embodiment, the electronic device may include an in-vehicle device.
[0172] Exemplary computer program products and storage media
[0173] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the audio processing methods according to various embodiments of this application as described in any of the above embodiments of this specification.
[0174] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0175] Furthermore, embodiments of this application may also be storage media storing a computer program, which is executed by a processor through which the steps of the audio processing method according to various embodiments of this application described in any of the above embodiments of this specification are performed, specifically implementing the following steps:
[0176] Step 101: Obtain user performance data.
[0177] Step 102: Based on the user's performance data, generate audio data corresponding to the track of the instrument played by the user.
[0178] Step 103: Divide the original audio data into tracks to obtain the audio data corresponding to each original track.
[0179] Step 104: Delete the audio data of the track to be muted from the audio data corresponding to each original track to obtain the audio data of the target track.
[0180] Step 105: Based on the audio data corresponding to the track of the instrument played by the user and the audio data of the target track, obtain the audio data corresponding to each track.
[0181] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0182] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system-type embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0183] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.
[0184] The components in the system of the various embodiments of this application can be merged, divided, and deleted according to actual needs.
[0185] It should be understood that the disclosed systems and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the system embodiments described above are merely illustrative; for instance, the division of components is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple components may be combined or integrated into another component, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between components may be electrical, mechanical, or other forms.
[0186] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.
[0187] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.
[0188] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0189] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0190] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0191] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An audio processing method, characterized in that, include: Obtain user performance data; Based on the user's performance data, audio data corresponding to the track of the instrument played by the user is generated; The original audio data is divided into tracks to obtain the audio data corresponding to each original track. From the audio data corresponding to each original audio track, delete the audio data of the track to be muted to obtain the audio data of the target audio track. Based on the audio data corresponding to the track of the instrument played by the user and the audio data of the target track, the audio data corresponding to each track is obtained.
2. The audio processing method according to claim 1, characterized in that, The method further comprises: Determine the mapping relationship between each audio track and each speaker; Based on the audio data corresponding to each audio track and the mapping relationship, the target audio data corresponding to each speaker is synthesized.
3. The audio processing method according to claim 1, characterized in that, The acquisition of user performance data includes: Obtain user performance data from the musical instruments connected to the target vehicle; and / or, A virtual musical instrument interface is displayed on the in-vehicle display screen of the target vehicle. User operations received by the virtual musical instrument interface are acquired, and user performance data is generated based on the user operations.
4. The audio processing method according to claim 1, characterized in that, The step of generating audio data corresponding to the track of the user's instrument based on the user's performance data includes: Extract the note number and velocity value from the user's performance data; Based on the note number and the velocity value, extract the target timbre from the preset timbre set; Based on the user's performance data and the target timbre, audio data corresponding to the track of the instrument played by the user is synthesized.
5. The audio processing method according to claim 2, characterized in that, Determining the mapping relationship between each audio track and each speaker includes: Based on the sound field positioning offset command, the preset mapping relationship between each audio track and each speaker is adjusted to obtain the mapping relationship between each audio track and each speaker.
6. The audio processing method according to claim 2, characterized in that, The process of synthesizing target audio data corresponding to each speaker based on the audio data corresponding to each audio track and the mapping relationship includes: Perform the following operations for any one speaker: Based on the mapping relationship, determine the audio adjustment parameters for each audio track when played on any one of the speakers; Based on the audio adjustment parameters of each audio track played on any one speaker, the audio data corresponding to each audio track is adjusted to obtain the adjusted audio data corresponding to each audio track. The adjusted audio data corresponding to each audio track are synthesized to obtain the target audio data corresponding to any one speaker.
7. The audio processing method according to any one of claims 2, 5, and 6, characterized in that, The method further comprises: Extract the audio data of the percussion instrument track from the target audio data corresponding to each speaker; Vibration parameters are determined based on audio data from percussion instrument tracks; While each speaker plays its corresponding target audio data, the seat vibration module of the target vehicle is controlled to vibrate according to the vibration parameters.
8. The audio processing method according to any one of claims 2, 5, and 6, characterized in that, The method further comprises: A sound field heatmap is rendered on the central control screen of the target vehicle; wherein the sound field heatmap includes the sound pressure level of each audio track and the mapping relationship.
9. An audio processing system, characterized in that, It includes input components, processing components, and output components; The input component is used to collect initial audio data; the initial audio data includes user performance data and original music audio data; The processing component is configured to execute the audio processing method as described in any one of claims 1 to 6 through a processor, convert the initial audio data collected by the input component into audio data corresponding to each audio track, and synthesize the target audio data corresponding to each speaker based on the audio data corresponding to each audio track. The output component is used to output the target audio data corresponding to each speaker to each speaker.
10. The audio processing system according to claim 9, characterized in that, The input component is specifically used to collect user performance data through an external musical instrument connected to the target vehicle, and / or to collect user performance data generated by user operation by displaying a virtual musical instrument interface on the vehicle's in-vehicle display screen; and to collect original audio data from the in-vehicle player.
11. The audio processing system according to claim 9 or 10, characterized in that, The output component is also used to extract audio data of percussion instrument tracks from the target audio data corresponding to each speaker; determine vibration parameters based on the audio data of the percussion instrument tracks; and control the seat vibration module of the target vehicle to vibrate according to the vibration parameters while each speaker plays its corresponding target audio data.
12. The audio processing system according to claim 11, characterized in that, The processing component is specifically used to determine the mapping relationship between each audio track and each speaker; and to synthesize the target audio data corresponding to each speaker based on the audio data corresponding to each audio track and the mapping relationship. The output component is also used to render a sound field heatmap on the central control screen of the target vehicle; wherein the sound field heatmap includes the sound pressure level of each audio track and the mapping relationship.
13. An electronic device, characterized in that, Including memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the audio processing method as described in any one of claims 1 to 8 by running a program in the memory.
Citation Information
Patent Citations
Musical instrument capable of audio track replacement
CN106652655A
Audio playing method and device, terminal and storage medium
CN113823250A
Three-dimensional audio system
CN118972776A
The method that the output of power for accompaniment
KR1020060006247A
A speaker for a car
KR2019980041229U