Voice data processing method based on vehicle-mounted voice AI and related equipment

By generating in-car voice AI singing data in the karaoke app and mixing it with user audio, the problem of in-car karaoke software being unable to sing along with the AI ​​assistant is solved, improving user interactivity and entertainment.

CN115938340BActive Publication Date: 2025-12-23VOYAH AUTOMOBILE TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202211335986.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2025-12-23
Estimated Expiration
2042-10-28

AI Technical Summary

Technical Problem

Existing in-car karaoke software cannot sing duets or sing in pairs with AI assistants, resulting in low user interaction with AI assistants and reducing the entertainment value and product competitiveness of smart cockpits.

Method used

By selecting the duet mode on a karaoke app, the system extracts the song's audio features and lyrics features to generate in-vehicle voice AI singing data, and collects the target user's audio data in real time. It then mixes and plays three audio formats to enable the user to sing along with the AI ​​assistant.

Benefits of technology

It enhances the interactivity between users and the AI ​​assistant, as well as the entertainment features of the in-car smart cockpit, thereby strengthening the product's competitiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115938340B_ABST
    Figure CN115938340B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a voice data processing method based on vehicle-mounted voice AI and related equipment, the main purpose is to solve the current in the vehicle, based on most vehicle cockpit has the mature K song software can make user voice through the microphone after effect processing and mixing to generate the effect of the voice with reverberation, and then mixed with the song accompaniment and then sing the voice, but can't sing with AI helper or sing back, only cut the original singing chorus, the driver and passenger can't sing the song with AI helper at the same time, the user likes the song, the interaction is low in the aspect of entertainment singing, thereby leading to the problem of low competitiveness of intelligent cockpit product. Among them, the above method comprises: obtaining song audio features and lyrics features of a target song; generating vehicle-mounted voice AI singing data of the target song based on the song audio features and the lyrics features; playing the target song and the vehicle-mounted voice AI singing data simultaneously; collecting and playing audio data of a target user in real time.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent voice technology, and in particular to a voice data processing method based on vehicle-mounted voice AI and related equipment. BACKGROUND

[0002] With the improvement of people's living standards and the pursuit of a better life, the demand for self-driving entertainment is increasing. At present, in the vehicle interior, based on the mature K song software of most vehicle-mounted cabins, users can sing by themselves or sing with other users through the microphone. This kind of K song software can make the user's voice pass through effect processing and mixing after the microphone to generate a reverberation effect voice, and then mix it with the song accompaniment to emit the singing sound. However, the disadvantage is that it cannot be sung with AI assistants or duet, and can only sing the original song, that is, the driver and passenger cannot sing the song they like at the same time with the AI assistant, which has low interactivity in entertainment singing, thereby leading to low product competitiveness of the intelligent cabin. SUMMARY

[0003] In view of the above problems, the present application provides a voice data processing method based on vehicle-mounted voice AI. In the K song application, the duet mode is selected, such as the duet mode and the follow-singing mode. The corresponding mode has a corresponding song, and the song to be sung is selected. The audio features and lyrics features of the selected song are extracted to generate vehicle-mounted voice AI singing data of the selected duet song. The selected song and the vehicle-mounted voice AI singing data are played, and the audio data of the target user is collected in real time. Finally, the above three kinds of audio are mixed and played, which improves the interactivity between the user and the AI assistant and the entertainment and product competitiveness of the in-vehicle intelligent cabin.

[0004] To solve the above at least one technical problem, in a first aspect, the present application provides a voice data processing method based on vehicle-mounted voice AI, which comprises:

[0005] obtaining song audio features and lyrics features of a target song;

[0006] generating vehicle-mounted voice AI singing data of the target song based on the song audio features and lyrics features;

[0007] playing the target song and the vehicle-mounted voice AI singing data at the same time;

[0008] collecting and playing audio data of a target user in real time.

[0009] Optionally, the obtained song audio features and lyrics features of the target song comprise:

[0010] obtaining the song audio features and lyrics features of the target song based on a K song application; or

[0011] acquire audio data of the target song based on a karaoke application;

[0012] determine song audio features and lyric features of the target song based on the audio data through a vehicle-mounted voice AI application.

[0013] Optionally, the method further comprises:

[0014] The vocal line feature of the AI singing voice of the vehicle-mounted voice AI singing voice data is the same as the current voice vocal line feature set by the vehicle-mounted voice AI application.

[0015] Optionally, the method further comprises:

[0016] determine the singing preference of the target user based on historical karaoke data of the target user;

[0017] adjust the vehicle-mounted voice AI singing voice data according to the singing preference.

[0018] Optionally, the vehicle-mounted voice AI singing voice data of the target song is generated based on the song audio features and the lyric features, comprising: transmitting the audio features and the lyric features to the karaoke application based on an AIDL cross-process communication mode.

[0019] Optionally, the method further comprises:

[0020] adjust the current voice vocal line set by the vehicle-mounted voice AI application based on the vocal line preference setting of the target user.

[0021] Optionally, the method further comprises:

[0022] based on the vocal line feature of the AI singing voice and / or

[0023] The song audio features of the target song are adaptively adjusted to the internal environment of the vehicle.

[0024] In a second aspect, the embodiments of the present application further provide a voice data processing device based on a vehicle-mounted voice AI, comprising:

[0025] an acquisition unit configured to acquire song audio features and lyric features of a target song;

[0026] a generation unit configured to generate vehicle-mounted voice AI singing voice data of the target song based on the song audio features and the lyric features;

[0027] a loudspeaker unit configured to simultaneously play the target song and the vehicle-mounted voice AI singing voice data;

[0028] a collection and playing unit configured to collect and play audio data of a target user in real time.

[0029] To achieve the above object, according to a third aspect of the present application, a computer readable storage medium is provided, which comprises a stored program, wherein the program, when executed by a processor, implements the above-mentioned voice data processing method based on vehicle-mounted voice AI.

[0030] To achieve the above object, according to a fourth aspect of the present application, an electronic device is provided, which comprises at least one processor and at least one memory connected to the processor; wherein the processor is configured to invoke program instructions in the memory to execute the above-mentioned voice data processing method based on vehicle-mounted voice AI.

[0031] By the above technical solution, the present application provides a voice data processing method based on vehicle-mounted voice AI. At present, in the vehicle interior, based on the mature K song software that most vehicle cabins have, the user's voice can be processed and mixed through the microphone to generate an effect voice with reverberation, and then mixed with the song accompaniment to emit the singing sound, but cannot be sung with the AI assistant or duet, and can only sing with the original singer, that is, the driver and the passenger cannot sing the song the user likes at the same time with the AI assistant, and the interactivity in the entertainment singing is low, thereby causing the problem that the product competitiveness of the intelligent cabin is not high. The present application selects a duet mode, such as a duet mode and a follow-singing mode, on the K song application, and the corresponding mode has a corresponding song. The song to be sung is selected, the audio features and the lyrics features of the selected song are extracted to generate vehicle-mounted voice AI singing data of the selected duet song, the selected song and the vehicle-mounted voice AI singing data are played, and the audio data of the target user is collected in real time. Finally, the above three kinds of audio are mixed and played, which improves the interactivity between the user and the AI assistant, and the entertainment and product competitiveness of the in-vehicle intelligent cabin.

[0032] The above description is only a summary of the technical solution of the present application. In order to more clearly understand the technical means of the present application, the content of the specification can be implemented, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS

[0033] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are intended to only illustrate preferred embodiments and are not considered limiting of the present application. Moreover, like reference numerals designate similar parts throughout the several views in the drawings. In the drawings:

[0034] Figure 1 A flowchart of a voice data processing method based on vehicle-mounted voice AI provided by an embodiment of the present application is shown;

[0035] Figure 2A schematic structural block diagram of a voice data processing device based on vehicle-mounted voice AI provided by an embodiment of the present application is shown.

[0036] Figure 3 A schematic structural block diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0037] Exemplary embodiments of the present application will be described in greater detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it is understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present application can be more thoroughly understood, and the scope of the present application can be accurately conveyed to those skilled in the art.

[0038] To solve the problem that the driver and the passenger cannot sing the song that the user likes at the same time with the AI assistant, can only cut the original song, the interaction with the vehicle in the aspect of entertainment singing is low, thereby leading to low product competitiveness of the smart cabin, embodiments of the present application provide a voice data processing method based on vehicle-mounted voice AI, as shown in Figure 1 The method comprises the following steps.

[0039] S101, obtaining song audio features and lyrics features of a target song.

[0040] Exemplarily, the above actual application scenario can be that the vehicle is in a starting state, and a vehicle-mounted infotainment product, i.e., a vehicle machine, installed in the car is operated. In terms of function, the AI assistant can select a duet mode according to the installed K song application, the corresponding duet mode has subordinate classified songs, the duet mode can be a counterpoint mode or a follow-singing mode, and a song to be sung is selected. According to the selected duet song, the audio features and the lyrics features of the selected song are obtained, the audio features can be phoneme information, tone information, prosodic boundary text information, note information, beat information, and tied note score information of the selected song, and the lyrics features can be built-in text lyrics information of the selected song, or lyrics information obtained by analyzing the audio of the selected song.

[0041] S102, generating vehicle-mounted voice AI singing data of the target song based on the song audio features and the lyrics features.

[0042] It should be noted that the AI assistant generates the vehicle-mounted voice AI singing data of the selected song by configuring an AI singing synthesis scheme model of a deep learning neural network algorithm according to the phoneme information, the tone information, the prosodic boundary text information, the note information, the beat information, the tied note score information, and the text lyrics information of the selected song obtained in S101, wherein the deep learning algorithm model is configured after model training using a large amount of data.

[0043] For example, according to the operation, the selected duet song by the car machine is "Good Luck", the lyric text information in the "Good Luck" song, the phoneme information corresponding to each font of the lyric text information, the tone information of each text sentence, the prosodic boundary text information of the lyric text segmentation, the note information of the selected song accompaniment audio, the beat information, and the tie note score information

[0044] Phoneme: is the smallest unit of speech according to the natural properties of speech, which is analyzed according to the pronunciation action in the syllable. A phoneme is composed of a single action. Phonemes are divided into two categories: vowels and consonants. For example, the Chinese syllable ah (ā) has only one phoneme, love (ài) has two phonemes, and generation (dài) has three phonemes.

[0045] Tone: refers to the change of tone in language. Tone refers to the inherent voice of a Chinese syllable that can distinguish meaning. The pitch of the tone is relative, not absolute; the change of the tone is sliding, not jumping from one scale to another. The high and low of the tone is usually marked by a five-degree mark: a vertical mark, divided into 5 degrees, the lowest is 1, and the highest is 5.

[0046] Prosodic boundary: Prosodic boundary plays an important role in naturalness and accuracy. In people's communication, the part of the pause between sentences is the prosodic boundary.

[0047] S103, play the target song and the car voice AI singing data simultaneously.

[0048] It can be understood that the karaoke application in the car machine will output the selected target song and the car voice AI singing data after audio processing through the loudspeaker device. The target song can be changed to output the song with the original song or only output the target song with the accompaniment by turning on the original song or turning off the original song. Simultaneous playback of the car voice AI singing and the original song target song or only outputting the target song with the accompaniment can lay a good foundation for subsequent user duet, making the overall audio output experience better, and freely turning on the original song can also realize three voice lines duet at the same time when the user inputs the sound source, and further improve the richness of singing.

[0049] S104, real-time collection and playback of target user audio data.

[0050] It should be noted that the car machine collects the sound information input by the user of the karaoke application through the external or built-in sound source collection device in real time while playing the selected song in duet mode, and outputs the audio through the playback device after processing the collected information through echo cancellation technology.

[0051] For example, when the user uses the KTV application, the external microphone is connected to the USB interface of the vehicle machine. After selecting the song corresponding to the preferred duet mode by the vehicle machine, the user sings and inputs the voice source within a reasonable angle and sound source collection range. At this time, the loudspeaker of the vehicle machine plays the song accompaniment and the vehicle-mounted voice AI singing data in S103. The vehicle machine system filters and processes the sound source output by the loudspeaker using sound echo cancellation technology, and then mixes the user input sound source with the song accompaniment and the vehicle-mounted voice AI singing data and outputs it again with low delay.

[0052] Echo cancellation: Echo canceller is a processing method that prevents the return of the far-end sound by eliminating or removing the far-end audio signal picked up by the local microphone. This removal of audio is done through digital signal processing. The basic principle of echo cancellation is to establish a speech model of the far-end signal based on the correlation between the loudspeaker signal and the multi-path echo generated by it, estimate the echo using it, and continuously modify the filter coefficients to make the estimated value closer to the true echo. Then, the echo estimate is subtracted from the input signal of the microphone to achieve the purpose of eliminating the echo.

[0053] In the above scheme, it can be ensured that for the current vehicle interior, based on the mature KTV software that most vehicle cabins already have, the user can sing by himself or with other users through the microphone. This kind of application software can make the user's voice pass through effect processing and mixing to generate a voice with reverb effect, and then mix it with the song accompaniment to emit the singing sound. However, it cannot be duet or sing with the AI assistant. Only the original song can be cut, and the driver and the passenger cannot sing the song the user likes at the same time with the AI assistant. The interaction in terms of entertainment singing is low, which leads to the problem that the product competitiveness of the intelligent cabin is not high. The above method of the present application can select a duet mode on the KTV application, and then select a song to be duet. The audio features and lyrics features of the selected song are extracted to generate vehicle-mounted voice AI singing data of the selected duet song. The selected song and the vehicle-mounted voice AI singing data are played, and the audio data of the target user is collected in real time. Finally, the above three kinds of audio are mixed and played, which realizes the effect of improving the interaction between the user and the AI assistant, the entertainment of the intelligent cabin in the vehicle, and the product competitiveness.

[0054] In some embodiments, the above method is executed, S101 obtains the song audio features and lyrics features of the target song, including:

[0055] S201-A, obtaining the song audio features and lyrics features of the target song based on the KTV application;

[0056] It should be noted that the audio features and lyrics features of the target song mentioned in S101 can be directly obtained through analysis by a karaoke application. For example, the karaoke application can directly call the karaoke audio file of the target song that has been cached internally or downloaded, and analyze it to obtain the audio features and lyrics features of the aforementioned audio file.

[0057] S201-B: Obtain the audio data of the target song based on the karaoke application; determine the song's audio features and lyrics features based on the audio data using an in-vehicle voice AI application.

[0058] It should be noted that the difference between S201-B and S201-A lies in that S201-B first transmits the internally cached or downloaded audio file of the target song from the karaoke application to the in-vehicle voice AI application. Then, the AI ​​assistant analyzes the audio file to obtain the song's audio and lyric features. The advantage of this two-step audio file analysis approach is that the system can determine whether the karaoke application or the in-vehicle voice AI application parses the audio file based on the system's workload, resulting in a smoother user experience.

[0059] This method can be understood as having two types, A and B. Either one can be executed during the process, and both can achieve the effect of obtaining the audio features and lyric features of the target song.

[0060] In some embodiments, the above method, when executed, includes:

[0061] S301, The vocal characteristics of the AI ​​singing data in the vehicle voice AI are the same as the currently set vocal characteristics of the vehicle voice AI application.

[0062] It's worth noting that the voice assistants in car infotainment systems offer a variety of voices, including male, female, and regional dialects such as Cantonese, Sichuanese, and Northeastern Mandarin. In daily use, users often become familiar with the voices of the AI ​​assistants based on their personalized settings, creating a natural sense of familiarity and companionship. By synchronizing the AI's vocal characteristics with those of the in-car voice assistant, the system can provide a sense of connection, like singing along with a familiar voice. This makes the AI's voice more than just a cold, impersonal presence, enhancing the overall harmony and user experience. It also prevents the use of different voices for certain notification sounds, which could disrupt the in-car entertainment and singing atmosphere.

[0063] Exemplarily, the voice sound line setting of the in-vehicle voice AI application is set as a soft female voice, and then the sound line feature of the AI singing voice is set to be synchronized to the sound line set in the in-vehicle voice AI application. After the setting is completed, the audio output by the loudspeaker according to the audio features and lyric features of the selected song processed based on the AI singing voice is the audio based on the soft female voice sound line.

[0064] In some embodiments, the above method, when executed, further comprises:

[0065] S401, based on the historical K song data of the target user, determining the singing preference of the target user.

[0066] It should be noted that during each singing process of the user, the car machine memorizes the singing habits and emotional fluctuations of the user according to the sound source input by the user through the sound source acquisition device, and saves the singing preference parameters of the user to the database after analysis by the AI singing synthesis scheme model of the neural network algorithm. The singing preference can be personalized and classified by setting different storage names.

[0067] Exemplarily, for example, driver A, passenger B, baby C. The car machine will remind the user to store, and the user can choose whether to perform classified storage. If classified storage is performed, when the K song application is started again, the user will automatically match the stored singing preference through the sound line feature when singing with the AI assistant through the K song application, so as to improve the singing experience.

[0068] S402, adjusting the in-vehicle voice AI singing voice data according to the singing preference.

[0069] It should be noted that the in-vehicle AI voice singing voice data will adapt to the singing preference of the user during the singing process according to the singing preference of the user determined in S401, so that the whole singing is more harmonious and pleasant. The above preference can be the emotional expression of the user singing, which can be high-pitched, excited, lost, sad, etc., and the emotional expression and volume of the AI assistant are adaptively adjusted according to the above preference.

[0070] Exemplarily, the historical singing data of user A in the database is mostly sad type songs, and the singing habit is mostly low sound volume and little pitch parameter fluctuation. When user A uses the K song application to select the singing mode and sing "coral sea" with the AI assistant, the AI assistant will adapt to the singing preference of user A, adjust the pitch parameter, rhythm boundary and other parameters similar to user A, so as to better complete the cooperative singing of the song.

[0071] In some embodiments, the above method, when executed, S102 generates in-vehicle voice AI singing voice data of the target song based on the song audio features and lyric features, comprising:

[0072] S501, transmit the audio features and the lyrics features to the K-song application based on an AIDL cross-process communication mode.

[0073] It can be understood that, according to the S102 step, after the AI assistant generates the singing data of the vehicle-mounted voice AI of the selected song based on the audio features and the lyrics features through the AI singing synthesis scheme model configured by the deep learning neural network algorithm, the singing data of the voice AI needs to be transmitted to the K-song application for mixing and playing together with the target song. The K-song application and the AI assistant belong to two independent processes in the car machine, so a way is needed to transmit the processed singing data of the AI across processes.

[0074] AIDL: Android Interface Definition Language, an interface definition language based on Android compilation. Because in Android, different applications run in their own independent processes, and applications cannot access each other's memory space. In order to realize inter-process communication, the PCI mechanism is used. Android supports the PCI mechanism, but needs to have serialized data readable by Android, and AIDL is used to describe the above data.

[0075] In some embodiments, the above method, when executed, further comprises:

[0076] S601, adjust the current voice line setting of the vehicle-mounted voice AI application based on the target user's voice line preference setting

[0077] It should be noted that different users have different preferences for the voice line set by the vehicle-mounted voice AI application. As can be seen from S301, the voice line feature of the AI singing data of the vehicle-mounted voice AI can be changed along with the current voice line feature set by the vehicle-mounted voice AI application. Therefore, the user can adjust the voice line setting of the current vehicle-mounted voice AI application according to his own preference, and then synchronize the voice line feature of the AI singing to his preferred voice line setting. The voice line can be a soft female voice, a deep male voice, etc.

[0078] For example, during the process of using the K-song application to sing, the car machine will identify user A as a male voice or a female voice, and adaptively switch the gender voice line feature of the AI singing. If the voice line feature of the AI singing of the voice AI singing data has been pre-set to be the same as the current voice line feature set by the vehicle-mounted voice AI application, but the voice line of the vehicle-mounted voice AI application is different from the adaptively switched gender voice line feature of the AI singing, the priority of this step is defaulted to be higher than that of step S301, and the voice line setting is overwritten to achieve the purpose of improving the singing experience.

[0079] In some embodiments, the above method, when executed, further comprises:

[0080] S701: Adaptively adjusts the interior environment atmosphere of the vehicle based on the vocal characteristics of the AI ​​singer and / or the audio characteristics of the target song.

[0081] It should be noted that the in-car lighting and shading devices can adaptively adjust the interior environment based on the vocal characteristics of the AI ​​singer and / or the audio characteristics of the target song, better integrating the scene with the atmosphere of the selected target song, so that the smart cockpit in the car is no longer just a carrier for singing songs, but part of the overall singing atmosphere.

[0082] For example, when a user selects the song "Hair Like Snow" through a karaoke app to sing along with the AI ​​assistant, the sunshade on the car window automatically closes, the grayscale of the vehicle glass is adaptively adjusted, and the light transmittance and saturation are reduced. The ambient lighting inside the car is adjusted to ice blue to simulate a snowy scene, and the speaker adjusts the low-frequency, mid-frequency, and high-frequency parameters according to the song to achieve a good singing environment.

[0083] It should be noted that, as a response to the above... Figure 1 In addition to the implementation of the methods shown in various related embodiments, this invention also provides a voice data processing device based on in-vehicle voice AI for processing the aforementioned... Figure 1 The methods described in the above embodiments are implemented accordingly. This device embodiment corresponds to the foregoing method embodiments. For ease of reading, this device embodiment will not repeat the details of the foregoing method embodiments one by one, but it should be clear that the device in this embodiment can implement all the contents of the foregoing method embodiments. Figure 2 As shown, the device includes:

[0084] The acquisition unit 21 is used to acquire the audio features and lyrics features of the target song;

[0085] Generation unit 22 is used to generate in-vehicle voice AI singing data for the target song based on the song's audio features and lyrics features;

[0086] Speaker unit 23 is used to simultaneously play the target song and the in-vehicle voice AI singing data;

[0087] The acquisition and playback unit 24 is used to acquire and play audio data from the target user in real time.

[0088] By employing the above technical solutions, this invention provides a voice data processing method based on in-vehicle voice AI. Currently, in vehicles, most mature karaoke software allows users to sing, generating a reverb-effected voice through microphone processing and mixing, which is then blended with the accompaniment. However, it doesn't allow for duets or duets with AI assistants; only the original vocals can be used. This means drivers and passengers cannot simultaneously sing along to the user's favorite songs with the AI ​​assistant, resulting in low interactivity and reduced competitiveness of the smart cockpit. This invention addresses this issue by allowing users to select a duet mode (e.g., duet or sing-along) within the karaoke application, with corresponding songs for each mode. The system extracts audio and lyric features from the selected song to generate in-vehicle voice AI vocal data for that song. It then plays the selected song and the in-vehicle voice AI vocal data while simultaneously collecting the target user's audio data. Finally, it mixes and plays all three types of audio, improving user-AI assistant interactivity and enhancing the entertainment value and competitiveness of the in-vehicle smart cockpit.

[0089] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured, and by adjusting kernel parameters, a voice data processing method based on in-vehicle voice AI can be implemented. This addresses the existing problems of limited interactivity with the vehicle's infotainment system, such as the inability to sing duets or duets with AI assistants, the inability to switch to the original vocals, and the inability for drivers and passengers to simultaneously sing along to their favorite songs with the AI ​​assistant.

[0090] This invention provides a storage medium storing a program that, when executed by a processor, implements a voice data processing method based on in-vehicle voice AI.

[0091] This invention provides a processor for running a program, wherein the program executes a voice data processing method based on in-vehicle voice AI during runtime.

[0092] This invention provides a device 30, such as... Figure 3 As shown, the device includes at least one processor 31, and at least one memory 32 and bus 33 connected to the processor; wherein, the processor 31 and the memory 32 communicate with each other through the bus 33; the processor 31 is used to call program instructions in the memory to execute the above-mentioned voice data processing method based on vehicle voice AI.

[0093] The devices mentioned in this article can be servers, PCs, tablets, mobile phones, etc.

[0094] The application also provides a computer program product suitable for executing the program steps of initializing the method when executed on a data processing device: obtaining song audio features and lyrics features of a target song; generating vehicle-mounted voice AI singing data of the target song based on the song audio features and lyrics features; playing the target song and the vehicle-mounted voice AI singing data simultaneously; and collecting and playing audio data of a target user in real time.

[0095] Further, the obtained song audio features and lyrics features of the target song include: obtaining the song audio features and lyrics features of the target song based on a K-song application; or,

[0096] Obtaining the audio data of the target song based on the K-song application;

[0097] Determining the song audio features and lyrics features of the target song based on the audio data through the vehicle-mounted voice AI application.

[0098] Further, the above method further includes:

[0099] The sound line features of the AI singing of the vehicle-mounted voice AI singing data are the same as the currently set voice sound line features of the vehicle-mounted voice AI application.

[0100] Further, the above method further includes:

[0101] Determining the singing preference of the target user based on the historical K-song data of the target user;

[0102] Adjusting the vehicle-mounted voice AI singing data according to the singing preference.

[0103] Further, the above generating the vehicle-mounted voice AI singing data of the target song based on the song audio features and lyrics features includes: transmitting the audio features and lyrics features to the K-song application based on an AIDL cross-process communication mode.

[0104] Further, the above method further includes:

[0105] Adjusting the currently set voice sound line of the vehicle-mounted voice AI application based on the target user sound line preference setting.

[0106] Further, the above method further includes:

[0107] Adjusting the currently set voice sound line of the vehicle-mounted voice AI application based on the target user sound line preference setting.

[0108] The song audio features of the target song are adaptively adjusted to the vehicle interior environment atmosphere.

[0109] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions specified in the flowchart block or blocks. Figure 1 The flowchart and / or block diagram in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments. In this regard, each block in the flowchart and / or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable Figure 1 The flowchart and / or block diagram in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments. In this regard, each block in the flowchart and / or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable

[0110] In one typical configuration, the device includes one or more processors (CPUs), memory, and a bus. The device can also include an input / output interface, a network interface, and the like.

[0111] The memory can include non-persistent memory and / or volatile memory, e.g., random access memory (RAM) comprising a number of memory locations that can be read and / or written on the fly. The memory can also include non-volatile memory, e.g., read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), flash memory (flash RAM), and the like. The memory includes at least one memory chip. The memory is an example of computer-readable media.

[0112] Computer-readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0113] It is also to be noted that the terms "comprising", "including", and any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0114] Those skilled in the art will appreciate that embodiments of the present application can be devised for a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer-readable program code.

[0115] The embodiments of the present application are only illustrative and are not intended to limit the present application. Various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A voice data processing method based on in-vehicle voice AI, characterized by, include: Obtain the audio and lyric features of the target song; The in-vehicle voice AI singing data of the target song is generated based on the song's audio features and lyrics features; Simultaneously play the target song and the in-vehicle voice AI singing data; Real-time acquisition and playback of audio data from the target user; Based on the target user's historical karaoke data, the target user's singing preferences are determined; The in-vehicle voice AI singing data is adjusted according to the singing preferences; Based on the vocal characteristics of the AI ​​singing and / or The target song's audio features adaptively adjust the atmosphere inside the vehicle. Audio features and lyrics features are transmitted to the karaoke application based on AIDL cross-process communication. The acquisition of the target song's audio features and lyric features includes: Based on karaoke applications, obtain the audio and lyric features of the target song; or, Obtain the audio data of the target song based on a karaoke application; Based on the audio data, the in-vehicle voice AI application determines the song's audio features and lyrics features. The vocal characteristics of the AI ​​singing voice data in the vehicle voice AI are the same as the currently set vocal characteristics of the vehicle voice AI application. The process of generating in-vehicle voice AI singing data for the target song based on the song's audio features and lyrics features includes: The audio features and lyrics features are transmitted to the karaoke application based on AIDL cross-process communication. Based on the target user's voice preference settings, adjust the currently set voice tone of the in-vehicle voice AI application. 2.A voice data processing apparatus based on in-vehicle voice AI, characterized by, include: The acquisition unit is used to acquire the audio features and lyric features of the target song. The generation unit is used to generate in-vehicle voice AI singing data of the target song based on the song's audio features and lyrics features; A speaker unit is used to simultaneously play the target song and the in-vehicle voice AI singing data; The acquisition and playback unit is used to acquire and play audio data from the target user in real time. Based on the target user's historical karaoke data, the target user's singing preferences are determined; The in-vehicle voice AI singing data is adjusted according to the singing preferences; Based on the vocal characteristics of the AI ​​singing and / or The target song's audio features adaptively adjust the atmosphere inside the vehicle. Audio features and lyrics features are transmitted to the karaoke application based on AIDL cross-process communication. The acquisition of the target song's audio features and lyric features includes: Based on karaoke applications, obtain the audio and lyric features of the target song; or, Obtain the audio data of the target song based on a karaoke application; Based on the audio data, the in-vehicle voice AI application determines the song's audio features and lyrics features. The vocal characteristics of the AI ​​singing voice data in the vehicle voice AI are the same as the currently set vocal characteristics of the vehicle voice AI application. The process of generating in-vehicle voice AI singing data for the target song based on the song's audio features and lyrics features includes: The audio features and lyrics features are transmitted to the karaoke application based on AIDL cross-process communication. Adjust the current voice line of the vehicle-mounted voice AI application based on the target user's voice line preference setting.

3. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a stored program, wherein the program, when executed by a processor, implements the voice data processing method based on the vehicle-mounted voice AI as claimed in claim 1.

4. An electronic device, comprising: The electronic device comprises at least one processor and at least one memory connected to the processor; wherein the processor is configured to invoke program instructions in the memory to execute the voice data processing method based on the vehicle-mounted voice AI as claimed in claim 1.

Citation Information

Patent Citations

  • Sound effect switching method and user terminal

    CN105025415A

  • Light control method and device of vehicle-mounted entertainment system and vehicle-mounted entertainment system

    CN111332197A

  • Voice processing method, terminal equipment and vehicle

    CN114071318A

  • Audio production method and device, terminal equipment and readable storage medium

    CN114974184A

  • Duet part singing generation system

    JP2009244607A