Vehicle-mounted music generation method and device, electronic equipment and storage medium
By acquiring user input and vehicle driving information, and using AI large-scale models to generate music that is highly relevant to the in-vehicle scene, the problem of low relevance between music and scene in in-vehicle entertainment systems is solved, providing a unique music experience.
Patent Information
- Application Number
- CN202410524448.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-29
- Publication Date
- 2025-10-31
AI Technical Summary
Existing in-vehicle entertainment systems cannot meet users' music needs in in-vehicle scenarios, and the correlation between music and in-vehicle scenarios is low.
By acquiring user input information and current vehicle driving information, the system uses a preset data model to generate audio segments that are highly relevant to the in-vehicle environment, including gear information, driving mode, steering wheel information, and pedal information. The system then uses a large AI model for recognition and processing to automatically generate music content that is relevant to the user and the in-vehicle environment.
It achieves a high degree of correlation between in-car music and in-car scenarios, providing a unique music experience, meeting users' music needs in in-car scenarios, and avoiding the homogenization of music experiences.
Smart Images

Figure CN120877685A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle technology, and in particular to a method, apparatus, electronic device, and storage medium for generating in-vehicle music. Background Technology
[0002] With the development of vehicle technology, vehicles are becoming increasingly intelligent. Among these technologies, vehicles are usually equipped with entertainment systems. The entertainment and interactivity of in-vehicle entertainment systems have become new growth points in user demand, providing music for drivers and passengers and keeping them in a pleasant mood.
[0003] Currently, many smart vehicle products have made numerous attempts around the concept of "music cockpit," including enriching music applications and content, providing a variety of sound effects suitable for in-vehicle spaces, and offering tuning options. However, current in-vehicle entertainment systems typically match and select available playlists based on user operations, or recommend music that users like based on their historical playback records. The in-vehicle music obtained through these matching methods has a low correlation with the in-vehicle scenario and cannot meet the user's music needs in the in-vehicle environment. Summary of the Invention
[0004] In view of this, the present invention aims to propose a method, device, electronic device and storage medium for generating in-vehicle music, to solve the problem that the correlation between in-vehicle music and the in-vehicle scene is low and cannot meet the music needs of users in the in-vehicle scene, and to achieve the effect of automatically generating music that is highly correlated with the in-vehicle scene.
[0005] According to a first aspect of the present invention, a method for generating in-vehicle music is provided, the method comprising:
[0006] In response to the activation of the vehicle's in-car music function, it obtains user input information and the vehicle's current driving information;
[0007] Based on the user input information and the current driving information, the target prompt words corresponding to the music features are determined;
[0008] The target prompt word is identified using a preset data model, audio parameters are determined based on the target prompt word, and at least one audio segment is generated using the audio parameters.
[0009] In response to a playback control command, the vehicle is controlled to play the audio clip.
[0010] Optionally, the step of acquiring user input information and current vehicle driving information in response to the activation of the vehicle's in-vehicle music function includes:
[0011] In response to the activation of the vehicle's in-vehicle music function, the system detects user input operations on the preset interactive interface and obtains user input information.
[0012] The vehicle's current driving status is monitored, and based on the driving status, current driving information is generated, including gear information, driving mode, steering wheel information, and pedal information.
[0013] Optionally, determining the target prompt word corresponding to the music feature based on the user input information and the current driving information includes:
[0014] The changes in the vehicle's driving actions are determined by using the gear information, driving mode, steering wheel information, and pedal information from the vehicle's current driving information.
[0015] Acquire road condition and environmental information when the vehicle's current driving action changes;
[0016] The user input information, changes in driving actions, road condition information, and environmental information are fused and processed to generate target prompt words.
[0017] Optionally, the step of fusing the user input information, changes in driving actions, road condition information, and environmental information to generate target prompt words includes:
[0018] The user input information, changes in driving actions, road condition information, and environmental information are used to extract features and generate at least two keywords.
[0019] At least two of the aforementioned feature prompt words are subjected to feature fusion processing to generate target prompt words; wherein, the target prompt words are used to characterize music features.
[0020] Optionally, the step of using a preset data model to identify the target prompt word, determining audio parameters based on the target prompt word, and generating at least one audio segment using the audio parameters includes:
[0021] The target prompt word is input into a preset data model, and the target prompt word is identified and processed using the preset data model;
[0022] Based on the correlation between natural language and music features in the preset data model, the audio parameters corresponding to the target prompt word are output;
[0023] At least one audio segment is generated using the audio parameters, and the audio segment is stored in the playlist according to the generation time.
[0024] Optionally, after generating at least one audio segment using the audio parameters and storing the audio segment in the playlist according to the generation time, the method further includes:
[0025] Monitor the number of audio segments in the playlist;
[0026] If the number of audio segments is greater than or equal to a preset storage threshold, the audio segments are deleted in rotation, and the playlist is updated.
[0027] According to a second aspect of the present invention, an in-vehicle music generation device is provided, the device comprising:
[0028] The acquisition module is used to acquire user input information and the vehicle's current driving information in response to the activation of the vehicle's in-vehicle music function.
[0029] The determination module is used to determine the target prompt word corresponding to the music feature based on the user input information and the current driving information;
[0030] The generation module is used to identify the target prompt word using a preset data model, determine audio parameters based on the target prompt word, and generate at least one audio segment using the audio parameters.
[0031] A control module is used to control the vehicle to play the audio segment in response to a playback control command.
[0032] Optionally, the acquisition module includes:
[0033] The first acquisition submodule is used to respond to the activation of the vehicle's in-vehicle music function, detect user input operations on the preset interactive interface, and acquire user input information.
[0034] The second acquisition submodule is used to monitor the current driving status of the vehicle and generate current driving information of the vehicle based on the driving status. The current driving information of the vehicle includes gear information, driving mode, steering wheel information, and pedal information.
[0035] Optionally, the determining module includes:
[0036] The determination submodule is used to determine the changes in the vehicle's driving actions by using the gear information, driving mode, steering wheel information, and pedal information in the vehicle's current driving information.
[0037] The third acquisition submodule is used to acquire road condition information and environmental information of the vehicle when the current driving action changes;
[0038] The fusion submodule is used to fuse the user input information, changes in driving actions, road condition information, and environmental information to generate target prompt words.
[0039] Optionally, the fusion submodule includes:
[0040] The generation unit is used to extract features from the user input information, changes in driving actions, road condition information, and environmental information to generate at least two keywords.
[0041] A fusion unit is used to perform feature fusion processing on at least two of the feature prompt words to generate a target prompt word; wherein the target prompt word is used to characterize music features.
[0042] Optionally, the generation module includes:
[0043] The recognition submodule is used to input the target prompt word into a preset data model and use the preset data model to recognize the target prompt word;
[0044] The output submodule is used to output the audio parameters corresponding to the target prompt word based on the correlation between natural language and music features in the preset data model.
[0045] A generation submodule is used to generate at least one audio segment using the audio parameters and store the audio segment in a playlist according to the generation time.
[0046] Optionally, the generation module further includes:
[0047] The monitoring submodule is used to monitor the number of audio segments in the playlist;
[0048] The update submodule is used to update the playlist by rotating and deleting the audio segments if the number of audio segments is greater than or equal to a preset storage threshold.
[0049] According to another aspect of the present invention, an electronic device is also provided, comprising:
[0050] processor;
[0051] Memory used to store the processor's executable instructions;
[0052] The processor is configured to execute the instructions to implement the in-vehicle music generation method described above.
[0053] According to another aspect of the present invention, a readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the in-vehicle music generation method as described above.
[0054] The in-vehicle music generation method provided in this invention, in response to the activation of the vehicle's in-vehicle music function, acquires user input information and the vehicle's current driving information. Based on the user input information and the current driving information, it determines target prompt words corresponding to music features, uses a preset data model to identify and process the target prompt words, determines audio parameters based on the target prompt words, generates at least one audio segment using the audio parameters, and controls the vehicle to play the audio segment in response to a playback control command. This invention obtains prompt words for audio generation by user active input or monitoring of vehicle driving information, and uses a preset data model to create and generate audio segments with a high degree of relevance to the user and the in-vehicle scenario based on the prompt words. This achieves automatic and intelligent generation of music content highly correlated with vehicle system information, providing users with a unique music experience in the in-vehicle scenario, further enhancing the uniqueness of in-vehicle music playback and meeting user needs.
[0055] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0056] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0057] Figure 1 This is a flowchart of the steps of an in-vehicle music generation method provided in an embodiment of the present invention;
[0058] Figure 2 yes Figure 1 A flowchart of step 101 in the in-vehicle music generation method provided in this embodiment of the invention;
[0059] Figure 3 yes Figure 1 A flowchart of step 102 in the in-vehicle music generation method provided in this embodiment of the invention;
[0060] Figure 4 yes Figure 1 A flowchart of step 103 in the in-vehicle music generation method provided in this embodiment of the invention;
[0061] Figure 5 This is a scenario illustration of the in-vehicle music generation method provided in this embodiment of the invention. Figure 1 ;
[0062] Figure 6 This is a scenario illustration of the in-vehicle music generation method provided in this embodiment of the invention. Figure 2 ;
[0063] Figure 7 This is a schematic diagram of the structure of an in-vehicle music generation device provided in an embodiment of the present invention;
[0064] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details are presented in the various embodiments of the present invention to facilitate a better understanding of this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for ease of description and should not constitute any limitation on the specific implementation of the present invention. The various embodiments can be combined with and referenced by each other without contradiction.
[0066] Reference Figure 1 The diagram illustrates a flowchart of the in-vehicle music generation method provided by an embodiment of the present invention. The method may include:
[0067] Step 101: In response to the activation of the vehicle's in-vehicle music function, obtain user input information and the vehicle's current driving information.
[0068] In this embodiment of the invention, to address the problem of low relevance between in-vehicle music and the in-vehicle environment, thus failing to meet users' music needs in in-vehicle scenarios, this embodiment automatically generates new and unique music content by considering user needs and the vehicle's driving status, achieving the effect of automatically generating music highly relevant to the in-vehicle environment. In this embodiment, the vehicle's in-vehicle entertainment system is equipped with intelligent in-vehicle audio software. The in-vehicle audio software is accessed via the central control screen or activated by the user using a voice assistant, thus activating the in-vehicle music function. When the in-vehicle music function is activated, it receives user input information or monitors the vehicle's current driving information to generate music content related to the user and the in-vehicle environment based on the user input information or current driving information.
[0069] Specifically, in response to the activation of the vehicle's in-vehicle music function, the system acquires user input information and current vehicle driving information. User input information includes initial prompts manually or selected by the user, user preferences, etc. The user can access the in-vehicle audio software through the vehicle's infotainment system. Users can download the in-vehicle audio software from the app store and open it in the application center on the central control screen, or they can open and access the software through a voice assistant. In response to the activation of the vehicle's in-vehicle music function, the vehicle's infotainment system triggers music content generation in at least two ways, primarily through user input and monitoring of current vehicle driving information. In some embodiments, the user can actively input or select one or more prompts through a simple and clear interactive method. The infotainment system can receive the input information entered by the user through the central control screen. Alternatively, after the vehicle's in-vehicle music function is activated, the system can monitor the vehicle's current driving information in real time. Specifically, it can monitor the vehicle's driving status through the bus to obtain current driving information, which will not be elaborated further here.
[0070] It should be noted that the in-vehicle audio software in this embodiment has AIGC (Artificial Intelligence Generated Content) functionality. AIGC is an artificial intelligence technology that can generate new content, such as images, text, or audio, based on input data. It uses large AI models as its foundation and intelligently generates music segments that are highly relevant to the text description and have good quality and completeness. In this embodiment, when a user enters the in-vehicle audio software, music segments can be intelligently generated through two triggering methods: receiving user input information or monitoring the vehicle's current driving information.
[0071] Step 102: Based on the user input information and current driving information, determine the target prompt word corresponding to the music feature.
[0072] In this embodiment of the invention, to improve the matching degree between the generated audio segments and the user and vehicle, the in-vehicle audio software uses the acquired user input information and current driving information to determine the target prompt words corresponding to the music features. Specifically, in response to the activation of the vehicle's in-vehicle music function, the software detects user input operations on a preset interactive interface, acquires user input information, monitors the vehicle's current driving status, and generates current vehicle driving information based on the driving status. This current driving information includes gear information, driving mode, steering wheel information, and pedal information. It should be noted that the in-vehicle audio software acquires information from various vehicle devices or applications via the CAN bus, including gear information from the vehicle's gear system, driving mode from the vehicle's chassis system, steering wheel information from the vehicle's steering wheel system, and pedal information.
[0073] Specifically, if in-vehicle music generation is triggered by user input, the in-vehicle audio software determines the target prompt word corresponding to the music feature based on the user input information. Specifically, the user can enter the interactive interface of the in-vehicle audio software, actively input or select an existing initial prompt word, and the in-vehicle audio software determines the target prompt word corresponding to the music feature based on at least one prompt word actively input or selected by the user. If in-vehicle music generation is triggered by monitoring the current driving information of the vehicle, the in-vehicle audio software obtains information from various devices or applications in the vehicle through the CAN bus, and determines the target prompt word corresponding to the music feature based on the current driving information.
[0074] For example, refer to Figure 5 This illustrates a scenario of the in-vehicle music generation method provided in an embodiment of the present invention. Figure 1 The in-vehicle audio software communicates with other devices and applications in the vehicle via the CAN bus, and can obtain information such as multimedia system, map application, weather, and time. Specifically, driving information can include gear position, vehicle speed, driving mode, steering wheel angle, pedal travel, etc. In addition, it can also obtain the vehicle's map application, time information, and weather information. Combined with user preference information in user input information, relevant data is collected through the CAN bus and cloud platform, and intelligently converted into target prompt words that can be used for AIGC to generate music content, thereby generating AI music content.
[0075] Step 103: Use a preset data model to identify the target prompt words, determine the audio parameters based on the target prompt words, and generate at least one audio segment using the audio parameters.
[0076] In this embodiment of the invention, after determining the target prompt word for generating a music segment, a preset data model is used to identify the target prompt word, determine audio parameters based on the target prompt word, and generate at least one audio segment using the audio parameters. The preset data model can be an AI large model, which is a complex artificial intelligence model built using a large number of parameters and deep learning technology. It can process and understand large-scale data and exhibit a higher level of intelligent behavior. In this embodiment, a preset data model is used to identify the target prompt word so as to automatically generate a music segment based on the prompt word.
[0077] Specifically, a preset data model is used to identify and process the target prompt words. Audio parameters are determined based on the target prompt words, and at least one audio segment is generated using the audio parameters. Specifically, the target prompt words are input into the preset data model, which then identifies and processes them. Based on the correlation between natural language and music features in the preset data model, the audio parameters corresponding to the target prompt words are output. At least one audio segment is generated using the audio parameters, and the audio segments are stored in the playlist according to the generation time.
[0078] It should be noted that audio parameters may include parameters such as tempo, scale, and rhythm. This embodiment does not specifically limit the duration of the generated audio segments. The number of audio segments to be generated can be preset, or at least one audio segment can be generated based on the correlation between the prompt words and musical features. This embodiment will not elaborate on these aspects.
[0079] Step 104: In response to the playback control command, control the vehicle to play an audio clip.
[0080] In this embodiment of the invention, the user can touch the play button on the central control screen or the human-machine interface of the in-vehicle audio software. Based on the user's touch operation, a playback control command is generated, and the in-vehicle audio software responds to the playback control command to control the playback of audio segments.
[0081] It should be noted that since at least one audio segment is generated, the playback order can be adjusted according to the user's selection. In this embodiment, the audio segment can be played using an in-vehicle multimedia device or an in-vehicle entertainment system. This embodiment does not specifically limit the device and method for playing the audio segment.
[0082] The in-vehicle music generation method provided in this invention, in response to the activation of the vehicle's in-vehicle music function, acquires user input information and the vehicle's current driving information. Based on the user input information and the current driving information, it determines target prompt words corresponding to music features, uses a preset data model to identify and process the target prompt words, determines audio parameters based on the target prompt words, generates at least one audio segment using the audio parameters, and controls the vehicle to play the audio segment in response to a playback control command. This invention obtains prompt words for audio generation by user active input or monitoring of vehicle driving information, and uses a preset data model to create and generate audio segments with a high degree of relevance to the user and the in-vehicle scenario based on the prompt words. This achieves automatic and intelligent generation of music content highly correlated with vehicle system information, providing users with a unique music experience in the in-vehicle scenario, further enhancing the uniqueness of in-vehicle music playback and meeting user needs.
[0083] Furthermore, refer to Figure 2 , showed Figure 1 The flowchart of step 101 in a method for generating in-vehicle music is provided. This method is basically the same as the method for generating in-vehicle music provided in the first embodiment of the present invention. Step 101, in response to the activation of the in-vehicle music function of the vehicle, specifically includes the following steps:
[0084] Step 201: In response to the activation of the vehicle's in-vehicle music function, detect user input operations on the preset interactive interface and obtain user input information.
[0085] In this embodiment of the invention, in response to the activation of the vehicle's in-vehicle music function, the system detects user input operations on a preset interactive interface, obtains user input information, and receives at least one initial prompt word input or selected by the user. The in-vehicle music function can be activated by voice or touch, and this embodiment does not specifically limit this.
[0086] Specifically, the system detects user input on a pre-defined interactive interface. Users can actively input prompts through a simple and clear interactive method. If the input prompt format is incorrect, the prompt format is converted, or prompts are provided in advance in the interactive interface so that users can actively select one or more prompts. The AI audio software intelligently generates multiple AI music contents that are highly related to the description based on the prompt phrases. Specifically, users select "Personalized Creation" on the main interface and enter the Personalized Creation interface. This interface should include an input box and music settings options. The input box or music settings options are used to determine the user's input information, which may include prompts, user preference information, etc.
[0087] Specifically, in the personalized creation interface, users can click the input box, which will bring up an input keyboard. Users can use keyboard characters or voice input to enter characteristic prompts for the AI music they wish to create. Each set of prompts is separated by a comma. The prompts can be related to music style, music scene, music form, and music emotion. For example, inputting "pop song, on the way home from get off work, recording studio, cheerful and relaxing" will allow the in-car audio software to receive the initial prompts and convert them into target prompts such as "pop music, home, recording studio effect, and relaxing melody". To ensure the effectiveness of the music content generated by the prompts, this embodiment can specify the number of target prompts. The number of target prompts can be preset or semantically expanded based on the initial prompts input by the user; no specific limitation is made here.
[0088] Step 202: Monitor the current driving status of the vehicle and generate current driving information based on the driving status. The current driving information includes gear information, driving mode, steering wheel information, and pedal information.
[0089] Specifically, in this embodiment, in response to the activation of the vehicle's in-vehicle music function, the current driving status of the vehicle is monitored, and the current driving information of the vehicle is generated based on the driving status. Specifically, the driving status of the vehicle is monitored through the bus, and the current driving information under the driving status is obtained. The current driving information of the vehicle includes gear information, driving mode, steering wheel information, and pedal information. Specifically, the vehicle driving mode information is determined through the chassis system, and the changes in driving actions are determined through the steering wheel and pedal change information.
[0090] For example, refer to Figure 6This illustrates a scenario of the in-vehicle music generation method provided in an embodiment of the present invention. Figure 2 Users can actively input or select one or more prompt words through a simple and clear interactive method. The AI audio software intelligently generates multiple AI music contents that are highly related to the description based on the prompt words. Specifically, users select "Personalized Creation" on the main interface. After selection, they enter the Personalized Creation interface, which should include: an input box, music setting options such as tempo, scale, rhythm, etc., a generate button, and a list display area to display the generated AI music content, namely the cover, title, audio, playback control, favorites, sharing, etc.
[0091] Specifically, users can click the input box on the personalized creation interface, which will bring up an input keyboard. Users can use keyboard characters or voice input to enter the characteristic prompts for the AI music they wish to create. The prompts can be related to music style, music scene, music form, and music emotion. Users can also click on music settings options, such as tempo, scale, and rhythm, for example, "tempo 80, C major, major, 4 / 4," to make more specific requirements for the structure of the generated AI music. Alternatively, users can generate audio parameters solely based on the prompts without setting any settings. After inputting and setting the parameters, clicking the "Generate" button will produce the AI music content, which will be displayed in a list with its cover, title, audio, playback controls, favorites, and sharing options.
[0092] It should be noted that the in-vehicle music function can be activated via voice or touch. The software can be opened and AI music generated through a voice assistant, or through user interaction with the multimedia system. In response to the activation of the in-vehicle music function, the system monitors the vehicle's driving status via the bus, acquiring current driving information. This information may specifically include determining the vehicle's current gear through the gear system, determining the driving mode through the chassis system, determining changes in driving actions through steering wheel and pedal movements, determining the vehicle's environment and road conditions through map applications, and collecting data and information such as time, weather, and user preferences through a cloud platform. Upon activation of the in-vehicle music function, the system receives at least one initial prompt word input or selection by the user, and information and control interactions can be performed through the multimedia system.
[0093] Furthermore, refer to Figure 3 , showed Figure 1 A flowchart of step 102 in a method for generating in-vehicle music is provided. This method is basically the same as the method for generating in-vehicle music provided in the first embodiment of the present invention. Step 102 may include:
[0094] Step 301: Use the gear information, driving mode, steering wheel information and pedal information in the current driving information of the vehicle to determine the changes in the vehicle's driving actions.
[0095] It should be noted that the in-vehicle audio software collects current driving information, using gear position information, driving mode, steering wheel information, and pedal information to determine changes in the vehicle's driving actions. For example, based on gear position information, it can be determined that the vehicle is currently traveling at low speed; based on driving mode, it can be determined whether the vehicle is in comfort mode, normal mode, sport mode, or eco mode; steering wheel information can determine the current steering angle and steering speed, thus determining the vehicle's steering action; and pedal information can determine the current acceleration and braking status, i.e., determining the vehicle's acceleration, deceleration, or stopping action. Therefore, based on gear position information, driving mode, steering wheel information, and pedal information, the changes in the vehicle's driving actions can be determined.
[0096] Step 302: Obtain road condition information and environmental information when the vehicle changes its current driving action.
[0097] It should be noted that, in order to generate audio segments that more accurately reflect the vehicle's driving status and user needs, this embodiment integrates environmental information and vehicle road condition information to increase the matching degree of audio segments. Specifically, it obtains road condition information and environmental information when the vehicle changes its current driving action. The road condition information is specifically collected through map applications, and the environmental information includes weather, temperature, and other information, which can be obtained through vehicle networking applications. These will be described in detail here.
[0098] Step 303: The user input information, changes in driving actions, road condition information, and environmental information are fused and processed to generate target prompt words.
[0099] In this embodiment of the invention, the in-vehicle audio software integrates information such as vehicle gear, speed, driving mode, steering wheel angle, pedal travel, map application, time information, weather information, and user preference information based on the acquired current driving information, user input information, road condition information, and environmental information. The software converts the current driving information into target prompt words corresponding to music features, such as "driving, in the city, low speed, comfort mode, frequent turning, smooth driving". The target prompt words are used to characterize the music attributes of the audio segment.
[0100] In this embodiment, the target prompt word corresponding to the music feature is obtained by processing at least one initial prompt word input or selected by the user or the current driving information obtained, thereby improving the effectiveness of the prompt words used to generate music content and enabling the generation of music content that is highly related to vehicle system information.
[0101] Specifically, step 303, which involves fusing user input information, changes in driving actions, road condition information, and environmental information to generate target prompt words, may include the following steps: extracting features from user input information, changes in driving actions, road condition information, and environmental information to generate at least two keywords; fusing the features of the at least two feature prompt words to generate target prompt words; wherein, the target prompt words are used to represent music features.
[0102] It should be noted that feature extraction is the process of extracting key information from the original data for subsequent processing and analysis. To improve the effectiveness of audio file generation, this embodiment generates at least two keywords to facilitate feature fusion. For example, based on changes in driving actions, features such as acceleration, deceleration, turning, and stopping are extracted to obtain the current state of the vehicle and the driver's driving behavior. The changes in driving actions and road condition information are then fused to generate the target prompt word "safe driving". Alternatively, user input information and environmental information can be fused to generate the target prompt word "driving in the city".
[0103] In this embodiment, by performing feature extraction and feature fusion processing on driving information, key information is extracted from complex raw data and target prompt words are generated to determine user needs and driving conditions. Based on the prompt words, audio segments with a high degree of relevance to the user and the in-vehicle scenario can be created, realizing the automatic and intelligent generation of music content that is highly related to vehicle system information, and providing users with a unique music experience in the in-vehicle scenario.
[0104] Furthermore, refer to Figure 4 , showed Figure 1 A flowchart of step 103 in a method for generating in-vehicle music is provided. This method is basically the same as the method for generating in-vehicle music provided in the first embodiment of the present invention. Step 103 may include:
[0105] Step 401: Input the target prompt word into the preset data model and use the preset data model to identify and process the target prompt word.
[0106] It should be noted that the preset data model in this embodiment can be pre-trained by collecting a large number of prompt words, and the text content of the target prompt words can be recognized by a deep learning network. This embodiment does not specifically limit the training process of the preset data model.
[0107] Step 402: Based on the correlation between natural language and music features in the preset data model, output the audio parameters corresponding to the target prompt word.
[0108] Specifically, a data model was pre-trained to understand the relationship between natural language and musical features. In this embodiment, the AI-driven model learns to associate specific music loops with text prompts, outputting audio parameters corresponding to the target prompt, thereby creating music corresponding to the input text. These audio parameters mainly include BPM, tempo, scale, and rhythm. BPM, or beats per minute, can be determined based on the content of the prompt. For example, if the prompt is "driving, city, low speed, comfort mode," the corresponding output audio parameters would be "tempo 80, C major, major, 4 / 4".
[0109] Step 403: Generate at least one audio segment using audio parameters, and store the audio segment in the playlist according to the generation time.
[0110] In this embodiment, by inputting the target prompt word into a preset data model, outputting the audio parameters corresponding to the target prompt word, and using the audio parameters to generate at least one audio segment, and storing the audio segment in a playlist, it is possible to create audio segments with a high degree of relevance to the user and the in-vehicle scenario based on the prompt word. This enables the automatic and intelligent generation of music content that is highly related to vehicle system information, providing users with a unique music experience in the in-vehicle scenario. Users can easily create unique music content that is highly relevant to the in-vehicle scenario, avoiding the serious problem of homogenization of music experience in the in-vehicle scenario.
[0111] Specifically, after generating at least one audio segment using the audio parameters and storing the audio segment in the playlist according to the generation time, the method further includes:
[0112] Monitor the number of audio segments in the playlist;
[0113] If the number of audio segments is greater than or equal to a preset storage threshold, the audio segments are deleted in rotation, and the playlist is updated.
[0114] In this embodiment, to ensure smooth music playback and the speed of music clip generation, the number of audio clips in the playlist is monitored. If the number of audio clips is greater than or equal to a preset storage threshold, the audio clips are rotated and deleted, and the playlist is updated. Specifically, the generated audio clips are placed in the historical playback record. The number of AI music contents saved in the historical record can be set. When the set number is exceeded, the oldest music content will be automatically deleted when new AI music is generated, completing the rotation. This avoids wasting storage space by storing too much worthless content in the software, ensuring smooth music playback and the speed of music clip generation.
[0115] In some embodiments, to enhance the enjoyment of in-car music, after input and settings are complete, the user clicks the "Generate" button. After a period of time, multiple AI music clips are generated, displayed in a list with their cover art, title, audio, playback controls, favorites, and sharing options. The user can click the play button to play the corresponding audio, drag the audio progress bar, click the favorite button to add the corresponding audio clip to their playlist, and click the share button to synchronize the target audio clip to the cloud platform for sharing on online communities. By having the user identify the target audio clip and synchronize it to the cloud platform, signals and data that were originally valueless can be transformed into information that can be recognized and used by software. This produces music content highly relevant to the in-car environment, providing a brand-new experience for drivers and passengers. Furthermore, the topicality and creativity inherent in AIGC (AI-generated content) can effectively stimulate communication and discussion among users, increasing long-term operational activity.
[0116] Reference Figure 7 The diagram shows a structural schematic of an in-vehicle music generation device 500 provided in an embodiment of the present invention. The device includes:
[0117] The acquisition module 501 is used to acquire user input information and current vehicle driving information in response to the activation of the vehicle's in-vehicle music function;
[0118] The determining module 502 is used to determine the target prompt word corresponding to the music feature based on the user input information and the current driving information;
[0119] The generation module 503 is used to identify the target prompt word using a preset data model, determine audio parameters based on the target prompt word, and generate at least one audio segment using the audio parameters.
[0120] The control module 504 is used to control the vehicle to play the audio segment in response to a playback control command.
[0121] Optionally, the acquisition module 501 includes:
[0122] The first acquisition submodule is used to respond to the activation of the vehicle's in-vehicle music function, detect user input operations on the preset interactive interface, and acquire user input information.
[0123] The second acquisition submodule is used to monitor the current driving status of the vehicle and generate current driving information of the vehicle based on the driving status. The current driving information of the vehicle includes gear information, driving mode, steering wheel information, and pedal information.
[0124] Optionally, the determining module 502 includes:
[0125] The determination submodule is used to determine the changes in the vehicle's driving actions by using the gear information, driving mode, steering wheel information, and pedal information in the vehicle's current driving information.
[0126] The third acquisition submodule is used to acquire road condition information and environmental information of the vehicle when the current driving action changes;
[0127] The fusion submodule is used to fuse the user input information, changes in driving actions, road condition information, and environmental information to generate target prompt words.
[0128] Optionally, the fusion submodule includes:
[0129] The generation unit is used to extract features from the user input information, changes in driving actions, road condition information, and environmental information to generate at least two keywords.
[0130] A fusion unit is used to perform feature fusion processing on at least two of the feature prompt words to generate a target prompt word; wherein the target prompt word is used to characterize music features.
[0131] Optionally, the generation module 503 includes:
[0132] The recognition submodule is used to input the target prompt word into a preset data model and use the preset data model to recognize the target prompt word;
[0133] The output submodule is used to output the audio parameters corresponding to the target prompt word based on the correlation between natural language and music features in the preset data model.
[0134] A generation submodule is used to generate at least one audio segment using the audio parameters and store the audio segment in a playlist according to the generation time.
[0135] Optionally, the generation module 503 further includes:
[0136] The monitoring submodule is used to monitor the number of audio segments in the playlist;
[0137] The update submodule is used to update the playlist by rotating and deleting the audio segments if the number of audio segments is greater than or equal to a preset storage threshold.
[0138] The in-vehicle music generation method provided in this invention, in response to the activation of the vehicle's in-vehicle music function, acquires user input information and the vehicle's current driving information. Based on the user input information and the current driving information, it determines target prompt words corresponding to music features, uses a preset data model to identify and process the target prompt words, determines audio parameters based on the target prompt words, generates at least one audio segment using the audio parameters, and controls the vehicle to play the audio segment in response to a playback control command. This invention obtains prompt words for audio generation by user active input or monitoring of vehicle driving information, and uses a preset data model to create and generate audio segments with a high degree of relevance to the user and the in-vehicle scenario based on the prompt words. This achieves automatic and intelligent generation of music content highly correlated with vehicle system information, providing users with a unique music experience in the in-vehicle scenario, further enhancing the uniqueness of in-vehicle music playback and meeting user needs.
[0139] Reference Figure 8 The present invention also provides an electronic device, such as... Figure 8 As shown, it includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.
[0140] Memory 603 is used to store computer programs;
[0141] When processor 601 executes a program stored in memory 603, it performs the following steps:
[0142] In response to the activation of the vehicle's in-car music function, it obtains user input information and the vehicle's current driving information;
[0143] Based on the user input information and the current driving information, the target prompt words corresponding to the music features are determined;
[0144] The target prompt word is identified using a preset data model, audio parameters are determined based on the target prompt word, and at least one audio segment is generated using the audio parameters.
[0145] In response to a playback control command, the vehicle is controlled to play the audio clip.
[0146] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0147] The communication interface is used for communication between the aforementioned terminal and other devices.
[0148] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0149] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0150] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform any of the in-vehicle music generation methods described in the above embodiments.
[0151] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0152] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0153] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0154] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A method for generating in-vehicle music, characterized in that, The method includes: In response to the activation of the vehicle's in-car music function, it obtains user input information and the vehicle's current driving information; Based on the user input information and the current driving information, the target prompt words corresponding to the music features are determined; The target prompt word is identified using a preset data model, audio parameters are determined based on the target prompt word, and at least one audio segment is generated using the audio parameters. In response to a playback control command, the vehicle is controlled to play the audio clip.
2. The method according to claim 1, characterized in that, The response to the activation of the vehicle's in-vehicle music function includes obtaining user input information and the vehicle's current driving information, including: In response to the activation of the vehicle's in-vehicle music function, the system detects user input operations on the preset interactive interface and obtains user input information. The vehicle's current driving status is monitored, and based on the driving status, current driving information is generated, including gear information, driving mode, steering wheel information, and pedal information.
3. The method according to claim 1, characterized in that, The step of determining the target prompt word corresponding to the music feature based on the user input information and the current driving information includes: The changes in the vehicle's driving actions are determined by using the gear information, driving mode, steering wheel information, and pedal information from the vehicle's current driving information. Acquire road condition and environmental information when the vehicle's current driving action changes; The user input information, changes in driving actions, road condition information, and environmental information are fused and processed to generate target prompt words.
4. The method according to claim 3, characterized in that, The process of fusing user input information, changes in driving actions, road condition information, and environmental information to generate target prompt words includes: The user input information, changes in driving actions, road condition information, and environmental information are used to extract features and generate at least two keywords. At least two of the aforementioned feature prompt words are subjected to feature fusion processing to generate target prompt words; wherein, the target prompt words are used to characterize music features.
5. The method according to claim 1, characterized in that, The process of identifying the target prompt word using a preset data model, determining audio parameters based on the target prompt word, and generating at least one audio segment using the audio parameters includes: The target prompt word is input into a preset data model, and the target prompt word is identified and processed using the preset data model; Based on the correlation between natural language and music features in the preset data model, the audio parameters corresponding to the target prompt word are output; At least one audio segment is generated using the audio parameters, and the audio segment is stored in the playlist according to the generation time.
6. The method according to claim 5, characterized in that, After generating at least one audio segment using the audio parameters and storing the audio segment in the playlist according to the generation time, the method further includes: Monitor the number of audio segments in the playlist; If the number of audio segments is greater than or equal to a preset storage threshold, the audio segments are deleted in rotation, and the playlist is updated.
7. A vehicle-mounted music generation device, characterized in that, The device includes: The acquisition module is used to acquire user input information and the vehicle's current driving information in response to the activation of the vehicle's in-vehicle music function. The determination module is used to determine the target prompt word corresponding to the music feature based on the user input information and the current driving information; The generation module is used to identify the target prompt word using a preset data model, determine audio parameters based on the target prompt word, and generate at least one audio segment using the audio parameters. A control module is used to control the vehicle to play the audio segment in response to a playback control command.
8. The apparatus according to claim 7, characterized in that, The acquisition module includes: The first acquisition submodule is used to respond to the activation of the vehicle's in-vehicle music function, detect user input operations on the preset interactive interface, and acquire user input information. The second acquisition submodule is used to monitor the current driving status of the vehicle and generate current driving information of the vehicle based on the driving status. The current driving information of the vehicle includes gear information, driving mode, steering wheel information, and pedal information.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the instructions to implement the in-vehicle music generation method as described in any one of claims 1 to 6.
10. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the in-vehicle music generation method as described in any one of claims 1 to 6.