Vehicle-mounted music recommendation method and device based on large language model and storage medium

By adopting a method based on a large language model in the car music recommendation system and combining multiple recommendation factors to generate music recommendations, the problem that the existing system cannot meet the music needs of driving scenes that are dynamically changed in real time is solved, and high-precision and personalized music recommendations are achieved, which improves the driving experience.

CN120086407APending Publication Date: 2025-06-03AISPEECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510134798.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing car music recommendation system cannot meet the music needs of real-time dynamically changing driving scenarios, and lacks comprehensive consideration of driving situation factors, resulting in insufficient recommendation results.

Method used

Using a car music recommendation method based on a large language model, by obtaining multiple recommendation factors, including driver music preference information, vehicle trip information, off-vehicle environment information and driver status information, fill it into a preset music recommendation prompt template, build a music recommendation prompt word, and enter a large language model to determine the music recommendation recommendation text, and finally generate a recommended music list and send it to the car terminal.

Benefits of technology

It realizes accurate response to the music needs of real-time dynamically changing driving scenes, improves the intelligence and dynamic level of music recommendations, provides a personalized and contextual music experience, and enhances driving comfort and pleasure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086407A_ABST
    Figure CN120086407A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle-mounted music recommendation method and device based on a large language model and a storage medium, and the method comprises the steps: obtaining a plurality of recommendation factors which comprise driver music preference information, and vehicle travel information, vehicle exterior environment information and driver state information which are collected in real time; filling each recommendation factor into a preset music recommendation prompt template to construct a music recommendation prompt word; inputting the music recommendation cue words into a large language model to determine a corresponding music recommendation suggestion text; and generating a recommended music list according to the music recommendation suggestion text, and sending the recommended music list to the vehicle-mounted terminal, so that the vehicle-mounted terminal can request music resources in the recommended music list from a music service provider. Therefore, music recommendation based on context awareness is realized, and the comfort and pleasure of driving can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice technology, and particularly to a vehicle-mounted music recommendation method, device, storage medium, and program product based on a large language model. Background Art

[0002] With the continuous development of artificial intelligence technology, in-vehicle information systems have become an important part of modern vehicles. Among them, the music recommendation system, as an important function in the in-vehicle information system, can provide personalized music experiences for drivers, improving driving comfort and safety.

[0003] Currently, in-vehicle music recommendation systems mainly rely on the music recommendation modules in the music streaming software (such as QQ Music, Apple Music, etc.) installed on the in-vehicle head unit. By analyzing static data such as the historical music play records and preference settings of software users, personalized music recommendations are provided using preset music recommendation algorithms. Although such music recommendation methods can meet the music needs of drivers to a certain extent, they lack comprehensive consideration of driving context factors, resulting in inaccurate recommendation results and difficulty in meeting the music needs of real-time and dynamically changing driving scenarios.

[0004] In response to the above problems, no better solutions have been proposed in the industry yet. Summary of the Invention

[0005] This application provides a vehicle-mounted music recommendation method, device, storage medium, and program product based on a large language model, which is used to at least solve the problem that the in-vehicle music recommendation method in the current related technology cannot meet the music needs of real-time and dynamically changing driving scenarios.

[0006] In a first aspect, an embodiment of this application provides a vehicle-mounted music recommendation method based on a large language model, including: obtaining a plurality of recommendation factors, where the plurality of recommendation factors include driver music preference information and vehicle travel information, external vehicle environment information, and driver status information collected in real time; filling each of the recommendation factors into a preset music recommendation prompt template to construct a music recommendation prompt; inputting the music recommendation prompt into the large language model to determine a corresponding music recommendation suggestion text; generating a recommended music list according to the music recommendation suggestion text, and sending the recommended music list to an in-vehicle terminal, so that the in-vehicle terminal can request music resources in the recommended music list from a music service provider.

[0007] In a second aspect, an embodiment of the present application provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the in-vehicle music recommendation method based on a large language model according to any embodiment of the present application.

[0008] In a third aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored, and characterized in that when the program is executed by a processor, it implements the steps of the in-vehicle music recommendation method based on a large language model according to any embodiment of the present application.

[0009] In a fourth aspect, an embodiment of the present application provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, they implement the steps of the in-vehicle music recommendation method based on a large language model according to any embodiment of the present application.

[0010] The beneficial effects of the embodiments of the present application are as follows: By integrating multi-dimensional in-vehicle real-time monitoring data and driver music preference information, and utilizing the powerful natural language processing ability of the large language model to comprehensively analyze various real-time recommendation factors, it is ensured that the music recommendation results can understand and respond to the music needs of the real-time dynamic driving scenarios in real time, realizing context-aware music recommendation and enhancing the comfort and pleasure of driving. In addition, by introducing a server middle layer between the in-vehicle terminal and the music service provider, the limitations of the music service provider's recommendation algorithm are fundamentally broken through. Independent of the music service provider's recommendation system, it can perform more accurate and dynamic personalized recommendations according to factors such as the driver's immediate needs and driving scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 Shows a flowchart of an example of the in-vehicle music recommendation method based on a large language model according to an embodiment of the present application; Figure 2 Shows an operation flow schematic diagram of an example of a multi-model recognition module deployed on a vehicle terminal according to an embodiment of the present application; Figure 3Shows an operation flowchart of an example of generating a recommended music list according to music recommendation suggestion text according to an embodiment of the present application; Figure 4 Shows an operation flowchart of an example of an in-vehicle music recommendation method based on a large language model according to an embodiment of the present application; Figure 5 Shows according to Figure 4 An effect schematic diagram of an example of steps ⑤ and ⑥ in Figure 6 Is a structural schematic diagram of an embodiment of an electronic device of the present application. Detailed implementation manners

[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0014] It should be noted that the in-vehicle music recommendation function in the current related technologies mainly relies on the recommendation function of the in-vehicle music streaming software, which performs personalized music recommendations by analyzing static data such as the historical music play records and preference settings of software users. However, on the one hand, the driver may have a different identity from that of the software user of the music streaming software. For example, the driver is a shared user of the software account, resulting in a large deviation in music recommendations. On the other hand, the vehicle driving journey may involve diverse scenarios, such as emotional scenarios, weather scenarios, etc. The types of music preferred by the driver may vary in different scenarios. Exemplarily, in extreme emotional scenarios, such as long-term driving, traffic congestion, etc., the weight of emotional factors should be increased in music recommendations, and effective emotional management should be provided to the driver. In other scenarios, the driver's music preferences can be considered first.

[0015] Therefore, the existing systems mainly rely on static recommendation weight parameters for music recommendations, lacking dynamic analysis and real-time response to various dimensional recommendation factors in a comprehensive scenario, resulting in inaccurate music recommendation effects.

[0016] It should be understood that the purpose of the above description of the current related technologies is only to facilitate the public's better understanding of the inventive spirit and motivation of the present application, and is not regarded as a limitation to the present application. In addition, the technical solutions described in the above current related technologies are not prior art, and they may also be unpublished technical solutions, such as those under research or in the laboratory stage.

[0017] In view of this,Figure 1 The figure shows a flowchart of an example of a vehicle-mounted music recommendation method based on a large language model according to an embodiment of the present application. The implementation subject of this method is a server or a vehicle cloud platform.

[0018] As Figure 1 shown, in step S110, a plurality of recommendation factors are obtained. The plurality of recommendation factors include driver music preference information and vehicle travel information, vehicle exterior environment information, and driver status information collected in real time.

[0019] Here, the driver music preference information may be information such as songs frequently played on the in-vehicle device, preferred music types (such as pop, rock, electronic, etc.), and music preferences set according to personal settings (such as singer preferences, play duration, etc.). It can be obtained through interface interaction with in-vehicle terminal devices or music service providers. In some examples, the driver music preference information includes music attribute preference information (such as music style, etc.) and singer preference information.

[0020] The vehicle travel information can be obtained through the vehicle's positioning information or navigation system, and it may include travel departure time (such as morning, morning, afternoon, evening), travel type (such as going to work, going home, traveling, business trips, etc.), and travel road conditions (such as smooth, slightly congested, moderately congested, severely congested, etc.), current location, speed, travel distance, driving mode (such as urban, highway, etc.). Through these data, it is helpful to judge the vehicle's real-time driving situation.

[0021] The vehicle exterior environment information can be used to obtain data on the current external environment through in-vehicle sensors (such as weather modules, external temperature and humidity sensors, etc.), such as the vehicle exterior weather state (sunny, rainy, foggy, etc.), temperature, humidity, etc. Different external environmental factors may affect the driver's mood and music preferences. For example, on rainy days, the driver may be more inclined to soothing music, while on sunny days, the driver may prefer energetic music.

[0022] The driver status information can be judged through biometric technologies (such as facial expression recognition, heart rate monitors, eye movement tracking, etc.) or data collected by in-vehicle devices (such as the driver's fatigue level, anxiety level, attention concentration, etc.). In some examples, the driver status information includes the driver's emotional state, such as calm, happy, angry, and sad, etc. In the case of fatigue driving, the system may recommend more relaxing and soothing music, while when the driver is excited or anxious, the system may recommend soothing and relaxing music.

[0023] In step S120, each recommendation factor is filled into a preset music recommendation prompt template to construct a music recommendation prompt word.

[0024] Here, for the convenience of processing in large language models, it is necessary to design a unified music recommendation prompt template, which reserves placeholders for each recommendation factor and indicates the corresponding recommendation factors to be filled. In this way, the recommendation factors of complex, multi-modal data sources are transformed into structured and unified prompt word inputs, providing standardized text for subsequent natural language processing (such as large language models).

[0025] In step S130, the music recommendation prompt words are input into the large language model to determine the corresponding music recommendation suggestion text.

[0026] It should be understood that the types of large language models can be diverse, such as open or closed large language models, and various stable and mature large models can be adopted, such as the GPT series, qwen, etc., which are not limited here for the time being. With the enhanced natural language processing ability of the large language model, it can understand the context and generate music recommendation suggestion text that conforms to the driving situation and driver preferences.

[0027] In an example of the embodiment of the present application, the recommendation suggestion text may include a specific sequence of song names, so that music recommendations can be directly made based on these song names. In another example of the embodiment of the present application, the recommendation suggestion text may also include various key information of the recommended music, such as the recommended music category (such as jazz, rock, etc.), the emotional tone of the song (such as calm, inspiring, etc.), and the suitable song playing order, providing a basis for the generation of subsequent recommended playlists.

[0028] Thus, by generating music recommendation text through the large language model, complex multi-factor associations and context analysis can be achieved, rather than just rule-based recommendations, improving the intelligence and dynamics of the recommendations and making each recommendation more in line with the driver's immediate needs.

[0029] In step S140, a recommended music list is generated according to the music recommendation suggestion text, and the recommended music list is sent to the in-vehicle terminal, so that the in-vehicle terminal can request music resources in the recommended music list from the music service provider.

[0030] In an example of the embodiment of the present application, when there is a sequence of recommended song names in the music recommendation suggestion text, the sequence of song names can be directly formatted and encapsulated to generate a recommended music list in a list format that conforms to the music service provider. In another example of the embodiment of the present application, when the music recommendation suggestion text contains various key information of the recommended music, a query is made in the preset music library, and then the corresponding recommended music list is constructed. Further, the in-vehicle terminal can call one or more service provider interfaces according to the recommended music list to request music resources from one or more music service providers.

[0031] Through the embodiments of the present application, multi-factor real-time data is introduced, music recommendation suggestions are generated using large language models, and personalized music recommendations are provided by making interface resource requests between the in-vehicle terminal and music service providers, improving the intelligence, dynamics, and personalization levels of music recommendations.

[0032] In some examples of the embodiments of the present application, the in-vehicle terminal is configured with multiple service provider interfaces, and each service provider interface is respectively used to access the corresponding music service provider. Specifically, after the in-vehicle terminal receives the recommended music list, the in-vehicle terminal calls each service provider interface in the order of the recommended music list to automatically schedule and request music resources from the corresponding music service providers.

[0033] In some embodiments, the call order of each service provider interface is preset in the in-vehicle terminal, so that the music resources of the corresponding music service providers are accessed and requested step by step through the call order. In some cases, if the first service provider interface fails to return the required resources (for example, the service provider interface is temporarily unavailable, or a certain song is not provided on this platform), the next interface, that is, the second service provider interface, can be called according to the call order to implement a fault tolerance mechanism, ensuring that as many recommended music resources as possible are obtained, and also providing a guarantee for the stable operation of the in-vehicle music recommendation system.

[0034] Through the embodiments of the present application, based on the multiple music service provider interfaces configured in the in-vehicle terminal, music resources from different platforms can be automatically obtained, and music resources are requested from the corresponding service providers according to the generated recommendation results, so that the driver is not limited to a single music platform and can enjoy music resources from multiple platforms, realizing a cross-service provider music recommendation mechanism and greatly optimizing the breadth of music recommendations.

[0035] Regarding the implementation details of obtaining multiple recommendation factors in the above step S110, in some embodiments, the server receives multiple recommendation factors from the in-vehicle terminal. In this way, an architecture design combining server cloud and in-vehicle terminal local computing is realized, enabling efficient allocation of computing resources. Complex reasoning tasks are handled by the cloud, while the in-vehicle system is responsible for real-time data collection and sensitive data processing (such as avoiding the leakage of driver information and vehicle information). In this way, through a distributed architecture, not only the computing pressure on in-vehicle hardware is reduced, but also the accuracy of recommendations can be ensured while protecting user privacy.

[0036] In some examples of the embodiments of the present application, the driver status information includes the driver's emotional state, and the out-of-vehicle environment information includes the out-of-vehicle weather state.

[0037] Figure 2 Shows a schematic diagram of the operation process of an example of a multi-model recognition module deployed on a vehicle terminal according to an embodiment of the present application.

[0038] As shown Figure 2 in the figure, the vehicle-mounted terminal is configured with a weather condition recognition model and a facial expression recognition model. The weather condition recognition model is used to recognize the external weather condition corresponding to the image captured by the external camera in real time, and the facial expression recognition model is used to recognize the driver's emotional state corresponding to the image captured by the internal camera in real time.

[0039] Specifically, by introducing supervised fine-tuning in the training stage, the weather condition recognition model and the facial expression recognition model can learn the data characteristics in the vehicle-mounted scenario (such as image features under different weather and different driving conditions), thereby improving the recognition accuracy of the model. The external environment image of the vehicle is collected in real time through the external camera, covering the weather condition information around the vehicle (such as light, visibility, rain and snow, etc.). The collected image will be used as input data and transmitted to the weather condition recognition model. In addition, the facial expression of the driver is monitored in real time through the internal camera. By capturing the expression features (such as eyes, mouth, eyebrows, etc.), the image information of the driver's emotional state is provided, and these data are transmitted to the facial expression recognition model for emotion analysis.

[0040] It should be understood that the model types of the weather condition recognition model and the facial expression recognition model can be diversified, such as convolutional neural network, model based on Vision Transformer, etc., and no restrictions are made here for the time being.

[0041] Through the embodiments of the present application, the weather condition recognition model and the facial expression recognition model are integrated into the vehicle-mounted terminal, and combined with real-time image acquisition and deep learning inference technology, it is possible to realize a comprehensive perception of the external driving environment and the driver's internal state, provide multi-dimensional context information for the vehicle-mounted system, and support more intelligent and accurate recommendation and auxiliary decision-making.

[0042] In some examples of the embodiments of the present application, the driver's music preference information is determined by the vehicle-mounted terminal according to the driver user profile and the vehicle-mounted music play record.

[0043] Specifically, the driver user profile includes user identification, driving habit information and personalized tags. The vehicle-mounted terminal can determine the driver's identity through the driver's personal account (such as owner registration information, mobile App account synchronization) or the driver's fingerprint / face recognition information, and then bind the personal profile. Driving habits related to music preferences are constructed based on the driving history (such as driving duration, driving route, common driving scenarios). For example, users with short-distance driving or daily commuting may prefer fast-paced music. The personalized tags can be generated by interacting with the driver (such as the in-vehicle screen or voice assistant), actively asking about the user's preferences, and generating preliminary personalized tags (such as music type preferences: pop, rock; or specific singer preferences).

[0044] Furthermore, based on the in-vehicle music play records, the attribute information of each piece of music or frequently played music can be analyzed in detail. For example, multi-dimensional music features such as music style, music rhythm, music emotion, singer information, etc. can be analyzed, and combined with the driver user profile to determine the driver's music preference information.

[0045] Exemplarily, during short-distance driving, if the frequently played music belongs to the pop style, it is marked as "Short-distance driving preference style: Pop"; during long-distance driving, if the user tends to play light music or classical music, it is marked as "Long-distance driving preference style: Light music / Classical". In addition, if the user prefers to play fast-paced music (such as BPM>120) during daily commuting, it is marked as "Commuting scenario rhythm preference: Fast rhythm"; if the user tends to play slow-paced music (such as BPM<80) during night driving, it is marked as "Night driving rhythm preference: Slow rhythm". Thus, the analysis method based on multi-dimensional features can significantly improve the matching degree between the recommended music and the driving situation, and provide a more suitable music experience for the driver according to the current needs.

[0046] Figure 3 FIG. shows an operation flowchart of an example of generating a recommended music list according to a music recommendation advice text according to an embodiment of the present application.

[0047] As Figure 3 shown, in step S310, a list of music names is screened from the music recommendation advice text.

[0048] In some embodiments, the music recommendation advice text is parsed by natural language processing (NLP) technology to extract the music names included in the text. Exemplarily, keywords related to music (such as song names, singer names, etc.) can be identified through keyword recognition, or the text can be segmented and semantically analyzed through syntactic analysis to ensure the separation of music names from other non-music-related information (such as scene descriptions), etc.

[0049] In step S320, a music search engine authorized by a music service provider is called to determine the song resource index information corresponding to each music name in the list of music names.

[0050] In some embodiments, one or more music search engines authorized by music service providers are configured in the server, and each music name in the list of music names is retrieved through these music search engines to obtain the corresponding song resource index information.

[0051] It should be noted that the song resource index information is used to locate the position of the corresponding song resource in the music service provider, which may include the unique identifier of the song, the storage address or access path of the song file, the copyright information of the song (whether it can be played), etc.

[0052] In some examples of the embodiments of the present application, based on the music search engine, it is retrieved whether the first music name belongs to an accessible song. Furthermore, in the case where it is determined that the first music name belongs to an accessible song, the song resource index information corresponding to the first music name is obtained.

[0053] On the other hand, in the case where it is retrieved that the first music name is an inaccessible song, for example, the copyright is not licensed or membership privileges are required, the first music name can be deleted to avoid invalid music resources appearing in the recommended music list, thereby ensuring the executability of the music recommendation result.

[0054] Preferably, multiple music search engines are set in the server, which respectively correspond to different music service providers. If it is retrieved through the first music search engine that the first music name is an inaccessible song in the song library of the first music service provider, the second music search engine is sequentially called to retrieve in the song library of the second music service provider, so as to realize the cross-library retrieval and integration of music resources of different music service providers.

[0055] In step S330, a recommended music list is generated according to each song resource index information.

[0056] In some embodiments, all the determined song service index information is sorted into a recommended music list, and the corresponding song description information is included, such as the song name, the singer name, the music service provider to which it belongs, the service provider song ID or access path, etc. In addition, if the recommended music list contains songs from different service providers, resource integration can also be performed in the music list so that the in-vehicle terminal can seamlessly call the interfaces of each service provider to obtain resources.

[0057] Through the embodiments of the present application, by extracting music names from the text and pre-using the music search engine for retrieval and matching, it can be ensured that the recommended music list covers the actually accessible song resources, and it can also integrate the music resources of different platforms, which can effectively address the problem of resource limitations.

[0058] Figure 4 Shows an operation flowchart of an example of the in-vehicle music recommendation method based on a large language model according to an embodiment of the present application.

[0059] As Figure 4As shown in the figure, by comprehensively analyzing the driver's music preferences, trip information, vehicle interior and exterior environment, and the driver's emotional state, intelligent music recommendations are realized to enhance the driving experience and emotional management. The system dynamically generates music recommendation prompts, uses large language models to generate appropriate music recommendation lists, and then provides personalized and contextual music experiences for drivers.

[0060] It should be noted that existing in-vehicle music recommendation systems mainly rely on static recommendation weight parameters for music recommendations, lacking dynamic analysis and real-time response to various dimensions of recommendation factors in comprehensive scenarios, and the recommendation effect is not precise enough.

[0061] In view of this, in the embodiments of this application, the scenario perception and context understanding capabilities of large language models are used to optimize the recommendation effect of the recommendation system under multi-dimensional features. The large model can analyze and understand more complex scenarios (such as "the driver is in a slightly anxious state and in a congested section"), and based on this, prioritize the incoming recommendation factors (such as the driver's preferences, emotions, weather, etc.). For example, during long drives, the system may consider the rhythm and style of music more to prevent driver fatigue, while during peak traffic, more consideration is given to emotion regulation.

[0062] The specific steps are as follows: Step ①: The environmental data collection module receives instructions and collects vehicle exterior environment and in-vehicle driver expression data through cameras inside and outside the vehicle.

[0063] Step ②: The edge multi-model recognition module (as shown in the figure) processes the vehicle exterior environment and expression data, and completes weather recognition and user emotion recognition respectively through two edge image recognition models. Figure 2 As shown in the figure

[0064] Step ③: The vehicle control terminal combines the driver's music preferences obtained by integrating the initialized user profile and historical feedback records, including music attribute preferences (such as music style) and artist preferences (such as singers), and the trip information obtained through system built-in components (settings, navigation), including departure time (morning, morning, afternoon, evening), travel type (going to work, going home, traveling, on business, etc.), overall road conditions (smooth, slightly congested, moderately congested, severely congested, etc.), and integrates them with the weather and user emotions as recommendation factors.

[0065] Step ④: Upload all the recommendation factor data to the cloud service system.

[0066] Step ⑤: The prompt constructor uses the recommendation factors to generate music recommendation prompts.

[0067] Step ⑥: The music recommendation prompt is input into the large language model to obtain the music recommendation suggestion text.

[0068] Step ⑦: The playlist constructor parses the music recommendation advice text to obtain a song list, and confirms the id information of the correct songs through the music search engine provided by the music service provider, and stores the matching song information item by item as a playlist.

[0069] Step ⑧: Send the playlist to the vehicle control terminal.

[0070] Step ⑨: The vehicle control terminal requests the music resources in the playlist from the music service provider.

[0071] Step ⑩: The vehicle control terminal receives the music resources and plays them.

[0072] Through the embodiments of the present application, based on the scenario analysis of the large language model and the generation of dynamic recommendation prompts, the large language model can adjust the recommendation strategy according to real-time data (such as emotional state, traffic conditions, etc.) by generating dynamic prompts for music recommendations. The model will evaluate the importance of each factor and accordingly prioritize the adjustment of the recommended content (such as emotion management or driving fatigue perception). This context-based intelligent recommendation method enables the system to dynamically adapt to the driver's needs and provide a more flexible and intelligent recommendation experience.

[0073] In addition, by real-time obtaining various data such as the driver's emotional state, trip information, external environment (such as weather, road conditions, etc.) and music preferences, the system can comprehensively understand the driver's current needs. The integration and analysis of this multi-modal and multi-source data ensure that the recommendation is not only based on static data (such as historical preferences), but combines real-time situations to provide more accurate and personalized music recommendations for the driver.

[0074] Furthermore, the architecture design combining cloud and local computing enables efficient allocation of computing resources. Complex reasoning tasks are handed over to the cloud for processing, while the in-vehicle system is responsible for real-time data collection and sensitive data processing (such as in-vehicle and out-of-vehicle videos). This distributed architecture not only reduces the computing pressure on in-vehicle hardware, but also ensures the accuracy of recommendations while protecting user privacy.

[0075] Figure 5 Shows an effect schematic diagram of an example according to Figure 4 Steps ⑤ and ⑥.

[0076] As Figure 5 shown, in the prompt constructor, the recommendation factors collected in Steps ① to ④ are integrated, including the driver's music preferences (such as music style, singer preferences, etc.), real-time context data (such as emotional state, weather, traffic congestion level, driving time, etc.), to generate a set of structured prompt words.

[0077] In this way, when generating prompts, the prompt constructor sorts the priorities of recommendation factors according to the driving situation. Exemplarily, when the driver's mood is anxious, music styles related to emotion management (such as soothing music, light music) will be given priority. However, in the long-term driving scenario, overly lyrical or slow music should be avoided, and music with a strong rhythm is more considered to prevent fatigue. However, the prompts generated by the prompt constructor are presented in natural language form, which is convenient for the large language model to understand and avoids model parsing errors caused by complex data formats. Thus, by generating prompts in natural language form, the multi-dimensional integration and adjustment of recommendation factors are realized, ensuring that the information input to the large language model is accurate, complete and logical, and making full use of the context understanding ability of the large language model to enable it to generate accurate recommendation suggestions according to complex scenarios.

[0078] After the large language model receives the prompt, it conducts a context analysis of the driving scenario according to the content of the prompt and comprehensively evaluates the driver's current needs. Specifically, in Example 1, when the driver's mood is in a "calm" state, the large model will give priority to matching the user's singer preferences (such as Zhou XX, Wang XX) and style preferences (pop, light music). In Example 2, when the driver's mood is in an "angry" state, the large model will give priority to recommending songs that can improve the user's mood, which can include music by some other singers besides the user's preferred singers (for example, Deng XX, etc.) to avoid overly stimulating or exacerbating the driver's bad mood and improve the driving mood.

[0079] The large language model utilizes its powerful context understanding ability to generate a logical and targeted music recommendation list according to complex recommendation factors, and can dynamically adjust the recommendation strategy according to the real-time situation, making the music recommendation more flexible and intelligent. In addition, the generated recommendation text not only includes the song list, but also attaches the recommendation reasons, enhancing the user's trust and acceptance of the recommendation results.

[0080] Through the above combination Figure 5 of examples, it more clearly describes how multi-dimensional recommendation factors are transformed into the input of the large language model through the prompt constructor, and how the large language model generates context-related recommendation text based on the prompt, realizing multi-dimensional scenario-driven music recommendation and providing dynamic and accurate personalized services for drivers.

[0081] Through the embodiments of the present application, real-time multi-source data fusion and dynamic music recommendation are achieved. By combining large language models (such as GPT, etc.) with various real-time and non-real-time data sources (such as the driver's music preferences, emotional state, trip information, weather, etc.), dynamic music recommendation results are generated. Additionally, in the generative music recommendation guided by large language models (such as GPT), by applying the large language model to the field of in-vehicle music recommendation, the model is guided by prompt words to dynamically evaluate the importance of various factors and generate recommendations that conform to the current driving situation. Different from traditional rule-based or historical data-based recommendation systems, the system can evaluate the importance of various factors in real time and intelligently, and generate personalized and context-aware recommendation results according to the current driving scenario.

[0082] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of actions combined. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be in other sequences or carried out simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application. In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0083] In some embodiments, the embodiments of the present application provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored. The execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any one of the above-mentioned in-vehicle music recommendation methods based on a large language model.

[0084] In some embodiments, the embodiments of the present application further provide a computer program product. The computer program product includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is enabled to execute any one of the above-mentioned in-vehicle music recommendation methods based on a large language model.

[0085] In some embodiments, the embodiments of the present application further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor. Among them, the memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to execute the in-vehicle music recommendation method based on a large language model.

[0086] Figure 6This is a schematic diagram of the hardware structure of an electronic device that executes an in-vehicle music recommendation method based on a large language model provided by another embodiment of the present application. As Figure 6 shown, the device includes: One or more processors 610 and a memory 620. Figure 6 Here, one processor 610 is taken as an example.

[0087] The device that executes the in-vehicle music recommendation method based on the large language model may further include: an input device 630 and an output device 640.

[0088] The processor 610, the memory 620, the input device 630, and the output device 640 may be connected through a bus or other means. Figure 6 Here, connection through a bus is taken as an example.

[0089] The memory 620, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the in-vehicle music recommendation method based on the large language model in the embodiments of the present application. The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 620, that is, implements the in-vehicle music recommendation method based on the large language model in the above method embodiments.

[0090] The memory 620 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 620 may include a high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 620 may optionally include a memory remotely set relative to the processor 610, and these remote memories can be connected to the electronic device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0091] The input device 630 can receive input digital or character information and generate signals related to the user settings and function control of the electronic device. The output device 640 may include a display device such as a display screen.

[0092] The one or more modules are stored in the memory 620 and, when executed by the one or more processors 610, execute the in-vehicle music recommendation method based on the large language model in any of the above method embodiments.

[0093] The above-mentioned products can execute the methods provided in the embodiments of the present application, and have the corresponding functional modules and beneficial effects for executing the methods. For technical details not described in detail in this embodiment, reference can be made to the methods provided in the embodiments of the present application.

[0094] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0095] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc.

[0096] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable in-vehicle navigation devices.

[0097] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.

[0098] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this embodiment.

[0099] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A method for recommending in-vehicle music based on a large language model, applied to a server, the method comprising: Acquire a plurality of recommendation factors, wherein the plurality of recommendation factors include the driver's music preference information and the vehicle travel information, the vehicle external environment information and the driver's status information collected in real time; Filling each of the recommendation factors into a preset music recommendation prompt template to construct a music recommendation prompt word; Inputting the music recommendation prompt words into a large language model to determine corresponding music recommendation suggestion text; A recommended music list is generated according to the music recommendation suggestion text, and the recommended music list is sent to the vehicle terminal, so that the vehicle terminal can request music resources in the recommended music list from a music service provider.

2. The method according to claim 1, wherein: The obtaining of multiple recommendation factors includes: The plurality of recommendation factors are received from the vehicle-mounted terminal.

3. The method according to claim 2, wherein: The driver status information includes the driver's emotional state, and the vehicle exterior environment information includes the vehicle exterior weather conditions; Among them, the vehicle-mounted terminal is configured with a weather condition recognition model and a facial expression recognition model. The weather condition recognition model is used to identify the outside weather condition corresponding to the image taken by the outside camera collected in real time, and the facial expression recognition model is used to identify the driver's emotional state corresponding to the image taken by the inside camera collected in real time.

4. The method according to claim 2, wherein: The driver's music preference information is determined by the vehicle-mounted terminal according to the driver's user portrait and the vehicle-mounted music playing record.

5. The method according to claim 1, wherein: The vehicle-mounted terminal is configured with a plurality of service provider interfaces, each of which is used to access a corresponding music service provider; The vehicle-mounted terminal is used to call each of the service provider interfaces in the order of the recommended music list to automatically schedule and request music resources from each music service provider.

6. The method according to any one of claims 1 to 5, wherein: The step of generating a recommended music list according to the music recommendation suggestion text comprises: Filtering a list of music titles from the music recommendation suggestion text; Calling a music search engine authorized by the music service provider to determine song resource index information corresponding to each music title in the music title list; Generate a recommended music list according to the index information of each song resource.

7. The method according to claim 6, wherein: The calling of a music search engine authorized by the music service provider to determine song resource index information corresponding to each music title in the music title list includes: Retrieve whether the first music title belongs to an accessible song based on the music search engine; When it is determined that the first music title belongs to an accessible song, the song resource index information corresponding to the first music title is obtained.

8. A storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.

9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Music recommendation method and device and XR equipment

    CN120744169A

  • Recommendation method and system based on big data tags

    CN120994907A

  • Multi-mode vehicle-mounted music interaction system and method based on driving situation

    CN121657862A

  • Sound effect generation method and device, vehicle, storage medium and product

    CN122493828A