Vehicle scene and audio matching association method and device, vehicle and storage medium
By acquiring multi-source data to analyze user preferences and vehicle scenarios, and combining historical data to match and recommend audio, the problem of in-vehicle systems being unable to provide personalized audio recommendations has been solved, achieving accurate audio recommendations and improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING LANDIAN AUTOMOBILE TECHNOLOGY CO LTD
- Filing Date
- 2025-12-04
- Publication Date
- 2026-05-12
AI Technical Summary
Existing in-vehicle infotainment systems cannot provide personalized audio recommendations that match user preferences and adapt to the current vehicle scenario.
By acquiring multi-source data, analyzing current audio preferences and vehicle scene patterns, and combining historical preference data, the current user profile is determined, and then recommended audio is matched.
It improves the accuracy of recommended audio, meets users' personalized needs in different vehicle scenarios, and enhances the driving experience.
Smart Images

Figure CN122019866A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of in-vehicle audio matching and recommendation technology, and in particular to a method, device, vehicle, and storage medium for associating vehicle scene with audio matching. Background Technology
[0002] Currently, in-vehicle infotainment systems generally have music playback functions and some preliminary personalized recommendation schemes have emerged. However, the current recommendation schemes generally rely on the user's historical behavior data of a single music application or make recommendations based on the current time (such as recommending soothing and relaxing music during the off-hours). They cannot provide users with a truly accurate and personalized audio experience that matches the current vehicle scenario. Summary of the Invention
[0003] This application provides a method, apparatus, vehicle, and storage medium for matching and associating vehicle scene with audio, in order to solve the technical problem of how to make recommended audio conform to user preferences and adapt to vehicle scene.
[0004] Firstly, this application provides a method for associating vehicle scene with audio, the method comprising: Acquire multi-source data; The audio preference analysis results and vehicle scene mode are determined based on the multi-source data. The current user profile is determined based on the current audio preference analysis results and the historical audio preference analysis results in the memory bank; The recommended audio is determined based on the current user profile and the scene pattern.
[0005] Optionally, acquire multi-source data, including: Collect raw data from multiple sources at preset intervals, and / or collect the raw data from multiple sources when a preset trigger event is detected; The multi-source raw data is preprocessed to obtain the multi-source data; wherein the data preprocessing includes at least one of data cleaning, data deduplication, and format unification.
[0006] Optionally, determining the audio preference analysis result based on the multi-source data includes: Preliminary audio preference information is determined based on music application data from the multi-source data; Extract explicit expression data from the speech data in the multi-source data; The results of the current audio preference analysis are determined based on the preliminary audio preference information and the explicit expression data.
[0007] Optionally, determining the current audio preference analysis result based on the preliminary audio preference information and the explicit expression data includes: When the explicit expression data contains audio preference features, a target preference analysis weight is obtained; wherein, the target preference analysis weight includes a first weight coefficient corresponding to the preliminary audio preference information and a second weight coefficient corresponding to the audio preference feature; the first weight coefficient is less than the second weight coefficient; The audio preference analysis result for this instance is determined based on the preliminary audio preference information, the first weighting coefficient, the audio preference features, and the second weighting coefficient. If the explicit expression data does not contain audio preference features, the result of the current audio preference analysis is determined based on the preliminary audio preference information.
[0008] Optionally, obtain the target preference analysis weights, including: Obtain the confidence level of the audio preference features; Obtain a preset preference mapping relationship; wherein the preference mapping relationship is a mapping relationship between confidence level and preference analysis weight; The target preference analysis weights are determined based on the confidence level and the preference mapping relationship; wherein, the higher the confidence level, the smaller the first weight coefficient and the larger the second weight coefficient.
[0009] Optionally, determining the vehicle's scene mode based on the multi-source data includes: Extract occupant status information, vehicle status information, and environmental perception data from the multi-source data; Based on the occupant status information, the vehicle status information, and the environmental perception data, the target scene mode with the highest matching degree is determined from the preset scene modes; The target scene mode is taken as the scene mode in which the vehicle is currently located.
[0010] Optionally, the current user profile is determined based on the current audio preference analysis results and the historical audio preference analysis results in the memory bank, including: Retrieve the results of historical audio preference analysis from the memory bank; The current audio preference analysis results and the historical audio preference analysis results are assigned time weight coefficients according to time; wherein, the earlier the time corresponding to the analysis result, the smaller the assigned time weight coefficient. The current user profile is determined based on the current audio preference analysis results, the historical audio preference analysis results, and the time weighting coefficient.
[0011] Optionally, each analysis result in the current audio preference analysis result and the historical audio preference analysis result includes audio preference features and corresponding preference degrees; determining the current user profile based on the current audio preference analysis result, the historical audio preference analysis result, and the time weight coefficient includes: For each audio preference feature, obtain the time weight coefficient corresponding to the analysis result containing the audio preference feature, and the preference degree of the audio preference feature in the corresponding analysis result. The sum of the products of each time weight coefficient and the corresponding preference degree is used as the preference feature weighted calculation value of the audio preference feature. Sort all the weighted calculated values of the preference features, and determine the current user profile based on the audio preference features corresponding to the top N weighted calculated values of the preference features; where N is greater than 1.
[0012] Optionally, determining matching recommended audio based on the current user profile and the scene pattern includes: The target audio type is determined based on the scene mode and the preset mapping relationship; wherein, the mapping relationship is the mapping relationship between scene mode and audio type; Determine the audio list corresponding to the target audio type; Based on the current user profile, a matching recommended audio is determined from the audio list; or, The current user profile and the scene pattern are input into a pre-trained artificial intelligence model; Obtain the recommended audio output by the artificial intelligence model.
[0013] Optionally, after determining the matching recommended audio based on the current user profile and the scene pattern, the method further includes: Obtain the recommendation reason corresponding to the recommended audio; Display the recommended audio and the reasons for the recommendation; Obtain user feedback on the recommended audio, and redetermine the latest scene mode and the latest current user profile; Based on the feedback information, the latest current user profile, and the latest scene pattern, the recommended audio for matching is redefined.
[0014] Secondly, this application provides a vehicle scene and audio matching and association device, the device comprising: The acquisition module is used to acquire data from multiple sources; The first determining module is used to determine the current audio preference analysis result and the vehicle's scene mode based on the multi-source data; The second determining module is used to determine the current user profile based on the current audio preference analysis result and the historical audio preference analysis results in the memory bank; The matching module is used to determine the recommended audio based on the current user profile and the scene mode.
[0015] Thirdly, this application provides a vehicle including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When the processor executes a program stored in memory, it implements the vehicle scene and audio matching association method described in any embodiment of the first aspect.
[0016] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the vehicle scene and audio matching association method as described in any embodiment of the first aspect.
[0017] Compared with the prior art, the technical solution provided in this application has the following advantages: The method provided in this application acquires multi-source data; determines the current audio preference analysis result and the vehicle's scene mode based on the multi-source data; determines the current user profile based on the current audio preference analysis result and historical audio preference analysis results in the memory bank; and determines matching recommended audio based on the current user profile and the scene mode. This method first determines the current audio preference analysis result and the vehicle's scene mode based on multi-source data, then determines the current user profile based on the current audio preference analysis result and historical audio preference analysis results in the memory bank. This current user profile accurately represents the current user's preferences. Finally, determining matching recommended audio based on the current user profile and the scene mode allows the recommended audio to both match user preferences and adapt to the vehicle's current scene mode, improving the accuracy of the recommended audio and the user's driving experience. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0021] Figure 1 A system architecture diagram of a vehicle scene and audio matching association method provided in one embodiment of this application; Figure 2 A flowchart illustrating a method for matching and associating vehicle scene with audio, provided as an embodiment of this application; Figure 3 This is a schematic diagram illustrating a vehicle scene and audio matching recommendation according to one embodiment of this application; Figure 4 A schematic diagram of the structure of a vehicle scene and audio matching association device provided in one embodiment of this application; Figure 5 This is a schematic diagram of the structure of a vehicle provided in one embodiment of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0023] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0024] To address the technical problem of how to make recommended audio conform to user preferences and adapt to vehicle scenarios in the prior art, this application provides a method, device, vehicle, and storage medium for matching and associating vehicle scenarios with audio, which enables recommended audio to not only conform to user preferences but also adapt to the current scenario mode of the vehicle, thereby improving the accuracy of recommended audio and the user's driving experience.
[0025] The first embodiment of this application provides a method for matching and associating vehicle scene with audio, which can be applied to, for example... Figure 1The system architecture shown includes either vehicle 101 or server 102. The type of vehicle 101 is not limited; it can be a gasoline-powered vehicle, a pure electric vehicle, a hybrid vehicle, or a fuel cell vehicle, etc. Server 102 can be a local server, a cloud server, or a server cluster. It should be understood that when this method is applied to vehicle 101, all execution processes can be completed within the vehicle itself. When this method is applied to server 102, server 102 can obtain multi-source data from vehicle 101, process it on the server, and then send recommended audio to vehicle 101 after processing.
[0026] Next, based on this system architecture, the method for associating the vehicle's scene with audio will be described in detail, such as... Figure 2 The method for associating the vehicle's scene with audio includes: Step 201: Obtain multi-source data.
[0027] It can collect multi-source data through vehicle bus system, sensors, application programming interface, voice acquisition device, etc. It can continuously collect multi-source data in real time, or collect data once at a time interval, or collect data when a trigger event is detected.
[0028] In one embodiment, acquiring multi-source data includes: collecting multi-source raw data at preset time intervals, and / or collecting multi-source raw data when a preset trigger event is detected; performing data preprocessing on the multi-source raw data to obtain multi-source data; wherein, data preprocessing includes at least one of data cleaning, data deduplication, and format unification.
[0029] In this embodiment, the acquisition of multi-source data can be automatically triggered, for example, by acquiring raw multi-source data at preset intervals, such as 5 minutes or 10 minutes. Alternatively, the acquisition can be event-driven, such as acquiring data when a preset trigger event is detected. These preset trigger events can be user commands, such as switching songs or adding songs to favorites, or events such as detecting user fatigue, a vehicle entering a highway, or the start of commuting time. Upon detecting a preset trigger event, the acquisition of multi-source data is immediately initiated. After acquiring the raw multi-source data, preprocessing can be performed, such as data cleaning, deduplication, and format standardization, to ensure the processed data meets the requirements for subsequent steps.
[0030] Specifically, multi-source data can include one or more of the following: music application data, voice data, and state-aware data, as well as other collected data, without limitation. Music application data can include local audio source data and third-party application data. Local audio source data can be music stored on a Universal Serial Bus Flash Drive (USB drive) connected to the vehicle, Bluetooth music connected via Bluetooth, or music from the car radio—audio source data locally on the vehicle. This data can be obtained by scanning the USB device with a local audio source data collector to acquire audio file metadata, record local music playback behavior logs, and store radio favorites. Third-party application data can be data read from third-party music applications via application programming interfaces (APIs), such as user listening history, favorite song lists, created playlists, playback behavior records (e.g., playback count, playback completion rate, manual selection records), and explicit data such as user likes, favorites, comments, ratings, shares, forwards, and dislikes. Voice data can include dialogue records stored in the vehicle, explicit preference expressions, or real-time voice streams collected by voice acquisition devices. Status perception data can be data on the perception of people, vehicles, and the environment, such as occupant status information, vehicle status information, and environmental perception data collected through in-vehicle bus systems and sensors. Specifically, vehicle status information such as vehicle speed and engine speed can be obtained through the Controller Area Network (CAN bus) or other in-vehicle buses; environmental perception data such as weather conditions and external environment can be obtained through environmental perception systems; and occupant status information can be obtained through in-vehicle sensors and cameras. Additionally, status perception data can also include geographical location information and time information; geographical location information can be obtained through a positioning module, and time information can be obtained through a clock module.
[0031] Step 202: Determine the audio preference analysis results and vehicle scene mode based on multi-source data.
[0032] The system can analyze audio preferences in a given moment by using music application data and voice data to obtain the audio preference analysis results, and analyze the vehicle's scene patterns by using state perception data.
[0033] In one embodiment, determining the current audio preference analysis result based on multi-source data includes: determining preliminary audio preference information based on music application data in the multi-source data; extracting explicit expression data from speech data in the multi-source data; and determining the current audio preference analysis result based on the preliminary audio preference information and the explicit expression data.
[0034] In this embodiment, preliminary audio preference information can be determined based on music application data. This preliminary audio preference information is used to characterize the audio type preferred by the user. For example, preliminary audio preference information can be inferred based on information such as playback records and single-track listening frequency in music application data. Explicit expression data is data that directly expresses preferences extracted from speech data. The audio preference analysis results determined by combining the preliminary audio preference information determined based on objective music application data and the explicit expression data of the user's subjective expression can accurately reflect the user's audio preferences.
[0035] Specifically, preliminary audio preference information is determined based on music application data from multi-source data, including: extracting audio file source data, music playback behavior logs, and favorite channel information from local audio source data in music application data; extracting historical listening records, favorite song lists, playlists, and playback behavior records from third-party application data in music application data; and determining preliminary audio preference information based on audio file source data, music playback behavior logs, favorite channel information, historical listening records, favorite song lists, playlists, and playback behavior records.
[0036] In this embodiment, audio file source data, music playback behavior logs, and favorite channel information can be extracted from local audio source data; historical listening records, favorite song lists, playlists, and playback behavior records can be extracted from third-party application data; and explicit expression data can be extracted from voice data. Then, preliminary audio preference information is determined based on the audio file source data, music playback behavior logs, favorite channel information, historical listening records, favorite song lists, playlists, and playback behavior records. For example, the song types are determined based on all songs appearing in the audio file source data, music playback behavior logs, favorite channel information, historical listening records, favorite song lists, playlists, and playback behavior records, and the frequency of song appearance is counted. Preliminary audio preference information, such as the user's preferred singer, singer style, genre information, etc., is obtained through weighted calculation based on song type and frequency. Then, the audio preference analysis result for the current session is determined based on the preliminary audio preference information and explicit expression data, thus accurately expressing the user's current audio preference analysis result and providing a data foundation for accurately representing the current user's preferences.
[0037] In one embodiment, determining the audio preference analysis result based on preliminary audio preference information and explicit expression data includes: if the explicit expression data contains audio preference features, obtaining a target preference analysis weight; wherein the target preference analysis weight includes a first weight coefficient corresponding to the preliminary audio preference information and a second weight coefficient corresponding to the audio preference features; the first weight coefficient is less than the second weight coefficient; determining the audio preference analysis result based on the preliminary audio preference information, the first weight coefficient, the audio preference features, and the second weight coefficient; if the explicit expression data does not contain audio preference features, determining the audio preference analysis result based on the preliminary audio preference information.
[0038] In this embodiment, the determination of the audio preference analysis result based on preliminary audio preference information and explicit expression data can be divided into two cases.
[0039] In the first scenario, when audio preference features are included in the explicit expression data, since the explicit expression data is extracted from the user's voice data and has a high degree of reliability, the explicit expression data can be used as a key reference to determine the results of the current audio preference analysis. For example, the target preference analysis weights containing the first weight coefficient and the second weight coefficient can be obtained. The first weight coefficient is the weight assigned to the initial audio preference information, and the second weight coefficient is the weight assigned to the audio preference features in the explicit expression data.
[0040] Both preliminary audio preference information and audio preference features represent preference characteristics. The difference between the two is that preliminary audio preference information is a preference feature determined based on objective music application data, while audio preference features are preference features extracted based on explicit expression data of users' subjective expressions.
[0041] The results of the current audio preference analysis are determined based on preliminary audio preference information, a first weighting coefficient, audio preference features, and a second weighting coefficient. This can include at least the following two methods.
[0042] Method 1: The preliminary audio preference information may only contain audio preference features. In this case, all audio preference features contained in the preliminary audio preference information can be assigned a first weight coefficient, and the audio preference features determined based on the explicit expression data can be assigned a second weight coefficient. If there are overlapping audio preference features in the preliminary audio preference information and the explicit expression data, the weight of the overlapping audio preference features is higher than the weight of audio preference features that exist only in the preliminary audio preference information or only in the explicit expression data. Therefore, the result of the current audio preference analysis can be determined based on the overlapping audio preference features. If there are no overlapping audio preference features in the preliminary audio preference information and the explicit expression data, since the first weight coefficient is smaller than the second weight coefficient, the weight of the audio preference features determined based on the explicit expression data will be higher than the weight of the audio preference features contained in the preliminary audio preference information. Therefore, the result of the current audio preference analysis can be determined based on the audio preference features determined by the explicit expression data.
[0043] Method 2: The preliminary audio preference information can include audio preference features and their corresponding preference degrees, and the explicit expression data can also contain audio preference features and their corresponding preference degrees. In this case, all audio preference features included in the preliminary audio preference information can be assigned a corresponding preference degree and a first weight coefficient, while all audio preference features in the explicit expression data can be assigned a corresponding preference degree and a second weight coefficient. That is, all audio preference features are weighted using their corresponding weight coefficients and preference degrees. The audio preference analysis result for each audio preference feature is determined based on the weighted result, for example, by determining the result based on a preset number of audio preference features with higher weighted values. Because the audio preference features contained in the explicit expression data extracted from the speech data are assigned higher weights, the resulting audio preference analysis result is more consistent with the user's actual preferences. The preference degree represents the user's degree of preference for the audio preference feature, which can be determined based on the frequency of the audio preference feature in the preliminary audio preference information, the frequency of the audio preference feature in the explicit expression data, or the context of the audio preference feature.
[0044] Specifically, obtaining the target preference analysis weights includes: obtaining the confidence level of audio preference features; obtaining a preset preference mapping relationship; wherein the preference mapping relationship is a mapping relationship between confidence level and preference analysis weights; determining the target preference analysis weights based on the confidence level and preference mapping relationship; wherein the higher the confidence level, the smaller the first weight coefficient and the larger the second weight coefficient.
[0045] The preference mapping relationship contains multiple preference analysis weights. Different confidence levels correspond to different preference analysis weights. The preference analysis weights determined from the preference mapping relationship based on confidence levels can be called target preference analysis weights. Confidence levels can be obtained based on explicit expressive data, such as extracting audio preference features from the context. For example, if the speech data contains phrases like "XXX's song is so beautiful, I haven't heard such gentle music in a long time," the extracted audio preference features include "XXX (singer)," "gentle music," and the song title determined based on song information. The confidence level can be quantified into a specific value based on descriptive keywords such as "so beautiful" and "so beautiful." For example, a mapping relationship between keywords and confidence levels can be pre-defined. After a keyword is detected, the quantified value corresponding to that keyword can be determined from the mapping relationship. Alternatively, confidence levels can be scored using a large language model. For example, the speech data can be input into a large language model, and the model can perform semantic analysis to obtain the confidence level. A higher confidence level indicates a higher credibility of the audio preference feature, thus assigning it a higher weight coefficient, while lowering the weight coefficient of the initial audio preference information; that is, the smaller the first weight coefficient and the larger the second weight coefficient. Conversely, a lower confidence level indicates a lower credibility of the audio preference feature, thus assigning it a lower weight coefficient, while lowering the weight coefficient of the initial audio preference information; that is, the larger the first weight coefficient and the smaller the second weight coefficient. It should be understood that "larger" and "smaller" here are relative concepts between different preference analysis weights; within the same preference analysis weight, the first weight coefficient is smaller than the second weight coefficient.
[0046] In the second scenario, where the explicit data does not include audio preference features, the results of the audio preference analysis can be determined solely based on preliminary audio preference information. Since this preliminary audio preference information is determined from audio file source data, music playback behavior logs, favorite channel information, historical listening records, favorite song lists, playlists, and playback behavior records, it can accurately represent the user's actual preferences.
[0047] Based on multi-source data, not only can the results of the current audio preference analysis be determined, but also the scene mode of the vehicle can be determined.
[0048] The system can determine the most suitable scenario mode from preset scenario modes based on state perception data. If the current scenario mode is determined to be a new scenario mode, it can be added to the preset scenario modes. There are no restrictions on the specific types of preset scenario modes, which can be preset as needed, such as commuting mode, road trip mode, lunch break mode, date mode, rainy driving mode, highway mode, suburban mode, etc. Each preset scenario mode can correspond to different state perception data.
[0049] Specifically, preset scene modes can be obtained by classifying historical scene data. For example, clustering algorithms can be used to classify driving scenarios. For instance, historical scene data includes multiple data points, each of which may contain vehicle status information such as speed, braking frequency, and turning angle, as well as environmental perception data such as road segment type. By analyzing the feature similarity of each data point through clustering algorithms, if there are five data points with the feature "high-speed constant speed driving, little braking", these five similar data points are automatically grouped into a cluster. If there are six data points with the feature "low speed, frequent braking, short distance", these six similar data points are automatically grouped into a cluster. If there are three data points with the feature "medium speed, occasional turning, complex road segment", these three similar data points are automatically grouped into a cluster. The feature consistency within the same cluster is high, while the feature differences between different clusters are obvious. For example, the above 14 data points were finally divided into three categories: high-speed mode, suburban mode, and commuting mode. These three categories are then used as preset scene modes. It should be noted that the above example uses historical scene data containing vehicle status information and environmental perception data. In actual classification, other factors such as driver and passenger status information can also be considered. When performing clustering algorithm classification, it can be classified into various preset scene modes according to the characteristics of the actual historical scene data, such as commuting mode, self-driving tour mode, lunch break mode, dating mode, rainy driving mode, highway mode, suburban mode, etc., without limitation.
[0050] In one embodiment, determining the vehicle's scene mode based on multi-source data includes: extracting occupant status information, vehicle status information, and environmental perception data from the multi-source data; determining the target scene mode with the highest matching degree from preset scene modes based on the occupant status information, vehicle status information, and environmental perception data; and using the target scene mode as the scene mode in which the vehicle is currently located.
[0051] In this embodiment, at least three dimensions, such as driver and passenger status information, vehicle status information, and environmental perception data, can be used to match the vehicle in a preset scene mode. For example, the matching degree of each dimension can be calculated. If the matching degree of all three dimensions is greater than a preset threshold (such as 80% or other preset values), or the sum of the matching degrees of the three dimensions is greater than a certain set value, the matching can be considered successful. The target scene mode with the highest matching degree can be determined from all successfully matched scene modes as the scene mode in which the vehicle is currently located.
[0052] Step 203: Determine the current user profile based on the current audio preference analysis results and the historical audio preference analysis results in the memory bank.
[0053] Each audio preference analysis result can be stored in the memory bank according to the analysis time. When determining the current user profile, it can be determined based on the current audio preference analysis result and the historical audio preference analysis results in the memory bank.
[0054] In one embodiment, determining the current user profile based on the current audio preference analysis result and the historical audio preference analysis result in the memory bank includes: obtaining the historical audio preference analysis result in the memory bank; assigning time weight coefficients to the current audio preference analysis result and the historical audio preference analysis result according to time; wherein, the earlier the time corresponding to the analysis result, the smaller the assigned time weight coefficient; and determining the current user profile based on the current audio preference analysis result, the historical audio preference analysis result, and the time weight coefficient.
[0055] In this embodiment, when determining the current user profile, time weight coefficients can be assigned to the current audio preference analysis results and historical audio preference analysis results according to time. The earlier the time corresponding to the analysis result, the smaller the assigned time weight coefficient, and the later the time corresponding to the analysis result, the larger the assigned time weight coefficient. That is, recent behavior is given higher weight, and the weight of historical behavior decays over time. The time weight coefficient assigned to the current audio preference analysis result is the largest, followed by the historical audio preference analysis result whose occurrence time is closest to the current time, and so on. The historical audio preference analysis result whose occurrence time is the earliest is assigned the smallest time weight coefficient. Therefore, when user preferences change, the current user profile determined based on the current audio preference analysis result, historical audio preference analysis result, and time weight coefficient is closer to the real situation.
[0056] In one embodiment, each analysis result in the current audio preference analysis and the historical audio preference analysis results contains audio preference features and corresponding preference degrees. The audio preference features in different analysis results may not be exactly the same. Determining the current user profile based on the current audio preference analysis result, the historical audio preference analysis result, and the time weight coefficient includes: for each audio preference feature, obtaining the time weight coefficient corresponding to the analysis result containing the audio preference feature, and the preference degree of the audio preference feature in the corresponding analysis result; using the sum of the products of each time weight coefficient and the corresponding preference degree as the weighted calculation value of the audio preference feature; sorting all the weighted calculation values of preference features, and determining the current user profile based on the audio preference features corresponding to the top N weighted calculation values of preference features; wherein, N is greater than 1.
[0057] In this embodiment, for each audio preference feature, the time weight coefficient corresponding to the analysis result containing the audio preference feature and the preference degree of the audio preference feature in the corresponding analysis result can be obtained. The preference feature weighted calculation value of the audio preference feature is calculated based on the time weight coefficient and preference degree corresponding to the audio preference feature. Then, all the preference feature weighted calculation values are sorted, and the current user profile is determined based on the audio preference features corresponding to the top N preference feature weighted calculation values.
[0058] Examples are given below: Let's refer to the current audio preference analysis result as Result 1. If the historical audio preference analysis results include Result 2 and Result 3, where Result 1 corresponds to the current time (i.e., the latest occurrence time), Result 2 occurred later, and Result 3 occurred earliest, then according to the principle of assigning higher weight to recent behavior and decreasing weight to historical behavior over time, we can assign a time weight coefficient of 0.5 to Result 1, a time weight coefficient of 0.3 to Result 2, and a time weight coefficient of 0.2 to Result 3.
[0059] Extract the preference feature information of result 1, such as audio preference feature 1, audio preference feature 2, and audio preference feature 3, where the preference degree of audio preference feature 1 is 80%, the preference degree of audio preference feature 2 is 70%, and the preference degree of audio preference feature 3 is 60%.
[0060] Extract the preference feature information of result 2, such as audio preference feature 1, audio preference feature 3, and audio preference feature 4. Among them, the preference degree of audio preference feature 1 is 90%, the preference degree of audio preference feature 3 is 60%, and the preference degree of audio preference feature 4 is 70%.
[0061] Extract the preference feature information of result 3, such as audio preference feature 1, audio preference feature 2, and audio preference feature 4. Among them, the preference degree of audio preference feature 1 is 80%, the preference degree of audio preference feature 2 is 70%, and the preference degree of audio preference feature 4 is 70%.
[0062] Then, the weighted calculated value of audio preference feature 1 is 0.5×80%+0.3×90%+0.2×80%=0.83, the weighted calculated value of audio preference feature 2 is 0.5×70%+0.2×70%=0.49, the weighted calculated value of audio preference feature 3 is 0.5×60%+0.3×60%=0.48, and the weighted calculated value of audio preference feature 4 is 0.3×70%+0.2×70%=0.35. After sorting, if N is set to 3, the current user profile is determined based on audio preference feature 1, audio preference feature 2, and audio preference feature 3. The preference feature information can be, for example, information representing preference features such as preferred singers, genres, and song types, without restrictions.
[0063] It should be noted that the number of results included in the above historical audio preference analysis results, and the specific time weight coefficients assigned to each result, are only illustrative examples of how to assign time weight coefficients according to the time corresponding to the analysis results. They do not represent any restrictions on the specific values of the time weight coefficients. Similarly, the number of audio preference features included in each result and the preference degree corresponding to the audio preference features are also only illustrative examples and do not represent any specific restrictions on them.
[0064] Step 204: Determine the matching recommended audio based on the current user profile and scene pattern.
[0065] This method first determines the current audio preference analysis results and vehicle scene mode based on multi-source data. Then, it determines the current user profile based on the current audio preference analysis results and historical audio preference analysis results in the memory bank. This current user profile can accurately represent the current user preferences. Finally, it determines the matching recommended audio based on the current user profile and scene mode. This can make the recommended audio not only in line with user preferences but also adapt to the current scene mode of the vehicle, thereby improving the accuracy of recommended audio and the user's driving experience.
[0066] In one embodiment, determining the matching recommended audio based on the current user profile and scene pattern can include at least the following methods: The first approach, a rule-based recommendation system, is suitable for scenarios with limited computing resources. For example, it involves determining the target audio type based on scene patterns and a preset mapping relationship, where the mapping relationship is between scene patterns and audio types; determining the audio list corresponding to the target audio type; and selecting matching recommended audio from the audio list based on the current user profile.
[0067] This method allows for the pre-setting of a mapping relationship between scene modes and audio types. Based on the scene mode and the pre-setting mapping relationship, the target audio type corresponding to the scene mode is determined. Then, based on the target audio type, its corresponding audio list is determined. Finally, based on the current user profile, matching recommended audio is selected from the audio list. For example, if the target audio type for the scene mode is determined to be gentle music based on the scene mode and the pre-setting mapping relationship, then the audio list can be determined to include all gentle music. If the current user profile indicates a preference for classical music and soothing music, then gentle classical music and gentle soothing music are matched from the audio list as the final recommended audio. This ensures that the recommended audio matches both user preferences and the current scene mode of the vehicle, improving the accuracy of the recommended audio and the user's driving experience.
[0068] The second approach, based on deep learning models, can provide more accurate recommendations. For example, it involves inputting the current user profile and scene pattern into a pre-trained artificial intelligence model and obtaining the recommended audio output by the model.
[0069] In this approach, a deep learning-based artificial intelligence model can be pre-trained. When in use, the current user profile and scene mode are input into the artificial intelligence model to obtain the recommended audio output by the model. It should be understood that the accuracy of the artificial intelligence model's output is closely related to the training effect. The artificial intelligence model can be trained based on calibrated training data to improve the accuracy of recommended audio.
[0070] Specifically, the AI model can represent input parameters such as the current user profile and scene pattern as a unified feature vector. Based on an attention mechanism, it calculates the weights of each factor and generates recommended audio. Furthermore, during audio generation, recommendation diversity control can be implemented. For example, a diversity control parameter can be set. If the weight of the diversity control parameter is set to 0, audio recommendations are based solely on the current user profile and scene pattern. If a weight is assigned to the diversity control parameter, such as 0.05, the final output recommended audio can include some music different from the user's preferences. This non-preferred music can bring novelty to the user and prevent the recommended audio from becoming too monotonous. The balance between recommendation accuracy and novelty can be adjusted by changing the weight of the diversity control parameter. It is important to note that non-preferred music is not necessarily music that the user dislikes. It may be that objective music application data does not include this type of music, nor does the user's explicit subjective expression data, resulting in the reflected audio preference not involving this type of music. By assigning weights to diversity control parameters, we can determine whether users prefer a particular type of music based on their feedback on non-preferred music in the recommended audio. This improves the preference recognition ability of the AI model, proactively expands users' acceptance of different types of music, and enhances the accuracy of subsequent recommendations.
[0071] The third approach is based on a combination of rule-based systems and deep learning models.
[0072] In this approach, a hybrid model combining rule systems and deep learning models can leverage the strengths of both to improve the accuracy and applicability of recommended audio.
[0073] In one embodiment, after determining the matching recommended audio based on the current user profile and scene mode, the method further includes: obtaining the recommendation reason corresponding to the recommended audio; displaying the recommended audio and the recommendation reason; obtaining user feedback information on the recommended audio; and re-determining the latest scene mode and the latest current user profile; and re-determining the matching recommended audio based on the feedback information, the latest current user profile, and the latest scene mode.
[0074] In this embodiment, the recommendation reason and the recommended audio can be displayed on the intelligent music recommendation assistant. The recommended audio can be a recommended playlist, and interactive options such as "Play Now," "Remind Me Later," and "Don't Recommend Again" can also be displayed. Users can provide feedback on the recommended audio and input feedback information. Based on the feedback information, the current user profile, and the scene mode, the matching recommended audio can be re-determined. Specifically, the illustration of vehicle scene and audio matching recommendation is as follows: Figure 3 When driving in rainy nighttime conditions, the system can recommend a tranquil playlist featuring soothing music designed to calm the driver. Below the recommended audio, interactive options include "Play Now," "Remind Me Later," and "Don't Recommend Again." For example, clicking "Play Now" will play the songs in the playlist in order. Clicking "Remind Me Later" will prompt for a re-recommendation after 5 or 10 minutes; this is understandable, as the recommended audio might change based on new status data. Clicking "Don't Recommend Again" will prevent further audio recommendations during the current driving session. In addition, the intelligent music recommendation assistant can display the reasons for recommending the current playlist, as well as a voice feedback interface. This interface allows for voice interaction with the assistant, facilitating the understanding of real-time user needs and enabling adjustments to the recommended audio. It's understandable that after obtaining user feedback, the assistant can redetermine the latest scene mode and the current user profile. For example, it can re-acquire multi-source data, determine the latest scene mode and current audio preference analysis results based on this re-acquired data, and determine the latest current user profile based on the current audio preference analysis results and historical audio preference analysis results in the memory bank. Then, new recommended audio can be re-matched based on the feedback, the latest scene mode, and the latest current user profile. After the recommendation is completed, both explicit and implicit feedback results can be collected and analyzed. Explicit feedback results can include direct user evaluations of the recommendation results (voice confirmation or rejection, or the option to no longer recommend), while implicit feedback results can include user actions during playback (e.g., playing the entire playlist, quickly skipping songs), or user satisfaction inferred from these actions. After obtaining explicit and implicit feedback results, the recommendation rule system and artificial intelligence model can be optimized. Multi-turn dialogues can also be used to verify the correctness of the understanding of user voice preferences, so that the recommended audio not only meets user preferences but also adapts to the current scenario mode of the vehicle, thereby improving the accuracy of recommended audio and the user's driving experience.
[0075] Based on the same technical concept, the second embodiment of this application provides a vehicle scene and audio matching and association device, such as... Figure 4 The device includes: Module 401 is used to acquire data from multiple sources; The first determining module 402 is used to determine the current audio preference analysis result and the vehicle scene mode based on the multi-source data; The second determining module 403 is used to determine the current user profile based on the current audio preference analysis result and the historical audio preference analysis result in the memory bank; The matching module 404 is used to determine the recommended audio based on the current user profile and the scene mode.
[0076] This device can first determine the current audio preference analysis results and vehicle scene mode based on multi-source data. Then, it can determine the current user profile based on the current audio preference analysis results and the historical audio preference analysis results in the memory bank. This current user profile can accurately represent the current user preferences. Then, it can determine the matching recommended audio based on the current user profile and scene mode. This can make the recommended audio not only in line with the user's preferences but also adapt to the current scene mode of the vehicle, thereby improving the accuracy of the recommended audio and the user's driving experience.
[0077] like Figure 5 As shown in the figure, this application embodiment provides a vehicle, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114. Memory 113 is used to store computer programs; In one embodiment of this application, the processor 111, when executing the program stored in the memory 113, implements the vehicle scene and audio matching association method provided in any of the foregoing method embodiments.
[0078] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0079] The communication interface is used for communication between the aforementioned terminal and other devices.
[0080] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0081] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0082] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the vehicle scene and audio matching association method provided in any of the foregoing method embodiments.
[0083] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0084] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0085] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0086] It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. In the description, suffixes such as "module," "part," or "unit" used to denote elements are used solely for illustrative purposes and have no specific meaning in themselves. Therefore, "module," "part," or "unit" may be used interchangeably.
[0087] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for matching and associating vehicle scene with audio, characterized in that, The method includes: Acquire multi-source data; The audio preference analysis results and vehicle scene mode are determined based on the multi-source data. The current user profile is determined based on the current audio preference analysis results and the historical audio preference analysis results in the memory bank; The recommended audio is determined based on the current user profile and the scene pattern.
2. The method according to claim 1, characterized in that, Acquire multi-source data, including: Collect raw data from multiple sources at preset intervals, and / or collect the raw data from multiple sources when a preset trigger event is detected; The multi-source raw data is preprocessed to obtain the multi-source data; wherein the data preprocessing includes at least one of data cleaning, data deduplication, and format unification.
3. The method according to claim 1, characterized in that, The results of the current audio preference analysis are determined based on the multi-source data, including: Preliminary audio preference information is determined based on music application data from the multi-source data; Extract explicit expression data from the speech data in the multi-source data; The results of the current audio preference analysis are determined based on the preliminary audio preference information and the explicit expression data.
4. The method according to claim 3, characterized in that, The results of the current audio preference analysis are determined based on the preliminary audio preference information and the explicit expression data, including: When the explicit expression data contains audio preference features, a target preference analysis weight is obtained; wherein, the target preference analysis weight includes a first weight coefficient corresponding to the preliminary audio preference information and a second weight coefficient corresponding to the audio preference feature; the first weight coefficient is less than the second weight coefficient; The audio preference analysis result for this instance is determined based on the preliminary audio preference information, the first weighting coefficient, the audio preference features, and the second weighting coefficient. If the explicit expression data does not contain audio preference features, the result of the current audio preference analysis is determined based on the preliminary audio preference information.
5. The method according to claim 4, characterized in that, Obtain the target preference analysis weights, including: Obtain the confidence level of the audio preference features; Obtain a preset preference mapping relationship; wherein the preference mapping relationship is a mapping relationship between confidence level and preference analysis weight; The target preference analysis weights are determined based on the confidence level and the preference mapping relationship; wherein, the higher the confidence level, the smaller the first weight coefficient and the larger the second weight coefficient.
6. The method according to claim 1, characterized in that, Determining the vehicle's scene mode based on the multi-source data includes: Extract occupant status information, vehicle status information, and environmental perception data from the multi-source data; Based on the occupant status information, the vehicle status information, and the environmental perception data, the target scene mode with the highest matching degree is determined from the preset scene modes; The target scene mode is taken as the scene mode in which the vehicle is currently located.
7. The method according to claim 1, characterized in that, The current user profile is determined based on the current audio preference analysis results and the historical audio preference analysis results in the memory bank, including: Retrieve the results of historical audio preference analysis from the memory bank; The current audio preference analysis results and the historical audio preference analysis results are assigned time weight coefficients according to time; wherein, the earlier the time corresponding to the analysis result, the smaller the assigned time weight coefficient. The current user profile is determined based on the current audio preference analysis results, the historical audio preference analysis results, and the time weighting coefficient.
8. The method according to claim 7, characterized in that, Each analysis result in the current audio preference analysis and the historical audio preference analysis includes audio preference features and corresponding preference degrees; The current user profile is determined based on the current audio preference analysis results, the historical audio preference analysis results, and the time weighting coefficient, including: For each audio preference feature, obtain the time weight coefficient corresponding to the analysis result containing the audio preference feature, and the preference degree of the audio preference feature in the corresponding analysis result. The sum of the products of each time weight coefficient and the corresponding preference degree is used as the preference feature weighted calculation value of the audio preference feature. Sort all the weighted calculated values of the preference features, and determine the current user profile based on the audio preference features corresponding to the top N weighted calculated values of the preference features; where N is greater than 1.
9. The method according to claim 1, characterized in that, Based on the current user profile and the scene pattern, the recommended audio is determined, including: The target audio type is determined based on the scene mode and the preset mapping relationship; wherein, the mapping relationship is the mapping relationship between scene mode and audio type; Determine the audio list corresponding to the target audio type; Based on the current user profile, a matching recommended audio is determined from the audio list; or, The current user profile and the scene pattern are input into a pre-trained artificial intelligence model; Obtain the recommended audio output by the artificial intelligence model.
10. The method according to claim 1, characterized in that, After determining the matching recommended audio based on the current user profile and the scene pattern, the method further includes: Obtain the recommendation reason corresponding to the recommended audio; Display the recommended audio and the reasons for the recommendation; Obtain user feedback on the recommended audio, and redetermine the latest scene mode and the latest current user profile; Based on the feedback information, the latest current user profile, and the latest scene pattern, the recommended audio for matching is redefined.
11. A vehicle scene and audio matching and association device, characterized in that, The device includes: The acquisition module is used to acquire data from multiple sources; The first determining module is used to determine the current audio preference analysis result and the vehicle's scene mode based on the multi-source data; The second determining module is used to determine the current user profile based on the current audio preference analysis result and the historical audio preference analysis results in the memory bank; The matching module is used to determine the recommended audio based on the current user profile and the scene mode.
12. A vehicle, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the scene and audio matching association method for a vehicle as described in any one of claims 1-10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the vehicle scene and audio matching association method as described in any one of claims 1-10.