Content recommendation method and device, vehicle and storage medium

By obtaining and analyzing the multimodal data of the target user in the vehicle, identifying their emotional state and personality characteristics, and providing personalized content recommendations, the problem of insufficient timeliness and accuracy of the recommended content of the existing vehicle computer system is solved, and the user experience and the intelligence of the recommendation system is improved.

CN120045728APending Publication Date: 2025-05-27GREAT WALL MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510123900.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When recommending content, existing vehicle and computer systems rely on fixed algorithms, resulting in insufficient timeliness and accuracy of recommended content, affecting user experience.

Method used

By obtaining modal data of multiple modals associated with the target user in the vehicle, the target user's current emotional state and personality characteristics are identified, and the recommended content matching the target user is determined based on these characteristics, and recommended to the user.

Benefits of technology

It improves the timeliness and accuracy of recommended content, enhances the user experience and intelligence of vehicle recommendation systems, and ensures the freshness and diversity of recommended content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045728A_ABST
    Figure CN120045728A_ABST
Patent Text Reader

Abstract

The invention provides a content recommendation method and device, a vehicle and a storage medium, the method is applied to the field of vehicles, and the method comprises the following steps: obtaining modal data of a plurality of modals associated with a target user in the vehicle; according to the modal data of the plurality of modals, identifying the current emotional state and character characteristics of the target user; according to the current emotional state and character characteristics of the target user, target recommendation content matched with the target user is determined, and the target recommendation content is recommended to the target user. According to the method, personalized recommendation can be carried out according to the emotion and character of the user, and the timeliness and accuracy of recommended content are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicles, and more specifically, to a method, device, vehicle, and storage medium for content recommendation in the field of vehicles. Background Art

[0002] With the development of in-vehicle intelligent technology, in-vehicle entertainment systems have become an important part of enhancing the driving experience. However, most existing in-vehicle systems rely on fixed algorithms for content recommendation, resulting in insufficient timeliness and accuracy of the recommended content, thus affecting the user experience. Summary of the Invention

[0003] The present application provides a method, device, vehicle, and storage medium for content recommendation. This method can perform personalized recommendation based on the emotions and personality of the user, improving the timeliness and accuracy of the recommended content.

[0004] In a first aspect, a method for content recommendation is provided. The method includes: obtaining modal data of multiple modalities associated with a target user in the vehicle; identifying the current emotional state and personality characteristics of the target user according to the modal data of the multiple modalities; determining target recommended content matching the target user according to the current emotional state and personality characteristics of the target user, and recommending the target recommended content to the target user.

[0005] In the above technical solution, by obtaining modal data of multiple modalities associated with a target user in the vehicle, it is more accurate and comprehensive compared to single-modal data. According to the modal data of multiple modalities, the current state of the target user can be comprehensively measured. Therefore, based on the modal data of multiple modalities, it is beneficial to accurately identify the emotional characteristics and personality characteristics of the user. Content recommendation based on the emotional characteristics and personality characteristics makes the recommended content conform to the current emotions and personality of the target user, improving the satisfaction of the target user. This method improves the user experience and the intelligence of the vehicle's recommendation system, ensuring the accuracy and timeliness of the recommended content.

[0006] Combined with the first aspect and the above implementation manners, in some possible implementation manners, the determining of the target recommended content that matches the target user according to the current emotional state and personality characteristics of the target user includes: determining a user group similar to the target user according to the current emotional state and personality characteristics of the target user; obtaining a set of multimedia content preferred by the user group, and screening out the multimedia content that the target user has not watched in the set of multimedia content as the first type of multimedia content; screening out the multimedia content that matches the current emotional state and personality characteristics of the target user in a preset multimedia content library as the second type of multimedia content; and determining the target recommended content that matches the target user according to the first type of multimedia content and the second type of multimedia content.

[0007] In the above technical solution, the content that the target user has not watched is screened out from the set of multimedia content preferred by the user group as the first type of multimedia content, ensuring that the recommended content has a strong sense of freshness for the target user. The content that matches the current emotions and personality characteristics of the target user is screened out in a preset multimedia content library as the second type of multimedia content, further expanding the source of the recommended content and increasing the diversity. According to the first type of multimedia content and the second type of multimedia content, the target recommended content that matches the target user is determined, which can capture more potential interest points and avoid recommendation biases caused by insufficient single-user data. By combining the preferences of the user group and the preset content library, it is ensured that the recommended content includes both highly relevant content verified by the user group and newly discovered matching content, balancing the diversity and relevance of the recommendation. The extensive preferences from the user group are introduced to avoid the information cocoon effect caused by over-reliance on personal historical behaviors, enabling the target user to access more diverse multimedia content.

[0008] Combined with the first aspect and the above implementation manners, in some possible implementation manners, the determining of the target recommended content that matches the target user according to the first type of multimedia content and the second type of multimedia content includes: extracting the content features of each multimedia content in the first type of multimedia content and the second type of multimedia content; determining the historical multimedia content with the content features in the historical multimedia content played by the vehicle; constructing an interaction matrix according to the historical multimedia content with the content features; where the interaction matrix is used to record the interest values of historical users in the vehicle for the content features; determining the target content features that match the target user according to the interaction matrix; and screening out the multimedia content with the target content features in the first type of multimedia content and the second type of multimedia content as the target recommended content.

[0009] In the above implementation, by extracting content features (such as emotion, style, creator, type, etc.) from the first type of multimedia content and the second type of multimedia content, the characteristics of the first type of multimedia content and the second type of multimedia content can be captured more comprehensively, providing a reference for subsequent personalized recommendations. Among the historical multimedia content, determine the historical multimedia content with content features to ensure that the recommended content conforms to the historical preferences of historical users. By constructing an interaction matrix to record the interest values of historical users for different content features, the interest preferences of historical users for different content features can be captured, which is beneficial to improving the attractiveness of the finally determined target recommended content to the target user.

[0010] Combined with the first aspect and the above implementation, in some possible implementations, the recommending the target recommended content to the target user includes: determining the current driving state of the vehicle; predicting the driving risk level corresponding to the target recommended content according to the driving state and the content features of the target recommended content; if the driving risk level is less than or equal to the preset level, recommending the target recommended content to the target user; if the driving risk level is greater than the preset level, reminding the driver of the vehicle to change the driving state to reduce the driving risk level of the target recommended content for the vehicle, and recommending the target recommended content to the target user when it is determined that the reduced driving risk level is less than the preset level.

[0011] In the above implementation, by real-time monitoring the driving state of the vehicle and predicting the driving risk level corresponding to the target recommended content according to the driving state and the content features of the target recommended content, and determining whether to directly recommend or recommend after reminding based on the driving risk level, it is beneficial to ensure that the recommended content will not have a negative impact on driving, improving the safety and stability of driving. By reminding the driver to adjust the driving state when the driving risk level is greater than the preset level, the risk of accidents caused by distracted driving due to recommending the target recommended content is effectively reduced, and the appropriate target recommended content can be recommended to the target user on the premise of ensuring driving safety. This method not only improves the safety of driving, but also optimizes the user experience, and through dynamic adjustment and continuous improvement, ensures the intelligence and reliability of content recommendation.

[0012] Combined with the first aspect, in some possible implementations, the identifying the current emotional state and personality characteristics of the target user according to the modal data of the multiple modalities includes: extracting the feature vectors of the target user in each modality according to the modal data of the multiple modalities; fusing the feature vectors of the target user in each modality to obtain fused data; and identifying the current emotional state and personality characteristics of the target user according to the fused data.

[0013] In the above implementation manner, by combining the modal data of multiple modalities, it is possible to more comprehensively capture the user's state information. Through reasonable feature fusion, the synergy between different modalities is promoted, enabling the advantages of each modality to be fully exerted, and improving the accuracy of emotion state and personality trait recognition.

[0014] Combined with the first aspect and the above implementation manner, in some possible implementation manners, the fusing of the feature vectors of the target user in each modality to obtain fusion data includes: transforming the feature vectors of the target user in each modality according to a target dimension to obtain the transformed feature vectors of each modality; performing normalization processing on the transformed feature vectors of each modality to obtain the normalized feature vectors of each modality; and performing weighted processing on the normalized feature vectors of each modality according to the weights corresponding to each modality to obtain the fusion data.

[0015] In the above implementation manner, by transforming the feature vectors of each modality according to the target dimension, it is ensured that the feature vectors of different modalities are compared and fused in the same dimension, enhancing the consistency and comparability of feature representation. Performing normalization processing on the feature vectors of each modality eliminates the influence of the numerical range difference between different modalities, avoiding a certain modality from dominating the fusion result due to a large numerical value. The normalization processing helps to eliminate the scale difference between different modalities, making each modality equally important in the fusion process, and reducing potential biases. Performing weighted processing according to the weights corresponding to each modality can highlight the information of important modalities, while suppressing the influence of noise or irrelevant modalities, improving the overall quality of the fusion data. The fusion data contains comprehensive information from multiple modalities, can more comprehensively describe the characteristics of the target user, reduces the information loss that may be brought by a single modality, provides a richer feature representation, thereby improving the accuracy of emotion recognition and personality recognition, and further improving the accuracy of content recommendation, making the final target recommended content conform to the current emotion state and personality characteristics of the target user, and improving the satisfaction of the target user with the target recommended content.

[0016] Combined with the first aspect and the above implementation manner, in some possible implementation manners, the modal data of the multiple modalities includes any of the following combinations: the visual modal data of the target user, the voice modal data of the target user, the text modal data of the target user, the physiological modal data of the target user.

[0017] In a second aspect, a content recommendation device is provided. The device includes: an acquisition module configured to acquire multimodal data associated with a target user in a vehicle; an identification module configured to identify the current emotional state and personality traits of the target user according to the multimodal data of the multiple modalities; and a recommendation module configured to determine target recommended content matching the target user according to the current emotional state and personality traits of the target user, and recommend the target recommended content to the target user.

[0018] In a third aspect, a vehicle is provided, including: a memory configured to store executable program code; and a processor configured to call and run the executable program code from the memory, so that the vehicle executes the method in the first aspect or any possible implementation manner of the first aspect.

[0019] In a fourth aspect, a computer program product is provided. The computer program product includes computer program code which, when running on a computer, causes the computer to execute the method in the first aspect or any possible implementation manner of the first aspect.

[0020] In a fifth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer program code which, when running on a computer, causes the computer to execute the method in the first aspect or any possible implementation manner of the first aspect. Description of the Drawings

[0021] Figure 1 is a schematic flowchart of a content recommendation method provided by an embodiment of the present application;

[0022] Figure 2 is a schematic flowchart of an implementation manner of step 102 provided by an embodiment of the present application;

[0023] Figure 3 is a schematic structural diagram of a content recommendation device provided by an embodiment of the present application;

[0024] Figure 4 is a schematic structural diagram of a vehicle provided by an embodiment of the present application. Detailed Embodiments

[0025] The technical solutions in the present application will be clearly and elaborately described below in conjunction with the accompanying drawings. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B. The "and / or" in the text is merely a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0026] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0027] With the development of intelligent vehicle technologies, in-vehicle entertainment systems have become an important part of enhancing the driving experience. However, most of the existing in-vehicle systems rely on fixed algorithms for content recommendation, resulting in insufficient timeliness and accuracy of the recommended content, thereby affecting the user experience. Based on this, this embodiment provides a content recommendation method, which is applied to a vehicle and can specifically be applied to the in-vehicle system in the vehicle. This method can perform personalized recommendations according to the user's mood and personality, and has a good recommendation effect, which is beneficial to improving the user experience.

[0028] Figure 1 It is a schematic flowchart of a content recommendation method provided by an embodiment of the present application.

[0029] Exemplarily, as Figure 1 shown, the method includes:

[0030] Step 101: Obtain multi-modal data of multiple modalities associated with the target user in the vehicle;

[0031] Step 102: Identify the current emotional state and personality characteristics of the target user according to the multi-modal data of multiple modalities;

[0032] Step 103: Determine the target recommended content that matches the target user according to the current emotional state and personality characteristics of the target user, and recommend the target recommended content to the target user.

[0033] In Figure 1In the illustrated embodiment, by obtaining multi-modal modality data associated with a target user in a vehicle, it is more accurate and comprehensive than single-modal data. Based on the multi-modal modality data, the current state of the target user can be comprehensively measured. Therefore, based on the multi-modal modality data, it is beneficial to accurately identify the user's emotional state and personality characteristics. Content recommendation is performed according to the emotional characteristics and personality characteristics, so that the recommended content conforms to the current emotions and personality of the target user, improving the satisfaction of the target user. This method improves the user experience and the intelligence of the vehicle's recommendation system, ensuring the accuracy and timeliness of the recommended content. Among them, timeliness mainly refers to that the recommended content conforms to the emotional state and personality characteristics of the target user at the current moment.

[0034] The following Figure 1 illustrates the specific implementation methods of each step in the illustrated embodiment:

[0035] In step 101, the target user in the vehicle can be the driver or other passengers, and this embodiment does not make specific limitations on this. The multi-modal modality data are the feature data of the target user in multiple dimensions. The multiple modalities may include: visual modality, sound modality, physiological modality, behavioral modality, text modality, etc. The data of multiple modalities can reflect the personality characteristics and emotional state of the target user.

[0036] Exemplarily, the multi-modal modality data includes any of the following combinations: the visual modality data of the target user, the sound modality data of the target user, the text modality data of the target user, the physiological modality data of the target user.

[0037] The visual modality data of the target user refers to: the image data of the target user collected by the in-vehicle camera, and the image data may include: data such as the facial expression, head posture, and behavioral actions of the target user.

[0038] The sound modality data of the target user refers to: the voice data related to the target user collected by the in-vehicle microphone, and the voice data may include: the speaking voice data of the target user and the ambient sound data of the environment where the target user is located, etc.

[0039] The text modality data of the target user refers to: the text data related to the target user obtained through the in-vehicle head unit system, and the text data may include: the text data input by the target user through the head unit system and the relevant text data obtained from the social media account of the target user when the head unit system has permission.

[0040] The physiological modal data of the target user refers to: the physiological data of the target user collected by physiological sensors such as a heart rate monitor, a skin conductivity sensor, and a temperature sensor in the vehicle. Among them, the heart rate monitor is used to monitor the heart rate changes of the target user. The skin electroresponse sensor is used to measure the skin conductivity of the target user, and the skin conductivity can reflect the emotional state to a certain extent. The temperature sensor is used to monitor the body temperature changes of the user.

[0041] Exemplarily, in the case where the target user is a driver, the modal data of multiple modalities may further include the modal data of the driving behavior modality, and the modal data of the driving behavior modality may include the usage data of the steering wheel, the accelerator, the brake, etc.

[0042] In step 102, based on the modal data of multiple modalities, a personality analysis and an emotion analysis are performed on the target user to obtain the personality characteristics and emotion characteristics of the target user. The emotion characteristics may be: positive emotion, negative emotion, neutral emotion, etc., or may be: happy, sad, melancholy, angry, etc. The personality characteristics may be, for example: openness, extraversion, conscientiousness, and neuroticism, etc.

[0043] Exemplarily, Figure 2 is a schematic diagram of the implementation manner of step 102 in this embodiment. As Figure 2 shown, the implementation manner of the above step 102 includes the following steps 1021 to 1023:

[0044] Step 1021: Based on the modal data of multiple modalities, extract the feature vectors of the target user in each modality.

[0045] Specifically, the modal data of multiple modalities can be preprocessed first, and then the feature vectors of the target user in each modality are extracted from the preprocessed modal data. Among them, the main purpose of the preprocessing is to ensure that the data of all modalities are aligned in time and to ensure that the data involving three-dimensional spatial information (such as facial expressions) are in the same spatial coordinate system. Optionally, the preprocessing method also includes using a filtering algorithm to remove the noise in the data and improve the data quality.

[0046] Exemplarily, different preprocessing methods can be adopted for the data of different modalities. The following are introduced separately:

[0047] For visual modal data, the visual modal data is usually image data, and the preprocessing methods for the image data include: graying, normalizing, and denoising the image data, etc., to obtain the preprocessed visual modal data.

[0048] For voice modality data, which is usually audio data, the ways to preprocess the audio data include: denoising the audio data, removing punctuation marks, numbers, stop words, etc., and then using speech recognition technology to convert the audio data into text data. The preprocessed voice modality data can include: the denoised audio data, and can also include the text data converted from the audio data.

[0049] For physiological modality data, which is usually physiological data such as heart rate and skin conductivity, the ways to preprocess the physiological data include: filtering and smoothing the physiological data such as heart rate and skin conductivity to obtain the preprocessed physiological modality data.

[0050] The ways to preprocess text modality data include: removing irrelevant characters and formatting, such as removing stop words, punctuation marks (such as full stops, commas, question marks, etc.) to reduce noise; converting the text modality data into a unified format to ensure consistency.

[0051] The ways to preprocess the modality data of driving behavior modality include: normalizing and synchronizing the driving behavior data, and removing invalid or incorrect data points.

[0052] After preprocessing the modality data of each modality, extract the feature vectors of the target user in each modality from the preprocessed modality data.

[0053] For example, use a deep learning model to extract the key features of facial expressions from the preprocessed visual modality data to obtain the feature vector of the visual modality (abbreviated as feature vector 1).

[0054] Use methods such as Mel Frequency Cepstral Coefficients, fundamental frequency, and energy to extract emotional and voice features from the preprocessed voice modality data to obtain the feature vector of the voice modality (abbreviated as feature vector 2).

[0055] Use Natural Language Processing (NLP) to perform sentiment analysis on the preprocessed text modality data, and extract sentiment labels such as positive, negative, and neutral. Use a topic model (such as LDA, Latent Dirichlet Allocation) to extract the topic distribution of the above text modality data to understand the topics that the target user is concerned about. Use a personality trait model (such as the Big Five personality model, MBTI, the Seven Personality model, etc.) to analyze the text modality data to obtain the personality labels of the target user. Combine the extracted sentiment labels of the target user, the topic distribution of the topics that the target user is concerned about, and the personality labels of the target user to obtain the feature vector of the text modality (abbreviated as feature vector 3).

[0056] Using time series analysis methods (such as wavelet transform, Fourier transform, etc.), extract features such as heart rate variability (HRV) and galvanic skin response (GSR) from the preprocessed physiological modality data to obtain the feature vector of the physiological modality (abbreviated as feature vector 4). High HRV and low GSR usually indicate that the target user is in a relaxed and calm state, with relatively stable emotions and has a steady or easy-going personality trait. Low HRV and high GSR usually indicate that the target user is in a tense and anxious state, with large emotional fluctuations and may have a neurotic or sensitive personality trait.

[0057] Use statistical methods to extract features such as acceleration change rate, steering angle, average speed, number of emergency brakes, and lane change frequency from the modality data of the driving behavior modality to obtain the feature vector of the driving behavior modality (abbreviated as feature vector 5). Sudden acceleration, sudden deceleration, large steering angles, frequent emergency brakes, and high lane change frequencies may indicate that the driver is in an anxious or angry emotional state and has an impulsive or neurotic personality trait. Smooth acceleration, small steering angles, fewer emergency brakes, and low lane change frequencies usually reflect that the driver is relatively calm, patient, and steady, and has a stable or easy-going personality trait.

[0058] Step 1022: Fuse the feature vectors of the target user in each modality to obtain fused data.

[0059] Among them, feature fusion means: combining data from different sensors (such as cameras, microphones, physiological sensors, etc.) to obtain more accurate and comprehensive information than single-modality data. The way of this feature fusion can be a two-stage fusion method, and this two-stage fusion method includes: early fusion and late fusion. Early fusion means: converting the feature vectors of each modality (such as the above-mentioned feature vectors 1 to 5) to the same dimension, and then performing normalization processing. Late fusion means: performing weighted summation on the feature vectors of each modality after early fusion to obtain the final fused data. It should be noted that the above example only takes the feature vectors of each modality including the above-mentioned feature vectors 1 to 5 as an example. In specific implementation, the feature vectors of each modality can also include any combination of the above-mentioned feature vectors 1 to 5.

[0060] In a possible implementation, the implementation of the above step 1022 includes the following S11 to S13:

[0061] S11: Convert the feature vectors of the target user in each modality according to the target dimension to obtain the converted feature vectors of each modality.

[0062] Specifically, dimensionality reduction methods such as PCA (Principal Component Analysis), t-SNE (t-distributed Stochastic Neighbor Embedding), and UMAP (Uniform Manifold Approximation and Projection), or autoencoders are used to convert the dimensionality of the feature vectors of each modality to the target dimension, so that the dimensionality of the transformed feature vectors of each modality is the same and is the target dimension. In specific implementation, the feature vectors of each modality can be transformed by padding or truncating to ensure that the feature vectors of each modality have the same length, that is, dimension.

[0063] S12: Normalize the transformed feature vectors of each modality to obtain the normalized feature vectors of each modality.

[0064] Specifically, the transformed feature vectors of each modality can be normalized by linear scaling to linearly transform each eigenvalue to a specific range. This specific range can be preset, for example, it can be [0, 1] or [-1, 1], etc. However, this embodiment does not make specific limitations in this regard.

[0065] The above S11 to S12 can be understood as the way of early fusion, and the way of late fusion can be seen in the following S13.

[0066] S13: According to the weights corresponding to each modality, perform weighted processing on the normalized feature vectors of each modality to obtain the fused data.

[0067] Among them, the weights corresponding to each modality can be preset, and the sum of the weights corresponding to each modality is 1. For example, the weight corresponding to the sound modality is 0.4, the weight corresponding to the text modality is 0.3, the weight corresponding to the visual modality is 0.1, and the weight corresponding to the physiological modality is 0.2.

[0068] Specifically, the fused data can be obtained through weighted processing by the following formula:

[0069]

[0070] where, N represents the total number of modalities, w i represents the weight of the i-th modality, and feature i represents the normalized feature vector of the i-th modality.

[0071] In the above implementation method, by transforming the feature vectors of each modality according to the target dimension, it is ensured that the feature vectors of different modalities are compared and fused under the same dimension, enhancing the consistency and comparability of feature representations. Normalizing the feature vectors of each modality eliminates the influence of the numerical range differences between different modalities, avoiding a certain modality dominating the fusion result due to a large numerical value. The normalization process helps to eliminate the scale differences between different modalities, making each modality equally important in the fusion process and reducing potential biases. Weighting according to the weights corresponding to each modality can highlight the information of important modalities while suppressing the influence of noise or irrelevant modalities, improving the overall quality of the fused data. The fused data contains comprehensive information from multiple modalities, can more comprehensively describe the characteristics of the target user, reduces the information loss that may be brought by a single modality, provides a richer feature representation, thereby improving the accuracy of emotion recognition and personality recognition, and further improving the accuracy of content recommendation, making the final target recommended content conform to the current emotional state and personality characteristics of the target user, and improving the satisfaction of the target user with the target recommended content.

[0072] Step 1023: Identify the current emotional state and personality characteristics of the target user according to the fused data.

[0073] Specifically, based on the fused data, an emotion recognition model can be used to identify the current emotional state of the target user, and a personality recognition model can be used to identify the personality characteristics of the target user. Combining the historical modality data and feedback data of the target user can also continuously optimize the emotion recognition model and the personality recognition model to continuously improve the recognition accuracy.

[0074] Considering that the recognition of emotional states and the recognition of personality characteristics are essentially classification problems, and SVM (Support Vector Machine) is suitable for classification problems, especially performing well when the feature dimension is high, and Random Forest is also suitable for classification problems and has strong anti-overfitting ability. Therefore, methods such as support vector machines or random forests can be used to extract the current emotional state and personality characteristics of the target user from the fused data.

[0075] In step 103, according to the current emotional state and personality characteristics of the target user, determine the target recommended content that matches the target user, and recommend the target recommended content to the target user.

[0076] Among them, the target recommended content is multimedia content that matches the current emotional state and personality traits of the target user, and the multimedia content can be music, video, news, blog, movie, etc. If the target recommended content is music, the target recommended content can be played through the music playback application in the in-vehicle system to recommend the target recommended content to the target user. If the target recommended content is video, the target recommended content can be played through the video playback application in the in-vehicle system to recommend the target recommended content to the target user.

[0077] For example, if the emotional state of the target user is sad and the personality trait of the target user is openness, the target recommended content can be: novel and relaxing music or video, etc. Neuroticism usually manifests as traits such as emotional instability, anxiety proneness, and sensitivity. If the emotional state of the target user is sad and the personality trait is neuroticism, the selection of the recommended content should aim to help the target user relieve negative emotions, provide emotional support, and avoid further exacerbating their uneasiness or anxiety. If the emotional state of the target user is sad and the personality trait of the target user is neuroticism, the target recommended content can be: music or video that soothes and stabilizes emotions.

[0078] In a specific implementation, a multimedia content library is preset. The multimedia content library includes several multimedia contents, and each multimedia content is labeled with the adapted emotional state and personality traits, indicating which emotional state and personality trait users this multimedia content is suitable for recommendation to. Thus, the current emotional state and personality traits of the target user can be matched with the emotional states and personality traits labeled by several multimedia contents in the multimedia content library, and the multimedia content with successful matching is determined as the target recommended content that matches the target user. Among them, the emotional state and personality traits labeled by the multimedia content with successful matching are the same as or the similarity is greater than the preset similarity threshold with the current emotional state and personality traits of the target user.

[0079] Exemplarily, the implementation manner of the above step 103 includes S21 to S24 as follows:

[0080] S21: Determine the user group similar to the target user according to the current emotional state and personality traits of the target user.

[0081] Among them, the user group similar to the target user refers to: the user group with similar or the same emotional state and personality traits as the target user. For example, assuming that the emotional state of the target user is sad and the personality trait is openness, the emotional state of the user group similar to the target user is sad and the personality trait is openness.

[0082] S22: Obtain the set of multimedia contents preferred by the user group, and screen out the multimedia contents that the target user has not watched in the set of multimedia contents as the first type of multimedia contents.

[0083] Among them, the multimedia content set consists of the multimedia content preferred by each user in the user group. The multimedia content preferred by the user is the multimedia content that the user likes. The multimedia content that the user likes is usually the multimedia content for which the user has performed operations such as giving a like, adding to favorites, or watching for a long time.

[0084] After obtaining the multimedia content set preferred by the user group, in this multimedia content set, filter out the multimedia content that the target user has not watched, and use the filtered multimedia content that the target user has not watched as the first type of multimedia content. This unwatched multimedia content is also the multimedia content that the target user has not come into contact with, and can also be understood as: the multimedia content that has not been played in the vehicle where the target user is currently located. The filtering method in S22 can be understood as: using collaborative filtering to filter out the first type of multimedia content.

[0085] S23: In the preset multimedia content library, filter out the multimedia content that matches the current emotional state and personality characteristics of the target user as the second type of multimedia content.

[0086] Among them, the multimedia content library includes a number of multimedia contents, and each multimedia content is labeled with the adapted emotional state and personality characteristics, indicating which kind of users with what emotional state and personality characteristics this multimedia content is suitable for recommendation. Thus, the current emotional state and personality characteristics of the target user can be matched with the emotional state and personality characteristics labeled in a number of multimedia contents in the multimedia content library, and the successfully matched multimedia content is used as the filtered multimedia content that matches the current emotional state and personality characteristics of the target user, that is, the second type of multimedia content. The filtering method in S23 can be understood as: using content filtering to filter out the second type of multimedia content.

[0087] S24: Determine the target recommended content that matches the target user according to the first type of multimedia content and the second type of multimedia content.

[0088] In this step, combine the multimedia content obtained by collaborative filtering (the multimedia content that the target user has not watched, that is, the first type of multimedia content) and the multimedia content obtained by content filtering (the multimedia content filtered in the multimedia content library, that is, the second type of multimedia content) to determine the target recommended content that matches the target user. That is to say, use the first type of multimedia content and the second type of multimedia content as the content to be recommended this time, and further filter out the target recommended content that matches the target user from the content to be recommended this time.

[0089] In a possible implementation manner, the implementation manner of the above S24 includes: determining the content preference of the target user, and then, according to the content preference of the target user, sorting the first type of multimedia content and the second type of multimedia content to obtain the top N multimedia contents in the sorting, and determining the top N multimedia contents in the sorting as the target recommended content matching the target user. Wherein, N is a natural number greater than or equal to 1. The content preference of the target user refers to the content characteristics of the multimedia content liked by the target user.

[0090] Exemplarily, the manner of sorting the first type of multimedia content and the second type of multimedia content according to the content preference of the target user may include: calculating the similarity between each multimedia content in the first type of multimedia content and the content preference of the target user, calculating the similarity between each multimedia content in the second type of multimedia content and the content preference of the target user, and sorting all the multimedia contents in the first type of multimedia content and the second type of multimedia content in descending order of the similarity to the content preference of the target user.

[0091] Exemplarily, the manner of sorting the first type of multimedia content and the second type of multimedia content according to the content preference of the target user may include: predicting the driving risk level corresponding to each multimedia content in the first type of multimedia content and the second type of multimedia content in the current driving state, and sorting each multimedia content in ascending order of the driving risk level. The lower the driving risk level corresponding to the multimedia content ranked higher, which is beneficial to reducing the impact of the recommended content on driving safety. The manner of predicting the risk level may include: determining the driving risk level corresponding to each multimedia content according to the driving state and the content characteristics of each multimedia content. Wherein, the combination of the driving state and the content characteristics has a corresponding relationship pre-calibrated with the driving risk level, so that the corresponding relationship can be queried according to the driving state and the content characteristics of each multimedia content to obtain the driving risk level corresponding to each multimedia content.

[0092] In the above technical solution, the content that the target user has not watched is screened out from the set of multimedia content preferred by the user group as the first type of multimedia content, ensuring that the recommended content has a strong sense of freshness for the target user. Content that matches the target user's current mood and personality characteristics is screened out from the preset multimedia content library as the second type of multimedia content, further expanding the source of the recommended content and increasing the diversity. Based on the first type of multimedia content and the second type of multimedia content, the target recommended content that matches the target user is determined, which can capture more potential interest points and avoid recommendation biases caused by insufficient single-user data. By combining the preferences of the user group and the preset content library, it is ensured that the recommended content includes both highly relevant content verified by the user group experience and newly discovered matching content, balancing the diversity and relevance of the recommendation. The broad preferences from the user group are introduced to avoid the information cocoon effect caused by over-reliance on personal historical behaviors, enabling the target user to be exposed to more diverse multimedia content.

[0093] In another possible implementation, the implementation of the above S24 includes the following S241 to S245:

[0094] S241: Extract the content features of each multimedia content in the first type of multimedia content and the second type of multimedia content.

[0095] Among them, the content features can include content features in multiple different dimensions, such as: content features in the emotional dimension (such as happy, sad, melancholy, angry, etc.), content features in the style dimension, content features in the creator dimension, content features in the type dimension (such as rock, classical, pop, etc.), etc. In S241, for each multimedia content in the first type of multimedia content and each multimedia content in the second type of multimedia content, the content features are extracted to obtain the content features of each multimedia content.

[0096] S242: In the historical multimedia content played by the vehicle, determine the historical multimedia content with the above content features.

[0097] Specifically, obtain the historical multimedia content played by the vehicle from the historical records of the vehicle's multimedia content playback system. Each historical multimedia content has its own content features. Based on this, the historical multimedia content with the content features extracted in S241 can be screened out from the historical multimedia content played by the vehicle.

[0098] S243: Construct an interaction matrix according to the historical multimedia content with the above content features.

[0099] Among them, the interaction matrix is used to record the interest values of historical users in the vehicle for the above content features. The historical user refers to a user who has ridden or driven the vehicle within a past historical time period. The interest value can be reflected by the number of clicks or the number of plays. The higher the number of clicks or the number of plays, the higher the interest value of the historical user in the multimedia content with this content feature. Each historical multimedia content has its own number of clicks or number of plays. Based on this, the interest values of historical users in the content features of historical multimedia content can be statistically obtained to obtain the interaction matrix.

[0100] In a possible implementation, if the content features extracted in S241 include content features of multiple dimensions, the interaction matrix can record the interest values of historical users in the vehicle for the combinations of content features of multiple dimensions. For example, the content features of multiple dimensions include: content features of the emotional dimension (such as happy, sad, melancholy, angry, etc.) and content features of the type dimension (such as rock, classical, pop, lyrical, etc.). On this basis, the form of the interaction matrix can be in the form of a two-dimensional table, and this two-dimensional table is used to record the interest values of historical users in the combinations of content features of the type dimension and the emotional dimension. The rows in this two-dimensional table represent a content feature of the emotional dimension (such as happy, sad, melancholy, angry, etc.), and the columns represent a content feature of the type dimension (such as rock, classical, pop, lyrical, etc.). The value corresponding to each row and column (such as C1 to C16 in Table 1) represents the interest value of the historical user in this type of content. Exemplarily, the form of this two-dimensional table can be as shown in Table 1:

[0101] Table 1

[0102] Rock Classical Pop Lyric Happy C1 C5 C9 C13 Sad C2 C6 C10 C14 Melancholy C3 C7 C11 C15 Angry C4 C8 C12 C16

[0103] It can be seen from Table 1 that the interest value of the historical user in the historical multimedia content with the content features of rock and happy is C1, and C1 can specifically be the number of clicks of the historical user on the historical multimedia content with the content features of rock and happy. The interest value of the historical user in the historical multimedia content with the content features of classical and sad is C6, and C6 can specifically be the number of clicks of the historical user on the historical multimedia content with the content features of classical and sad. The interest value of the historical user in the historical multimedia content with the content features of pop and melancholy is C11, and C11 can specifically be the number of clicks of the historical user on the historical multimedia content with the content features of pop and melancholy. The interest value of the historical user in the historical multimedia content with the content features of lyrical and angry is C16, and C16 can specifically be the number of clicks of the historical user on the historical multimedia content with the content features of lyrical and angry. The meanings of the remaining rows and columns in Table 1 can be obtained by referring to the above description, and will not be repeated here to avoid redundancy.

[0104] S244: Determine the target content features that match the target user according to the above interaction matrix.

[0105] Specifically, since the interaction matrix records the interest values of historical users for content features, based on this, the content features with the highest interest values can be determined as the target content features that match the target user. For example, referring to Table 1, if the highest value among C1 to C16 in Table 1 is C11, then the two content features of popularity and melancholy corresponding to C11 are determined as the target content features that match the target user.

[0106] In a possible implementation, when determining the target content features, the current driving state of the vehicle can also be combined. According to the driving state, invalid interest values and valid interest values are determined in the interaction matrix. According to the valid interest values, the target content features that match the target user are determined. When determining the target recommended content, the multimedia content with content features having invalid interest values will not be determined as the target recommended content, but the multimedia content with content features having valid interest values will be determined as the target recommended content.

[0107] Specifically, the driving risk level corresponding to each content feature can be determined according to the driving state and each content feature in the interaction matrix. Among them, there is a pre-calibrated corresponding relationship between the combination of the driving state and the content feature and the driving risk level. Therefore, the driving risk level corresponding to each content feature in the interaction matrix can be queried according to the driving state and each content feature in the interaction matrix. The interest values under the content features with a driving risk level greater than the preset level are determined as invalid interest values, and the interest values under the content features with a driving risk level less than or equal to the preset level are determined as valid interest values.

[0108] For example, in combination with Table 1, in the current driving state, if the vehicle speed exceeds 120 km / h, it is determined that the driving risk level corresponding to the content feature of rock in the interaction matrix is relatively high. Thus, it is determined that C1 to C4 under the column of rock are all invalid interest values, and C5 to C16 are all valid interest values. Furthermore, the content feature corresponding to the maximum valid interest value among C5 to C16 can be determined as the target content feature. If the maximum value among C5 to C16 is C11, it is determined that the target content features include: popularity and melancholy. Considering that multimedia content of the rock type may affect driving safety when the vehicle speed exceeds 120 km / h, by setting C1 to C4 under the column of rock as invalid interest values, it is possible to avoid recommending multimedia content of the rock type to the target user, thereby ensuring driving safety while making recommendations.

[0109] S245: Screen out the multimedia content with the target content features from the first type of multimedia content and the second type of multimedia content as the target recommended content.

[0110] It can be understood that in the above S241, the content features of each multimedia content in the first type of multimedia content and the second type of multimedia content are extracted. Based on this, the target content features can be compared with the content features of the first type of multimedia content and the second type of multimedia content to find the multimedia content that best matches the target content features as the target recommended content. Specifically, the multimedia content with the target content features can be screened out from the first type of multimedia content and the second type of multimedia content as the target recommended content. For example, if the target content features include popularity and melancholy, the multimedia content with the content features of these two dimensions of popularity and melancholy in the first type of multimedia content and the second type of multimedia content is used as the target recommended content.

[0111] In a specific implementation, if the number of multimedia content with the target content features screened out is multiple, then according to the content preferences of the target user, the screened multimedia content with the target content features can be sorted to obtain the top N multimedia content in the ranking, and the top N multimedia content in the ranking is determined as the target recommended content that matches the target user.

[0112] In the above implementation method, by extracting content features (such as emotions, styles, creators, types, etc.) from the first type of multimedia content and the second type of multimedia content, the characteristics of the first type of multimedia content and the second type of multimedia content can be captured more comprehensively, providing a reference for subsequent personalized recommendations. Among the historical multimedia content, the historical multimedia content with content features is determined to ensure that the recommended content conforms to the historical preferences of historical users. By constructing an interaction matrix to record the interest values of historical users for different content features, the interest preferences of historical users for different content features can be captured, which is beneficial to improving the attractiveness of the finally determined target recommended content to the target user.

[0113] In a possible implementation method, the above-mentioned recommending the target recommended content to the target user includes: determining the current driving state of the vehicle; predicting the driving risk level corresponding to the target recommended content according to the driving state and the content features of the target recommended content; if the driving risk level is less than or equal to the preset level, then recommending the target recommended content to the target user; if the driving risk level is greater than the preset level, then reminding the driver of the vehicle to change the driving state to reduce the driving risk level of the target recommended content for the vehicle, and when it is determined that the reduced driving risk level is less than the preset level, recommending the target recommended content to the target user.

[0114] Among them, the current driving state of the vehicle includes: the position information of the vehicle, the driving route, the speed information, the surrounding environment information of the driving route, etc. Under different driving states, the driving risk levels corresponding to the predicted target recommended content may vary. For example, in the low-speed driving state (the speed is lower than a certain threshold such as 30 km / h), the driving risk level corresponding to the target recommended content is a low risk level. In the medium-speed driving state (the speed is between two thresholds such as 30 km / h - 80 km / h) and the high-speed driving state (the speed is higher than a certain threshold such as 80 km / h), the driving risk level corresponding to the target recommended content can be determined by combining the content characteristics of the target recommended content. If the content characteristics of the target recommended content include the rock type, in the medium-high speed driving state, the content of the rock type may interfere with the driving safety of the driver. Therefore, in the medium-speed driving state, the driving risk level corresponding to the target recommended content can be a medium risk level, and in the high-speed driving state, the driving risk level corresponding to the target recommended content can be a high risk level.

[0115] In a possible implementation manner, predicting the driving risk level corresponding to the target recommended content according to the driving state and the content characteristics of the target recommended content can be understood as: predicting the driving risk level that will be caused if the target recommended content is recommended to the target user in the current driving state according to the driving state and the content characteristics of the target recommended content. In a specific implementation, there can be a pre-calibrated corresponding relationship between the combination of the driving state and the content characteristics and the driving risk level, so that the corresponding driving risk level can be queried according to the driving state and the content characteristics of the target recommended content.

[0116] The above preset level can be a medium risk level. If the driving risk level corresponding to the target recommended content is less than or equal to the preset level, it means that the risk caused by recommending the target recommended content to the target user is relatively small. Therefore, the target recommended content can be directly recommended to the target user. If the driving risk level corresponding to the target recommended content is greater than the preset level, it means that the risk caused by recommending the target recommended content to the target user is relatively large. Therefore, the target recommended content will not be directly recommended to the target user, but the driver of the vehicle will be reminded to change the driving state in a voice or dashboard display manner (such as suggesting the driver to slow down appropriately and enter a more stable driving state, suggesting the driver to choose a section with less traffic flow and better road conditions to continue driving, etc.) to reduce the driving risk level of the target recommended content for the vehicle. After the driver adjusts the driving state, the driving state of the vehicle can be re-detected, and the driving risk level can be re-predicted. If it is confirmed that the driving risk level has dropped below the preset level, the target recommended content can be recommended to the target user.

[0117] In the above implementation method, by monitoring the driving state of the vehicle in real time and predicting the driving risk level corresponding to the target recommended content based on the driving state and the content characteristics of the target recommended content, and determining whether to directly recommend or recommend after reminding based on the driving risk level, it is beneficial to ensure that the recommended content will not have a negative impact on driving, and improves the safety and stability of driving. By reminding the driver to adjust the driving state when the driving risk level is greater than the preset level, the risk of accidents caused by distracted driving due to recommending the target recommended content is effectively reduced, and the appropriate target recommended content can be recommended to the target user on the premise of ensuring driving safety. This method not only improves the driving safety, but also optimizes the user experience, and through dynamic adjustment and continuous improvement, ensures the intelligence and reliability of content recommendation.

[0118] Figure 3 It is a schematic structural diagram of a content recommendation device provided by an embodiment of the present application.

[0119] Exemplarily, as Figure 3 shown, the content recommendation device 300 includes: an acquisition module 301, configured to acquire multi-modal modal data associated with a target user in a vehicle; an identification module 302, configured to identify the current emotional state and personality characteristics of the target user according to the multi-modal modal data; a recommendation module 303, configured to determine target recommended content matching the target user according to the current emotional state and personality characteristics of the target user, and recommend the target recommended content to the target user.

[0120] In a possible implementation manner, the identification module 302 is specifically configured to: extract the feature vectors of the target user in each modality according to the multi-modal modal data; fuse the feature vectors of the target user in each modality to obtain fused data; and identify the current emotional state and personality characteristics of the target user according to the fused data.

[0121] In a possible implementation manner, the identification module 302 is specifically configured to: transform the feature vectors of the target user in each modality according to a target dimension to obtain the transformed feature vectors of each modality; perform normalization processing on the transformed feature vectors of each modality to obtain the normalized feature vectors of each modality; and perform weighted processing on the normalized feature vectors of each modality according to the weights corresponding to each modality to obtain the fused data.

[0122] In a possible implementation, the recommendation module 303 is specifically configured to: determine a user group similar to the target user according to the current emotional state and personality characteristics of the target user; obtain a set of multimedia content preferred by the user group, and screen out the multimedia content that the target user has not viewed as the first type of multimedia content from the set of multimedia content; screen out the multimedia content that matches the current emotional state and personality characteristics of the target user in a preset multimedia content library as the second type of multimedia content; and determine the target recommended content that matches the target user according to the first type of multimedia content and the second type of multimedia content.

[0123] In a possible implementation, the recommendation module 303 is specifically configured to: extract the content features of each multimedia content in the first type of multimedia content and the second type of multimedia content; determine the historical multimedia content with the content features in the historical multimedia content played by the vehicle; construct an interaction matrix according to the historical multimedia content with the content features, where the interaction matrix is used to record the interest values of the historical users in the vehicle for the content features; determine the target content features that match the target user according to the interaction matrix; and screen out the multimedia content with the target content features as the target recommended content from the first type of multimedia content and the second type of multimedia content.

[0124] In a possible implementation, the recommendation module 303 is specifically configured to: determine the current driving state of the vehicle; predict the driving risk level corresponding to the target recommended content according to the driving state and the content features of the target recommended content; if the driving risk level is less than or equal to a preset level, recommend the target recommended content to the target user; if the driving risk level is greater than the preset level, remind the driver of the vehicle to change the driving state to reduce the driving risk level of the target recommended content for the vehicle, and recommend the target recommended content to the target user when it is determined that the reduced driving risk level is less than the preset level.

[0125] In a possible implementation, the modal data of the multiple modalities includes any of the following combinations: visual modal data of the target user, voice modal data of the target user, text modal data of the target user, physiological modal data of the target user.

[0126] Figure 4 It is a schematic structural diagram of a vehicle provided by an embodiment of the present application.

[0127] Exemplarily, such as Figure 4As shown, the vehicle 400 includes: a memory 401 and a processor 402. Among them, an executable program code 4011 is stored in the memory 401, and the processor 402 is used to call and execute the executable program code 4011 to execute a content recommendation method.

[0128] In addition, an embodiment of the present application also protects a device, which may include a memory and a processor. Among them, an executable program code is stored in the memory, and the processor is used to call and execute the executable program code to execute a content recommendation method provided by the embodiment of the present application.

[0129] In this embodiment, the device can be divided into functional modules according to the above method example. For example, it can correspond to each functional module, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0130] In the case of dividing each functional module corresponding to each function, the device may further include an acquisition module, an identification module, a recommendation module, etc. It should be noted that all relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be repeated here.

[0131] It should be understood that the device provided in this embodiment is used to execute the above content recommendation method, so it can achieve the same effect as the above implementation method.

[0132] In the case of adopting an integrated unit, the device may include a processing module and a storage module. Among them, when the device is applied to a vehicle, the processing module can be used to control and manage the actions of the vehicle. The storage module can be used to support the vehicle to execute relevant program codes, etc.

[0133] Among them, the processing module can be a processor or a controller, which can implement or execute various exemplary logical blocks, modules and circuits shown in combination with the disclosure content of the present application. The processor can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc. The storage module can be a memory.

[0134] In addition, the device provided by the embodiment of the present application may specifically be a chip, a component or a module. The chip may include a connected processor and a memory; among them, the memory is used to store instructions, and when the processor calls and executes the instructions, the chip can execute a content recommendation method provided by the above embodiment.

[0135] This embodiment also provides a computer-readable storage medium, in which computer program code is stored. When the computer program code runs on a computer, the computer is enabled to execute the above-related method steps to implement a content recommendation method provided in the above embodiment.

[0136] This embodiment also provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute the above-related steps to implement a content recommendation method provided in the above embodiment.

[0137] Among them, the device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.

[0138] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0139] In the embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0140] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A method for content recommendation, characterized in that: The method comprises: acquiring modal data of a plurality of modalities associated with a target user in a vehicle; Identifying the current emotional state and personality characteristics of the target user based on the modal data of the multiple modalities; According to the current emotional state and personality characteristics of the target user, target recommended content matching the target user is determined, and the target recommended content is recommended to the target user.

2. The method according to claim 1, characterized in that The step of determining target recommended content matching the target user according to the current emotional state and personality characteristics of the target user includes: Determine a user group similar to the target user based on the target user's current emotional state and personality traits; Acquire a multimedia content set preferred by the user group, and select multimedia content that has not been viewed by the target user from the multimedia content set as the first category of multimedia content; In a preset multimedia content library, multimedia content matching the current emotional state and personality characteristics of the target user is selected as the second type of multimedia content; Determine target recommended content matching the target user according to the first category of multimedia content and the second category of multimedia content.

3. The method according to claim 2, characterized in that The determining, according to the first category of multimedia content and the second category of multimedia content, target recommended content matching the target user includes: extracting content features of each multimedia content in the first category of multimedia content and the second category of multimedia content; Determining historical multimedia content having the content feature among historical multimedia content played by the vehicle; Constructing an interaction matrix according to historical multimedia content having the content feature; wherein the interaction matrix is ​​used to record the interest values ​​of historical users in the vehicle for the content feature; Determining target content features matching the target user according to the interaction matrix; Among the first category of multimedia content and the second category of multimedia content, multimedia content having the target content feature is screened out as target recommended content.

4. The method according to claim 1, characterized in that: The recommending the target recommended content to the target user includes: Determining a current driving state of the vehicle; predicting a driving risk level corresponding to the target recommended content according to the driving state and content features of the target recommended content; If the driving risk level is less than or equal to a preset level, recommending the target recommended content to the target user; If the driving risk level is greater than the preset level, the driver of the vehicle is reminded to change the driving state to reduce the driving risk level of the target recommended content to the vehicle, and when it is determined that the reduced driving risk level is less than the preset level, the target recommended content is recommended to the target user.

5. The method according to claim 1, characterized in that: The identifying the current emotional state and personality characteristics of the target user according to the modal data of the multiple modalities includes: Extracting a feature vector of the target user in each modality according to the modality data of the multiple modalities; Fusing the feature vectors of the target user in each mode to obtain fused data; The current emotional state and personality traits of the target user are identified based on the fused data.

6. The method according to claim 5, characterized in that The step of fusing the feature vectors of the target user in each mode to obtain fused data includes: Transforming the feature vectors of the target user in each modality according to the target dimension to obtain transformed feature vectors of each modality; Normalizing the transformed feature vectors of each mode to obtain the normalized feature vectors of each mode; According to the weights corresponding to the various modes, the normalized feature vectors of the various modes are weighted to obtain the fused data.

7. The method according to any one of claims 1 to 6, characterized in that: The modal data of the multiple modalities include any combination of the following: Visual modal data of the target user, sound modal data of the target user, text modal data of the target user, and physiological modal data of the target user.

8. A content recommendation device, characterized in that: The device comprises: An acquisition module, configured to acquire modal data of multiple modalities associated with a target user in a vehicle; An identification module, configured to identify the current emotional state and personality characteristics of the target user based on the modal data of the multiple modalities; The recommendation module is used to determine target recommended content that matches the target user according to the current emotional state and personality characteristics of the target user, and recommend the target recommended content to the target user.

9. A vehicle, characterized in that: The vehicle comprises: A memory for storing executable program codes; A processor, configured to call and run the executable program code from the memory, so that the vehicle executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 is implemented.