Vehicle-mounted interaction method and device based on interpretable recommendation
By introducing an interpretable recommendation model into the in-vehicle recommendation system, which combines multiple data sources to generate recommendation results and provide explanations, the problem of insufficient transparency in traditional in-vehicle recommendation systems is solved, thereby improving user trust and recommendation accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-03-20
AI Technical Summary
Traditional in-vehicle recommendation systems lack transparency, making it difficult for users to understand the reasons behind the recommendation results, leading to a decrease in trust.
An interpretable recommendation model is adopted, which combines vehicle sensor data, user behavior data, multimedia data and historical driving data to generate recommendation results and provide explanations. The reasons for the generation are shown to users through multimodal explanation.
It improves the transparency of recommendation results and user trust, enhances user experience satisfaction, and achieves personalized and accurate recommendations.
Smart Images

Figure CN118824239B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a vehicle-mounted interaction method and device based on explainable recommendation. BACKGROUND
[0002] With the continuous development of artificial intelligence, speech recognition, natural language processing and other technologies, vehicle-mounted voice interaction systems have gradually become the mainstream interaction method for in-vehicle entertainment and information services. Vehicle-mounted voice interaction systems can convert user voice input into text through speech recognition technology and understand user intent through natural language processing technology. Subsequently, the system will perform corresponding operations according to the user's intent, such as playing music, navigation planning, weather query, etc., greatly improving the convenience during driving.
[0003] Under the background of the widespread application of vehicle-mounted voice interaction systems, recommendation systems, as a key link to improve user experience, have also gradually integrated and played an important role. However, traditional recommendation systems often adopt a "black box" operation mode, and the recommended results lack transparency, making it difficult for users to understand the logic and basis behind them, which to some extent reduces the trust of users in the recommendation system. SUMMARY
[0004] The present application provides a vehicle-mounted interaction method and device based on explainable recommendation to solve the defect that the recommended results lack explicit explanation and users cannot understand the reasons for generating the recommended results in related technologies.
[0005] The present application provides a vehicle-mounted interaction method based on explainable recommendation, comprising:
[0006] receiving user voice in the vehicle;
[0007] recognizing the user voice to obtain recognized text, and performing intent understanding on the recognized text to obtain intent information;
[0008] based on an explainable recommendation model, applying at least one of vehicle-mounted sensor data, user behavior data, multimedia data, historical driving data and the intent information to generate a recommended result and an explanation result;
[0009] based on the explanation result, applying a multi-modal explanation method to explain the reasons for generating the recommended result.
[0010] According to the vehicle-mounted interaction method based on explainable recommendation provided by the present application, the training steps of the explainable recommendation model include:
[0011] An initial model is acquired, and sample training data is collected, the sample training data including at least one of sample vehicle-mounted sensor data, sample user behavior data, sample multimedia data, sample user feedback data, and sample intent information and a sample recommendation result corresponding to the sample intent information;
[0012] Based on the initial model, at least one of sample vehicle-mounted sensor data, sample user behavior data, sample multimedia data, sample user feedback data and sample intent information in the sample training data is applied to generate a predicted recommendation result and a predicted explanation result;
[0013] Based on the difference between the predicted recommendation result and the sample recommendation result, a recommendation loss is determined, and the accuracy of the predicted explanation result and the user satisfaction are evaluated to obtain an explanation loss;
[0014] Based on the recommendation loss and the explanation loss, the initial model is iterated in parameters to obtain the interpretable recommendation model.
[0015] According to the car-mounted interaction method based on the interpretable recommendation provided by the application, the multi-modal explanation mode includes at least two of the voice explanation mode, the text explanation mode and the visual explanation mode.
[0016] According to the car-mounted interaction method based on the interpretable recommendation provided by the application, based on the explanation result, a multi-modal explanation mode is applied to explain the generation reason of the recommendation result, including:
[0017] In the case where the multi-modal explanation mode includes the voice explanation mode, voice synthesis is performed based on the explanation result to generate voice explanation, and the voice explanation is played to explain the generation reason of the recommendation result to the user;
[0018] In the case where the multi-modal explanation mode includes the text explanation mode, text explanation is generated based on the explanation result, and the text explanation is displayed on the head-up display or displayed on the car-mounted instrument panel to explain the generation reason of the recommendation result to the user;
[0019] In the case where the multi-modal explanation mode includes the visual explanation mode, visual explanation content is generated based on the explanation result, and the visual explanation content is displayed on the head-up display or displayed on the car-mounted instrument panel to explain the generation reason of the recommendation result to the user.
[0020] According to the car-mounted interaction method based on the interpretable recommendation provided by the application, based on the explanation result, a multi-modal explanation mode is applied to explain the generation reason of the recommendation result, including:
[0021] determine a target explanation mode from the multi-modal explanation modes based on the satisfaction evaluation result and / or the understanding evaluation result of each explanation mode in the multi-modal explanation modes;
[0022] based on the explanation result, applying the target explanation mode to explain the generation reason of the recommendation result.
[0023] According to the vehicle-mounted interaction method based on the interpretable recommendation provided by the application, further comprising:
[0024] collecting a user image in the vehicle and performing emotion recognition based on the user image to obtain an emotion recognition result, and / or receiving user feedback data, the user feedback data including at least one of recommendation result feedback data, explanation result feedback data, and explanation mode feedback data;
[0025] based on the emotion recognition result and / or the user feedback data, adjusting at least one of the recommendation result, the explanation result, and the multi-modal explanation mode.
[0026] The application further provides a vehicle-mounted interaction device based on interpretable recommendation, comprising:
[0027] a receiving unit configured to receive a user voice in the vehicle;
[0028] an identification unit configured to identify the user voice to obtain an identified text, and perform intent understanding on the identified text to obtain intent information;
[0029] a recommendation unit configured to generate a recommendation result and an explanation result based on an interpretable recommendation model, and apply at least one of vehicle-mounted sensor data, user behavior data, multimedia data, and historical driving data and the intent information;
[0030] an explanation unit configured to explain the generation reason of the recommendation result based on the explanation result and apply a multi-modal explanation mode.
[0031] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the vehicle-mounted interaction method based on interpretable recommendation according to any one of the above when executing the computer program.
[0032] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the vehicle-mounted interaction method based on interpretable recommendation according to any one of the above.
[0033] The application further provides a computer program product comprising a computer program which, when executed by a processor, implements the vehicle-mounted interaction method based on interpretable recommendation as described above.
[0034] The vehicle-mounted interaction method and device based on interpretable recommendation provided by the application can generate corresponding explanation results while generating recommendation results, and explain the reasons for generating the recommendation results based on the explanation results, so that users can better understand how the recommendation results are generated, increasing transparency and trust. The interpretable recommendation model can more accurately capture real-time needs and preferences of users by combining vehicle-mounted sensor data, user behavior data, multimedia data, historical driving data, and intent information, etc. multi-source information, to achieve more personalized recommendations. In addition, the multi-modal explanation method makes the explanation process more intuitive and easy to understand, meeting the needs of different users. Furthermore, the interpretable recommendation model considers various data sources and user intent information when generating recommendation results, which helps to reduce uncertainty in the recommendation process and improve the accuracy and relevance of the recommendations. BRIEF DESCRIPTION OF DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the application or related art, the following will briefly introduce the drawings needed to be used in the embodiments or related art descriptions. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0036] Figure 1 is a flowchart of the vehicle-mounted interaction method based on interpretable recommendation provided by the application;
[0037] Figure 2 is a flowchart of the training steps of the interpretable recommendation model provided by the application;
[0038] Figure 3 is a structural diagram of the vehicle-mounted interaction device based on interpretable recommendation provided by the application;
[0039] Figure 4 is a structural diagram of the electronic device provided by the application. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solutions and advantages of the application clearer, the technical solutions in the application will be described clearly and completely below in combination with the drawings in the application. Obviously, the described embodiments are some embodiments of the application, not all embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0041] With the rapid progress of technology, artificial intelligence, speech recognition and natural language processing technologies have been widely applied in the automotive industry, driving the innovation and development of in-vehicle voice interaction systems. This system not only greatly improves the convenience and safety during driving, but also provides users with more intelligent and personalized in-car entertainment and information service experiences through seamless voice interaction. The in-vehicle voice interaction system can capture and analyze user voice commands in real time, convert speech into text using advanced speech recognition technology, and then analyze user intentions and needs through complex natural language processing mechanisms. Based on these accurate understanding information, the system can quickly respond and perform corresponding operations, such as playing favorite music, planning travel routes for navigation, querying and broadcasting weather conditions in real time, etc., realizing highly intelligent interaction between the vehicle and the passenger.
[0042] Under the background of the wide application of in-vehicle voice interaction systems, the recommendation system, as a key link to improve user experience, has also gradually integrated and played an important role. However, traditional recommendation systems often adopt a "black box" operation mode, that is, the internal working mechanism of the system is not transparent to users. This opacity leads to a lack of explicit explanation of the recommended results, and users are difficult to understand why the system will recommend specific content or services, thereby causing a lack of trust in the recommendation system. In order to solve this problem, it is particularly important to build an interpretable in-vehicle voice interaction recommendation system.
[0043] To this end, the present application provides an in-vehicle interaction method based on interpretable recommendation, which can be applied to an interpretable in-vehicle voice interaction recommendation system. Not only does it inherit the convenience of in-vehicle voice interaction, but also makes the recommendation process more transparent by introducing an interpretable recommendation model. When the system makes recommendations according to user voice commands, it will provide clear explanations and reveal the reasons for generating recommended results, not only enhancing user trust in the recommendation system, but also improving user experience satisfaction, thereby overcoming the above-mentioned defects.
[0044] Figure 1 The present application provides a flowchart of an in-vehicle interaction method based on interpretable recommendation, as shown in Figure 1 The method comprises:
[0045] Step 110, receiving user voice in the vehicle;
[0046] Specifically, the method provided by the embodiments of the present application can be applied to an interpretable vehicle-mounted voice interaction recommendation system (hereinafter referred to as a vehicle-mounted system), which can be obtained by integrating an interpretable recommendation system with a vehicle-mounted intelligent voice assistant. For example, a user can inquire about the recommendation reason through voice, and the vehicle-mounted intelligent voice assistant can explain in a voice manner. The system can be equipped with a microphone or a microphone array composed of multiple microphones, through which the user voice in the vehicle can be received. Here, the user voice in the vehicle refers to the sound emitted by the driver or passenger in the vehicle cabin, which is used to interact with the system. When the user emits voice, the microphone converts the sound into an electrical signal and then transmits it to the voice processing module of the system.
[0047] In step 120, the user voice is recognized to obtain recognized text, and the recognized text is subjected to intent understanding to obtain intent information.
[0048] Specifically, after receiving the user voice, the voice recognition can be performed to obtain the recognized text. Here, voice recognition is a technology that converts human voice into text. Specifically, the vehicle-mounted system can transmit the received user voice signal to a voice recognition engine, which analyzes the acoustic characteristics of the voice signal, determines the words or phrases spoken by the user according to the acoustic characteristics, and converts them into text form, thereby obtaining the recognized text. Here, the recognized text is the output of the voice recognition process, i.e., the text form into which the user voice is converted. It contains the words, phrases or sentences spoken by the user, and is the basis for subsequent intent understanding and interaction.
[0049] Subsequently, the recognized text can be subjected to intent understanding to obtain intent information. Here, intent understanding refers to inferring the user's intent or requirement through semantic analysis of the recognized text to determine the information of the operation or query that the user wants to perform. It should be understood that the vehicle-mounted system can use NLP (Natural Language Processing) algorithms to perform word segmentation, part-of-speech tagging, syntax analysis, etc. on the recognized text to extract key information in the text, and input these key information into a pre-trained intent understanding model to determine the user's intent. Here, the intent understanding model can be a classifier or other machine learning model, such as CNN (Convolutional Neural Networks), LSTM (Long Short-Term Memory), Transformer model, etc., which is not specifically limited by the embodiments of the present application.
[0050] It can be understood that the above-mentioned intention information refers to a specific description of the user's intention or demand obtained through the intention understanding process. It can be one or more predefined intention labels, or a structured description containing the user's specific demand.
[0051] In step 130, based on the interpretable recommendation model, at least one of the vehicle-mounted sensor data, user behavior data, multimedia data, and historical driving data and the intention information are applied to generate a recommendation result and an explanation result.
[0052] Specifically, after obtaining the user's intention information, at least one of the vehicle-mounted sensor data, user behavior data, multimedia data, and historical driving data and the intention information can be input into the interpretable recommendation model to obtain the recommendation result and the explanation result output by the model. Here, the interpretable recommendation model refers to a recommendation system that can generate a recommendation result while providing the reasons or basis behind these recommendation results. The interpretable recommendation model aims to improve the transparency and user trust of the recommendation result. By providing explanations for recommendations, users can better understand the recommended content, thereby enhancing user experience and satisfaction.
[0053] The above-mentioned vehicle-mounted sensor data refers to the data collected by various sensors installed on the vehicle, such as speed sensors, acceleration sensors, gyroscopes, cameras, etc. These data can be used to analyze the driving state of the vehicle, environmental information, and driver state, etc., to provide real-time and dynamic context information for the recommendation system. User behavior data refers to records of activities and choices made by users in the vehicle or vehicle-mounted system, such as music playback records, navigation history, application usage, etc. These data reflect the user's preferences and habits and are an important basis for the recommendation system to understand user demand. Historical driving data includes the user's past driving behavior, route selection, speed pattern, etc. These data help the recommendation system understand the user's driving habits and possible driving needs, thereby providing more personalized recommendations.
[0054] The multimedia data refers to information data containing multiple media types, such as audio data, video data, image data, text data, etc. Among them, the audio data can include voice instructions input by the user through the vehicle-mounted voice assistant, as well as audio content such as music, audio books, navigation voice prompts, etc. played by the system. The video data can include video content played by the vehicle-mounted display screen, external environment video (such as a dashcam video) captured by the vehicle-mounted camera, etc. The image data can include static images (such as road signs, traffic conditions, etc.) captured by the vehicle-mounted camera, images (such as images of points of interest on a map) uploaded by the user or obtained by the vehicle-mounted system, etc. The text data can include text content input by the user (such as navigation target, search query, etc.) and text content generated by the vehicle-mounted system (such as navigation prompts, information notifications, etc.). In addition, the multimedia data can also include interaction data, such as touch screen interaction records of the user with the vehicle-mounted system, interaction records of the user with the system through voice or gestures, etc.
[0055] It can be understood that the interpretable recommendation model comprehensively processes and analyzes the above-mentioned multi-source data and intent information by combining various algorithms and technologies, such as machine learning, deep learning, rule engine, etc. Specifically, first, the obtained vehicle-mounted sensor data, user behavior data, multimedia data, historical driving data, etc. can be cleaned, integrated and feature extracted for subsequent model use; after obtaining the user's intent information, the pre-processed multi-source data and intent information are input into the interpretable recommendation model, and the model uses built-in algorithms and rules to comprehensively analyze these data and evaluate the applicability and rationality of different recommendation options; based on the results of model reasoning, the most suitable recommendation item is selected as the recommendation result, and at the same time, according to the internal logic and decision path of the model, the explanation of the recommendation result is generated, i.e. why these items are recommended. It should be understood that the recommendation result refers to the specific content or suggestion recommended by the interpretable recommendation model for the user, such as a music track, a navigation route, a driving mode, etc. The explanation result refers to a detailed explanation of the recommendation result, explaining why these specific contents are recommended. By providing the explanation result, the user can better understand the source and rationality of the recommended content, thereby improving the trust and satisfaction of the recommendation system.
[0056] For example, for the intent information that the user wants to play music, the interpretable recommendation model can combine the vehicle-mounted sensor data (such as location, speed, road conditions, etc.) to recommend appropriate music according to the real-time road conditions, improving the user experience. For example, for the intent information that the user wants to plan a navigation, the interpretable recommendation model can combine user behavior data, multimedia data and historical driving data to determine the user's preferred scenery or points of interest, and plan a route accordingly, so as to recommend a personalized route to the user.
[0057] At step 140, based on the interpretation result, a multi-modal interpretation method is applied to interpret the generation reason of the recommendation result.
[0058] Specifically, after generating the recommendation result and the interpretation result, the vehicle-mounted system can deeply analyze the interpretation result, extract key information points and logical chains, which can include the user's historical behavior, the current context, the logical rules of the recommendation algorithm, etc. According to the content and characteristics of the interpretation result, the system can select the most suitable interpretation mode or mode combination from the multi-modal interpretation method, for example, for user behavior-based recommendation, the user's behavior trajectory and preference distribution can be displayed in the form of text and charts; for the current context-based recommendation, the current driving environment and road condition information can be described in the form of voice and image. After determining the interpretation mode, the system can present the interpretation content to the user in the form of multi-modal through the display screen, audio and other hardware devices of the vehicle-mounted device. The user can receive information through various sensory channels such as vision and hearing, so as to better understand the generation reason of the recommendation result. It should be understood that the above multi-modal interpretation method refers to a method of using multiple different information presentation forms (such as text, voice, visualization, etc.) to explain and describe a phenomenon or result. The generation reason of the recommendation result refers to various factors and information that the interpretable recommendation model relies on when generating the recommendation content.
[0059] The method provided by the embodiment of the present application can generate an explanation result corresponding to the recommendation result by introducing an interpretable recommendation model, and explain the generation reason of the recommendation result based on the explanation result, so that the user can better understand how the recommendation result is generated, and the transparency and trust are increased. The interpretable recommendation model can more accurately capture the real-time needs and preferences of the user by combining vehicle-mounted sensor data, user behavior data, multimedia data, historical driving data, and intent information, etc. multi-source information, to achieve more personalized recommendation. At the same time, the multi-modal interpretation method makes the interpretation process more intuitive and easy to understand, and meets the needs of different users. In addition, the interpretable recommendation model considers various data sources and user intent information when generating the recommendation result, which helps to reduce the uncertainty in the recommendation process and improve the accuracy and relevance of the recommendation.
[0060] Based on the above embodiment, Figure 2 is a flowchart of the training steps of the interpretable recommendation model provided by the present application, as Figure 2 shown, the training steps of the interpretable recommendation model include:
[0061] At step 210, an initial model is obtained, and sample training data is collected, the sample training data including at least one of sample vehicle-mounted sensor data, sample user behavior data, sample multimedia data, sample user feedback data, and sample intent information and a sample recommendation result corresponding to the sample intent information;
[0062] Specifically, to build an interpretable recommendation model, the initial model can be selected as an attention-based model, which uses an attention mechanism to explain the recommendation result. With the attention mechanism, users can see the factors that the model focuses on when making a recommendation decision. For example, a model based on Self-Attention or Transformer can be used to let users know which factors the model focuses on when recommending a movie, such as the user's historical viewing records, the type and theme of the movie, the rating and evaluation of the movie, etc.
[0063] In addition to attention-based models, other methods can also be used to build interpretable recommendation models. For example, a method based on LRP (Layer-wise Relevance Propagation) can be used, which is a backpropagation algorithm that can calculate the influence of each input feature on the model output. A method based on SHAP (SHapley Additive exPlanations) can also be used, which is a Shapley value-based explanation method that can calculate the contribution of each feature to the model's prediction result. Methods based on LRP or SHAP can explain the recommendation result by calculating the contribution of each feature to the model's prediction result. In addition, GNN (Graph Neural Networks) can be used to build a recommendation model based on social networks, which can be used to explain how the model uses social relationships for recommendations. It should be understood that by introducing attention-based models and other interpretability methods into the recommendation system, not only can user experience and satisfaction be improved, but also the transparency and credibility of the model can be enhanced.
[0064] Before training the interpretable recommendation model, sample training data also needs to be collected, which can include at least one of sample vehicle sensor data, sample user behavior data, sample multimedia data, sample user feedback data, and sample intent information and sample recommendation results corresponding to the sample intent information. Here, sample user feedback data refers to direct or indirect feedback made by a user on the recommendation results given by a recommendation system. Such feedback can be explicit (such as user behaviors such as rating, liking, collecting, commenting on recommended content, etc.) or implicit (such as user dwell time, click rate, scroll depth, etc. on recommended content). It should be understood that in order to collect sample training data, data sources such as vehicle sensor data, user behavior logs, multimedia content information library, user feedback system, etc. can be determined first; then relevant data samples such as user's historical behavior records, vehicle sensor recorded driving data, user's rating or click records on recommended content, etc. are extracted from the data sources; then the extracted data is cleaned to remove duplicate, incorrect or invalid data, ensuring the accuracy and consistency of the data; finally, data from different sources is integrated into a unified data set for model training.
[0065] Step 220, based on the initial model, applying at least one of sample vehicle sensor data, sample user behavior data, sample multimedia data, sample user feedback data in the sample training data and sample intent information to generate a predicted recommendation result and a predicted explanation result;
[0066] Specifically, after collecting sample training data, a feature vector can be constructed from the sample training data. These features can include features of vehicle sensor data, features of user behavior data, features of multimedia data, and features of user feedback data. At the same time, the sample intent information also needs to be encoded into a format that the model can understand. Then, the preprocessed feature vector and sample intent information are input into the initial model, and the model performs forward propagation based on the input data to generate a predicted recommendation result. At the same time of generating the predicted recommendation result, the model also generates a corresponding predicted explanation result. The method of explanation generation depends on the design of the model's interpretability, such as attention, LRP, SHAP or GNN, etc.
[0067] Step 230, based on the difference between the predicted recommendation result and the sample recommendation result, determining a recommendation loss and evaluating the accuracy of the predicted explanation result and user satisfaction to obtain an explanation loss;
[0068] Specifically, the recommendation loss refers to the difference between the predicted recommendation result generated by the recommendation system and the sample recommendation result (i.e., the true recommendation result). This difference can be measured by a certain loss function, such as mean square error, cross-entropy loss, etc. The size of the recommendation loss reflects the strength of the prediction ability of the recommendation system, and is an important basis for optimizing the recommendation model.
[0069] The explanation loss refers to the loss used to measure the quality (including accuracy and user satisfaction) of the predicted explanation result in an explainable recommendation model. By optimizing the explanation loss, the explainability of the recommendation system can be improved, making it easier for users to understand and accept the recommendation result, thereby enhancing the credibility and user satisfaction of the recommendation system.
[0070] It can be understood that in order to obtain the explanation loss, the predicted explanation result output by the model can be evaluated, which includes two aspects: the accuracy of the explanation result, i.e., whether the model explanation result can reflect the actual recommendation decision-making process; and the user satisfaction of the explanation result, i.e., whether the user can understand the model explanation result. However, directly evaluating the user satisfaction of the explanation result often involves user research or subjective evaluation, which may be difficult to quantify directly in an automated system. In order to convert the evaluation of the predicted explanation result into a quantifiable explanation loss, some specific evaluation indicators can be designed, for example, the evaluation indicators can include the consistency between the explanation and the recommendation result (i.e., whether the explanation correctly reflects the recommendation logic), the conciseness of the explanation (i.e., whether the explanation is too long or difficult to understand), and the diversity of the explanation (i.e., whether multiple possible explanation paths are provided), etc. Through quantitative evaluation of these indicators, the explanation loss can be finally determined.
[0071] Specifically, the accuracy of the model explanation result can be evaluated by the following methods, for example, domain experts can review the explanation result to determine whether the explanation is reasonable and accurate, and the experts can analyze whether the explanation reflects the actual recommendation decision-making process by comparing the recommendation result and the input features. For example, the explanation result can be compared and analyzed with the actual recommendation logic, using known standard data sets and decision-making logic to verify the accuracy of the model explanation. For example, visualization tools can be used to display the decision-making process and explanation result of the model, and charts, heat maps, etc. can be used to help evaluate the rationality of the explanation, such as using SHAP value charts to show the contribution of each feature to the prediction result and checking its rationality. In addition, quantitative indicators such as explanation consistency and explanation reliability can be calculated to evaluate the stability and reliability of the explanation, for example, by comparing the consistency of the explanation results of similar inputs in multiple recommendations to evaluate the explanation stability of the model.
[0072] The user's satisfaction evaluation on the model explanation result can be achieved by one or more combinations of user survey, user experience test, user feedback analysis, behavior data analysis, sentiment analysis, etc. Among them, the user survey refers to designing a questionnaire to ask the user about the understanding degree and satisfaction of the explanation result. The questionnaire can include Likert scale, open-ended questions, etc. to quantify the user's satisfaction and understanding degree. The user experience test refers to letting the user experience the recommendation and explanation function in the actual use scene and collecting user feedback. The user feedback analysis refers to collecting and analyzing the feedback provided by the user during use, such as likes, comments, suggestions, etc. The behavior change of the user after using the explanation function is analyzed, such as the click rate after explanation, the recommendation acceptance rate, etc. The behavior data analysis refers to analyzing the interaction behavior data of the user and the explanation result, such as the duration of the user viewing the explanation result, the frequency of clicking detailed explanation, etc. These data can reflect the user's attention and understanding degree of the explanation result. The sentiment analysis refers to using sentiment analysis tools to analyze the sentiment tendency of the user feedback to judge the user's satisfaction with the explanation result, for example, by analyzing the sentiment polarity of the user's comments to evaluate the positive or negative sentiment of the user on the explanation result.
[0073] In step 240, based on the recommendation loss and the explanation loss, the initial model is iterated in parameters to obtain the interpretable recommendation model.
[0074] Specifically, after obtaining the recommendation loss and the explanation loss, a total loss value can be obtained based on the recommendation loss and the explanation loss, for example, the sum or weighted sum of the recommendation loss and the explanation loss is taken as the total loss value, so that the initial model is iterated in parameters based on the total loss value, that is, the interpretable recommendation model is trained.
[0075] Based on any of the above embodiments, the multi-modal explanation mode includes at least two of a voice explanation mode, a text explanation mode, and a visual explanation mode.
[0076] Specifically, the voice explanation mode refers to generating a voice explanation through voice synthesis technology to explain the reason for generating the recommendation result to the user in the form of voice. The text explanation mode refers to generating a text explanation to show the user the reason for generating the recommendation result in the form of text. The visual explanation refers to using charts, graphs, and other visual methods to show the user the reason for generating the recommendation result. In the embodiments of the present application, by providing multiple explanation modes, the needs of different users can be met.
[0077] Based on any of the above embodiments, step 140 specifically includes:
[0078] In step 141, in the case where the multi-modal explanation mode includes the voice explanation mode, voice synthesis is performed based on the explanation result to generate a voice explanation, and the voice explanation is played to explain the reason for generating the recommendation result to the user.
[0079] Specifically, the system can employ a speech synthesis technique to generate a speech explanation for the generation reason of the recommendation result, and then play the speech explanation to the user, which is suitable for driving or other scenarios that require visual concentration but can receive information through hearing. Specifically, first, the system can convert the generated explanation result into a format suitable for speech expression, convert the processed explanation result into human understandable speech using a speech synthesis technique, and play the speech explanation through a car audio system or other audio output device, so that the user can understand the generation reason of the recommendation result without looking at the screen.
[0080] Step 142, in the case that the multi-modal explanation mode includes the text explanation mode, generating a text explanation based on the explanation result, and displaying the text explanation on the head-up display, or displaying the text explanation on the car instrument panel, to explain the generation reason of the recommendation result to the user.
[0081] Specifically, when the system chooses to explain the recommendation result in text form, it can first generate a text explanation according to the explanation result and display it on a suitable interface, such as a head-up display (HUD) or a car instrument panel, so that the user can quickly read and understand the generation reason of the recommendation result. Specifically, first, the explanation result can be directly converted into a readable text format to ensure accurate information and clear expression; then, according to user preferences and system configuration, it is determined whether to display the text explanation on the head-up display or the car instrument panel; finally, the text explanation is displayed at the selected location to ensure that it is within the user's field of view and easy to recognize.
[0082] It can be understood that projecting the explanation information into the driver's field of view using the HUD can reduce the driver's distraction, for example, by displaying the reason for recommending a song or the basis for route selection. By displaying the explanation information on the car instrument panel, it is convenient for the driver to view, for example, the ratings and reviews of recommended restaurants, or the road condition information and estimated arrival time, etc. In addition, augmented reality (AR) technology can be used to superimpose the explanation information onto the real-world scene, for example, displaying the ratings and reviews of a restaurant on the restaurant building, to improve the user experience.
[0083] Step 143, in the case that the multi-modal explanation mode includes the visualization explanation mode, generating visualization explanation content based on the explanation result, and displaying the visualization explanation content on the head-up display, or displaying the visualization explanation content on the car instrument panel, to explain the generation reason of the recommendation result to the user.
[0084] Specifically, for scenarios that require more intuitive and graphical display of explanation content, the system can adopt a visual explanation method by generating visual content (such as charts, graphs, animations, etc.) based on the explanation result and displaying it on the head-up display or the vehicle-mounted dashboard. Specifically, first, the corresponding visual elements such as flowcharts, pie charts, bar charts, or dynamic demonstrations are designed and generated according to the explanation result to intuitively display the generation process and basis of the recommended result. Similar to text explanation, according to specific requirements and user preferences, the visual content is selected to be displayed on the head-up display or the vehicle-mounted dashboard. The visual explanation content is displayed on the selected display position to ensure that the user can easily understand the generation reason of the recommended result and the logic behind it.
[0085] Based on any of the above embodiments, step 140 specifically includes:
[0086] Based on the satisfaction evaluation result and / or the understanding evaluation result of each explanation method in the multi-modal explanation method, a target explanation method is determined from the multi-modal explanation method;
[0087] Based on the explanation result, the target explanation method is applied to explain the generation reason of the recommended result.
[0088] It should be noted that for each user, when explaining the recommended result, the target explanation method that is most suitable for the current situation and user preferences can be determined according to the satisfaction evaluation result and / or the understanding evaluation result of the user for different explanation methods (including voice explanation, text explanation, visual explanation, etc.) to explain the generation reason of the recommended result. This method can improve user experience and ensure that users can obtain and understand the required information in the most effective way.
[0089] Specifically, the satisfaction evaluation result of each explanation method refers to the user's subjective satisfaction with each explanation method, for example, the user may prefer voice explanation because it does not require distraction to look at the screen, or prefer visual explanation because it provides more intuitive information display. The understanding evaluation result of each explanation method refers to the degree to which the user can understand and master the generation reason of the recommended result through different explanation methods, which reflects the performance of the explanation method in terms of information transmission effect and clarity. If the user can quickly and accurately obtain the required information from the explanation, the understanding degree of the explanation method is high.
[0090] After obtaining the user's satisfaction evaluation results and understanding evaluation results of different explanation methods, the collected satisfaction evaluation and understanding evaluation results can be quantitatively processed, such as converting the scores into numerical values or converting the behavior data into relative indicators. Statistical analysis methods (such as mean, median, standard deviation, etc.) are used to summarize and analyze the evaluation data of each explanation method to understand the overall evaluation and preference of users for different explanation methods. According to the summary analysis results, each explanation method is comprehensively evaluated, and according to the results of the comprehensive evaluation, the priority of each explanation method can be set. The explanation method with higher priority usually has higher user satisfaction and understanding, and is more suitable as the target explanation method. Therefore, after determining the priority, the target explanation method can be selected from the multi-modal explanation methods. Here, the target explanation method refers to the explanation method that is most suitable for the current situation and user needs determined comprehensively according to the satisfaction evaluation results and understanding evaluation results, which aims to maximize user satisfaction and understanding, and ensures that users can obtain and understand the generation reasons of the recommended results in the most effective way.
[0091] Further, the selected target explanation method is applied to the actual scene to explain the generation reasons of the recommended results to the user. User feedback is continuously collected in the actual application process to understand the effect and user satisfaction of the target explanation method. According to user feedback and actual situation, the target explanation method can be iteratively optimized, such as adjusting the explanation content, improving the explanation method or introducing new explanation techniques, etc. By dynamically adjusting the target explanation method, the system can adapt to the changing needs of different users and different scenarios, and provide more personalized and intelligent explanation services.
[0092] Based on any of the above embodiments, the method further comprises:
[0093] Collecting user images in the vehicle and performing emotion recognition based on the user images to obtain emotion recognition results, and / or receiving user feedback data, the user feedback data including at least one of recommended result feedback data, explanation result feedback data, and explanation method feedback data;
[0094] Based on the emotion recognition results and / or the user feedback data, at least one of the recommended results, the explanation results, and the multi-modal explanation methods is adjusted.
[0095] Specifically, in an embodiment, the user's image can be collected by the camera or sensor in the vehicle and emotion recognition is performed, and the recommended results, explanation results, and explanation methods, etc. are adjusted according to the emotional state. For example, when the user shows negative emotions such as impatience to the system output results, the output content or explanation output logic can be automatically adjusted, and the driver is guided to adjust (such as recommending soothing music when the driver is anxious, and explaining in a more concise way).
[0096] Specifically, first, the user's images can be collected, which contain sufficient facial details so that the subsequent emotion recognition algorithm can accurately analyze; preprocessing operations are performed on the collected images to improve the accuracy and efficiency of subsequent emotion recognition. The preprocessing steps here can include face detection, image cropping, grayscale / color processing, noise removal, etc. On the preprocessed images, feature extraction algorithms are used to extract emotion-related facial features, which can include geometric features (such as the shape and position of eyes, eyebrows, and mouth, as well as their relative distances and angles), texture features (such as changes in skin texture, such as wrinkles, luster, etc.), dynamic features (such as the speed of expression changes, etc.), etc. The extracted features are input into an emotion classifier for classification and judgment, and the emotion recognition result is obtained. Here, the emotion classifier can be a trained machine learning model, such as a support vector machine, random forest, neural network, etc. The model can predict the user's current emotional state (such as happy, sad, angry, surprised, calm, etc.) based on the input features. The emotion recognition result output by the emotion classifier is used for subsequent processing and analysis, and based on the emotion recognition result, the system can adjust the recommended results, explanation results, or multi-modal explanation methods accordingly. For example, if the user shows positive emotions, the system can recommend more similar content or explanation methods; if the user shows negative emotions, the system can try to change the explanation method or provide more detailed explanations to help the user better understand the recommended results.
[0097] In another embodiment, the system can provide a user feedback mechanism, allowing users to provide feedback through voice commands, for example, users can directly feedback "I don't like this song" through voice, or directly ask "why recommend this route" through voice to improve interaction efficiency. Users can also provide feedback through gesture control functions, such as liking or skipping recommended content through gestures. Implicit feedback from users, such as song listening duration and route selection, can also be analyzed to further understand user preferences and improve recommendation accuracy. Through the above methods, user feedback data can be obtained, and based on the user feedback data, the system can adjust the recommended results, explanation results, or multi-modal explanation methods accordingly.
[0098] It is understood that the aforementioned user feedback data may include recommendation result feedback data, explanation result feedback data, and explanation method feedback data. Recommendation result feedback data refers to user feedback on the system's output recommendation results, which may include the accuracy of the recommendation results and user satisfaction with the recommendation system. Explanation result feedback data refers to user feedback on the explanation information provided by the system, which may include the accuracy of the explanation results and user satisfaction with the explanation results. Explanation method feedback data refers to user feedback on the explanation method used to interpret the recommendation results, which may include user satisfaction with the explanation method and their level of understanding.
[0099] The methods for obtaining evaluation feedback data for interpretation results have been described in the above embodiments and will not be repeated here. The following describes the methods for obtaining evaluation feedback data for recommendation results and interpretation method feedback data. Specifically, the accuracy of recommendation results can be evaluated in the following ways: ① Offline evaluation: Precision and recall of recommendation results can be calculated to measure the accuracy and coverage of the recommendations; F1 scores can be calculated by combining precision and recall to comprehensively evaluate the performance of the recommendation results; ROC curves can be plotted and AUC (area under the curve) can be calculated to evaluate the classification effect of the recommendation model. ② Online evaluation: Click-through rate of recommendation results can be statistically analyzed to measure the attractiveness of the recommendations; the proportion of users who perform target operations (such as purchase or registration) after clicking on recommendation results can be calculated to evaluate the effectiveness of the recommendations; user behavior on recommendation results (such as dwell time, interaction frequency, etc.) can be tracked to determine whether the recommendations match user preferences. ③ User feedback: Display feedback from users on recommendation results (such as likes, ratings, comments, etc.) can be collected to evaluate the accuracy of the recommendations; implicit user behavior (such as viewing time, skip frequency, etc.) can be analyzed to indirectly evaluate the accuracy of the recommendations. ④ Offline testing: Stratified sampling and cross-validation methods can be used to test the recommendation model offline and evaluate its accuracy on different datasets.
[0100] User satisfaction with a recommendation system can be assessed through the following methods: ① User surveys: Detailed satisfaction questionnaires, including Likert scales, can be designed to ask users about their overall satisfaction with the recommendation system; users can also freely express their opinions and suggestions. ② Behavioral data analysis: The frequency and duration of user usage of the recommendation system can be statistically analyzed to measure user acceptance of the system; the proportion of users who use the recommendation system multiple times within a certain period can be calculated to assess long-term user satisfaction. ③ User feedback analysis: Direct user feedback on the recommendation system (such as ratings and reviews) can be collected and analyzed to assess user satisfaction; sentiment analysis tools can be used to analyze the sentiment trends in user feedback to determine user satisfaction with the recommendation system.
[0101] In addition, the accuracy of the recommendation result and the user satisfaction can be comprehensively evaluated. For example, different weights can be given to the accuracy and satisfaction indicators according to actual needs, and a comprehensive score can be calculated. Preferably, the user groups can be subdivided according to user characteristics (such as age, gender, interests, etc.), and the recommendation accuracy and satisfaction of different groups can be evaluated to analyze the differences in accuracy and satisfaction of different user groups, and the recommendation system can be optimized to adapt to different user needs. Through the above methods, the accuracy of the recommendation result and the user satisfaction with the recommendation system can be comprehensively evaluated, and the recommendation system can be optimized and improved to improve the user experience and system performance.
[0102] It can be understood that the satisfaction and understanding of the user with different explanation methods can also be achieved by user surveys, user experience tests, sentiment analysis, user behavior data analysis, user feedback analysis, and user group analysis. For details, refer to the above embodiments, which will not be repeated here.
[0103] Based on any of the above embodiments, Figure 3 is a structural schematic diagram of a vehicle-mounted interaction device based on interpretable recommendation provided by the present application, as Figure 3 shown, the device comprises:
[0104] The receiving unit 310 is configured to receive a user voice in the vehicle.
[0105] The recognition unit 320 is configured to recognize the user voice to obtain a recognized text, and perform intent understanding on the recognized text to obtain intent information.
[0106] The recommendation unit 330 is configured to apply at least one of vehicle-mounted sensor data, user behavior data, multimedia data, and historical driving data and the intent information based on an interpretable recommendation model to generate a recommendation result and an explanation result.
[0107] The explanation unit 340 is configured to apply a multi-modal explanation method based on the explanation result to explain the generation reason of the recommendation result.
[0108] The device provided by the embodiment of the present application can generate corresponding explanation results while generating the recommendation results by introducing an interpretable recommendation model, and the generation reason of the recommendation results can be explained based on the explanation results, so that the user can better understand how the recommendation results are generated, and the transparency and trust are increased. The interpretable recommendation model can more accurately capture the real-time needs and preferences of the user by combining vehicle sensor data, user behavior data, multimedia data, historical driving data, intention information and other multi-source information, so as to realize more personalized recommendation. Meanwhile, the multi-modal explanation mode makes the explanation process more intuitive and easy to understand, and meets the needs of different users. In addition, the interpretable recommendation model considers various data sources and intention information of the user when generating the recommendation results, which helps to reduce the uncertainty in the recommendation process and improve the accuracy and relevance of the recommendation.
[0109] Based on any of the above embodiments, the device further includes a model training unit, which is configured to:
[0110] obtain an initial model and collect sample training data, the sample training data including at least one of sample vehicle sensor data, sample user behavior data, sample multimedia data, sample user feedback data and sample intention information and a sample recommendation result corresponding to the sample intention information;
[0111] based on the initial model, applying at least one of sample vehicle sensor data, sample user behavior data, sample multimedia data, sample user feedback data and sample intention information in the sample training data to generate a predicted recommendation result and a predicted explanation result;
[0112] based on the difference between the predicted recommendation result and the sample recommendation result, determining a recommendation loss and evaluating the accuracy and user satisfaction of the predicted explanation result to obtain an explanation loss;
[0113] based on the recommendation loss and the explanation loss, performing parameter iteration on the initial model to obtain the interpretable recommendation model.
[0114] Based on any of the above embodiments, the multi-modal explanation mode includes at least two of a voice explanation mode, a text explanation mode and a visual explanation mode.
[0115] Based on any of the above embodiments, the explanation unit 340 is specifically configured to:
[0116] in a case where the multi-modal explanation mode includes the voice explanation mode, performing voice synthesis based on the explanation result to generate a voice explanation, and playing the voice explanation to explain the generation reason of the recommendation result to the user;
[0117] In a case where the multi-modal explanation mode includes the text explanation mode, a text explanation is generated based on the explanation result, and the text explanation is displayed on a head-up display or a vehicle-mounted instrument panel to explain the recommended result to the user.
[0118] In a case where the multi-modal explanation mode includes the visual explanation mode, visual explanation content is generated based on the explanation result, and the visual explanation content is displayed on a head-up display or a vehicle-mounted instrument panel to explain the recommended result to the user.
[0119] Based on any of the above embodiments, the explanation unit 340 is specifically configured to:
[0120] Based on the satisfaction evaluation result and / or the understanding evaluation result of each explanation mode in the multi-modal explanation mode, a target explanation mode is determined from the multi-modal explanation mode;
[0121] Based on the explanation result, the target explanation mode is applied to explain the generation reason of the recommended result.
[0122] Based on any of the above embodiments, the device further includes a feedback optimization unit, which is configured to:
[0123] Collect a user image in the vehicle, and perform emotion recognition based on the user image to obtain an emotion recognition result, and / or receive user feedback data, the user feedback data including at least one of recommended result feedback data, explanation result feedback data, and explanation mode feedback data;
[0124] Based on the emotion recognition result and / or the user feedback data, at least one of the recommended result, the explanation result, and the multi-modal explanation mode is adjusted.
[0125] Figure 4 An example of an entity structure diagram of an electronic device is shown in FIG. 1. Figure 4As shown, the electronic device can include a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 complete mutual communication through the communication bus 440. The processor 410 can invoke the logical instructions in the memory 430 to execute the in-vehicle interaction method based on the interpretable recommendation, which includes: receiving a user voice in the vehicle; identifying the user voice to obtain identified text, and performing intent understanding on the identified text to obtain intent information; based on an interpretable recommendation model, applying at least one of vehicle sensor data, user behavior data, multimedia data, and historical driving data and the intent information to generate a recommendation result and an explanation result; and based on the explanation result, applying a multi-modal explanation method to explain the generation reason of the recommendation result.
[0126] In addition, the logical instructions in the memory 430 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application or parts of the related art that essentially contribute or the parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0127] On the other hand, the present application also provides a computer program product, which includes a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program is executed by a processor, and the computer can execute the in-vehicle interaction method based on the interpretable recommendation provided by the above-mentioned method, which includes: receiving a user voice in the vehicle; identifying the user voice to obtain identified text, and performing intent understanding on the identified text to obtain intent information; based on an interpretable recommendation model, applying at least one of vehicle sensor data, user behavior data, multimedia data, and historical driving data and the intent information to generate a recommendation result and an explanation result; and based on the explanation result, applying a multi-modal explanation method to explain the generation reason of the recommendation result.
[0128] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the vehicle interaction method based on interpretable recommendation provided by the above method, which comprises: receiving a user voice in a vehicle; recognizing the user voice to obtain recognized text, and performing intent understanding on the recognized text to obtain intent information; based on an interpretable recommendation model, applying at least one of vehicle sensor data, user behavior data, multimedia data, historical driving data and the intent information to generate a recommendation result and an explanation result; and based on the explanation result, applying a multi-modal explanation method to explain the generation reason of the recommendation result.
[0129] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.
[0130] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary universal hardware platforms, and of course, can also be realized by hardware. Based on such understanding, the above technical solutions, essentially or in terms of related art, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0131] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A vehicle interaction method based on interpretable recommendations, characterized in that, include: Receive user voice commands from inside the vehicle; The user's voice is recognized to obtain recognized text, and the recognized text is subjected to intent understanding to obtain intent information; Based on the interpretable recommendation model, at least one of vehicle sensor data, user behavior data, multimedia data, and historical driving data, along with the intent information, is used to generate recommendation results and explanation results. Based on the explanation results, a multimodal explanation method is applied to explain the reasons for the generation of the recommendation results; The training steps of the interpretable recommendation model include: Obtain the initial model and collect sample training data; Based on the initial model, the sample training data is used to generate prediction recommendation results and prediction explanation results; Based on the difference between the predicted recommendation results and the sample recommendation results in the sample training data, the recommendation loss is determined, and the accuracy of the predicted explanation results and user satisfaction are evaluated to obtain the explanation loss. Based on the recommendation loss and the explanation loss, the parameters of the initial model are iterated to obtain the interpretable recommendation model.
2. The in-vehicle interaction method based on interpretable recommendation according to claim 1, characterized in that, The sample training data includes at least one of sample vehicle sensor data, sample user behavior data, sample multimedia data, and sample user feedback data, as well as sample intent information and sample recommendation results corresponding to the sample intent information.
3. The in-vehicle interaction method based on interpretable recommendation according to claim 1, characterized in that, The multimodal interpretation method includes at least two of the following: voice interpretation method, text interpretation method, and visual interpretation method.
4. The in-vehicle interaction method based on interpretable recommendation according to claim 3, characterized in that, Based on the explanation results, a multimodal explanation method is applied to explain the reasons for the generation of the recommendation results, including: When the multimodal explanation method includes the voice explanation method, voice synthesis is performed based on the explanation result to generate a voice explanation, and the voice explanation is played to explain to the user the reason for the generation of the recommendation result; When the multimodal interpretation method includes the text interpretation method, a text interpretation is generated based on the interpretation result, and the text interpretation is displayed in a head-up display, or the text interpretation is displayed on the vehicle dashboard to explain to the user the reason for the generation of the recommendation result; When the multimodal explanation method includes the visualization explanation method, visualization explanation content is generated based on the explanation result, and the visualization explanation content is displayed in a head-up display, or the visualization explanation content is displayed on the vehicle dashboard to explain to the user the reason for the generation of the recommendation result.
5. The in-vehicle interaction method based on interpretable recommendation according to claim 3, characterized in that, Based on the explanation results, a multimodal explanation method is applied to explain the reasons for the generation of the recommendation results, including: Based on the satisfaction evaluation results and / or comprehension evaluation results of each interpretation method in the multimodal interpretation method, the target interpretation method is determined from the multimodal interpretation method; Based on the explanation results, the target explanation method is applied to explain the reasons for the generation of the recommendation results.
6. The in-vehicle interaction method based on interpretable recommendations according to any one of claims 1 to 5, characterized in that, Also includes: The system acquires user images inside the vehicle and performs emotion recognition based on the user images to obtain emotion recognition results, and / or receives user feedback data, wherein the user feedback data includes at least one of recommendation result feedback data, explanation result feedback data, and explanation method feedback data. Based on the emotion recognition results and / or the user feedback data, at least one of the recommendation results, the explanation results, and the multimodal explanation methods is adjusted.
7. A vehicle-mounted interactive device based on interpretable recommendations, characterized in that, include: The receiving unit is used to receive user voice messages inside the vehicle. The recognition unit is used to recognize the user's voice to obtain recognized text, and to perform intent understanding on the recognized text to obtain intent information; The recommendation unit is used to generate recommendation results and explanation results based on an interpretable recommendation model, applying at least one of vehicle sensor data, user behavior data, multimedia data, historical driving data, and the intent information. An explanation unit is used to explain the reasons for the generation of the recommendation results by applying a multimodal explanation method based on the explanation results. The training steps of the interpretable recommendation model include: Obtain the initial model and collect sample training data; Based on the initial model, the sample training data is used to generate prediction recommendation results and prediction explanation results; Based on the difference between the predicted recommendation results and the sample recommendation results in the sample training data, the recommendation loss is determined, and the accuracy of the predicted explanation results and user satisfaction are evaluated to obtain the explanation loss. Based on the recommendation loss and the explanation loss, the parameters of the initial model are iterated to obtain the interpretable recommendation model.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the in-vehicle interaction method based on interpretable recommendations as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the in-vehicle interaction method based on interpretable recommendations as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the in-vehicle interaction method based on interpretable recommendations as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Voice interactive recommendation method and system, storage medium and vehicle-mounted equipment
CN116052658A
Writing human readable interpretations for user navigation recommendations
CN117256001A