Vehicle function recommendation method and device, equipment and storage medium
By integrating voice and facial expression data into the vehicle function recommendation system, a multimodal intent vector is generated, which solves the problem of inaccurate intent recognition in traditional systems, enables accurate recommendations of personalized marketing strategies, and improves sales performance.
Patent Information
- Application Number
- CN202511392146.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-01-09
AI Technical Summary
Existing vehicle feature recommendation systems cannot fully and accurately identify users' true intentions, resulting in low adaptability to marketing strategies and poor recommendation effectiveness.
By acquiring users' voice and facial expression data, multi-level fusion is performed to generate multimodal target intent vectors. Combined with users' historical behavior and real-time emotions, targeted marketing strategies are generated to provide personalized feature recommendations.
It improved the accuracy of user intent recognition and the adaptability of marketing strategies, thereby increasing sales conversion rates.
Smart Images

Figure CN121304285A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of service recommendation technology, and in particular to methods, apparatus, devices and storage media for recommending vehicle functions. Background Technology
[0002] In special scenarios such as auto shows, media days, and customer self-driving experiences, where the focus is on recommending features of vehicles for sale, there is a lack of effective and comprehensive marketing-oriented intelligent guidance within the vehicle when users are freely experiencing it.
[0003] Current vehicle function recommendations are mainly achieved through user intent recognition. Traditional in-vehicle systems rely primarily on text input for intent recognition, which cannot accurately and comprehensively identify the user's true intent, resulting in poor recommendation performance and low adaptability to marketing strategies. Summary of the Invention
[0004] The main purpose of this application is to provide a vehicle function recommendation method, device, equipment and storage medium, which aims to solve the technical problem that current vehicle function recommendations cannot fully and accurately identify the user's true intentions, resulting in poor recommendation effect and low adaptability to marketing strategies.
[0005] To achieve the above objectives, this application proposes a vehicle function recommendation method, which includes:
[0006] During the display of vehicles for sale, user voice and facial expression data are collected;
[0007] The voice data and facial expression data are fused at multiple levels to obtain a multimodal target intent vector;
[0008] A target marketing strategy is generated based on the multimodal target intent vector, and the target marketing strategy is used to recommend features to the vehicles to be sold.
[0009] In one embodiment, the step of performing multi-level fusion of the voice data and the facial expression data to obtain a multimodal target intent vector includes:
[0010] Semantic content extraction and acoustic feature extraction are performed on the speech data to obtain the user's text content and acoustic features;
[0011] The facial expression data is subjected to feature extraction and facial expression classification to obtain facial expression features and facial expression categories;
[0012] A multimodal target intent vector is obtained by multi-level fusion based on the text content, acoustic features, facial expression features, and facial expression category.
[0013] In one embodiment, the step of performing multi-level fusion based on the text content, the acoustic features, the facial expression features, and the facial expression category to obtain a multimodal target intent vector includes:
[0014] The text content is subjected to text intent extraction to obtain an initial text intent vector;
[0015] Based on the acoustic features, speech emotion detection is performed to obtain the current speech emotion;
[0016] The initial text intent vector is adjusted based on the current voice emotion to obtain a voice-enhanced intent vector;
[0017] The facial expression features and the facial expression category are compared with the current voice emotion to obtain the comparison result;
[0018] The speech enhancement intent vector is adjusted based on the comparison results to obtain a multimodal target intent vector.
[0019] In one embodiment, the step of adjusting the speech enhancement intent vector based on the comparison result to obtain a multimodal target intent vector includes:
[0020] When the comparison result shows that the facial expression features and the facial expression category are consistent with the current speech emotion, the confidence of the speech enhancement intent vector is adjusted according to the current speech emotion to obtain a multimodal target intent vector;
[0021] When the comparison result shows that the facial expression features and the facial expression category are inconsistent with the current voice emotion, infrared detection is performed on the user to obtain the user's target emotional state.
[0022] The speech enhancement intent vector is adjusted based on the target emotional state to obtain a multimodal target intent vector.
[0023] In one embodiment, the step of adjusting the speech enhancement intent vector based on the comparison result to obtain a multimodal target intent vector includes:
[0024] When the comparison result shows that the facial expression features and the facial expression category are inconsistent with the current voice emotion, the user's historical interaction data is obtained;
[0025] Historical emotional states are determined based on the historical interaction data.
[0026] The speech enhancement intent vector is adjusted based on the historical emotional state to obtain a multimodal target intent vector.
[0027] In one embodiment, the step of generating a targeted marketing strategy based on the multimodal target intent vector includes:
[0028] Acquire users' historical behavior and real-time emotions;
[0029] The target intent is determined based on the multimodal target intent vector;
[0030] A targeted marketing strategy is generated based on the user's historical behavior, the user's real-time emotions, and the target intent.
[0031] In one embodiment, after generating a target marketing strategy based on the multimodal target intent vector and recommending features to the vehicle to be sold using the target marketing strategy, the method further includes:
[0032] Obtain user operation information regarding feature recommendations for the vehicles to be sold;
[0033] An optimization strategy is generated based on the operation information;
[0034] The target marketing strategy is adjusted using the optimization strategy described above.
[0035] Furthermore, to achieve the above objectives, this application also proposes a vehicle function recommendation device, which includes:
[0036] The acquisition module is used to acquire users' voice and facial expression data during the display of vehicles for sale;
[0037] The fusion module is used to perform multi-level fusion of the voice data and the facial expression data to obtain a multimodal target intent vector;
[0038] The generation module is used to generate a target marketing strategy based on the multimodal target intent vector, and to make functional recommendations for the vehicles to be sold through the target marketing strategy.
[0039] In addition, to achieve the above objectives, this application also proposes a vehicle function recommendation device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the vehicle function recommendation method as described above.
[0040] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the vehicle function recommendation method as described above.
[0041] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the vehicle function recommendation method as described above.
[0042] This application proposes one or more technical solutions that, during the display of a vehicle for sale, acquire user voice and facial expression data; perform multi-level fusion of the voice and facial expression data to obtain a multimodal target intent vector; generate a target marketing strategy based on the multimodal target intent vector; and recommend features to the vehicle for sale using the target marketing strategy. Building upon traditional text intent recognition, this application innovatively introduces hierarchical integration of voice and facial expression features. By comprehensively analyzing user voice and facial expression information, it can capture users' potential needs and preferences, thereby providing users with more personalized vehicle feature recommendations. This allows for a more accurate understanding of users' car-buying intentions, improves the adaptability of marketing strategies, and ultimately increases sales conversion rates. Attached Figure Description
[0043] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0044] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is a flowchart illustrating an embodiment of the vehicle function recommendation method of this application.
[0046] Figure 2 This is a flowchart illustrating Embodiment 2 of the vehicle function recommendation method of this application;
[0047] Figure 3 A simplified flowchart illustrating the vehicle function recommendation method provided in Embodiment 2 of this application;
[0048] Figure 4 This is a schematic diagram of the module structure of the vehicle function recommendation device according to an embodiment of this application;
[0049] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the vehicle function recommendation method in this application embodiment.
[0050] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0051] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0052] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0053] The main solution of this application embodiment is: during the display of a vehicle to be sold, acquiring the user's voice data and facial expression data; performing multi-level fusion on the voice data and facial expression data to obtain a multimodal target intent vector; generating a target marketing strategy based on the multimodal target intent vector; and recommending functions to the vehicle to be sold through the target marketing strategy.
[0054] Because current technologies primarily rely on text input for user intent recognition (such as semantic analysis after speech-to-text conversion), emotional information is lost. After speech-to-text conversion, emotional features such as tone and speech rate are ignored. Plain text cannot capture user facial expressions (such as frowning or smiling), and speech emotion recognition and facial expression recognition typically operate independently, lacking deep integration with text intent recognition and dynamic decision-making mechanisms. Usually, after speech-to-text conversion, the data is fused with a large model based on plain text and relies on post-conversion manual fine-tuning, resulting in low adaptability to marketing strategies.
[0055] This application provides a solution, proposing a three-level fusion intent recognition link. Building upon traditional text intent recognition, this link significantly improves the accuracy of intent recognition and the depth of emotion understanding by injecting speech and facial expression features in a hierarchical manner. First, at the semantic level, the link extracts core information from the text to construct a basic intent framework. This layer primarily focuses on the text content itself, identifying the user's main intent and needs through natural language processing technology. Second, at the emotion level, a speech feature extraction module analyzes the emotional nuances in the speech signal, such as tone, speed, and pauses. The purpose of this layer is to extract emotional information from the speech, further enriching the intent expression and enabling the system to perceive the user's emotional state. Finally, at the emotional level, combined with facial expression feature recognition technology, it captures non-verbal information such as facial micro-expressions and eye movements. The purpose of this layer is to deeply analyze the user's emotional state and understand their underlying intent and emotional fluctuations. Through this layered and progressive fusion approach, the system can more accurately identify the user's explicit and implicit intents while capturing the impact of emotional fluctuations on intent. This method not only improves the accuracy and comprehensiveness of intent recognition, but also provides richer emotion perception capabilities for intelligent interaction systems. It is applicable to scenarios such as intelligent customer service and virtual assistants, helping to achieve more natural and humanized dialogue interaction.
[0056] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, such as a vehicle function recommendation device. The following description uses a vehicle function recommendation device as an example to illustrate this embodiment and the subsequent embodiments.
[0057] Based on this, embodiments of this application provide a vehicle function recommendation method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the vehicle function recommendation method of this application.
[0058] In this embodiment, the vehicle function recommendation method includes steps S10 to S30:
[0059] Step S10: During the display of the vehicle to be sold, acquire the user's voice data and facial expression data.
[0060] It should be noted that the vehicles for sale are those displayed and sold in car showrooms and car dealerships. When users freely experience these vehicles, the lack of effective and comprehensive feature recommendations can reduce their willingness to purchase. Therefore, to improve the user's visit and test drive experience, user intent can be identified to generate specific marketing strategies that recommend features to the vehicles for sale, thereby enhancing the user's test drive and visit experience.
[0061] The user's voice and facial expression data can be collected by a voice acquisition device and an image acquisition device installed on the vehicle to be sold. The voice acquisition device can be a microphone, and the image acquisition device can be a camera, or other devices or apparatus capable of voice and image acquisition.
[0062] It should be noted that all user-related data involved in this application (such as user attribute data, user behavior data, and user geographical location, etc., the data types here should be modified according to the adaptability of the solution content) were obtained after obtaining the user's permission or consent; that is to say, when this application is applied to specific products or technologies, user permission is required to obtain and process the relevant data, and the processing of the relevant data must comply with the relevant laws, regulations and regulatory standards of the relevant countries and regions.
[0063] Understandably, voice data includes the language content of users interacting with the vehicle or talking to other people, while facial expression data is the facial expressions of users when test driving or experiencing the vehicle to be sold, which can be obtained by capturing the user's facial images through the in-vehicle camera.
[0064] Step S20: Perform multi-level fusion on the voice data and the facial expression data to obtain a multimodal target intent vector.
[0065] Understandably, speech data and facial expression data can be fused at multiple levels to gradually enhance the accuracy of intent recognition and obtain a multimodal target intent vector.
[0066] In practice, the multimodal target intent vector can most accurately represent the user's current intent. Specifically, features can be extracted from voice data and facial expression data to perform multi-level intent fusion and obtain the multimodal target intent vector.
[0067] Step S30: Generate a target marketing strategy based on the multimodal target intent vector, and recommend functions to the vehicle to be sold through the target marketing strategy.
[0068] It should be noted that after determining the user's intent, the corresponding marketing strategy can be triggered. For example, if the multimodal target intent is to query functions, the target marketing strategy can be determined to be to play a function demonstration video for the user, thereby recommending the corresponding functions of the vehicle to be sold through playing the function demonstration video. For example, if the multimodal target intent is to inquire about prices, the target marketing strategy is to pop up limited-time offer information, so the limited-time offer information can be used to recommend the price of the vehicle to be sold.
[0069] In practice, to improve the accuracy of targeted marketing strategies, it is possible to combine users' historical behavior data with their current emotions and multimodal target intentions to generate targeted marketing strategies.
[0070] In one feasible implementation, step S30 may include steps A11 to A13:
[0071] Step A11: Obtain user's historical behavior and real-time user sentiment;
[0072] It's important to note that user history refers to information the user has previously queried about vehicles for sale, such as frequently searching for "autonomous driving." Real-time user emotion refers to the user's actual emotional state, which can be determined through infrared liveness detection. Real-time user emotions can include curiosity, hesitation, excitement, and frustration. Combining user history and real-time emotion helps prevent policy recommendations from becoming disconnected from the user's current interests.
[0073] Step A12: Determine the target intent based on the multimodal target intent vector;
[0074] In practice, target intent can be determined through multimodal target intent vectors. Since multimodal target intent vectors have integrated non-textual information such as voice and facial expressions, they can more accurately identify users’ explicit intent (such as “querying battery life”) and implicit intent (such as inferring price sensitivity through hesitation), providing precise direction for strategy generation.
[0075] Step A13: Generate a target marketing strategy based on the user's historical behavior, the user's real-time emotions, and the target intent.
[0076] It's important to note that marketing strategies can be generated by combining user history, real-time user sentiment, and target intent. This is an intent-driven, dynamic marketing strategy that enables personalized recommendations. Historical behavior ensures the strategy aligns with user habits, real-time sentiment adjusts communication methods, and target intent focuses on core needs, ultimately improving user acceptance and conversion rates.
[0077] For example, if the target intent is to search for a function, the user's emotion is curiosity, and their historical behavior shows multiple searches for the functions or performance information of the vehicle for sale, then the target marketing strategy is to play a function demonstration video. If the target intent is unknown, the current emotion is annoyance, and there is no historical behavior, then the target marketing strategy is to terminate the marketing and switch to a simplified mode to avoid alienating the user.
[0078] Understandably, to improve the accuracy of intent recognition and marketing strategies, user action data can be collected in real time after the target marketing strategy is recommended to the user, thereby continuously optimizing the target marketing strategy. Therefore, after step S30, steps S31 to S33 are also included:
[0079] Step S31: Obtain user operation information regarding the function recommendation content for the vehicle to be sold;
[0080] It should be understood that after recommending content corresponding to the target marketing strategy to users, it is possible to obtain in real time the user's operation information related to the content recommended to the vehicle to be sold, such as whether the operation information is skip or click.
[0081] Step S32: Generate an optimization strategy based on the operation information;
[0082] In practice, optimization strategies can be generated based on user action information. For example, if the user's action information is to skip, the recommended content will be adjusted.
[0083] Step S33: Adjust the target marketing strategy using the optimization strategy.
[0084] In practice, the previous target marketing strategy can be adjusted by optimizing the strategy, thereby improving the accuracy of intent recognition and marketing strategy.
[0085] In practice, a multimodal fusion model can be generated through training, and then used to identify user intent and obtain a multimodal target intent vector. Therefore, after obtaining the optimization strategy, the target marketing strategy can be adjusted through the optimization strategy, thereby optimizing the multimodal fusion model and improving the robustness of the model.
[0086] This embodiment provides a vehicle feature recommendation method. During the display of a vehicle for sale, user voice and facial expression data are acquired. The voice and facial expression data are then fused at multiple levels to obtain a multimodal target intent vector. A target marketing strategy is generated based on the multimodal target intent vector, and features are recommended to the vehicle for sale using this strategy. Building upon traditional text intent recognition, this method innovatively introduces hierarchical integration of voice and facial expression features. By comprehensively analyzing user voice and facial expression information, it can capture users' potential needs and preferences, thereby providing more personalized vehicle feature recommendations. This allows for a more accurate understanding of users' car-buying intentions, improves the adaptability of marketing strategies, and ultimately increases sales conversion rates.
[0087] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Step S20 includes steps S201 to S203:
[0088] Step S201: Extract semantic content and acoustic features from the speech data to obtain the user's text content and acoustic features.
[0089] It should be noted that automatic speech recognition technology can be used to extract semantic content from speech data, and acoustic features can be extracted by analyzing the acoustic characteristics of the speech. For example, ASR technology can be used to convert speech into text to obtain the user's textual language content, and the acoustic features of the speech, including pitch, intonation, speech rate, pauses, and spectral energy, can be analyzed to obtain the user's acoustic characteristics. Natural language processing technology is then used to identify the user's main intentions and needs for subsequent intent fusion.
[0090] Step S202: Perform feature extraction and expression classification on the expression data to obtain expression features and expression categories.
[0091] In practice, facial expression recognition technology can be used to extract and classify features from users' facial expression data, such as classifying users' facial expressions into categories like happy, confused, and impatient.
[0092] By extracting features from facial expression data, such as eye movement frequency and mouth tremors, users' emotional states can be analyzed in more detail.
[0093] Step S203: Perform multi-level fusion based on the text content, the acoustic features, the facial features, and the facial expression category to obtain a multimodal target intent vector.
[0094] In practice, multi-level intent fusion can be performed using text content, acoustic features, facial features, and facial expression categories to obtain a multimodal target intent vector.
[0095] It should be noted that in this embodiment, a three-level fusion method is mainly used: first, primary fusion; then, intermediate fusion; and finally, advanced fusion, to obtain multimodal target intent adjacency.
[0096] In one feasible implementation, step S203 may include steps B11 to B15:
[0097] Step B11: Extract text intent from the text content to obtain an initial text intent vector;
[0098] It should be noted that text intent can be extracted from text content. Specifically, the basic intent of the text can be extracted using a pre-trained natural language processing model. For example, if the input text content from the user is "What other functions are available?", then by extracting the text intent from this text content, an initial text intent vector is obtained, which is the intent of "query function". After obtaining the initial text intent vector Intent_Text, the initial text intent vector can be used as the basis for subsequent fusion.
[0099] Step B12: Detect speech emotion based on the acoustic features to obtain the current speech emotion;
[0100] It should be noted that acoustic features can be used to detect speech emotion, thereby obtaining the user's current speech emotion. Specifically, the acoustic features of the speech can be input into the speech emotion detection model to obtain the current speech emotion. For example, if the acoustic features are faster speech rate and higher pitch, the current speech emotion may be irritability or anger.
[0101] Step B13: Adjust the initial text intent vector according to the current voice emotion to obtain a voice enhancement intent vector;
[0102] Understandably, intermediate-level fusion can be performed based on the current speech emotion, that is, the initial text intent vector is adjusted based on the current speech emotion to obtain a speech-enhanced intent vector.
[0103] For example, if the current voice emotion detected is positive (e.g., excitement), the initial text intent vector (Intent_Text) is weighted to increase the priority of related intents (e.g., recommending more features); if the current voice emotion detected is negative (e.g., impatience), the text intent vector is deweighted and corresponding strategies are triggered (e.g., shortening the reply or changing the topic). By weighting or deweighting the initial text intent vector, the voice enhancement intent vector Intent_Voice is obtained.
[0104] Step B14: Compare the facial expression features and the facial expression category with the current voice emotion to obtain the comparison result;
[0105] In practice, facial expression features and categories can be compared with the current voice emotion to determine whether the emotions expressed by the facial expression features and categories are consistent with the current voice emotion, thus obtaining specific comparison results.
[0106] Specifically, for example, if a user's facial expression is a frown and the expression category is impatience, then the user's frowning facial expression and impatience expression category are compared with the current voice emotion. If the current voice emotion is negative, then the comparison result is consistent; if the current voice emotion is positive, then the comparison result is inconsistent.
[0107] Step B15: Adjust the speech enhancement intent vector according to the comparison results to obtain a multimodal target intent vector.
[0108] It should be noted that the speech enhancement intent vector can be adjusted according to the specific comparison results, thereby increasing or decreasing the confidence of the intent, achieving advanced fusion, and obtaining a multimodal target intent vector.
[0109] In one feasible implementation, step B15 may include: when the comparison result shows that the facial expression features and the facial expression category are consistent with the current voice emotion, adjusting the confidence of the voice enhancement intent vector according to the current voice emotion to obtain a multimodal target intent vector; when the comparison result shows that the facial expression features and the facial expression category are inconsistent with the current voice emotion, performing infrared detection on the user to obtain the user's target emotional state; and adjusting the voice enhancement intent vector based on the target emotional state to obtain a multimodal target intent vector.
[0110] It should be noted that if the comparison result is consistent, it means that the current voice emotion can accurately represent the user's intent. Therefore, the confidence of the voice enhancement intent vector can be directly adjusted according to the current voice emotion to obtain the multimodal target intent vector.
[0111] For example, if the current voice emotion is anger, the confidence of the voice enhancement intent vector will be reduced, and emergency strategies may be triggered, such as stopping marketing or asking about user needs. If the current voice emotion is happiness, the confidence of the voice enhancement intent vector will be increased, thereby generating a targeted marketing strategy.
[0112] It should be noted that if the user's comparison results are inconsistent—for example, the current voice emotion is calm, but the facial expression features disgust and the expression category is anger—it indicates a conflict between the voice emotion and the facial expression. To eliminate environmental influences or misjudgments, infrared liveness detection can be activated to eliminate environmental interference (such as changes in lighting) and confirm the true emotional state. The target emotional state is the true emotional state obtained through infrared liveness detection of the user. This allows for adjustments to the weights of voice and facial expression based on the interaction scenario (e.g., reducing the weight of voice and increasing the weight of facial expression in noisy environments).
[0113] In practice, the confidence level of the voice-enhanced intent vector can be adjusted according to the target's emotional state. For example, if the target's emotional state is anger, the confidence level of the voice-enhanced intent vector can be reduced, and an emergency strategy can be triggered. If the target's emotional state is happiness, the confidence level of the voice-enhanced intent vector can be increased, and a corresponding target marketing strategy can be generated.
[0114] It should be noted that, to avoid misjudgment, the user's historical interaction data can also be used to judge the voice emotion. For example, if a user was previously misjudged as angry because of "speaking too fast", the voice enhancement intent vector can be adjusted by combining the user's historical interaction data on the vehicles being sold. Therefore, step B15 may also include: when the comparison result is that the facial expression features and the facial expression category are inconsistent with the current voice emotion, obtaining the user's historical interaction data; determining the historical emotional state based on the historical interaction data; and adjusting the voice enhancement intent vector based on the historical emotional state to obtain a multimodal target intent vector.
[0115] It should be noted that if the comparison result shows a discrepancy, historical interaction data of the user regarding the vehicle being sold can be obtained. This historical interaction data can then be used to extract features, revealing the user's historical emotional state. The voice enhancement intent vector can then be adjusted based on this historical emotional state; for example, if the historical emotional state is "happy," the confidence level of the voice enhancement intent vector will be increased, thereby generating a targeted marketing strategy. By using historical interaction data, misjudgments caused by signal inconsistencies can be significantly reduced, improving the reliability of the interaction and the user experience.
[0116] This embodiment extracts semantic content and acoustic features from the speech data to obtain the user's text content and acoustic features; it also extracts features and classifies facial expressions from the facial expression data to obtain facial expression features and categories; and performs multi-level fusion based on the text content, acoustic features, facial expression features, and facial expression categories to obtain a multimodal target intent vector. By extracting and fusing information at three levels—semantic, emotional, and affective—a deeper understanding of the user's intent is achieved. This multi-level fusion approach effectively improves the accuracy of intent recognition, capturing both explicit and implicit user needs, and surpasses the limitations of traditional single-text analysis.
[0117] For example, to help understand the implementation process of the vehicle function recommendation method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 3 , Figure 3 A simplified flowchart of a vehicle function recommendation method is provided, specifically:
[0118] The system collects user voice and facial expression signals. Text and acoustic features are extracted from the voice signals to obtain text content and acoustic features. Facial expressions and micro-expressions are extracted from the facial expression signals to obtain expression categories and features. Then, a primary fusion is performed to extract the text intent baseline, resulting in an initial text intent vector. This initial text intent vector is then fused at a secondary level to obtain a voice-enhanced intent vector. Finally, a high-level fusion is performed on the voice-enhanced intent vector to dynamically correct facial expressions, resulting in a multimodal target intent vector. This allows for the generation of intent-driven dynamic marketing strategies. For example, if the multimodal target intent vector represents a query intent with the emotion of curiosity, the target marketing strategy is to play a function demonstration video. If the multimodal target intent vector represents a price inquiry intent with the emotion of hesitation, the target marketing strategy is to pop up a limited-time offer. If the multimodal target intent vector represents an unknown intent with the emotion of frustration, the target marketing strategy is to terminate the marketing and switch to a simplified mode.
[0119] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the vehicle function recommendation method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0120] This application also provides a vehicle function recommendation device, please refer to... Figure 4 The vehicle function recommendation device includes:
[0121] The acquisition module 10 is used to acquire the user's voice data and facial expression data during the display of vehicles to be sold.
[0122] The fusion module 20 is used to perform multi-level fusion of the voice data and the facial expression data to obtain a multimodal target intent vector.
[0123] The generation module 30 is used to generate a target marketing strategy based on the multimodal target intent vector, and to make functional recommendations for the vehicles to be sold through the target marketing strategy.
[0124] The vehicle function recommendation device provided in this application, employing the vehicle function recommendation method described in the above embodiments, can solve the technical problem that current vehicle function recommendations cannot comprehensively and accurately identify the user's true intent, resulting in poor recommendation effects and low adaptability to marketing strategies. Compared with the prior art, the beneficial effects of the vehicle function recommendation device provided in this application are the same as those of the vehicle function recommendation method provided in the above embodiments, and other technical features in the vehicle function recommendation device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0125] In one embodiment, the fusion module 20 is further configured to extract semantic content and acoustic features from the speech data to obtain the user's text content and acoustic features; extract features and classify expressions from the facial expression data to obtain facial expression features and facial expression categories; and perform multi-level fusion based on the text content, the acoustic features, the facial expression features, and the facial expression categories to obtain a multimodal target intent vector.
[0126] In one embodiment, the fusion module 20 is further configured to: extract text intent from the text content to obtain an initial text intent vector; detect speech emotion based on the acoustic features to obtain the current speech emotion; adjust the initial text intent vector based on the current speech emotion to obtain a speech enhancement intent vector; compare the facial expression features and the facial expression category with the current speech emotion to obtain a comparison result; and adjust the speech enhancement intent vector based on the comparison result to obtain a multimodal target intent vector.
[0127] In one embodiment, the fusion module 20 is further configured to: when the comparison result indicates that the facial expression features and the facial expression category are consistent with the current voice emotion, adjust the confidence of the voice enhancement intent vector according to the current voice emotion to obtain a multimodal target intent vector; when the comparison result indicates that the facial expression features and the facial expression category are inconsistent with the current voice emotion, perform infrared detection on the user to obtain the user's target emotional state; and adjust the voice enhancement intent vector based on the target emotional state to obtain a multimodal target intent vector.
[0128] In one embodiment, the fusion module 20 is further configured to: acquire the user's historical interaction data when the comparison result shows that the facial expression features and the facial expression category are inconsistent with the current voice emotion; determine the historical emotional state based on the historical interaction data; and adjust the voice enhancement intent vector based on the historical emotional state to obtain a multimodal target intent vector.
[0129] In one embodiment, the generation module 30 is further configured to acquire user historical behavior and user real-time emotions; determine target intent based on the multimodal target intent vector; and generate a target marketing strategy based on the user historical behavior, the user real-time emotions, and the target intent.
[0130] In one embodiment, the generation module 30 is further configured to acquire user operation information regarding the function recommendation content of the vehicle to be sold; generate an optimization strategy based on the operation information; and adjust the target marketing strategy through the optimization strategy.
[0131] This application provides a vehicle function recommendation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the vehicle function recommendation method in the above embodiment 1.
[0132] The following is for reference. Figure 5 The diagram illustrates a structural schematic suitable for implementing a vehicle function recommendation device according to embodiments of this application. The vehicle function recommendation device in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The vehicle function recommendation device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0133] like Figure 5As shown, the vehicle function recommendation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the vehicle function recommendation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touch screens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, LCDs (Liquid Crystal Displays), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the vehicle function recommendation device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show vehicle function recommendation devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0134] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0135] The vehicle function recommendation device provided in this application, employing the vehicle function recommendation method described in the above embodiments, can solve the technical problem that current vehicle function recommendations cannot comprehensively and accurately identify the user's true intent, resulting in poor recommendation effects and low adaptability to marketing strategies. Compared with the prior art, the beneficial effects of the vehicle function recommendation device provided in this application are the same as those of the vehicle function recommendation method provided in the above embodiments, and other technical features in this vehicle function recommendation device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0136] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0137] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0138] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to perform the vehicle function recommendation method in the above embodiments.
[0139] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or Flash Memory), optical fibers, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0140] The aforementioned computer-readable storage medium may be included in the vehicle function recommendation device; or it may exist independently and not be installed in the vehicle function recommendation device.
[0141] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the vehicle function recommendation device, cause the vehicle function recommendation device to: acquire user voice data and facial expression data during the display of a vehicle to be sold; perform multi-level fusion of the voice data and the facial expression data to obtain a multimodal target intent vector; generate a target marketing strategy based on the multimodal target intent vector; and recommend functions to the vehicle to be sold through the target marketing strategy.
[0142] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LAN (Local Area Network) or WAN (Wide Area Network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0144] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0145] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described vehicle function recommendation method. This solves the technical problem that current vehicle function recommendations cannot fully and accurately identify the user's true intent, resulting in poor recommendation performance and low adaptability to marketing strategies. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the vehicle function recommendation method provided in the above embodiments, and will not be repeated here.
[0146] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the vehicle function recommendation method described above.
[0147] The computer program product provided in this application can solve the technical problem that current vehicle function recommendations cannot fully and accurately identify the user's true intent, resulting in poor recommendation performance and low adaptability to marketing strategies. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the vehicle function recommendation method provided in the above embodiments, and will not be repeated here.
[0148] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for recommending vehicle functions, characterized in that, The vehicle function recommendation method includes: During the display of vehicles for sale, user voice and facial expression data are collected; The voice data and facial expression data are fused at multiple levels to obtain a multimodal target intent vector; A target marketing strategy is generated based on the multimodal target intent vector, and the target marketing strategy is used to recommend features to the vehicles to be sold.
2. The method as described in claim 1, characterized in that, The step of performing multi-level fusion of the speech data and the facial expression data to obtain a multimodal target intent vector includes: Semantic content extraction and acoustic feature extraction are performed on the speech data to obtain the user's text content and acoustic features; The facial expression data is subjected to feature extraction and facial expression classification to obtain facial expression features and facial expression categories; A multimodal target intent vector is obtained by multi-level fusion based on the text content, acoustic features, facial expression features, and facial expression category.
3. The method as described in claim 2, characterized in that, The step of performing multi-level fusion based on the text content, the acoustic features, the facial expression features, and the facial expression category to obtain a multimodal target intent vector includes: The text content is subjected to text intent extraction to obtain an initial text intent vector; Based on the acoustic features, speech emotion detection is performed to obtain the current speech emotion; The initial text intent vector is adjusted based on the current voice emotion to obtain a voice-enhanced intent vector; The facial expression features and the facial expression category are compared with the current voice emotion to obtain the comparison result; The speech enhancement intent vector is adjusted based on the comparison results to obtain a multimodal target intent vector.
4. The method as described in claim 3, characterized in that, The step of adjusting the speech enhancement intent vector based on the comparison result to obtain the multimodal target intent vector includes: When the comparison result shows that the facial expression features and the facial expression category are consistent with the current speech emotion, the confidence of the speech enhancement intent vector is adjusted according to the current speech emotion to obtain a multimodal target intent vector; When the comparison result shows that the facial expression features and the facial expression category are inconsistent with the current voice emotion, infrared detection is performed on the user to obtain the user's target emotional state. The speech enhancement intent vector is adjusted based on the target emotional state to obtain a multimodal target intent vector.
5. The method as described in claim 3, characterized in that, The step of adjusting the speech enhancement intent vector based on the comparison result to obtain the multimodal target intent vector includes: When the comparison result shows that the facial expression features and the facial expression category are inconsistent with the current voice emotion, the user's historical interaction data is obtained; Historical emotional states are determined based on the historical interaction data. The speech enhancement intent vector is adjusted based on the historical emotional state to obtain a multimodal target intent vector.
6. The method as described in claim 1, characterized in that, The steps for generating a targeted marketing strategy based on the multimodal target intent vector include: Acquire users' historical behavior and real-time emotions; The target intent is determined based on the multimodal target intent vector; A targeted marketing strategy is generated based on the user's historical behavior, the user's real-time emotions, and the target intent.
7. The method according to any one of claims 1 to 6, characterized in that, After generating a target marketing strategy based on the multimodal target intent vector and recommending features to the vehicle to be sold using the target marketing strategy, the process further includes: Obtain user operation information regarding feature recommendations for the vehicles to be sold; An optimization strategy is generated based on the operation information; The target marketing strategy is adjusted using the optimization strategy described above.
8. A vehicle function recommendation device, characterized in that, The device includes: The acquisition module is used to acquire users' voice and facial expression data during the display of vehicles for sale; The fusion module is used to perform multi-level fusion of the voice data and the facial expression data to obtain a multimodal target intent vector; The generation module is used to generate a target marketing strategy based on the multimodal target intent vector, and to make functional recommendations for the vehicles to be sold through the target marketing strategy.
9. A vehicle function recommendation device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the vehicle function recommendation method as claimed in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the vehicle function recommendation method as described in any one of claims 1 to 7.