Vehicle-mounted voice control method, electronic equipment and vehicle

By fusing multimodal data and processing dynamic user profiles of in-vehicle voice systems, personalized voice command recommendations are generated, solving the problem that existing systems cannot adapt to user preferences and driving environments, and improving interaction efficiency and user experience.

CN121996072APending Publication Date: 2026-05-08GREAT WALL MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GREAT WALL MOTOR CO LTD
Filing Date
2026-01-23
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing in-vehicle voice systems cannot intelligently recommend voice commands based on user preferences and driving environment, resulting in poor interaction efficiency and user experience.

Method used

By fusing multimodal data, a dynamic user profile is constructed, and a personalized voice command recommendation list is generated based on intent prediction. The display format is determined in conjunction with driving environment assessment.

Benefits of technology

It improves the personalization and contextual relevance of voice commands, thereby increasing click-through rates and user interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121996072A_ABST
    Figure CN121996072A_ABST
Patent Text Reader

Abstract

The invention relates to a vehicle-mounted voice control method, electronic equipment and a vehicle, is applied to the electronic equipment, and relates to the technical field of vehicle control. The vehicle-mounted voice control method comprises the following steps: when a user uses a voice interaction function in a vehicle, performing intention prediction through multi-modal fusion data obtained by performing data fusion on multi-modal data including user-related data and vehicle-related data, and a constructed dynamic user portrait; the voice instruction recommendation list is generated based on the obtained potential intention set and the preset recommendation strategy, the instruction can be intelligently recommended according to the user preference and the driving environment, the recommendation individuation degree and the scene fitting degree are improved, and therefore the click rate of the voice instruction and the user interaction efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of vehicle control technology, and in particular to an in-vehicle voice control method, electronic equipment, and vehicle. Background Technology

[0002] With the rapid development of vehicle-to-everything (V2X) technology, in-vehicle voice interaction has become an important human-machine interaction method for improving driving safety and convenience. In-vehicle voice interaction enables voice communication between the driver and the vehicle, providing functions such as navigation, music playback, and phone calls. Among the components of in-vehicle voice interaction, the voice command recommendation function plays a crucial role in helping users quickly get started, guiding them to discover and understand functions, and enhancing the user experience.

[0003] Most mainstream in-vehicle voice systems employ a fixed and static approach to voice command recommendation. Common solutions include: carousel recommendations based on preset hot word lists, providing limited related commands based on the current application interface, or simply sorting recommendations based on global usage frequency. Therefore, these systems cannot adapt to the personalized usage habits and preferences of different users, nor can they recommend appropriate commands based on the real-time driving environment, thus affecting voice interaction efficiency and user experience. Therefore, how to intelligently recommend commands based on user preferences and the driving environment is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure provides an in-vehicle voice control method, an electronic device, and a vehicle.

[0005] A first aspect of this disclosure provides an in-vehicle voice control method, applied to an electronic device, comprising: In response to the activation of the in-vehicle voice interaction function, the acquired multimodal data is subjected to data fusion processing to obtain corresponding multimodal fused data, which includes user-related data and vehicle-related data; Construct a dynamic user profile, which includes multiple dimensional attributes representing the user's real-time status; Based on the multimodal fusion data and the dynamic user profile, intent prediction processing is performed to obtain a set of potential intents; A voice command recommendation list is generated based on the set of potential intents and a preset recommendation strategy, and the voice command recommendation list is displayed.

[0006] In some embodiments of this disclosure, the step of performing data fusion processing on the acquired multimodal data to obtain corresponding multimodal fused data includes: Feature extraction is performed on the acquired multimodal data to obtain corresponding multimodal feature vectors, which include user preference feature vectors, context feature vectors, vehicle state feature vectors, and historical interaction feature vectors. The multimodal feature vectors are weighted and fused based on the target weights to obtain the multimodal fused data, and the target weights are dynamically adjusted according to the multimodal fused data.

[0007] In some embodiments of this disclosure, the dynamic user profile includes static attributes, dynamic attributes, and contextual attributes; The construction of dynamic user profiles includes: The static attributes are constructed based on the user preference feature vector; The dynamic attributes are constructed based on the historical interaction feature vectors. The context attributes are constructed based on the context feature vector and the vehicle state feature vector; The dynamic user profile is constructed based on the static attributes, the dynamic attributes, and the contextual attributes.

[0008] In some embodiments of this disclosure, the process of performing intent prediction based on the multimodal fusion data and the dynamic user profile to obtain a set of potential intents includes: Intent recognition is performed on the multimodal fusion data and the dynamic user profile to obtain the actual instruction intent; The actual instruction intent is predicted based on a pre-constructed intent association tree to obtain the potential intent set. The intent association tree is pre-trained from the multimodal fusion data, and the potential intent set includes multiple potential instruction intents.

[0009] In some embodiments of this disclosure, the step of predicting the actual instruction intent based on a pre-built intent association tree to obtain the potential intent set includes: Calculate the dynamic correlation degree between the actual instruction intent and each potential instruction intent in the intent association tree; Based on the dynamic correlation, the top-ranked potential instruction intentions are determined to obtain the potential intention set.

[0010] In some embodiments of this disclosure, generating a voice command recommendation list based on the potential intent set and a preset recommendation strategy includes: Based on the preset recommendation strategy, the potential intent set is recommended to obtain multiple candidate voice commands; The multiple candidate voice commands are scored and sorted to generate a recommended list of voice commands.

[0011] In some embodiments of this disclosure, before displaying the recommended list of voice commands, the method further includes: The driving load level of the vehicle is assessed in real time based on the multimodal fusion data; The target display format is determined based on the driving load level.

[0012] In some embodiments of this disclosure, displaying the recommended list of voice commands includes: If the driving load level is high, the target display format is determined to be a partial display of the voice command recommendation list; If the driving load level is low, the target display format is determined to be a complete display of the recommended list of voice commands.

[0013] In some embodiments of this disclosure, the method further includes: The system receives user interaction feedback on the voice command recommendation list and updates the dynamic user profile and optimizes the potential intent set in real time based on the interaction feedback.

[0014] A second aspect of this disclosure provides an in-vehicle voice control device, comprising: The first processing module is used to respond to the activation of the voice interaction function in the vehicle, perform data fusion processing on the acquired multimodal data, and obtain corresponding multimodal fusion data, wherein the multimodal data includes user-related data and vehicle-related data. A user profile building module is used to build dynamic user profiles, which include multiple dimensional attributes representing the user's real-time status. The second processing module is used to perform intent prediction processing based on the multimodal fusion data and the dynamic user profile to obtain a set of potential intents; The instruction recommendation module is used to generate a voice instruction recommendation list based on the potential intent set and a preset recommendation strategy, and to display the voice instruction recommendation list.

[0015] In some embodiments of this disclosure, the first processing module includes: The first processing unit is used to extract features from the acquired multimodal data to obtain corresponding multimodal feature vectors, wherein the multimodal feature vectors include user preference feature vectors, context feature vectors, vehicle state feature vectors, and historical interaction feature vectors. The second processing unit is used to perform weighted fusion of the multimodal feature vectors based on the target weights to obtain the multimodal fused data, wherein the target weights are dynamically adjusted according to the multimodal fused data.

[0016] In some embodiments of this disclosure, the dynamic user profile includes static attributes, dynamic attributes, and contextual attributes; The portrait construction module includes: The first construction unit is used to construct the static attributes based on the user preference feature vector; The second construction unit is used to construct the dynamic attribute based on the historical interaction feature vector; The third construction unit is used to construct the context attributes based on the context feature vector and the vehicle state feature vector; The fourth construction unit is used to construct the dynamic user profile based on the static attributes, the dynamic attributes, and the contextual attributes.

[0017] In some embodiments of this disclosure, the second processing module includes: The third processing unit is used to perform intent recognition on the multimodal fusion data and the dynamic user profile to obtain the actual instruction intent; The fourth processing unit is used to predict the actual instruction intent based on the pre-constructed intent association tree to obtain the potential intent set. The intent association tree is obtained by pre-training the multimodal fusion data, and the potential intent set includes multiple potential instruction intents.

[0018] In some embodiments of this disclosure, the fourth processing unit may also be specifically used to calculate the dynamic correlation degree between the actual instruction intent and each potential instruction intent in the intent association tree; and determine a number of potential instruction intents that are ranked first according to the dynamic correlation degree to obtain the potential intent set.

[0019] In some embodiments of this disclosure, the instruction recommendation module includes: The fifth processing unit is used to recommend instructions to the potential intent set based on the preset recommendation strategy, and obtain multiple candidate voice instructions; The sixth processing unit is used to score and sort the multiple candidate voice commands to generate a recommended list of voice commands.

[0020] In some embodiments of this disclosure, the apparatus further includes: The rating assessment module is used to assess the driving load level of the vehicle in real time based on the multimodal fusion data; The display determination module is used to determine the target display format based on the driving load level.

[0021] In some embodiments of this disclosure, the instruction recommendation module includes: The first display unit is used to determine, if the driving load level is high, that the target display format is to partially display the voice command recommendation list; The second display unit is used to determine that if the driving load level is low, the target display format is to fully display the recommended list of voice commands.

[0022] In some embodiments of this disclosure, the apparatus further includes: The real-time optimization module is used to receive user interaction feedback behavior on the voice command recommendation list, and update the dynamic user profile and optimize the potential intent set in real time based on the interaction feedback behavior.

[0023] A third aspect of this disclosure provides a vehicle, including: processor; Memory, used to store executable instructions; The processor is used to read executable instructions from memory and execute the executable instructions to implement the in-vehicle voice control method provided in the first aspect above.

[0024] A fourth aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the in-vehicle voice control method provided in the first aspect.

[0025] A fifth aspect of this disclosure provides a computer program product comprising a computer program or instructions that, when executed by a processor, implement the in-vehicle voice control method of the first aspect described above.

[0026] The technical solution provided in this disclosure has the following advantages: The in-vehicle voice control method, electronic device, and vehicle provided in this disclosure, in response to the activation of the in-vehicle voice interaction function, perform data fusion processing on the acquired multimodal data to obtain corresponding multimodal fusion data. Then, a dynamic user profile is constructed. Based on the multimodal fusion data and the dynamic user profile, intent prediction processing is performed to obtain a set of potential intents. Finally, a voice command recommendation list is generated based on the set of potential intents and a preset recommendation strategy, and the voice command recommendation list is displayed. Thus, when a user uses the in-vehicle voice interaction function, by using multimodal fusion data (including user-related data and vehicle-related data) and the constructed dynamic user profile for intent prediction, and generating a voice command recommendation list based on the obtained set of potential intents and a preset recommendation strategy, intelligent commands can be recommended according to user preferences and the driving environment, improving the personalization and scenario fit of the recommendations, thereby increasing the click-through rate of voice commands and user interaction efficiency. Attached Figure Description

[0027] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0028] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a flowchart of an in-vehicle voice control method provided in an embodiment of this disclosure; Figure 2 This is a flowchart of another in-vehicle voice control method provided in this embodiment of the disclosure; Figure 3 This is a schematic diagram of the structure of an in-vehicle voice control device provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0030] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0031] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0032] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0033] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0034] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0035] With the rapid development of vehicle-to-everything (V2X) technology, in-vehicle voice interaction has become an important human-machine interaction method for improving driving safety and convenience. In-vehicle voice interaction enables voice communication between the driver and the vehicle, providing functions such as navigation, music playback, and phone calls. Among the components of in-vehicle voice interaction, the voice command recommendation function plays a crucial role in helping users quickly get started, guiding them to discover and understand functions, and enhancing the user experience.

[0036] Most mainstream in-vehicle voice systems employ a fixed and static approach to voice command recommendation. A common solution involves determining an optimal frequency set based on the in-vehicle scenario and corresponding recommendation rules, identifying the semantic intent corresponding to this set as the optimal semantic intent, and then determining the voice command for the target text data based on this optimal semantic intent—essentially, simply ranking and recommending commands based on global usage frequency. Therefore, this approach fails to adapt to the personalized usage habits and preferences of different users, and cannot recommend appropriate commands based on the real-time driving environment, thus impacting voice interaction efficiency and user experience. Therefore, how to intelligently recommend commands based on user preferences and the driving environment is a pressing technical problem that those skilled in the art urgently need to solve.

[0037] Therefore, this disclosure provides an in-vehicle voice control method. In response to the activation of the in-vehicle voice interaction function, the method performs data fusion processing on the acquired multimodal data to obtain corresponding multimodal fusion data. Then, a dynamic user profile is constructed. Based on the multimodal fusion data and the dynamic user profile, intent prediction processing is performed to obtain a set of potential intents. Finally, a voice command recommendation list is generated based on the set of potential intents and a preset recommendation strategy, and the voice command recommendation list is displayed. Thus, when a user uses the in-vehicle voice interaction function, the method uses multimodal fusion data (including user-related data and vehicle-related data) and the constructed dynamic user profile to perform intent prediction. Based on the obtained set of potential intents and the preset recommendation strategy, a voice command recommendation list is generated. This allows for intelligent recommendation of commands based on user preferences and the driving environment, improving the personalization and scenario fit of the recommendations, thereby increasing the click-through rate of voice commands and the efficiency of user interaction.

[0038] Figure 1 This is a flowchart of an in-vehicle voice control method provided in an embodiment of the present disclosure. The method can be applied to an electronic device, which can be a device integrating multiple processors, such as a human-machine user terminal (HUT).

[0039] like Figure 1 As shown in the embodiments of this disclosure, the in-vehicle voice control method can be applied to the field of vehicle control technology. The in-vehicle voice control method may include the following steps.

[0040] S110. In response to the activation of the voice interaction function in the vehicle, the acquired multimodal data is subjected to data fusion processing to obtain corresponding multimodal fusion data, wherein the multimodal data includes user-related data and vehicle-related data.

[0041] In this embodiment of the disclosure, the electronic device responds to the activation of the voice interaction function in the vehicle by performing data fusion processing on the acquired multimodal data to obtain the corresponding multimodal fused data.

[0042] Optionally, the voice interaction function allows users to interact with the vehicle's in-vehicle equipment via voice, enabling the vehicle to understand, process, and respond to human voice commands, thereby achieving convenient control and information acquisition.

[0043] Optionally, multimodal data may include user-related data and vehicle-related data. User-related data may include historical valid audio data (such as high-frequency commands (including colloquial expressions), domain classification labels, command execution success rate, skip / cancel behavior), user-related account data (such as mobile app favorites (music / addresses), calendar trips, frequently used contacts, consumption preferences), and real-time user status (such as voice emotion characteristics, driving fatigue (assessed through steering wheel operation frequency)). Vehicle-related data may include real-time vehicle data (such as vehicle speed, battery / fuel level, air conditioning status, window status, lighting mode, fault codes, seat position), current interface information (such as active applications (music / navigation / phone, etc.), interface hierarchy, current operation progress), and environmental perception data (such as time (commuting / weekend), geographical location, weather, road conditions), etc., without limitation. Multimodal data can be a dataset composed of data from two or more different modalities. Modality refers to the form or type of data, such as text, images, audio, video, sensor data, etc.

[0044] Optionally, multimodal fusion data can be fused data obtained by integrating and analyzing multimodal data (such as images, text, audio, video, sensor signals, etc.) from different sources and in different forms.

[0045] Specifically, the electronic device can detect and respond to the activation of the voice interaction function within the vehicle, acquiring multimodal data, such as user-related data and vehicle-related data. The acquisition methods can include, but are not limited to, local storage log acquisition, sensor acquisition, vehicle-to-everything (V2X) acquisition, etc., without limitation here. The electronic device can then perform data fusion processing on the acquired multimodal data, such as data preprocessing, data feature extraction, and weighted data fusion processing, without limitation here, to obtain the corresponding multimodal fused data.

[0046] S120. Construct a dynamic user profile, wherein the dynamic user profile includes multiple dimensional attributes that characterize the user's real-time status.

[0047] In this embodiment of the disclosure, the electronic device can construct a dynamic user profile.

[0048] Optionally, a dynamic user profile may include multiple dimensional attributes that characterize the user's real-time state. Here, a dynamic user profile refers to a user model built based on real-time user behavior and change data, capable of continuous updates to reflect the user's current interests, state, or intentions, etc., without limitation here.

[0049] Specifically, electronic devices can build dynamic user profiles based on acquired multimodal data, such as user-related data (e.g., user's historical valid audio data, user's associated account data, user's real-time status, etc.) to continuously update and reflect the user's current interests, status, or intentions.

[0050] S130. Based on the multimodal fusion data and the dynamic user profile, perform intent prediction processing to obtain a potential intent set.

[0051] In this embodiment of the disclosure, the electronic device can perform intent prediction processing based on the multimodal fusion data and the dynamic user profile to obtain a set of potential intents.

[0052] Optionally, the set of potential intents can be a collection of multiple potential instruction intents. These potential instruction intents can be intents associated with the user's actual instruction intent. For example, if the user's actual instruction intent is "turn on music," the potential instruction intents could be "switch artists," "adjust volume," etc., without limitation.

[0053] Specifically, after obtaining multimodal fusion data and dynamic user profiles, electronic devices can perform intent prediction processing on the multimodal fusion data and dynamic user profiles to obtain a potential intent set that includes multiple potential instruction intents.

[0054] S140. Generate a voice command recommendation list based on the potential intent set and the preset recommendation strategy, and display the voice command recommendation list.

[0055] In this embodiment of the disclosure, the electronic device can generate a voice command recommendation list based on the potential intent set and a preset recommendation strategy, and display the voice command recommendation list.

[0056] Optionally, the preset recommendation strategy can be a pre-defined instruction recommendation logic. For example, the preset recommendation strategy can be content-based recommendation logic, collaborative filtering-based recommendation logic, deep reinforcement learning-based recommendation logic, etc., and there is no limitation here.

[0057] Optionally, the voice command recommendation list can be a list that includes multiple recommended voice commands.

[0058] Specifically, after obtaining a set of potential intentions, the electronic device can generate a voice command recommendation list according to a preset recommendation strategy, such as content-based recommendation logic, collaborative filtering-based recommendation logic, or deep reinforcement learning-based recommendation logic. The voice command recommendation list can include multiple recommended voice commands, and the voice command recommendation list is displayed for the user to select by touch screen or voice.

[0059] Therefore, in response to the activation of the in-vehicle voice interaction function, the acquired multimodal data is fused to obtain corresponding multimodal fused data. A dynamic user profile is then constructed. Based on the multimodal fused data and the dynamic user profile, intent prediction is performed to obtain a set of potential intents. Finally, a voice command recommendation list is generated based on the set of potential intents and a preset recommendation strategy, and this list is displayed. Thus, when a user uses the in-vehicle voice interaction function, the multimodal fused data (including user-related and vehicle-related data) and the constructed dynamic user profile are used for intent prediction. Based on the obtained set of potential intents and the preset recommendation strategy, a voice command recommendation list is generated. This allows for intelligent recommendation of commands based on user preferences and the driving environment, improving the personalization and contextual relevance of the recommendations, thereby increasing the click-through rate of voice commands and the efficiency of user interaction.

[0060] Optionally, S110 may specifically include: extracting features from the acquired multimodal data to obtain corresponding multimodal feature vectors, wherein the multimodal feature vectors include user preference feature vectors, context feature vectors, vehicle state feature vectors, and historical interaction feature vectors; and performing weighted fusion of the multimodal feature vectors based on target weights to obtain the multimodal fused data, wherein the target weights are dynamically adjusted according to the multimodal fused data.

[0061] In this embodiment of the disclosure, the electronic device can perform feature extraction on the acquired multimodal data to obtain the corresponding multimodal feature vector.

[0062] Optionally, the multimodal feature vector may include user preference feature vector, context feature vector, vehicle state feature vector, and historical interaction feature vector.

[0063] Specifically, after acquiring multimodal data, electronic devices can first perform data preprocessing on the multimodal data. For example, for historical valid audio data of users, the original audio stream is preprocessed based on spectral subtraction and then input into a lightweight CNN-BLSTM hybrid model. This model simultaneously extracts MFCC features and Mel spectrogram features, improving the robustness of command endpoint detection and speech recognition in a typical in-vehicle noise environment of 60-80dB. The NLP module uses an attention-enhanced Bi-LSTM model to accurately map colloquial expressions such as "cool down the temperature" to standardized commands such as "cool down the air conditioner."

[0064] Furthermore, after data preprocessing, the electronic device can perform feature extraction on the acquired multimodal data to obtain the corresponding multimodal feature vector. For example, a 128-dimensional multimodal feature vector can be constructed, including user preference feature vector (30-dimensional), context feature vector (48-dimensional), vehicle state feature vector (30-dimensional), and historical interaction feature vector (20-dimensional).

[0065] Furthermore, the electronic device can perform weighted fusion of the multimodal feature vectors based on target weights to obtain the multimodal fusion data. For example, the electronic device can use a lightweight cross-modal attention network to perform weighted fusion of the multimodal feature vectors, wherein the target weights are dynamically adjusted according to the multimodal fusion data. For example, if it is detected that a vehicle is driving on a highway (from GPS and vehicle speed) and the rain sensor is triggered, the attention weights of the context feature vector and the vehicle state feature vector will automatically increase to above 0.6, while the weight of the user preference feature vector may be temporarily reduced, making the recommendation logic more inclined to generate instructions that are strongly related to safety and comfort, such as "close the windows", "turn on the wipers" or "switch to a clearer radio station".

[0066] Therefore, by performing fine-grained feature extraction on multimodal data, feature vectors such as user preferences, context, vehicle status, and historical interactions are obtained. Weighted fusion is then performed using weights dynamically adjusted based on the multimodal fusion data, making the data fusion process more than a simple stitching. This mechanism adaptively enhances the influence of features most relevant to the current scene (for example, automatically increasing the weights related to windows and wipers in vehicle status and environmental features when rain is detected), thereby generating multimodal fusion data that better reflects core needs and laying a solid foundation for subsequent accurate intent prediction.

[0067] Optionally, the dynamic user profile includes static attributes, dynamic attributes, and contextual attributes. Static attributes include frequently used account preferences (music genre, navigation habits), fixed itineraries (commuting routes), and basic vehicle control preferences (default air conditioning temperature), etc.; dynamic attributes include recent high-frequency needs (such as navigation / music preferences related to weekend camping) and temporary interests (such as recently searched attractions); contextual attributes include current driving status (highway / traffic jam), environmental conditions (rainy day / night), and in-car scenarios (with children / alone), etc.

[0068] Optionally, S120 may specifically include: constructing the static attribute based on the user preference feature vector; constructing the dynamic attribute based on the historical interaction feature vector; constructing the contextual attribute based on the contextual feature vector and the vehicle state feature vector; and constructing the dynamic user profile based on the static attribute, the dynamic attribute, and the contextual attribute.

[0069] In this embodiment of the disclosure, the electronic device can construct the static attributes based on the user preference feature vector, such as generating "preferred music type is classical music" or "has gym navigation records every Wednesday night" based on long-term statistical user preference feature vectors, and construct the static attributes (including frequently used account preferences (music type, navigation habits), fixed trips (commuting routes), basic vehicle control preferences (default air conditioning temperature), etc.); the electronic device can construct the dynamic attributes based on historical interaction feature vectors, such as generating them based on recent (e.g., within 24 hours) historical interaction feature vectors, such as "queried coffee shops multiple times this morning" or "just skipped music recommendations", and construct the dynamic attributes (including recent high-frequency needs (e.g., navigation / music preferences related to weekend camping) or temporary interests (e.g., recently searched attractions)); the electronic device can construct the contextual attributes based on the contextual feature vector and the vehicle state feature vector, such as generating them based on real-time generated contextual feature vectors and vehicle state feature vectors, such as "currently in highway driving mode" or "outside temperature is 35 degrees Celsius", and construct the contextual attributes (including current driving state (highway / congestion), environmental conditions (rainy day / night), in-car scenario (with child / alone)), etc.

[0070] Furthermore, the electronic device can obtain a corresponding dynamic user profile based on the static attributes, dynamic attributes, and contextual attributes constructed above. This dynamic user profile is a user model built based on real-time user behavior and change data, and it can be continuously updated to reflect the user's current interests, state, or intentions.

[0071] Therefore, by clearly dividing dynamic user profiles into static attributes, dynamic attributes, and contextual attributes, and constructing them based on feature vectors of different categories, user profiles are no longer flat sets of labels, but three-dimensional models with time and state dimensions. Static attributes reflect long-term stable preferences, dynamic attributes capture recent changes in interests, and contextual attributes characterize instantaneous states. The combination of these three aspects enables the system to have a deep and timely understanding of users, providing an accurate model basis for generating personalized recommendation instructions.

[0072] Optionally, S130 may specifically include: performing intent recognition on the multimodal fusion data and the dynamic user profile to obtain the actual instruction intent; performing association prediction on the actual instruction intent based on a pre-built intent association tree to obtain the potential intent set, wherein the intent association tree is pre-trained from the multimodal fusion data, and the potential intent set includes multiple potential instruction intents.

[0073] In this embodiment of the disclosure, the electronic device can first perform intent recognition on the multimodal fusion data and the dynamic user profile to obtain the actual instruction intent. For example, the multimodal fusion data and the dynamic user profile can be input into a pre-trained intent prediction model (such as a Long Short-Term Memory (LSTM) model) so that the intent prediction model can perform intent recognition to obtain the corresponding actual instruction intent (such as recognizing "I am a little hot" as "adjust the air conditioner" intent).

[0074] Furthermore, the electronic device can predict the actual instruction intent based on a pre-built intent association tree to obtain the potential intent set.

[0075] Optionally, the intent association tree can be pre-trained from multimodal fusion data. For example, the intent association tree can be trained by mining the sequence and co-occurrence relationships between different intents through massive amounts of anonymized driving data. In the intent association tree, the intent "adjust the air conditioning" is often associated with the intent "close the windows" and the intent "turn on the seat ventilation".

[0076] Specifically, the electronic device performs intent recognition on the multimodal fusion data and the dynamic user profile based on the intent prediction model. After obtaining the actual instruction intent, it can continue to perform association prediction on the actual instruction intent based on the pre-built intent association tree through the intent prediction model to obtain the potential intent set. For example, if the actual instruction intent is "adjust the air conditioner", the corresponding potential intent set may include the intent to "close the car window", the intent to "turn on the seat ventilation", etc., which are not limited here.

[0077] This not only improves the accuracy of services for users' explicit needs, but more importantly, it enables in-depth exploration and proactive fulfillment of users' implicit and continuous needs, greatly enhancing the system's intelligence and user-friendliness.

[0078] Optionally, the potential intent set is obtained by predicting the association between the actual instruction intent and each potential instruction intent in the intent association tree based on the pre-constructed intent association tree, including: calculating the dynamic association degree between the actual instruction intent and each potential instruction intent in the intent association tree; and determining the top-ranked potential instruction intents based on the dynamic association degree to obtain the potential intent set.

[0079] In this embodiment of the disclosure, the electronic device can calculate the dynamic correlation between the actual instruction intent and each potential instruction intent in the intent association tree.

[0080] Optionally, dynamic correlation is used to characterize the correlation score between the actual instruction intent and each potential instruction intent.

[0081] Specifically, after receiving the actual instruction intent, the electronic device can calculate the dynamic correlation degree between the actual instruction intent and each potential instruction intent in the intent association tree. For example, if the actual instruction intent is "adjust the air conditioner", according to the preset calculation formula, the correlation score with potential instruction intent A (quickly reduce the temperature to 22°C and increase the fan speed) is 92 points; the correlation score with potential instruction intent B (close all windows) is 70 points; the correlation score with potential instruction intent B (open the sunroof sunshade) is 40 points, etc., which is to obtain the dynamic correlation degree.

[0082] Furthermore, the electronic device can determine the top-ranked potential instruction intents based on the dynamic correlation, thereby obtaining the potential intent set.

[0083] For example, electronic devices can sort potential instruction intentions from high to low according to their dynamic correlation scores. For instance, they can identify the top-ranked potential instruction intentions, such as selecting the top 5 potential instruction intentions, thereby obtaining multiple potential instruction intentions with high correlation scores that include multiple dynamic correlations.

[0084] Therefore, by taking into account the real-time driving environment, vehicle status, and the user's ever-changing personal state, the system ensures that the recommended potential intent is the most relevant and valuable option for the specific user in the current context, thereby improving the overall in-vehicle voice interaction experience and safety.

[0085] Optionally, S140 may include: recommending instructions to the potential intent set based on the preset recommendation strategy to obtain multiple candidate voice instructions; and performing scoring and sorting processing on the multiple candidate voice instructions to generate the voice instruction recommendation list.

[0086] In this embodiment of the disclosure, the electronic device can recommend instructions to the set of potential intentions based on the preset recommendation strategy to obtain multiple candidate voice instructions.

[0087] Optionally, the preset recommendation strategy may include content-based recommendation logic, collaborative filtering-based recommendation logic, deep reinforcement learning-based recommendation logic, etc.

[0088] Specifically, electronic devices can recommend instructions for each potential instruction intent in the potential intent set based on the preset recommendation strategy, such as content-based recommendation logic: matching the similarity between the user's historical high-frequency instructions and the current scene (e.g., recommending commonly used navigation instructions during commuting); collaborative filtering-based recommendation logic: mining the contextual instruction preferences of "similar user groups" (e.g., most users choose "play soothing music" during traffic jams); and deep reinforcement learning-based recommendation logic: optimizing long-term interaction rewards through the DQN model, i.e., the Deep Q-Network model, to avoid recommendation homogenization (e.g., automatically switching recommendation domains after a user skips similar instructions consecutively), and obtaining multiple candidate voice instructions.

[0089] Furthermore, the electronic device can score and rank the multiple candidate voice commands, such as scoring the multiple candidate voice commands separately to obtain recommendation scores for the multiple candidate voice commands, and then ranking the multiple candidate voice commands based on the recommendation scores, such as using a weighted ranking model (such as LambdaMART) to fuse the scores of the three channels to generate a ranked recommendation list, that is, a voice command recommendation list, which displays multiple voice commands in order of different recommendation scores.

[0090] Therefore, after generating candidate commands through a preset recommendation strategy, a unified scoring and sorting process is performed to ensure that the final list of recommended voice commands presented to the user is optimized and prioritized. This avoids presenting the user with a large number of possible options in an unordered manner, instead placing the most likely to be adopted and most suitable commands for the current scenario at the forefront, effectively reducing the user's choice cost, improving decision-making efficiency, and making the interaction smoother and more natural.

[0091] Optionally, the in-vehicle voice control method may further include: assessing the vehicle's driving load level in real time based on the multimodal fusion data; and determining the target display format based on the driving load level.

[0092] In this embodiment, the electronic device can assess the driving load level of the vehicle in real time based on the multimodal fusion data. For example, it can assess the driving load level of the vehicle in real time based on vehicle speed, acceleration, steering angle rate, etc. When the vehicle speed is >80km / h or a sudden acceleration / sudden braking is detected, the driving load level is determined to be high; when the vehicle is stopped or in a low-speed congested state, the driving load level is determined to be low, etc., without limitation.

[0093] Furthermore, the electronic device can determine the target display format based on the driving load level. For example, when the driving load level is determined to be high, the highest priority recommended instruction is output in a concise voice broadcast format; when the driving load level is determined to be low, multiple candidate voice instructions are displayed on the vehicle screen in the form of a visual list, etc., without limitation.

[0094] Therefore, by determining the target display format through real-time assessment of driving load levels, driving safety is fundamentally guaranteed while improving interactive intelligence, achieving a balance between efficiency and safety.

[0095] Optionally, S140 may specifically include: if the driving load level is high, determining that the target display format is to partially display the voice command recommendation list; if the driving load level is low, determining that the target display format is to fully display the voice command recommendation list.

[0096] In some embodiments of this disclosure, if the driving load level is high, the electronic device may determine that the target display format is to partially display the recommended list of voice commands. For example, when the driving load level is determined to be high, only the first-ranked command may be spoken in the voice broadcast (e.g., "Battery level is below 20%, do you need to navigate to the nearest charging station?"), etc., which is not limited here.

[0097] In other embodiments of this disclosure, if the driving load level is low, the electronic device can determine that the target display format is to fully display the recommended list of voice commands. For example, if the driving load level is determined to be low, 2-4 candidate commands with icons are displayed as a foldable "Smart Suggestion" card at the bottom or sidebar of the vehicle screen, along with a brief recommendation reason, such as "You often listen to the news at this time," etc., without limitation.

[0098] Therefore, in high-risk driving scenarios (such as highways and traffic jams), the system automatically simplifies information output and provides core recommendations via voice broadcast, minimizing deviations in the driver's line of sight and cognition. In safe scenarios (such as waiting in a park), it provides a richer visual list for users to browse. This mechanism enhances interactive intelligence while fundamentally ensuring driving safety, achieving a balance between efficiency and safety.

[0099] Optionally, the in-vehicle voice control method may further include: receiving user interaction feedback behavior on the voice command recommendation list, and updating the dynamic user profile and optimizing the potential intent set in real time based on the interaction feedback behavior.

[0100] In this embodiment, the electronic device can receive user interaction feedback on the voice command recommendation list and update the dynamic user profile and optimize the potential intent set in real time based on the interaction feedback. For example, the electronic device can receive user interaction feedback on the voice command recommendation list, which includes all explicit user behaviors (clicking, voice confirmation) and implicit behaviors (card exposure for 3 seconds without operation, directly speaking other commands), etc. The electronic device can update the dynamic user profile and optimize the potential intent set in real time based on the interaction feedback, that is, fine-tune the attributes of the dynamic user profile and the parameters of the intent prediction model in real time to achieve closed-loop optimization.

[0101] Therefore, the real-time optimization loop based on user interaction feedback improves the accuracy of recommended instructions and aligns with users' personalized habits.

[0102] Figure 2 This is a flowchart illustrating another in-vehicle voice control method provided in this embodiment.

[0103] like Figure 2 As shown, the in-vehicle voice control method may include the following steps.

[0104] S210, in response to the activation of the voice interaction function in the vehicle, performs data fusion processing on the acquired multimodal data to obtain the corresponding multimodal fusion data.

[0105] In this embodiment of the disclosure, in response to the activation of the voice interaction function in the vehicle, the electronic device can preprocess the multimodal data after acquiring it. Then, feature extraction is performed on the acquired multimodal data to obtain corresponding multimodal feature vectors. For example, a 128-dimensional multimodal feature vector can be constructed, including a user preference feature vector (30 dimensions), a contextual feature vector (48 dimensions), a vehicle state feature vector (30 dimensions), and a historical interaction feature vector (20 dimensions). Finally, the electronic device can perform weighted fusion of the multimodal feature vectors based on the target weights to obtain the multimodal fusion data. For example, the electronic device can use a lightweight cross-modal attention network to perform weighted fusion of the multimodal feature vectors, wherein the target weights are dynamically adjusted according to the multimodal fusion data. For example, if it is detected that a vehicle is driving on a highway (from GPS and vehicle speed) and the rain sensor is triggered, the attention weights of the context feature vector and the vehicle state feature vector will automatically increase to above 0.6, while the weight of the user preference feature vector may be temporarily reduced, making the recommendation logic more inclined to generate instructions that are strongly related to safety and comfort, such as "close the windows", "turn on the wipers" or "switch to a clearer radio station".

[0106] S220. Construct static attributes based on user preference feature vectors, construct dynamic attributes based on historical interaction feature vectors, construct contextual attributes based on contextual feature vectors and vehicle state feature vectors, and construct a dynamic user profile based on static attributes, dynamic attributes, and contextual attributes.

[0107] In this embodiment of the disclosure, the electronic device can construct the static attributes based on the user preference feature vector, such as generating "preferred music type is classical music" or "has gym navigation records every Wednesday night" based on long-term statistical user preference feature vectors, and construct the static attributes (including frequently used account preferences (music type, navigation habits), fixed trips (commuting routes), basic vehicle control preferences (default air conditioning temperature), etc.); the electronic device can construct the dynamic attributes based on historical interaction feature vectors, such as generating them based on recent (e.g., within 24 hours) historical interaction feature vectors, such as "queried coffee shops multiple times this morning" or "just skipped music recommendations", and construct the dynamic attributes (including recent high-frequency needs (e.g., navigation / music preferences related to weekend camping) or temporary interests (e.g., recently searched attractions)); the electronic device can construct the contextual attributes based on the contextual feature vector and the vehicle state feature vector, such as generating them based on real-time generated contextual feature vectors and vehicle state feature vectors, such as "currently in highway driving mode" or "outside temperature is 35 degrees Celsius", and construct the contextual attributes (including current driving state (highway / congestion), environmental conditions (rainy day / night), in-car scenario (with child / alone)), etc.

[0108] Furthermore, the electronic device can obtain a corresponding dynamic user profile based on the static attributes, dynamic attributes, and contextual attributes constructed above. This dynamic user profile is a user model built based on real-time user behavior and change data, and it can be continuously updated to reflect the user's current interests, state, or intentions.

[0109] S230. Perform intent recognition on multimodal fusion data and dynamic user profiles to obtain actual instruction intents. Based on the pre-built intent association tree, perform association prediction on the actual instruction intents to obtain a set of potential intents.

[0110] In this embodiment, the electronic device can first perform intent recognition on the multimodal fusion data and the dynamic user profile to obtain the actual instruction intent, such as recognizing "I'm a little hot" as the intent to "adjust the air conditioning". Further, the electronic device can continue to use an intent prediction model to predict the actual instruction intent based on a pre-built intent association tree to obtain the potential intent set. For example, if the actual instruction intent is "adjust the air conditioning", the corresponding potential intent set may include intents such as "close the windows" and "turn on the seat ventilation", etc., without limitation.

[0111] After receiving the actual instruction intent, the electronic device can calculate the dynamic correlation degree between the actual instruction intent and each potential instruction intent in the intent association tree. For example, if the actual instruction intent is "adjust the air conditioning," according to a preset calculation formula, the correlation score with potential instruction intent A (quickly lower the temperature to 22°C and increase the fan speed) is 92 points; the correlation score with potential instruction intent B (close all windows) is 70 points; and the correlation score with potential instruction intent B (open the sunroof sunshade) is 40 points, etc., thus obtaining the dynamic correlation degree. Further, the potential instruction intents are sorted from high to low according to their dynamic correlation scores. For example, the top 5 potential instruction intents are selected, thus obtaining multiple potential instruction intents with high correlation scores and multiple dynamic correlation degrees.

[0112] S240. Based on a preset recommendation strategy, recommend instructions to the potential intent set to obtain multiple candidate voice instructions. Then, score and sort the multiple candidate voice instructions to generate a voice instruction recommendation list.

[0113] In this embodiment, the electronic device can recommend commands for each potential command intent in the potential intent set based on the preset recommendation strategy. This includes content-based recommendation logic: matching the similarity between a user's historical high-frequency commands and the current scenario (e.g., recommending commonly used navigation commands during commuting); collaborative filtering-based recommendation logic: mining the scenario-based command preferences of "similar user groups" (e.g., most users choose "play soothing music" during traffic jams); and deep reinforcement learning-based recommendation logic: optimizing long-term interaction rewards through a DQN model (Deep Q-Network) to avoid recommendation homogenization (e.g., automatically switching recommendation domains after a user skips similar commands consecutively), resulting in multiple candidate voice commands. Furthermore, the electronic device can perform scoring and ranking processing on these multiple candidate voice commands. For example, it can score each candidate voice command separately to obtain recommendation scores, and then rank the multiple candidate voice commands based on these scores. For instance, it can fuse the scores from three channels using a weighted ranking model (e.g., LambdaMART) to generate a ranked recommendation list, thus obtaining a voice command recommendation list. This voice command recommendation list displays multiple voice commands sorted according to different recommendation scores.

[0114] S250: Real-time assessment of vehicle driving load level based on multimodal fusion data, and determination of target display format based on driving load level.

[0115] In this embodiment, the electronic device can assess the vehicle's driving load level in real time based on the multimodal fusion data. For example, it can assess the driving load level based on vehicle speed, acceleration, steering angle rate, etc. When the vehicle speed is >80km / h or sudden acceleration / braking is detected, the driving load level is determined to be high; when the vehicle is stopped or in a low-speed congested state, the driving load level is determined to be low, etc., without limitation. Furthermore, the electronic device can determine the target display format based on the driving load level. For example, when the driving load level is determined to be high, the highest priority recommended command is output in a concise voice broadcast format; when the driving load level is determined to be low, multiple candidate voice commands are displayed on the vehicle screen in the form of a visual list, etc., without limitation.

[0116] S260. If the driving load level is high, the target display format is to partially display the recommended list of voice commands; if the driving load level is low, the target display format is to fully display the recommended list of voice commands.

[0117] In some embodiments of this disclosure, if the driving load level is high, the electronic device may determine that the target display format is to partially display the recommended list of voice commands. For example, when the driving load level is determined to be high, only the first-ranked command may be spoken in the voice broadcast (e.g., "Battery level is below 20%, do you need to navigate to the nearest charging station?"), etc., which is not limited here.

[0118] In other embodiments of this disclosure, if the driving load level is low, the electronic device can determine that the target display format is to fully display the recommended list of voice commands. For example, if the driving load level is determined to be low, 2-4 candidate commands with icons are displayed as a foldable "Smart Suggestion" card at the bottom or sidebar of the vehicle screen, along with a brief recommendation reason, such as "You often listen to the news at this time," etc., without limitation.

[0119] S270: Receive user interaction feedback on the voice command recommendation list, and update the dynamic user profile and optimize the potential intent set in real time based on the interaction feedback.

[0120] In this embodiment, the electronic device can receive user interaction feedback on the voice command recommendation list and update the dynamic user profile and optimize the potential intent set in real time based on the interaction feedback. For example, the electronic device can receive user interaction feedback on the voice command recommendation list, which includes all explicit user behaviors (clicking, voice confirmation) and implicit behaviors (card exposure for 3 seconds without operation, directly speaking other commands), etc. The electronic device can update the dynamic user profile and optimize the potential intent set in real time based on the interaction feedback, that is, fine-tune the attributes of the dynamic user profile and the parameters of the intent prediction model in real time to achieve closed-loop optimization.

[0121] Figure 3 This is a schematic diagram of the structure of an in-vehicle voice control device provided in an embodiment of this disclosure.

[0122] In this embodiment, the in-vehicle voice control device can be installed within an electronic device and is understood as a functional module within the aforementioned electronic device. Specifically, the electronic device can be a server or a terminal, wherein the terminal specifically includes an in-vehicle terminal, a computer, or a tablet computer, etc., without limitation.

[0123] like Figure 3 As shown, the in-vehicle voice control device 300 may include a first processing module 310, a profile construction module 320, a second processing module 330, and an instruction recommendation module 340.

[0124] The first processing module 310 is used to respond to the activation of the voice interaction function in the vehicle by performing data fusion processing on the acquired multimodal data to obtain corresponding multimodal fusion data, wherein the multimodal data includes user-related data and vehicle-related data. The profile building module 320 is used to build a dynamic user profile, which includes multiple dimensional attributes that represent the real-time status of the user. The second processing module 330 is used to perform intent prediction processing based on the multimodal fusion data and the dynamic user profile to obtain a set of potential intents; The instruction recommendation module 340 is used to generate a voice instruction recommendation list based on the potential intent set and the preset recommendation strategy, and to display the voice instruction recommendation list.

[0125] In some embodiments of this disclosure, the first processing module 310 includes: The first processing unit is used to extract features from the acquired multimodal data to obtain corresponding multimodal feature vectors, wherein the multimodal feature vectors include user preference feature vectors, context feature vectors, vehicle state feature vectors, and historical interaction feature vectors. The second processing unit is used to perform weighted fusion of the multimodal feature vectors based on the target weights to obtain the multimodal fused data, wherein the target weights are dynamically adjusted according to the multimodal fused data.

[0126] In some embodiments of this disclosure, the dynamic user profile includes static attributes, dynamic attributes, and contextual attributes; The portrait construction module 320 includes: The first construction unit is used to construct the static attributes based on the user preference feature vector; The second construction unit is used to construct the dynamic attribute based on the historical interaction feature vector; The third construction unit is used to construct the context attributes based on the context feature vector and the vehicle state feature vector; The fourth construction unit is used to construct the dynamic user profile based on the static attributes, the dynamic attributes, and the contextual attributes.

[0127] In some embodiments of this disclosure, the second processing module 330 includes: The third processing unit is used to perform intent recognition on the multimodal fusion data and the dynamic user profile to obtain the actual instruction intent; The fourth processing unit is used to predict the actual instruction intent based on the pre-constructed intent association tree to obtain the potential intent set. The intent association tree is obtained by pre-training the multimodal fusion data, and the potential intent set includes multiple potential instruction intents.

[0128] In some embodiments of this disclosure, the fourth processing unit may also be specifically used to calculate the dynamic correlation degree between the actual instruction intent and each potential instruction intent in the intent association tree; and determine a number of potential instruction intents that are ranked first according to the dynamic correlation degree to obtain the potential intent set.

[0129] In some embodiments of this disclosure, the instruction recommendation module 340 includes: The fifth processing unit is used to recommend instructions to the potential intent set based on the preset recommendation strategy, and obtain multiple candidate voice instructions; The sixth processing unit is used to score and sort the multiple candidate voice commands to generate a recommended list of voice commands.

[0130] In some embodiments of this disclosure, the device 300 further includes: The rating assessment module is used to assess the driving load level of the vehicle in real time based on the multimodal fusion data; The display determination module is used to determine the target display format based on the driving load level.

[0131] In some embodiments of this disclosure, the instruction recommendation module 340 includes: The first display unit is used to determine, if the driving load level is high, that the target display format is to partially display the voice command recommendation list; The second display unit is used to determine that if the driving load level is low, the target display format is to fully display the recommended list of voice commands.

[0132] In some embodiments of this disclosure, the device 300 further includes: The real-time optimization module is used to receive user interaction feedback behavior on the voice command recommendation list, and update the dynamic user profile and optimize the potential intent set in real time based on the interaction feedback behavior.

[0133] It should be noted that, Figure 3 The in-vehicle voice control device 300 shown can execute the various steps in the above method embodiments and realize the various processes and effects in the above method embodiments, which will not be elaborated here.

[0134] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.

[0135] In this embodiment of the disclosure, Figure 4 The electronic device shown can be a server or a terminal. Specifically, the terminal includes in-vehicle terminals, computers, or tablets, etc., without limitation.

[0136] like Figure 4 As shown, the electronic device may include a processor 410 and a memory 420 storing computer program instructions.

[0137] Specifically, the processor 410 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this disclosure.

[0138] Memory 420 may include a large-capacity storage device for information or instructions. For example, and not limitingly, memory 420 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 420 may include removable or non-removable (or fixed) media. Where appropriate, memory 420 may be internal or external to the integrated gateway device. In a particular embodiment, memory 420 is a non-volatile solid-state memory. In a particular embodiment, memory 420 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (Electrically Programmable ROM, EPROM), an electrically erasable programmable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0139] The processor 410 reads and executes computer program instructions stored in the memory 420 to perform the steps of the in-vehicle voice control method provided in this embodiment of the present disclosure.

[0140] In one example, the electronic device may also include a transceiver 430 and a bus 440. Wherein, as... Figure 4 As shown, the processor 410, memory 420 and transceiver 430 are connected via bus 440 and communicate with each other.

[0141] Bus 440 may include hardware, software, or both. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 440 may include one or more buses.

[0142] This disclosure also provides a computer-readable storage medium that can store a computer program, which, when executed by a processor, enables the processor to implement the in-vehicle voice control method provided in this disclosure.

[0143] When the computer program is executed by the processor, it can perform the following steps: In response to the activation of the voice interaction function in the vehicle, it performs data fusion processing on the acquired multimodal data to obtain corresponding multimodal fusion data; then, it constructs a dynamic user profile; based on the multimodal fusion data and the dynamic user profile, it performs intent prediction processing to obtain a set of potential intents; finally, it generates a voice command recommendation list based on the set of potential intents and a preset recommendation strategy, and displays the voice command recommendation list. Thus, when a user uses the voice interaction function in the vehicle, by using multimodal fusion data (including user-related data and vehicle-related data) and the constructed dynamic user profile to predict intents, and generating a voice command recommendation list based on the obtained set of potential intents and a preset recommendation strategy, it can intelligently recommend commands according to user preferences and driving environment, improving the personalization and scenario fit of the recommendations, thereby increasing the click-through rate of voice commands and user interaction efficiency.

[0144] The aforementioned storage medium may, for example, include a memory 420 containing computer program instructions, which can be executed by a processor 410 of an electronic device to complete the in-vehicle voice control method provided in this embodiment. Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as read-only memory (ROM), random access memory (RAM), external cache memory, compact disc ROM (CD-ROM), magnetic tape, floppy disk, flash memory, and optical data storage devices. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM).

[0145] This disclosure also provides a vehicle that includes electronic devices that can implement the various processes and effects described in the above embodiments of this disclosure, which will not be elaborated here.

[0146] This disclosure also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, they implement the in-vehicle voice control method provided in this disclosure and can achieve the various processes and effects in the above embodiments of this disclosure, which will not be elaborated here.

[0147] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A vehicle-mounted voice control method, characterized in that, The method includes: In response to the activation of the in-vehicle voice interaction function, the acquired multimodal data is subjected to data fusion processing to obtain corresponding multimodal fused data, which includes user-related data and vehicle-related data; Construct a dynamic user profile, which includes multiple dimensional attributes representing the user's real-time status; Based on the multimodal fusion data and the dynamic user profile, intent prediction processing is performed to obtain a set of potential intents; A voice command recommendation list is generated based on the set of potential intents and a preset recommendation strategy, and the voice command recommendation list is displayed.

2. The method according to claim 1, characterized in that, The process of fusing the acquired multimodal data to obtain corresponding multimodal fused data includes: Feature extraction is performed on the acquired multimodal data to obtain corresponding multimodal feature vectors, which include user preference feature vectors, context feature vectors, vehicle state feature vectors, and historical interaction feature vectors. The multimodal feature vectors are weighted and fused based on the target weights to obtain the multimodal fused data, and the target weights are dynamically adjusted according to the multimodal fused data.

3. The method according to claim 2, characterized in that, The dynamic user profile includes static attributes, dynamic attributes, and contextual attributes; The construction of dynamic user profiles includes: The static attributes are constructed based on the user preference feature vector; The dynamic attributes are constructed based on the historical interaction feature vectors. The context attributes are constructed based on the context feature vector and the vehicle state feature vector; The dynamic user profile is constructed based on the static attributes, the dynamic attributes, and the contextual attributes.

4. The method according to claim 1, characterized in that, The intent prediction process based on the multimodal fusion data and the dynamic user profile yields a set of potential intents, including: Intent recognition is performed on the multimodal fusion data and the dynamic user profile to obtain the actual instruction intent; The actual instruction intent is predicted based on a pre-constructed intent association tree to obtain the potential intent set. The intent association tree is pre-trained from the multimodal fusion data, and the potential intent set includes multiple potential instruction intents.

5. The method according to claim 4, characterized in that, The process of predicting the actual instruction intent based on a pre-built intent association tree to obtain the potential intent set includes: Calculate the dynamic correlation degree between the actual instruction intent and each potential instruction intent in the intent association tree; Based on the dynamic correlation, the top-ranked potential instruction intentions are determined to obtain the potential intention set.

6. The method according to claim 1, characterized in that, The step of generating a voice command recommendation list based on the set of potential intents and a preset recommendation strategy includes: Based on the preset recommendation strategy, the potential intent set is recommended to obtain multiple candidate voice commands; The multiple candidate voice commands are scored and sorted to generate a recommended list of voice commands.

7. The method according to claim 1, characterized in that, Before displaying the recommended list of voice commands, the method further includes: The driving load level of the vehicle is assessed in real time based on the multimodal fusion data; The target display format is determined based on the driving load level.

8. The method according to claim 7, characterized in that, The display of the recommended list of voice commands includes: If the driving load level is high, the target display format is determined to be a partial display of the voice command recommendation list; If the driving load level is low, the target display format is determined to be a complete display of the recommended list of voice commands.

9. The method according to claim 1, characterized in that, The method further includes: The system receives user interaction feedback on the voice command recommendation list and updates the dynamic user profile and optimizes the potential intent set in real time based on the interaction feedback.

10. A vehicle, characterized in that, include: processor; Memory, used to store executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the vehicle voice control method according to any one of claims 1-9.