Conversation processing method and device, storage medium and electronic equipment
By monitoring user behavior and driving context in the in-vehicle voice system, building an interest graph and using large models for topic reasoning, personalized conversation content is generated, which solves the problem of lack of active guidance in the in-vehicle voice system and improves the interactivity and sense of companionship during the driving process.
Patent Information
- Application Number
- CN202510906593.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-23
AI Technical Summary
Existing in-vehicle voice systems lack autonomous perception and active guidance capabilities, making it difficult to meet users' needs for active companionship and continuity of human-computer interaction in scenarios such as driving fatigue or long commutes.
By monitoring user behavior information and driving context information in vehicle driving scenarios, a user interest map is constructed, and a large driving dialogue model is used to perform topic context reasoning, generate scenario dialogue plans and opening dialogue content, and realize active dialogue processing.
It realizes the intelligent proactive dialogue function in the in-vehicle driving scenario, improves the relevance, interactivity and acceptability of the dialogue content, enhances the emotional companionship ability and human-computer interaction intelligence level of the in-vehicle voice assistant, and improves the experience quality of the driving process.
Smart Images

Figure CN120690198A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a conversation processing method, device, storage medium, and electronic device. Background Art
[0002] With the development of intelligent connected technology, in-vehicle voice interaction systems have gradually become an important component of modern vehicles. By interacting with the in-vehicle voice assistant, drivers can perform operations such as navigation control, multimedia playback, and communication scheduling, thereby reducing manual operation while driving and improving driving safety.
[0003] Most existing in-vehicle voice systems rely on a passive wake-up mechanism, requiring users to actively activate the system and issue specific requests using a wake-up word or trigger command. In this model, voice assistants primarily respond to predefined commands, lacking autonomous perception and proactive guidance. This passive interaction model limits the frequency of voice system use and the intelligent experience, particularly in scenarios like driver fatigue or long commutes, making it difficult to meet users' demands for active companionship and continuous human-machine interaction. Summary of the Invention
[0004] The embodiments of this specification provide a conversation processing method, device, storage medium, and electronic device. The technical solutions are as follows:
[0005] In a first aspect, an embodiment of this specification provides a method for processing a conversation, the method comprising:
[0006] In vehicle driving scenarios, monitor the target user's in-vehicle behavior information and driving context information;
[0007] Maintaining a user interest graph based on the in-vehicle behavior information;
[0008] Based on the user interest graph and the driving context information, a driving conversation model is used to perform topic context reasoning to determine a target topic context for the target user, and a scenario conversation plan and opening conversation content are generated based on the target topic context;
[0009] Actively conduct a dialogue with the target user based on the scenario dialogue plan and the opening dialogue content.
[0010] In a feasible implementation, maintaining a user interest graph based on the in-vehicle behavior information includes:
[0011] Obtaining in-vehicle behavior information of the target user during driving, wherein the in-vehicle behavior information includes voice interaction behavior, media selection behavior, driving behavior, or user feedback behavior on conversation content;
[0012] Identify the target interest tag based on the vehicle-borne behavior information, and determine whether the target interest tag already has a target interest tag node in the user interest graph:
[0013] If so, update the graph weight value corresponding to the target interest tag node in the user interest graph;
[0014] If it does not exist, a target interest tag node corresponding to the target interest tag is added to the user interest graph, and a target graph weight value is set for the target interest tag node.
[0015] In a feasible implementation manner, updating the graph weight value corresponding to the target interest tag in the user interest graph includes:
[0016] Determine feedback impact information and behavior occurrence time interval for the target interest tag node based on the vehicle-borne behavior information, determine a weight change based on the feedback impact information, and determine a time attenuation factor based on the behavior occurrence time interval;
[0017] Obtaining the original interest weight value of the target interest tag node, and performing weight update using a weight update calculation formula based on the weight change and the time attenuation factor to obtain an updated interest weight value;
[0018] The weight update calculation formula satisfies the following formula:
[0019] W_new=W_old×λ+ΔW
[0020] Wherein, the W_new represents the updated interest weight value, the W_old represents the original interest weight value, the λ represents the time decay factor, and the ΔW represents the weight change amount.
[0021] In a feasible implementation, the determining a target topic context for the target user by performing topic context reasoning based on the user interest graph and the driving context information using a large driving conversation model includes:
[0022] Inputting the user interest graph and the driving context information into a driving dialogue macromodel, encoding the user interest graph using the driving dialogue macromodel to generate a user interest embedding representation, encoding the driving context information to generate a driving context representation, and fusing the user interest embedding representation with the driving context representation to obtain a candidate context reference representation;
[0023] Topic reasoning is performed based on the candidate context reference representation to obtain a target topic context for the target user.
[0024] In a feasible implementation, performing topic reasoning based on the candidate context reference representation to obtain a target topic context for the target user includes:
[0025] Determine multiple candidate topic contexts for target users;
[0026] Performing adaptive reasoning based on the candidate context reference representation and the candidate topic context to obtain user interest similarity, context adaptation and external event relevance, and determining a candidate topic score for each candidate topic context based on the user interest similarity, context adaptation and external event relevance;
[0027] A target topic context for the target user is determined based on the candidate topic scores.
[0028] In a feasible implementation, generating a scenario dialogue plan and opening dialogue content based on the target topic context includes:
[0029] Determining field information based on the target topic context, the field information including a target topic category field, a content anchor keyword field, a suggested tone information field, an expected user interaction target field, and a dialogue branch planning strategy field;
[0030] Field matching and filling are performed based on a preset dialogue strategy template and the field information to generate a structured scenario dialogue plan, and the opening dialogue content is determined based on the scenario dialogue plan.
[0031] In a feasible implementation, the active dialogue processing with the target user based on the scenario dialogue plan and the opening dialogue content includes:
[0032] Outputting the opening dialogue content to the target user;
[0033] Monitor the target user's feedback behavior and determine the user response type based on the user feedback behavior;
[0034] The scenario dialogue plan is used to perform dialogue processing based on the user response type.
[0035] In a feasible implementation, the performing dialogue processing based on the user response type and using the scenario dialogue plan includes:
[0036] If the user response type is a positive response type, obtaining the user's answer content to the user feedback behavior, calling the multi-round dialogue planning strategy and preset interaction goals in the scenario dialogue plan based on the user's answer content, selecting the next round of dialogue path and generating the next round of dialogue content through the driving dialogue macro model based on the user's answer content and the preset interaction goals, and outputting the next round of dialogue content;
[0037] If the user response type is a non-response type or a negative response type, the dialogue response information is determined according to the dialogue termination strategy in the scenario dialogue plan, and the dialogue processing is performed based on the dialogue response information.
[0038] In a second aspect, an embodiment of this specification provides a conversation processing device, the device comprising:
[0039] The information monitoring module is used to monitor the target user's in-vehicle behavior information and driving context information in the vehicle driving scenario;
[0040] An information processing module, configured to maintain a user interest graph based on the in-vehicle behavior information;
[0041] The information processing module is configured to perform topic context reasoning using a large driving conversation model based on the user interest graph and the driving context information to determine a target topic context for the target user, and generate a scenario conversation plan and opening conversation content based on the target topic context;
[0042] An active dialogue module is used to conduct active dialogue processing with the target user based on the scenario dialogue plan and the opening dialogue content.
[0043] In a feasible implementation, it is characterized in that maintaining the user interest graph based on the in-vehicle behavior information includes:
[0044] Obtaining in-vehicle behavior information of the target user during driving, wherein the in-vehicle behavior information includes voice interaction behavior, media selection behavior, driving behavior, or user feedback behavior on conversation content;
[0045] Identify the target interest tag based on the vehicle-borne behavior information, and determine whether the target interest tag already has a target interest tag node in the user interest graph:
[0046] If so, update the graph weight value corresponding to the target interest tag node in the user interest graph;
[0047] If it does not exist, a target interest tag node corresponding to the target interest tag is added to the user interest graph, and a target graph weight value is set for the target interest tag node.
[0048] In a feasible implementation, it is characterized in that updating the graph weight value corresponding to the target interest tag in the user interest graph includes:
[0049] Determine feedback impact information and behavior occurrence time interval for the target interest tag node based on the vehicle-borne behavior information, determine a weight change based on the feedback impact information, and determine a time attenuation factor based on the behavior occurrence time interval;
[0050] Obtaining the original interest weight value of the target interest tag node, and performing weight update using a weight update calculation formula based on the weight change and the time attenuation factor to obtain an updated interest weight value;
[0051] The weight update calculation formula satisfies the following formula:
[0052] W_new=W_old×λ+ΔW
[0053] Wherein, the W_new represents the updated interest weight value, the W_old represents the original interest weight value, the λ represents the time decay factor, and the ΔW represents the weight change amount.
[0054] In a feasible implementation, it is characterized in that the determining the target topic context for the target user by performing topic context reasoning based on the user interest graph and the driving context information using a large driving conversation model includes:
[0055] Inputting the user interest graph and the driving context information into a driving dialogue macromodel, encoding the user interest graph using the driving dialogue macromodel to generate a user interest embedding representation, encoding the driving context information to generate a driving context representation, and fusing the user interest embedding representation with the driving context representation to obtain a candidate context reference representation;
[0056] Topic reasoning is performed based on the candidate context reference representation to obtain a target topic context for the target user.
[0057] In a feasible implementation, performing topic reasoning based on the candidate context reference representation to obtain a target topic context for the target user includes:
[0058] Determine multiple candidate topic contexts for target users;
[0059] Performing adaptive reasoning based on the candidate context reference representation and the candidate topic context to obtain user interest similarity, context adaptation and external event relevance, and determining a candidate topic score for each candidate topic context based on the user interest similarity, context adaptation and external event relevance;
[0060] A target topic context for the target user is determined based on the candidate topic scores.
[0061] In a feasible implementation, generating a scenario dialogue plan and opening dialogue content based on the target topic context includes:
[0062] Determining field information based on the target topic context, the field information including a target topic category field, a content anchor keyword field, a suggested tone information field, an expected user interaction target field, and a dialogue branch planning strategy field;
[0063] Field matching and filling are performed based on a preset dialogue strategy template and the field information to generate a structured scenario dialogue plan, and the opening dialogue content is determined based on the scenario dialogue plan.
[0064] In a feasible implementation, the active dialogue processing with the target user based on the scenario dialogue plan and the opening dialogue content includes:
[0065] Outputting the opening dialogue content to the target user;
[0066] Monitor the target user's feedback behavior and determine the user response type based on the user feedback behavior;
[0067] The scenario dialogue plan is used to perform dialogue processing based on the user response type.
[0068] In a feasible implementation, the performing dialogue processing based on the user response type and using the scenario dialogue plan includes:
[0069] If the user response type is a positive response type, obtaining the user's answer content to the user feedback behavior, calling the multi-round dialogue planning strategy and preset interaction goals in the scenario dialogue plan based on the user's answer content, selecting the next round of dialogue path and generating the next round of dialogue content through the driving dialogue macro model based on the user's answer content and the preset interaction goals, and outputting the next round of dialogue content;
[0070] If the user response type is a non-response type or a negative response type, the dialogue response information is determined according to the dialogue termination strategy in the scenario dialogue plan, and the dialogue processing is performed based on the dialogue response information.
[0071] In a third aspect, an embodiment of this specification provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.
[0072] In a fourth aspect, an embodiment of this specification provides an electronic device, which may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.
[0073] The beneficial effects of the technical solutions provided by some embodiments of this specification include at least:
[0074] In one or more embodiments of the present specification, in a vehicle driving scenario, the electronic device monitors the target user's in-vehicle behavior information and driving context information, maintains the user interest map based on the in-vehicle behavior information, uses a driving dialogue model to perform topic context reasoning based on the user interest map and driving context information to determine the target topic context for the target user, generates a scenario dialogue plan and opening dialogue content based on the target topic context, and conducts active dialogue processing with the target user based on the scenario dialogue plan and opening dialogue content. This enables an intelligent proactive dialogue function driven by user behavior and contextual information in in-vehicle driving scenarios. On the one hand, by dynamically monitoring in-vehicle behavior information and driving context information, an interest graph that fits the user's real preferences is constructed and maintained, enabling continuous modeling and iterative optimization of user interests. On the other hand, by integrating the user's interest graph with the current driving context, and utilizing the driving dialogue large model for topic context reasoning and natural language generation, it is able to actively output topic content that is scene-adaptive and personalized, and conduct multiple rounds of natural interactions with users based on structured dialogue plans, thereby improving the relevance, interactivity, and acceptance of the dialogue content, significantly enhancing the emotional companionship capabilities and human-computer interaction intelligence level of the in-vehicle voice assistant, and improving the experience quality and sense of companionship during the driving process. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0076] Figure 1 This is a flow chart of a conversation processing method provided in an embodiment of this specification;
[0077] Figure 2 This is a flowchart of maintaining a user interest graph provided by an embodiment of this specification;
[0078] Figure 3 This is a flowchart of a topic context reasoning method provided by an embodiment of this specification;
[0079] Figure 4 This is a flowchart of determining a topic context provided by an embodiment of this specification;
[0080] Figure 5 This is a flow chart of a dialogue generation process provided by an embodiment of this specification;
[0081] Figure 6 This is a flowchart of an active dialogue provided by an embodiment of this specification;
[0082] Figure 7 This is a structural diagram of a conversation processing device provided in an embodiment of this specification;
[0083] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification;
[0084] Figure 9 This is a schematic diagram of the structure of the operating system and user space provided in the embodiments of this specification;
[0085] Figure 10 yes Figure 9 The architecture diagram of the Android operating system;
[0086] Figure 11 yes Figure 9 Architecture diagram of the IOS operating system. DETAILED DESCRIPTION
[0087] The following will be combined with the drawings in the embodiments of this specification to clearly and completely describe the technical solutions in the embodiments of this specification. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.
[0088] In the description of this specification, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise expressly specified and limited, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood according to the specific circumstances. In addition, in the description of this specification, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0089] The present specification is described in detail below with reference to specific embodiments.
[0090] In one embodiment, Figure 1 As shown, a conversation processing method is proposed. This method can be implemented using a computer program and run on a conversation processing device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone tool application. The conversation processing device can be an electronic device, including but not limited to a service platform, a vehicle computer, a tablet computer, a handheld device, an in-vehicle device, a computing device, or other processing device connected to a wireless modem.
[0091] Specifically, the dialogue processing method includes:
[0092] S102: In a vehicle driving scenario, monitoring the target user's in-vehicle behavior information and driving context information;
[0093] In-vehicle behavior information: refers to the perceptible behavior data generated by the target user through the in-vehicle interactive system while driving the vehicle. This behavior information can generally be categorized into the following types:
[0094] Voice interaction behavior: questions and answers, interruptions, emotional tone, etc. between users and voice assistants;
[0095] Media usage behavior: play, pause, switch, and favorite certain types of audio content (such as music, podcasts, and news);
[0096] Navigation and travel behavior: Users set navigation destinations, select route types, and prefer travel time periods;
[0097] Driving behavior characteristics: such as driving time, acceleration and deceleration frequency, frequent braking behavior, driving style (steady type, aggressive type), etc.
[0098] Driving context information: refers to the comprehensive operating status and environmental context of the vehicle, which is used to assist in determining the dialogue scenario. This information includes but is not limited to:
[0099] Time information: for example, whether it is during the morning rush hour, noon, evening, late night, etc.
[0100] Weather and environmental information: including real-time weather (such as sunny / rainy / snowy), temperature, and lighting conditions;
[0101] Vehicle status information: such as whether it is moving, stopped, on a highway or in a city;
[0102] User status information: such as user emotions or fatigue levels inferred through cameras, voice waveforms, etc.
[0103] Schematically, the electronic device system synchronously collects vehicle behavior information and driving context information by integrating multimodal sensors and vehicle operation interfaces.
[0104] S104: Maintaining a user interest graph based on the in-vehicle behavior information;
[0105] User interest graph: refers to a data structure organized in the form of a graph that is used to express user interest tendencies and their correlation strength. User interest graphs usually include:
[0106] Interest tag nodes: represent different topic types, content areas, or media preferences, such as "technology," "entertainment," "interview," and "nighttime chat."
[0107] User node: corresponds to the target user identified by the system;
[0108] Connection edge and edge weight: connects the user node and the interest tag node, used to represent the user's preference for a certain interest tag. The edge weight is usually a floating point value in the range of [0,1].
[0109] Label attributes: Nodes can be accompanied by attributes such as label type, last active time, and behavior source.
[0110] Schematically, the maintenance process includes adding, updating, and adjusting the weights of graph nodes, and mainly includes the following steps:
[0111] 1) Behavioral event analysis
[0112] Based on the vehicle behavior information collected in step S102, the system extracts the behavior event type, behavior content label, and behavior occurrence time. For example:
[0113] Media Playback Behavior → Tags: “Talk show” and “Humor”;
[0114] Voice response behavior → Tags: “News”, “Technology”;
[0115] Driving preference behavior → Tag: “Night tag” and “Light music.”
[0116] 2) Interest tag matching and node positioning
[0117] For each newly generated tag, determine whether there is a corresponding node in the existing user interest graph;
[0118] If it exists, locate the node of interest tag and prepare to perform edge weight update operation;
[0119] If it does not exist, a new interest tag node is added to the graph, and a connection edge is established with the user node, and the edge weight is initialized.
[0120] 3) Weight update strategy
[0121] For existing interest tag nodes, the system updates the interest weight based on the positive or negative behavior feedback and the time interval between behaviors. The update process uses a time decay and incremental correction mechanism. The update formula is as follows:
[0122] W_new=W_old×λ+ΔW
[0123] in:
[0124] W_old: the original edge weight of the interest tag in the current graph;
[0125] λ: is the time decay factor, ranging from (0,1), which is used to indicate the decay of the influence of old behaviors;
[0126] ΔW: The weight increment (positive or negative) calculated based on the current behavior feedback strength.
[0127] In an alternative embodiment, the system can assign different impact factors to different types of behaviors. For example, ΔW for voice response behavior is set to 0.1, while ΔW for media playback behavior is set to 0.05, to reflect the differences in the ability of different behaviors to express interest tendencies.
[0128] Optionally, graph optimization and sparsity control steps are also included: To prevent the graph from growing indefinitely, the following maintenance mechanisms are set up:
[0129] Weight lower limit pruning: If the weight of a node is lower than the set threshold (such as 0.05), the node will be deleted;
[0130] Temporal inactivity pruning: If a node has no update behavior for T consecutive periods, pruning is performed;
[0131] Node normalization: Periodically renormalize all edge weights to maintain the stability of the overall preference distribution.
[0132] [Example 1]: The user made multiple voice requests to play technology news this week.
[0133] Tag extraction: tag = technology;
[0134] Behavior type = voice active command;
[0135] Graph maintenance: If the node already exists, update the weight of the interest node from 0.65 to 0.74; if it does not exist, add a "Technology" node and set the initial weight to 0.3.
[0136] [Example 2]: A user plays meditation music at night without any voice interaction.
[0137] Tag extraction: "light music" "night preference";
[0138] Behavior type = passive media selection;
[0139] Graph processing: Update the edge weight of the "night" label, generate the "light music" node and form a contextual collaborative weight with the "night" node.
[0140] Furthermore, through this maintenance mechanism, the user interest graph accurately reflects the evolving content preferences of target users across different time periods and driving conditions. Compared to static user profile modeling, this approach offers: enhanced real-time performance driven by behavioral responses; a combination of decay and increment for better representation of long-term trends; and an extensible tag structure for greater scenario generalization. (Step S106) This provides a precise foundation for individual interest expression, enabling subsequent topic contextual reasoning.
[0141] S106: Performing topic context reasoning using a driving conversation model based on the user interest graph and the driving context information to determine a target topic context for the target user, and generating a scenario conversation plan and opening conversation content based on the target topic context;
[0142] Driving Dialogue Large Model (LLM): refers to a specially trained or fine-tuned multimodal large language model that has the ability to fuse user interest semantics with current scene information and infer and generate topic content. It is used to perform topic matching and natural language generation tasks.
[0143] Target topic context: refers to candidate topics and their semantic description structures that are highly consistent with user interests and suitable for triggering interaction at the current time and driving environment.
[0144] Dialogue Plan: A structured guide plan for the target topic, defining the opening tone, content anchors, interaction goals, and possible subsequent branching strategies.
[0145] In a feasible implementation, topic context inference is performed using a driving conversation macro model based on the user interest graph and the driving context information. This may be performed by using user interests and driving context as search queries to retrieve topic fragments or event summaries, which serve as semantic references for the driving conversation macro model to generate a target topic context. The processing process is as follows:
[0146] 1. Use interest graph labels and context labels to jointly construct query vectors;
[0147] 2. Retrieve relevant topics from the preset topic library and hot news summaries (e.g., "Commuting + Entertainment" → "This morning's hottest celebrity jokes");
[0148] 3. Input the search results and embedding vectors into the driving dialogue model as a reference for generating dialogue intent.
[0149] 4. Use the driving dialogue model to determine a topic context that is closer to the real background content, and the driving dialogue model generates a scenario dialogue plan and opening dialogue content based on the target topic context.
[0150] S108: Actively conduct a dialogue with the target user based on the scenario dialogue plan and the opening dialogue content.
[0151] Active dialogue processing: refers to the process in which the on-board intelligent agent automatically detects the appropriate interaction opportunity without the user actively waking up, starts the topic and conducts multiple rounds of natural language dialogue with the user.
[0152] This step can be broken down into the following sub-processes:
[0153] 1) Opening statement output (S108-1)
[0154] The system determines whether it is an appropriate time to initiate a conversation (e.g., stable driving, quiet environment) based on contextual information such as the current time period and driving status. If so, the system calls up the opening conversation content and outputs it via in-car voice broadcast.
[0155] For example: "I just saw on the news that the most popular movie last month was 'Drive'. Have you been to the theater recently?"
[0156] 2) Monitoring and identifying user response behavior (S108-2)
[0157] The system monitors the user's voice feedback and uses the speech recognition and semantic parsing modules to extract the user's response content and further identify the response type, mainly including:
[0158] Positive response: such as actively answering, asking questions, and actively extending the topic;
[0159] No response: such as long periods of silence and no clear language input;
[0160] Negative response: such as rejection, denial or interruption of the system.
[0161] 3) Dialogue Response Strategy Processing (S108-3)
[0162] Based on the user's response type, the system matches the subsequent strategy in the scenario dialogue plan:
[0163] If it is a positive response, then: the multi-round dialogue planning logic in the plan is called, the user's answer is input into the driving dialogue model, and the next round of dialogue content is generated based on the preset interaction goals;
[0164] Output a new round of dialogue sentences and enter the next round of monitoring.
[0165] If there is no response or a negative response, then: call the termination strategy in the plan;
[0166] Output a concise response (such as: "It's okay, I'll talk to you later.") or enter a dormant state and wait for the next interaction opportunity.
[0167] 4) Interaction cycle management (S108-4)
[0168] The system can set policies such as a maximum number of interaction rounds and a maximum response wait time within a task cycle to avoid interrupting users or causing fatigue. It can also set different levels of guidance intensity for task-oriented topics (such as "Did you navigate home?") and companion-oriented topics (such as "Did you watch the game today?").
[0169] In an embodiment of the present specification, in a vehicle driving scenario, the electronic device monitors the target user's in-vehicle behavior information and driving context information, maintains a user interest map based on the in-vehicle behavior information, uses a driving dialogue model to perform topic context reasoning based on the user interest map and driving context information to determine a target topic context for the target user, generates a scenario dialogue plan and opening dialogue content based on the target topic context, and conducts active dialogue processing with the target user based on the scenario dialogue plan and opening dialogue content. This enables an intelligent proactive dialogue function driven by user behavior and contextual information in in-vehicle driving scenarios. On the one hand, by dynamically monitoring in-vehicle behavior information and driving context information, an interest graph that fits the user's real preferences is constructed and maintained, enabling continuous modeling and iterative optimization of user interests. On the other hand, by integrating the user's interest graph with the current driving context, and utilizing the driving dialogue large model for topic context reasoning and natural language generation, it is able to actively output topic content that is scene-adaptive and personalized, and conduct multiple rounds of natural interactions with users based on structured dialogue plans, thereby improving the relevance, interactivity, and acceptance of the dialogue content, significantly enhancing the emotional companionship capabilities and human-computer interaction intelligence level of the in-vehicle voice assistant, and improving the experience quality and sense of companionship during the driving process.
[0170] Optional, see Figure 2 , Figure 2 This is a flowchart of a user interest graph maintenance process proposed in this specification. The specific implementation of the maintenance of the user interest graph based on the vehicle behavior information can refer to the following methods:
[0171] S202: Acquiring in-vehicle behavior information of the target user during driving;
[0172] Illustratively, behavioral data related to user interactions is collected in real time or periodically during each driving cycle. The in-vehicle behavioral information includes but is not limited to:
[0173] User-initiated voice commands (such as "play a certain type of music" or "ask about the weather");
[0174] Multimedia content playback records (such as listening to certain podcasts, watching short videos, etc.);
[0175] Interactive operation behaviors (such as selecting a certain type of application or skipping certain content through the central control screen);
[0176] Driving preference behavior (such as maintaining a certain content preference in a certain driving situation for a long time);
[0177] This step can be performed through the behavior perception module in the vehicle system, and the output behavior event stream is used as the basis for subsequent label extraction.
[0178] S204: Identify a target interest tag based on the vehicle-mounted behavior information, and determine whether the target interest tag already has a target interest tag node in the user interest graph:
[0179] In this step, semantic parsing is performed on the vehicle behavior information to extract the user content preferences reflected and generate several interest tags. The interest tags may include:
[0180] Content-related tags (such as "technology information", "crosstalk", "car review");
[0181] Tone and style labels (e.g., “relaxed,” “information-dense,” “emotional”);
[0182] Time-related tags (such as "evening entertainment" and "morning news").
[0183] Next, a node search is performed in the current user interest graph based on the interest tag to determine whether the tag already exists in the graph structure. The interest graph can be in the form of a graph structure, where user nodes and interest tag nodes are associated through weighted edges.
[0184] S206: If so, updating the graph weight value corresponding to the target interest tag node in the user interest graph;
[0185] Indicatively, if the interest tag node already exists in the graph, the system will update the weight value of the corresponding edge based on the characteristics of this round of behavior.
[0186] In one feasible implementation, for target interest tag nodes already in the user interest graph, the system performs a graph weight update process to dynamically reflect the activity of the interest tag in the current driving cycle and the user's response preferences. Specifically, the following method can be used to update the graph weight value corresponding to the target interest tag in the user interest graph:
[0187] 1) Determining feedback impact information and behavior occurrence time interval for the target interest tag node based on the vehicle-borne behavior information, determining a weight change based on the feedback impact information, and determining a time attenuation factor based on the behavior occurrence time interval;
[0188] Schematically, after receiving the newly added vehicle behavior information, the feedback elements and time features related to the target interest tag are first extracted from the behavior data, including:
[0189] Feedback Impact Info: This indicates the degree to which the behavior event strengthens the target interest tag. This information can be quantified based on the behavior type, for example: active voice question: +1.0; active click-to-play: +0.7; passive play with high completion rate: +0.5; quick skip: –0.3.
[0190] Behavior occurrence time interval (ΔT): represents the time difference between the current behavior and the last active time of the interest tag, which is used to model the "heat decay" of interest.
[0191] Time decay factor (λ): indicates the degree to which the effect of the interest tag being strengthened decreases over time, λ∈(0,1), which can be calculated by exponential function, linear function, etc., such as:
[0192] λ=e -kΔT
[0193] Where k is an adjustable time decay parameter.
[0194] Weight change (ΔW): This can be determined based on feedback impact information and other contextual factors, such as:
[0195] ΔW=ImpactScore×α
[0196] Among them, ImpactScore is the above feedback impact information, and α is the adjustment factor.
[0197] 2) obtaining the original interest weight value of the target interest tag node, and performing weight update based on the weight change and the time attenuation factor using a weight update calculation formula to obtain an updated interest weight value;
[0198] The weight update calculation formula satisfies the following formula:
[0199] W_new=W_old×λ+ΔW,
[0200] Wherein, the W_new represents the updated interest weight value, the W_old represents the original interest weight value, the λ represents the time decay factor, and the ΔW represents the weight change amount.
[0201] This weight update calculation formula ensures that: recent active behaviors have a greater positive impact; interest nodes that have been inactive for a long time gradually fade out; and the impact of different types of behaviors on the weight of interest tags is dynamically adjusted in a differentiated manner.
[0202] In principle, through the above-mentioned weight update mechanism, the system can realize dynamic adjustment of the weights of interest nodes, taking into account both time sensitivity and behavioral response intensity, thereby improving the expression accuracy of the user interest map in the time dimension, enhancing the real-time relevance and personalized matching capabilities of topic selection in subsequent dialogue generation, and realizing the modeling of behavioral change trends such as interest migration and topic decay.
[0203] S208: If it does not exist, then add a target interest tag node corresponding to the target interest tag in the user interest graph, and set a target graph weight value for the target interest tag node.
[0204] If it is determined that the interest tag node does not exist in the graph, then: add a target interest tag node to the interest graph, establish a connecting edge between the node and the user node, initialize the edge weight to a preset value (such as 0.3) or directly set it according to the first behavior weight; the node can be given meta-attributes such as label category and creation time for subsequent judgment on whether to remove or merge.
[0205] The above operations enable the system to continuously expand the scope of user interests and dynamically adapt to changes in user content preferences.
[0206] In this embodiment of the specification, the processing flow from steps S202 to S208 continuously collects and analyzes the target user's in-vehicle behavior information while driving, identifies and dynamically maintains interest tag nodes and their weights in the user's interest graph, and implements structured and sustainable modeling of user interests. This mechanism not only supports the adaptive evolution and long-term memory of user interest preferences, but also possesses the ability to sensitively capture emerging interests, significantly improving the personalized recommendation capabilities and topic matching accuracy of subsequent conversation content, providing a high-quality semantic foundation and user profiling support for intelligent and proactive conversations.
[0207] Optional, see Figure 3 , Figure 3 This is a flowchart of topic context reasoning. Specifically, the topic context reasoning based on the user interest graph and the driving context information is performed using the driving conversation model to determine the target topic context for the target user. The following method can be used:
[0208] S302: Inputting the user interest graph and the driving context information into a driving dialogue macromodel, encoding the user interest graph using the driving dialogue macromodel to generate a user interest embedding representation, encoding the driving context information to generate a driving context representation, and fusing the user interest embedding representation with the driving context representation to obtain a candidate context reference representation;
[0209] (User Interest) Embedding Representation: refers to converting structured or unstructured information into a fixed-length, high-dimensional, differentiable vector form to facilitate subsequent unified processing in the model.
[0210] Candidate context reference (embedding) representation: refers to the joint vector obtained by fusing the user interest embedding representation and the driving context representation, which serves as the semantic decision basis for topic selection and content generation.
[0211] Schematically, the driving conversation model encodes the user interest graph to generate a user interest embedding representation. The driving conversation model first models the structure of the user interest graph, and then inputs the node type (such as "music", "news", "technology"), node weight, label category and its connection edge structure into the graph modeling module to encode the user interest embedding vector V user .
[0212] The driving context information is embedded and encoded to generate the driving context representation V context ;
[0213] Re-integrate user interest embedding vector V user and the driving context representation V context to generate candidate context reference representations, and generate the fused candidate context reference representation V fused , used for downstream topic context reasoning. The fusion mechanism can be adopted: first embed the user interest vector V user The vector is concatenated with the driving context representation V context, and then processed using a fully connected network (MLP):
[0214] Alternatively, the attention fusion mechanism can be used: using the context information in the driving context representation Vcontext as the query vector, and embedding the user interest vector V user The user interest features in the dataset are extracted with attention weighting to highlight the interest dimensions with higher relevance to the current scene.
[0215] S304: Perform topic reasoning based on the candidate context reference representation to obtain a target topic context for the target user.
[0216] Candidate context reference representation: A unified semantic representation vector obtained by fusing the user interest embedding representation and the driving context representation, denoted as V fused , which is used to reflect the user's comprehensive interest status and external context perception in the current driving scenario.
[0217] Candidate topic context: refers to a set of structured topic units that are pre-defined or dynamically generated. Each topic context contains elements such as topic type, content anchor, tone label, interaction goal, adaptation scenario, etc., and can serve as a candidate entry point for active dialogue.
[0218] Target topic context: refers to the structured topic that best matches the current user preferences and context characteristics, selected from the candidate topic context set based on the current context reference representation.
[0219] In principle, we first construct or obtain a candidate topic context set through a large driving dialogue model. Specifically, we can extract recent hot topics from the knowledge base, generate multi-category topic templates based on user group preference data, and filter the candidate topic context set. All candidate topic contexts in the candidate topic context set are converted into embedded vectors. We perform similarity calculations on the candidate context reference representation and each topic context embedding vector to obtain the semantic similarity Sim, and then calculate the final topic score Score by combining external relevance and interactive adaptability. i :
[0220] Score i =α□Sim(V fused ,V topic i)+β□EventRel i +γ□ContextFit i
[0221] Among them, Sim() represents semantic similarity, EventRel i Indicates the relevance of the current topic to external hot events or current affairs (such as based on news API or trending searches); ContextFit i : Indicates the degree of adaptation between the current driving situation and the topic (such as the technical interpretation of the adaptation of static driving scenes); α, β, and γ are weighting coefficients used to control the contribution ratio of different factors.
[0222] Furthermore, all candidate topics are ranked by the driving dialogue model, and one or more topic contexts with the highest scores or meeting the preset threshold conditions are selected as the final output, which is the target topic context.
[0223] In this specification, steps S302 through S304 construct a semantically fused representation based on the user's interest graph and driving context. Multi-factor topic reasoning is then performed on this basis to generate a target topic context that closely matches the current user state. This mechanism effectively improves the relevance and personalization of topic content, enabling a transition from static recommendations to dynamic, proactive interaction. This enhances the contextual awareness and dialogue adaptability of the in-vehicle intelligent system, providing key support for natural and fluid human-machine proactive dialogue.
[0224] In one possible implementation, see Figure 4 , Figure 4 This is a flowchart of determining a topic context. The topic reasoning based on the candidate context reference representation is performed to obtain the target topic context for the target user. The following method can be used:
[0225] S402: Determine multiple candidate topic contexts for the target user;
[0226] Multiple candidate topic contexts are obtained from a predefined knowledge base or an online topic generation service. The candidate topic contexts are structured semantic units. Each topic context includes but is not limited to the following fields:
[0227] Topic type (e.g., technology, entertainment, weather, commuting tips, etc.);
[0228] Content anchors (keywords or summaries used to locate topics);
[0229] Tone style (e.g., informative, teasing, companionable);
[0230] Suggest interaction goals (e.g., relieving fatigue, livening up the atmosphere, delivering information);
[0231] Adaptation tags (such as recommendation scenarios, user preference matching conditions, etc.).
[0232] The candidate topic context can be filled in through a preset template, generated by a large driving dialogue model trained based on user group data, and can also be updated in real time in combination with current external events.
[0233] S404: performing adaptive reasoning based on the candidate context reference representation and the candidate topic context to obtain user interest similarity, context adaptation and external event relevance, and determining a candidate topic score for each candidate topic context based on the user interest similarity, context adaptation and external event relevance;
[0234] In one optional implementation, the driving conversation model first obtains multiple candidate topic contexts based on a pre-set topic resource library or an online generation mechanism. Each candidate topic context corresponds to a structured topic unit, which includes at least fields such as topic type, content anchor, recommendation tone, applicable scenario label, and interaction target information.
[0235] Subsequently, the driving dialogue model performs semantic matching between the current candidate context reference representation and each candidate topic context. Specifically, the semantic matching process includes reasoning and judgment in the following three dimensions:
[0236] First, the driving conversation model determines user interest similarity. This similarity measures the degree of match between the user interest characteristics reflected in the candidate context reference representation and the candidate topic context. This similarity is typically measured by analyzing interest tags, node associations, and semantic distance to reflect the user's preference for the topic content.
[0237] Second, contextual compatibility is determined through the driving dialogue model. This compatibility measures the degree of adaptability between the candidate topic context and the current vehicle driving situation. This adaptability can be calculated based on pre-set scenario adaptation rules or by comparing similarities between context labels to determine whether the topic is appropriate for output at the current time, location, and driving state.
[0238] Third, the driving conversation model determines the relevance of external events. This relevance measures whether the content associated with the candidate topic context is significantly related to current social hot topics, news events, or social trends. This determination can be made by comparing news summaries, popular tags on social platforms, or predefined real-time event feature vectors.
[0239] Based on the judgment results of the above three dimensions, a corresponding candidate topic score is generated for each candidate topic context. This candidate topic score is used to indicate the recommendation priority of the candidate topic based on the current user status and environment. The higher the score, the more appropriate it is for the current scenario and the more suitable it is for initiating a conversation.
[0240] Ultimately, the score will be used to screen out the optimal target topic context in subsequent steps.
[0241] S406: Determine a target topic context for the target user based on the candidate topic scores.
[0242] After obtaining candidate topic scores for multiple candidate topic contexts, the driving conversation model first sorts all candidate topic scores according to a preset priority ranking rule. This priority ranking rule can be dynamically adjusted based on the numerical value of the candidate topic score, the policy parameters defined in the user interaction preference model, or the current in-vehicle system's conversation frequency control policy.
[0243] After the ranking is completed, the driving dialogue model selects one or more candidate topic contexts with the highest scores from the ranked candidate topic context set as the current target topic context. The selection operation may include one of the following two methods:
[0244] First, when the candidate topic scoring results meet the set uniqueness judgment conditions, the candidate topic context with the highest score can be directly selected as the only target topic context for subsequent dialogue content generation and decision-making guidance.
[0245] Secondly, when there are multiple candidate topic scoring results that are all higher than the preset adaptability threshold, polynomial sorting and strategy selection can be further performed based on factors such as the long-term activity of label nodes in the user interest graph, the interaction sensitivity of recent driving status, and the priority of historical dialogue feedback, so as to screen out a better target topic context from multiple candidate topic contexts.
[0246] Once the target topic context is determined, it will be used as the input basis for the subsequent scenario dialogue plan construction and opening dialogue generation, ensuring that the topic content involved in the active dialogue can be highly matched with the user's interest status in the current driving context, thereby improving the interactive experience quality and user acceptance of the in-vehicle dialogue system.
[0247] In this specification, steps S402 through S406, based on the integration of user interests and driving context, accurately infer and select the target topic context that best matches the current user state from multiple candidate topic contexts. This process comprehensively considers multiple factors such as user interest similarity, contextual adaptability, and relevance to external events, dynamically optimizing conversation content in terms of personalization, contextual relevance, and real-time performance, effectively improving the quality of proactive interaction within the in-vehicle intelligent system and user experience satisfaction.
[0248] Optional, see Figure 5 , Figure 5 This is a flow chart of dialogue generation. The following methods can be used to generate scenario dialogue plans and opening dialogue content based on the target topic context:
[0249] S502: Determine field information based on the target topic context, the field information including a target topic category field, a content anchor keyword field, a suggested tone information field, an expected user interaction target field, and a dialogue branch planning strategy field;
[0250] First, based on the structured content of the target topic context, extract the field information required to generate the scenario dialogue plan. The field information may include but is not limited to the following types:
[0251] Target topic category field: used to identify the subject type of the current topic, such as "Technology News", "Life Interesting Stories", "Weather Reminders", etc.
[0252] Content anchor keyword field: used to refer to the core information points or topic entry points of this round of conversation, such as "Commute congestion this morning" and "Hot talk show jokes".
[0253] Suggested Tone Information field: This field is used to guide the tone of voice used in conversational sentences, such as "lighthearted banter," "neutral reporting," and "gentle caring," ensuring that the human-computer conversation style is appropriate for the user's current emotional state.
[0254] Expected user interaction goal field: used to clarify the interaction intention goal between the system and the user, such as "obtaining feedback emotions", "providing information broadcast", "relieving fatigue emotions", etc.
[0255] Dialogue branch planning strategy field: used to indicate the interaction path planning method that the system can adopt in this round of dialogue, such as "open question guidance", "continuous content exploration", "rapid information feedback", etc.
[0256] The above field information can be automatically extracted and classified by the system by parsing the tag information in the context of the target topic and combining it with the context knowledge base rules.
[0257] S504: Perform field matching and filling based on the preset dialogue strategy template and the field information to generate a structured scenario dialogue plan, and determine the opening dialogue content based on the scenario dialogue plan.
[0258] After the field information is determined, the driving dialogue model selects a template that matches the field information from the preset dialogue strategy template library. Each template includes sentence structure, tone pattern, dialogue guidance logic, user response mapping path, and other content.
[0259] The driving conversation model extracts field information and populates it into the corresponding placeholders in the template, generating a structured scenario-based conversation plan with semantic integrity and appropriate tone. This conversation plan can define multiple interaction paths and trigger conditions, supporting branching planning based on different user response types.
[0260] Based on the scenario-based dialogue plan, the driving dialogue model automatically generates the opening dialogue content for initiating the first round of dialogue with the target user. This opening statement is topic-oriented, natural in tone, and context-appropriate, ensuring that the user can enter the conversation without being interruptive or abrupt.
[0261] In this specification, through the processing of steps S502 to S504, the system automatically extracts key field information based on the target topic context and, in combination with pre-set dialogue strategy templates, generates structured scenario dialogue plans and natural and smooth opening dialogue content, thereby achieving personalized dialogue content and precise control of tone and style. This mechanism not only improves the contextual fit and user acceptance of proactive dialogue initiated by the in-vehicle intelligent system, but also provides a clear and controllable planning foundation for subsequent multiple rounds of interaction, effectively enhancing the naturalness, flexibility, and user satisfaction of vehicle-wide human-machine dialogue.
[0262] Optional, see Figure 6 , Figure 6 This is a flowchart of an active dialogue process, which specifically implements the active dialogue process with the target user based on the scenario dialogue plan and the opening dialogue content. The following methods can be used as a reference:
[0263] S602: Outputting the opening dialogue content to the target user;
[0264] The system generates an opening dialogue based on the target topic context and outputs it to the target user via a voice broadcast module, an in-vehicle display terminal, or other human-computer interaction interface. This opening dialogue typically uses an anthropomorphic, low-intrusive, and gentle language structure to attract the user's attention and naturally enter the interaction state.
[0265] In actual implementation, the output method can be adapted according to the driving scenario, for example:
[0266] Use only voice announcements when driving on highways to avoid distractions;
[0267] When the vehicle is stationary or parked, the screen can be used for graphic display.
[0268] S604: Monitoring the user feedback behavior of the target user, and determining the user response type according to the user feedback behavior;
[0269] After outputting the opening content, the system starts the monitoring mechanism to perceive and collect the target user's feedback behavior. The user feedback behavior includes but is not limited to:
[0270] Voice response content;
[0271] changes in facial expressions (e.g., smiling, frowning);
[0272] Gestures (e.g., nodding, waving);
[0273] The direction of gaze;
[0274] Prolonged unresponsive behavior.
[0275] Based on the above perception signals, the system determines the user's response type through a multimodal fusion recognition model. The response type may include:
[0276] Positive response type: such as actively replying to voice messages, showing interest, and actively asking questions;
[0277] Non-response type: such as silence, no clear feedback;
[0278] Negative response types: such as rejection, denial, showing disinterest, etc.
[0279] The judgment result will be used to guide subsequent interaction path selection and dialogue rhythm control.
[0280] S606: Based on the user response type, the scenario dialogue plan is used to perform dialogue processing.
[0281] The system implements personalized dialogue strategies based on the user's response type and the preset scenario dialogue plan, including:
[0282] In the case of a positive response, the system can call multiple dialogue branches related to user feedback in the scenario dialogue plan, extract the keyword content in the current user's answer, combine the preset interaction goals and context status, and generate and output subsequent dialogue content through a large language model or a dedicated dialogue engine to form a continuous and natural dialogue chain.
[0283] In the event of no response or a negative response, the system can determine the situation based on the termination strategy or alternative topic guidance strategy defined in the scenario dialogue plan. If it determines that the user is unwilling to interact in the short term, the system can choose to play light music, switch to a lighter topic, or end the current conversation to minimize disruption to driving status or user emotions.
[0284] In this specification, through the active dialogue execution process from S602 to S606 mentioned above, a closed-loop interaction mechanism from topic generation to user response is completed, so that the in-vehicle dialogue system has the comprehensive capabilities of active initiation, understanding feedback, and dynamic response; effectively improves the adaptability and response flexibility of dialogue, ensures driving safety while enhancing the user experience; through multimodal perception fusion and branch planning strategies, personalized guidance and fault-tolerant processing of dialogue content are achieved, and the stability and naturalness of in-vehicle companionship interaction are enhanced.
[0285] In a feasible implementation, the dialogue processing based on the user response type and the scenario dialogue plan may be performed in the following manner:
[0286] B2: If the user response type is a positive response type, obtain the user's answer content to the user feedback behavior, invoke the multi-round dialogue planning strategy and preset interaction goals in the scenario dialogue plan based on the user's answer content, select the next round of dialogue path and generate the next round of dialogue content based on the user's answer content and the preset interaction goals through the driving dialogue macro model, and output the next round of dialogue content;
[0287] In this case, when the target user makes a clear and positive response, the system determines the behavior as a positive response type. In this case, the system performs the following processing steps:
[0288] First, parse the user's answer content, parse the voice text in the user feedback, extract keywords, emotional features or implicit intentions, and form a semantic representation of user feedback.
[0289] Then, the multi-round dialogue planning strategy in the scenario dialogue plan is called: according to the current dialogue stage, the system searches for the corresponding dialogue branch planning path, and determines the subsequent interaction direction based on the preset interaction goals (for example: expanding the topic, obtaining attitudes, completing recommendations).
[0290] Then, the driving dialogue model generates subsequent content, takes the semantic representation of user feedback and the current interaction goal as joint input, and inputs them into the large language model trained exclusively on board the vehicle to generate candidate dialogue paths and corresponding natural language response content for the next round of interaction.
[0291] Finally, the content of the next round of dialogue is output: the system selects the content corresponding to the optimal dialogue path, and continuously interacts with the user through voice broadcast or other human-computer interaction methods to achieve natural and smooth multi-round dialogue.
[0292] This processing mechanism ensures that the system can provide high-quality responses that are scalable, coherent, and interactive when users express interest, thereby enhancing user engagement.
[0293] B4: If the user response type is a non-response type or a negative response type, determine the dialogue response information according to the dialogue termination strategy in the scenario dialogue plan, and perform dialogue processing based on the dialogue response information.
[0294] In an illustrative example, when the system determines that the user does not provide obvious feedback or provides negative rejection feedback, the system classifies the behavior as a non-response type or a negative response type and executes the following processing flow:
[0295] First, the system matches the termination strategy in the scenario dialogue plan. Based on the current topic category, driving situation and the user's past interaction behavior, it searches for the corresponding topic exit strategy or switching strategy. The strategy can be pre-defined into multiple modes such as "flexible termination", "attempt to turn", and "end the conversation".
[0296] Then, determine the dialogue response information and generate appropriate response content based on the selected termination strategy, such as "I won't bother you anymore" or "Let's talk about this topic again when we have time", to ensure the naturalness of exiting the dialogue and user comfort.
[0297] Finally, the system outputs the content of terminating the conversation or switching topics. The system outputs the response information to the user through voice or interface prompts, and chooses whether to end the interaction, reduce the active frequency, or start the next candidate topic awakening process based on the termination strategy.
[0298] This approach can effectively avoid user fatigue or annoyance caused by frequent or inappropriate proactive conversations, ensuring the consistency of the interactive experience and the establishment of a trusting relationship between humans and machines.
[0299] In this specification, the aforementioned approach allows for flexible judgment of interaction intent based on user feedback and dynamic adjustment of conversation strategies. When the user responds positively, the system leverages feedback and pre-set interaction goals to drive deeper conversations through multi-round planning and language model generation. When the user doesn't respond or responds negatively, the system uses a conversation termination strategy to naturally end the conversation or reduce interaction frequency, avoiding interruptions or emotional burdens. This mechanism significantly enhances the proactive conversation system's adaptability and interactive affinity, achieving a closed-loop process for human-machine conversations from "topic initiation" to "response adaptation" to "strategy convergence," enhancing the coherence, contextual adaptability, and user satisfaction of the in-car companionship experience.
[0300] The following will be combined Figure 7 , the dialogue processing device provided by the embodiment of this specification is introduced in detail. It should be noted that, Figure 7 The dialogue processing device shown is used to execute the Figures 1 to 6 For the convenience of explanation, only the part related to the embodiment of this specification is shown. For the specific technical details not disclosed, please refer to this specification. Figures 1 to 6 The embodiment shown.
[0301] See Figure 7 , which shows a schematic diagram of the structure of a conversation processing device according to an embodiment of this specification. The conversation processing device 1 can be implemented as all or part of a user terminal through software, hardware, or a combination of both. According to some embodiments, the conversation processing device 1 includes an information monitoring module 11, an information processing module 12, and an active conversation module 13, specifically configured to:
[0302] The information monitoring module 11 is used to monitor the target user's in-vehicle behavior information and driving context information in a vehicle driving scenario;
[0303] An information processing module 12, configured to maintain a user interest graph based on the vehicle-borne behavior information;
[0304] The information processing module 12 is configured to perform topic context reasoning using a large driving conversation model based on the user interest graph and the driving context information to determine a target topic context for the target user, and generate a scenario conversation plan and opening conversation content based on the target topic context;
[0305] The active dialogue module 13 is used to conduct active dialogue processing with the target user based on the scenario dialogue plan and the opening dialogue content.
[0306] In a feasible implementation, it is characterized in that maintaining the user interest graph based on the in-vehicle behavior information includes:
[0307] Obtaining in-vehicle behavior information of the target user during driving, wherein the in-vehicle behavior information includes voice interaction behavior, media selection behavior, driving behavior, or user feedback behavior on conversation content;
[0308] Identify the target interest tag based on the vehicle-borne behavior information, and determine whether the target interest tag already has a target interest tag node in the user interest graph:
[0309] If so, update the graph weight value corresponding to the target interest tag node in the user interest graph;
[0310] If it does not exist, a target interest tag node corresponding to the target interest tag is added to the user interest graph, and a target graph weight value is set for the target interest tag node.
[0311] In a feasible implementation, it is characterized in that updating the graph weight value corresponding to the target interest tag in the user interest graph includes:
[0312] Determine feedback impact information and behavior occurrence time interval for the target interest tag node based on the vehicle-borne behavior information, determine a weight change based on the feedback impact information, and determine a time attenuation factor based on the behavior occurrence time interval;
[0313] Obtaining the original interest weight value of the target interest tag node, and performing weight update using a weight update calculation formula based on the weight change and the time attenuation factor to obtain an updated interest weight value;
[0314] The weight update calculation formula satisfies the following formula:
[0315] W_new=W_old×λ+ΔW,
[0316] Wherein, the W_new represents the updated interest weight value, the W_old represents the original interest weight value, the λ represents the time decay factor, and the ΔW represents the weight change amount.
[0317] In a feasible implementation, it is characterized in that the determining the target topic context for the target user by performing topic context reasoning based on the user interest graph and the driving context information using a large driving conversation model includes:
[0318] Inputting the user interest graph and the driving context information into a driving dialogue macromodel, encoding the user interest graph using the driving dialogue macromodel to generate a user interest embedding representation, encoding the driving context information to generate a driving context representation, and fusing the user interest embedding representation with the driving context representation to obtain a candidate context reference representation;
[0319] Topic reasoning is performed based on the candidate context reference representation to obtain a target topic context for the target user.
[0320] In a feasible implementation, performing topic reasoning based on the candidate context reference representation to obtain a target topic context for the target user includes:
[0321] Determine multiple candidate topic contexts for target users;
[0322] Performing adaptive reasoning based on the candidate context reference representation and the candidate topic context to obtain user interest similarity, context adaptation and external event relevance, and determining a candidate topic score for each candidate topic context based on the user interest similarity, context adaptation and external event relevance;
[0323] A target topic context for the target user is determined based on the candidate topic scores.
[0324] In a feasible implementation, generating a scenario dialogue plan and opening dialogue content based on the target topic context includes:
[0325] Determining field information based on the target topic context, the field information including a target topic category field, a content anchor keyword field, a suggested tone information field, an expected user interaction target field, and a dialogue branch planning strategy field;
[0326] Field matching and filling are performed based on a preset dialogue strategy template and the field information to generate a structured scenario dialogue plan, and the opening dialogue content is determined based on the scenario dialogue plan.
[0327] In a feasible implementation, the active dialogue processing with the target user based on the scenario dialogue plan and the opening dialogue content includes:
[0328] Outputting the opening dialogue content to the target user;
[0329] Monitor the target user's feedback behavior and determine the user response type based on the user feedback behavior;
[0330] The scenario dialogue plan is used to perform dialogue processing based on the user response type.
[0331] In a feasible implementation, the performing dialogue processing based on the user response type and using the scenario dialogue plan includes:
[0332] If the user response type is a positive response type, obtaining the user's answer content to the user feedback behavior, calling the multi-round dialogue planning strategy and preset interaction goals in the scenario dialogue plan based on the user's answer content, selecting the next round of dialogue path and generating the next round of dialogue content through the driving dialogue macro model based on the user's answer content and the preset interaction goals, and outputting the next round of dialogue content;
[0333] If the user response type is a non-response type or a negative response type, the dialogue response information is determined according to the dialogue termination strategy in the scenario dialogue plan, and the dialogue processing is performed based on the dialogue response information.
[0334] It should be noted that the above-described embodiments of the conversation processing device, when executing the conversation processing method, illustrate the division of the aforementioned functional modules only as an example. In actual applications, the aforementioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the conversation processing device and the conversation processing method embodiments provided in the above-described embodiments share the same concept. Their implementation process is detailed in the method embodiments and will not be further elaborated here.
[0335] The serial numbers of the embodiments in the present specification are for description only and do not represent the advantages or disadvantages of the embodiments.
[0336] The embodiment of this specification also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor as described above. Figures 1 to 6 The specific execution process of the dialogue processing method in the embodiment shown can be found in Figures 1 to 6 The detailed description of the illustrated embodiment will not be repeated here.
[0337] This specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figures 1 to 6 The specific execution process of the dialogue processing method in the embodiment shown can be found in Figures 1 to 6 The detailed description of the illustrated embodiment will not be repeated here.
[0338] Please refer to Figure 8 , which shows a block diagram of the structure of an electronic device provided by an exemplary embodiment of this specification. The electronic device described in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.
[0339] The processor 110 may include one or more processing cores. The processor 110 utilizes various interfaces and circuits to connect various components within the electronic device. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, as well as accesses data stored in the memory 120, to perform various functions of the electronic device and process data. Optionally, the processor 110 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 110 and may be implemented separately via a communications chip.
[0340] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The operating system may be an Android system, including a system deeply developed based on the Android system, an IOS system developed by Apple, including a system deeply developed based on the IOS system or other systems. The data storage area may also store data created by the electronic device during use, such as a phone book, audio and video data, chat record data, etc.
[0341] See also Figure 9 As shown, the memory 120 can be divided into operating system space and user space. The operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve better operating results, the operating system allocates corresponding system resources to different third-party applications. However, the requirements for system resources in different application scenarios in the same third-party application are also different. For example, in the local resource loading scenario, the third-party application has higher requirements for disk reading speed; in the animation rendering scenario, the third-party application has higher requirements for GPU performance. The operating system and the third-party application are independent of each other, and the operating system often cannot perceive the current application scenario of the third-party application in a timely manner, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application.
[0342] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to open up data communication between third-party applications and the operating system so that the operating system can obtain the current scenario information of third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.
[0343] Taking the Android operating system as an example, the programs and data stored in the memory 120 are as follows: Figure 10As shown, the memory 120 may store a Linux kernel layer 320, a system runtime library layer 340, an application framework layer 360, and an application layer 380. The Linux kernel layer 320, the system runtime library layer 340, and the application framework layer 360 belong to the operating system space, and the application layer 380 belongs to the user space. The Linux kernel layer 320 provides underlying drivers for various hardware components of electronic devices, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, power management, etc. The system runtime library layer 340 provides major feature support for the Android system through some C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D drawing support, and the Webkit library provides browser kernel support. The system runtime library layer 340 also provides the Android runtime library (Android runtime), which mainly provides some core libraries that allow developers to write Android applications using the Java language. The application framework layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider management, package management, call management, resource management, and location management. The application layer 380 runs at least one application. These applications can be native applications that come with the operating system, such as contacts, SMS, clock, and camera applications, or third-party applications developed by third-party developers, such as games, instant messaging programs, and photo enhancement programs.
[0344] Taking the operating system as the IOS system as an example, the programs and data stored in the memory 120 are as follows: Figure 11As shown, the IOS system includes: a core operating system layer 420 (Core OS layer), a core service layer 440 (CoreServices layer), a media layer 460 (Media layer), and a touchable layer 480 (Cocoa Touch Layer). The core operating system layer 420 includes the operating system kernel, drivers, and underlying program frameworks. These underlying program frameworks provide functions closer to the hardware for use by the program framework located in the core service layer 440. The core service layer 440 provides system services and / or program frameworks required by applications, such as the foundation framework, account framework, advertising framework, data storage framework, network connection framework, geographic location framework, motion framework, etc. The media layer 460 provides applications with audio-visual interfaces, such as graphics and image-related interfaces, audio technology-related interfaces, video technology-related interfaces, and wireless playback (AirPlay) interfaces for audio and video transmission technologies. The touchable layer 480 provides various commonly used interface-related frameworks for application development. The touchable layer 480 is responsible for user touch interaction operations on electronic devices. For example, local notification service, remote push service, advertising framework, game tool framework, message user interface (UI) framework, user interface UIKit framework, map framework, etc.
[0345] exist Figure 11 Among the frameworks shown, those relevant to most applications include, but are not limited to, the Foundation framework in the core services layer 440 and the UIKit framework in the touchable layer 480. The Foundation framework provides many basic object classes and data types, offering fundamental system services for all applications and having nothing to do with the UI. The classes provided by the UIKit framework are the foundational UI class library for creating touch-based user interfaces. iOS applications can use the UIKit framework to provide their UIs, providing the application infrastructure for building user interfaces, drawing, handling user interaction events, responding to gestures, and so on.
[0346] Among them, the method and principle of implementing data communication between third-party applications and the operating system in the IOS system can be referred to the Android system, and this manual will not go into details here.
[0347] Among them, the input device 130 is used to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone or a touch device. The output device 140 is used to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are touch screen displays, which are used to receive touch operations on or near the user using any suitable objects such as fingers and touch pens, and to display the user interface of each application. The touch screen display is usually provided on the front panel of the electronic device. The touch screen display can be designed as a full screen, a curved screen or a special-shaped screen. The touch screen display can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in the embodiments of this specification.
[0348] In addition, those skilled in the art will understand that the structures of the electronic devices shown in the above figures do not limit the electronic devices. The electronic devices may include more or fewer components than shown, or may combine certain components, or arrange the components differently. For example, the electronic devices may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, and other components, which are not described in detail here.
[0349] In the embodiments of this specification, the execution entity of each step can be the electronic device described above. Optionally, the execution entity of each step is the operating system of the electronic device. The operating system can be Android, iOS, or other operating systems, and this embodiment of this specification does not limit this.
[0350] The electronic device of the embodiment of this specification may also be equipped with a display device, which may be any device capable of realizing a display function, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), a plasma display panel (PDP), etc. The user may use the display device on the electronic device to view displayed text, images, videos and other information. The electronic device may be a smart phone, a tablet computer, a gaming device, an AR (Augmented Reality) device, a car, a data storage device, an audio playback device, a video playback device, a notebook, a desktop computing device, a wearable device such as an electronic watch, electronic glasses, an electronic helmet, an electronic bracelet, an electronic necklace, electronic clothing and the like.
[0351] exist Figure 8 In the electronic device shown, the processor 110 can be used to call the application stored in the memory 120 and specifically execute the method steps in one or more embodiments of this specification.
[0352] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0353] The above disclosure is only a preferred embodiment of this specification, and certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.
Claims
1. A method for processing a conversation, characterized in that: The method comprises: In vehicle driving scenarios, monitor the target user's in-vehicle behavior information and driving context information; Maintaining a user interest graph based on the in-vehicle behavior information; Based on the user interest graph and the driving context information, a driving conversation model is used to perform topic context reasoning to determine a target topic context for the target user, and a scenario conversation plan and opening conversation content are generated based on the target topic context; Actively conduct a dialogue with the target user based on the scenario dialogue plan and the opening dialogue content.
2. The method according to claim 1, characterized in that Maintaining the user interest graph based on the in-vehicle behavior information includes: Obtain the target user's in-vehicle behavior information during driving; Identify the target interest tag based on the vehicle-borne behavior information, and determine whether the target interest tag already has a target interest tag node in the user interest graph: If so, update the graph weight value corresponding to the target interest tag node in the user interest graph; If it does not exist, a target interest tag node corresponding to the target interest tag is added to the user interest graph, and a target graph weight value is set for the target interest tag node.
3. The method according to claim 2, characterized in that The updating of the graph weight value corresponding to the target interest tag in the user interest graph includes: Determine feedback impact information and behavior occurrence time interval for the target interest tag node based on the vehicle-borne behavior information, determine a weight change based on the feedback impact information, and determine a time attenuation factor based on the behavior occurrence time interval; Obtaining the original interest weight value of the target interest tag node, and performing weight update using a weight update calculation formula based on the weight change and the time attenuation factor to obtain an updated interest weight value; The weight update calculation formula satisfies the following formula: W_new=W_old×λ+ΔW Wherein, the W_new represents the updated interest weight value, the W_old represents the original interest weight value, the λ represents the time decay factor, and the ΔW represents the weight change amount.
4. The method according to claim 1, wherein The determining of a target topic context for the target user by performing topic context reasoning based on the user interest graph and the driving context information using a driving conversation macro model includes: Inputting the user interest graph and the driving context information into a driving dialogue macromodel, encoding the user interest graph using the driving dialogue macromodel to generate a user interest embedding representation, encoding the driving context information to generate a driving context representation, and fusing the user interest embedding representation with the driving context representation to obtain a candidate context reference representation; Topic reasoning is performed based on the candidate context reference representation to obtain a target topic context for the target user.
5. The method according to claim 4, characterized in that The performing topic reasoning based on the candidate context reference representation to obtain a target topic context for the target user includes: Determine multiple candidate topic contexts for target users; Performing adaptive reasoning based on the candidate context reference representation and the candidate topic context to obtain user interest similarity, context adaptation and external event relevance, and determining a candidate topic score for each candidate topic context based on the user interest similarity, context adaptation and external event relevance; A target topic context for the target user is determined based on the candidate topic scores.
6. The method according to claim 1, characterized in that The generating of the scenario dialogue plan and the opening dialogue content based on the target topic context includes: Determining field information based on the target topic context, the field information including a target topic category field, a content anchor keyword field, a suggested tone information field, an expected user interaction target field, and a dialogue branch planning strategy field; Field matching and filling are performed based on a preset dialogue strategy template and the field information to generate a structured scenario dialogue plan, and the opening dialogue content is determined based on the scenario dialogue plan.
7. The method according to claim 1, characterized in that The active dialogue processing with the target user based on the scenario dialogue plan and the opening dialogue content includes: Outputting the opening dialogue content to the target user; Monitor the target user's feedback behavior and determine the user response type based on the user feedback behavior; The scenario dialogue plan is used to perform dialogue processing based on the user response type.
8. The method according to claim 7, characterized in that The process of performing dialogue processing based on the user response type and using the scenario dialogue plan includes: If the user response type is a positive response type, obtaining the user's answer content to the user feedback behavior, calling the multi-round dialogue planning strategy and preset interaction goals in the scenario dialogue plan based on the user's answer content, selecting the next round of dialogue path and generating the next round of dialogue content through the driving dialogue macro model based on the user's answer content and the preset interaction goals, and outputting the next round of dialogue content; If the user response type is a non-response type or a negative response type, the dialogue response information is determined according to the dialogue termination strategy in the scenario dialogue plan, and the dialogue processing is performed based on the dialogue response information.
9. A dialogue processing device, characterized in that: The device comprises: The information monitoring module is used to monitor the target user's in-vehicle behavior information and driving context information in the vehicle driving scenario; An information processing module, configured to maintain a user interest graph based on the in-vehicle behavior information; The information processing module is configured to perform topic context reasoning using a large driving conversation model based on the user interest graph and the driving context information to determine a target topic context for the target user, and generate a scenario conversation plan and opening conversation content based on the target topic context; An active dialogue module is used to conduct active dialogue processing with the target user based on the scenario dialogue plan and the opening dialogue content.
10. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and apparatus for obtaining Web browsing interest of user
CN105930507A
Voice output method and device, storage medium and electronic equipment
CN111951787A
Active dialogue triggering method, device and equipment, vehicle and storage medium
CN116701581A
Advertisement recommendation method and system based on fusion neural network
CN119273408A
Topic content recommendation method, electronic device and storage medium based on large model
CN119782636A