Driving processing method and device, storage medium and electronic equipment

By acquiring and analyzing user interaction data through the in-vehicle interactive large model, the action and voice responses of the in-vehicle companion virtual image are generated, which solves the problem of insufficient interactive functions of the in-vehicle companion robot and realizes real emotional interaction and rich companionship experience.

CN120803266APending Publication Date: 2025-10-17BEIJING VISION WORLD TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510907586.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

How to enrich the interactive functions of in-vehicle companion robots so that they can better understand the user's emotional state during driving and respond interactively with real emotions.

Method used

The user's companionship interaction data is obtained through the in-vehicle interaction large model, emotional state inference is performed, the response emotion category and dialogue response content are determined, companionship action response data is generated, and the in-vehicle companionship virtual image is controlled to perform corresponding actions and voice output.

Benefits of technology

The interactive methods of the in-car companion virtual image have been enriched, making its interaction method more realistic and emotional, increasing the sense of companionship.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803266A_ABST
    Figure CN120803266A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a driving processing method and device, a storage medium and electronic equipment, and the method comprises the steps: obtaining accompanying interaction data inputted by a user for a vehicle-mounted accompanying virtual image in a vehicle driving scene, on the basis of the accompanying interaction data, a vehicle-mounted interaction large model is adopted to conduct emotional state reasoning to obtain a response emotion category aiming at the user and determine dialogue response content aiming at the user, and accompanying action response data aiming at the vehicle-mounted accompanying virtual image is determined on the basis of the response emotion category and the dialogue response content; and controlling the vehicle-mounted accompanying virtual image to perform user accompanying response processing based on the accompanying action response data. Therefore, the interaction mode of the vehicle-mounted accompanying virtual image is enriched, the interaction mode of the vehicle-mounted accompanying virtual image has real emotion, and the accompanying feeling of the vehicle-mounted accompanying virtual image is increased.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computers, and in particular to a driving processing method and device, a storage medium, and an electronic device. BACKGROUND

[0002] With the development of intelligent driving technology and artificial intelligence technology, a vehicle-mounted companion robot can accurately understand a user's voice instruction, and can easily control various functions of a vehicle or obtain real-time information such as navigation, weather, news, and the like when a driver is inconvenient to perform manual operation, thereby improving safety and convenience during driving.

[0003] In related technologies, the functions of a vehicle-mounted companion robot are constantly expanding, and therefore, how to enrich the interactive functions of a vehicle-mounted companion robot is a technical problem that needs to be solved by those skilled in the art. SUMMARY

[0004] Embodiments of the present application provide a driving processing method and device, a storage medium, and an electronic device. The technical solution is as follows:

[0005] In a first aspect, the embodiments of the present application provide a driving processing method, which comprises:

[0006] In a vehicle driving scenario, companion interaction data input by a user for a vehicle-mounted companion virtual image is acquired;

[0007] Based on the companion interaction data, a vehicle-mounted interactive large model is used to infer an emotional state to obtain a response emotion category for the user and determine a dialogue response content for the user;

[0008] Based on the response emotion category and the dialogue response content, companion action response data for the vehicle-mounted companion virtual image is determined;

[0009] Based on the companion action response data, the vehicle-mounted companion virtual image is controlled to perform user companion response processing.

[0010] In combination with the first aspect, in some possible implementations, the companion interaction data is input into a vehicle-mounted interactive large model, voice acoustic features, facial expression features, and dialogue input text are determined through the vehicle-mounted interactive large model, emotion recognition processing is performed based on the voice acoustic features, the facial expression features, and the dialogue input text to obtain a user emotion category, and emotion matching processing is performed based on the user emotion category to obtain a response emotion category for the user.

[0011] In combination with the first aspect, in some possible implementations, the companion interaction data is input into a vehicle-mounted interactive large model, voice acoustic features, facial expression features, and dialogue input text are determined through the vehicle-mounted interactive large model, emotion recognition processing is performed based on the voice acoustic features, the facial expression features, and the dialogue input text to obtain a user emotion category, and emotion matching processing is performed based on the user emotion category to obtain a response emotion category for the user.

[0012] match a dialogue response content based on the dialogue input text and the response sentiment category by the vehicle-mounted interactive large model.

[0013] In some possible implementation manners, the matching the dialogue response content based on the dialogue input text and the response sentiment category by the vehicle-mounted interactive large model includes:

[0014] determining a dialogue intent based on the dialogue input text by the vehicle-mounted interactive large model, determining dialogue response materials based on the dialogue intent, and performing dialogue expansion processing based on the dialogue response materials to obtain initial dialogue response content;

[0015] performing dialogue matching processing on the initial dialogue response content based on the response sentiment category by the vehicle-mounted interactive large model to obtain the dialogue response content.

[0016] In some possible implementation manners, the determining the dialogue response materials based on the dialogue intent includes:

[0017] determining a topic material preference corresponding to the user, and performing dialogue material association processing based on the dialogue intent and the topic material preference to obtain the dialogue response materials.

[0018] In some possible implementation manners, the determining the dialogue response content based on the response sentiment category and the dialogue intent includes:

[0019] determining a voice output mode and an image demonstration script of the vehicle-mounted companion virtual image based on the response sentiment category and the dialogue response content;

[0020] determining companion action data corresponding to the vehicle-mounted companion virtual image based on the image demonstration script;

[0021] generating companion voice data corresponding to the vehicle-mounted companion virtual image in the voice output mode based on the dialogue response content;

[0022] generating companion action response data based on the companion action data and the companion voice data.

[0023] In some possible implementation manners, the determining the dialogue response content based on the response sentiment category and the dialogue intent includes:

[0024] determining a target response strategy based on the response sentiment category and the dialogue response content by performing response strategy reasoning;

[0025] perform voice feature matching based on the target response strategy to determine a voice output mode;

[0026] perform demonstration expression matching based on the target response strategy to obtain expression design information, perform demonstration action matching based on the target response strategy to obtain action design information, and perform demonstration integration processing based on the expression design information and the action design information to obtain an image demonstration script of the in-vehicle companion virtual image.

[0027] In combination with the above embodiments, in some possible embodiments, the determining, based on the image demonstration script, of the companion action data corresponding to the in-vehicle companion virtual image includes:

[0028] generating, based on the image demonstration script, expression demonstration data of the in-vehicle companion virtual image, generating, based on the image demonstration script, action demonstration data of the in-vehicle companion virtual image, and synthesizing the companion action data based on the expression demonstration data and the action demonstration data.

[0029] In a second aspect, an embodiment of the present application provides a driving processing apparatus, the apparatus comprising:

[0030] a data acquisition module configured to acquire, in a vehicle driving scene, companion interaction data input by a user for an in-vehicle companion virtual image;

[0031] a data processing module configured to perform, based on the companion interaction data, emotion state reasoning by using an in-vehicle interactive large model to obtain a response emotion category for the user and determine a dialogue response content for the user;

[0032] a data generation module configured to determine, based on the response emotion category and the dialogue response content, companion action response data for the in-vehicle companion virtual image;

[0033] a response processing module configured to control the in-vehicle companion virtual image to perform user companion response processing based on the companion action response data.

[0034] Optionally, the data processing module comprises:

[0035] an emotion recognition unit configured to input the companion interaction data to the in-vehicle interactive large model, determine, by using the in-vehicle interactive large model, voice acoustic features, facial expression features, and dialogue input text, perform emotion recognition processing based on the voice acoustic features, the facial expression features, and the dialogue input text to obtain a user emotion category, and perform emotion matching processing based on the user emotion category to obtain a response emotion category for the user;

[0036] a dialogue matching unit configured to match, by using the in-vehicle interactive large model, dialogue response content based on the dialogue input text and the response emotion category.

[0037] Optionally, the dialogue matching unit comprises:

[0038] The first dialogue generation subunit is configured to determine a dialogue intention based on the dialogue input text by using the vehicle-mounted interactive large model, determine dialogue response materials based on the dialogue intention, and perform dialogue expansion processing based on the dialogue response materials to obtain initial dialogue response content.

[0039] The second dialogue generation subunit is configured to perform dialogue matching processing on the initial dialogue response content based on the response emotion category by using the vehicle-mounted interactive large model to obtain dialogue response content.

[0040] Optionally, the first dialogue generation subunit is specifically configured to:

[0041] determine the topic material preference corresponding to the user, perform dialogue material association processing based on the dialogue intention and the topic material preference to obtain dialogue response materials.

[0042] Optionally, the data generation module comprises:

[0043] The first data generation unit is configured to determine a voice output mode and an image demonstration script of the vehicle-mounted companion virtual image based on the response emotion category and the dialogue response content.

[0044] The second data generation unit is configured to determine companion action data corresponding to the vehicle-mounted companion virtual image based on the image demonstration script.

[0045] The third data generation unit is configured to generate companion voice data corresponding to the vehicle-mounted companion virtual image based on the dialogue response content and using the voice output mode.

[0046] The fourth data generation unit is configured to generate companion action response data based on the companion action data and the companion voice data.

[0047] Optionally, the first data generation unit is specifically configured to:

[0048] determine a target response strategy based on the response emotion category and the dialogue response content;

[0049] perform voice feature matching based on the target response strategy to determine a voice output mode;

[0050] perform demonstration expression matching based on the target response strategy to obtain expression design information, perform demonstration action matching based on the target response strategy to obtain action design information, and perform demonstration integration processing based on the expression design information and the action design information to obtain an image demonstration script of the vehicle-mounted companion virtual image.

[0051] Optionally, a second data generation unit, specifically configured to:

[0052] generate expression demonstration data of the vehicle-mounted companion virtual image based on the image demonstration script, generate action demonstration data of the vehicle-mounted companion virtual image based on the image demonstration script, and synthesize companion action data based on the expression demonstration data and the action demonstration data.

[0053] In a third aspect, embodiments of the present application provide a computer storage medium, which has a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method described above.

[0054] In a fourth aspect, embodiments of the present application provide an electronic device, which can include a memory and a processor, wherein the memory stores a computer program, and the computer program is suitable for being loaded by the memory and executing the method described above.

[0055] The technical solutions provided by the embodiments of the present application have at least the following beneficial effects:

[0056] The driving processing method provided by the embodiments of the present application can obtain companion interaction data input by a user for a vehicle-mounted companion virtual image in a vehicle driving scene, infer a response emotion category of the user and determine a dialogue response content of the user based on the companion interaction data and using a vehicle-mounted interactive large model, determine companion action response data of the vehicle-mounted companion virtual image based on the response emotion category and the dialogue response content, and control the vehicle-mounted companion virtual image to perform user companion response processing based on the companion action response data. Thus, the response emotion category and the dialogue response content of the vehicle-mounted companion virtual image responding to the user can be determined according to the companion interaction data input by the user, the companion action response data of the vehicle-mounted companion virtual image responding to the user can be determined according to the response emotion category and the dialogue response content, the vehicle-mounted companion virtual image is controlled to perform companion response processing on the user, the interactive mode of the vehicle-mounted companion virtual image is enriched, the interactive mode of the vehicle-mounted companion virtual image has real emotions, and the companion feeling of the vehicle-mounted companion virtual image is increased. BRIEF DESCRIPTION OF DRAWINGS

[0057] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0058] Figure 1 is a flowchart of a driving processing method provided by an embodiment of the present application;

[0059] Figure 2 is a flowchart of another driving processing method provided by an embodiment of the present application;

[0060] Figure 3 is a structural diagram of a driving processing device provided by an embodiment of the present application;

[0061] Figure 4 is a structural diagram of a data processing module provided by an embodiment of the present application;

[0062] Figure 5 is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0063] In order to make the purpose, characteristics and advantages of the embodiments of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0064] In the description of the present application, it should be understood that the terms "first", "second" and the like are used only for the purpose of description, and should not be understood as indicating or implying relative importance. In the description of the present application, it should be noted that, unless otherwise specified and limited, "including" and "having" and any variants thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device. Those skilled in the art can understand the specific meaning of the above terms in the present application according to the specific circumstances. In addition, in the description of the present application, "multiple" means two or more, unless otherwise specified. "And / or" describes the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects.

[0065] The present application will be described in detail below with reference to specific embodiments.

[0066] In one embodiment, as Figure 1As shown, a driving processing method is particularly proposed, which can be implemented by a computer program and run on a driving processing device based on a von Neumann architecture. The computer program can be integrated in an application or run as a standalone tool application. The driving processing method can be applied to an electronic device, which can be a vehicle-mounted device, a vehicle-mounted computer, a smart terminal, a computing device, etc.

[0067] Specifically, the driving processing method includes:

[0068] S101, in a vehicle driving scenario, obtaining companion interaction data input by a user for a vehicle-mounted companion virtual image.

[0069] It can be understood that the vehicle-mounted companion virtual image can refer to a virtual character or a virtual assistant used in a vehicle. The vehicle-mounted companion virtual image can be configured in a vehicle-mounted system and can interact with a driver or a passenger to provide companionship, entertainment, information query, navigation, etc.

[0070] The companion interaction data can refer to voice data or video data of the user collected in the process of the user conversing with the vehicle-mounted companion virtual image.

[0071] In the vehicle driving scenario, the user can wake up the vehicle-mounted companion virtual image through voice instructions or touch instructions. When it is monitored that the user inputs voice instructions or touch instructions for waking up the vehicle-mounted companion virtual image, the collected companion interaction data is obtained. The companion interaction data can include user voice data, the companion interaction data can also include user video data, and the companion interaction data can further include user voice data and user video data.

[0072] S102, based on the companion interaction data, using a vehicle-mounted interactive large model to infer a response emotion category for the user and determine a dialogue response content for the user.

[0073] It can be understood that the response emotion category refers to the emotion category exhibited by the vehicle-mounted companion virtual image in the process of the vehicle-mounted companion virtual image conversing with the user. The dialogue response content refers to the dialogue content responded by the vehicle-mounted companion virtual image to the user in the process of the vehicle-mounted companion virtual image conversing with the user.

[0074] In some embodiments, performing step S102 can include: inputting the companion interaction data to the vehicle-mounted interactive large model, the vehicle-mounted interactive large model determining voice acoustic features and dialogue input text, performing emotion recognition based on the voice acoustic features and the dialogue input text to determine a user emotion category, matching a response emotion category for the user based on the user emotion category, and matching a dialogue response content for the user based on the response emotion category and the dialogue input text.

[0075] In some embodiments, the step S102 can include: inputting the number of companion interaction users into the vehicle interactive large model, the vehicle interactive large model determining the voice acoustic features, the facial expression features and the dialogue input text, performing emotion recognition processing based on the voice acoustic features, the facial expression features and the dialogue input text to obtain the user emotion category, performing emotion matching processing based on the user emotion category to obtain the response emotion category for the user, and matching the dialogue response content based on the dialogue input text and the response emotion category.

[0076] In the embodiments of the present application, the vehicle interactive large model can be obtained based on a multi-modal large language model.

[0077] Optionally, the training process of the vehicle interactive large model can include: creating an initial vehicle interactive large model based on a multi-modal large model; determining sample companion interaction data input by a sample user for a vehicle companion virtual image, and labeling the sample companion interaction data with a response emotion category label, a dialogue response content label and a companion action response data label; performing at least one round of model training on the initial vehicle interactive large model using the sample companion interaction data; in the forward propagation training process of the model, calling the initial vehicle interactive large model based on the sample companion interaction data to perform emotion state reasoning to obtain a predicted response emotion category for the sample user and determine a predicted dialogue response content for the sample user, and determining predicted companion action response data for the vehicle companion virtual image based on the predicted response emotion category and the predicted dialogue response content; in the backward propagation training process of the model, determining a first loss value based on the predicted response emotion category and the response emotion category label, determining a second loss value based on the predicted dialogue response content and the dialogue response content label, determining a third loss value based on the predicted companion action response data and the companion action response data label, and determining a model comprehensive loss based on the first loss value, the second loss value and the third loss value; and performing model parameter adjustment on the initial vehicle interactive large model based on the model comprehensive loss to obtain the vehicle interactive large model after model training.

[0078] Optionally, the first loss value, the second loss value and the third loss value can be determined using any one of hinge loss functions, contrast loss functions, Euclidean distance loss functions or cross-entropy loss functions in related technologies.

[0079] Optionally, the model end training condition for obtaining the vehicle interactive large model can include that the value of the loss function is less than or equal to a preset loss function threshold, the number of iterations reaches a preset number threshold, etc. The model end training condition can be determined based on actual conditions, which is not limited here.

[0080] Optionally, the creation of an initial in-vehicle interaction big model based on the multimodal big model can be performed as follows: obtaining the multimodal big model, creating an initial in-vehicle companionship interaction scene adaptation module for the in-vehicle companionship interaction scene and a big language generation module based on the multimodal big model, and forming an initial multimodal big model based on the big language generation module and the initial in-vehicle companionship interaction scene adaptation module.

[0081] Parameter adjustment of the initial in-vehicle interaction model based on the model comprehensive loss is performed to obtain the trained in-vehicle interaction model. This can include adjusting the model parameters of the initial in-vehicle companionship interaction scenario adaptation module within the initial in-vehicle interaction model based on the model comprehensive loss, while maintaining the model parameters of the large language generation module unchanged. This process continues until the model training termination conditions are met, resulting in the large language generation module and the in-vehicle companionship interaction scenario adaptation module. This completes the model fusion of the large language generation module and the in-vehicle companionship interaction scenario adaptation module, resulting in the trained in-vehicle interaction model.

[0082] Optionally, the model fusion of the large language generation module and the in-vehicle companionship interaction scene adaptation module can be: weight fusion of the model structure layer weight of the in-vehicle companionship interaction scene adaptation module and the large language generation module, by determining the target model structure layer corresponding to the model structure layer weight in the large language generation module, and performing parameter fusion of the model structure layer parameters of the target model structure layer and the model structure layer weight. The model structure layer weight of the in-vehicle companionship interaction scene adaptation module may only partially correspond to and have model structure layer weights in all model structure layers in the multimodal large model. By completing the parameter update of the model structure layer based on the model structure layer weight for this part of the target model structure layer, and so on, the reference update process of all model structure layer weights is completed, thereby obtaining the in-vehicle interaction large model.

[0083] S103: Determine companion action response data for the in-vehicle companion virtual image based on the response emotion category and the dialogue response content.

[0084] It can be understood that the companion action response data refers to configuration data that records the expressions, actions and output voice of the in-vehicle companion virtual image when responding to the user.

[0085] In some embodiments, executing step S103 may include: determining companion action data and companion voice data for the in-vehicle companion virtual image based on the response emotion category and the dialogue response content, and generating companion action response data based on the companion action data and companion voice data.

[0086] S104: Control the in-vehicle companion virtual image to perform user companionship response processing based on the companionship action response data.

[0087] Specifically, the in-vehicle companion virtual image is controlled to synchronously output voice, display expressions and body movements according to the companion action response data.

[0088] The driving processing method provided by the embodiments of the present application can obtain companion interaction data input by a user for an in-vehicle companion virtual image in a vehicle driving scenario, infer a response emotion category for the user and determine a dialogue response content for the user based on the companion interaction data by using an in-vehicle interactive large model, determine companion action response data for the in-vehicle companion virtual image based on the response emotion category and the dialogue response content, and control the in-vehicle companion virtual image to perform user companion response processing based on the companion action response data. Thus, the response emotion category and the dialogue response content of the in-vehicle companion virtual image in response to the user can be determined according to the companion interaction data input by the user, the companion action response data of the in-vehicle companion virtual image in response to the user can be determined according to the response emotion category and the dialogue response content, the in-vehicle companion virtual image can be controlled to perform companion response processing for the user, the interactive mode of the in-vehicle companion virtual image is enriched, the interactive mode of the in-vehicle companion virtual image has real emotions, and the companionship of the in-vehicle companion virtual image is improved.

[0089] Next, please refer to Figure 2 , a flowchart of another embodiment of a driving processing method proposed by the present application.

[0090] Specifically, the driving processing method comprises:

[0091] S201, in a vehicle driving scenario, obtaining companion interaction data input by a user for an in-vehicle companion virtual image.

[0092] In the vehicle driving scenario, the user can wake up the in-vehicle companion virtual image through a voice instruction or a touch instruction. When it is monitored that the user inputs a voice instruction or a touch instruction for waking up the in-vehicle companion virtual image, the collected companion interaction data is obtained. The companion interaction data can include user voice data and user video data. The user video data is video data of a user facial expression collected in the process of collecting the user voice data.

[0093] S202, inputting the companion interaction data to an in-vehicle interactive large model, determining voice acoustic features, facial expression features and dialogue input text by the in-vehicle interactive large model, performing emotion recognition processing to obtain a user emotion category based on the voice acoustic features, the facial expression features and the dialogue input text, and performing emotion matching processing to obtain a response emotion category for the user based on the user emotion category.

[0094] Specifically, the accompanying interaction data includes user voice data and user video data, and the vehicle-mounted interactive large model extracts speech acoustic features from the user voice data, the speech acoustic features including prosodic features, spectral features and voice quality features, the prosodic features can include fundamental frequency, intensity, duration, pitch, pause, speed, and length, the spectral features can include spectral energy distribution (formant), linear predictive cepstral coefficient, mel frequency cepstral coefficient, and the voice quality features can include glottal parameters, frequency perturbation, amplitude perturbation, formant frequency machine bandwidth, etc. The vehicle-mounted interactive large model extracts user facial images from the user video data, and extracts facial expression features from the user facial images, the facial expression features including facial key point features (such as eyebrows, eyes, mouth, etc.). The vehicle-mounted interactive large model performs speech-to-text processing on the user voice data to obtain dialogue input text, and extracts text features (such as keywords expressing emotions) from the dialogue input text by the vehicle-mounted interactive large model. The vehicle-mounted interactive large model performs multi-modal feature fusion on the speech acoustic features, the facial expression features and the text features to obtain target emotion features, and performs emotion recognition on the target emotion features to determine the user emotion category.

[0095] The emotion matching processing based on the user emotion category is performed to obtain a response emotion category for the user, specifically, an emotion response mapping table can be obtained, the emotion response mapping table is configured with reference response emotion categories corresponding to different user emotion categories, a reference response emotion category corresponding to the user emotion category is queried in the emotion response mapping table, and the reference response emotion category corresponding to the user emotion category is determined as the response emotion category for the user.

[0096] For example, when the user emotion category is happy, the response emotion category can also be happy; when the user emotion category is sad or sad, the response emotion category can be at least one of comfort, concern and gentleness.

[0097] S203, the dialogue response content is matched by the vehicle-mounted interactive large model based on the dialogue input text and the response emotion category.

[0098] In some embodiments, performing step S203 can include: A1 determining a dialogue intent based on the dialogue input text by the vehicle-mounted interactive large model, determining a dialogue response material based on the dialogue intent, and performing dialogue expansion processing based on the dialogue response material to obtain initial dialogue response content; A2: performing dialogue matching processing on the initial dialogue response content based on the response emotion category by the vehicle-mounted interactive large model to obtain the dialogue response content.

[0099] The dialogue intent is used to represent the core demand or purpose expressed by the user in the dialogue. The dialogue response material refers to the selectable answer content prepared according to the dialogue intent.

[0100] In step A1, the dialogue input text is input to the vehicle-mounted interactive large model, and the vehicle-mounted interactive large model determines the dialogue intent according to semantic analysis processing of the dialogue input text.

[0101] The dialogue response material is determined based on the dialogue intent, which can be: querying the dialogue response material corresponding to the dialogue intent in the dialogue response library. The response library can pre-configure response materials corresponding to different kinds of intents.

[0102] The dialogue response material is determined based on the dialogue intent, which can also be: determining the topic material preference of the user, and determining the dialogue response material based on the dialogue intent and the topic material preference. Specifically, in determining the topic material preference of the user, the target identity information of the user is identified, and when the target identity information indicates that the user is a preset driver, the topic material preference is determined according to the historical dialogue content of the user. When the target identity information indicates that the user is not a preset driver, the age information and gender information of the user are identified according to the user voice data or user video data, and the topic material preference is inferred according to the age information and gender information of the user. It can be understood that the topic material preference can include the topic material that the user is interested in or the topic material that the user does not like. The topic material that the user is interested in can be selected from the response material associated with the dialogue intent to determine the dialogue response material, or the topic material that the user does not like can be excluded from the response material associated with the dialogue intent to obtain the dialogue response material.

[0103] In step A2, the vehicle-mounted interactive large model can adjust the words and / or sentence patterns of the initial dialogue response content according to the response emotion category to obtain the dialogue response content, such as generating the dialogue response content with the appropriate words and / or sentence patterns that match the response emotion category.

[0104] S204, determining the voice output mode and image demonstration script of the vehicle-mounted companion virtual image based on the response emotion category and the dialogue response content.

[0105] It can be understood that the voice output mode refers to the voice expression mode of the vehicle-mounted companion virtual image when interacting with the user. The voice output mode can specifically include tone, pitch, speed, emotional color, etc.

[0106] The image demonstration script refers to the action, expression or animation design of the vehicle-mounted companion virtual image when interacting with the user. The image demonstration script can specifically include facial expression, body movement, posture, eye design, etc.

[0107] In some embodiments, step S204 is performed, which can specifically include: B1: determining a target response strategy based on the response emotion category and the dialogue response content; B2: determining a voice output mode based on the target response strategy; B3: matching a demonstration expression based on the target response strategy to obtain expression design information, matching a demonstration action based on the target response strategy to obtain action design information, and integrating the expression design information and the action design information to obtain a demonstration script of the image of the in-vehicle companion virtual image.

[0108] It can be understood that the target response strategy refers to a strategy for determining the performance behavior of the in-vehicle companion virtual image when responding to the user. The expression design information refers to design information of facial movements made by the in-vehicle companion virtual image when responding to the user. The expression design information can include expression type, duration, change process, and the like. The action design information refers to design information of body movements made by the in-vehicle companion virtual image when responding to the user. The action design information can include gestures, postures, action sequences, and the like.

[0109] Specifically, the response emotion category and the dialogue response content are input into the in-vehicle interactive large model. The in-vehicle interactive large model queries a preset response strategy knowledge base to determine a target response strategy that is adapted to the response emotion category and the dialogue response content.

[0110] The in-vehicle interactive large model queries a preset voice feature library to determine voice features, such as tone, speech rate, tone of voice, and the like, that match the target response strategy, and generates a voice output mode including the voice features.

[0111] The in-vehicle interactive large model queries a preset expression design library to determine expression design information that matches the target response strategy, queries a preset action design library to determine action design information that matches the target response strategy, and generates a demonstration script of the image in combination with the expression design information and the action design information. The demonstration script of the image can specifically include time points, durations of each action and expression, and synchronization modes of actions and expressions, and the like.

[0112] For example, a target response strategy can be: according to the dialogue response content, a pleasant and relaxed dialogue atmosphere needs to be created; the voice output mode corresponding to the target response strategy can include voice features such as tone of voice, tone, and speech rate, wherein the tone of voice is lively and infectious, the tone is bright, and the speech rate is moderate or slightly fast; and the demonstration script of the image corresponding to the target response strategy can be: showing a smile, eyes curved, eyebrows slightly raised, a relaxed body posture, and a body movement of nodding or waving hands.

[0113] For example, a target response strategy can be: according to the dialogue response content, a comforting and gentle dialogue context needs to be created; the voice output mode corresponding to the target response strategy can include voice features such as tone, pitch, and speech rate, wherein the tone is a concerned and slightly calm tone, the pitch is a low and steady pitch, and the speech rate is a slower speech rate; the image demonstration script corresponding to the target response content can be: a gentle and concerned expression is shown, the eyebrows are slightly wrinkled, the eyes are soft, and the body movement is a slight nod or a comforting gesture.

[0114] In some embodiments, the step S205 can include: generating expression demonstration data of the in-vehicle companion virtual image based on the image demonstration script, generating action demonstration data of the in-vehicle companion virtual image based on the image demonstration script, and synthesizing the companion action data based on the expression demonstration data and the action demonstration data.

[0115] In some embodiments, the step S205 can include: generating expression demonstration data of the in-vehicle companion virtual image based on the image demonstration script, generating action demonstration data of the in-vehicle companion virtual image based on the image demonstration script, and synthesizing the companion action data based on the expression demonstration data and the action demonstration data.

[0116] It can be understood that the expression demonstration data refers to configuration data for configuring the expression of the in-vehicle companion virtual image. The action demonstration data refers to configuration data for configuring the action of the in-vehicle companion virtual image. The companion action data is configuration data in which the in-vehicle companion virtual image performs the expression according to the expression demonstration data and performs the action according to the action demonstration data.

[0117] Specifically, expression design information is extracted from the image demonstration script, expression demonstration data of the in-vehicle companion virtual image is generated according to the expression design information, action design information is extracted from the image demonstration script, action demonstration data of the in-vehicle companion virtual image is generated according to the action design information, and the expression demonstration data and the action demonstration data are spliced to synthesize the companion action data.

[0118] S206, generating companion voice data corresponding to the in-vehicle companion virtual image based on the dialogue response content using the voice output mode.

[0119] Specifically, the dialogue response content is processed by voice synthesis using the voice output mode to generate companion voice data corresponding to the in-vehicle companion virtual image. The companion voice data is voice data in which the in-vehicle companion virtual image speaks the dialogue response content in the voice output mode indicated by the voice output mode.

[0120] S207, generating companion action response data based on the companion action data and the companion voice data.

[0121] In some embodiments, step S207 is performed, which can specifically include: determining the action execution time and the expression execution time based on the accompanying action data, determining the speech output time based on the accompanying speech data, determining the action execution time period, the expression execution time period, and the speech output time period based on the action execution time, the expression execution time, and the speech output time, performing time configuration processing on the accompanying action data based on the action execution time period and the expression execution time period to obtain target accompanying action data, performing time configuration processing on the accompanying speech data to obtain target accompanying speech data, and generating accompanying action response data including the target accompanying action data and the target accompanying speech data.

[0122] In step S208, the vehicle-mounted accompanying virtual image is controlled to perform user accompanying response processing based on the accompanying action response data.

[0123] In some embodiments, step S208 is performed, which can specifically include: determining a target timeline of the accompanying action response data, and controlling the vehicle-mounted accompanying virtual image to output speech, display expressions, and body actions according to the target timeline.

[0124] In the driving processing method provided in the embodiments of the present application, in a vehicle driving scene, accompanying interaction data input by a user for a vehicle-mounted accompanying virtual image is obtained, the accompanying interaction data is input to a vehicle-mounted interactive large model, speech acoustic features, facial expression features, and dialogue input text are determined by the vehicle-mounted interactive large model, a user emotion category is obtained by performing emotion recognition processing based on the speech acoustic features, the facial expression features, and the dialogue input text, and a response emotion category for the user is obtained by performing emotion matching processing based on the user emotion category. In this way, the user emotion category is accurately identified by the vehicle-mounted interactive large model based on the comprehensive analysis of various types of features, so as to accurately identify the response emotion category. Subsequently, dialogue response content is matched based on the dialogue input text and the response emotion category by the vehicle-mounted interactive large model, a speech output mode and an image demonstration script of the vehicle-mounted accompanying virtual image are determined based on the response emotion category and the dialogue response content, accompanying action data corresponding to the vehicle-mounted accompanying virtual image is determined based on the image demonstration script, accompanying speech data corresponding to the vehicle-mounted accompanying virtual image is generated in the speech output mode based on the dialogue response content, and accompanying action response data is generated based on the accompanying action data and the accompanying speech data. In this way, the accompanying action response data with a multi-modal interaction mode is generated by the vehicle-mounted interactive large model. Then, the vehicle-mounted accompanying virtual image is controlled to perform user accompanying response processing based on the accompanying action response data. In this way, the vehicle-mounted accompanying virtual image can have a rich interaction mode, and the interaction mode of the vehicle-mounted accompanying virtual image has real emotions, thereby increasing the accompanying feeling of the vehicle-mounted accompanying virtual image.

[0125] The embodiments of the present application will be described below in conjunction with Figure 3 The driving processing device provided in the embodiments of the present application will be described in detail. It should be noted that, Figure 3The driving processing apparatus shown is used to execute the application Figures 1-2 The method of the embodiment shown only shows parts related to the embodiment of the application for ease of illustration, and specific technical details not disclosed are referred to the application Figures 1-2 The embodiment shown.

[0126] See Figure 3 , which shows a structural schematic diagram of the driving processing apparatus of the embodiment of the application. The driving processing apparatus 1 can be realized by software, hardware or a combination of both to become all or part of the apparatus. According to some embodiments, the driving processing apparatus 1 includes a data acquisition module 11, a data processing module 12, a data generation module 13 and a response processing module 14, specifically used for:

[0127] The data acquisition module 11 is used to acquire companion interaction data input by a user for an in-vehicle companion virtual image in a vehicle driving scene;

[0128] The data processing module 12 is used to infer a response emotion category for the user and determine a dialogue response content for the user based on the companion interaction data using an in-vehicle interactive large model;

[0129] The data generation module 13 is used to determine companion action response data for the in-vehicle companion virtual image based on the response emotion category and the dialogue response content;

[0130] The response processing module 14 is used to control the in-vehicle companion virtual image to perform user companion response processing based on the companion action response data.

[0131] Optionally, see Figure 4 The structural schematic diagram of the data processing module 12 shown, the data processing module 12 includes an emotion recognition unit 121 and a dialogue matching unit 122, specifically used for:

[0132] The emotion recognition unit 121 is used to input the companion interaction data to the in-vehicle interactive large model, determine voice acoustic features, facial expression features and dialogue input text through the in-vehicle interactive large model, perform emotion recognition processing based on the voice acoustic features, the facial expression features and the dialogue input text to obtain a user emotion category, and perform emotion matching processing based on the user emotion category to obtain a response emotion category for the user;

[0133] The dialogue matching unit 122 is used to match dialogue response content based on the dialogue input text and the response emotion category through the in-vehicle interactive large model.

[0134] Optionally, the dialogue matching unit 122 includes:

[0135] The first dialogue generation subunit is configured to determine a dialogue intention based on the dialogue input text by using the vehicle-mounted interactive large model, determine dialogue response materials based on the dialogue intention, and perform dialogue expansion processing based on the dialogue response materials to obtain initial dialogue response content.

[0136] The second dialogue generation subunit is configured to perform dialogue matching processing on the initial dialogue response content based on the response emotion category by using the vehicle-mounted interactive large model to obtain dialogue response content.

[0137] Optionally, the first dialogue generation subunit is specifically configured to:

[0138] determine a topic material preference corresponding to the user, and perform dialogue material association processing based on the dialogue intention and the topic material preference to obtain dialogue response materials.

[0139] Optionally, the data generation module 13 comprises:

[0140] The first data generation unit is configured to determine a voice output mode and an image demonstration script of the vehicle-mounted companion virtual image based on the response emotion category and the dialogue response content.

[0141] The second data generation unit is configured to determine companion action data corresponding to the vehicle-mounted companion virtual image based on the image demonstration script.

[0142] The third data generation unit is configured to generate companion voice data corresponding to the vehicle-mounted companion virtual image based on the dialogue response content and using the voice output mode.

[0143] The fourth data generation unit is configured to generate companion action response data based on the companion action data and the companion voice data.

[0144] Optionally, the first data generation unit is specifically configured to:

[0145] determine a target response strategy based on the response emotion category and the dialogue response content;

[0146] perform voice feature matching based on the target response strategy to determine a voice output mode;

[0147] perform demonstration expression matching based on the target response strategy to obtain expression design information, perform demonstration action matching based on the target response strategy to obtain action design information, and perform demonstration integration processing based on the expression design information and the action design information to obtain the image demonstration script of the vehicle-mounted companion virtual image.

[0148] Optionally, the second data generation unit is specifically configured to:

[0149] Generate expression demonstration data of the in-vehicle companion virtual image based on the image demonstration script, generate action demonstration data of the in-vehicle companion virtual image based on the image demonstration script, and synthesize companion action data based on the expression demonstration data and the action demonstration data.

[0150] The driving processing apparatus provided by the embodiments of the present application can obtain companion interaction data input by a user for an in-vehicle companion virtual image in a vehicle driving scenario, infer a response emotion category for the user and determine a dialogue response content for the user based on the companion interaction data and using an in-vehicle interactive large model, determine companion action response data for the in-vehicle companion virtual image based on the response emotion category and the dialogue response content, and control the in-vehicle companion virtual image to perform user companion response processing based on the companion action response data. In this way, the response emotion category and the dialogue response content of the in-vehicle companion virtual image in response to the user can be determined according to the companion interaction data input by the user, the companion action response data of the in-vehicle companion virtual image in response to the user can be determined according to the response emotion category and the dialogue response content, the in-vehicle companion virtual image can be controlled to perform companion response processing for the user according to the companion action response data, the interactive mode of the in-vehicle companion virtual image is enriched, the interactive mode of the in-vehicle companion virtual image has real emotions, and the companion feeling of the in-vehicle companion virtual image is increased.

[0151] Please refer to Figure 5 , Figure 5 A structural schematic diagram of an electronic device is provided for the embodiments of the present application. The electronic device can include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 can be connected through the bus 150.

[0152] The processor 110 can include one or more processing cores. The processor 110 connects various parts within the entire electronic device by various interfaces and lines, performs various functions of the electronic device and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 120, and calling data stored in the memory 120. Alternatively, the processor 110 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA). The processor 110 can integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes an operating system, a user interface, and an application program, etc.; the GPU is responsible for rendering and drawing display content; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 110, but be implemented by a separate communication chip.

[0153] The memory 120 can include a random access memory (RAM) and can also include a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 120 can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing each of the following method embodiments, etc., and the operating system can be an Android system, an IOS system developed by Apple Inc., a system developed based on the Android system or other systems.

[0154] In order to enable the operating system to distinguish the specific application scenarios of the third-party application, it is necessary to open up the data communication between the third-party application and the operating system, so that the operating system can obtain the current scenario information of the third-party application at any time, and then perform targeted system resource adaptation based on the current scenario.

[0155] The input device 130 is configured to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 140 is configured to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In an example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are a touch display screen.

[0156] The touch display screen can be designed as a full screen, a curved screen, or a special-shaped screen. The touch display screen can also be designed as a combination of a full screen and a curved screen, a combination of a special-shaped screen and a curved screen, and the present application does not limit this.

[0157] In addition, those skilled in the art can understand that the structure of the electronic device shown in the above-mentioned drawings does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than the drawings, or combine certain components, or different component arrangements. For example, the electronic device also includes radio frequency circuitry, an input unit, a sensor, audio circuitry, a wireless fidelity (WiFi) module, a power supply, a Bluetooth module, and the like, which are not described here.

[0158] In Figure 5 In the electronic device shown, the processor 110 can be configured to invoke the program of the driving processing method stored in the memory 120, and specifically perform the following operations:

[0159] In a vehicle driving scenario, obtaining companion interaction data input by a user for a vehicle-mounted companion virtual image;

[0160] Based on the companion interaction data, a vehicle-mounted interactive large model is used to infer the emotional state to obtain a response emotion category for the user and determine a dialogue response content for the user;

[0161] Based on the response emotion category and the dialogue response content, companion action response data for the vehicle-mounted companion virtual image is determined;

[0162] Based on the companion action response data, the vehicle-mounted companion virtual image is controlled to perform user companion response processing.

[0163] In some embodiments, when performing the step of obtaining a response emotion category for the user based on the companion interaction data using a vehicle-mounted interactive large model to infer the emotional state and determining a dialogue response content for the user, the processor 110 specifically performs the following operations:

[0164] inputting the companion interaction data into a vehicle-mounted interactive large model, determining, by the vehicle-mounted interactive large model, voice acoustic features, facial expression features, and dialogue input texts, performing emotion recognition processing based on the voice acoustic features, the facial expression features, and the dialogue input texts to obtain a user emotion category, and performing emotion matching processing based on the user emotion category to obtain a response emotion category for the user;

[0165] matching, by the vehicle-mounted interactive large model, dialogue response content based on the dialogue input texts and the response emotion category.

[0166] In some embodiments, when the processor 110 performs the step of matching, by the vehicle-mounted interactive large model, dialogue response content based on the dialogue input texts and the response emotion category, the processor 110 specifically performs the following operations:

[0167] determining, by the vehicle-mounted interactive large model, a dialogue intent based on the dialogue input texts, determining dialogue response materials based on the dialogue intent, and performing dialogue expansion processing based on the dialogue response materials to obtain initial dialogue response content;

[0168] performing, by the vehicle-mounted interactive large model, dialogue matching processing on the initial dialogue response content based on the response emotion category to obtain dialogue response content.

[0169] In some embodiments, when the processor 110 performs the step of determining dialogue response materials based on the dialogue intent, the processor 110 specifically performs the following operations:

[0170] determining a topic material preference corresponding to the user, and performing dialogue material association processing based on the dialogue intent and the topic material preference to obtain dialogue response materials.

[0171] In some embodiments, when the processor 110 performs the step of determining companion action response data for the vehicle-mounted companion virtual image based on the response emotion category and the dialogue response content, the processor 110 specifically performs the following operations:

[0172] determining a voice output mode and an image demonstration script of the vehicle-mounted companion virtual image based on the response emotion category and the dialogue response content;

[0173] determining companion action data corresponding to the vehicle-mounted companion virtual image based on the image demonstration script;

[0174] generating companion voice data corresponding to the vehicle-mounted companion virtual image based on the dialogue response content and using the voice output mode;

[0175] generating companion action response data based on the companion action data and the companion voice data.

[0176] In some embodiments, the processor 110, when performing the step of determining the voice output mode and the image demonstration script of the in-vehicle companion virtual image based on the response sentiment category and the dialogue response content, specifically performs the following operations:

[0177] determining a target response strategy based on the response sentiment category and the dialogue response content;

[0178] determining a voice output mode based on the target response strategy;

[0179] obtaining expression design information based on the target response strategy, obtaining action design information based on the target response strategy, and obtaining the image demonstration script of the in-vehicle companion virtual image based on the expression design information and the action design information.

[0180] In some embodiments, the processor 110, when performing the step of determining the companion action data corresponding to the in-vehicle companion virtual image based on the image demonstration script, specifically performs the following operations:

[0181] generating expression demonstration data of the in-vehicle companion virtual image based on the image demonstration script, generating action demonstration data of the in-vehicle companion virtual image based on the image demonstration script, and synthesizing the companion action data based on the expression demonstration data and the action demonstration data.

[0182] The embodiments of the present application also provide a computer readable storage medium, which stores at least one instruction, and the at least one instruction is used to be executed by a processor to implement the driving processing method according to the above various embodiments.

[0183] The embodiments of the present application also provide a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the driving processing method according to the above various embodiments.

[0184] Those skilled in the art should be aware that the functions described in the above one or more examples can be implemented in hardware, software, firmware or any combination thereof. When implemented in software, the functions can be stored in a computer readable medium or transmitted as one or more instructions or code on a computer readable medium. The computer readable medium includes computer storage medium and communication medium, and the communication medium includes any medium that facilitates the transfer of computer program from one place to another. The storage medium can be any available medium accessible by a general or special purpose computer.

[0185] The above merely provides the optional embodiments of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A driving processing method, characterized in that: The method comprises: In a vehicle driving scenario, obtaining companionship interaction data input by the user to the in-vehicle companion virtual image; Based on the companionship interaction data, the in-vehicle interaction big model is used to perform emotional state inference to obtain a response emotional category for the user and determine the content of the dialogue response for the user; determining companion action response data for the in-vehicle companion virtual image based on the response emotion category and the dialogue response content; The vehicle-mounted companionship virtual image is controlled to perform user companionship response processing based on the companionship action response data.

2. The method according to claim 1, characterized in that The step of performing emotional state inference based on the companionship interaction data using the in-vehicle interaction large model to obtain a response emotional category for the user and determine the content of the dialogue response for the user includes: Inputting the companionship interaction data into an in-vehicle interaction large model, determining speech acoustic features, facial expression features, and conversation input text through the in-vehicle interaction large model, performing emotion recognition processing based on the speech acoustic features, facial expression features, and conversation input text to obtain a user emotion category, and performing emotion matching processing based on the user emotion category to obtain a response emotion category for the user; The in-vehicle interaction model is used to match the dialogue response content based on the dialogue input text and the response emotion category.

3. The method according to claim 2, characterized in that The matching of the dialogue response content based on the dialogue input text and the response emotion category by the in-vehicle interaction macro model includes: Determining a conversation intent based on the conversation input text using the in-vehicle interaction macro model, determining a conversation response material based on the conversation intent, and performing conversation expansion processing based on the conversation response material to obtain initial conversation response content; The in-vehicle interaction large model performs dialogue matching processing on the initial dialogue response content based on the response emotion category to obtain the dialogue response content.

4. The method according to claim 3, characterized in that The determining of the dialogue response material based on the dialogue intention includes: The topic material preference corresponding to the user is determined, and a dialogue material association process is performed based on the dialogue intention and the topic material preference to obtain a dialogue response material.

5. The method according to claim 1, wherein The determining of the companion action response data for the in-vehicle companion virtual image based on the response emotion category and the dialogue response content includes: Determining a voice output mode and an image demonstration script of the in-vehicle companion virtual image based on the response emotion category and the dialogue response content; Determining the accompanying action data corresponding to the in-vehicle accompanying virtual image based on the image demonstration script; Generate accompanying voice data corresponding to the in-vehicle accompanying virtual image using the voice output mode based on the dialogue response content; Accompanying motion response data is generated based on the accompanying motion data and the accompanying voice data.

6. The method according to claim 5, characterized in that The step of determining the voice output mode and image demonstration script of the in-vehicle companion virtual image based on the response emotion category and the dialogue response content includes: Performing response strategy reasoning based on the response emotion category and the dialogue response content to determine a target response strategy; Performing voice feature matching based on the target response strategy to determine a voice output mode; Demonstration expression matching is performed based on the target response strategy to obtain expression design information, demonstration action matching is performed based on the target response strategy to obtain action design information, and demonstration integration processing is performed based on the expression design information and the action design information to obtain the image demonstration script of the in-vehicle companion virtual image.

7. The method according to claim 5, characterized in that The determining of the accompanying action data corresponding to the in-vehicle accompanying virtual image based on the image demonstration script includes: Generate expression demonstration data of the in-vehicle companion virtual image based on the image demonstration script, generate action demonstration data of the in-vehicle companion virtual image based on the image demonstration script, and synthesize companion action data based on the expression demonstration data and the action demonstration data.

8. A driving processing device, characterized in that: The device comprises: A data acquisition module, used to acquire the companion interaction data input by the user to the in-vehicle companion virtual image in the vehicle driving scene; A data processing module, configured to use a large vehicle-mounted interaction model to perform emotional state inference based on the companionship interaction data to obtain a response emotional category for the user and determine a dialogue response content for the user; a data generating module, configured to determine companion action response data for the in-vehicle companion virtual image based on the response emotion category and the dialogue response content; The response processing module is used to control the vehicle-mounted companion virtual image to perform user companion response processing based on the companion action response data.

9. A computer storage medium, characterized in that The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Interaction method, interaction device, electronic equipment and storage medium

    CN114357135A

  • Interaction method and device based on vehicle-mounted virtual image, vehicle and readable medium

    CN117724674A

  • Interaction method and device for emotion accompanying, vehicle and storage medium

    CN117935795A

  • Emotionally intelligent companion device

    IN201621035955A