Large model-based interaction method, vehicle terminal, electronic device, and storage medium

By detecting when a person is seated stably in a vehicle and collecting images and attribute features, and interacting with a large cloud model, the problem of poor accuracy of emotional value information in existing technologies is solved, achieving more efficient and user-friendly emotional value information interaction.

CN122450286APending Publication Date: 2026-07-24BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2026-03-10
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing in-vehicle terminals and smart terminals mainly rely on fixed templates or simple semantic recognition plus fixed scripts to provide emotional value information, resulting in poor accuracy of emotional value information.

Method used

By installing pressure sensors and cameras in vehicles, once a person is seated and stable, images and attribute characteristics of the person are collected and interacted with a large cloud model to provide interactive information carrying emotional value based on the person's information.

Benefits of technology

It improves the accuracy and efficiency of interaction, enhances the user experience, and provides more natural, credible, and objective emotional value information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122450286A_ABST
    Figure CN122450286A_ABST
Patent Text Reader

Abstract

The present disclosure discloses a large model-based interaction method, a vehicle terminal, an electronic device and a storage medium, specifically relates to the fields of intelligent interaction and large model and the like artificial intelligence technology, and can be applied to the scenes of vehicle voice interaction, intelligent cockpit, intelligent terminal and the like. The specific implementation scheme is as follows: determining that a person in a vehicle is stably seated; obtaining person information of the person who is seated in the vehicle; and based on the person information, interacting with a cloud end to provide, by a cloud end large model of the cloud end, interaction information carrying emotional value information for the person in the vehicle based on the person information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, specifically to the fields of intelligent interaction and artificial intelligence technology such as large models, and can be applied to scenarios such as in-vehicle voice interaction, smart cockpit, and smart terminal; in particular, it relates to an interaction method based on a large model, an in-vehicle terminal, an electronic device, and a storage medium. Background Technology

[0002] With the development of artificial intelligence (AI) technology, existing in-vehicle terminals or other smart terminals are becoming increasingly intelligent in human-computer interaction. For example, existing in-vehicle terminals or other smart terminals can be configured with interactions that provide emotional value information, such as praise or greetings.

[0003] Existing in-vehicle terminals or other smart terminals mainly achieve the interaction of emotional value information such as praise or greetings through fixed template broadcasts or simple semantic recognition plus fixed script feedback. Summary of the Invention

[0004] This disclosure provides an interaction method based on a large model, an in-vehicle terminal, an electronic device, and a storage medium.

[0005] According to one aspect of this disclosure, a large-model-based interaction method is provided for use in an in-vehicle terminal, including:

[0006] Ensure that the people inside the vehicle are seated and stable;

[0007] Obtain information about the people seated in the vehicle;

[0008] Based on the person's information, the system interacts with the cloud, so that the cloud-based big data model provides interactive information carrying emotional value information to the person in the vehicle based on the person's information.

[0009] According to another aspect of this disclosure, a vehicle-mounted terminal is provided, comprising:

[0010] The determination module is used to determine the stability of people sitting inside the vehicle;

[0011] The acquisition module is used to acquire information about the people sitting in the vehicle.

[0012] The interaction module is used to interact with the cloud based on the person information, so that the cloud big model provides interactive information carrying emotional value information to the person in the vehicle based on the person information.

[0013] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described above and any possible implementations.

[0017] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described above and any possible implementation thereof.

[0018] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the aspects and any possible implementations described above.

[0019] According to the technology disclosed herein, it is possible to provide more accurate emotional value information for people in a vehicle during interaction, which can effectively improve the accuracy and efficiency of interaction and enhance the user experience.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0021] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0022] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;

[0023] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;

[0024] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure;

[0025] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure;

[0026] Figure 5 This is a schematic diagram according to the fifth embodiment of the present disclosure;

[0027] Figure 6 This is a schematic diagram according to the sixth embodiment of the present disclosure;

[0028] Figure 7This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation

[0029] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0030] Obviously, the described embodiments are only some, not all, of the embodiments disclosed herein. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0031] It should be noted that the terminal devices involved in the embodiments of this disclosure may include, but are not limited to, smart devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, and tablet computers; the display devices may include, but are not limited to, personal computers, televisions, and other devices with display functions.

[0032] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0033] Existing in-vehicle terminals primarily rely on fixed templates or scripts to deliver emotional value information such as praise or greetings, resulting in poor accuracy of the provided emotional value information.

[0034] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure; as shown Figure 1 As shown, this embodiment provides an interaction method based on a large model, applied in an in-vehicle terminal, which may specifically include the following steps:

[0035] S101. Ensure that the people inside the vehicle are seated stably;

[0036] In practical applications, interaction will only be triggered after the people inside the vehicle are seated and stable. In this embodiment, determining that the people inside the vehicle are seated and stable can be used as a prerequisite for interaction.

[0037] In this embodiment, "stable seating" means that the people in the driver's seat and all other seats in the vehicle are in a stable state after they have sat down.

[0038] S102. Obtain the information of the people sitting in the vehicle;

[0039] The information about the people in the vehicle can include various aspects such as their appearance, clothing, and expressions. This information can reflect the state of the people in the vehicle to a certain extent.

[0040] S103. Based on the person information, interact with the cloud, so that the cloud-based big model provides interactive information carrying emotional value information to the people in the vehicle based on the person information.

[0041] In this embodiment, interaction with the cloud can be based on the information of people in the vehicle. In this way, the large model in the cloud can refer to the information of people in the vehicle to provide more accurate emotional value information for the people in the vehicle, thereby realizing interaction with the people in the vehicle that carries emotional value information.

[0042] The emotional value information in this embodiment may include positive and uplifting words such as praise, appreciation, and / or encouragement, which can promote a pleasant mood.

[0043] The large-model-based interaction method in this embodiment, after determining that the people inside the vehicle are seated and stable, can interact with the cloud based on the information of the people inside the vehicle. The cloud-based large model provides interactive information carrying emotional value information to the people inside the vehicle based on the information of the people inside the vehicle. In this scheme, the cloud-based large model can provide more accurate emotional value information to the people inside the vehicle during the interaction, which can effectively improve the accuracy and efficiency of the interaction and enhance the user experience.

[0044] Figure 2 This is a schematic diagram according to the second embodiment of this disclosure; the interaction method based on a large model in this embodiment, in the above... Figure 1 The technical solutions of the illustrated embodiments will be further described in more detail below. For example... Figure 2 As shown, the interaction method based on a large model in this embodiment may specifically include the following steps:

[0045] S201. Check and confirm that the vehicle doors are closed;

[0046] Specifically, the in-vehicle terminal can detect whether the current door is open or closed. If it is open, subsequent interactive operations cannot be triggered. If it is closed, further steps can be executed.

[0047] S202. Detect and determine that the value of the pressure sensor of the seat in the vehicle is greater than or equal to a first preset threshold, and the fluctuation value of the pressure sensor is less than a second preset threshold within a preset time period.

[0048] In this embodiment, the first preset threshold is greater than the second preset threshold.

[0049] Specifically, a pressure sensor is installed under each seat in the vehicle. When a person sits down, the pressure sensor registers a certain value. For example, a first preset threshold can be configured based on experience. If the value of the pressure sensor at a particular seat exceeds the first preset threshold, it can be determined that someone is sitting in that seat. Based on this, the pressure sensor at each seat in the vehicle can be monitored to see if its value is greater than or equal to the first preset threshold. If so, it can be determined that someone is sitting in the corresponding seat; otherwise, it can be determined that no one is sitting in the corresponding seat.

[0050] Furthermore, for any given seat, within a short period after a person sits down, they may adjust their seat or posture. During this time, the pressure sensor readings for that seat will fluctuate significantly, indicating an unstable seating state. Once the person is settled, the pressure sensor readings will no longer show significant fluctuations. Based on this, the fluctuation values ​​of the pressure sensors on seats occupied within a preset time period can be detected to see if they are less than a second preset threshold. If so, the seating is considered stable; otherwise, it is determined that the seating is not stable.

[0051] In practical scenarios, interactions are not triggered immediately when people in a vehicle are not yet firmly seated. Therefore, in this embodiment, setting a condition for interaction as a stable seated person in the vehicle can improve the accuracy of the interaction.

[0052] The preset time length, the first preset threshold, and the second preset threshold in this embodiment can all be set based on experience, and are not limited here.

[0053] In this embodiment, by detecting and determining that the car door is closed, the value of the pressure sensor of the seat inside the vehicle is greater than or equal to a first preset threshold, and the fluctuation value of the pressure sensor is less than a second preset threshold within a preset time period, it is determined that a person has sat down in the vehicle and is sitting stably, which can effectively improve the detection accuracy and efficiency of sitting stability.

[0054] S203. Control the camera to capture images of people sitting in the vehicle;

[0055] In the vehicle of this embodiment, a corresponding camera and voice collector can be installed on each seat. The camera can use the image of the person in the corresponding seat, and the voice collector can collect the voice commands of the person in the corresponding seat.

[0056] The seat in this embodiment can be implemented using a smart cockpit, which includes a camera, a voice acquisition device, and a pressure sensor, etc., to achieve all the functions described above in this embodiment.

[0057] Specifically, based on step S202 above, the seat position of the person can be determined, and then the camera at the corresponding seat position can be controlled to capture an image of the person in that seat. For example, the captured image of the person can include their appearance, clothing, and expression. Among them, clothing can include clothes and jewelry.

[0058] For example, if it is determined that someone is sitting in the driver's seat, the camera in the driver's seat is controlled to capture an image of the person in the driver's seat;

[0059] If it is determined that a passenger is seated in any of the multiple passenger seats, the camera at the corresponding passenger seat will be controlled to capture an image of the passenger.

[0060] In summary, once the people in the vehicle are detected to be seated and stable, images of each seated person in the vehicle can be accurately and comprehensively captured.

[0061] S204. Store the collected images of seated people according to their seating positions.

[0062] In this embodiment, taking the collected person information, including images of people, as an example, the image of each seated person is stored according to their corresponding seat position, that is, a mapping relationship between the person image and the corresponding seat position is stored. For example, driver's seat 01 → the image of the person in the driver's seat; passenger seat 02 → the image of the person in the passenger seat; passenger seat 03 → the image of the person in the passenger seat, and so on. The person images stored in this way are very convenient for subsequent retrieval of images of people in any seat position.

[0063] It should be noted that steps S201-S204 are for the accurate operation of the large model-based interaction in this embodiment. In practical applications, after storing and caching the images of the seated people according to their seat positions in step S204, it is also necessary to perform real-time detection on each seat so that it can be updated in a timely manner when changes occur.

[0064] For example, if a vehicle door is detected to be open, it is necessary to further check whether the seating signal for each seat is true. This involves checking whether the pressure sensor value for each seat is greater than or equal to a first preset threshold. If so, the seating signal for the corresponding seat is determined to be true; otherwise, it is determined to be false. If, after the door opens, the seating signal for a certain seat changes from true to false, the stored image of the person sitting at that location needs to be deleted. If, at this time, someone sits down in that seat, and the seating signal changes from false to true again, the stored image of the newly seated person needs to be added back to the stored location.

[0065] If the seat occupancy signal changes from false to true after the car door opens, it is also necessary to store the image of the person at that location.

[0066] In summary, every time the door opens or closes, it is necessary to detect the seating signal of each seat in a timely manner, and update the stored image of the person in the corresponding seat in a timely manner when there is a change. This can effectively improve the accuracy and efficiency of subsequent interactions between the people in each seat and the vehicle terminal.

[0067] S205. Obtain the image of the driver;

[0068] S206. Send an image of the driver to the cloud so that the cloud-based large model can extract the driver's attribute features based on the image. The driver's attribute features include at least one of the driver's appearance, clothing, age, and emotional features. Based on the driver's attribute features, generate first interaction information. The first interaction information carries emotional value information.

[0069] S207, Receive the first interactive information returned from the cloud;

[0070] S208. Using a speech synthesis module, the first interactive speech is generated based on the first interactive information;

[0071] S209. Play the first interactive voice to provide the driver in the vehicle with an interaction carrying emotional value information.

[0072] In this embodiment, by employing steps S205-S209, the in-vehicle terminal can proactively initiate an interaction providing emotional value information. This proactive interaction is primarily based on the information of the driver in the vehicle. In this embodiment, the vehicle-side information includes the driver's image. Specifically, on the cloud side, a large cloud model extracts the driver's attribute features based on the driver's image. For example, this may include at least one of the driver's appearance, clothing, age, and emotional characteristics. The driver's appearance, clothing, age, and emotional characteristics can all serve as supporting information for emotional value information, providing the driver with more natural, credible, and objective praise, appreciation, and / or encouragement.

[0073] For example, using the above-described scheme in this embodiment, the cloud-based big data model generates first interactive information based on the attribute characteristics of the driver, which may include: "This young lady in the driver's seat looks so stylish in her glasses and black T-shirt! She manages to make even a casual look look so sophisticated; she really has a great sense of style." In this way, after the driver is seated and settled, the system proactively provides the driver with such emotionally resonant praise, effectively improving the interaction efficiency between the driver and the in-vehicle terminal and increasing user satisfaction.

[0074] The large-model-based interaction method in this embodiment, by adopting the above-mentioned technical solution, can timely, accurately, and comprehensively acquire images of all people in the vehicle after they are seated and stable. Based on the image of the driver, it actively interacts with the cloud, so that the cloud-based large model can extract the attribute features of the driver based on the image of the driver, and then generate second interactive information carrying emotional value information based on the attribute features of the driver. This can provide the driver with more natural, more credible, and more objective emotional value information, effectively improving the accuracy and efficiency of the interaction and enhancing the user experience.

[0075] Figure 3 This is a schematic diagram according to the third embodiment of this disclosure; the interaction method based on a large model in this embodiment, in the above... Figure 1 The technical solutions of the illustrated embodiments will be further described in more detail below. For example... Figure 3 As shown, the interaction method based on a large model in this embodiment may specifically include the following steps:

[0076] S301. Check and confirm that the vehicle doors are closed;

[0077] S302. Detect and determine that the value of the pressure sensor of the seat in the vehicle is greater than or equal to a first preset threshold, and the fluctuation value of the pressure sensor is less than a second preset threshold within a preset time period.

[0078] S303, Control the camera to capture images of people sitting in the vehicle;

[0079] S304. Using a large vehicle-side model, extract the attribute features of the person based on the image of the person. The attribute features of the person include at least one of the following: appearance features, clothing features, age features, and emotional features.

[0080] Among these, physical features can include the facial features and other external characteristics of the person in the image.

[0081] Clothing features can include at least one feature such as the color and style of the clothing worn by the person in the image, and at least one feature such as accessories such as earrings and glasses. For example, the more vibrant the color and the more exquisite the style of the clothing, and the more exquisite the earrings worn, the better the person's mood and the better their condition; while the more monotonous the color and the simpler the style of the clothing, and the less earrings worn, the worse the person's mood and the worse their condition.

[0082] Age features can include age predictions of a person based on their facial features.

[0083] Emotional characteristics can include the emotional features represented by the facial expressions of people in an image. Specifically, the emotions of a person can be inferred based on the state of their eyebrows, eyes, and mouth. For example, if the eyebrows and eyes are curved and the corners of the mouth are turned up, the person's emotional characteristic can be inferred to be happy; if the eyebrows and eyes are furrowed and the corners of the mouth are turned down, the person's emotional characteristic can be inferred to be sad; and so on.

[0084] The aforementioned physical features, clothing features, age features, and emotional features can be considered as multimodal features of a person. Each modality of these multimodal features can, to some extent, characterize the person's state and be used to generate more accurate emotional value information in subsequent processing. The more multimodal features included in the image-based perception of a person's attributes, the more accurate and relevant the generated emotional value information will be.

[0085] In this embodiment, by using a large model on the vehicle side to accurately and effectively extract the attribute features of the person based on the image of the person, the vehicle side can be used.

[0086] S305. Store the collected attribute characteristics of seated people according to their seat positions.

[0087] With the above Figure 2 The difference between steps S201-S204 in the illustrated embodiment is that the above... Figure 2In the illustrated embodiment, the collected person information includes the person's image as an example. In this embodiment, the collected person information includes the person's attribute features as an example. Specifically, on the vehicle terminal side, a large vehicle model is also set up, which accurately and effectively extracts the person's attribute features based on the person's image.

[0088] In addition, the attribute characteristics of the seated individuals collected are stored according to their seating position, which is consistent with the above. Figure 2 In step S204 of the illustrated embodiment, the method of storing the image of the person according to seat position is the same. For example, driver's seat 01 → attribute features of the driver; passenger seat 02 → attribute features of the passenger; passenger seat 03 → attribute features of the passenger, and so on. Storing the attribute features of the person in each seat position in this way makes it very convenient to obtain the attribute features of the person in any seat position later.

[0089] S306. Obtain the attribute characteristics of the main driver character;

[0090] S307. Send the attribute characteristics of the main driver to the cloud so that the cloud-based large model can generate second interaction information based on the attribute characteristics of the main driver; the second interaction information carries emotional value information.

[0091] S308, Receive the second interactive information returned from the cloud;

[0092] S309. A speech synthesis module is used to generate second interactive speech based on the second interactive information;

[0093] S310, Play the second interactive voice to provide the driver in the vehicle with an interaction carrying emotional value information.

[0094] In this embodiment, steps S306-S310 are the same as those described above. Figure 2 The difference between steps S205-S209 in the illustrated embodiment is that... Figure 2 In the embodiment shown, when the vehicle terminal interacts with the cloud, it sends an image of the driver to the cloud. The cloud-based big model extracts the driver's attribute features based on the image of the driver, and then generates the first interaction information based on the driver's attribute features.

[0095] In this embodiment, on the vehicle terminal side, a large vehicle model is used to extract the attribute features of the person based on the image of the person. When the vehicle terminal interacts with the cloud, the attribute features of the driver are sent directly to the cloud. In this way, the large cloud model does not need to perform any feature extraction and can directly generate the second interaction information based on the attribute features of the driver.

[0096] The large-model-based interaction method in this embodiment, by adopting the above-mentioned technical solution, can also provide the active driver with more natural, credible, and objective emotional value information, effectively improving the accuracy and efficiency of the interaction and enhancing the user experience.

[0097] Figure 4 This is a schematic diagram according to the fourth embodiment of this disclosure; the interaction method based on a large model in this embodiment, in the above... Figure 1 The technical solutions of the illustrated embodiments will be further described in more detail below. For example... Figure 4 As shown, the interaction method based on a large model in this embodiment may specifically include the following steps:

[0098] S401-S405, and the above Figure 3 The steps S301-S305 are the same;

[0099] S406. When a person in the vehicle issues a voice command, collect the voice command.

[0100] Specifically, each seat in the vehicle is equipped with a corresponding voice acquisition device to collect the voice commands of the person in that seat.

[0101] S407. Determine the target location of the person initiating the voice command in the vehicle; the target location is the driver's seat or any of the multiple passenger seats.

[0102] Since there is a one-to-one mapping between the position of each seat in the vehicle and the voice acquisition device in that seat, when a voice command is detected, the corresponding voice acquisition device collects the corresponding voice command. The vehicle terminal can determine the identifier of the voice acquisition device that collected the voice command, and then determine the corresponding seat position based on the identifier of the voice acquisition device, which is the target position of the person who initiated the voice command in the vehicle.

[0103] S408. Based on the target location, obtain the stored attribute characteristics of the target person;

[0104] Specifically, based on the attribute features of the person stored according to seat location in the previous step S405, the target person information corresponding to the target location can be obtained, that is, the attribute features of the target person. Specifically, the attribute features of the target person may include at least one of the target person's appearance features, clothing features, age features, and emotional features. Specifically, the more modal features the target person's attribute features include, the more accurate the subsequently generated emotional value information will be.

[0105] S409. Send voice commands and the target person's attribute characteristics to the cloud so that the cloud-based large model can generate third interactive information in response to the voice commands based on the voice commands and the target person's attribute characteristics; the third interactive information carries emotional value information.

[0106] Optionally, before sending this step, the voice command can be recognized as text, and then the text voice command and the target person's attribute characteristics can be sent to the cloud. The cloud big model can then generate a third interactive message that responds to the voice command and carries emotional value information based on the content of the voice command and the target person's attribute characteristics.

[0107] S410, Receive third interactive information returned from the cloud;

[0108] S411. A speech synthesis module is used to generate third-interactive speech based on third-interactive information;

[0109] S412. Play the third interactive voice message to provide the target person with interactive information carrying emotional value information.

[0110] It should be noted that in this embodiment, step S409 is a specific implementation of sending voice commands and target person information to the cloud so that the cloud-based large model can generate third-party interactive information based on the voice commands and target person information. In this implementation, the target person information is taken as the attribute features of the target person. At this time, the person information stored according to the seat position is taken as an example of storing the attribute features of the person.

[0111] In practical applications, optionally, the person information stored according to seat location can also be stored as in step S204, storing the person's image. At this point, a voice command and target person information are sent to the cloud so that the cloud-based large model can generate third interactive information based on the voice command and target person information. Specifically, this can be done by: sending a voice command and an image of the target person to the cloud so that the cloud-based large model can extract the target person's attribute features based on the image. The target person's attribute features include at least one of the target person's appearance, clothing, age, and emotional characteristics; and generating third interactive information in response to the voice command based on the target person's attribute features and the voice command. This third interactive information carries emotional value information.

[0112] With the above Figure 2 and Figure 3 The difference between the embodiments shown is that the above-described embodiments are... Figure 2 and Figure 3 The illustrated embodiment is an active interaction initiated by the in-vehicle terminal based on the driver's information. This embodiment, however, is a passive interaction initiated by the in-vehicle terminal based on voice commands from people within the vehicle. Although this embodiment differs from the one described above... Figure 2 and Figure 3 In the embodiments shown, the mechanisms for initiating the interaction are different, but the interaction carries emotional value information that can be generated based on the person's information.

[0113] The large-model-based interaction method in this embodiment, by adopting the above-mentioned technical solution, can provide more natural, more credible, and more objective emotional value information while responding to voice commands during interaction, effectively improving the accuracy and efficiency of interaction and enhancing user experience.

[0114] The large-model-based interaction method in this embodiment, by adopting the above-mentioned technical solution, can accurately perceive the attribute characteristics of a person. Through the interaction between the vehicle terminal and the cloud, the large-model in the cloud generates interactive messages that respond to voice commands in a timely and accurate manner based on the attribute characteristics and voice commands of the person. Moreover, the interactive messages carry natural and credible emotional value information, which can solve the technical problem of poor accuracy when emotional value information is generated by using fixed templates or fixed language in the prior art. This effectively improves the accuracy and efficiency of human-computer interaction, and can also significantly enhance the user experience and user satisfaction.

[0115] Figure 5 This is a schematic diagram according to the fifth embodiment of this disclosure; as shown Figure 5 As shown, this embodiment provides a vehicle-mounted terminal 500, including:

[0116] Module 501 is used to determine whether the person inside the vehicle is seated stably.

[0117] The acquisition module 502 is used to acquire information about the people sitting in the vehicle.

[0118] The interaction module 503 is used to interact with the cloud based on the person information, so that the cloud big model in the cloud can provide interactive information carrying emotional value information to the person in the vehicle based on the person information.

[0119] The vehicle terminal 500 in this embodiment achieves the same implementation principle and technical effect of interaction based on a large model by adopting the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.

[0120] Figure 6 This is a schematic diagram according to the sixth embodiment of this disclosure; as shown Figure 6 As shown, the vehicle terminal 600 in this embodiment, in the above-mentioned... Figure 5 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be further described in more detail. For example... Figure 6 As shown, the vehicle-mounted terminal 600 in this embodiment includes the above-mentioned... Figure 5The modules with the same name and function shown are: determination module 601, acquisition module 602, and interaction module 603.

[0121] The determining module 601 is used for:

[0122] Detect and confirm that the vehicle doors are closed;

[0123] The system detects and determines that the value of the pressure sensor of the seat in the vehicle is greater than or equal to a first preset threshold, and the fluctuation value of the pressure sensor is less than a second preset threshold within a preset time period; the first preset threshold is greater than the second preset threshold.

[0124] Further optional, such as Figure 6 As shown, in one embodiment of this disclosure, the acquisition module 602 includes:

[0125] The image acquisition unit 6021 is used to control the camera to acquire images of the people sitting in the vehicle.

[0126] Further optionally, in one embodiment of this disclosure, the image acquisition unit 6021 is used for:

[0127] Once it is confirmed that someone is seated in the driver's seat, the camera in the driver's seat is controlled to capture an image of the person in the driver's seat.

[0128] Further optionally, in one embodiment of this disclosure, the interaction module 603 is used for:

[0129] The image of the driver is sent to the cloud so that the cloud-based big data model can extract the driver's attribute features based on the image. The driver's attribute features include at least one of the driver's appearance, clothing, age, and emotional features. Based on the driver's attribute features, first interactive information is generated. The first interactive information carries emotional value information.

[0130] Receive the first interactive information returned by the cloud;

[0131] A speech synthesis module is used to generate the first interactive speech based on the first interactive information;

[0132] Play the first interactive voice message to provide the driver in the vehicle with an interaction carrying emotional value information.

[0133] Further optional, such as Figure 6 As shown, in one embodiment of this disclosure, the acquisition module 602 further includes:

[0134] The feature extraction unit 6022 is used to extract the attribute features of the person based on the image of the person using a large vehicle-side model. The attribute features of the person include at least one of appearance features, clothing features, age features, and emotional features.

[0135] Further optionally, in one embodiment of this disclosure, the image acquisition unit 6021 is also used for:

[0136] If it is determined that a passenger is seated in any of the multiple passenger seats, the camera at the corresponding passenger seat is controlled to capture an image of the passenger in that passenger seat.

[0137] Further optional, such as Figure 6 As shown, in one embodiment of this disclosure, the vehicle terminal 600 further includes:

[0138] Storage module 604 is used to store the collected information about the person according to their seat position.

[0139] Further optionally, in one embodiment of this disclosure, the interaction module 603 is used for:

[0140] The attribute features of the person are sent to the cloud so that the cloud-based big model can generate second interactive information based on the attribute features of the person; the second interactive information carries emotional value information.

[0141] Receive the second interactive information returned by the cloud;

[0142] A speech synthesis module is used to generate a second interactive voice based on the second interactive information;

[0143] Play the second interactive voice to provide the person in the vehicle with an interaction carrying emotional value information.

[0144] Further optional, such as Figure 6 As shown, in one embodiment of this disclosure, the vehicle terminal 600 further includes:

[0145] The voice acquisition module 605 is used to acquire the voice command when it detects that a person in the vehicle is issuing a voice command.

[0146] The determining module 601 is further configured to determine the target location of the person initiating the voice command in the vehicle; the target location is the driver's seat or any one of the multiple passenger seats.

[0147] The acquisition module 602 is also used to acquire the stored target person information based on the target location.

[0148] Further optionally, in one embodiment of this disclosure, the interaction module 603 is used for:

[0149] The voice command and the target person information are sent to the cloud so that the cloud-based big data model can generate third interactive information in response to the voice command based on the voice command and the target person information; the third interactive information carries emotional value information.

[0150] Receive the third interactive information returned by the cloud;

[0151] A speech synthesis module is used to generate third interactive speech based on the third interactive information;

[0152] Play the third interactive voice message to provide the target person with interactive information carrying emotional value information.

[0153] The vehicle terminal 600 in this embodiment achieves the same implementation principle and technical effect of interaction based on a large model by adopting the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.

[0154] The acquisition, storage, and application of any type of information, such as user personal information, involved in the technical solutions disclosed herein comply with relevant laws and regulations and do not violate public order and good morals.

[0155] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0156] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0157] like Figure 7As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded into random access memory (RAM) 703 from storage unit 708. The RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0158] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0159] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the methods of this disclosure. For example, in some embodiments, the methods of this disclosure may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the methods of this disclosure described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the methods of this disclosure by any other suitable means (e.g., by means of firmware).

[0160] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0161] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0163] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0164] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0165] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0166] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0167] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A large-model-based interaction method applied in an in-vehicle terminal, comprising: Ensure that the people inside the vehicle are seated and stable; Obtain information about the people seated in the vehicle; Based on the person's information, the system interacts with the cloud, so that the cloud-based big data model provides interactive information carrying emotional value information to the person in the vehicle based on the person's information.

2. The method according to claim 1, wherein, Ensure that the people inside the vehicle are seated and stable, including: Detect and confirm that the vehicle doors are closed; The system detects and determines that the value of the pressure sensor of the seat in the vehicle is greater than or equal to a first preset threshold, and the fluctuation value of the pressure sensor is less than a second preset threshold within a preset time period; the first preset threshold is greater than the second preset threshold.

3. The method according to claim 1, wherein, Obtain information about the people seated in the vehicle, including: The camera is controlled to capture images of the people seated in the vehicle.

4. The method according to claim 3, wherein, The camera is controlled to capture images of people seated in the vehicle, including: Once it is confirmed that someone is seated in the driver's seat, the camera in the driver's seat is controlled to capture an image of the person in the driver's seat.

5. The method according to claim 4, wherein, Based on the person information, interaction is performed with the cloud, so that the cloud-based big data model provides interactive information carrying emotional value information to the person in the vehicle based on the person information, including: The image of the driver is sent to the cloud so that the cloud-based big data model can extract the driver's attribute features based on the image. The driver's attribute features include at least one of the driver's appearance, clothing, age, and emotional features. Based on the driver's attribute features, first interactive information is generated. The first interactive information carries emotional value information. Receive the first interactive information returned by the cloud; A speech synthesis module is used to generate the first interactive speech based on the first interactive information; Play the first interactive voice message to provide the driver in the vehicle with an interaction carrying emotional value information.

6. The method according to claim 3, wherein, Obtaining information about the people seated in the vehicle also includes: Using a large vehicle-side model, the attribute features of the person are extracted based on the image of the person. The attribute features of the person include at least one of the following: appearance features, clothing features, age features, and emotional features.

7. The method according to claim 4 or 6, wherein, Controlling the camera to capture images of people seated in the vehicle also includes: If it is determined that a passenger is seated in any of the multiple passenger seats, the camera at the corresponding passenger seat is controlled to capture an image of the passenger in that passenger seat.

8. The method according to claim 6, wherein, The method further includes: The collected information about the individuals is stored according to their seating position.

9. The method according to claim 6, wherein, Based on the person information, interaction is performed with the cloud, so that the cloud-based big data model provides interactive information carrying emotional value information to the person in the vehicle based on the person information, including: The attribute features of the person are sent to the cloud so that the cloud-based big model can generate second interactive information based on the attribute features of the person; the second interactive information carries emotional value information. Receive the second interactive information returned by the cloud; A speech synthesis module is used to generate a second interactive voice based on the second interactive information; Play the second interactive voice to provide the person in the vehicle with an interaction carrying emotional value information.

10. The method according to claim 8, wherein, Before interacting with the cloud based on the aforementioned person information, and before the cloud-based big data model provides interactive information carrying emotional value information to the person in the vehicle based on the aforementioned person information, the process further includes: When a person is detected issuing a voice command in the vehicle, the voice command is collected; Determine the target location of the person initiating the voice command within the vehicle; the target location is either the driver's seat or any of the multiple passenger seats. Based on the target location, obtain the stored information about the target person.

11. The method according to claim 10, wherein, Based on the person information, interaction is performed with the cloud, so that the cloud-based big data model provides interactive information carrying emotional value information to the person in the vehicle based on the person information, including: The voice command and the target person information are sent to the cloud so that the cloud-based big data model can generate third interactive information in response to the voice command based on the voice command and the target person information; the third interactive information carries emotional value information. Receive the third interactive information returned by the cloud; A speech synthesis module is used to generate third interactive speech based on the third interactive information; Play the third interactive voice message to provide the target person with interactive information carrying emotional value information.

12. A vehicle-mounted terminal, comprising: The determination module is used to determine the stability of people sitting inside the vehicle; The acquisition module is used to acquire information about the people sitting in the vehicle. The interaction module is used to interact with the cloud based on the person information, so that the cloud big model provides interactive information carrying emotional value information to the person in the vehicle based on the person information.

13. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-11.

14. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.

15. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-11.