Human-vehicle interaction method and apparatus
By combining real-time environmental information processing and large language models, personalized and proactive human-vehicle interaction is achieved, solving the problems of monotonous and mechanical interaction methods in existing technologies and enhancing the naturalness and entertainment value of the interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PATEO CONNECT (NANJING) CO LTD
- Filing Date
- 2024-11-28
- Publication Date
- 2026-05-29
Smart Images

Figure CN122116892A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to human-vehicle interaction methods and apparatus. Background Technology
[0002] Human-vehicle interaction is essentially a process of information exchange between humans and machines. Information exchange efficiency is the most important indicator for measuring human-vehicle dialogue interaction. Human-vehicle interaction will follow the evolutionary path of human-to-human interaction; dialogue interaction is the most efficient form of human-to-human interaction and will also become the most efficient form of human-vehicle dialogue interaction.
[0003] In the era of in-vehicle digitalization, with the maturity of vehicle hardware and software technologies, voice interaction technology between humans and vehicles is also developing rapidly. More complex and intelligent voice interactions are now possible between people and cars. However, current in-vehicle human-machine interaction suffers from a mechanical feel, with monotonous and repetitive voice timbre and tone, resulting in generic and formulaic responses that fail to meet the diverse service needs of different car owners in different scenarios. Furthermore, existing human-machine interaction operates on a passive wake-up mode, unable to provide drivers with comprehensive and seamless services, and unable to respond proactively to unexpected situations. Summary of the Invention
[0004] Embodiments of this disclosure present a human-vehicle interaction method and apparatus.
[0005] In a first aspect, embodiments of this disclosure provide a human-vehicle interaction method, comprising: acquiring real-time environmental information; inputting the environmental information into a pre-trained scene recognition model and outputting scene information; inputting the scene information into a pre-trained large language model and outputting a guiding statement, wherein the guiding statement is used to guide the user to perform human-vehicle interaction; and terminating the human-vehicle interaction in response to detecting the end of the scene through real-time environmental information or the user issuing an end signal.
[0006] In some embodiments, inputting the scene information into a pre-trained large language model and outputting guiding statements includes: determining the language style of human-vehicle interaction based on the user's profile; inputting the scene information and the language style into a pre-trained large language model and outputting guiding statements.
[0007] In some embodiments, the method further includes: in response to receiving negative feedback from the user, re-identifying scene information; inputting the re-identified scene information into a pre-trained large language model, and outputting guiding statements.
[0008] In some embodiments, obtaining real-time environmental information includes: capturing a user's facial image through an in-vehicle camera; and identifying the user's emotional state and the number of users based on the facial image.
[0009] In some embodiments, obtaining real-time environmental information includes: determining points of interest of a predetermined type within a predetermined range based on a map and location information; and querying activity information related to the points of interest.
[0010] In some embodiments, acquiring real-time environmental information includes: collecting in-vehicle sounds via a microphone; identifying the in-vehicle sounds to determine the user's emotional state.
[0011] In some embodiments, obtaining real-time environmental information includes: obtaining music played in the vehicle; in response to detecting that the music is not played by the in-vehicle player, identifying the music and obtaining its attribute information; in response to detecting that the music is played by the in-vehicle player, directly obtaining the music's attribute information from the in-vehicle player; and querying the artist's activities and related music culture based on the attribute information.
[0012] In some embodiments, obtaining real-time environmental information includes obtaining abnormal alarm information of the vehicle.
[0013] Secondly, embodiments of this disclosure provide a human-vehicle interaction device, comprising: an acquisition unit configured to acquire real-time environmental information; a recognition unit configured to input the environmental information into a pre-trained scene recognition model and output scene information; a guidance unit configured to input the scene information into a pre-trained large language model and output guidance statements, wherein the guidance statements are used to guide a user to perform human-vehicle interaction; and an termination unit configured to terminate human-vehicle interaction in response to detecting the end of the scene through real-time environmental information or the user issuing an termination signal.
[0014] In some embodiments, the guidance unit is further configured to: determine the language style of human-vehicle interaction based on the user's profile; input the scene information and the language style into a pre-trained large language model; and output guidance statements.
[0015] In some embodiments, the apparatus further includes a feedback unit configured to re-identify scene information in response to receiving negative feedback from a user; input the re-identified scene information into a pre-trained large language model; and output guiding statements.
[0016] In some embodiments, the acquisition unit is further configured to: acquire a user's facial image via an in-vehicle camera; and identify the user's emotional state and the number of users based on the facial image.
[0017] In some embodiments, the acquisition unit is further configured to: determine points of interest of a predetermined type within a predetermined range based on a map and location information; and query activity information related to the points of interest.
[0018] In some embodiments, the acquisition unit is further configured to: collect in-vehicle sounds via a microphone; identify the in-vehicle sounds to determine the user's emotional state.
[0019] In some embodiments, the acquisition unit is further configured to: acquire music played in the vehicle; in response to detecting that the music is not played by the in-vehicle player, identify the music and acquire the attribute information of the music; in response to detecting that the music is played by the in-vehicle player, directly acquire the attribute information of the music from the in-vehicle player; and query the artist's activities and related music culture based on the attribute information.
[0020] In some embodiments, the acquisition unit is further configured to acquire abnormal alarm information of the vehicle.
[0021] Thirdly, embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more computer programs stored thereon, wherein when the one or more computer programs are executed by the one or more processors, the one or more processors perform the method as described in any one of the first aspects.
[0022] Fourthly, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method as described in any one of the first aspects.
[0023] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, implements the method as described in any one of the first or second aspects.
[0024] The human-vehicle interaction method and apparatus provided in this disclosure achieve intelligent chat triggering by combining various scene information, thereby improving the initiative, entertainment value, and response speed of the AI assistant. It offers a rich variety of character options to meet the personalized needs of different users and create atmospheres for different scenarios. Leveraging the advantages of large models, it generates more natural and creative dialogue content.
[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0026] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0027] Figure 1This is an exemplary system architecture diagram to which one embodiment of this disclosure can be applied;
[0028] Figure 2 This is a flowchart of an embodiment of the human-vehicle interaction method according to the present disclosure;
[0029] Figure 3 This is a schematic diagram of an application scenario of the human-vehicle interaction method disclosed herein;
[0030] Figure 4 This is a flowchart of yet another embodiment of the human-vehicle interaction method according to the present disclosure;
[0031] Figure 5 This is a schematic diagram of the structure of one embodiment of the human-vehicle interaction device according to the present disclosure;
[0032] Figure 6 This is a schematic diagram of the structure of a computer system suitable for implementing embodiments of the present disclosure. Detailed Implementation
[0033] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0034] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0035] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "over," and "on top" of the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.
[0036] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0037] Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those embodiments or examples, without contradiction. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0038] Figure 1 An exemplary system architecture is shown that can be applied to embodiments of the human-vehicle interaction method or human-vehicle interaction device disclosed herein.
[0039] like Figure 1 As shown, the system architecture may include a dashcam 110, an in-vehicle camera 120, an in-vehicle display screen 130, a touchpad 140, a cloud server 150, a microphone, a speaker, and a radar unit (not shown in the figures). The system architecture may also include an AR HUD (not shown in the figures) and controllers (not shown in the figures) that establish communication connections with the in-vehicle display screen and the AR HUD, respectively. The radar unit is used to detect objects around the vehicle (vehicles, pedestrians, green belts, etc.) and measure the distance between the vehicle and the objects. Figure 1 The touchpad 140 shown is just an example of one mounted on the steering wheel; it can also be mounted on other vehicle components such as the dashboard and armrest. The windshield can be used as a projection screen for the AR HUD to display augmented reality information.
[0040] The dashcam 110 is used to record video images and sound of the entire driving process of a car.
[0041] The vehicle-mounted camera 120 can be a roof-mounted panoramic camera or cameras mounted on each side of the vehicle body. In some embodiments, the position and angle of the vehicle-mounted camera 120 can be adjusted as needed, and can be adjusted via voice commands, button commands, touch commands, etc. For example, a user can send a voice command such as "adjust the angle of the front camera upwards by 10 degrees" or "adjust the position of the front camera downwards by 1 centimeter." In other embodiments, such as a roof-mounted panoramic camera, it can acquire panoramic images around the vehicle body, and can crop images within a specific angle range from the panoramic images for image display or corner recognition according to instructions.
[0042] The touchpad 140 can be mounted on vehicle components such as the dashboard or steering wheel, allowing users to input touch commands to adjust the position of calibration points. For example, the touchpad consists of multiple piezoelectric vibrators, which can be mounted on the back of the touchpad. When pressure is applied to the surface of the touchpad, elastic waves are generated. These elastic waves are then transmitted to different piezoelectric vibrators, where corresponding elastic waveforms are picked up. These elastic waveforms have essentially the same shape, differing only in their arrival time and amplitude.
[0043] In some implementations, the touch position can be identified based on the TOF (Time of Flight) principle to obtain the touch command received by the touchpad 140: receiving voltage signals collected by multiple piezoelectric vibrators; determining a characteristic time point corresponding to each voltage signal based on at least one voltage point with similarity among the multiple voltage signals; the characteristic time point being determined based on at least one time point corresponding to the at least one voltage point; determining at least three characteristic time point pairs among the multiple characteristic time points; and determining the touch position information based on the relative time difference corresponding to the at least three characteristic time point pairs and the position information of a pair of piezoelectric vibrators corresponding to each relative time difference.
[0044] In other embodiments, the touch position can be identified and the touch command received by the touchpad 140 can be obtained by geometric calculation based on the relationship between the reciprocal of the detected voltage value and the distance: by receiving electrical signals collected by multiple piezoelectric vibrators respectively, and determining a first electrical signal point corresponding to each of the multiple electrical signals based on at least one electrical signal point with similarity among the multiple electrical signals, the first electrical signal point obtained can characterize the signal value of the touch point at the first time point collected by the piezoelectric vibrator; at the same time, by determining the proportional relationship between multiple first touch distances based on the first electrical signal point corresponding to each of the multiple electrical signals, the first touch distance is the distance between the piezoelectric vibrator and the touch point at the first time point. Based on the position information of multiple piezoelectric vibrators, the position information of the touch point at the first time point is determined. This takes into account the principle that the farther away from the touch point, the greater the attenuation of the mechanical elastic wave, the smaller the pressure sensing of the mechanical elastic wave on the piezoelectric vibrator, and the smaller the corresponding electrical signal output. By combining the position information of each piezoelectric vibrator, the touch point can be located through geometric relationships. At the same time, since the piezoelectric vibrator can be perfectly integrated with the surface material of the vehicle and has the characteristics of sun exposure resistance, it has stable performance when facing the complex usage scenarios of the vehicle. By determining the proportional relationship between the first touch distance through the first electrical signal point, and then combining the position information of the piezoelectric vibrator, the position information of the touch point can be accurately determined.
[0045] The vehicle-mounted display screen 130 can be various types of displays, such as the display screen of a DVR (Digital Video Recorder), or a central control screen, instrument panel screen, or passenger-side screen. It can also be an electronic device display screen that establishes a communication connection with the vehicle. The vehicle-mounted camera 120 and the vehicle-mounted display screen 130 can be connected via wired or wireless communication. For example, images captured by the vehicle-mounted camera 120 can be transmitted to the vehicle-mounted display screen 130 for display via WiFi, Bluetooth, or satellite imagery technology.
[0046] In some embodiments, the vehicle display 130 may be a touch screen for receiving instructions to adjust the displayed image, such as zooming the displayed image by swiping with a finger.
[0047] AR HUD is configured to project content from in-vehicle displays.
[0048] The controller is configured to receive signals sent by the user via an in-vehicle display, touchpad, microphone, or buttons.
[0049] Cloud servers can provide map and navigation data.
[0050] The vehicle camera 120 can be a 360-degree panoramic camera. It is connected to a processor at the vehicle's infotainment port, which can read and process vehicle data. Radar sensors connected to the processor are installed at both ends of the front and rear bumpers on the vehicle, and the turn signal switch wires are connected to the processor.
[0051] The vehicle-mounted camera 120 collects image data from around the vehicle, creating a 360-degree panoramic overhead view of the vehicle's surroundings. The processor reads this image data. The processor reads data from the vehicle's infotainment system via a chip, then processes the vehicle's electronic power steering data along with the vehicle's track width and wheelbase data to obtain vehicle trajectory prediction data. This trajectory prediction data is then merged with the 360-degree panoramic overhead view to produce a panoramic overhead image with predicted driving trajectory.
[0052] The vehicle camera 120 may also include a camera located inside the vehicle for capturing facial images of passengers inside the vehicle and identifying the number of passengers and their emotional state through the images.
[0053] It should be noted that the human-vehicle interaction method provided in the embodiments of this disclosure is generally executed by a controller.
[0054] Continue to refer to Figure 2 The diagram illustrates a flow 200 of an embodiment of a human-vehicle interaction method according to the present disclosure. This human-vehicle interaction method includes the following steps:
[0055] Step 201: Obtain real-time environmental information.
[0056] In this embodiment, the executing entity of the human-vehicle interaction method (e.g.) Figure 1 The controller shown can acquire real-time environmental information via wired or wireless connection. This environmental information can include both in-vehicle and external environmental information. In-vehicle environmental information may include the number of passengers, passenger mood, in-vehicle music, temperature, and humidity. External environmental information may include weather, traffic information, and surrounding POIs. For example, it can acquire real-time information such as the number of people in the vehicle, their status, the music currently playing, and POIs passed by during the vehicle's journey. By using in-vehicle sensors and cameras to accurately determine the number of people in the vehicle, their status, and by connecting to the music playback device to obtain music type information, it can identify surrounding POIs using the in-vehicle navigation system and map data, providing basic information for scene analysis.
[0057] Step 202: Input the environmental information into the pre-trained scene recognition model and output the scene information.
[0058] In this embodiment, multimodal environmental information can be used for scene recognition, such as in-vehicle voice, traffic information, and surrounding POIs. Using as much environmental information as possible can improve the accuracy of scene recognition.
[0059] Scene recognition models can be various existing neural network models created based on machine learning techniques. These neural network models can have various existing neural network architectures (e.g., DenseBox, VGGNet, ResNet, SegNet, etc.).
[0060] Scene recognition models can first extract keywords from environmental information, and then determine scene information based on these keywords. Scene information can include scene category and scene triggering conditions. For example, if a child is detected crying inside the car, "Why haven't we arrived yet?", combined with traffic information from the external environment, it can be determined that the current scene is a traffic jam. Triggering conditions refer to the timing when the AI assistant initiates a conversation. For example, when the number of people in the car increases, it is judged as an enhanced social scene, and topics suitable for multi-person interaction can be initiated; when playing a specific type of music, related music culture or artist updates can be introduced; when passing by POIs such as concert venues, concert information can be proactively introduced.
[0061] The training process of the scene recognition model is as follows:
[0062] 1. Obtain the sample set.
[0063] Sample sets can be obtained in various ways. For example, the executing entity can retrieve an existing sample set stored in a database server via a wired or wireless connection. Alternatively, a user can collect samples via a terminal. In this way, the executing entity can receive the samples collected by the terminal and store them locally, thereby generating a sample set.
[0064] Here, the sample set may include at least one sample. The sample may include environmental information and scene information labels.
[0065] 2. Select samples from the sample set.
[0066] The method and number of samples are not limited in this disclosure. For example, at least one sample may be selected randomly, or a sample with a wide variety of environmental information may be selected.
[0067] 3. Input the environmental information from the selected samples into the initial scene recognition model to obtain the predicted scene information.
[0068] The initial scene recognition model can be any existing neural network model created based on machine learning techniques. This neural network model can have various existing neural network architectures (e.g., DenseBox, VGGNet, ResNet, SegNet, etc.). The storage location of the initial scene recognition model is also not limited in this disclosure.
[0069] 4. Analyze the predicted scene information with the scene information labels in the sample to determine the loss value.
[0070] The predicted scene information and the scene information labels in the samples can be used as parameters and input into the specified loss function to calculate the loss value between the two.
[0071] In this embodiment, the loss function is typically used to estimate the degree of inconsistency between the model's predicted values (such as the predicted scene category) and the true values (such as scene information labels in the samples). It is a non-negative real-valued function. Generally, the smaller the loss function, the better the robustness of the model. The loss function can be set according to actual needs.
[0072] 5. Compare the loss value with the target value.
[0073] A target value is generally used to represent the ideal situation where the predicted value (such as predicted scene information) is inconsistent with the true value (such as scene information labels in the sample). In other words, when the loss value reaches the target value, the predicted value can be considered close to or approximately equal to the true value. The target value can be set according to actual needs.
[0074] It should be noted that if multiple (at least two) samples are selected, the executing entity can compare the loss value of each sample with the target value separately. This allows it to determine whether the loss value of each sample has reached the target value.
[0075] 6. Determine whether the initial model has been trained based on the comparison results.
[0076] In this embodiment, based on the comparison results, the executing entity can determine whether the initial model training is complete. For example, if multiple samples are selected, the executing entity can determine that the initial model training is complete if the loss value of each sample reaches the target value. Alternatively, the executing entity can calculate the proportion of samples whose loss values reach the target value out of the selected samples. If this proportion reaches a preset sample proportion (e.g., 95%), the initial model training can be determined to be complete.
[0077] In this embodiment, if the executing entity determines that the initial model has been successfully trained, it can use the initial model as the scene recognition model. If the executing entity determines that the initial model has not been successfully trained, it can adjust the relevant parameters in the initial model. For example, it can use backpropagation to modify the weights in each convolutional layer of the initial model. It can also reselect samples from the sample set. Thus, the above training steps can continue.
[0078] It should be noted that the selection method is not limited in this disclosure. For example, if there are a large number of samples in the sample set, the executing entity can select samples that have not been selected before.
[0079] Step 203: Input the scene information into the pre-trained large language model and output the guiding statement.
[0080] In this embodiment, the timing for the AI assistant to initiate a chat is determined based on scene information. For example, when the number of people in the car increases, it is determined that the social scene is enhanced, and topics suitable for multi-person interaction can be initiated; when playing a specific type of music, related music culture or singer updates are introduced; when passing by POIs such as concert venues, concert information is proactively introduced.
[0081] Based on the triggering timing and the selected persona, generate engaging dialogue content. This content can include questions, suggestions, anecdotes, etc., aiming to stimulate user interest and engagement. Dialogue can be generated using existing large language models.
[0082] Large Language Models (LLMs) are deep learning models trained on large amounts of text data that can generate natural language text or understand the meaning of language text. LLMs can handle various natural language tasks, such as text classification, question answering, and dialogue, and are an important pathway to artificial intelligence. Currently, LLMs employ a Transformer architecture and pre-training objectives (such as Language Modeling) similar to small models, differing only in increased model size, training data, and computational resources.
[0083] The scene information is input into the large language model as a prompt. The large language model will actively initiate human-vehicle interaction dialogue through guiding statements. The guiding statements are used to guide the user to interact with the vehicle.
[0084] For example, if a novice driver is driving normally on the highway and suddenly experiences an abnormal drop in tire pressure, the vehicle's AI assistant accurately identifies the situation, reassures the user, provides solutions, and guides the user through the problem.
[0085] The vehicle was driving normally on the highway when the tire pressure dropped abnormally (the vehicle monitoring system monitors the vehicle's tire pressure).
[0086] When the vehicle's dashboard monitoring lights illuminate, the vehicle's intelligent driving assistant automatically wakes up (based on the scenario level, it automatically determines the wake-up timing and actively triggers it).
[0087] "Car owner, we have detected an abnormally low tire pressure. Please remain calm and do not panic. Please follow my instructions." (Based on user profile and scenario-specific guidance to soothe the car owner's emotions)
[0088] ...
[0089] "I have already turned on the hazard warning lights. Now please gently release the accelerator pedal to gradually reduce the vehicle speed." (Analyze the scenario, propose a solution, and guide the driver to follow the instructions.)
[0090] ...
[0091] "Well done. There is an emergency lane 500 meters ahead. Please observe the surrounding vehicles and carefully move your vehicle to the emergency lane."
[0092] ...
[0093] "OK, we have safely reached the emergency lane. Please get out of your car and place a warning triangle 150 meters behind you. This is the nearest roadside assistance information: xxxxxxxx. You can contact them and wait patiently in your car." (Analyze the scenario, propose a solution based on the driver's characteristics, and guide the driver to follow the instructions.)
[0094] ...
[0095] Step 204: In response to the detection of the end of the scene through real-time environmental information or the user issuing an end signal, the human-vehicle interaction ends.
[0096] In this embodiment, the current environmental information is detected in real time using a scene recognition model. If the environmental information changes, the scene recognition result will change accordingly, indicating the end of the scene. Users can also actively initiate a command to end the scene. This can be an explicit command, such as "End conversation," or a hint. The system then uses semantic understanding to determine if the user's intention is to end the conversation. For example, if a user says "You're so noisy," it means the user wants to end the human-vehicle interaction. When ending the human-vehicle interaction, a voice prompt can be output to inform the user. For example, in a traffic jam scenario, the changed traffic information is input into the scene recognition model, and no more traffic jam information is obtained. Therefore, the traffic jam scenario ends, and the vehicle system can output, "The road ahead is clear. That's enough for today. Fasten your seatbelts; we're about to continue," to prompt the user to end the traffic jam scenario.
[0097] The method provided in the above embodiments of this disclosure combines various scenario information to intelligently trigger chat, improving the initiative, entertainment value, and response speed of the AI assistant. It offers a rich variety of character options to meet the personalized needs of different users and create atmospheres for different scenarios. Leveraging the advantages of large models, it generates more natural and creative dialogue content.
[0098] In some optional implementations of this embodiment, the step of inputting the scene information into a pre-trained large language model and outputting guiding statements includes: determining the language style of human-vehicle interaction based on the user's profile; inputting the scene information and the language style into the pre-trained large language model and outputting guiding statements. Multiple different personas are provided, such as humorous, gentle and cute, and knowledgeable. Personas are automatically switched according to user settings or scenarios, with each persona having a unique tone of voice, personality traits, and knowledge reserves. Appropriate personas can be matched based on the user's age, gender, etc. For example, an 8-year-old girl can be spoken to using an animated character's voice. Users can also be asked beforehand what personas they prefer, allowing for the use of a corresponding tone of voice. The dialogue content should also consider age; for example, simple and easy-to-understand vocabulary should be used when speaking with children.
[0099] The following example illustrates how to generate dialogue based on character settings in an entertainment scenario.
[0100] For example, in a recreational scenario, a family drives to camping on the weekend. On the way to the campsite, a major traffic jam caused by an accident ahead leaves them stuck, unable to move forward or backward. The children and adults in the car gradually become anxious.
[0101] Scene recognition information elements:
[0102] The navigation system showed a traffic jam.
[0103] The system detected repeated use of phrases such as "How much longer until we can leave?" and "Why can't we set off yet?"; facial expressions such as furrowed brows and downcast eyes; and actions such as repeatedly looking out the window and getting on and off vehicles.
[0104] The AI assistant said, "This traffic jam is so annoying! It looks like it'll be a while before it clears. How about we play a game together?"
[0105] (Waiting for the passengers inside the vehicle to react)
[0106] The little girl's attention was drawn, and her gaze was fixed on the smart assistant.
[0107] (Detecting the interests of the people in the car, continue the conversation)
[0108] "It seems Duoduo is very interested. Let me explain the rules of the game. I will display the silhouette of an object on the screen. You can ask me questions to guess what it is. If you guess correctly, you can collect it and I will reward you with a small accessory that you can use to dress up Princess Xixi! Let's work together for the beautiful Xixi!"
[0109] ...She started playing the game, but Duoduo encountered difficulties.
[0110] "Duoduo, let's invite Mom and Dad to join us! It'll be even more fun that way!"
[0111] (By engaging everyone in the car, their mood gradually improved.)
[0112] ...
[0113] (The road ahead is gradually becoming clear)
[0114] "The road ahead is clear now, let's call it a day! Duoduo, fasten your seatbelts, we're about to set off again. The weather's perfect, our camping trip today is sure to be great!"
[0115] In some optional implementations of this embodiment, the method further includes: in response to receiving negative feedback from the user, re-identifying the scene information; inputting the re-identified scene information into a pre-trained large language model, and outputting guiding statements.
[0116] Users can input feedback via voice or the central control screen. If the AI assistant does not provide feedback within the predetermined time, it is considered that the user has implicitly agreed, which is positive feedback. Users can also explicitly express positive feedback. For example, if the AI assistant says, "Would you like to listen to xx's song?", and the user says "Yes," it is positive feedback. If the user says, "Change it," it is negative feedback.
[0117] By inputting current environmental information and negative feedback information into a pre-trained scene recognition model, scene information can be re-identified, thereby improving the accuracy of scene recognition.
[0118] In some optional implementations of this embodiment, obtaining real-time environmental information includes: capturing the user's facial image through an in-vehicle camera; and identifying the user's emotional state and the number of users based on the facial image.
[0119] A pre-trained facial recognition model can be used to identify a user's emotional state and the number of users in a vehicle by analyzing facial images. This model, a neural network, first segments faces in an image. The number of bounding boxes determines the number of faces in the image, i.e., the number of people in the vehicle. Then, for each segmented face, facial expressions are identified. The model detects key facial features such as eyes, eyebrows, and mouth. The position of these key features, such as the distance between eyebrows and the direction of the corners of the mouth, can be used to determine facial expressions. Facial expressions can then be analyzed to discern emotions, such as happiness or annoyance. Based on the user's emotional state and the number of users, the current scenario can be determined, allowing an AI assistant to generate appropriate guidance statements and provide human-computer interaction based on user preferences, thereby increasing user interest and satisfaction.
[0120] The emotions detected by facial recognition can also be combined with other environmental information to generate more accurate emotions. For example, the user's voice can be recognized and combined with the user's facial expressions to further determine the user's emotions.
[0121] For example, if the system detects repeated phrases like "How much longer until we can leave?" or "Why can't we set off yet?"; facial expressions such as furrowed brows and downcast eyes; and actions like repeatedly looking out the window or getting on and off vehicles, the AI assistant will say: "This traffic jam is really annoying! It looks like it'll be a while before it clears. How about we play a game together?"
[0122] In some optional implementations of this embodiment, obtaining real-time environmental information includes: determining points of interest (POIs) of a predetermined type within a predetermined range based on maps and location information; and querying activity information related to the POIs. Vehicles can obtain location information in real time, which, combined with maps, allows them to find POIs that nearby users might be interested in, such as stadiums hosting concerts or newly opened shopping malls. Users can query the official websites, public accounts, or review apps of the POIs to obtain information on events held at the venues corresponding to the POIs, and filter content that users might be interested in, which can then be read aloud by an AI assistant. User interests can be determined through historical behavior, such as frequently visited locations or locations frequently searched. Large language models can also answer various user questions, such as the concert schedule of a certain celebrity. User interests can be determined through user query records. A pre-trained interest recognition model can be used, inputting user behavior information into the model to identify the user's interest categories. For example, liking a certain celebrity or liking hot pot.
[0123] In some optional implementations of this embodiment, obtaining real-time environmental information includes: collecting in-vehicle sounds through a microphone; identifying the in-vehicle sounds to determine the user's emotional state.
[0124] The sounds inside the car can include the user's voice and music. The sound collected by the microphone can be input into a pre-trained speech recognition model to identify the text information. This text information is then input into a pre-trained semantic understanding model to understand the user's intent. For example, if a user says "Why aren't we there yet?", after text recognition and language understanding, the model can interpret the user's intent as complaining about the car's slow speed, being late, and feeling anxious.
[0125] It can also perform voiceprint recognition to identify the user's identity, thereby determining the number of people making the voice and performing semantic understanding of the dialogue, rather than mistakenly believing it to be the voice of the same person.
[0126] In addition, it can identify timbre information such as sound intensity and frequency, thereby more accurately judging the user's emotions. For example, if a user shouts hysterically, it indicates an emotional breakdown.
[0127] The user's emotional state affects the prompting statements generated by the large language model, which then uses these prompting statements to soothe the user's emotions.
[0128] In some optional implementations of this embodiment, obtaining real-time environmental information includes: obtaining music played in the vehicle; in response to detecting that the music is not played by the in-vehicle player, identifying the music and obtaining its attribute information; in response to detecting that the music is played by the in-vehicle player, directly obtaining the music's attribute information from the in-vehicle player; and querying the artist's activities and related music culture based on the attribute information.
[0129] The system can capture in-car audio via microphone to determine whether it's user voice or music. Music can originate from two sources: the car's built-in player or other mobile devices, such as the user's phone or tablet. Since the car's infotainment system (the controller mentioned earlier) communicates with the player, it can read the attribute information of the music playing. Based on this information, it can then query artist information and related music culture. For example, if artist A's song is playing, it can query artist A's recent concert information and upcoming movie releases. If folk music from country X is playing, it can introduce the music culture of country X. The system can also engage in multi-round dialogues with the user, encouraging them to learn more about related music culture.
[0130] The controller can also collect music played by other players through the microphone, perform voice recognition, obtain lyrics, and then search for song titles, singers, and other attribute information based on the lyrics. It can then query singer updates and related music culture knowledge based on the attribute information.
[0131] In some optional implementations of this embodiment, obtaining real-time environmental information includes: obtaining abnormal alarm information from the vehicle. Vehicle sensors (e.g., lidar, accelerometer, gyroscope, etc.) can report vehicle driving data to the controller. The controller can detect whether the data is within the normal range; if it is not within the normal range, it can output alarm information. For example, if abnormal tire pressure is detected, an alarm is output on the central control screen to instruct the user to stop and check. For example, if the air conditioning is not cooling, an alarm is output on the central control screen to instruct the user to drive the car to a 4S shop for repair. The controller can use the abnormal alarm information as environmental information, along with environmental information from other sources, to identify scene information.
[0132] See also Figure 3 , Figure 3 This is a schematic diagram illustrating an application scenario of the human-vehicle interaction method according to this embodiment. Figure 3In this application scenario, a vehicle plays a song by artist XX while driving. Using location and map information, it detects that a concert by artist XX is being held at a stadium ahead. The song title and the stadium name are input into a pre-trained scene recognition model, which outputs a concert recommendation scenario. A large language model generates recommendation information and proactively announces via voice, "A concert by artist XX is being held at the stadium to your right..." Users can interact with the AI assistant to learn more about the concert, such as where the next show is being held and who the guest performers are. Once the vehicle has moved a certain distance from the stadium, the concert recommendation scenario ends, concluding the human-vehicle interaction.
[0133] Further reference Figure 4 This illustrates a flow 400 of another embodiment of the human-vehicle interaction method. Flow 400 of this human-vehicle interaction method includes the following steps:
[0134] Step 401: Obtain real-time environmental information.
[0135] Step 402: Input the environmental information into the pre-trained scene recognition model and output the scene information.
[0136] Step 403: Input the scene information into the pre-trained large language model and output guiding statements, which are used to guide the user to interact with the vehicle.
[0137] Step 404: In response to the detection of the end of the scene through real-time environmental information or the user issuing an end signal, the human-vehicle interaction ends.
[0138] Steps 401-404 are basically the same as steps 201-204, so they will not be described again.
[0139] Step 405: In response to receiving negative feedback from the user, re-identify the scene information.
[0140] In this embodiment, the user can input feedback information via voice or the central control screen. If the user does not respond within a predetermined time after the AI assistant finishes speaking, it is considered that the user has implicitly agreed, i.e., positive feedback. The user can also explicitly express positive feedback; for example, if the AI assistant says, "Would you like to listen to xx's song?", and the user says "Yes," it represents positive feedback; if the user says "Change it," it represents negative feedback. The current environmental information and negative feedback information are input into a pre-trained scene recognition model to re-recognize the scene information. This can improve the accuracy of scene recognition.
[0141] Step 406: Input the re-identified scene information into the pre-trained large language model and output guiding statements.
[0142] In this embodiment, the timing for the AI assistant to initiate a chat is determined based on scene information. For example, when the number of people in the car increases, it is determined that the social scene is enhanced, and topics suitable for multi-person interaction can be initiated; when playing a specific type of music, related music culture or singer updates are introduced; when passing by POIs such as concert venues, concert information is proactively introduced.
[0143] Based on the triggering timing and the selected persona, generate engaging dialogue content. This content can include questions, suggestions, anecdotes, etc., aiming to stimulate user interest and engagement. Dialogue can be generated using existing large language models.
[0144] Further reference Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a human-vehicle interaction device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0145] like Figure 5 As shown, the human-vehicle interaction device 500 of this embodiment includes: an acquisition unit 501, a recognition unit 502, a guidance unit 503, and an termination unit 504. The acquisition unit 501 is configured to acquire real-time environmental information; the recognition unit 502 is configured to input the environmental information into a pre-trained scene recognition model and output scene information; the guidance unit 503 is configured to input the scene information into a pre-trained large language model and output guidance statements, wherein the guidance statements are used to guide the user in human-vehicle interaction; and the termination unit 504 is configured to terminate the human-vehicle interaction in response to the detection of scene termination through real-time environmental information or the user issuing an termination signal.
[0146] In this embodiment, the specific processing of the acquisition unit 501, recognition unit 502, guidance unit 503, and termination unit 504 of the human-vehicle interaction device 500 can be referred to Figure 2 The corresponding steps are 201, 202, 203, and 204 in the embodiment.
[0147] In some optional implementations of this embodiment, the guidance unit 503 is further configured to: determine the language style of human-vehicle interaction based on the user's profile; input the scene information and the language style into a pre-trained large language model, and output guidance statements.
[0148] In some optional implementations of this embodiment, the device 500 further includes a feedback unit (not shown in the figures), configured to re-identify scene information in response to receiving negative feedback information from the user; input the re-identified scene information into a pre-trained large language model; and output guiding statements.
[0149] In some optional implementations of this embodiment, the acquisition unit 501 is further configured to: acquire the user's facial image through the vehicle-mounted camera; and identify the user's emotional state and the number of users based on the facial image.
[0150] In some optional implementations of this embodiment, the acquisition unit 501 is further configured to: determine points of interest of a predetermined type within a predetermined range based on the map and location information; and query activity information related to the points of interest.
[0151] In some optional implementations of this embodiment, the acquisition unit 501 is further configured to: collect in-vehicle sounds through a microphone; identify the in-vehicle sounds to determine the user's emotional state.
[0152] In some optional implementations of this embodiment, the acquisition unit 501 is further configured to: acquire music played in the vehicle; in response to detecting that the music is not played by the in-vehicle player, identify the music and acquire the attribute information of the music; in response to detecting that the music is played by the in-vehicle player, directly acquire the attribute information of the music from the in-vehicle player; and query the singer's activities and related music culture based on the attribute information.
[0153] In some optional implementations of this embodiment, the acquisition unit 501 is further configured to: acquire abnormal alarm information of the vehicle.
[0154] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in this disclosed technical solution all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.
[0155] According to embodiments of this disclosure, this disclosure also provides an electronic device and a readable storage medium.
[0156] An electronic device includes: one or more processors; and a storage device having one or more computer programs stored thereon, which, when executed by the one or more processors, cause the one or more processors to implement the method described in process 200 or 400.
[0157] A computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in process 200 or 400.
[0158] Figure 6A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0159] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0160] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0161] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as road planning methods. For example, in some embodiments, the road planning method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the road planning method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the road planning method by any other suitable means (e.g., by means of firmware).
[0162] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0163] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0164] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0165] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0166] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0167] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be servers in distributed systems or servers incorporating blockchain technology. Servers can also be cloud servers, or intelligent cloud computing servers or intelligent cloud hosts with artificial intelligence technology.
[0168] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0169] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A human-vehicle interaction method, comprising: Obtain real-time environmental information; The environmental information is input into a pre-trained scene recognition model, which outputs scene information. The scene information is input into a pre-trained large language model, which outputs guiding statements, which are used to guide users to interact with the vehicle. The human-vehicle interaction ends when the scene ends or the user issues an end signal based on real-time environmental information.
2. The method according to claim 1, wherein, The step of inputting the scene information into a pre-trained large language model and outputting guiding statements includes: Determine the language style for human-vehicle interaction based on the user profile; The scene information and the language style are input into a pre-trained large language model, which outputs guiding statements.
3. The method according to claim 1, wherein, The method further includes: In response to receiving negative feedback from the user, the scene information is re-identified; The re-identified scene information is input into a pre-trained large language model, which outputs guiding statements.
4. The method according to claim 1, wherein, The acquisition of real-time environmental information includes: The user's facial image is captured using an in-vehicle camera; The system identifies the user's emotional state and the number of users based on the facial images.
5. The method according to claim 1, wherein, The acquisition of real-time environmental information includes: Determine the points of interest of the specified type within the specified area based on the map and location information; Query activity information related to the points of interest.
6. The method according to claim 1, wherein, The acquisition of real-time environmental information includes: The sound inside the car is collected using a microphone; The system identifies the sounds inside the vehicle to determine the user's emotional state.
7. The method according to claim 1, wherein, The acquisition of real-time environmental information includes: Get the music playing in the car; In response to detecting that the music is not played by the car player, the music is identified and its attribute information is obtained; In response to the detection that the music is music played by a car player, the attribute information of the music is obtained directly from the car player; Based on the attribute information, you can query the singer's activities and related music culture.
8. The method according to claim 1, wherein, The acquisition of real-time environmental information includes: Obtain abnormal alarm information from the vehicle.
9. A human-vehicle interaction device, comprising: The acquisition unit is configured to acquire real-time environmental information; The recognition unit is configured to input the environmental information into a pre-trained scene recognition model and output scene information; The guidance unit is configured to input the scene information into a pre-trained large language model and output guidance statements, wherein the guidance statements are used to guide the user to perform human-vehicle interaction. The termination unit is configured to terminate human-vehicle interaction in response to the detection of scene termination through real-time environmental information or the user issuing a termination signal.
10. An electronic device, comprising: One or more processors; Storage device, on which one or more computer programs are stored, When the one or more computer programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-8.
11. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.