Human-computer interaction method, device, related equipment and computer program product
By obtaining the user's line of sight information in the on-board driving assistance system, adjusting the camera's viewing angle and calling a large multimodal model, the problem of inaccurate driver intention recognition in the existing system is solved, more efficient interaction and environmental perception are achieved, and driving safety and user experience are improved.
Patent Information
- Application Number
- CN202411881172.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-19
AI Technical Summary
Existing in-vehicle driving assistance systems lack effective information fusion and interaction understanding mechanisms, resulting in inaccurate recognition of driver intentions. In particular, it is difficult to understand the driver's gaze direction and interaction intentions in complex traffic scenarios, which limits environmental perception and interactive feedback capabilities.
By obtaining the interaction questions of the target user in the car, using the in-car camera to determine the line of sight information, adjusting the viewing angle of the on-board camera to obtain the image outside the car, and calling the multimodal large model to generate the response result in combination with the image outside the car, we can achieve accurate understanding and response to the user's intention.
It improves the efficiency and quality of interaction, enhances the system's perception of the external environment, provides a more personalized and intuitive interactive experience, and improves driving safety and system intelligence.
Smart Images

Figure CN119428733B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a human-computer interaction method, apparatus, related equipment, and computer program product. Background Art
[0002] With the development of intelligent transportation systems, in-vehicle driving assistance systems have improved driving and comfort by integrating sensors, cameras and advanced algorithms, making in-vehicle driving assistance systems play an increasingly important role.
[0003] However, existing in-vehicle driver assistance systems primarily rely on onboard DMS (Driver Monitoring System) cameras to monitor the driver's status and fixed external cameras to perceive the external environment. These systems typically operate independently and lack effective information fusion and interaction understanding mechanisms, resulting in inaccurate recognition of driver intent. Especially in complex traffic scenarios, in-vehicle driver assistance systems struggle to accurately understand the driver's gaze direction and interaction intentions, limiting their environmental perception and interactive feedback capabilities. Summary of the Invention
[0004] In view of the above problems, this application is proposed to provide a human-computer interaction method, device, equipment and storage medium to enhance the interactive capabilities of the vehicle-mounted driving assistance system. The specific solution is as follows:
[0005] In a first aspect, a human-computer interaction method is provided, comprising:
[0006] Obtain interactive questions raised by target users in the car;
[0007] Obtaining an image of the target user captured by a first camera in the vehicle, and determining sight direction information of the target user based on the image;
[0008] Adjusting the viewing angle of a second camera on the vehicle according to the sight direction information so that the viewing angle of the second camera is consistent with the sight direction of the target user, and acquiring an image outside the vehicle captured by the second camera;
[0009] The multimodal large model is called to instruct the multimodal large model to generate a response result corresponding to the interactive question in combination with the external vehicle image, and output the response result.
[0010] In one possible design, in another implementation of the first aspect of the embodiments of the present application, calling the multimodal large model to instruct the multimodal large model to generate a response result corresponding to the interactive question based on the external vehicle image and output the response result includes:
[0011] Obtaining a first prompt instruction prompt format template, the first prompt instruction prompt format template including a first task instruction, an interactive question slot, and an external vehicle image slot, the first task instruction being used to instruct the multimodal large model to combine the external vehicle image in the external vehicle image slot to generate the response result corresponding to the interactive question in the interactive question slot;
[0012] Fill the interactive question into the interactive question slot, fill the outside-vehicle image into the outside-vehicle image slot, obtain a first prompt instruction prompt, input the first prompt instruction prompt into the pre-trained multimodal large model, and obtain the reply result output by the multimodal large model.
[0013] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the target user is a driver. Before calling the multimodal large model, the method further includes:
[0014] Obtaining a task perception result of at least one configured active safety detection task, where the task perception result of the active safety detection task is determined based on the driver's state information, external environment information, and / or vehicle driving state information;
[0015] The multimodal large model is called to instruct the multimodal large model to generate a response result corresponding to the interactive question in combination with the image outside the vehicle, including:
[0016] The multimodal large model is called to instruct the multimodal large model to comprehensively consider the task perception result and the image outside the vehicle to generate the response result corresponding to the interactive question.
[0017] In one possible design, in another implementation of the first aspect of the embodiments of the present application, calling the multimodal large model to instruct the multimodal large model to comprehensively consider the task perception result and the external vehicle image to generate the response result corresponding to the interactive question includes:
[0018] Obtain a second prompt instruction prompt format template, where the second prompt instruction prompt format template includes a second task instruction, an interaction question slot, a task perception result slot, and an external vehicle image slot, where the second task instruction is used to instruct the multimodal large model to comprehensively consider the external vehicle image in the external vehicle image slot and the task perception result in the task perception result slot to generate the response result corresponding to the interaction question in the interaction question slot;
[0019] Fill the interactive question into the interactive question slot, fill the outside-vehicle image into the outside-vehicle image slot, fill the task perception result into the task perception result slot, obtain a second prompt instruction prompt, input the second prompt instruction prompt into the pre-trained multimodal large model, and obtain the reply result output by the multimodal large model.
[0020] In one possible design, in another implementation of the first aspect of the embodiments of the present application, before acquiring the image outside the vehicle captured by the second camera, the method further includes:
[0021] The interaction question is parsed, and if distance information can be extracted from the interaction question, a shooting focal length is determined according to the distance information, and the focus of the second camera is adjusted to the shooting focal length.
[0022] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the process of determining the sight direction information of the target user based on the image includes:
[0023] Locating the eye area of the target user in the image;
[0024] Obtaining the spatial position of the human eye origin and the eye movement direction in the eye area;
[0025] The sight line direction information of the target person is determined according to the spatial position of the human eye origin and the eye movement direction.
[0026] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the line of sight direction information is obtained based on a first coordinate system where the first camera is located;
[0027] The process of adjusting the viewing angle of the second vehicle-mounted camera according to the line of sight direction information includes:
[0028] Using a coordinate system mapping relationship between the first camera and the second camera, the sight line direction information is converted into a second coordinate system where the second camera is located to obtain target sight line direction information;
[0029] The rotation angle of the second camera is calculated according to the target sight direction information, and the rotation of the second camera is controlled according to the rotation angle.
[0030] In a second aspect, a human-computer interaction device is provided, comprising:
[0031] A question acquisition unit, used to acquire interactive questions raised by the target user in the car;
[0032] an image acquisition unit, configured to acquire an image of the target user captured by a first camera in the vehicle, and determine the sight direction information of the target user based on the image;
[0033] a viewing angle determining unit, configured to adjust the viewing angle of a second onboard camera according to the sight line direction information so that the viewing angle of the second camera is consistent with the sight line direction of the target user, and to obtain an image outside the vehicle captured by the second camera;
[0034] The reply output unit is used to call the multimodal large model to instruct the multimodal large model to generate a reply result corresponding to the interactive question in combination with the image outside the vehicle, and output the reply result.
[0035] In a third aspect, an electronic device is provided, comprising: a memory and a processor;
[0036] The memory is used to store programs;
[0037] The processor is used to execute the program to implement each step of the human-computer interaction method described in any one of the first aspects of the present application.
[0038] In a fourth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the various steps of the human-computer interaction method described in any one of the first aspects of the present application are implemented.
[0039] In a fifth aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the various steps of the human-computer interaction method described in any one of the aforementioned first aspects of the present application.
[0040] By leveraging the above technical solution, the present application proposes a human-computer interaction method that obtains interactive questions raised by a target user in the vehicle, uses a first camera in the vehicle to capture an image of the target user, and determines the target user's gaze direction information based on the image, thereby better understanding the user's intentions and needs and providing a more personalized and intuitive interactive experience. The first camera is used to determine the target user's gaze direction information, and the perspective of the second camera is adjusted accordingly. The second camera then captures the image outside the vehicle, allowing the system to see the object or scene of interest to the user, enhancing the system's perception of the external environment. Furthermore, by invoking a large multimodal model and combining the external image with the user's interactive questions, deeper information processing and understanding can be achieved. By understanding the user's interactive questions and generating corresponding responses based on the external image, the responses are more accurate and relevant, improving the efficiency and quality of the interaction. The present application's human-computer interaction method accurately captures and responds to user intentions, improving the naturalness, accuracy, and efficiency of the interaction, while also enhancing driving safety and the intelligence of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0042] Figure 1 A schematic diagram of the system architecture of a human-computer interaction method provided in an embodiment of the present application;
[0043] Figure 2 A flowchart of a human-computer interaction method provided in an embodiment of the present application;
[0044] Figure 3 A flowchart of another human-computer interaction method provided in an embodiment of the present application;
[0045] Figure 4 A schematic diagram of the working principle of a human-computer interaction system provided in an embodiment of the present application;
[0046] Figure 5 A schematic diagram of the working principle of another human-computer interaction system provided in an embodiment of the present application;
[0047] Figure 6 A schematic diagram of the structure of a human-computer interaction device provided in an embodiment of the present application;
[0048] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] Before introducing this application plan, let me first explain the English terms involved in this article:
[0050] DMS cameras, short for Driver Monitoring System cameras, are in-vehicle camera systems specifically designed to monitor driver behavior and status. The primary purpose of DMS cameras is to improve road safety by monitoring the driver's physiological and behavioral status in real time, ensuring they remain alert and focused behind the wheel.
[0051] Prompt: Instructions. When interacting with an AI (such as an artificial intelligence model), you need to send instructions to the AI. This can be a text description, such as "Please recommend me a pop song" when interacting with the AI, or it can be a parameter description in a specific format, such as describing the relevant drawing parameters to ask the AI to draw a drawing according to a certain format.
[0052] Multimodal large models: Multimodal large models are an advanced AI technology that can process and understand data from different sources, such as text, images, sound, and video. By integrating these different types of information, the model can more comprehensively understand the context and improve the performance of tasks such as recognition, classification, prediction, and generation. These models typically use complex neural network architectures, such as Transformers, which are pre-trained on large-scale multimodal datasets and then fine-tuned for specific tasks to achieve more accurate understanding and more natural interactions. Multimodal large models have a wide range of applications, from autonomous driving to smart assistants, from content recommendation to health diagnosis, and they all play an important role in improving robustness, enhancing learning, and interactive learning.
[0053] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0054] See also Figure 1 , Figure 1 This is a schematic diagram of the system architecture of a human-computer interaction method provided in an embodiment of this application. The system structure described in this embodiment is a human-computer interaction system that integrates a first camera 101, a second camera 102, an onboard processor (CPU) 103, and a control module 104, aiming to improve driving safety and the interactivity of the assistance system. The following describes each part:
[0055] The first camera 101 can be a vehicle-mounted DMS camera. The first camera can be installed inside the vehicle and is responsible for real-time monitoring of the real-time image of the target user, which can be not only the driver but also other passengers in the vehicle.
[0056] Second camera 102: This can be a rotatable camera. As a physical device, the second camera can be installed inside or outside the vehicle. The second camera can adjust its viewing angle based on the target user's gaze direction and interaction intent, focusing near the target user's gaze point to track the target user's attention and capture an image.
[0057] The vehicle-mounted processor (CPU) 103: as the brain of the system, the vehicle-mounted processor is responsible for running algorithms, which can process and integrate data from the first camera and the second camera, and analyze the target user's line of sight direction and intention. The vehicle-mounted processor 103 can also include a multi-modal large model, so as to process and analyze the out-of-vehicle images captured by the second camera and the questions raised by the target user, and generate the reply results corresponding to the questions.
[0058] The control module 104: according to the analysis result of the vehicle-mounted processor 103, the rotation and zoom of the second camera are controlled to ensure that the second camera can track the target user's line of sight direction or the object of interest.
[0059] One possible implementation is that the multi-modal large model can not be configured in the vehicle-mounted processor (CPU) 103, but placed in the remote server 105. In this case, the system architecture can also include a server 105, and the processing of the interactive task is realized through the communication between the vehicle-mounted processor and the server. The vehicle-mounted processor transmits the collected data to the server, and the server calls the multi-modal large model to realize the processing of the interactive task by using its powerful processing capability. In order to enable the system to utilize more powerful computing resources for data processing and analysis, it can also provide scalability, allowing the system to dynamically adjust resources as needed, and can more easily perform software updates and maintenance.
[0060] The present application provides a human-computer interaction method, which acquires the interactive question raised by the target user in the vehicle, acquires the image of the target user by using the first camera in the vehicle, and determines the line of sight direction information of the target user according to the image, so as to better understand the intention and demand of the user. The line of sight direction information of the target user is determined by using the first camera, and the angle of view of the second camera is adjusted accordingly, and the out-of-vehicle image captured by the second camera is acquired, so that the system can see the object or scene concerned by the user, and through calling the multi-modal large model, the out-of-vehicle image and the interactive question of the user can be combined to perform deeper information processing and understanding. By understanding the interactive question of the user, the corresponding reply result is generated according to the out-of-vehicle image, so that the reply is more accurate and relevant, and the efficiency and quality of the interaction are improved.
[0061] Next, refer to Figure 2 , Figure 2 A flowchart of a human-computer interaction method provided by an embodiment of the present application is provided, which is applied to a computer device, which can be specifically a vehicle-mounted processor 103 in Figure 1 or a system composed of the vehicle-mounted processor 103 and the server 105. As described below, it specifically includes the following steps:
[0062] Step S100: Obtain interactive questions raised by the target user in the car.
[0063] Specifically, the vehicle's voice recognition module can actively capture the target user's voice commands or questions in the vehicle. This voice data can be converted into text information for further processing and analysis. The voice data can also be kept in its state and processed and analyzed. The following is an example of this process:
[0064] Questions asked by the target user in the car, for example, when the target user sees an attractive building outside the car and curiously asks "What is this building?", this question can be captured by the in-car microphone.
[0065] Interactive questions are not limited to specific types. They can be inquiries about the external environment or knowledge questions. For example, assuming that the target user wants to know the current speed of the vehicle or the external temperature, he or she may ask: "How fast are we going now?" or "What is the temperature outside?" The present application can capture these questions through the on-board microphone and combine them with the vehicle's sensor data through the on-board processor to provide corresponding speed or temperature information.
[0066] During these interactions, the design of this application ensures flexible response to the diverse interaction needs of the target users in the vehicle, providing timely and accurate feedback and enhancing the driving experience. This allows both the driver and other passengers to communicate with the vehicle system through natural language to obtain the information they need or perform specific operations, which not only improves driving convenience but also provides passengers with a richer interactive experience.
[0067] Step S110: Acquire an image of the target user captured by a first camera in the vehicle, and determine the sight direction information of the target user based on the image.
[0068] Specifically, the first camera can be a vehicle-mounted DMS camera. This application uses the first camera to capture facial images of the target user in the car, especially detailed images of the eye area, and can extract the target user's line of sight direction information from these images.
[0069] This application can analyze the gaze direction of target users in the vehicle to identify their interests or needs, thereby providing more thoughtful services. For example, if a passenger's gaze lingers on a landmark outside the vehicle window, relevant information about that landmark can be proactively provided, enhancing the passenger's riding experience. In this way, this application can monitor and understand the behavior of target users in the vehicle, achieving more comprehensive in-vehicle interaction and environmental awareness.
[0070] Step S120: Adjust the viewing angle of the second vehicle-mounted camera according to the sight direction information so that the viewing angle of the second camera is consistent with the sight direction of the target user, and obtain an image outside the vehicle captured by the second camera.
[0071] Specifically, the present application uses a first camera to monitor the gaze direction of a target user in the vehicle, whether the driver or another passenger. Once the target user's gaze direction is determined using an eye tracking algorithm, the control module adjusts the viewing angle of a second camera to match the target user's gaze direction, ensuring that the second camera's viewing angle is consistent with the target user's gaze direction in order to capture the corresponding image outside the vehicle.
[0072] The second camera can be a rotatable camera, which offers a high degree of flexibility and configurability. The second camera's design allows it to be installed either inside or outside the vehicle. This flexibility enables the second camera to be optimally configured based on actual application requirements and vehicle design.
[0073] When a second camera is placed inside the vehicle, such as a dashcam, and its lens faces outward, it can capture the user's gaze direction based on the first camera (the onboard DMS camera) to identify the target object of interest. Alternatively, when the second camera is mounted outside the vehicle, it can serve as an external sensor, enhancing the vehicle's environmental awareness.
[0074] Whether inside or outside the car, the second camera can be linked with the first camera and adjust the viewing angle according to the driver's or passenger's line of sight information to obtain images of the outside environment. This is crucial for improving driving safety and providing richer outside-car information. For example, if the target user is interested in a landmark or event outside the car, the second camera can rotate and focus outside the car to capture relevant images in real time. This makes it possible to quickly respond to the target user's line of sight and request and provide timely visual information feedback. Regardless of whether the second camera is installed inside or outside the car, the present application can be controlled by the on-board processor and control module to ensure that the camera's viewing angle is consistent with the target user's line of sight and obtain images outside the car when needed, which not only improves the response speed and accuracy, but also enhances the user experience, allowing the vehicle to interact more intelligently with the target user in the car.
[0075] One possible implementation involves keeping the second camera in a standby state during normal driving, continuously tracking the target user's gaze without actually capturing images. This design reduces unnecessary data processing and storage requirements while maintaining real-time monitoring of the target user's gaze. Only when the voice recognition module captures a question from the target user is the second camera activated, capturing the most relevant exterior imagery at the critical moment.
[0076] For example, if a target user asks, "What's that building?" the second camera is activated immediately upon receiving the question, capturing an image of the building in the target user's line of sight. This image is taken at the moment the question is asked or a short time later. This design makes the second camera more efficient and purposeful, capturing images only when the target user needs to obtain external information. This not only meets user needs but also optimizes resource usage.
[0077] Another possible implementation involves the second camera continuously tracking and capturing images in the direction of the user's gaze during normal driving. These images can be stored only as previews or cached, without further processing. However, if a user asks a specific question, such as "What does that sign mean?" or "What's wrong with the car in front?", through voice recognition or other interactive methods, the second camera is triggered to capture and process the target image within a preset timeframe. This image can be the instantaneous target image captured at the moment the user asks the question. Furthermore, given that the process of asking a question is a continuous time period, from the start of the question to the end of the question, multiple images within that timeframe can be captured. Furthermore, to account for the potential delay between the user seeing the target and asking the question, this embodiment can also extend the timeframe from the time the user begins asking the question by a set amount of time, capturing multiple images from the extended start time to the end of the question. This ensures that the captured images are most relevant to the user's question. This design makes the second camera more efficient, responding to the user's gaze and question in real time, ensuring that the captured images are most relevant to the user's question.
[0078] In another possible implementation, the angle of view of the second camera can remain unchanged in the regular state until the user raises a question. Once the user raises a question, the second camera will be activated quickly to obtain the visual direction information of the target user and adjust the angle of view of the second camera to accurately align with the line-of-sight direction of the target user, and start shooting the out-of-vehicle image. For example, if the driver notices traffic congestion ahead during driving and asks, “What's going on ahead?” After receiving this request, the second camera will be immediately directed to rotate to the corresponding direction to capture real-time images of the road ahead. In this way, the second camera can provide a detailed view of the scene that the user is interested in, helping the driver better understand the situation. This design makes the use of the second camera more efficient and on-demand, and only when the user actually raises a question or needs it, the angle of view will be adjusted and shooting will be performed, thereby ensuring reasonable allocation of resources and timely response to the user's needs.
[0079] In step S130, a multimodal large model is called to instruct the multimodal large model to generate a reply result corresponding to the interactive question in combination with the out-of-vehicle image, and output the reply result.
[0080] Specifically, after obtaining the out-of-vehicle image, the multimodal large model is further used to process and understand the images and the interactive questions related thereto. When the second camera captures the out-of-vehicle image and the user raises a specific interactive question, the multimodal large model is called. The multimodal large model can combine the out-of-vehicle image and the user's voice instruction or question to conduct in-depth analysis and understanding, and generate a reply result corresponding to the question. For example, if the user asks, “What does this sign mean?” after seeing an unfamiliar traffic sign, the image of the sign and the question will be input into the multimodal large model.
[0081] After the corresponding reply result is generated, the vehicle-mounted processor outputs the information to the user. This can include the explanation of the sign, the relevant traffic rules, or the direct answer to the user's question. The vehicle-mounted processor can display the information through the vehicle-mounted display system or directly tell the user through the voice interaction interface, ensuring that the user can obtain the required answer in the most intuitive and convenient way.
[0082] The above method of the present application not only can capture and understand the out-of-vehicle environment, but also can provide timely and accurate feedback, enhancing the interactivity and practicality of the driving assistance system. By calling the multimodal large model, the present application can process complex interactive questions and provide more rich and in-depth information for the user, improving the intelligent level of the driving experience.
[0083] This embodiment provides a human-computer interaction method that obtains interactive questions posed by a target user in a vehicle, captures an image of the target user using a first camera inside the vehicle, and determines the target user's gaze direction information based on the image, thereby better understanding the user's intentions and needs. The first camera is used to determine the target user's gaze direction information, and the perspective of the second camera is adjusted accordingly. The second camera captures an image outside the vehicle, allowing the user to see the object or scene of interest. By invoking a large multimodal model, combining the image outside the vehicle with the user's interactive questions, deeper information processing and understanding can be achieved. By understanding the user's interactive questions, a corresponding response result is generated based on the image outside the vehicle, making the response more accurate and relevant, thereby improving the efficiency and quality of interaction.
[0084] Furthermore, the process of calling the multimodal large model in step S130 in the aforementioned embodiment to instruct the multimodal large model to generate a response result corresponding to the interactive question in combination with the image outside the vehicle and output the response result can be described in detail.
[0085] First, a predefined first prompt format template is obtained. This template consists of three main parts: the first task instruction, the interaction question slot, and the external image slot. The first task instruction is a key component of the template, explicitly instructing the multimodal large model how to combine the external image in the external image slot to generate a response corresponding to the interaction question in the interaction question slot.
[0086] Filling the Interaction Question Slot: When the target user asks an interaction question, you can fill this question into the Interaction Question Slot in the template. For example, if the target user asks, "What does the traffic sign at the intersection ahead mean?" this question will be placed in the Interaction Question Slot.
[0087] Filling the vehicle exterior image slot: At the same time, the vehicle exterior image captured by the second camera is filled into the vehicle exterior image slot in the template, so that the template contains real-time visual information related to the target user's question.
[0088] Task instructions: These are instructions or instructions that direct the multimodal model to perform a specific task. They tell the multimodal model how to combine the external image and the interactive question to generate an accurate response. For example, "Please combine the image in the external image with the interactive question and output the response."
[0089] Generate the first prompt: After completing the above filling, a complete first prompt can be generated, which includes task instructions, interaction questions and corresponding external vehicle images.
[0090] Input to the multi-modal large model: input this first prompt instruction into the pre-trained multi-modal large model. The multi-modal large model will analyze and understand the content of the question according to this structured input, while referring to the image outside the vehicle, to generate an accurate reply result. Finally, the multi-modal large model outputs a reply result to the interactive question raised by the target user.
[0091] The embodiment can ensure that the multi-modal large model receives clear, comprehensive and structured input, thereby improving the accuracy and relevance of the reply result, so that the target user's question can be processed and responded to more efficiently while maintaining the smoothness and naturalness of the interaction.
[0092] Further, in some embodiments of the present application, for the case where the driver is the target user, in addition to the foregoing scheme, another interactive scheme is provided during the interaction with the driver to further focus on driving safety. Referring to Figure 3 , Figure 3 The flowchart of another human-computer interaction method provided by the embodiment of the present application is shown in the figure, and the following part will be introduced in detail.
[0093] Step S200, obtaining an interactive question raised by a target user in the vehicle.
[0094] Step S210, obtaining an image of the target user taken by a first camera in the vehicle, and determining the line-of-sight direction information of the target user according to the image.
[0095] Step S220, adjusting the viewing angle of a second camera in the vehicle according to the line-of-sight direction information, so that the viewing angle of the second camera is consistent with the line-of-sight direction of the target user, and obtaining an image outside the vehicle taken by the second camera.
[0096] Specifically, the steps S200-S220 above correspond one by one to the steps S100-S120 in the foregoing embodiment, and detailed reference is made to the foregoing description, which will not be repeated here.
[0097] Step S230, obtaining the task perception result of at least one active safety detection task configured.
[0098] Specifically, when the target user is a driver, before the multi-modal large model is called to process the interactive question of the driver, at least one active safety detection task configured is first executed. These tasks are designed to monitor and analyze the state information of the driver, the environment information outside the vehicle and the driving state information of the vehicle in real time, so as to ensure driving safety.
[0099] Driver status information: The on-board DMS camera can monitor the driver's physiological and behavioral status, such as fatigue level, concentration, head and eye movements, to assess whether the driver is able to drive safely.
[0100] External environmental information: External sensors and cameras can be used to collect environmental data, including traffic flow, weather conditions, road conditions, etc., to help determine the impact of the external environment on driving safety.
[0101] Vehicle driving status information: It can also monitor the vehicle's real-time driving status, such as speed, acceleration, braking status, etc., which helps to understand the vehicle's dynamic behavior and prevent potential driving risks.
[0102] This application integrates this information to determine perception results for active safety detection tasks. These results reflect the current driving environment and the driver's state, providing a basis for determining whether safety measures should be taken. Active safety detection tasks may include fatigue detection, speed monitoring and warnings, pedestrian detection, and blind spot monitoring. For example, if signs of driver fatigue are detected, the system may remind the driver to rest or take other preventative measures.
[0103] After acquiring these task perception results, they can be combined with the interactive questions raised by the driver and invoked to conduct in-depth analysis and processing using a multimodal large model. The multimodal large model can comprehensively consider the user's interactive questions and the perception results of the active safety detection tasks. In this way, this application can not only respond to the driver's information needs, but also ensure that the answers provided are also considered in consideration of driving safety, thereby providing a more comprehensive and reliable driving assistance experience.
[0104] Step S240: calling the multimodal large model to instruct the multimodal large model to comprehensively consider the task perception result and the image outside the vehicle to generate the response result corresponding to the interactive question.
[0105] Specifically, when dealing with interactive questions raised by the driver, multiple factors can be comprehensively considered. Before preparing to call the multimodal large model, the perception results of at least one active safety detection task can be obtained and analyzed. These results are based on the driver's status information, external environment information and vehicle driving status information. These perception results are input into the multimodal large model together with the external image of the vehicle taken by the second camera, and the multimodal large model is instructed to comprehensively consider this information to generate a response result corresponding to the driver's interactive question. Based on the comprehensive consideration of the perception results of the safety detection task, the multimodal large model can output a response result while ensuring driving safety. This response result not only takes into account accuracy, but also driving safety factors. There can be many situations for the response results. The following are examples of some situations:
[0106] Handling safety issues: The multimodal model will determine whether to output an answer to the interaction question based on the input information. If the multimodal model's analysis of the task perception results reveals safety issues, such as driver fatigue or distraction, these safety issues will be addressed first. In this case, the multimodal model will not output the specific question posed by the user, but will instead issue a safety reminder requiring the driver to take immediate action, such as "Attention, the system has detected a pedestrian ahead. Please concentrate on driving. We will interact with you later."
[0107] Continuous safety monitoring: During periods of unresolved safety issues, such as persistent driver fatigue, the multimodal large model will suspend outputting answers to user questions until the safety issue is resolved. This prevents distracted drivers from ignoring safety warnings and ensures safe driving.
[0108] Response when there are no safety issues: When the task perception results do not contain safety issues, the multimodal large model will process the user's question normally and output the corresponding answer. For example, if the driver asks about the landmark ahead, the name of the landmark and related information can be provided by combining the external image and the analysis results of the multimodal large model.
[0109] This embodiment implements a comprehensive human-computer interaction system that improves driving safety and interactive experience by analyzing the driver's status, the external environment, and the vehicle's driving state. It first executes at least one configured active safety detection task, collecting information about the driver's status, the external environment, and the vehicle's driving state to determine the task perception results. These perception results provide the system with important data about the driver's behavior and the external environment, forming the basis for the system's next steps. Next, a multimodal large model is invoked to combine the task perception results with the external vehicle image to generate specific responses to the driver's interactive questions. This embodiment not only understands the driver's questions but also provides accurate feedback based on the real-time environment and vehicle status. In this way, this embodiment not only provides information assistance but also promptly alerts the driver when potential safety issues are detected, ensuring safe driving. This design reflects the emphasis on driving safety while ensuring that the driver receives timely and effective information feedback when it is safe to do so, enhancing the practicality and reliability of the driving assistance system.
[0110] Furthermore, in some embodiments of the present application, the process of calling the multimodal large model in step S240 in the aforementioned embodiment to instruct the multimodal large model to comprehensively consider the task perception results and the external vehicle image to generate the response result corresponding to the interactive question is explained in detail.
[0111] One possible implementation involves using a more detailed second prompt format template when the multimodal large model is invoked and the target user is detected as a driver. This template not only includes the interaction question and the external vehicle image, but also integrates the task perception results. First, a second prompt format template is obtained. This template is more complex because it includes four main components: the second task instruction, the interaction question slot, the task perception result slot, and the external vehicle image slot. The second task instruction is the guiding part of the template, explicitly instructing the multimodal large model how to comprehensively consider the external vehicle image and the task perception results to generate a response corresponding to the interaction question.
[0112] The specific steps are as follows:
[0113] Fill the Interaction Question Slot: Fill the Interaction Question Slot in the template with the specific question the driver asks. For example, if the driver asks, "What is the building ahead?" the question will be placed in the Interaction Question Slot.
[0114] Filling the exterior image slot: At the same time, the exterior image captured by the second camera can be filled into the exterior image slot in the template to provide visual information for model analysis.
[0115] Filling the Task Perception Result Slot: You can also fill the Task Perception Result slot in the template with the Task Perception Result results obtained from the active safety detection task. These results may include key information such as the driver's fatigue state and the vehicle's current speed.
[0116] Generate a second prompt: After completing the above filling, a complete second prompt will be generated, which includes task instructions, specific interaction questions, relevant external images, and task perception results.
[0117] Input to the multimodal large model: This second prompt is fed into the pre-trained multimodal large model. Based on this structured input, the model analyzes and understands the question, while also considering external imagery and task perception results to generate an accurate response. Ultimately, the multimodal large model outputs a response to the driver's interactive question, taking into account the external environment and the driver's status to provide safer and more accurate driving assistance.
[0118] This embodiment further designs a second prompt template based on the specific identity of the target user (driver), ensuring driving safety while completing the interaction. This second prompt template not only includes the interaction question and the exterior image, but also integrates task perception results, such as the driver's fatigue level and the vehicle's current speed, among other key information. This comprehensively considers the exterior environment and the driver's state, generating safer and more accurate driving assistance responses. This not only improves the driving assistance system's responsiveness and interaction quality, but also enhances its adaptability and interactivity, providing users with a more intelligent, efficient, and safe driving experience.
[0119] Furthermore, in some embodiments of the present application, to address the issue of the second camera capturing the target object of interest to the target user with accuracy and clarity, a method can be employed to determine the focal length of the second camera by extracting distance information, thereby obtaining a clearer image outside the vehicle. This aspect is described in detail below.
[0120] Specifically, after obtaining an interactive question from a target user in the vehicle, the interactive question is parsed to determine whether it contains distance information. Upon receiving the user's interactive question, speech recognition technology is used to convert the user's voice command into text. Natural language processing technology is then used to parse the specific content of the question to determine whether it contains distance information. If so, the distance information is extracted and used to determine the focal length of the second camera. If no distance information is present, a default or preset focal length setting is used, and the control module directs the second camera to adjust its viewing angle to match the target user's line of sight.
[0121] One possible implementation involves extracting distance information using a variety of techniques, some examples of which are provided below. Natural language processing (NLP) techniques can be used to identify and extract distance information from text, rules can be designed to match and identify common distance expressions, and machine learning models can be trained to identify and extract distance information. Distance information can take various forms, including specific numerical distances such as "100 meters" or "200 kilometers," relative distance descriptions such as "nearby" or "faraway," distances relative to a reference object such as "at the next traffic light" or "across from the supermarket," time-based distance estimates such as "a 5-minute drive" or "a 10-minute walk," and common navigation system phrases such as "3 kilometers to the destination." Voice recognition and natural language processing technologies can parse these different types of distance information and convert them into specific numerical values, which are then used to adjust the focus and viewing angle of the secondary camera to ensure a clear image of the user's target.
[0122] Once distance information is successfully obtained from the interaction question, the appropriate focal length for the shot is determined based on this distance information. The system's onboard processor calculates a focal length value that matches the distance information and uses the control module to control the focus of the second camera. The second camera then adjusts its focus to ensure that the captured image is both clear and within the user's desired distance range. The system also obtains the target user's gaze direction information. The system's onboard processor combines these two pieces of information and uses the control module to adjust the second camera.
[0123] For example, if a user asks, "What's the traffic sign 100 meters ahead?" the system can analyze the line of sight and the distance "100 meters." Based on this information, the second camera's viewing angle and focal length are adjusted to focus on the target 100 meters in the direction of the user's line of sight. This allows the second camera to capture the specific target the user is looking at and provide a clear image.
[0124] This embodiment analyzes the user's interactive question to determine and extract distance information. If the interactive question contains distance information, it can interpret this information and adjust the second camera's shooting focus accordingly to ensure that the captured image matches the target distance of the driver's attention. This not only enhances understanding and response capabilities, but also integrates the user's gaze direction and distance information to adjust the second camera's viewing angle and focus, thereby more comprehensively understanding the user's intention and providing more accurate visual feedback.
[0125] Furthermore, the present application can utilize the image acquired by the first camera and the eye tracking method to determine the gaze direction information of the target user, that is, the gaze direction information of the driver or the passenger in the vehicle. The following is a detailed description of this process:
[0126] Locating the Eye Region: First, the target user's eyes are identified and located in the captured image. This step utilizes image processing techniques to detect and track facial features, then determine the position of the eyes, laying the foundation for further analysis of gaze direction.
[0127] Obtaining the spatial location of the eye origin and the direction of eye movement: Once the eye region is located, eye tracking algorithms can be used to identify and track the target user's eye movements, including eye rotations, blinks, and other eye movements. By analyzing these eye movements, the algorithm can determine the origin of the target user's eyes, which can be the center of the pupil or the location of the pupils of both eyes.
[0128] Determining gaze direction: Using data on the spatial location of the eye's origin and the direction of eye movement, the onboard processor can calculate the target user's gaze direction. This is typically located on a straight line between the eye and the object being gazed. By calculating the relative position between the eye's origin and the point of gaze, the algorithm can construct a virtual ray representing the target user's gaze direction. This ray not only indicates the approximate direction of the target user's gaze but also provides information on the angle between the gaze and the horizontal and vertical directions, which is crucial for understanding the target user's intentions and attention allocation. This calculation may involve in-depth analysis of eye image data, as well as possible geometric and trigonometric calculations, to determine the precise direction of the gaze.
[0129] This embodiment can accurately capture and analyze the target user's gaze direction, providing important input information for subsequent vehicle operations. The application of this gaze tracking technology enhances the interaction between the vehicle and the driver, making the driving experience more personalized and responsive.
[0130] Furthermore, in some embodiments of the present application, the first camera (the onboard DMS camera) can be used to obtain the driver's or passenger's line of sight direction information, and the viewing angle of the second camera (the rotatable camera) can be adjusted based on the line of sight direction information. The following is a detailed description of this process:
[0131] First, the first camera is used to capture images of the target user's eye area. These images are analyzed in the first coordinate system where the first camera resides to determine the target user's gaze direction information. Because the first and second cameras are located in different locations within the vehicle, they each have independent coordinate systems. The system's onboard processor uses known coordinate system mapping relationships to convert the gaze direction information in the first coordinate system to the second coordinate system where the second camera resides. This step may involve spatial geometry calculations to ensure the accurate transmission and adaptation of gaze direction information.
[0132] The system's onboard processor calculates the required rotation angle for the second camera based on the converted target gaze direction information. Based on the spatial relationship between the second camera's current position and the target gaze direction, the system ensures that the second camera's viewing angle aligns with the target user's gaze. After determining the rotation angle, the system's control module controls the rotation of the second camera according to the calculated rotation angle.
[0133] This embodiment ensures that the second camera's viewing angle aligns with the target user's line of sight, whether in the first coordinate system or the transformed second coordinate system. This viewing angle adjustment capability captures the user's focused object or scene and provides a corresponding exterior image, enhancing interactivity and functionality.
[0134] Further, see Figure 4, Figure 4 This is a schematic diagram of the working principle of a human-computer interaction system provided in an embodiment of the present application. The following is a detailed introduction to the picture.
[0135] Figure 4 The figure below illustrates the working principle of the human-computer interaction system. First, the first camera installed inside the vehicle, the onboard DMS camera, captures a real-time image of the target user. The system's onboard processor uses an eye-tracking algorithm to capture and analyze the driver's eye movements, determining the driver's eye origin and line of sight, enabling real-time monitoring of the driver's focus.
[0136] The system's onboard processor then sends instructions to a control module based on the driver's gaze direction. The control module then adjusts a second camera mounted on the vehicle's exterior, rotating it to capture real-time images of the exterior. These images are then transmitted to the onboard processor for further analysis.
[0137] When a user asks an interactive question, such as "What is the building 100 meters ahead?", the question and the image outside the vehicle are fed into a pre-trained multimodal large model. This large model comprehensively considers multiple data types, such as text and images, and uses deep learning algorithms to analyze and understand the content of the question, while also referencing the image outside the vehicle to generate an accurate response.
[0138] Finally, the response results output by the multimodal large model, such as "The building 100 meters ahead is the Guangzhou Tower", can be fed back to the driver through the on-board display system or voice interaction interface.
[0139] This embodiment not only monitors the driver's status and the external environment in real time, but also enhances the driving experience by providing timely and accurate feedback to help the driver better understand. Through this precise gaze tracking and analysis, the driver's intentions can be accurately understood, thereby improving the interactive quality of the driving assistance system.
[0140] Further, see Figure 5 , Figure 5 This is a schematic diagram of the working principle of another human-computer interaction system provided in an embodiment of the present application. The following is a detailed introduction to the picture.
[0141] Figure 5The diagram shows the system's workflow when the target person is the driver. The first camera, the onboard DMS camera, captures the driver's image, and the onboard processor applies an eye-tracking algorithm to determine the spatial position of the driver's eye origin and the direction of their gaze. This information is used to guide the second camera, a rotating camera, to adjust its viewing angle to capture real-time images of the vehicle's exterior. The system's onboard processor then sends instructions to the control module based on the driver's gaze direction. The control module then adjusts the second camera, mounted on the vehicle's exterior, to capture real-time images of the exterior. These images are then transmitted to the onboard processor for further analysis.
[0142] When the driver asks an interactive question, such as "What is the building 100 meters ahead?", the question, the external image, and the task perception results are input into a pre-trained multimodal large model. This model comprehensively processes text, image, and other data, analyzes the question content, and combines it with external image information to generate an accurate response. The task perception result is determined based on the driver's status information, the external environment information, and the vehicle's driving status information. This information may include key data such as the driver's fatigue level and the vehicle's current speed.
[0143] Ultimately, if the response output by the multimodal large model indicates a safety issue in the task perception result, it will stop responding to the user's questions and instead send a prompt message to the user, such as "Please note, the system has detected a pedestrian ahead. Please concentrate on driving and we will communicate with you later." This can be fed back to the driver through the in-vehicle display system or voice interaction interface.
[0144] This embodiment not only improves driving safety by enabling real-time monitoring of the driver's status and the external environment, but also enhances the driving experience, helping drivers better understand and navigate complex traffic situations. Through this precise gaze tracking and analysis of task perception results, the driver's intent is accurately understood, thereby improving the interactive quality and safety of the driver assistance system.
[0145] The human-computer interaction device provided in an embodiment of the present application is described below. The human-computer interaction device described below and the human-computer interaction method described above can be referenced to each other.
[0146] See also Figure 6 , Figure 6 A schematic diagram of the structure of a human-computer interaction device provided in an embodiment of the present application.
[0147] like Figure 6 As shown, the device may include:
[0148] A question acquisition unit 11 is used to acquire interactive questions raised by a target user in the vehicle;
[0149] An image acquisition unit 12 is configured to acquire an image of the target user captured by a first camera in the vehicle, and determine the sight direction information of the target user based on the image;
[0150] a viewing angle determining unit 13, configured to adjust the viewing angle of a second vehicle-mounted camera according to the sight line direction information so that the viewing angle of the second camera is consistent with the sight line direction of the target user, and to obtain an image outside the vehicle captured by the second camera;
[0151] The reply output unit 14 is used to call the multimodal large model to instruct the multimodal large model to generate a reply result corresponding to the interactive question in combination with the external vehicle image, and output the reply result.
[0152] In one possible implementation, the response output unit 14 calls the multimodal large model to instruct the multimodal large model to generate a response result corresponding to the interactive question based on the image outside the vehicle, and outputs the response result, including:
[0153] Obtaining a first prompt instruction prompt format template, the first prompt instruction prompt format template including a first task instruction, an interactive question slot, and an external vehicle image slot, the first task instruction being used to instruct the multimodal large model to combine the external vehicle image in the external vehicle image slot to generate the response result corresponding to the interactive question in the interactive question slot;
[0154] Fill the interactive question into the interactive question slot, fill the outside-vehicle image into the outside-vehicle image slot, obtain a first prompt instruction prompt, input the first prompt instruction prompt into the pre-trained multimodal large model, and obtain the reply result output by the multimodal large model.
[0155] In one possible implementation, when the target user is a driver, the human-computer interaction device of the embodiment of the present application further includes:
[0156] The task perception result acquisition unit is configured to acquire, before processing by the response output unit 14, a task perception result of at least one configured active safety detection task, wherein the task perception result of the active safety detection task is determined based on the driver's state information, the external vehicle environment information, and / or the vehicle driving state information. On this basis, the response output unit 14 calls the multimodal large model to instruct the multimodal large model to generate a response result corresponding to the interactive question in combination with the external vehicle image, including:
[0157] The multimodal large model is called to instruct the multimodal large model to comprehensively consider the task perception result and the image outside the vehicle to generate the response result corresponding to the interactive question.
[0158] In one possible implementation, the response output unit 14 calls the multimodal large model to instruct the multimodal large model to comprehensively consider the task perception result and the external vehicle image to generate the response result corresponding to the interaction question, including:
[0159] Obtain a second prompt instruction prompt format template, where the second prompt instruction prompt format template includes a second task instruction, an interaction question slot, a task perception result slot, and an external vehicle image slot, where the second task instruction is used to instruct the multimodal large model to comprehensively consider the external vehicle image in the external vehicle image slot and the task perception result in the task perception result slot to generate the response result corresponding to the interaction question in the interaction question slot;
[0160] Fill the interactive question into the interactive question slot, fill the outside-vehicle image into the outside-vehicle image slot, fill the task perception result into the task perception result slot, obtain a second prompt instruction prompt, input the second prompt instruction prompt into the pre-trained multimodal large model, and obtain the reply result output by the multimodal large model.
[0161] In one possible implementation, a human-computer interaction device according to an embodiment of the present application further includes:
[0162] The focal length determination unit is configured to analyze the interaction question before the viewing angle determination unit 13 processes the question. If distance information can be extracted from the interaction question, the focal length is determined based on the distance information, and the focus of the second camera is adjusted to the focal length.
[0163] In one possible implementation, the process of the image acquisition unit 12 determining the sight direction information of the target user based on the image includes:
[0164] Locating the eye area of the target user in the image;
[0165] Obtaining the spatial position of the human eye origin and the eye movement direction in the eye area;
[0166] The sight line direction information of the target person is determined according to the spatial position of the human eye origin and the eye movement direction.
[0167] In one possible implementation, the sight direction information is obtained based on a first coordinate system in which the first camera is located; and the viewing angle determination unit 13 adjusts the viewing angle of the second vehicle-mounted camera according to the sight direction information, including:
[0168] The line-of-sight direction information is converted into a second coordinate system in which the second camera is located by using a coordinate system mapping relationship between the first camera and the second camera, to obtain target line-of-sight direction information.
[0169] According to the target line-of-sight direction information, a rotation angle of the second camera is calculated, and the second camera is controlled to rotate according to the rotation angle.
[0170] An electronic device is also provided in the embodiments of the present application. As shown in Figure 7 The electronic device in the embodiments of the present application can include, but is not limited to, a fixed terminal such as a vehicle central processor, etc. Figure 7 The electronic device shown is merely an example, and should not impose any limitation on the functions and use range of the embodiments of the present application.
[0171] As shown in Figure 7 The electronic device can include a processing device (for example, a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or loaded from a storage device 608 into a random access memory (RAM) 603, to implement the aforementioned human-computer interaction method of the embodiments of the present application. In a powered-on state of the electronic device, various programs and data required for operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0172] Generally, the following devices can be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 608 including, for example, a memory card, a hard disk, etc.; and communication devices 609. The communication devices 609 can allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 The electronic device is shown with various devices, but it should be understood that all the shown devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.
[0173] A computer readable storage medium is also provided in the embodiments of the present application, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the aforementioned human-computer interaction method of the embodiments of the present application.
[0174] The present application also provides a computer program product in accordance with an embodiment of the present application. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the computer program instructions fully or partially generate the process or function described in accordance with the embodiment of the present application.
[0175] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0176] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0177] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0178] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referenced to each other.
Claims
1. A human-computer interaction method, characterized in that: include: Obtain interactive questions raised by target users in the car; Obtaining an image of the target user captured by a first camera in the vehicle, and determining sight direction information of the target user based on the image; Adjusting the viewing angle of a second camera on the vehicle according to the sight direction information so that the viewing angle of the second camera is consistent with the sight direction of the target user, and acquiring an image outside the vehicle captured by the second camera; Calling the multimodal large model to instruct the multimodal large model to generate a response result corresponding to the interactive question in combination with the external vehicle image, and output the response result; The calling of the multimodal large model to instruct the multimodal large model to generate a response result corresponding to the interactive question in combination with the external vehicle image and output the response result includes: Obtaining a first prompt instruction prompt format template, the first prompt instruction prompt format template including a first task instruction, an interactive question slot, and an external vehicle image slot, the first task instruction being used to instruct the multimodal large model to combine the external vehicle image in the external vehicle image slot to generate the response result corresponding to the interactive question in the interactive question slot; Fill the interactive question into the interactive question slot, fill the outside-vehicle image into the outside-vehicle image slot, obtain a first prompt instruction prompt, input the first prompt instruction prompt into the pre-trained multimodal large model, and obtain the reply result output by the multimodal large model.
2. The method according to claim 1, characterized in that If the target user is a driver, before calling the multimodal large model, the following steps are also included: Obtaining a task perception result of at least one configured active safety detection task, where the task perception result of the active safety detection task is determined based on the driver's state information, external environment information, and / or vehicle driving state information; The multimodal large model is called to instruct the multimodal large model to generate a response result corresponding to the interactive question in combination with the image outside the vehicle, including: The multimodal large model is called to instruct the multimodal large model to comprehensively consider the task perception result and the image outside the vehicle to generate the response result corresponding to the interactive question.
3. The method according to claim 2, characterized in that The calling of the multimodal large model to instruct the multimodal large model to comprehensively consider the task perception result and the external vehicle image to generate the response result corresponding to the interactive question includes: Obtain a second prompt instruction prompt format template, where the second prompt instruction prompt format template includes a second task instruction, an interaction question slot, a task perception result slot, and an external vehicle image slot, where the second task instruction is used to instruct the multimodal large model to comprehensively consider the external vehicle image in the external vehicle image slot and the task perception result in the task perception result slot to generate the response result corresponding to the interaction question in the interaction question slot; Fill the interactive question into the interactive question slot, fill the outside-vehicle image into the outside-vehicle image slot, fill the task perception result into the task perception result slot, obtain a second prompt instruction prompt, input the second prompt instruction prompt into the pre-trained multimodal large model, and obtain the reply result output by the multimodal large model.
4. The method according to claim 1, wherein Before acquiring the image outside the vehicle captured by the second camera, the method further includes: The interaction question is parsed, and if distance information can be extracted from the interaction question, a shooting focal length is determined according to the distance information, and the focus of the second camera is adjusted to the shooting focal length.
5. The method according to claim 1, wherein The process of determining the sight direction information of the target user according to the image includes: Locating the eye area of the target user in the image; Obtaining the spatial position of the human eye origin and the eye movement direction in the eye area; The sight line direction information of the target person is determined according to the spatial position of the human eye origin and the eye movement direction.
6. The method according to any one of claims 1 to 5, characterized in that The sight direction information is obtained based on the first coordinate system where the first camera is located; The process of adjusting the viewing angle of the second vehicle-mounted camera according to the line of sight direction information includes: Using a coordinate system mapping relationship between the first camera and the second camera, the sight line direction information is converted into a second coordinate system where the second camera is located to obtain target sight line direction information; The rotation angle of the second camera is calculated according to the target sight direction information, and the rotation of the second camera is controlled according to the rotation angle.
7. A human-computer interaction device, characterized in that: include: A question acquisition unit, used to acquire interactive questions raised by the target user in the car; an image acquisition unit, configured to acquire an image of the target user captured by a first camera in the vehicle, and determine the sight direction information of the target user based on the image; a viewing angle determining unit, configured to adjust the viewing angle of a second onboard camera according to the sight line direction information so that the viewing angle of the second camera is consistent with the sight line direction of the target user, and to obtain an image outside the vehicle captured by the second camera; a response output unit, configured to call the multimodal large model to instruct the multimodal large model to generate a response result corresponding to the interactive question in combination with the external vehicle image, and output the response result; The process in which the response output unit calls the multimodal large model to instruct the multimodal large model to generate a response result corresponding to the interactive question in combination with the external vehicle image and output the response result includes: Obtaining a first prompt instruction prompt format template, the first prompt instruction prompt format template including a first task instruction, an interactive question slot, and an external vehicle image slot, the first task instruction being used to instruct the multimodal large model to combine the external vehicle image in the external vehicle image slot to generate the response result corresponding to the interactive question in the interactive question slot; Fill the interactive question into the interactive question slot, fill the outside-vehicle image into the outside-vehicle image slot, obtain a first prompt instruction prompt, input the first prompt instruction prompt into the pre-trained multimodal large model, and obtain the reply result output by the multimodal large model.
8. An electronic device, characterized in that: include: memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the human-computer interaction method according to any one of claims 1 to 6.
9. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the human-computer interaction method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, each step of the human-computer interaction method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Target object information pushing method and device
CN110784523A
Man-machine interaction system and method for intelligent driving
CN112115797A