Human-computer interaction method and related device, system and storage medium

By combining voice and gaze data to recognize user commands in intelligent vehicles, the problem of insufficient freedom and accuracy in human-computer interaction has been solved, enabling wake-word-free switching and gaze-based supplementation, thereby improving the system's freedom and accuracy in interaction.

CN116080672BActive Publication Date: 2026-04-10IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2022-09-08
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing intelligent vehicle systems, the freedom to switch between human-machine interaction and human-to-human interaction is low, and the accuracy and timeliness of voice interaction are insufficient, especially when voice commands are incomplete, which affects system response.

Method used

By combining voice and image data when the gaze interaction function is enabled, user commands are recognized and gaze status is detected. The target of the command is determined by the gaze position and duration, thus realizing wake-word-free switching of human-computer interaction and supplementation of voice commands.

Benefits of technology

It enhances the freedom of human-computer interaction and human-to-human interaction, improves the accuracy and timeliness of interaction, reduces interference when gazing at multiple points, and enhances the system's response accuracy and speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116080672B_ABST
    Figure CN116080672B_ABST
Patent Text Reader

Abstract

The application discloses a human-computer interaction method and related devices, systems and storage media, wherein the human-computer interaction method comprises: in response to the gaze interaction function being in an open state, when collecting voice data of a user, image data containing the face of the user is photographed; based on the voice data, a user instruction is obtained, and based on a vehicle cabin model and the image data, a line-of-sight gaze condition of the user in the vehicle is detected; wherein the line-of-sight gaze condition comprises whether a gaze position is detected, and the duration of each gaze position when the gaze position is detected; in response to the line-of-sight gaze condition comprising the detection of the gaze position, it is determined that the user instruction needs to be executed, and based on the duration of each gaze position, a first executed object of the user instruction is determined, and the user instruction is executed on the first executed object. The above scheme can improve the degree of freedom of switching between human-computer interaction and human-human interaction, and improve the accuracy and timeliness of human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent interaction, in particular to a human-computer interaction method and related device, system and storage medium. BACKGROUND

[0002] With the continuous improvement of products and technologies in the field of intelligent vehicles, users have increasingly high requirements for intelligent vehicles. In particular, with the continuous popularization of voice interaction applications, users have increasingly high requirements for voice interaction in intelligent vehicles.

[0003] However, users use voice interaction more and more casually, and users expect the switching between human-computer interaction and human-human interaction to be more free. Currently, systems usually rely on a wake-up word to start human-computer interaction. On the one hand, mainstream systems can only achieve short-time wake-up-free after one wake-up, and cannot achieve long-time wake-up-free, thereby greatly limiting the degree of freedom of switching between human-computer interaction and human-human interaction. On the other hand, due to the casualness of users when using voice interaction, the voice instruction may not be complete (for example, the voice instruction "open the window" does not specify which window to open), thereby affecting the accuracy and timeliness of voice interaction. Therefore, how to improve the degree of freedom of switching between human-computer interaction and human-human interaction, and improve the accuracy and timeliness of human-computer interaction, has become a problem to be solved. SUMMARY

[0004] The technical problem solved by the present application is to provide a human-computer interaction method and related device, system and storage medium, which can improve the degree of freedom of switching between human-computer interaction and human-human interaction, and improve the accuracy and timeliness of human-computer interaction.

[0005] To solve the above technical problem, the present application provides a human-computer interaction method in the first aspect, comprising: in response to the gaze interaction function being in an open state, when collecting voice data of a user, shooting image data containing the face of the user; based on the voice data, identifying to obtain a user instruction, and based on a vehicle cabin model and the image data, detecting to obtain a line of sight gaze condition of the user in the vehicle; wherein the line of sight gaze condition includes whether a gaze position is detected, and the duration of each gaze position when the gaze position is detected; in response to the line of sight gaze condition including detecting the gaze position, determining that the user instruction needs to be executed, and based on the duration of each gaze position, determining a first executed object of the user instruction, and executing the user instruction on the first executed object.

[0006] To solve the above technical problems, the second aspect of the present application provides a human-computer interaction device, comprising: a data acquisition module, a voice recognition module, a line-of-sight detection module, and an instruction execution module. The data acquisition module is configured to, in response to the gaze interaction function being in an open state, capture image data containing a user's face when collecting the user's voice data. The voice recognition module is configured to recognize based on the voice data to obtain a user instruction. The line-of-sight detection module is configured to detect based on the vehicle cabin model and the image data to obtain the user's line-of-sight gaze condition in the vehicle. The line-of-sight gaze condition includes whether a gaze position is detected and the duration of each gaze position when the gaze position is detected. The instruction execution module is configured to, in response to the line-of-sight gaze condition including the detection of the gaze position, determine that the user instruction needs to be executed, determine a first executed object of the user instruction based on the duration of each gaze position, and execute the user instruction on the first executed object.

[0007] To solve the above technical problems, the third aspect of the present application provides a human-computer interaction system, comprising a microphone, a camera, and a vehicle machine. The microphone and the camera are coupled to the vehicle machine. The microphone is configured to collect voice data, the camera is configured to capture image data, and the vehicle machine is configured to execute the human-computer interaction method of the first aspect.

[0008] To solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions executable by a processor, the program instructions being configured to implement the human-computer interaction method of the first aspect.

[0009] The above scheme, in response to the gaze interaction function being in an open state, captures image data containing a user's face when collecting the user's voice data, thereby recognizing based on the voice data to obtain a user instruction, and detecting based on the vehicle cabin model and the image data to obtain the user's line-of-sight gaze condition in the vehicle. The line-of-sight gaze condition includes whether a gaze position is detected and the duration of each gaze position when the gaze position is detected. Based on this, in response to the line-of-sight gaze condition including the detection of the gaze position, it is determined that the user instruction needs to be executed, and a first executed object of the user instruction is determined based on the duration of each gaze position. The user instruction is executed on the first executed object. On the one hand, human-computer interaction is realized by combining line-of-sight gaze and user voice, so that it is not necessary to rely on a wake-up word to switch to human-computer interaction, thereby improving the degree of freedom of switching between human-computer interaction and human-human interaction. On the other hand, the incompleteness of the voice instruction can be compensated for by line-of-sight gaze, and the first executed object of the user instruction can be determined with the help of the duration of each gaze position, which can reduce the interference with the machine response as much as possible when there are multiple gaze positions, thereby helping to improve the accuracy and timeliness of human-computer interaction. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1This is a flowchart illustrating an embodiment of the applicant's human-computer interaction method;

[0011] Figure 2 This is a schematic diagram illustrating an embodiment of recognizing user commands;

[0012] Figure 3 This is a schematic diagram illustrating another embodiment of the process for recognizing user instructions;

[0013] Figure 4 This is a schematic diagram of a process in one embodiment of the gaze interaction preparation stage;

[0014] Figure 5 This is a schematic diagram of the framework of an embodiment of the applicant's human-computer interaction device;

[0015] Figure 6 This is a schematic diagram of the framework of an embodiment of the applicant's human-computer interaction system;

[0016] Figure 7 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0017] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0018] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0019] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the slash " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper indicates two or more objects.

[0020] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the human-computer interaction method of this applicant. It should be noted that the human-computer interaction method in this embodiment can be executed by a vehicle-mounted system, which can run a vehicle-mounted system based on embedded systems such as Android or WinCE; this is not limited thereto. Specifically, this embodiment may include the following steps:

[0021] Step S11: In response to the gaze interaction function being enabled, capture image data containing the user's face while collecting the user's voice data.

[0022] In an implementation scenario, the car machine can include a touch screen, and a user can set to turn on or turn off the gaze interaction function through the touch screen. Specifically, in the case of turning on the gaze interaction function, the human-computer interaction can be realized through the steps in the embodiments of the present disclosure, and details can be referred to the embodiments of the present disclosure; and in the case of turning off the gaze interaction function, the human-computer interaction can be realized by relying on the wake-up word and the like, and details can be referred to the related technical details of the wake-up word, which will not be repeated here.

[0023] In an implementation scenario, in the case of turning on the gaze interaction function, the microphone and the camera can always be in a working state, the microphone can be used to collect audio, and the camera can be used to capture images of the user's head. At the same time, the car machine can perform voice activity detection on the collected audio to detect the start time and the end time of the user's speech, and intercept the audio from the start time to the end time to obtain the voice data of the user, and intercept the images from the start time to the end time to obtain the image data.

[0024] In another implementation scenario, different from the foregoing manner, in the case of turning on the gaze interaction function, the microphone can always be in a working state, and the car machine can perform voice activity detection on the collected audio while collecting audio to detect the start time of the user's speech, and in response to detecting the start time, the camera is turned on to capture images of the user's head. When the car machine detects the end time of the user's speech, the camera can be turned off, so that the audio from the start time to the end time can be used as the voice data of the user, and the images from the start time to the end time can be used as the image data.

[0025] It should be noted that when the user triggers to turn on the gaze interaction function, a prompt can be given that "the gaze interaction function requires microphone, camera and other hardware support, and in the process of turning on the gaze interaction function, your voice, image and other information will be collected. If you confirm to turn on the gaze interaction function, it is considered that you have known and agreed to the information collection", in the case of user's confirmation and agreement, the gaze interaction function can be kept in the on state, otherwise, if the user disagrees, the gaze interaction function can be switched to the off state.

[0026] Step S12: performing recognition based on the voice data to obtain a user instruction, and performing detection based on the car cabin model and the image data to obtain a gaze condition of the user in the car.

[0027] In one implementation scenario, only the voice data can be recognized to obtain the user instruction. In order to improve the recognition efficiency, the voice recognition model can be pre-trained, which can include but is not limited to a recurrent neural network, a time delay neural network, etc., and of course, an encoder-decoder (i.e., Encoder-Decoder), a CTC (i.e., Connectionist Temporal Classification) model framework can also be used, and the network structure of the voice recognition model is not limited here. Illustratively, sample voice can be pre-collected, and the sample voice is labeled with corresponding sample text. On this basis, the sample voice can be input into the voice recognition model for recognition to obtain the recognized text corresponding to the sample voice. Based on this, the network parameters of the voice recognition model can be adjusted based on the difference between the sample text and the recognized text. After the voice recognition model is trained and converged, the voice data can be input into the voice recognition model to recognize the user instruction.

[0028] In another implementation scenario, please refer to Figure 2 , Figure 2 is a process schematic diagram of an embodiment of recognizing the user instruction. As Figure 2 indicated, in order to further improve the recognition accuracy, the user instruction can also be recognized based on the voice data and the lip image extracted from the image data, so that the user instruction can be recognized based on the multi-modal data, which helps to improve the recognition accuracy. Specifically, the first feature extraction can be performed on the current lip image to obtain the first image feature, and the first image feature includes a plurality of channel first sub-features. The first sub-feature located in the target channel in the first image feature extracted from the current lip image is replaced with the first sub-feature located in the target channel in the first image feature extracted from the reference lip image to obtain the second image feature, and the reference lip image is located before the current lip image. On this basis, the second feature extraction is performed based on the second image feature to obtain the image feature of the current lip image, so that the user instruction can be obtained based on the voice feature extracted from the voice data of the image feature of each lip image. The above-mentioned manner replaces the first sub-feature located in the target channel in the first image feature extracted from the current lip image with the first sub-feature located in the target channel in the first image feature extracted from the reference lip image to obtain the second image feature, so that the temporal context information of the previous frame can be fully utilized to assist the feature extraction of the current frame, which helps to build the association between the video frames and further combine the image feature and the voice feature for recognition, which helps to improve the recognition accuracy, especially in complex scenarios such as noisy environment, which can greatly improve the accuracy of voice recognition.

[0029] In one specific implementation scenario, please refer to Figure 3 , Figure 3This is a schematic diagram illustrating another embodiment of the process for recognizing user instructions. For example... Figure 3 As shown, the image data can include several video frames. For each video frame, face detection can be performed to detect the face region in the video frame. Then, feature point detection can be performed on the face region to detect the feature points in the face region (e.g., 68 key points of the face). Based on this, the lip image can be extracted from the video frame based on the feature points related to the lips.

[0030] In a specific implementation scenario, the target channel can be the first 1 / 8 of the channels. For example, if the first image feature contains a first sub-feature with 16 channels, then the first 2 channels can be selected as the target channels. Of course, in practical applications, it is not limited to this. For example, the first 1 / 4 or the first 1 / 2 of the channels can also be selected as the target channels. There is no limitation here.

[0031] In a specific implementation scenario, please refer to the following: Figure 3 The first and second feature extractions mentioned above can be performed based on the MobileNet backbone network. Of course, the backbone network is not limited to MobileNet and can also include, but is not limited to, ResNet, etc., which are not specified here. The backbone network can include several sequentially connected network blocks for feature extraction. Based on the backbone network, temporal convolution and gated recurrent units (GRUs) can be further combined. Taking a backbone network consisting of two network blocks and N frames of lip images extracted from image data as an example, the first feature extraction of the lip images from frame 1 to frame N is performed through the first network block, which can extract the first image features of the lip images from frame 1 to frame N respectively. For ease of distinction, these can be denoted as f. 1_1 ... f 1_N Furthermore, for the lip image in the i-th frame, its first image feature f can be added. 1_i The first sub-feature of the first 1 / 8 channels is replaced with the first image feature f of the lip image in the (i-1)th frame. 1_i-1 The first sub-feature of the first 1 / 8 channel is used to obtain the second image feature f of the lip image in the i-th frame. 2_i When the backbone network contains a different number of network blocks or the lip image has a different number of frames, the same principle applies, and examples will not be given here.

[0032] In a specific implementation scenario, please refer to the following: Figure 3The Fb40 feature of the voice data can be extracted as the audio feature. Of course, in actual application, it is not limited to this, and other acoustic features such as MFCC (Mel Frequency Cepstral Coefficient) can also be extracted as the audio feature, which is not limited herein.

[0033] In a specific implementation scenario, please continue to refer to Figure 3 In order to combine the audio feature and the image feature, the image feature and the audio feature can also be input into a convolutional neural network (CNN), a deep neural network (DNN), a long short-term memory mapping layer (LSTMP), etc. for processing to map to the same feature space and be fused, and then a speech recognition network is used to process the fused feature to obtain the user instruction. The specific structure of the speech recognition network can be referred to the foregoing related description, which will not be repeated here.

[0034] In the embodiments of the present disclosure, the line of sight gaze case can include whether the gaze position is detected, and the duration of each gaze position when the gaze position is detected. It should be noted that the duration specifically represents the duration of the user's gaze at the gaze position. Specifically, the spatial position of the pupil and the attitude information of the head can be obtained based on the image data, and the eye image can be extracted based on the image data. On this basis, the attitude information and the eye image can be detected based on the line of sight detection model to obtain the line of sight direction, so that the line of sight gaze case can be obtained based on the spatial position, the line of sight direction and the vehicle cabin model, and the gaze position is the landing position of the vehicle cabin model along the line of sight direction from the spatial position. The above-mentioned manner determines the line of sight gaze case by combining the spatial position of the pupil, the line of sight direction of the user and the vehicle cabin model, which helps to improve the accuracy of the gaze interaction.

[0035] In an implementation scenario, please refer to Figure 4 , Figure 4 is a process schematic diagram of an embodiment of the gaze interaction preparation stage. As Figure 4As shown, in order to further improve the accuracy of gaze detection, the camera can be calibrated in advance to obtain the internal parameters of the camera. Further, a calibration board integrated with a laser range finder can be used to model the cabin in three dimensions, and the camera coordinate system and the cabin coordinate system are related. Finally, after the vehicle is delivered, due to installation errors and user adjustments, a fixed reference object (such as an instrument panel) in the cabin can be used as a reference point of the cabin coordinate system, and the angle of the camera relative to the cabin coordinate system can be calculated using the parallax principle. It should be noted that after the above processing, the gaze perception area can be customized according to the project vehicle layout to meet different needs, and gaze perception of positions such as the inside rearview mirror, left / right rearview mirror, center control screen, instrument panel, road surface, HUB, left / right window, etc. can be realized.

[0036] In one implementation scenario, as previously described, the camera coordinate system can be calibrated in advance, and on this basis, the face can be modeled in three dimensions based on image data in combination with the camera coordinate system, so that the spatial position of the pupil can be located. For specific process of three-dimensional modeling, reference can be made to technical details of modeling methods such as SFM (Structure From Motion), which will not be described here.

[0037] In one implementation scenario, as previously described, feature point detection (such as 68 key points of a face) can be performed on image data, so that an eye image can be extracted based on feature points related to the human eye. For specific extraction process of the eye image, reference can be made to the aforementioned extraction process of the lip image, which will not be described here.

[0038] In one implementation scenario, the pose information of the head can include but is not limited to yaw (i.e., yaw angle), roll (i.e., roll angle), pitch (i.e., pitch angle), etc., which will not be limited here.

[0039] In one implementation scenario, the gaze detection model can include but is not limited to a convolutional neural network, etc., and the network structure of the gaze detection model will not be limited here. On this basis, the gaze detection model can be used to analyze the eye image and fuse the pose information of the head, so as to detect the gaze direction.

[0040] In one implementation scenario, as previously described, the image data can include a plurality of video frames, so after the spatial position of the pupil and the gaze direction of the user in each video frame are detected, the intersection on the cabin model can be determined by extending along the gaze direction from the spatial position of the pupil, and if the intersection points corresponding to the last N video frames are all on the same vehicle component, the position of the vehicle component can be taken as the landing position of the user's gaze in the last N frames, that is, the gaze position, and the duration of the last N frames can be taken as the duration of the gaze position.

[0041] Step S13: in response to the line-of-sight gaze case including the detected gaze position, determining that the user instruction needs to be executed, and determining a first executed object of the user instruction based on the duration of each gaze position, and executing the user instruction on the first executed object.

[0042] In one implementation scenario, in the case that the line-of-sight gaze case includes the detected gaze position, the gaze position in the line-of-sight gaze case can be selected as the first target position, and in response to the duration of the first target position satisfying a first condition, the vehicle component located at the first target position is determined as the first executed object. Specifically, the first condition can be set to include a time duration not less than a first duration. Exemplarily, the first duration can be set to 0.5 seconds, 1 second, 2 seconds, etc., which is not limited herein. In the above manner, by selecting the gaze position in the line-of-sight gaze case as the first target position, and in response to the duration of the first target position satisfying the first condition, the vehicle component located at the first target position is determined as the first executed object, which can improve the accuracy of determining the executed object.

[0043] In one specific implementation scenario, in the case that the line-of-sight gaze case includes multiple gaze positions, the last detected gaze position in the line-of-sight gaze case can be selected as the first target position. In the above manner, in the case that the line-of-sight gaze case includes multiple gaze positions, the last detected gaze position in the line-of-sight gaze case is selected as the first target position, which can reduce the pre-gaze interference as much as possible and improve the accuracy of gaze interaction. In addition, in the case that the line-of-sight gaze case includes only one gaze position, the gaze position can be directly selected as the first target position.

[0044] In one specific implementation scenario, in response to the duration of the first target position not satisfying the first condition, it can be determined that the user instruction does not need to be executed. At this time, no response can be made. In the above manner, in response to the duration of the first target position not satisfying the first condition, it is determined that the user instruction does not need to be executed, which can exclude the interference of short line-of-sight stay on gaze interaction and help to further improve the accuracy of gaze interaction.

[0045] In one implementation scenario, in actual application, the line-of-sight gaze case can also include a case where no gaze position is detected, and in response to the line-of-sight gaze case including no detected gaze position, it can be determined that the user instruction does not need to be executed. At this time, no response can be made. In the above manner, in response to the line-of-sight gaze case including no detected gaze position, it is determined that the user instruction does not need to be executed, which can exclude the interference of voice data generated when the line-of-sight does not stay on human-computer interaction and help to further improve the accuracy of human-computer interaction.

[0046] In an implementation scenario, in the case that the driver is staring at the co-driver window and saying "open the window", voice data "open the window" can be collected, and image data containing the face of the driver can be captured. On this basis, the user instruction (i.e., open the window) can be obtained based on the voice data, and the line of sight gaze condition of the user in the vehicle can be obtained based on the vehicle cabin model and the image data, and at this time the line of sight gaze condition includes detecting a gaze position, and the gaze position is located at the co-driver window, and the duration is 1 second. At this time, the gaze position can be taken as the first target position in response to the line of sight gaze condition including detecting the gaze position, and since there is only one gaze position, the vehicle component (i.e., the co-driver window) of the first target position can be taken as the first executed object, and the user instruction can be executed on it, so as to achieve opening the co-driver window. Other cases can be similarly deduced, which will not be repeated here. Through actual testing, by the above-mentioned manner, the accuracy rate of human-computer interaction can reach more than 90%, the response time is shortened to within 600 milliseconds, and the accuracy error of the gaze position is within 3 degrees.

[0047] In an implementation scenario, in the case that voice data is not collected and it is detected that the user triggers the control button, it can be detected whether the control button has been associated with an adjustment object. In response to the control button not being associated with the adjustment object and the line of sight gaze condition including detecting the gaze position, the gaze position with a duration satisfying a second condition can be selected as the second target position, and the vehicle component located at the second target position can be determined as the second executed object, and the control instruction defined by the control button can be executed on the second executed object. The above-mentioned manner, in the case that voice data is not collected and it is detected that the user triggers the control button, by detecting whether the control button has been associated with the adjustment object, in the case that the control button is not associated with the adjustment object and the line of sight gaze condition includes detecting the gaze position, the gaze position with a duration satisfying a second condition is further selected as the second target position, so as to perform subsequent instruction execution operation based on this, so as to be able to preferentially execute in the case that there is gaze and vehicle control, and thus it is possible to simultaneously satisfy the multi-modal scene of existing voice and image, and the single-modal scene of only existing voice, which helps to improve the use range of human-computer interaction.

[0048] In a specific implementation scenario, the control button can be a physical button of the vehicle or a virtual button of the vehicle. For example, a physical button defined as "cooling" can be integrated on the vehicle, or a virtual button defined as "navigation" can be provided on the vehicle machine touch screen, which will not be repeated here.

[0049] In a specific implementation scenario, it is to be noted that the adjustment object is a vehicle component such as the co-driver window, the driver window, the air conditioner, etc. In addition, the control button can be pre-associated with the adjustment object. For example, the aforementioned entity button defined as “cooling” is pre-associated with the adjustment object “air conditioner”, and other cases can be similarly deduced and will not be listed one by one. In addition, the control button can also be pre-associated with the adjustment object. For example, a virtual button defined as “open window” can be provided on the car machine touch screen, but it is not associated with the adjustment object, which is specifically the “driver window” or the “co-driver window” or the “rear seat window”. Other cases can be similarly deduced and will not be listed one by one.

[0050] In a specific implementation scenario, the second condition can include a duration of no less than a second duration, which can be set to 0.5 seconds, 1 second, 2 seconds, etc., which is not limited here. That is, if the duration of the gaze position is less than the second duration, it can be considered that the gaze position is an invalid position, and it will not be used as the second target position.

[0051] In a specific implementation scenario, if the gaze position in the line of sight gaze case contains multiple gaze positions with a duration satisfying the second condition, the last detected gaze position (i.e., the last detected gaze position among the gaze positions with a duration satisfying the second condition) can be selected as the second target position, and the vehicle component located at the second target position can be determined as the second executed object. For example, the user gazes at the left rearview mirror for a second duration and then switches to gaze at the right rearview mirror for a second duration. At this time, the right rearview mirror can be selected as the second executed object, and other cases can be similarly deduced and will not be listed one by one. The above-mentioned method, in the case where the line of sight gaze case contains multiple gaze positions with a duration satisfying the second condition, selects the last detected gaze position as the second target position, which can eliminate ambiguity as much as possible in the presence of multiple gaze positions with a duration satisfying the second condition, and improve the accuracy of human-computer interaction response.

[0052] In a specific implementation scenario, unlike the foregoing, in the case where it is detected that the control button is not associated with the adjustment object, the adjustment object associated with the control button can be directly determined as the third executed object, and the control instruction defined by the control button can be executed on the third executed object. The above-mentioned method, in response to the control button not being associated with the adjustment object, directly determines the adjustment object associated with the control button as the third executed object, and executes the control instruction defined by the control button on the third executed object, thereby enabling direct human-computer interaction response in the case where the control button is associated with the adjustment object, which helps to improve the response speed.

[0053] In one specific implementation scenario, in order to facilitate subsequent use, in the case that the control button is not associated with the adjustment object, if the gaze fixation condition includes detecting the fixation position, as described above, the fixation position that satisfies the second condition for the duration can be selected as the second target position, and the vehicle component located at the second target position is determined as the second executed object, at the same time, the second executed object can be associated as the adjustment object of the control button, so that when the user triggers the control button next time, the adjustment object associated with it can be directly taken as the third executed object, and the control instruction defined by the control button is executed on the third executed object. Further, when the user needs to cancel the association, the adjustment object associated with the control button can be canceled through the car machine touch screen. Of course, after the cancellation is successful, a new adjustment object can be associated again through the above steps, and the specific steps can be referred to the foregoing description, which will not be described here.

[0054] In one specific implementation scenario, in response to the control button not being associated with the adjustment object and the gaze fixation condition including not detecting the fixation position, at this time, the control button can be considered as invalid triggering, and the control instruction defined by the control button can not be executed. In the above manner, in the case that the control button is not associated with the adjustment object and the gaze fixation condition includes not detecting the fixation position, the control instruction defined by the control button is not executed, which can effectively filter invalid triggering and help improve the accuracy of human-computer interaction response.

[0055] In one specific implementation scenario, in response to the control button not being associated with the adjustment object and the gaze fixation condition including detecting the fixation position, in the case that the duration of each fixation position does not satisfy the second condition, at this time, the control button can be considered as invalid triggering, and the control instruction defined by the control button can not be executed. In the above manner, in the case that the control button is not associated with the adjustment object and the gaze fixation condition includes detecting the fixation position, but the duration of each fixation position does not satisfy the second condition, the control instruction defined by the control button is not executed, which can effectively filter invalid triggering and help improve the accuracy of human-computer interaction response.

[0056] In a specific implementation scenario, taking the control button defined as the control instruction "adjust the rearview mirror inward" as an example, when it is detected that the user triggers the control button, it can be detected first whether the control button has associated adjustment objects. For example, if the control button has associated adjustment objects "right rearview mirror", the adjustment objects associated with the control button "right rearview mirror" can be directly determined as the third executed object, and the control instruction "adjust the rearview mirror inward" defined by the control button can be executed on the third executed object "right rearview mirror". In the case where the control button is associated with other adjustment objects (for example, left rearview mirror), the same can be applied, which will not be exemplified one by one here. Conversely, if the control button is not associated with adjustment objects, the gaze position can be detected in response to the line of sight gaze condition, the gaze position whose duration satisfies the second condition (for example, not less than 2 seconds) is selected as the second target position, and the vehicle component at the second target position, such as the left rearview mirror, is determined. The "left rearview mirror" can be taken as the second executed object, and the control instruction defined by the control button, i.e., "adjust the rearview mirror inward", can be executed on the second executed object "left rearview mirror", so as to realize the human-computer interaction "adjust the left rearview mirror inward". In the case where the vehicle component at the second target position is the right rearview mirror, the "right rearview mirror" can be taken as the second executed object, and the control instruction defined by the control button, i.e., "adjust the rearview mirror inward", can be executed on the second executed object "right rearview mirror", so as to realize the human-computer interaction "adjust the right rearview mirror inward". Other cases can be applied in the same way, which will not be exemplified one by one here.

[0057] In an implementation scenario, when the gaze interaction function is in an open state, it can be further detected whether the vehicle gear position satisfies a third condition. Specifically, the third condition can be set as the vehicle gear position being in P or N. If the vehicle gear position satisfies the third condition, the step of "capturing image data containing the user's face when collecting the user's voice data" and subsequent steps can be executed. Conversely, if the vehicle gear position does not satisfy the third condition, it can be prompted that "the current driving state does not support the gaze interaction function", which helps to improve driving safety. In addition, in order to further improve driving safety, after the user triggers the gaze interaction function and it is in an open state, it can be prompted that "you have turned on the gaze interaction function. In order to ensure driving safety, the gaze interaction function is only valid when the vehicle is in P or N. Please pay attention to road safety during driving".

[0058] In the above scheme, in response to the gaze interaction function being in an open state, when collecting voice data of a user, image data containing a face of the user is photographed, so as to identify based on the voice data to obtain a user instruction, and detect based on a vehicle cabin model and the image data to obtain a line of sight gaze condition of the user in the vehicle. The line of sight gaze condition includes whether a gaze position is detected, and a duration of each gaze position when the gaze position is detected. In response to the line of sight gaze condition including that the gaze position is detected, it is determined that the user instruction needs to be executed, and based on the duration of each gaze position, a first executed object of the user instruction is determined, and the user instruction is executed on the first executed object. On the one hand, human-computer interaction is realized in combination with the line of sight gaze and the voice of the user, so that the human-computer interaction is switched without relying on a wake-up word, thereby improving the degree of freedom of switching between human-computer interaction and human-human interaction. On the other hand, the line of sight gaze can make up for the incompleteness of the voice instruction, and the first executed object of the user instruction is determined by means of the duration of each gaze position, which can reduce the interference of the machine response as much as possible when there are multiple gaze positions, thereby helping to improve the accuracy and timeliness of human-computer interaction.

[0059] Please refer to Figure 5 , Figure 5 is a frame schematic diagram of an embodiment of the human-computer interaction device 50 of the present application. The human-computer interaction device 50 comprises a data acquisition module 51, a voice recognition module 52, a line of sight detection module 53, and an instruction execution module 54. The data acquisition module 51 is configured to, in response to the gaze interaction function being in an open state, when collecting voice data of a user, photograph image data containing a face of the user. The voice recognition module 52 is configured to identify based on the voice data to obtain a user instruction. The line of sight detection module 53 is configured to detect based on a vehicle cabin model and the image data to obtain a line of sight gaze condition of the user in the vehicle. The line of sight gaze condition includes whether a gaze position is detected, and a duration of each gaze position when the gaze position is detected. The instruction execution module 54 is configured to, in response to the line of sight gaze condition including that the gaze position is detected, determine that the user instruction needs to be executed, and based on the duration of each gaze position, determine a first executed object of the user instruction, and execute the user instruction on the first executed object.

[0060] The above scheme, in response to the gaze interaction function being in an open state, when collecting voice data of the user, image data containing the face of the user is photographed, so as to obtain the user instruction based on the voice data, and the gaze condition of the user in the vehicle is obtained based on the vehicle cabin model and the image data, and the gaze condition includes whether the gaze position is detected, and the duration of each gaze position when the gaze position is detected, and based on this, in response to the gaze condition including detecting the gaze position, it is determined that the user instruction needs to be executed, and based on the duration of each gaze position, the first executed object of the user instruction is determined, and the user instruction is executed on the first executed object. Therefore, on the one hand, the line of sight and the user voice are combined to realize human-computer interaction, so as to switch to human-computer interaction without relying on the wake-up word, thereby improving the degree of freedom of switching between human-computer interaction and human-human interaction. On the other hand, the gaze can make up for the incompleteness of the voice instruction, and the duration of each gaze position is used to determine the first executed object of the user instruction, which can reduce the interference of the machine response as much as possible when there are multiple gaze positions, and help to improve the accuracy and timeliness of human-computer interaction.

[0061] In some disclosed embodiments, the instruction execution module 54 includes a position selection sub-module for selecting a gaze position in the gaze condition as a first target position, and the instruction execution module 54 includes an object determination sub-module for determining, in response to the duration of the first target position satisfying a first condition, a vehicle component located at the first target position as a first executed object.

[0062] Therefore, by selecting a gaze position in the gaze condition as a first target position, and determining, in response to the duration of the first target position satisfying a first condition, a vehicle component located at the first target position as a first executed object, the accuracy of determining the executed object can be improved.

[0063] In some disclosed embodiments, in the case where the gaze condition includes multiple gaze positions, the first target position is the last detected gaze position in the gaze condition.

[0064] Therefore, in the case where the gaze condition includes multiple gaze positions, the last detected gaze position in the gaze condition is selected as the first target position, which can reduce the pre-gaze interference as much as possible and improve the accuracy of the gaze interaction.

[0065] In some disclosed embodiments, the instruction execution module 54 includes a silent processing sub-module for determining, in response to the duration of the first target position not satisfying the first condition, that the user instruction does not need to be executed.

[0066] Therefore, in response to the duration of the first target position not satisfying the first condition, it is determined that the user instruction does not need to be executed, which can help to further improve the accuracy of the gaze interaction by excluding the interference of the temporary line-of-sight stay on the gaze interaction.

[0067] In some disclosed embodiments, the human-computer interaction apparatus 50 further comprises an object detection module configured to, in a case where the voice data is not collected and the user trigger control button is detected, detect whether the control button is associated with an adjustment object, and the human-computer interaction apparatus 50 further comprises a first control module configured to, in a case where the control button is not associated with the adjustment object and the line-of-sight gaze condition comprises a detected gaze position, select the gaze position with a duration satisfying a second condition as a second target position, determine a vehicle component located at the second target position as a second executed object, and execute a control instruction defined by the control button on the second executed object.

[0068] Therefore, in a case where the voice data is not collected and the user trigger control button is detected, by detecting whether the control button is associated with an adjustment object, in a case where the control button is not associated with the adjustment object and the line-of-sight gaze condition comprises a detected gaze position, the gaze position with a duration satisfying a second condition is further selected as a second target position, and subsequent instruction execution operations are performed based on this, so that the execution can be prioritized in the presence of gaze and vehicle control, thereby being able to simultaneously meet the multi-modal scenario of the presence of voice and image and the single-modal scenario of the presence of voice only, and helping to improve the use range of human-computer interaction.

[0069] In some disclosed embodiments, the human-computer interaction apparatus 50 further comprises a second control module configured to, in a case where the control button is associated with the adjustment object, directly determine the adjustment object associated with the control button as a third executed object, and execute the control instruction defined by the control button on the third executed object.

[0070] Therefore, in a case where the control button is not associated with the adjustment object, the adjustment object associated with the control button is directly determined as a third executed object, and the control instruction defined by the control button is executed on the third executed object, so that in a case where the control button is associated with the adjustment object, the human-computer interaction response can be directly performed, which helps to improve the response speed.

[0071] In some disclosed embodiments, the human-computer interaction apparatus 50 further comprises a first mute module configured to, in a case where the control button is not associated with the adjustment object and the line-of-sight gaze condition comprises a non-detected gaze position, not execute the control instruction defined by the control button.

[0072] Therefore, in a case where the control button is not associated with the adjustment object and the line-of-sight gaze condition comprises a non-detected gaze position, the control instruction defined by the control button is not executed, which can effectively filter invalid triggers and help to improve the accuracy of human-computer interaction response.

[0073] In some disclosed embodiments, the human-computer interaction apparatus 50 further comprises a second muting module configured to, in response to the control button being not associated with the adjustment object and the gaze condition comprising detecting the gaze position, and in the case that the duration of each gaze position does not satisfy a second condition, not executing the control instruction defined by the control button.

[0074] Therefore, in the case that the control button is not associated with the adjustment object and the gaze condition comprises detecting the gaze position, but the duration of each gaze position does not satisfy the second condition, the control instruction defined by the control button is not executed, which can effectively filter invalid triggers and help improve the accuracy of human-computer interaction response.

[0075] In some disclosed embodiments, the human-computer interaction apparatus 50 further comprises a muting response module configured to, in response to the gaze condition comprising not detecting the gaze position, determine that the user instruction does not need to be executed.

[0076] Therefore, in response to the gaze condition comprising not detecting the gaze position, determining that the user instruction does not need to be executed can exclude the interference of voice data generated when the gaze does not stay on the human-computer interaction, and help further improve the accuracy of human-computer interaction.

[0077] In some disclosed embodiments, the user instruction is recognized based on the voice data and the lip image extracted from the image data.

[0078] Therefore, the user instruction is recognized based on the voice data and the lip image extracted from the image data, so that the user instruction can be recognized in combination with multi-modal data, which helps improve the recognition accuracy.

[0079] In some disclosed embodiments, the voice recognition module 52 comprises a first extraction sub-module configured to perform first feature extraction on the current lip image to obtain first image features; wherein the first image features comprise first sub-features of a plurality of channels; the voice recognition module 52 comprises a feature replacement sub-module configured to replace the first sub-features located in a target channel in the first image features extracted from the current lip image with the first sub-features located in the target channel in the first image features extracted from a reference lip image to obtain second image features; wherein the reference lip image is located before the current lip image; the voice recognition module 52 comprises a second extraction sub-module configured to perform second feature extraction based on the second image features to obtain image features of the current lip image; the voice recognition module 52 comprises an instruction recognition sub-module configured to obtain the user instruction based on the image features of each lip image and the voice features extracted from the voice data.

[0080] Therefore, by replacing the first sub-feature located in the target channel in the first image feature extracted from the current lip image with the first sub-feature located in the target channel in the first image feature extracted from the reference lip image, the second image feature is obtained, so that the temporal context information of the previous frame can be fully utilized to assist the feature extraction of the current frame, which helps to build the association between video frames, and further combines the image features and speech features for recognition, which helps to improve the recognition accuracy, especially in complex scenes such as noisy environments, which can greatly improve the accuracy of speech recognition.

[0081] In some disclosed embodiments, the gaze detection module 53 includes an image analysis submodule for analyzing based on image data to obtain spatial position of the pupil and attitude information of the head, and extracting an eye image based on the image data; the gaze detection module 53 includes a direction detection submodule for detecting the attitude information and the eye image based on a gaze detection model to obtain a gaze direction; the gaze detection module 53 includes a situation determination submodule for determining a gaze situation based on the spatial position, the gaze direction, and a vehicle cabin model; wherein the gaze position is a landing position of the vehicle cabin model along the gaze direction from the spatial position.

[0082] Therefore, by combining the spatial position of the pupil, the gaze direction of the user, and the vehicle cabin model to determine the gaze situation, the accuracy of the gaze interaction can be improved.

[0083] Please refer to Figure 6 , Figure 6 is a schematic diagram of an embodiment of a human-computer interaction system 60 of the present application. The human-computer interaction system 60 includes a microphone 61, a camera 62, and a vehicle machine 63 coupled with each other, the microphone 61 is used to collect speech data, the camera 62 is used to shoot image data, and the vehicle machine 63 is used to execute the steps in any of the above human-computer interaction method embodiments.

[0084] Specifically, the car machine 63 is configured to control itself and the microphone 61 and the camera 62 to implement the steps in any of the above-mentioned human-computer interaction method embodiments. The car machine 63 can be integrated with a processor (not shown), which can be referred to as a CPU (Central Processing Unit). The processor can be an integrated circuit chip with processing capability. The processor can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor can be implemented by an integrated circuit chip together.

[0085] The above-mentioned scheme, on the one hand, implements human-computer interaction in combination with line of sight gaze and user voice, so that it is not necessary to rely on a wake-up word to switch to human-computer interaction, thereby being able to improve the degree of freedom of switching between human-computer interaction and human-human interaction, and on the other hand, the incompleteness of the voice instruction can be made up by means of line of sight gaze, and the first executed object of the user instruction can be determined by means of the duration of each gaze position, so that the interference with the machine response can be reduced as much as possible when there are multiple gazes, which helps to improve the accuracy and timeliness of human-computer interaction.

[0086] Please refer to Figure 7 , Figure 7 is a frame diagram of an embodiment of the computer readable storage medium 70 of the present application. The computer readable storage medium 70 stores program instructions 71 capable of being executed by a processor, which are used to implement the steps in any of the above-mentioned human-computer interaction method embodiments.

[0087] The above-mentioned scheme, on the one hand, implements human-computer interaction in combination with line of sight gaze and user voice, so that it is not necessary to rely on a wake-up word to switch to human-computer interaction, thereby being able to improve the degree of freedom of switching between human-computer interaction and human-human interaction, and on the other hand, the incompleteness of the voice instruction can be made up by means of line of sight gaze, and the first executed object of the user instruction can be determined by means of the duration of each gaze position, so that the interference with the machine response can be reduced as much as possible when there are multiple gazes, which helps to improve the accuracy and timeliness of human-computer interaction.

[0088] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, they will not be described here again.

[0089] The above description of the various embodiments tends to emphasize differences between the various embodiments, and the same or similar elements can be referred to each other, and will not be described herein for the sake of brevity.

[0090] In several embodiments provided in the present application, it should be understood that the disclosed method and device can be implemented in other ways. For example, the above-described device implementation is only schematic, and the division of the modules or units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0091] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.

[0092] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0093] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that makes a contribution or the whole or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0094] If the technical solution of the present application involves personal information, the product applying the technical solution of the present application has clearly informed the personal information processing rules before processing the personal information and obtained the personal independent consent. If the technical solution of the present application involves sensitive personal information, the product applying the technical solution of the present application has obtained the personal independent consent before processing the sensitive personal information and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent mark is set to inform that it has entered the personal information collection range and will collect personal information. If the individual voluntarily enters the collection range, it is considered to agree to collect personal information. Or on the device for processing personal information, through the pop-up information or by asking the individual to upload his / her personal information, the individual's authorization is obtained under the condition of using obvious mark / information to inform the personal information processing rules. The personal information processing rules can include personal information processor, personal information processing purpose, processing method and personal information type, etc.

Claims

1. A human-machine interaction method, characterized in that, The method comprises: in response to the gaze interaction function being in an open state, capturing image data containing a face of a user while capturing voice data of the user; based on the voice data, identifying a user instruction, and based on a vehicle cabin model and the image data, detecting a line-of-sight gaze condition of the user in the vehicle; wherein the line-of-sight gaze condition comprises whether a gaze position is detected, and a duration of each of the gaze positions when the gaze position is detected; in response to the line-of-sight gaze condition comprising the gaze position being detected, determining that the user instruction needs to be executed, and based on the duration of each of the gaze positions, determining a first executed object of the user instruction, and executing the user instruction on the first executed object; wherein, in a case where the voice data is not captured and a user trigger control button is detected, the method further comprises: detecting whether the control button is associated with an adjustment object; in response to the control button not being associated with the adjustment object and the line-of-sight gaze condition comprising the gaze position being detected, selecting a gaze position with a duration satisfying a second condition as a second target position, and determining a vehicle component located at the second target position as a second executed object, and executing a control instruction defined by the control button on the second executed object.

2. The method of claim 1, wherein, The method further comprises: in response to the duration of the first target position satisfying a first condition, determining that a vehicle component located at the first target position is the first executed object. In a case where the line-of-sight gaze condition comprises a plurality of the gaze positions, the first target position is a last detected gaze position in the line-of-sight gaze condition.

3. The method of claim 2, wherein, The method further comprises:

4. The method of claim 2, wherein, in response to the duration of the first target position not satisfying the first condition, determining that the user instruction does not need to be executed. The method further comprises:

5. The method of claim 1, wherein, in response to the control button being associated with the adjustment object, directly determining the adjustment object associated with the control button as a third executed object, and executing the control instruction defined by the control button on the third executed object. The method further comprises at least one of:

6. The method of claim 1, wherein, in response to the control button not being associated with the adjustment object and the line-of-sight gaze condition comprising the gaze position not being detected, not executing the control instruction defined by the control button; in response to the control button not being associated with the adjustment object and the line-of-sight gaze condition comprising the gaze position being detected, in a case where the duration of each of the gaze positions does not satisfy the second condition, not executing the control instruction defined by the control button. The method further comprises:

7. The method of claim 1, wherein, in response to the line-of-sight gaze condition comprising the gaze position not being detected, determining that the user instruction does not need to be executed. The user instruction is identified based on the voice data and lip images extracted from the image data.

8. The method of claim 1, wherein, A plurality of the lip images are extracted from the image data, and the identification of the user instruction comprises:

9. The method of claim 8, wherein, ​ performing first feature extraction on the current lip image to obtain first image features; wherein the first image features include first sub-features of a plurality of channels; replacing first sub-features of a target channel in the first image features extracted from the current lip image with first sub-features of the target channel extracted from a reference lip image to obtain second image features; wherein the reference lip image is located before the current lip image; performing second feature extraction based on the second image features to obtain image features of the current lip image; obtaining the user instruction based on the image features of each of the lip images and speech features extracted from the speech data.

10. The method of claim 1, wherein, the detection based on the cabin model and the image data to obtain the line-of-sight gaze of the user in the vehicle, includes: analyzing based on the image data to obtain spatial positions of pupils and posture information of a head, and extracting an eye image based on the image data; detecting based on a line-of-sight detection model to obtain a line-of-sight direction based on the posture information and the eye image; obtaining the line-of-sight gaze based on the spatial positions, the line-of-sight direction, and the cabin model; wherein the gaze position is a landing position on the cabin model along the line-of-sight direction from the spatial position.

11. A human-machine interaction device, characterized in that, comprising: a data acquisition module configured to, in response to the gaze interaction function being in an open state, capture image data containing a user's face while capturing speech data of the user; a speech recognition module configured to recognize the user instruction based on the speech data; a line-of-sight detection module configured to detect, based on a cabin model and the image data, a line-of-sight gaze of the user in the vehicle; wherein the line-of-sight gaze includes whether a gaze position is detected, and a duration of each of the gaze positions when the gaze position is detected; an instruction execution module configured to, in response to the line-of-sight gaze including the detection of the gaze position, determine that the user instruction needs to be executed, and determine a first executed object of the user instruction based on the duration of each of the gaze positions, and execute the user instruction on the first executed object; wherein, in the case that the speech data is not captured and the user triggers a control button, further comprising: detecting whether the control button is associated with an adjustment object; in response to the control button not being associated with the adjustment object and the line-of-sight gaze including the detection of the gaze position, selecting a gaze position whose duration satisfies a second condition as a second target position, and determining a vehicle component located at the second target position as a second executed object, and executing a control instruction defined by the control button on the second executed object.

12. A human-machine interaction system, characterized by a microphone, a camera, and a vehicle machine, the microphone and the camera are coupled to the vehicle machine, the microphone is configured to capture speech data, the camera is configured to capture image data, and the vehicle machine is configured to execute the human-computer interaction method of any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The storage stores program instructions capable of being run by the processor, and the program instructions are used for realizing the human-computer interaction method in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Human-vehicle interaction device and method

    CN110857067A

  • Human-vehicle interaction system

    CN113442941A