Man-machine interaction method and device and carrying tool
Patent Information
- Application Number
- CN202280100738.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-22
- Publication Date
- 2025-05-13
AI Technical Summary
The voice interaction method of existing vehicle virtual assistants is single and lacks a sense of interaction, which affects the user's human-computer interaction experience.
By obtaining the sight direction and posture information of the user in the vehicle, the virtual assistant is controlled to interact with the user, including face pinching, stroking, summoning and throwing gestures, etc., to enhance the sense of interaction, and save the gestures, postures and interactive actions on the cloud server or locally. Mapping relationships to achieve flexible interaction.
It improves the human-computer interaction experience when the user interacts with the virtual assistant, saves the computing overhead of the vehicle, reduces power consumption, and increases the intimacy of the virtual assistant and the user's interaction satisfaction.
Smart Images

Figure CN119998760A_ABST
Abstract
Description
Human-computer interaction method, device and vehicle Technical Field
[0001] The embodiments of the present application relate to the field of human-computer interaction, and more specifically, to a human-computer interaction method, device, and vehicle. Background Art
[0002] With the development of intelligent vehicles, more and more vehicles are equipped with virtual assistants. For the sake of convenience, most interactions between users and virtual assistants are voice-based. This interaction method is relatively simple and lacks a sense of interactivity, which affects the user's human-computer interaction experience.
[0003] Summary of the Invention
[0004] The embodiments of the present application provide a human-computer interaction method, device, and vehicle, which help to enhance the sense of interaction when a user interacts with a virtual assistant, thereby helping to enhance the user's human-computer interaction experience.
[0005] The vehicles in this application may include road vehicles, water vehicles, air vehicles, industrial equipment, agricultural equipment, or recreational equipment. For example, the vehicle may be a vehicle, which is a vehicle in a broad sense, and may be a vehicle (such as a commercial vehicle, a passenger car, a motorcycle, a flying car, a train, etc.), an industrial vehicle (such as a forklift, a trailer, a tractor, etc.), an engineering vehicle (such as an excavator, a bulldozer, a crane, etc.), agricultural equipment (such as a mower, a harvester, etc.), amusement equipment, a toy vehicle, etc. The embodiments of this application do not specifically limit the type of vehicle. For another example, the vehicle may be a vehicle such as an airplane or a ship.
[0006] In a first aspect, a human-computer interaction method is provided, which is applied to a cockpit of a vehicle, wherein the cockpit includes a virtual assistant, the method comprising: obtaining the sight direction of a first user in the cockpit; when the sight direction of the first user is on the virtual assistant, obtaining the posture information of the first user; and controlling the virtual assistant to interact with the first user according to the posture information of the first user.
[0007] In this embodiment of the present application, when the user's gaze is directed toward the virtual assistant, the virtual assistant can be controlled to interact with the user based on the user's posture information. This helps enhance the user's human-computer interaction experience when interacting with the virtual assistant by having the virtual assistant respond to the user's posture. Furthermore, obtaining the user's posture information when the user's gaze is directed toward the virtual assistant helps reduce the vehicle's computing overhead, thereby reducing the vehicle's power consumption.
[0008] In combination with the first aspect, in certain implementations of the first aspect, the posture information of the first user includes a first gesture posture, and controlling the virtual assistant to interact with the first user based on the posture information of the first user includes: controlling the virtual assistant to interact with the first user based on the first gesture posture.
[0009] In an embodiment of the present application, when the first user's line of sight is directed at the virtual assistant, the virtual assistant can be controlled to respond to the user's gestures, which helps to enhance the user's human-computer interaction experience when interacting with the virtual assistant.
[0010] In some possible implementations, the vehicle stores a mapping relationship between gesture postures and interactive actions, and controls the virtual assistant to interact with the first user according to the first gesture posture, including: controlling the virtual assistant to interact with the first user through a first interactive action according to the first gesture posture and the mapping relationship.
[0011] In some possible implementations, the mapping relationship may also be stored in a cloud server. After obtaining the human posture information of the first user, the vehicle may send the human posture information to the cloud server. The cloud server may determine a first interaction action based on the human posture information and the mapping relationship. The cloud server may send information about the first interaction action to the vehicle. The vehicle may then control the virtual assistant to interact with the first user through the first interaction action.
[0012] In some possible implementations, before controlling the virtual assistant to interact with the first user according to the first gesture, the method further includes: detecting an operation of the user setting a mapping relationship between the first gesture and the first interaction action.
[0013] In an embodiment of the present application, a user can set a mapping relationship between their desired gesture posture and the interactive action of the virtual assistant in the vehicle. When a user's first gesture posture is detected, the virtual assistant can be controlled to interact with the first user through the corresponding interactive action based on the mapping relationship between the gesture posture set by the user and the interactive action of the virtual assistant. While improving the user's human-computer interaction experience, it also helps to increase the flexibility in setting the mapping relationship between gesture posture and the interactive action of the virtual assistant.
[0014] In combination with the first aspect, in some implementations of the first aspect, the posture information includes at least one of a face pinching gesture, a stroking gesture, a summoning gesture, and a throwing gesture.
[0015] In some possible implementations, the posture information further includes a head pinching gesture or a body pinching gesture.
[0016] In some possible implementations, the gesture information also includes a finger rotation gesture. For example, the forward direction of the vehicle is the x-axis, and the direction perpendicular to the plane of the vehicle is the z-axis. When the user's finger is detected rotating around the x-axis, the virtual assistant can be controlled to do a somersault or flip. For another example, when the user's finger is detected rotating around the z-axis, the virtual assistant can be controlled to rotate 360° along the z-axis.
[0017] In some possible implementations, the virtual assistant may be a virtual pet.
[0018] In an embodiment of the present application, when the user's face pinching gesture, stroking gesture, summoning gesture and throwing gesture are detected, the virtual assistant can respond to the corresponding gesture posture, which helps to improve the human-computer interaction experience of the user interacting with the virtual assistant, and also helps to increase the user's intimacy with the virtual assistant.
[0019] In some possible implementations, the method further includes: acquiring image information of the pet; and generating the virtual pet based on the image information of the pet.
[0020] The embodiment of the present application can be divided into a modeling phase and an interaction phase. In the modeling phase, a three-dimensional model of the virtual pet can be generated based on the user's input; in the interaction phase, the virtual assistant can be controlled to respond to the user's gestures.
[0021] In some possible implementations, before generating the virtual pet based on the pet's image information, the method further includes: prompting the user to select a material for the virtual pet, where the material of the virtual pet includes a restorable material (e.g., elastic material) and an irreversible material (e.g., a clay figurine); and generating the virtual pet based on the material selected by the user.
[0022] In some possible implementations, the virtual assistant is displayed through the first display device, and the material of the virtual assistant is an irreversible material. Before controlling the virtual assistant to interact with the first user, the method includes: controlling the first display device to display the first interactive image of the virtual assistant; wherein, controlling the virtual assistant to interact with the first user includes: controlling the first display device to display the second interactive image of the virtual assistant; and keeping the second interactive image displayed when the interaction between the virtual assistant and the first user ends.
[0023] The above-mentioned end of the interaction between the virtual assistant and the first user can also be understood as the first user's line of sight is no longer on the virtual assistant.
[0024] For example, during the modeling phase, it is detected that the material of the virtual pet selected by the user is an unrecoverable material. During the interaction phase, when the user's posture information is not detected, the normal three-dimensional model of the virtual pet can be displayed (for example, the virtual pet's face is not pinched and remains smiling). When the user makes a face-pinching gesture, the virtual assistant can be controlled to perform an animation of the face being pinched. After the user stops making the face-pinching gesture, the virtual assistant can be controlled to keep the face in the pinched state. In this way, the virtual assistant does not return to its original image after the user interacts with the virtual assistant, which can make the virtual assistant serve as the user's punching bag, helping the user to vent negative emotions when the user is in a bad mood.
[0025] For example, if the material of the virtual pet selected by the user is detected as restorable during the modeling phase, during the interaction phase, when the user performs a face-pinching gesture, the virtual assistant can be controlled to perform an animation of the face being pinched; and after the user stops pinching the face, the virtual assistant can be controlled to restore the normal 3D model.
[0026] In some possible implementations, generating the virtual pet based on the image information of the pet includes: generating a three-dimensional model of the virtual pet corresponding to the pet based on the image information of the pet; and repairing the first organ in the three-dimensional model based on the user's operation on the first organ, thereby generating a repaired three-dimensional model of the virtual pet.
[0027] In some possible implementations, the user's operation on the first organ in the 3D model includes an air gesture. During the modeling process, the air gesture can be used to pinch the model. The air gesture can include a rotation gesture or a flip gesture. By pinching or patting, the image of the virtual pet's organs or limbs can be modified until the virtual pet presents the desired appearance.
[0028] For example, during the modeling phase, an original three-dimensional model of the pet can be generated based on the pet's image information through a display device, and a virtual hand can be displayed through the display device, wherein the virtual hand can move with the movement of the user's hand. When it is detected that the user's hand is moving toward the nose in the original three-dimensional model, the virtual hand on the display device is controlled to move to the nose of the original three-dimensional model. When a gesture of pinching the nose of the user's hand is detected and the user's hand moves in a direction away from the display device, the nose in the original three-dimensional model can be repaired to obtain a repaired three-dimensional model. For example, the farther the user's hand moves, the pointeder the nose of the virtual pet becomes.
[0029] In this embodiment of the present application, users can upload photos or videos of their pets to a vehicle, which can then generate a virtual pet based on the pet's photos or videos. This helps further enhance the user experience when interacting with the virtual pet. Furthermore, users can select the material of the virtual assistant or change the virtual assistant's appearance to their desired image through operation, which helps increase the user's flexibility in setting up the virtual assistant and thus improve the user experience.
[0030] In combination with the first aspect, in certain implementations of the first aspect, the posture information includes a throwing gesture, and controlling the virtual assistant to interact with the first user includes: controlling the virtual assistant to respond to the throwing gesture based on at least one of the first user's line of sight direction, the throwing angle corresponding to the throwing gesture, and the throwing speed corresponding to the throwing gesture.
[0031] In an embodiment of the present application, when a throwing gesture of a user is detected, the virtual assistant can be controlled to respond to the throwing gesture based on at least one of the user's line of sight direction, the throwing angle corresponding to the throwing gesture, and the throwing speed corresponding to the throwing gesture, so that the virtual assistant's feedback on the throwing gesture is more in line with the user's expectations, which helps to improve the user's human-computer interaction experience.
[0032] In some possible implementations, controlling the virtual assistant to respond to the throwing gesture includes: controlling the virtual assistant to perform an operation of picking up the ball.
[0033] In some possible implementations, when the throwing speed corresponding to the throwing gesture is greater than or equal to a preset speed threshold, the speed at which the virtual assistant picks up the ball can be controlled to be a first speed; or, when the throwing speed corresponding to the throwing gesture is less than the preset speed threshold, the speed at which the virtual assistant picks up the ball can be controlled to be a second speed, wherein the first speed is greater than the second speed.
[0034] In some possible implementations, the vehicle is a vehicle, which includes a main driving area and a non-main driving area, and the first user is a user in the main driving area. The method also includes: when the line of sight of the first user and the line of sight of a third user in the cockpit are both on the virtual assistant and the vehicle is in driving state, controlling the virtual assistant to interact with the third user.
[0035] In this embodiment of the present application, if it is detected that both the user in the driver's seat and the user in the non-driver's seat are looking at the virtual assistant and the vehicle is currently in motion, the virtual assistant can be controlled to interact with the third user. This prevents the driver's seat from being distracted by the virtual assistant's interaction with the user in the driver's seat. Furthermore, by not responding to the user's gestures in the driver's seat, the user's attention can be diverted to driving the vehicle, thereby helping to improve the user's driving safety.
[0036] In some possible implementations, the method further includes: prompting the first user to concentrate on driving the vehicle.
[0037] In some possible implementations, the method further includes: obtaining the line of sight direction of the fourth user in the cabin; when the line of sight direction of the first user and the line of sight direction of the fourth user are both on the virtual assistant and the information of the first user is saved in the vehicle but the information of the fourth user is not saved, controlling the virtual assistant to interact with the first user.
[0038] In this embodiment of the present application, when multiple users wish to interact with the virtual assistant, the virtual assistant can be controlled to interact with the first user based on the user information stored in the vehicle. This helps improve the human-computer interaction experience between the virtual assistant and the users who have entered their user information into the vehicle, and also helps prevent the virtual assistant's interactive actions from disturbing strangers when they are looking at the virtual assistant.
[0039] The above user information may include the user's biometric information, such as face information, iris information, etc.
[0040] In combination with the first aspect, in certain implementations of the first aspect, the posture information includes human body posture, and controlling the virtual assistant to interact with the first user based on the posture information of the first user includes: controlling the virtual assistant to interact with the first user based on the human body posture.
[0041] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: controlling the multimedia device in the vehicle to play a multimedia file; wherein, based on the posture information of the first user, controlling the virtual assistant to interact with the first user includes: during the playback of the multimedia file, controlling the virtual assistant to imitate the posture of the first user.
[0042] In an embodiment of the present application, during the playback of multimedia files, the virtual assistant can be controlled to imitate the posture of the first user, which can create a feeling of dancing with the virtual assistant for the user, helping to enhance the user's human-computer interaction experience.
[0043] In combination with the first aspect, in certain implementations of the first aspect, the posture information of the first user is obtained when the first user's line of sight is on the virtual assistant, including: obtaining the posture information of the first user when the first user's line of sight is on the virtual assistant for a duration greater than or equal to a preset duration.
[0044] In an embodiment of the present application, when the first user's line of sight on the virtual assistant lasts for a period greater than or equal to a preset period, obtaining the first user's posture information helps to further improve the accuracy of the virtual assistant's response to the user's posture.
[0045] Exemplarily, the preset duration is 1 second (s).
[0046] In combination with the first aspect, in some implementations of the first aspect, the virtual assistant is a physical virtual assistant.
[0047] In some possible implementations, the vehicle is a vehicle, which includes a front area and a rear area, and the physical virtual assistant is located in the front area of the vehicle.
[0048] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: obtaining the line of sight direction of the second user in the back row area; when the line of sight direction of the second user is on the physical virtual assistant, controlling the display device in the back row area to display a first image and obtaining the body posture information of the second user, the first image including the image of the physical virtual assistant; based on the body posture information of the second user, controlling the display device in the back row area to display a second image, the second image including an image of the virtual assistant responding to the body posture of the second user.
[0049] In an embodiment of the present application, when a second user in the back row wishes to interact with a physical virtual assistant in the front row, the image information of the physical virtual assistant can be displayed on a display device in the back row, and the display device in the back row can be controlled to display an image of the physical virtual assistant interacting with the second user. This can avoid the feeling of alienation that may arise when the back row user interacts with the physical virtual assistant due to the long distance, thereby improving the human-computer interaction experience of the back row user when interacting with the physical virtual assistant.
[0050] In combination with the first aspect, in certain implementations of the first aspect, the virtual assistant is displayed via a display device in the cockpit.
[0051] In some possible implementations, the user's gaze on the virtual assistant may include the user's gaze on a display device that displays the virtual assistant.
[0052] In some possible implementations, controlling the virtual assistant to interact with the first user includes: controlling a first display device in the cockpit to display a video or animation of the virtual assistant interacting with the first user.
[0053] In some possible implementations, the first display device is associated with an area where the first user is located.
[0054] In some possible implementations, the method further includes: obtaining the line of sight of a fifth user in the cockpit; when the line of sight of the fifth user is on the virtual assistant, switching from controlling the first display device to display the virtual assistant to controlling the second display device to display the virtual assistant and obtaining the posture information of the fifth user, the second display device being associated with the area where the fifth user is located; and according to the posture information of the fifth user, controlling the second display device to display an image of the interaction between the virtual assistant and the fifth user.
[0055] The second display device is associated with the area where the fifth user is located, including: the second display device is a display device in the area where the fifth user is located. For example, if the fifth user is located in the passenger seat area, the second display device may be a passenger seat entertainment screen.
[0056] In this embodiment of the present application, when the fifth user in the cabin is looking at the virtual assistant, the control of displaying the virtual assistant on the first display device can be switched to displaying the virtual assistant on the second display device. This ensures that the virtual assistant is always displayed on the display device closest to the user, increasing the user's sense of familiarity when interacting with the virtual assistant, helping to enhance the user's human-computer interaction experience, and also contributing to the intelligent level of the vehicle.
[0057] In some possible implementations, the method further includes: when the fifth user's line of sight is directed toward the virtual assistant, prompting the virtual assistant to be displayed on a display device in an area where the fifth user is located.
[0058] In some possible implementations, prompting the virtual assistant to be displayed on a display device in the area where the fifth user is located includes: emitting a prompt sound through a sound-emitting device to prompt the virtual assistant to be displayed on the display device in the area where the fifth user is located.
[0059] In some possible implementations, the sound-emitting device may be located near the second display device. This allows the fifth user to quickly switch their gaze from the virtual assistant on the first display device to the virtual assistant on the second display device, allowing the fifth user to quickly discover that the virtual assistant is displayed on the display device closest to them, thereby enhancing their human-computer interaction experience.
[0060] In some possible implementations, the physical virtual assistant and the virtual assistant displayed by the display device may respond differently to user interactions. For example, the physical virtual assistant is located in a plane area in the front row, and the physical virtual assistant can move in the plane area. When a user's face-pinching gesture is detected, the expression of the physical virtual assistant can be controlled to change. If the virtual assistant is displayed by a display device, when a user's face-pinching gesture is detected, an animation effect of the virtual assistant's face being pinched can be displayed through the display device.
[0061] For another example, when a throwing gesture of the user is detected and the throwing direction corresponding to the throwing gesture is from the main driver's seat to the passenger seat, the physical virtual assistant can be controlled to move from a position close to the main driver's seat in the plane area to a position close to the passenger seat; and if the virtual assistant is displayed through a display device, when a throwing gesture of the user is detected, the display device can display an animation effect of the virtual assistant moving from the main driver's seat to the passenger seat to pick up the ball.
[0062] In combination with the first aspect, in certain implementations of the first aspect, obtaining the sight line direction of the first user in the cockpit includes: obtaining the sight line direction of the first user when the virtual assistant is in an awake state.
[0063] In a second aspect, a human-computer interaction device is provided, which includes: an acquisition unit for acquiring the line of sight direction of a first user in a cabin of a vehicle, wherein the cabin includes a virtual assistant; the acquisition unit is also used to acquire posture information of the first user when the line of sight of the first user is on the virtual assistant; and a control unit is used to control the virtual assistant to interact with the first user according to the posture information of the first user.
[0064] In combination with the second aspect, in some implementations of the second aspect, the posture information includes at least one of a face pinching gesture, a stroking gesture, a summoning gesture, and a throwing gesture.
[0065] In combination with the second aspect, in certain implementations of the second aspect, the posture information includes a throwing gesture, and the control unit is used to control the virtual assistant to respond to the throwing gesture based on at least one of the first user's line of sight direction, the throwing angle corresponding to the throwing gesture, and the throwing speed corresponding to the throwing gesture.
[0066] In combination with the second aspect, in certain implementations of the second aspect, the control unit is used to: control the multimedia device in the vehicle to play a multimedia file; and during the playback of the multimedia file, control the virtual assistant to imitate the posture of the first user.
[0067] In combination with the second aspect, in certain implementations of the second aspect, the acquisition unit is used to: acquire the posture information of the first user when the duration of the first user's line of sight on the virtual assistant is greater than or equal to a preset duration.
[0068] In combination with the second aspect, in some implementations of the second aspect, the virtual assistant is a physical virtual assistant, or the virtual assistant is displayed through a display device in the cockpit.
[0069] In combination with the second aspect, in certain implementations of the second aspect, the vehicle includes a front area and a rear area, the virtual assistant is a physical virtual assistant located in the front area, and the acquisition unit is further used to obtain the line of sight direction of the second user in the rear area; the control unit is further used to control the display device in the rear area to display a first image and obtain the body posture information of the second user when the line of sight of the second user is on the physical virtual assistant, the first image including the image of the physical virtual assistant; the control unit is further used to control the display device in the rear area to display a second image based on the body posture information of the second user, the second image including an image of the virtual assistant responding to the body posture of the second user.
[0070] In combination with the second aspect, in certain implementations of the second aspect, the virtual assistant is displayed through the first display device, the first display device is associated with the first area where the first user is located, and the acquisition unit is further used to obtain the line of sight direction of the third user in the second area; the control unit is further used to switch from controlling the first display device to display the virtual assistant to controlling the second display device to display the virtual assistant when the line of sight of the third user is on the virtual assistant, and the second display device is associated with the second area; the acquisition unit is also used to obtain the human body posture information of the third user; the control unit is also used to control the second display device to display a third image based on the human body posture information of the third user, and the third image includes an image of the virtual assistant responding to the human body posture of the third user.
[0071] In combination with the second aspect, in certain implementations of the second aspect, the virtual assistant is displayed through the first display device, and the material of the virtual assistant is an irreversible material. The control unit is used to control the first display device to display the first interactive image of the virtual assistant before controlling the virtual assistant to interact with the first user; control the first display device to display the second interactive image of the virtual assistant according to the human posture information of the first user; and keep controlling the first display device to display the second interactive image when the interaction between the virtual assistant and the first user ends.
[0072] In combination with the second aspect, in some implementations of the second aspect, the acquisition unit is used to: acquire the sight direction of the first user when the virtual assistant is in an awake state.
[0073] In a third aspect, a control device is provided, which includes a processing unit and a storage unit, wherein the storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to enable the control device to perform any possible method in the first aspect.
[0074] In a fourth aspect, a control system is provided, which includes a physical virtual assistant and a computing platform, and the computing platform includes any possible device in the second aspect or the third aspect above.
[0075] In a fifth aspect, a control system is provided, which includes a display device and a computing platform, wherein the display device is used to display a virtual assistant, and the computing platform includes any possible device in the second aspect or the third aspect mentioned above.
[0076] In a sixth aspect, a vehicle is provided, which includes any possible device in the second aspect, or includes the device described in the third aspect, or includes the system described in the fourth aspect, or includes the system described in the fifth aspect.
[0077] In some possible implementations, the vehicle is a vehicle.
[0078] In a seventh aspect, a computer program product is provided, comprising: a computer program code, which, when executed on a computer, enables the computer to execute any possible method in the first aspect.
[0079] It should be noted that the above-mentioned computer program code can be stored in whole or in part on the first storage medium, wherein the first storage medium can be packaged together with the processor or separately packaged with the processor, and the embodiments of the present application do not specifically limit this.
[0080] In an eighth aspect, a computer-readable medium is provided, wherein the computer-readable medium stores a program code, and when the computer program code is run on a computer, the computer is enabled to execute any possible method in the first aspect.
[0081] In the ninth aspect, an embodiment of the present application provides a chip system, which includes a processor for calling a computer program or computer instructions stored in a memory so that the processor executes any possible method in the above-mentioned first aspect.
[0082] In combination with the ninth aspect, in one possible implementation, the processor is coupled to the memory through an interface.
[0083] In combination with the ninth aspect, in one possible implementation, the chip system also includes a memory, in which a computer program or computer instructions are stored.
[0084] In an embodiment of the present application, responding to the user's posture by the virtual assistant helps to improve the user's human-computer interaction experience when interacting with the virtual assistant; at the same time, obtaining the user's posture information when the user's line of sight is on the virtual assistant helps to save the computing overhead of the vehicle, thereby helping to reduce the power consumption of the vehicle.
[0085] Users can set the mapping relationship between their desired gesture postures and the interactive actions of the virtual assistant, which helps to improve the user's human-computer interaction experience and also helps to improve the flexibility in setting the mapping relationship between gesture postures and the interactive actions of the virtual assistant.
[0086] Virtual pets can be created from photos or videos of pets. This helps further enhance the user experience when interacting with the virtual pet. Furthermore, users can select the material of the virtual assistant or change the appearance of the virtual assistant to their desired image, which increases the user's flexibility in setting up the virtual assistant and thus improves the user experience.
[0087] By controlling at least one of the user's line of sight direction, the throwing angle corresponding to the throwing gesture, and the throwing speed corresponding to the throwing gesture, the virtual assistant is controlled to respond to the throwing gesture, so that the virtual assistant's feedback on the throwing gesture is more in line with the user's expectations, which helps to improve the user's human-computer interaction experience.
[0088] When multiple users want to interact with the virtual assistant, by not responding to the gestures of the user in the main driving area, the attention of the user in the main driving area can be diverted to driving the vehicle, thereby helping to improve the safety of the user driving the vehicle.
[0089] When multiple users want to interact with the virtual assistant, they can control the virtual assistant to interact with users who have entered user information in the vehicle based on the user information saved in the vehicle, which helps to avoid the virtual assistant's interactive actions interfering with strangers when they are looking at the virtual assistant.
[0090] During the playback of multimedia files, the virtual assistant can be controlled to imitate the first user's posture, which can create a feeling of dancing with the virtual assistant for the user, helping to enhance the user's human-computer interaction experience.
[0091] When a second user in the back row wishes to interact with the virtual assistant in the front row, the virtual assistant's image information can be displayed on the display device in the back row. This can avoid the feeling of alienation between the back row user and the virtual assistant due to the long distance, and help improve the human-computer interaction experience when the back row user interacts with the virtual assistant. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] FIG1 is a functional block diagram of a vehicle provided in an embodiment of the present application.
[0093] FIG2 is a schematic diagram of the distribution of display screens in a vehicle cabin according to an embodiment of the present application.
[0094] FIG3 is a set of GUIs provided in an embodiment of the present application.
[0095] FIG4 is another set of GUIs provided in an embodiment of the present application.
[0096] FIG5 is another set of GUIs provided in an embodiment of the present application.
[0097] FIG6 is another set of GUIs provided in an embodiment of the present application.
[0098] FIG7 is another set of GUIs provided in an embodiment of the present application.
[0099] FIG8 is a schematic diagram of estimating the user's line of sight provided in an embodiment of the present application.
[0100] FIG9 is a schematic diagram of estimating a user's gesture posture provided by an embodiment of the application.
[0101] FIG10 is a schematic flowchart of the human-computer interaction method provided in an embodiment of the present application.
[0102] FIG11 is another set of GUIs provided by an embodiment of the present application.
[0103] FIG12 is a schematic block diagram of a human-computer interaction device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0104] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is only a way to describe the association relationship of associated objects, indicating that there can be three kinds of relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. "At least one" means one or more. For example, "at least one of A and B" is similar to "A and / or B", describing the association relationship of associated objects, indicating that there can be three kinds of relationships, for example, at least one of A and B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0105] In the embodiments of the present application, prefixes such as "first" and "second" are used only to distinguish different description objects and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of prefixes such as ordinal numbers to distinguish description objects in the embodiments of the present application does not constitute a restriction on the described objects. For the statement of the described objects, please refer to the description in the context of the claims or embodiments, and the use of such prefixes should not constitute an unnecessary restriction. In addition, in the description of this embodiment, unless otherwise specified, the meaning of "plurality" is two or more.
[0106] Figure 1 is a functional block diagram of a vehicle 100 provided in an embodiment of the present application. The vehicle 100 may include a perception system 120, a display device 130, and a computing platform 150, wherein the perception system 120 may include one or more sensors for sensing information about the environment surrounding the vehicle 100. For example, the perception system 120 may include a positioning system, which may be a global positioning system (GPS), a BeiDou system, or other positioning systems. The perception system 120 may also include one or more of an inertial measurement unit (IMU), a laser radar, a millimeter-wave radar, an ultrasonic radar, and a camera device.
[0107] Some or all functions of the vehicle 100 can be controlled by the computing platform 150. The computing platform 150 may include one or more processors, such as processors 151 to 15n (n is a positive integer). A processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationship of a hardware circuit. The logical relationship of the hardware circuit is fixed or reconfigurable. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as a field programmable gate array (FPGA). In a reconfigurable hardware circuit, the process of the processor loading a configuration file to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, the processor may also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc. In addition, the computing platform 150 may also include a memory for storing instructions. Some or all of the processors 151 to 15n may call the instructions in the memory and execute the instructions to implement corresponding functions.
[0108] The display devices 130 in the cockpit are mainly divided into two categories: the first is the vehicle-mounted display screen; the second is a projection display screen, such as a head-up display (HUD). The vehicle-mounted display screen is a physical display screen and a key component of the in-vehicle infotainment system. The cockpit can be equipped with multiple displays, such as the digital instrument panel, the central control screen, the display in front of the front passenger (also known as the front passenger), the display in front of the left rear passenger, and the display in front of the right rear passenger. Even the vehicle windows can serve as display screens. A head-up display, also known as a head-up display system, is primarily used to display driving information such as speed and navigation on a display device in front of the driver (such as the windshield). This reduces the driver's gaze shift time, avoids pupil changes caused by the driver's gaze shift, and improves driving safety and comfort. HUDs include, for example, combiner-HUD (C-HUD), windshield-HUD (W-HUD), and augmented reality HUD (AR-HUD). It should be understood that other types of HUD systems may appear as technology evolves, and this application is not limited to this.
[0109] The above display device 130 is described by taking a vehicle-mounted display screen and a projection display screen as examples, and the embodiments of the present application are not limited thereto. For example, the display device 130 can also be a light display screen or a projection screen.
[0110] Figure 2 shows a schematic diagram of an exemplary display screen layout within a vehicle cabin, as provided in an embodiment of the present application. As shown in Figure 2 , the vehicle cabin may include display screen 201 (or, alternatively, a central control screen), display screen 202 (or, alternatively, a passenger entertainment screen), display screen 203 (or, alternatively, a second-row left entertainment screen), display screen 204 (or, alternatively, a second-row right entertainment screen), and an instrument panel.
[0111] It should be understood that the graphical user interface (GUI) in the following embodiments is described using the five-seater vehicle shown in FIG2 as an example, and the embodiments of the present application are not limited thereto. For example, for a seven-seater sport utility vehicle (SUV), the cockpit may include a central control screen, a co-pilot entertainment screen, an entertainment screen in the second row left area, an entertainment screen in the second row right area, an entertainment screen in the third row left area, and an entertainment screen in the third row right area. For another example, for a passenger car, the cockpit may include a front row entertainment screen and a rear row entertainment screen; or, the cockpit may include a display screen in the driving area and an entertainment screen in the passenger area. In one implementation, the entertainment screen in the passenger area may also be set on the top of the cockpit.
[0112] Exemplarily, FIG3 shows a set of GUIs provided in an embodiment of the present application.
[0113] As shown in (a) of FIG. 3 , the vehicle may display a virtual assistant (eg, a virtual pet Xiao A) through the display screen 201 .
[0114] In one embodiment, when it is detected that user 1 in the main driving area issues a voice wake-up command "Xiao A Xiao A", the vehicle can display the virtual pet Xiao A through the display screen 201.
[0115] In one embodiment, the virtual pet Xiao A may be set up when the vehicle leaves the factory; or, the virtual assistant Xiao A may be selected by user 1 from multiple virtual assistants. For example, the vehicle includes multiple virtual assistants, and when the operation of user 1 selecting virtual pet Xiao A from multiple virtual assistants is detected, an association relationship between virtual pet Xiao A and user 1 may be established. When it is detected that user 1 issues a voice wake-up command, the area where user 1 is located may be determined through image information collected by a camera in the cabin (for example, a camera of a driver monitor system (DMS) or a camera of a cabin monitor system (CMS)). For example, when it is detected that the area where user 1 is located is the main driving area, the virtual pet Xiao A may be displayed on the display screen 201 based on the association relationship.
[0116] In one embodiment, the virtual pet Xiao A can also be generated from image data. For example, when a user uploads a favorite pet image (e.g., a photo or video of a pet) to the vehicle and chooses to generate a corresponding three-dimensional (3D) model of the virtual pet, the vehicle can generate a 3D model of the virtual pet based on the pet image. At the same time, the user can also choose to name the 3D model of the virtual pet (e.g., Xiao A). In this way, when the user detects the voice wake-up command "Xiao A Xiao A", the vehicle can display the virtual pet Xiao A through the display screen 201.
[0117] As shown in (b) and (c) of FIG3 , when the face pinching gesture of user 1 is detected, the vehicle can display an animation effect of the face of the virtual pet Xiao A being pinched through the display screen 201 .
[0118] In one embodiment, when a face-pinching gesture of user 1 is detected, the vehicle can display an animation effect of the face of the virtual pet Xiao A being pinched through the display screen 201, including: when it is detected that the line of sight of user 1 is on the virtual pet Xiao A and a face-pinching gesture of user 1 is detected, the vehicle can display an animation effect of the face of the virtual pet Xiao A being pinched through the display screen 201.
[0119] In one embodiment, when it is detected that the direction of the user 1's gaze is on the virtual pet Xiao A and the user 1's face-pinching gesture is detected, the vehicle can display the animation effect of the virtual pet Xiao A's face being pinched through the display screen 201, including: when it is detected that the direction of the user 1's gaze is on the virtual pet Xiao A for a duration greater than or equal to a preset duration and the user 1's face-pinching gesture is detected, the vehicle can display the animation effect of the virtual pet Xiao A's face being pinched through the display screen 201.
[0120] Exemplarily, the preset duration is 1 second.
[0121] In one embodiment, upon detecting a pinching gesture by user 1, the vehicle may also respond to the pinching gesture by sending a voice signal through the virtual pet Xiao A. For example, the vehicle may send a voice message "Master, what's up?" through the virtual pet Xiao A.
[0122] In an embodiment of the present application, when a pinching gesture is detected, the virtual pet can respond to the pinching gesture by pinching its face. The interaction through body movements increases the sense of interaction between the user and the virtual assistant, helping to enhance the user's human-computer interaction experience.
[0123] Exemplarily, FIG4 shows a set of GUIs provided in an embodiment of the present application.
[0124] As shown in (a) of FIG. 4 , the vehicle may display a virtual assistant (eg, a virtual pet Xiao A) through the display screen 201 .
[0125] In one embodiment, when it is detected that user 1 in the main driving area issues a voice wake-up command "Xiao A Xiao A", the vehicle can display the virtual pet Xiao A through the display screen 201.
[0126] As shown in (b) and (c) of FIG. 4 , when a stroking gesture of user 1 is detected, the vehicle may display an animation effect of the virtual pet Little A enjoying the stroking through the display screen 201 .
[0127] In one embodiment, when a stroking gesture of user 1 is detected, the vehicle can display an animation effect of the virtual pet Xiao A enjoying the stroking through the display screen 201, including: when it is detected that the direction of user 1's sight is on the virtual pet Xiao A and the stroking gesture of user 1 is detected, the vehicle can display an animation effect of the virtual pet Xiao A enjoying the stroking through the display screen 201.
[0128] In one embodiment, when it is detected that the direction of the user 1's gaze is on the display screen 201 and the user 1's stroking gesture is detected, the vehicle can display the animation effect of the virtual pet Xiao A enjoying the stroking through the display screen 201, including: when it is detected that the direction of the user 1's gaze is on the virtual pet Xiao A for a duration greater than or equal to a preset duration and the user 1's stroking gesture is detected, the vehicle can display the animation effect of the virtual pet Xiao A enjoying the stroking through the display screen 201.
[0129] In one embodiment, when a stroking gesture of user 1 is detected, the vehicle may also send a voice signal through the virtual pet Xiao A to respond to the stroking gesture of user 1. For example, the vehicle may send a voice message "Master, I feel so comfortable" through the virtual pet Xiao A.
[0130] In the embodiment of the present application, when a user's stroking gesture is detected, the virtual pet can respond to the stroking gesture by enjoying the stroking. The interaction through physical movements increases the sense of interaction between the user and the virtual assistant, which helps to enhance the user's human-computer interaction experience.
[0131] Exemplarily, FIG5 shows a set of GUIs provided in an embodiment of the present application.
[0132] As shown in (a) of FIG. 5 , the vehicle may display a virtual assistant (eg, a virtual pet Little A) through the display screen 201 , where the display area where the virtual assistant is located is area 501 .
[0133] In one embodiment, when it is detected that user 1 in the main driving area issues a voice wake-up command "Xiao A Xiao A", the vehicle can display the virtual pet Xiao A through the display screen 201.
[0134] As shown in (b) and (c) of FIG. 5 , when detecting that the user 1 has made a summoning gesture, the vehicle may display an enlarged virtual pet Xiao A through the display screen 201 .
[0135] In one embodiment, when it is detected that the user 1 makes a summoning gesture, the vehicle can display an animation effect of the virtual pet Xiao A running towards the direction of the user through the display screen 201.
[0136] In one embodiment, the summoning gesture may be a gesture with the palm facing upward and four fingers waving toward the side of the user's body.
[0137] For example, when user 1 is detected to have made a summoning gesture, the virtual pet A can be displayed in area 502, where the size of area 502 is larger than that of area 501. In this way, by enlarging the size of the virtual assistant display area, the user can feel that the virtual pet is close to the user.
[0138] In one embodiment, when it is detected that user 1 makes a summoning gesture, the vehicle can display an enlarged virtual pet Xiao A through the display screen 201, including: when it is detected that the direction of user 1's line of sight is on the virtual pet Xiao A and it is detected that user 1 makes a summoning gesture, the vehicle can display an enlarged virtual pet Xiao A through the display screen 201.
[0139] In one embodiment, when it is detected that the direction of the user 1's line of sight is on the virtual pet Xiao A and it is detected that the user 1 makes a summoning gesture, the vehicle can display the enlarged virtual pet Xiao A through the display screen 201, including: when it is detected that the direction of the user 1's line of sight is on the virtual pet Xiao A for a duration greater than or equal to a preset duration and it is detected that the user 1 makes a summoning gesture, the vehicle can display the enlarged virtual pet Xiao A through the display screen 201.
[0140] In one embodiment, when the vehicle detects that user 1 has made a summoning gesture, the vehicle may also send a voice signal through the virtual pet Xiao A to respond to the summoning gesture made by user 1. For example, the vehicle may send a voice message "Master, I'm here" through the virtual pet Xiao A.
[0141] In the embodiment of the present application, when the user makes a summoning gesture, the virtual pet can respond to the gesture by moving closer to the user. The interaction through physical movements increases the sense of interaction between the user and the virtual assistant, which helps to enhance the user's human-computer interaction experience.
[0142] Exemplarily, FIG6 shows a set of GUIs provided in an embodiment of the present application.
[0143] As shown in (a) of FIG. 6 , the vehicle may display a virtual assistant (eg, a virtual pet Xiao A) through the display screen 201 .
[0144] As shown in (b) and (c) of FIG6 , when the throwing gesture (or the hair ball throwing gesture) of the user 1 is detected, the vehicle can display the animation effect of the virtual pet Xiao A picking up the hair ball through the display screen 201 .
[0145] In one embodiment, when a throwing gesture of user 1 is detected, the vehicle can display an animation effect of virtual pet Xiao A picking up a hair ball through the display screen 201, including: when it is detected that the direction of user 1's line of sight is on the virtual pet Xiao A and a throwing gesture of user 1 is detected, the vehicle can display an animation effect of virtual pet Xiao A picking up a hair ball through the display screen 201.
[0146] In one embodiment, when it is detected that the duration of the user 1's line of sight on the virtual pet Xiao A is greater than or equal to a preset duration and the throwing gesture of the user 1 is detected, the vehicle can display the animation effect of the virtual pet Xiao A picking up the ball of hair through the display screen 201.
[0147] In one embodiment, the vehicle displays the animation effect of the virtual pet Xiao A picking up a ball of fur through the display screen 201, including: displaying the animation effect of the virtual pet Xiao A looking for the ball of fur based on at least one of the user's line of sight direction, the throwing angle corresponding to the throwing gesture, and the throwing speed corresponding to the throwing gesture.
[0148] For example, when it is detected that user 1 throws a wool ball toward the right front of the vehicle, an animation effect of the virtual pet Xiao A picking up the wool ball toward the right front of the vehicle may be displayed.
[0149] For another example, when it is detected that user 1 throws a wool ball toward the left front of the vehicle, an animation effect of the virtual pet Xiao A picking up the wool ball toward the left front of the vehicle may be displayed.
[0150] For another example, when it is detected that the direction of the user 1's sight is in the right front of the vehicle and the throwing gesture of the user 1 is detected, an animation effect of the virtual pet Xiao A picking up the ball of hair towards the right front of the vehicle can be displayed.
[0151] For another example, when it is detected that the throwing speed corresponding to the throwing gesture of user 1 is greater than or equal to the preset speed, an animation effect of the virtual pet picking up the ball of hair at speed 1 can be displayed; or, when it is detected that the throwing speed corresponding to the throwing gesture of user 1 is less than the preset speed, an animation effect of the virtual pet picking up the ball of hair at speed 2 can be displayed, where speed 1 is greater than speed 2.
[0152] In one embodiment, when it is detected that the line of sight of user 1 in the main driver's seat is not directed at the virtual pet Xiao A and the line of sight of user 2 in the passenger seat is directed at the virtual pet Xiao A, an animation effect of the virtual pet Xiao A looking for a fur ball can be displayed through the display screen 201 according to the throwing gesture of user 2.
[0153] In one embodiment, when it is detected that the gaze direction of user 1 in the main driver's seat is not directed at virtual pet Xiao A and the gaze direction of user 2 in the passenger seat is directed at virtual pet Xiao A, virtual pet Xiao A can be switched to be displayed on display screen 202. In response to the throwing gesture of user 2, the vehicle can display an animation effect of virtual pet Xiao A searching for a fur ball on display screen 202.
[0154] In one embodiment, the front area of the vehicle may include a long screen (or, it may also be referred to as a long continuous screen). For example, the central control screen and the co-pilot screen in the vehicle cockpit may be the same screen. The screen may be divided into two display areas, namely Area 1 and Area 2, where Area 1 may be the display area near the main driver's seat, and Area 2 may be the display area near the co-pilot's seat. When the user in the main driver's seat looks at the virtual assistant, an animation of the virtual assistant responding to the gesture of the user in the main driver's seat may be displayed in Area 1 of the screen.
[0155] When the user in the driver's seat isn't looking at the virtual assistant, and the user in the passenger seat is looking at it, the screen can display an animation of the virtual assistant moving from area 1 to area 2. This can increase the sense of familiarity when the user in the passenger seat interacts with the virtual assistant, and also help improve the vehicle's intelligence.
[0156] After the virtual assistant moves from area 1 to area 2, the vehicle can control the virtual assistant to respond to the posture of the user in the passenger area based on the posture information of the user in the passenger area.
[0157] In the embodiment of the present application, when a throwing gesture of the user is detected, the virtual pet can respond to the throwing gesture by picking up the hairball. The interaction through body movements increases the sense of interaction between the user and the virtual assistant, which helps to enhance the user's human-computer interaction experience.
[0158] In one embodiment, when it is detected that a multimedia device in the cockpit is playing a multimedia file (e.g., music or video) and a change in the user's posture (e.g., body dancing) is detected, an animation effect of the virtual pet Xiao A imitating the user's posture can be displayed.
[0159] In one embodiment, the vehicle may display an animation effect of the virtual pet Little A imitating the user's posture through a display screen corresponding to the area where the user is located.
[0160] For example, when the user 1 in the main driving area dances to the music, the display screen 201 may display an animation effect of the virtual pet Little A imitating the posture of the user in the main driving area.
[0161] For example, when user 2 in the passenger seat dances to music, an animation effect of virtual pet Xiao A imitating the user's posture in the passenger seat can be displayed on display screen 202. Alternatively, if user 2 has an association with virtual assistant Xiao B, an animation effect of virtual assistant Xiao B imitating the user's posture can be displayed on display screen 202.
[0162] Exemplarily, FIG7 shows a set of GUIs provided in an embodiment of the present application.
[0163] For example, as shown in (a) and (b) of FIG7 , when it is detected that user 3 in the right area of the second row issues a voice command “Xiao A Xiao A”, the display screen 204 may be used to display the virtual pet Xiao A.
[0164] In one embodiment, when a pinching gesture, a stroking gesture, a summoning gesture, or a throwing gesture of the user 3 is detected, the display screen 204 may display the interactive action of the virtual pet Little A in response to the corresponding gesture.
[0165] The above Figures 3 to 7 illustrate a virtual assistant displayed on the screen as an example, but the embodiments of the present application are not limited thereto. For example, the virtual assistant can also be a physical virtual assistant. For another example, the virtual assistant can also be displayed through a HUD.
[0166] The above describes several GUIs provided by the embodiments of the present application in conjunction with Figures 3 to 7. The following describes the process of estimating the user's gaze direction and the user's gesture posture provided by the embodiments of the present application in conjunction with Figures 8 and 9.
[0167] Figure 8 shows a schematic diagram of estimating the user's line of sight direction provided by an embodiment of the present application. When a user is detected in the main driving area of the vehicle cabin, the image data 1 of the main driving area collected by a time of flight (TOF) camera and the image data 2 of the main driving area collected by a red, green, blue (RGB) camera can be input into a neural network (NN) 1 to output the user's eye frame and eye position, wherein the image data 1 and the image data 2 include the user's face image. The neural network here can use a convolutional network such as Convolutional Networks. By inputting the eye frame and eye position into the line of sight direction estimation module, the line of sight direction of the user in the main driving area can be estimated.
[0168] The above NN1 can be trained using labeled data. For example, a face image 1 captured by a TOF camera and an RGB camera with the eye frame set to eye frame 1 and the eye position set to eye position 1 can be collected. This face image 1 can then be labeled with eye frame 1 and eye position 1. This labeled face image 1 can then be used as labeled data. Another example is a face image 2 captured by a TOF camera and an RGB camera with the eye frame set to eye frame 2 and the eye position set to eye position 2. This labeled face image 2 can then be used as labeled data. This NN1 can be trained using labeled data.
[0169] Figure 9 shows a schematic diagram of estimating user gesture posture provided by an embodiment of the present application. As shown in Figure 9, when it is detected that there is a user in the main driving area in the vehicle cabin, the image data 3 of the main driving area collected by the RGB camera can be input into NN2, so that the 2-dimensional (2-dimension, 2D) result of the human hand and the human hand model can be output. The 2.5-dimensional (2.5-dimension, 2.5D) result of the key points of the human hand can be determined by the human hand model. According to the 2.5D results of the key points of the human hand and the 2D results of the human hand, the 3D key point position of the entire human hand can be obtained by the perspective-n-point (PNP) algorithm. By inputting the 3D key point position into the gesture posture estimation module, the gesture posture of the user in the main driving area can be obtained.
[0170] Similarly, the user's body posture can be determined by referring to the process shown in Figure 9 above. For example, image data 4 of the main driving area, captured by an RGB camera, can be input into NN3, which can output a 2D human body result and a human body model. The human body model can be used to determine the 2.5D results of the human body's key points. Based on the 2.5D results of the human body's key points and the 2D results of the human body, the PNP algorithm can be used to determine the 3D key point positions of the entire human body. By inputting the 3D key point positions of the entire human body into the human posture estimation module, the human body posture of the user in the main driving area can be determined.
[0171] It should be understood that the processes of estimating the line of sight direction and estimating the user's gesture posture shown in Figures 8 and 9 above are merely illustrative, and the embodiments of the present application are not limited thereto. The line of sight direction, user gesture posture, and user body posture can also be determined in other ways. For example, the user's line of sight direction can also be determined based on a line of sight tracking algorithm based on a three-dimensional model or a line of sight tracking algorithm based on a two-dimensional appearance, wherein the line of sight tracking algorithm based on a three-dimensional model can include a pupil (iris)-eye canthus method and a pupil (iris)-corneal reflection method. For another example, the user's gesture posture can also be estimated by a convolutional pose machine (CPM) algorithm.
[0172] Figure 10 shows a schematic flow chart of a human-computer interaction method 1000 provided in an embodiment of the present application. The method 1000 can be performed by a vehicle (e.g., a vehicle); or, the method 1000 can be performed by the above-mentioned computing platform; or, the method 1000 can be performed by a system consisting of a computing platform and a display device; or, the method 1000 can be performed by a system consisting of a computing platform and a physical virtual device; or, the method 1000 can be performed by a system-on-a-chip (SoC) in the above-mentioned computing platform; or, the method 1000 can be performed by a processor in the computing platform. The method 1000 is applied to a cockpit of a vehicle, which includes a virtual assistant, and the method 1000 includes:
[0173] S1010: Obtain the sight direction of the first user in the cabin.
[0174] Exemplarily, the sight direction of the first user in the cabin can be determined based on image data collected by the TOF camera and the RGB camera in the cabin.
[0175] Optionally, obtaining the sight line direction of the first user in the cockpit includes: obtaining the sight line direction of the first user when the virtual assistant is in an awake state.
[0176] S1020: When the first user's line of sight is directed toward the virtual assistant, obtain posture information of the first user.
[0177] Optionally, the posture information of the first user includes hand gestures and / or body postures of the first user.
[0178] Optionally, the virtual assistant may be a virtual pet. The virtual pet may be factory-set when the vehicle leaves the factory; or, the vehicle may be factory-set with multiple virtual pets, from which the user can select a favorite virtual pet.
[0179] Optionally, method 1000 further includes: obtaining image information of the pet; and generating the virtual pet based on the image information of the pet. In this way, a user can upload photos or videos of their pet to the vehicle, and the vehicle can then generate a virtual pet based on the photos or videos of the pet. When the user interacts with the virtual pet, this helps enhance the user's sense of familiarity with the virtual pet, thereby improving the user's human-computer interaction experience.
[0180] S1030: Control the virtual assistant to interact with the first user according to the posture information of the first user.
[0181] Optionally, the posture information of the first user includes a first hand gesture, and controlling the virtual assistant to interact with the first user according to the posture information of the first user includes: controlling the virtual assistant to interact with the first user according to the first hand gesture.
[0182] For example, the first gesture posture includes at least one of a face pinching gesture, a stroking gesture, a summoning gesture, and a throwing gesture.
[0183] Optionally, the vehicle stores a mapping relationship between gesture postures and interactive actions, and controls the virtual assistant to interact with the first user according to the first gesture posture, including: controlling the virtual assistant to interact with the first user through a first interactive action according to the first gesture posture and the mapping relationship.
[0184] For example, Table 1 shows a mapping relationship between hand gestures and interactive actions of a virtual assistant.
[0185] Table 1
[0186] Gestures, gestures, interactive actions, face pinching gesture, face being pinched, caressing gesture, enjoying the caress, summoning gesture, virtual assistant approaching the user, throwing gesture, picking up a wool ball...
[0187] The mapping relationship between the above gestures and interactive actions is merely illustrative and is not specifically limited in the embodiments of the present application.
[0188] In some possible implementations, before controlling the virtual assistant to interact with the first user according to the first gesture, the method 1000 further includes: detecting an operation in which the user sets a mapping relationship between the first gesture and the first interaction action.
[0189] Exemplarily, FIG11 shows another set of GUIs provided by an embodiment of the present application.
[0190] As shown in FIG11(a), the vehicle displays a desktop on a display screen. The desktop includes a card 1101, which displays icons for setting functions. When the vehicle detects that a user clicks on card 1101, the display screen may display a GUI as shown in FIG11(b).
[0191] As shown in FIG11(b), in response to detecting a user clicking on card 1101, the vehicle may display a display interface for settings functions on the display screen, where the settings functions include language and date, display function, sports mode, battery, and update functions. When the vehicle detects a user clicking on a display function, the vehicle may display a GUI as shown in FIG11(c).
[0192] As shown in FIG11(c), in response to detecting a user clicking a display function, the vehicle may display a display interface for the display function on the display screen, wherein the display interface includes settings for screen brightness, function bars, cards, and a virtual assistant. When the vehicle detects a user clicking control 1102, the GUI shown in FIG11(d) may be displayed.
[0193] As shown in (d) of Figure 11, the display interface is the settings interface of the virtual assistant. The settings interface includes settings for the appearance, sound, and correspondence between air gestures and interactive actions of the virtual assistant. When a user clicks on control 1103, a GUI as shown in (e) of Figure 11 can be displayed.
[0194] As shown in (e) in Figure 11, in response to detecting the user clicking on the control 1103, the vehicle can display a setting interface for the correspondence between air gestures and interactive actions through the display screen. On this setting interface, the user can set the correspondence between air gestures and interactive actions of the virtual assistant. For example, when the vehicle detects the user clicking in the rectangular frame 1103, the correspondence between picking up a hairball and the throwing gesture can be set. In this way, as shown in (b) and (c) in Figure 6, when it is detected that the user is looking at the display screen and making a throwing gesture, the vehicle can display the animation effect of the virtual pet picking up a hairball through the display screen.
[0195] In an embodiment of the present application, the user can set the mapping relationship between the desired gesture posture and the interactive action of the virtual assistant in the vehicle, so that when the user's first gesture posture is detected, the virtual assistant can be controlled to interact with the first user through the corresponding interactive action according to the mapping relationship between the gesture posture set by the user and the interactive action of the virtual assistant. While improving the user's human-computer interaction experience, it also helps to improve the flexibility in setting the mapping relationship between gesture posture and the interactive action of the virtual assistant.
[0196] Optionally, the posture information includes a throwing gesture, and controlling the virtual assistant to interact with the first user includes: controlling the virtual assistant to respond to the throwing gesture based on at least one of the first user's line of sight direction, the throwing angle corresponding to the throwing gesture, and the throwing speed corresponding to the throwing gesture.
[0197] Optionally, controlling the virtual assistant to respond to the throwing gesture includes: controlling the virtual assistant to perform an operation of picking up the ball, or displaying an animation effect of the virtual assistant picking up the ball through a display screen.
[0198] Optionally, when the throwing speed corresponding to the throwing gesture is greater than or equal to a preset speed threshold, the speed of the virtual assistant when picking up the ball can be controlled to be a first speed; or, when the throwing speed corresponding to the throwing gesture is less than the preset speed threshold, the speed of the virtual assistant when picking up the ball can be controlled to be a second speed, wherein the first speed is greater than the second speed.
[0199] Exemplarily, the preset speed threshold is 1 meter per second (m / s).
[0200] For example, the process for determining the throwing speed corresponding to the above throwing gesture can be as follows: at time T1, the user's throwing gesture is detected. At this time, coordinate 1 of a certain point in the throwing gesture (for example, a point on the thumb) in the vehicle coordinate system can be determined. At time T2, coordinate 2 of this point in the vehicle coordinate system can be determined. Based on coordinates 1 and 2, the distance moved by the point from time T1 to time T2 can be determined. Based on this distance and the duration between time T1 and time T2, the throwing speed corresponding to the throwing gesture can be determined.
[0201] Optionally, the vehicle is a vehicle, which includes a main driving area and a non-main driving area, and the first user is a user located in the main driving area. The method 1000 also includes: when the line of sight of the first user and the line of sight of the third user are both on the virtual assistant, according to the vehicle being in a driving state, controlling the virtual assistant to interact with the third user, and the third user is a user located in the non-main driving area.
[0202] In this way, if it is detected that both the driver's seat user and the non-driver's seat user are looking at the virtual assistant and the vehicle is currently in motion, the virtual assistant can be controlled to interact with the third user. This prevents distraction to the driver's seat caused by the virtual assistant interacting with the driver's seat user. Furthermore, by not responding to the driver's seat user's gestures, the driver's seat user's attention can be diverted to driving the vehicle, thereby helping to improve the user's driving safety.
[0203] Optionally, the method 1000 further includes: prompting the first user to concentrate on driving the vehicle.
[0204] Optionally, the method 1000 further includes: when the line of sight of the first user and the line of sight of the fourth user are both on the virtual assistant and the information of the first user is saved in the vehicle but the information of the fourth user is not saved, controlling the virtual assistant to interact with the first user.
[0205] In this way, when multiple users wish to interact with the virtual assistant, the virtual assistant can be controlled to interact with the first user based on the user information stored in the vehicle. This helps improve the human-computer interaction experience between the virtual assistant and the users who have entered their user information in the vehicle, and also helps avoid the virtual assistant's interactive actions disturbing strangers when they are looking at the virtual assistant.
[0206] The above user information may include the user's biometric information, such as face information, iris information, etc.
[0207] Optionally, the posture information includes human body posture, and controlling the virtual assistant to interact with the first user according to the posture information of the first user includes: controlling the virtual assistant to interact with the first user according to the human body posture.
[0208] Optionally, the method 1000 further includes: controlling the multimedia device in the vehicle to play a multimedia file; wherein, based on the posture information of the first user, controlling the virtual assistant to interact with the first user includes: during the playback of the multimedia file, controlling the virtual assistant to imitate the posture of the first user.
[0209] In this way, during the playback of multimedia files, the virtual assistant can imitate the first user's posture, which can create a feeling of dancing with the virtual assistant for the user, helping to enhance the user's human-computer interaction experience.
[0210] For example, the multimedia file may include music or video.
[0211] Optionally, when the first user's line of sight is on the virtual assistant, the posture information of the first user is obtained, including: when the first user's line of sight is on the virtual assistant for a duration greater than or equal to a preset duration, the posture information of the first user is obtained.
[0212] In this way, when the first user's line of sight on the virtual assistant lasts longer than or equal to the preset time, obtaining the first user's posture information helps to further improve the accuracy of the virtual assistant's response to the user's posture.
[0213] Exemplarily, the preset duration is 1 second.
[0214] Optionally, the virtual assistant is a physical virtual assistant.
[0215] Optionally, the vehicle is a vehicle including a front area and a rear area, and the physical virtual assistant is located in the front area of the vehicle.
[0216] Optionally, the method 1000 also includes: obtaining the sight line direction of the second user in the back row area; when the sight line direction of the second user is on the physical virtual assistant, controlling the display device in the back row area to display a first image and obtaining the body posture information of the second user, the first image including the image of the physical virtual assistant; according to the body posture information of the second user, controlling the display device in the back row area to display a second image, the second image including an image of the virtual assistant responding to the body posture of the second user.
[0217] In this way, when a second user in the back row wishes to interact with the physical virtual assistant in the front row, the image information of the physical virtual assistant can be displayed on the display device in the back row, and the display device in the back row can be controlled to display the image of the physical virtual assistant interacting with the second user. This can avoid the sense of alienation between the users in the back row and the physical virtual assistant due to the long distance, and help improve the human-computer interaction experience when the users in the back row interact with the physical virtual assistant.
[0218] Optionally, the virtual assistant is displayed via a display device in the cockpit.
[0219] Optionally, controlling the virtual assistant to interact with the first user includes: controlling a first display device in the cockpit to display a video or animation of the virtual assistant interacting with the first user.
[0220] Optionally, the first display device is associated with an area where the first user is located.
[0221] The first display device is associated with the area where the first user is located, including: the first display device is a display device in the area where the first user is located. For example, if the first user is located in the main driving area, the first display device can be a HUD or the display screen 201 in Figure 2 above.
[0222] Optionally, the first display device is a display device located in an area where the first user is located.
[0223] Optionally, the method 1000 also includes: obtaining the line of sight direction of a fifth user in the cockpit; when the line of sight direction of the fifth user is on the virtual assistant, switching from controlling the first display device to display the virtual assistant to controlling the second display device to display the virtual assistant and obtaining the posture information of the fifth user, the second display device being associated with the area where the fifth user is located; and according to the posture information of the fifth user, controlling the second display device to display an image of the virtual assistant interacting with the fifth user.
[0224] Optionally, the second display device is associated with the area where the fifth user is located, including: the second display device is a display device in the area where the fifth user is located. For example, if the fifth user is a user in the passenger seat area, the second display device may be the display screen 202 in FIG. 2 .
[0225] In this way, when the fifth user in the cabin focuses their gaze on the virtual assistant, the control of displaying the virtual assistant on the first display device can be switched to displaying the virtual assistant on the second display device. This ensures that the virtual assistant is always displayed on the display device closest to the user, increasing the user's sense of familiarity when interacting with the virtual assistant, helping to enhance the user's human-computer interaction experience, and also helping to improve the intelligence of the vehicle.
[0226] Optionally, the method 1000 further includes: when the fifth user's line of sight is directed towards the virtual assistant, prompting the virtual assistant to be displayed on a display device in an area where the fifth user is located.
[0227] Optionally, prompting the virtual assistant to display on a display device in the area where the fifth user is located includes: prompting the virtual assistant to display on the display device in the area where the fifth user is located by emitting a prompt sound through a sound-emitting device.
[0228] Optionally, the sound-emitting device may be located near the second display device. This allows the fifth user to quickly switch their gaze from the virtual assistant on the first display device to the virtual assistant on the second display device, allowing the fifth user to quickly discover that the virtual assistant is displayed on the display device closest to them, thereby enhancing their human-computer interaction experience.
[0229] In the embodiment of the present application, when the user's gaze is directed toward the virtual assistant, the virtual assistant can be controlled to interact with the user based on the user's posture information. This helps enhance the user's human-computer interaction experience when interacting with the virtual assistant by having the virtual assistant respond to the user's posture. Furthermore, obtaining the user's posture information when the user's gaze is directed toward the virtual assistant helps reduce the vehicle's computing overhead, thereby reducing the vehicle's power consumption.
[0230] The embodiments of the present application also provide a human-computer interaction device for implementing any of the above methods. For example, a device is provided that includes units (or means) for implementing each step performed by a vehicle or computer in any of the above methods.
[0231] Figure 12 shows a schematic block diagram of a human-computer interaction device 1200 provided in an embodiment of the present application. As shown in Figure 12, device 1200 includes: an acquisition unit 1210 for acquiring the line of sight of a first user in a vehicle cabin, where a virtual assistant is located; acquisition unit 1210 for acquiring posture information of the first user when the first user's line of sight is directed at the virtual assistant; and a control unit 1220 for controlling the virtual assistant to interact with the first user based on the first user's posture information.
[0232] Optionally, the gesture information includes at least one of a face pinching gesture, a stroking gesture, a summoning gesture, and a throwing gesture.
[0233] Optionally, the posture information includes a throwing gesture, and the control unit 1220 is used to control the virtual assistant to respond to the throwing gesture according to at least one of the line of sight direction of the first user, the throwing angle corresponding to the throwing gesture, and the throwing speed corresponding to the throwing gesture.
[0234] Optionally, the control unit 1220 is used to: control the multimedia device in the vehicle to play a multimedia file; and during the playback of the multimedia file, control the virtual assistant to imitate the posture of the first user.
[0235] Optionally, the acquisition unit 1210 is configured to acquire the posture information of the first user when the duration of the first user's line of sight on the virtual assistant is greater than or equal to a preset duration.
[0236] Optionally, the virtual assistant is a physical virtual assistant, or the virtual assistant is displayed through a display device in the cockpit.
[0237] Optionally, the vehicle includes a front area and a rear area, the virtual assistant is a physical virtual assistant located in the front area, and the acquisition unit 1210 is further used to obtain the line of sight direction of the second user in the rear area; the control unit 1220 is further used to control the display device in the rear area to display a first image and obtain the body posture information of the second user when the line of sight of the second user is on the physical virtual assistant, the first image including the image of the physical virtual assistant; the control unit 1220 is further used to control the display device in the rear area to display a second image based on the body posture information of the second user, the second image including the image of the virtual assistant responding to the body posture of the second user.
[0238] Optionally, the virtual assistant is displayed through the first display device, and the first display device is associated with the first area where the first user is located. The acquisition unit 1210 is also used to obtain the line of sight direction of a third user in the second area; the control unit 1220 is also used to switch from controlling the first display device to display the virtual assistant to controlling the second display device to display the virtual assistant when the line of sight of the third user is on the virtual assistant, and the second display device is associated with the second area; the acquisition unit 1210 is also used to obtain human body posture information of the third user; the control unit 1220 is also used to control the second display device to display a third image based on the human body posture information of the third user, and the third image includes an image of the virtual assistant responding to the human body posture information of the third user.
[0239] Optionally, the virtual assistant is displayed through the first display device, and the material of the virtual assistant is an irreversible material. The control unit 1220 is used to control the first display device to display the first interactive image of the virtual assistant before controlling the virtual assistant to interact with the first user; control the first display device to display the second interactive image of the virtual assistant according to the human posture information of the first user; and keep controlling the first display device to display the second interactive image when the interaction between the virtual assistant and the first user ends.
[0240] Optionally, the acquisition unit 1210 is used to: acquire the sight direction of the first user when the virtual assistant is in an awake state.
[0241] For example, acquisition unit 1210 may be the computing platform shown in FIG1 , or a processing circuit, processor, or controller within the computing platform. For example, if acquisition unit 1210 is processor 151 within the computing platform, processor 151 may acquire pressure data collected by a seat pressure sensor. Processor 151 may determine the presence of a user in the driver's seat area based on this pressure data. Processor 151 may control an in-cabin camera to capture image information of the driver's seat area. Based on this image information, processor 151 may determine the direction of the user's gaze in the driver's seat area.
[0242] For another example, when the processor 151 determines that the user's gaze is directed at the virtual assistant, it may obtain the user's posture information based on the image information. Alternatively, when the processor 151 determines that the user's gaze is directed at the virtual assistant, it may obtain data collected by a radar in the cabin and determine the user's posture information based on the radar data.
[0243] For another example, control unit 1220 may be the computing platform in Figure 1 or a processing circuit, processor, or controller within the computing platform. For example, if control unit 1220 is processor 152 within the computing platform, processor 152 may control the virtual assistant in the cockpit to interact with the user based on the posture information acquired by processor 151.
[0244] The functions implemented by the above-mentioned acquisition unit 1210 and the functions implemented by the control unit 1220 can be implemented by different processors, or can also be implemented by the same processor, which is not limited in the embodiment of the present application.
[0245] It should be understood that the division of the various units in the above device is merely a division of logical functions. In actual implementation, they may be fully or partially integrated into a single physical entity, or they may be physically separated. Furthermore, the units in the device may be implemented in the form of a processor calling software; for example, the device includes a processor connected to a memory storing instructions, and the processor calls the instructions stored in the memory to implement any of the above methods or the functions of the various units in the device, where the processor is, for example, a general-purpose processor such as a CPU or a microprocessor, and the memory is a memory within the device or a memory external to the device. Alternatively, the units in the device may be implemented in the form of hardware circuits, and the functions of some or all of the units may be implemented through the design of the hardware circuits. The hardware circuits may be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units may be implemented through the design of the logical relationships between the components within the circuits. In another implementation, the hardware circuit may be implemented using a PLD, such as an FPGA, which may include a large number of logic gate circuits, and the connections between the logic gate circuits may be configured using a configuration file to implement the functions of some or all of the above units. All units of the above apparatus may be implemented entirely in the form of software called by a processor, or entirely in the form of hardware circuits, or partially in the form of software called by a processor and the rest in the form of hardware circuits.
[0246] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and execute instructions, such as a CPU, a microprocessor, a GPU, or a DSP. In another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit may be fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.
[0247] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0248] In addition, the various units in the above apparatus may be fully or partially integrated together, or may be implemented independently. In one implementation, these units are integrated together and implemented in the form of a system-on-chip (SoC). The SoC may include at least one processor for implementing any of the above methods or implementing the functions of the various units of the apparatus. The at least one processor may be of different types, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0249] An embodiment of the present application also provides a device, which includes a processing unit and a storage unit, wherein the storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit so that the device executes the method or steps performed by the above embodiment.
[0250] Alternatively, if the device is located in a vehicle, the processing unit may be the processors 151 - 15n shown in FIG. 1 .
[0251] An embodiment of the present application further provides a human-computer interaction system, which may include a computing platform and a physical virtual device, and the computing platform may include the above-mentioned human-computer interaction device 1200.
[0252] An embodiment of the present application also provides a human-computer interaction system, which may include a computing platform and a display device. The computing platform may include the above-mentioned human-computer interaction device 1200, and the display device is used to display a virtual assistant.
[0253] Exemplarily, the display device may include an in-vehicle display screen, such as the display device 130 in FIG. 1 ; or one or more of the display screen 201 , the display screen 202 , the display screen 203 , or the display screen 204 in FIG. 2 .
[0254] Optionally, the human-computer interaction system further includes one or more sensors.
[0255] An embodiment of the present application also provides a vehicle, which may include the above-mentioned human-computer interaction device or human-computer interaction system.
[0256] Optionally, the vehicle may be a vehicle.
[0257] An embodiment of the present application further provides a computer program product, which includes: a computer program code, which, when executed on a computer, enables the computer to execute the human-computer interaction method in the above embodiment.
[0258] An embodiment of the present application further provides a computer-readable medium, wherein the computer-readable medium stores a program code. When the computer program code runs on a computer, the computer executes the human-computer interaction method in the above embodiment.
[0259] An embodiment of the present application further provides a chip, which includes a circuit, and the circuit is used to execute the human-computer interaction method in the above embodiment.
[0260] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or a power-on erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.
[0261] It should be understood that in the embodiment of the present application, the memory may include a read-only memory and a random access memory, and provide instructions and data to the processor.
[0262] It should also be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0263] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0264] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0265] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0266] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0267] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0268] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0269] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be covered and fall within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A human-computer interaction method, characterized in that: Applied to a vehicle cockpit, the cockpit including a virtual assistant, the method comprising: Obtaining a sight line direction of a first user in the cabin; When the first user's sight line is directed toward the virtual assistant, acquiring posture information of the first user; Control the virtual assistant to interact with the first user according to the posture information of the first user.
2. The method according to claim 1, wherein The gesture information includes at least one of a face pinching gesture, a stroking gesture, a summoning gesture, and a throwing gesture.
3. The method according to claim 2, wherein The posture information includes a throwing gesture, and controlling the virtual assistant to interact with the first user includes: The virtual assistant is controlled to respond to the throwing gesture according to at least one of the line of sight direction of the first user, the throwing angle corresponding to the throwing gesture, and the throwing speed corresponding to the throwing gesture.
4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: controlling a multimedia device in the vehicle to play multimedia files; The controlling the virtual assistant to interact with the first user according to the posture information of the first user includes: During playback of the multimedia file, the virtual assistant is controlled to imitate the gesture of the first user.
5. The method according to any one of claims 1 to 4, characterized in that When the first user's sight line is directed toward the virtual assistant, acquiring the first user's posture information includes: When the duration of the first user's sight line direction on the virtual assistant is greater than or equal to a preset duration, the posture information of the first user is obtained.
6. The method according to any one of claims 1 to 5, characterized in that The virtual assistant is a physical virtual assistant, or the virtual assistant is displayed through a first display device in the cockpit.
7. The method according to claim 6, wherein The vehicle includes a front row area and a rear row area, the virtual assistant is a physical virtual assistant located in the front row area, and the method further includes: Obtaining the sight line direction of the second user in the rear area; When the second user's line of sight is directed toward the physical virtual assistant, controlling the display device in the rear area to display a first image and acquiring body posture information of the second user, wherein the first image includes an image of the physical virtual assistant; According to the body posture information of the second user, the display device in the rear area is controlled to display a second image, where the second image includes an image of the virtual assistant responding to the body posture of the second user.
8. The method according to claim 6, wherein The virtual assistant is displayed via the first display device, where the first display device is associated with a first area where the first user is located. The method further includes: Obtaining a sight line direction of a third user in the second area; When the third user's line of sight is directed toward the virtual assistant, switching from controlling the first display device to display the virtual assistant to controlling the second display device to display the virtual assistant and obtaining the third user's body posture information, wherein the second display device is associated with the second area; According to the body posture information of the third user, the second display device is controlled to display a third image, where the third image includes an image of the virtual assistant responding to the body posture of the third user.
9. The method according to claim 6, wherein The virtual assistant is displayed through the first display device, the material of the virtual assistant is an irreversible material, and before controlling the virtual assistant to interact with the first user, the method includes: controlling the first display device to display a first interactive image of the virtual assistant; The controlling the virtual assistant to interact with the first user includes: controlling the first display device to display a second interactive image of the virtual assistant; When the virtual assistant ends the interaction with the first user, the first display device is controlled to display the second interactive image.
10. The method according to any one of claims 1 to 9, characterized in that The obtaining of the sight direction of the first user in the cabin includes: When the virtual assistant is in an awake state, the sight direction of the first user is obtained.
11. A human-computer interaction device, characterized in that: include: an acquisition unit, configured to acquire a sight line direction of a first user in a vehicle cabin, wherein the cabin includes a virtual assistant; The acquiring unit is further configured to acquire the posture information of the first user when the sight line direction of the first user is on the virtual assistant; A control unit is used to control the virtual assistant to interact with the first user based on the posture information of the first user.
12. The device according to claim 11, wherein The gesture information includes at least one of a face pinching gesture, a stroking gesture, a summoning gesture, and a throwing gesture.
13. The device according to claim 12, wherein The posture information includes a throwing gesture, and the control unit is configured to: The virtual assistant is controlled to respond to the throwing gesture according to at least one of the line of sight direction of the first user, the throwing angle corresponding to the throwing gesture, and the throwing speed corresponding to the throwing gesture.
14. The device according to any one of claims 11 to 13, characterized in that The control unit is used to: controlling a multimedia device in the vehicle to play multimedia files; During playback of the multimedia file, the virtual assistant is controlled to imitate the gesture of the first user.
15. The device according to any one of claims 11 to 14, characterized in that The acquisition unit is configured to: When the duration of the first user's sight line direction on the virtual assistant is greater than or equal to a preset duration, the posture information of the first user is obtained.
16. The device according to any one of claims 11 to 15, characterized in that The virtual assistant is a physical virtual assistant, or the virtual assistant is displayed through a first display device in the cockpit.
17. The device according to claim 16, wherein The vehicle includes a front row area and a back row area, and the virtual assistant is a physical virtual assistant located in the front row area. The acquisition unit is further configured to acquire a sight line direction of a second user in the rear row area; The control unit is further configured to control the display device in the rear area to display a first image and obtain body posture information of the second user when the second user's line of sight is directed toward the physical virtual assistant, wherein the first image includes an image of the physical virtual assistant; The control unit is further used to control the display device in the rear area to display a second image based on the body posture information of the second user, where the second image includes an image of the virtual assistant responding to the body posture of the second user.
18. The device according to claim 16, wherein The virtual assistant is displayed via the first display device, and the first display device is associated with a first area where the first user is located. The acquisition unit is further configured to acquire a sight line direction of a third user in the second area; The control unit is further configured to switch from controlling the first display device to display the virtual assistant to controlling the second display device to display the virtual assistant when the third user's line of sight is directed toward the virtual assistant, where the second display device is associated with the second area; The acquiring unit is further configured to acquire the body posture information of the third user; The control unit is further configured to control the second display device to display a third image based on the body posture information of the third user, wherein the third image includes an image of the virtual assistant responding to the body posture of the third user.
19. The device according to claim 16, wherein The virtual assistant is displayed via the first display device, and the material of the virtual assistant is irreversible. The control unit is configured to control the first display device to display a first interactive image of the virtual assistant before controlling the virtual assistant to interact with the first user; controlling the first display device to display a second interactive image of the virtual assistant according to the human posture information of the first user; When the virtual assistant finishes interacting with the first user, the first display device is controlled to display the second interactive image.
20. The device according to any one of claims 11 to 19, characterized in that The acquisition unit is configured to: When the virtual assistant is in an awake state, the sight direction of the first user is obtained.
21. A control device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program stored in the memory, so that the apparatus performs the method according to any one of claims 1 to 10.
22. A control system, characterized in that: It comprises a physical virtual assistant and a computing platform, wherein the computing platform comprises the apparatus as claimed in any one of claims 11 to 21.
23. A control system, characterized in that: It comprises a display device and a computing platform, wherein the display device is used to display a virtual assistant, and the computing platform comprises the device as described in any one of claims 11 to 21.
24. A vehicle, characterized in that: The method comprises a control device according to any one of claims 11 to 21, or a control system according to claim 22 or 23.
25. The vehicle according to claim 24, characterized in that The transport vehicle is a vehicle.
26. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a computer, the method according to any one of claims 1 to 10 is implemented.
27. A chip, characterized in that: The chip includes a processor and a data interface, and the processor reads instructions stored in a memory through the data interface to execute the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
A vehicle multi-screen control system and a vehicle multi-screen control method
CN109828655A
Voice interaction method and device, vehicle and machine readable medium
CN110211586A
Robot interaction method and device, electronic equipment and storage medium
CN110737335A
Vehicle-mounted robot and human-machine interaction method thereof
CN110871447A
Non-verbal engagement of a virtual assistant
CN111492328A