Interaction device, interaction method and vehicle

By combining a depth camera and an ultrasonic transmitter array, the system recognizes gesture semantics and spatial location, and drives the ultrasonic transmitter array to provide tactile feedback. This solves the problem of lack of tactile feedback in rear passenger interactions, and enables reliable operation and enhanced immersion in complex environments.

CN121918700APending Publication Date: 2026-04-24BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-05
Publication Date
2026-04-24

Smart Images

  • Figure CN121918700A_ABST
    Figure CN121918700A_ABST
Patent Text Reader

Abstract

The invention relates to an interaction device, an interaction method and a vehicle, and relates to the technical field of vehicles, and the interaction device comprises a depth camera which is used for collecting image data of a hand of a user; the ultrasonic transmitting array comprises a plurality of ultrasonic transmitters and is used for transmitting ultrasonic waves; the controller is connected with the ultrasonic transmitter and the depth camera, and the controller is configured to receive the image data and determine gesture semantics and the three-dimensional space position of the hand target area based on the image data; generating a tactile feedback control instruction based on the gesture semantics; and according to the three-dimensional space position and the tactile feedback control instruction, each ultrasonic transmitter in an ultrasonic transmitting array is driven to transmit focused ultrasonic waves to the hand target area so as to generate tactile feedback corresponding to the gesture semantics. Through the method, the feedback of the gesture operation can be focused and projected to the target area of the hand through the ultrasonic waves, and accurate interaction confirmation tactile feedback is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of vehicle technology, and more specifically, to an interactive device, an interactive method, and a vehicle. Background Technology

[0002] In the current era of rapid development of intelligent vehicles, human-machine interaction has become a core battleground for automakers vying for the high-end market. From early physical button remote controls to today's headrest touchscreens and ceiling-mounted voice control, the industry's exploration of visual and auditory experiences has reached near saturation. However, the application of tactile feedback, a key dimension in the human perception system for "confirming the validity of operations," has long lagged behind. When adjusting seat temperature or switching audio-visual content, rear passengers still rely on the traditional mode of "looking at the screen + manual operation," unable to obtain operational confirmation through tactile perception, and finding it difficult to operate blindly in complex lighting or dynamic driving environments. This current state of interaction, which "emphasizes hearing and neglects touch," has resulted in rear-seat interaction remaining at the level of "one-way command output," lacking deep perceptual interaction between humans and vehicles. Summary of the Invention

[0003] The purpose of this disclosure is to provide a vehicle control method, vehicle, vehicle controller, and storage medium that can provide accurate interactive confirmation tactile feedback by focusing ultrasonic waves onto a target area of ​​the hand to provide feedback for gesture operations.

[0004] To achieve the above objectives, this disclosure provides an interactive device, the interactive device comprising: A depth camera is used to capture image data of the user's hand; An ultrasonic transmitting array, comprising multiple ultrasonic transmitters, for emitting ultrasonic waves; A controller, connected to both the ultrasonic transmitter and the depth camera, is configured to: The system receives the image data and determines the gesture semantics and the three-dimensional spatial position of the hand target area based on the image data; it generates tactile feedback control commands based on the gesture semantics; and it drives each ultrasonic transmitter in the ultrasonic transmitting array to emit focused ultrasonic waves toward the hand target area according to the three-dimensional spatial position and the tactile feedback control commands, so as to generate tactile feedback corresponding to the gesture semantics.

[0005] Optionally, the gesture semantics are determined based on continuous image data acquired in real time by the depth camera, and the gesture semantics include the feature changes of the gesture operation; the tactile feedback corresponding to the gesture semantics includes fixed-point tactile feedback and / or sliding tactile feedback; The fixed-point tactile feedback is used to instruct the ultrasonic emission array to form a sound pressure focus at a specific location in the target area of ​​the hand, and to dynamically adjust the intensity of the sound pressure focus according to the characteristic change. The sliding tactile feedback is used to instruct the ultrasonic emitting array to control the sound pressure focus formed by the ultrasonic emitting array to move along a specified direction or path on the target area of ​​the hand according to the characteristic change.

[0006] Optionally, the fixed-point haptic feedback is generated in the following manner: The controller is configured to: determine that the gesture semantics are a first type of gesture semantics for continuous intensity adjustment, and determine the target intensity value of the sound pressure focus based on the feature change amount; Based on the three-dimensional spatial position of the target area of ​​the hand, determine the required phase control parameters for each ultrasonic transmitter in the ultrasonic transmitting array; Based on the target intensity value, determine the amplitude control parameters for driving the ultrasonic transmitting array; Based on the phase control parameters and amplitude control parameters, the ultrasonic transmitting array is driven to emit ultrasonic waves to form a sound pressure focus in the target area of ​​the hand, and the intensity of the sound pressure focus corresponds to the target intensity value.

[0007] Optionally, the haptic feedback is generated in the following manner: The controller is configured to: determine that the gesture semantics are a second type of gesture semantics used to indicate a specified direction or path, and determine the movement vector information of the sound pressure focus based on the feature change amount; The three-dimensional spatial position of the hand target area is determined based on the image data acquired in real time by the depth camera; Based on the three-dimensional spatial position of the hand target area and the movement vector information, the three-dimensional spatial position of the sound pressure focus is determined; Based on the three-dimensional spatial position of the sound pressure focus, the required phase control parameters for each ultrasonic transmitter in the ultrasonic transmitting array are determined; Based on the phase control parameters, the ultrasonic transmitting array is driven to emit ultrasonic waves.

[0008] Optionally, the controller is connected to multiple execution devices, and the controller is further configured to: Based on the gesture semantics, a device control command is generated, and the device control command is sent to the execution device corresponding to the gesture semantics, so that the execution device executes the device control command.

[0009] Optionally, the depth camera and the ultrasonic transmitting array are integrated into a housing according to a specified positional relationship.

[0010] Optionally, the housing is cylindrical, and the depth camera is located in the central region of one end face of the cylindrical structure; the ultrasonic transmitting array includes multiple ultrasonic transmitters that are circumferentially distributed around the depth camera.

[0011] Optionally, the diaphragm and resonant cavity of the ultrasonic transmitter are manufactured using MEMS technology.

[0012] A second aspect of this disclosure provides an interaction method applied to a controller in an interaction device, comprising: Receive image data captured by a depth camera; Based on the image data, determine the semantics of the gesture and the three-dimensional spatial position of the target area of ​​the hand; Based on the gesture semantics, haptic feedback control commands are generated; Based on the three-dimensional spatial position and the tactile feedback control command, the ultrasonic wave emitting array is driven to emit focused ultrasonic waves toward the target area of ​​the hand to generate tactile feedback corresponding to the gesture semantics.

[0013] A third aspect of this disclosure provides a vehicle including a body and the aforementioned interactive device.

[0014] Optionally, the interactive device is located on the rear armrest of the vehicle body.

[0015] Through the above technical solution, the interactive device can accurately identify the semantic meaning of the user's gesture and the three-dimensional spatial position of the target area of ​​the hand through image data collected by the depth camera. Then, based on the semantic meaning of the gesture and the three-dimensional spatial position, it drives the ultrasonic wave transmitting array to emit focused ultrasonic waves toward the target area of ​​the hand, thereby providing the user with realistic tactile feedback corresponding to the semantic meaning of the gesture without having to touch the physical object.

[0016] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a connection block diagram of an interactive device according to an exemplary embodiment of the present disclosure.

[0018] Figure 2 This is a schematic diagram of the structure of an interactive device according to an exemplary embodiment of the present disclosure.

[0019] Figure 3This is a schematic diagram of the structure of a depth camera according to an exemplary embodiment of the present disclosure.

[0020] Figure 4 This is another structural schematic diagram of an interactive device according to an exemplary embodiment of the present disclosure.

[0021] Figure 5 This is a schematic diagram illustrating the correspondence between gestures for audio-visual control and semantic meanings according to an exemplary embodiment of this disclosure.

[0022] Figure 6 This is a schematic diagram illustrating the correspondence between game control gestures and semantic meanings according to an exemplary embodiment of the present disclosure.

[0023] Figure 7 This is a flowchart illustrating an interaction method according to an exemplary embodiment of the present disclosure.

[0024] Figure 8 This is a connection block diagram of a vehicle according to an exemplary embodiment of the present disclosure.

[0025] Figure 9 Another connection block diagram of a vehicle is shown according to an exemplary embodiment of the present disclosure.

[0026] Figure 10 This is a schematic diagram of the structure of a vehicle according to an exemplary embodiment.

[0027] Figure 11 This is a schematic block diagram illustrating the structure of a vehicle according to an exemplary embodiment.

[0028] Figure 12 This is a flowchart illustrating the interactive control process of a vehicle according to an exemplary embodiment.

[0029] Explanation of reference numerals in the attached figures Interactive device-100; Depth camera-110; Ultrasonic emission array-120; Controller-130; Housing-140; Body-200; Rear armrest-210; Window-220; Entertainment display screen-230; Vehicle-20. Detailed Implementation

[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0032] The term "comprising" and its variations as used herein are open-ended inclusion, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0033] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0034] Human-computer interaction between vehicles and users in related technologies typically suffers from the following three defects: First, the interaction modality is limited, and each solution is confined to a single interaction dimension: Existing rear-seat human-computer interaction solutions generally suffer from a single interaction modality. Each interaction method is limited to a specific dimension, making it difficult to meet the flexible operation needs in complex scenarios. Traditional remote controls and physical buttons rely entirely on physical contact; users must press physical buttons to complete operations, making it impossible to control without hand touch. Touchscreen operation (such as armrest screens and headrest screens) relies solely on the touch perception modality; all commands must be triggered by physical contact between the finger and the screen, lacking other dimensions of auxiliary interaction. Voice control relies solely on the sound modality, entirely relying on voice commands to convey operational intentions, unable to combine visual or tactile information to enhance accuracy. Simple gesture recognition relies solely on the visual tracking modality, only using a camera to capture hand movements for control, lacking supplementary auditory or tactile feedback. This limitation of a single modality makes it difficult for each solution to cope with the user's interaction needs in different states; for example, when wanting to relax and raise an arm, voice cannot replace touchscreen, or in noisy environments, gestures cannot replace voice.

[0035] Second, it has poor environmental adaptability and is easily affected by environmental factors such as light, sound, and space. The environmental adaptability of the rear-seat human-machine interaction system is particularly problematic, with different solutions significantly affected by environmental factors such as light, sound, and space. In well-lit environments, strong light severely interferes with touchscreen operation, and screen glare significantly reduces visibility, making it difficult to recognize operation commands. In low-light environments, the accuracy of simple gesture recognition is weakened, as the camera cannot clearly capture hand movements due to insufficient light, resulting in a sharp drop in recognition accuracy. Traditional remote controls and physical buttons are also affected by light; in low-light conditions at night, users need to bend down and fumble for buttons, greatly reducing operational efficiency. In sound environments, voice control is significantly affected by ambient noise (such as tire noise and passenger conversations), drastically reducing command recognition accuracy. In spatial environments, traditional remote controls are also prone to being lost; rear passengers, in a relaxed posture, may place them carelessly, easily leading to inefficient "searching for the remote" scenarios, further highlighting the lack of environmental adaptability.

[0036] Third, the lack of immersive operation disrupts the relaxed posture and the smoothness of interaction: Existing solutions suffer from significant deficiencies in immersive operation, failing to meet the core need of rear-seat passengers for precise control in a relaxed posture. Touchscreen operation (such as armrest and headrest screens) requires the arm to be continuously raised and suspended in the air, which can easily lead to muscle fatigue over time, forcing passengers to change their relaxed posture. Traditional remote controls and physical buttons require users to look down to find the buttons, disrupting the relaxed state of the rear seats and distracting their attention. The interaction chain of simple gesture recognition is broken, providing only visual feedback. Users need to frequently look at the screen to confirm whether the operation has taken effect, unable to perceive the operation status through touch, and lacking key tactile experiences such as "pressing" and "sliding resistance," weakening the realism of the operation. Voice control lacks privacy in multi-passenger scenarios. Commands such as "play music" can be accidentally triggered by adjacent passengers, and the operation process is completely exposed, destroying the immersive experience of personal interaction. These problems collectively make it difficult for rear-seat passengers to complete control operations naturally and smoothly in a relaxed state, resulting in a fragmented interaction experience.

[0037] Based on this, this disclosure provides an interactive device that can accurately identify the semantic meaning of a user's gestures and the three-dimensional spatial position of the target area of ​​the hand using image data acquired by a depth camera. This drives an ultrasonic wave emission array to emit focused ultrasonic waves towards the hand, providing the user with realistic tactile feedback corresponding to the interactive semantics without physical contact. This allows users to operate in a completely relaxed state using natural gestures, without relying on visual confirmation as with pure gesture recognition, and without needing to raise their voice in noisy environments as with voice recognition. Furthermore, because the controller can process the image data acquired by the depth camera using stereo vision and deep learning algorithms to determine the three-dimensional spatial position, it maintains a high recognition rate even in low-light environments. Due to the tactile feedback, users can perform accurate blind operations even when the screen is highly reflective and the interface is completely obscured, completely eliminating absolute dependence on the visual interface. Moreover, tire noise, wind noise, or passenger conversations during vehicle operation do not affect the accuracy of gesture recognition or the stability of tactile feedback, maintaining smooth and reliable interaction even in noisy environments.

[0038] Please refer to the following: Figure 1 and Figure 2 As shown, the interactive device 100 provided in this disclosure specifically includes a depth camera 110, an ultrasonic emitting array, and a controller 130, with the depth camera 110 and the ultrasonic emitting array 120 respectively connected to the controller 130.

[0039] The depth camera 110 is used to acquire image data of the user's hand; the ultrasonic transmitting array 120 includes multiple ultrasonic transmitters for emitting ultrasonic waves; the controller 130 is configured to: receive the image data and determine the gesture semantics and the three-dimensional spatial position of the hand target area based on the image data; generate tactile feedback control commands based on the gesture semantics; and drive each ultrasonic transmitter in the ultrasonic transmitting array 120 to emit focused ultrasonic waves toward the hand target area according to the three-dimensional spatial position and the tactile feedback control commands, so as to generate tactile feedback corresponding to the gesture semantics.

[0040] By employing the aforementioned interactive device 100, the device can accurately identify the semantic meaning of the user's gestures and the three-dimensional spatial position of the target area of ​​the hand using image data collected by the depth camera 110. This drives the ultrasonic wave emission array 120 to emit focused ultrasonic waves towards the target area of ​​the hand, providing the user with realistic tactile feedback corresponding to the interactive semantics without physical contact. This ensures that every user gesture is accompanied by a tactile response, eliminating the need for the user to look at confirmation prompts on the screen. The user can be certain the command has been executed simply by touch, allowing their gaze to remain focused on the road surface, fundamentally reducing the risk of distracted driving.

[0041] Furthermore, when a user does not perceive the expected tactile feedback, they can immediately realize that the operation is invalid or the position is inaccurate, thus facilitating immediate adjustment. This instant negative feedback enables users to quickly correct their operating posture.

[0042] The depth camera 110 can be any camera capable of collecting depth data, such as a point cloud image acquisition device or a binocular camera.

[0043] Taking the Depth Camera 110 as an example of a binocular camera, a binocular camera is an imaging device designed to mimic the principle of human binocular vision, such as... Figure 3 As shown, a binocular camera consists of two identical camera modules. The two cameras capture images of the same object from different angles, resulting in two images. Using computer vision algorithms, the same feature points (i.e., any pixel in the target area of ​​the hand in the image data) are found in the two images. The disparity is calculated by calculating the difference in pixel position of the same feature point in the left and right camera images. Based on the known precise distance (baseline distance) between the two cameras and the disparity, the depth distance between the feature point and the camera can be accurately calculated using the principle of triangulation, thereby obtaining the three-dimensional spatial information of the feature point.

[0044] In this embodiment, the depth camera 110 can collect image data in real time, namely image data of the user's hand. The aforementioned image data specifically includes image data collected by the two cameras in the binocular camera.

[0045] The ultrasonic transmitting array 120 includes multiple ultrasonic transmitters arranged in an array. The energy of a single ultrasonic transmitter is very weak, insufficient to produce a clear sensation on the skin. The energy of the ultrasonic waves emitted by the multiple ultrasonic transmitters in the ultrasonic transmitting array 120 is converged at a single point (i.e., the sound pressure focal point), and the resulting sound pressure intensity is sufficient to be perceived by the user.

[0046] In one possible implementation, the diaphragm and resonant cavity of the ultrasonic transmitter are manufactured using MEMS technology.

[0047] Because MEMS (Micro-Electro-Mechanical Systems) technology can precisely fabricate diaphragms and resonant cavities at the micrometer scale, the size of a single ultrasonic transmitter can be significantly reduced (e.g., the diameter can be as small as 0.4 cm). This allows for the integration of up to a greater number of ultrasonic transmitters within the same housing surface area 140, forming a high-density ultrasonic transmission array 120, thereby making ultrasonic feedback more accurate. Furthermore, the low driving voltage and low power consumption of MEMS devices result in lower energy consumption for the interactive device 100.

[0048] The depth camera 110 and the ultrasonic transmitting array 120 are integrated into a housing 140 according to a specified positional relationship.

[0049] The aforementioned designated positional relationship can be that the depth camera 110 is centered and multiple ultrasonic transmitters in the ultrasonic transmitting array 120 surround the depth camera 110; or the depth camera 110 and the ultrasonic transmitting array 120 are on the same plane, side by side or top to bottom. The aforementioned designated positional relationship is only illustrative and there can be many other settings, which are not specifically limited here.

[0050] It is worth mentioning that by adopting the above configuration, the depth camera 110 and the ultrasonic transmitter array 120 are integrated into a housing 140 with a fixed positional relationship. This allows the depth camera 110 and each ultrasonic transmitter to have a unique and stable spatial positional relationship. In subsequent control, only a set of pre-calibrated spatial transformation parameters are needed to directly map the three-dimensional spatial position of the hand target area determined by the depth camera 110 to the control parameters required to drive the ultrasonic transmitter array 120. This greatly simplifies the coordinate transformation calculation inside the controller 130, reduces the calculation delay, and provides a physical basis for real-time interaction.

[0051] In one possible implementation, please refer to Figure 4 The housing 140 is cylindrical in shape, and the depth camera 110 is disposed in the central region of one end face of the cylindrical shape; the ultrasonic transmitting array 120 includes a plurality of ultrasonic transmitters, which are circumferentially distributed around the depth camera 110.

[0052] In this implementation, because the depth camera 110 is centrally located and has a 360° unobstructed omnidirectional field of view, it can clearly and accurately capture the user's hand no matter what angle the hand enters the interaction area from, fundamentally eliminating blind spots and significantly improving the range and reliability of gesture recognition. Furthermore, the circumferentially distributed ultrasonic transmitters form an omnidirectional sound wave emission surface. This allows the system to emit ultrasonic waves towards the target area of ​​the hand from multiple directions. Through beamforming technology, the sound pressure focus energy formed within the target area of ​​the hand is more concentrated and evenly distributed. Regardless of the direction from which the user performs hand operations, they can obtain consistent intensity and precise positioning tactile feedback, experiencing no directional differences.

[0053] The controller 130 can be a multimodal fusion controller, integrating multiple dedicated computing units, such as one or more of a vision processing unit, a control unit, and a coordination unit. The vision processing unit is used to run a high-speed deep learning model, process the continuous image data stream from the depth camera 110, and perform real-time hand detection, key point recognition, and 3D spatial positioning. The control unit is used to execute a high-precision ultrasonic phased array beamforming algorithm, calculating the phase and amplitude parameters required by hundreds of ultrasonic transmitters in real time. The coordination unit is used to send control commands corresponding to the gestures to the execution device.

[0054] The process by which the controller 130 determines the gesture semantics and the three-dimensional spatial position of the hand target area based on image data can specifically be as follows: Machine learning models (such as convolutional neural networks CNN) are used to segment images and detect key points, and to accurately locate hand target areas such as joints (fingertips, finger roots, and wrists). Based on the collected image data, the pixel position difference of the same feature point (e.g., feature points in the hand target area) in the left and right camera images can be calculated, i.e., parallax. And based on the precise distance (baseline distance) between the two cameras in the depth camera 110, the camera intrinsic parameters, and the parallax, the depth distance between the feature point and the camera can be accurately calculated through the principle of triangulation, thereby obtaining the three-dimensional spatial position of the hand feature point.

[0055] Subsequently, the motion trajectory of hand key points can be obtained based on the three-dimensional spatial positions of hand feature points in multiple consecutive images. Furthermore, the motion velocity and acceleration of the key points can be calculated by analyzing the positional changes between consecutive frames. Based on the motion trajectory, velocity, and acceleration of multiple key points, the user's gesture semantics are determined. This gesture semantics refers to abstract instruction information with a clear interactive purpose, parsed from the user's continuous hand movements. Gesture semantics includes not only the gesture type but also the feature changes corresponding to the hand key points and the corresponding operation instructions (e.g., brightness adjustment, volume adjustment, and window opening / closing).

[0056] Different feature changes can correspond to different tactile feedback changes (e.g., changes in pressure, vibration amplitude, or vibration frequency generated on the skin by the sound pressure focus, or changes in the position of the pressure focus on the user's hand), so as to accurately map the user's physical gestures into a perceptible tactile experience, thereby achieving intuitive interactive control.

[0057] Among them, gesture types and corresponding operation commands can be applied in different interaction scenarios, specifically, such as Figure 5 As shown: In the audio-visual control scenario, the gestures are as follows: a normal relaxed palm waving downwards (corresponding to the operation command of turning on the rear screen), a gesture with five fingers spread open and held (corresponding to the operation command of turning on the screen cursor), a gesture of changing from an open gesture to a fist (corresponding to the operation command of opening the video), a gesture of moving the fist left or right (corresponding to the operation command of rewinding or fast-forwarding the video), and a gesture of moving the fist downwards or upwards (corresponding to the operation command of increasing or decreasing the volume).

[0058] like Figure 6 As shown, in the game scene, the gestures are: pointing the index finger downwards to make a click gesture (corresponding to the operation command is "interact"); moving the index finger in all directions on the palm plane (corresponding to the operation command is "move"); tapping the wrist with a single finger and waving it upwards or downwards (corresponding to the operation command is "jump" or "crouch"); waving the index finger to the left or right on the palm plane (corresponding to the operation command is "move left or right to dodge"); rotating the index finger around the wrist (corresponding to the operation command is "turn around"); and pointing the index and middle fingers downwards simultaneously to make a click gesture (corresponding to the operation command is "accelerate").

[0059] It should be understood that there can be many more interaction scenarios and more gestures. Different interaction scenarios may correspond to the same gestures, and the operation instructions corresponding to the same gestures in different interaction scenarios may be the same or different. No specific limitations are made here.

[0060] The tactile feedback corresponding to different gestures can be applied to different parts of the hand. For examples, please refer to [link / reference]. Figure 5 and Figure 6 As shown, with right-hand operation as the core, targeted gesture operations and corresponding tactile feedback positions were designed around audio-visual control and game control scenarios. Each gesture operation provides feedback in the target area of ​​the hand (the target area of ​​the hand is the area where the hand circle is located, i.e., the tactile feedback position), so that users can perceive that the gesture operation has been executed through touch, enhancing the immersion and accuracy of the interaction.

[0061] Based on the aforementioned gesture semantics, the specific method for generating haptic feedback control commands can be to determine the haptic feedback effect from a preset rule base according to the gesture type. Different gesture types can correspond to different haptic feedback effects, such as fixed-point haptic feedback (press feedback or vibration feedback, etc.) and sliding haptic feedback, etc.

[0062] In one possible implementation, the tactile feedback corresponding to the gesture semantics includes fixed-point tactile feedback and sliding tactile feedback, wherein the fixed-point tactile feedback is used to instruct the ultrasonic emitting array 120 to form a sound pressure focus at a specific position in the target area of ​​the hand, and dynamically adjust the intensity of the sound pressure focus according to the characteristic change amount; The sliding tactile feedback is used to instruct the ultrasonic emitting array 120 to control the sound pressure focus formed by the ultrasonic emitting array 120 to move along a specified direction or path on the target area of ​​the hand according to the characteristic change amount.

[0063] It's worth noting that, based on the interaction intent, gesture semantics are mainly divided into two categories (i.e., the first type of gesture semantics and the second type of gesture semantics), each corresponding to different haptic feedback implementation methods. Specifically, the distinction between the first and second types of gesture semantics can be based on the interaction intent: whether it involves continuous, analog adjustment or discrete, navigable switching. Alternatively, it can be pre-set according to actual needs.

[0064] The first type of gesture semantics is used for continuous, analog quantity adjustment, and its operation result is a continuously changing quantity (such as volume or brightness). The corresponding gestures can be, in audio-visual scenarios, clenching a fist and moving it downwards or upwards (adjusting volume), or in game scenarios, rotating the index finger around the wrist (continuous turning), or moving the index finger continuously in a certain direction (controlling movement speed). The second type of gesture semantics is used for discrete, navigable switching or triggering, and its operation result is the execution of a specific command (such as switching or confirming). The corresponding gestures can be, in audio-visual scenarios, waving the palm downwards (turning on the screen), clenching a fist and moving it left or right (changing songs), or changing from an open hand to a clenched fist (selecting and confirming); in game scenarios, clicking the index finger downwards (interaction), waving the wrist up and down (jumping / crouching), or waving the index finger left and right (dodge), etc.

[0065] The fixed-point tactile feedback is generated in the following manner: the controller 130 is configured to: determine that the gesture semantics are a first type of gesture semantics for continuous intensity adjustment; determine the target intensity value of the sound pressure focus based on the feature change; determine the phase control parameters required for each ultrasonic transmitter in the ultrasonic transmitting array 120 according to the three-dimensional spatial position of the hand target area; determine the amplitude control parameters for driving the ultrasonic transmitting array 120 according to the target intensity value; and drive the ultrasonic transmitting array 120 to emit ultrasonic waves to form a sound pressure focus in the hand target area according to the phase control parameters and amplitude control parameters, wherein the intensity of the sound pressure focus corresponds to the target intensity value.

[0066] In this implementation, the controller 130 calculates the target intensity value of the sound pressure focus based on the characteristic changes of the gesture (such as vertical displacement ΔY) using a mapping function (e.g., target intensity = base value + k*|ΔY|). Subsequently, based on the real-time three-dimensional coordinates of the hand target area and the geometric model M of the ultrasonic transmitting array 120 (containing the three-dimensional coordinates (Xn, Yn, Zn) of each ultrasonic transmitter Tn in the ultrasonic transmitting array 120), a beamforming algorithm is used to determine the required phase control parameters for each ultrasonic transmitter in the ultrasonic transmitting array 120 to form a positionally stable sound pressure focus. Then, the actual driving voltage of each ultrasonic transmitter in the ultrasonic transmitting array 120 can be determined based on the target intensity value; this voltage is the amplitude control parameter. Combining the phase control parameter and the amplitude control parameter, the array is driven to emit ultrasonic waves, forming a fixed-point tactile sensation in the hand target area whose intensity dynamically changes with the operation. This fixed-point tactile sensation can simulate the resistance of a physical knob or slider.

[0067] This implementation method allows the sound pressure focus to be stably locked onto a specific target area of ​​the hand (such as the fingertips), and the physical motion of the user's gestures (such as vertical displacement ΔY) is converted into perceptible tactile intensity in real time through a mapping function. This enables users to continuously and steplessly adjust parameters such as volume and brightness with precision, just like rotating a physical knob or pushing a physical slider, enhancing the realism and immersion of the operation.

[0068] In one possible implementation, the haptic feedback is generated in the following manner: the controller 130 is configured to: determine that the gesture semantics are a second type of gesture semantics used to indicate a specified direction or path; determine the movement vector information of the sound pressure focus based on the feature change amount; determine the three-dimensional spatial position of the target position in the hand target area based on the image data acquired in real time by the depth camera 110; determine the three-dimensional spatial position of the sound pressure focus based on the three-dimensional spatial position of the hand target area and the movement vector information; determine the phase control parameters required for each ultrasonic transmitter in the ultrasonic transmitting array 120 based on the three-dimensional spatial position of the sound pressure focus; and drive the ultrasonic transmitting array 120 to emit ultrasonic waves according to the phase control parameters to update the three-dimensional spatial position of the sound pressure focus.

[0069] Specifically, the controller 130 calculates the motion vector information (including direction and speed) of the sound pressure focus based on the characteristic changes of the gesture (such as displacement vector and velocity). Then, starting from the hand target position acquired in real time by the depth camera 110, the controller plans a sequence of three-dimensional spatial path points that the sound pressure focus needs to continuously pass through, in conjunction with the motion vector. Based on the coordinates of each path point and the geometric model of the ultrasonic transmitting array 120, the controller calculates the required phase control parameter sequence for each transmitter in the array in real time using a beamforming algorithm. Finally, the controller drives the ultrasonic transmitting array 120 according to these rapidly updated phase parameters, so that the sound pressure focus produces a stable and continuous sliding sensation on the user's hand skin along the predetermined path, thereby simulating a directional feedback effect as if brushing against a physical interface.

[0070] By employing the above settings, when a user performs actions such as waving their hand horizontally to change songs or turn pages, a clear sound pressure focus will glide across the skin of the hand in a specified direction or along a specified path. This vector-based feedback provides the user with a clear command execution signal without visual confirmation, significantly improving the certainty and reliability of the operation.

[0071] It is worth mentioning that the aforementioned interactive device 100 can be associated with execution devices, such as gaming devices, windows (e.g., car windows), audio playback devices, and displays (in-vehicle entertainment screens). The controller 130 can also obtain corresponding control commands and execution devices for executing the control commands based on gesture semantics, and send the control commands to the execution devices to cause the execution devices to execute the control commands.

[0072] This application provides an interactive device 100 that utilizes a depth camera 110 and an ultrasonic transmitter array 120 to achieve a deep fusion of "visual gesture recognition + ultrasonic tactile feedback," forming a multi-dimensional collaborative interactive system. The depth camera 110 captures the three-dimensional trajectory of gestures to achieve visual input, while multiple miniature ultrasonic transmitters in the ultrasonic transmitter array 120 generate spatially focused sound waves through phase modulation, with the sound pressure focus located in the target area of ​​the hand, providing tactile feedback such as pressing and sliding resistance. The two are integrated through a hardware-level multimodal controller 130 to achieve millisecond-level data interaction. This closed-loop "input-response-feedback" mechanism not only overcomes the single-modality limitation of touchscreens and physical buttons that rely on physical contact, but also compensates for the shortcomings of simple gesture recognition lacking tactile confirmation and voice control being limited by sound. For example, when adjusting the volume, the visual system captures the gesture trajectory, while the tactile feedback provides progressive vibrations as the volume changes. Only by combining this deep hardware integration of vision and tactile feedback (such as a shared data bus design) with software collaborative algorithms (dynamically adjusting focus parameters) can a multimodal experience of "being able to perceive operations without touching" be achieved. Such precise collaborative effects cannot be achieved by relying on a single module.

[0073] Please see Figure 7 As shown, another embodiment of this application provides an interaction method, which can be applied to the controller 130 in the above-described interaction device 100. The method includes: Step S310: Receive image data acquired by depth camera 110; Step S320: Determine the gesture semantics and the three-dimensional spatial position of the hand target area based on the image data; Step S330: Generate haptic feedback control commands based on the gesture semantics; Step S340: Based on the three-dimensional spatial position and the tactile feedback control command, drive the ultrasonic wave emitting array 120 to emit focused ultrasonic waves toward the hand target area to generate tactile feedback corresponding to the gesture semantics.

[0074] For details on the specific implementation principles of the above steps, please refer to the specific description of the interactive device 100 in the foregoing embodiments, which will not be repeated here.

[0075] In one possible implementation, the tactile feedback control command includes phase control parameters and amplitude control parameters; step S330 includes: determining that the gesture semantics is a first type of gesture semantics for continuous intensity adjustment; determining the target intensity value of the sound pressure focus based on the feature change amount; determining the phase control parameters required for each ultrasonic transmitter in the ultrasonic transmitting array 120 according to the three-dimensional spatial position of the hand target area; and determining the amplitude control parameters for driving the ultrasonic transmitting array 120 according to the target intensity value; step S340 includes: driving the ultrasonic transmitting array 120 to emit ultrasonic waves to form a sound pressure focus in the hand target area according to the phase control parameters and amplitude control parameters, wherein the intensity of the sound pressure focus corresponds to the target intensity value.

[0076] In one possible implementation, the haptic feedback control command includes phase control parameters; step S330 includes: determining that the gesture semantics are a second type of gesture semantics used to indicate a specified direction or path; determining the movement vector information of the sound pressure focus based on the feature change amount; determining the three-dimensional spatial position of the sound pressure focus based on the three-dimensional spatial position of the target position in the hand target area determined in real time based on the image data collected by the depth camera 110, and the movement vector information; determining the phase control parameters required for each ultrasonic transmitter in the ultrasonic transmitting array 120 based on the three-dimensional spatial position of the sound pressure focus; step S340 includes: driving the ultrasonic transmitting array 120 according to the phase control parameters to update the three-dimensional spatial position of the sound pressure focus.

[0077] Please refer to the following: Figure 8 , Figure 9 as well as Figure 10 As shown, another embodiment of this application provides a vehicle 20, which includes a body 200 and an interaction device 100, wherein the interaction device 100 is disposed on the rear armrest 210 of the body 200.

[0078] Specifically, when the interactive device 100 is installed in the rear armrest 210 of the vehicle body 200, it can be placed inside the armrest box of the rear armrest 210.

[0079] In one possible implementation, the interactive device 100 adopts a highly integrated cylindrical configuration. One end of the cylindrical configuration has a MEMS-manufactured ultrasonic transmitting array 120 distributed in a ring around its perimeter, with a depth camera 110 supporting autofocus embedded in the center. The other end of the cylindrical configuration has a power interface and a communication interface, responsible for connecting to the vehicle's power supply and data interaction, respectively. Internally, it achieves efficient collaboration through a modular layout, integrating the ultrasonic transmitting array 120 and the depth camera 110 at the front end of the cylindrical configuration, arranging a controller 130 in the middle, and integrating the power and communication modules at the rear, thus completing sensing, processing, feedback, and communication functions within a limited space.

[0080] When the interactive device 100 is installed on the rear armrest 210 of the vehicle body 200, the interactive device 100 can be connected to the display screen 230 (e.g., ceiling-mounted entertainment display screen, headrest entertainment display screen) and the rear window 220 and other execution devices. The ceiling-mounted entertainment display screen and the headrest entertainment display screen provide interactive display interfaces, and the rear window 220 acts as the controlled object and responds to device control commands.

[0081] like Figure 11 As shown, when a user interacts with the vehicle 20, after the vehicle 20 is powered on, it enters the initialization phase. The interaction device 100 can initiate self-testing and parameter calibration, and the rear entertainment screen and window 220 control synchronization is ready and in a wake-up state. Rear users can control the rear entertainment screen or window 220 with gestures. When the user performs gesture operation, the interaction device 100 can capture the hand movement trajectory with millimeter-level precision to operate the entertainment display screen 230 or window 220, and focus sound waves according to the three-dimensional coordinates of the hand target area to enable the user to synchronously perceive tactile feedback, forming a complete closed loop of "perception-response-feedback". For example, such as Figure 12 As shown, the interactive device 100 includes an ultrasonic transmitting array 120, a depth camera 110, a controller 130, a communication interface, and a power module; the window 220 specifically includes a motor drive component, a position detection component, and a control component; and the entertainment display screen 230 specifically includes a display module, an audio component, and a touch interaction component.

[0082] When opening and closing the window via the interactive device 100, the user makes a "pushing forward with the palm" gesture within the image acquisition range of the interactive device 100 located on the rear armrest 210. The depth camera 110 captures continuous images of this hand movement and sends the image data stream to the controller 130. The controller 130 can recognize the gesture semantics and the three-dimensional coordinates of the palm in space based on the image data, generate tactile feedback commands based on the gesture semantics and three-dimensional coordinates, and generate device control commands based on the gesture semantics. The controller 130 sends the device control commands to the control components of the window 220 and the drive ultrasonic wave transmitter array 120 via a communication interface. Based on the spatial coordinates of the palm, it emits focused ultrasonic waves, generating a clear "click" tactile sensation on the user's palm. Upon receiving the command, the control components of the window 220 immediately activate the motor drive component to begin raising the window 220. The position detection component monitors the position of the window 220 in real time, and after the window 220 is fully closed, it sends a feedback signal to the control components, and the motor stops working. Throughout the process, users do not need to look down or turn their heads to look at the car window 220; they can confirm that the operation has taken effect simply by relying on the tactile feedback of their hands, achieving perfect blind operation.

[0083] When playing audio via the interactive device 100, the user places their hand naturally in front of the armrest and makes a rotating gesture (simulating rotating a knob) by pinching their index finger and thumb together and then moving it upwards, or a simple "palm lift" gesture. The depth camera 110 captures the trajectory and speed of this fine gesture. The controller 130 can perform gesture semantic recognition based on the real-time acquired image data to determine that the gesture is "increase the volume," and analyze the change in the target volume based on the movement distance or speed of the gesture, thereby generating a haptic feedback command and a device control command. The controller 130 sends the device control command to the audio component of the display screen 230 through the communication interface. At the same time, based on the haptic feedback command, the ultrasonic transmitting array 120 is driven to form a continuous sound pressure focus in the area between the user's thumb and index finger, and the intensity of this focus increases linearly with the increase in volume. After receiving the command, the audio component smoothly adjusts the volume to the target value. The display module can also simultaneously show an animation of increasing volume (such as a progress bar lengthening), while the user's tactile feedback gradually intensifies. This perfectly synchronizes with the auditory and visual feedback of increasing volume, creating an immersive experience "as if turning a real volume knob." Users can precisely control the volume adjustment without looking at the screen, relying solely on touch.

[0084] By placing the interactive device 100 in the rear armrest 210, the problem of easily lost traditional remote controls can be avoided, ensuring that the interactive device 100 is always in a stable and usable state. Furthermore, by integrating the ultrasonic transmitting array 120 and the depth camera 110 at the front of the cylindrical configuration, the interactive device 100 is less susceptible to strong light, avoiding the problems of poor visibility and operational obstruction caused by strong light in traditional touchscreens. Moreover, since the interactive device 100 of this application adopts a "visual + tactile" non-voice control mode, it completely eliminates the influence of ambient noise on the interaction. For insufficient light, the high performance and hardware-level integrated design of the depth camera 110 ensure the stability of gesture recognition. For vehicle vibration, the close proximity of the human hand and the device significantly reduces the impact of vehicle vibration. Combined with the high-speed data transmission of the multimodal controller 130, it ensures that gesture recognition and tactile feedback remain stable during driving, solving problems such as inaccurate recognition and operational failure caused by environmental factors in traditional solutions. This comprehensive optimization from fixed layout to technical performance greatly improves the system's adaptability in complex in-vehicle environments.

[0085] Furthermore, when using the interactive device 100 located on the rear armrest 210 to collect user gestures, the distance between the user's gestures and the interactive device 100 can be 10-40cm to conform to the natural arm placement posture, avoiding the fatigue caused by the continuous arm raising required by traditional touchscreens or gesture recognition. This combination of ergonomic design and multimodal interaction allows passengers to operate without changing their relaxed posture. In terms of feedback mechanism, the depth camera 110 captures image data of gestures, and the controller 130 processes the image data to generate corresponding instructions (haptic feedback control instructions; the ultrasonic transmitter generates haptic feedback, allowing the user to instantly perceive the operation taking effect without needing to look at the screen for confirmation, solving the problem of broken interaction chains in simple gesture recognition). In addition, the multimodal integrated controller 130 ensures the consistency of operation response through collaborative algorithms (such as synchronous changes in vibration intensity when adjusting volume), and the retractable cover reduces accidental touches. In multi-person scenarios, it avoids instruction conflicts through area division. This combination of "hardware layout + feedback mechanism + algorithm optimization" not only meets the needs of a relaxed posture but also enhances the realism of operation through haptic feedback, providing an immersive experience that cannot be achieved by improving a single module.

[0086] It is worth mentioning that the interactive device 100 has flexible vehicle model adaptation capabilities, such as... Figure 8 As shown, for 5-seater models, the kit is deployed in the rear center armrest, as... Figure 9 As shown, for 6 / 7-seater models, the interactive device 100 can be deployed independently on the left and right sides of the second row; that is, the interactive device 100 can be adjusted in position according to the spatial configuration, structural characteristics and passenger operating habits of the rear row of different models to adapt to the differences of various models.

[0087] The smooth operation of the interaction process involved in the interaction method relies on the close interconnection of various components in the vehicle 20. The controller in the interaction device 100 can serve as the core, realizing data interaction with the entertainment screen and the window 220 through the communication interface. The interaction device 100 processes gestures and ultrasonic signals, the entertainment screen is responsible for display, and the window 220 control module controls the raising and lowering of the window 220. All components cooperate efficiently under the coordination of the controller. In summary, the ultrasonic visual control solution of this patented technology, in terms of hardware architecture, takes a highly integrated interactive device 100 as its core, combined with a reasonable overall layout and flexible vehicle adaptation, providing a solid foundation for interaction. In terms of interaction methods, it constructs an efficient interaction logic through a complete interaction process, tight module interconnection, and optimized experience design. The organic integration of hardware architecture and interaction methods achieves a comprehensive effect of efficient hardware collaboration, flexible scenario adaptation, and precise and smooth interaction, providing a practical solution for rear-seat human-computer interaction. Based on the same technical concept, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the interaction method as shown in the above embodiments of this disclosure.

[0088] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.

[0089] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.

[0090] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.

Claims

1. An interactive device, characterized in that, The interactive device includes: A depth camera is used to capture image data of the user's hand; An ultrasonic transmitting array, comprising multiple ultrasonic transmitters, for emitting ultrasonic waves; A controller, connected to both the ultrasonic transmitter and the depth camera, is configured to: The system receives the image data and determines the gesture semantics and the three-dimensional spatial position of the hand target area based on the image data; it generates tactile feedback control commands based on the gesture semantics; and it drives each ultrasonic transmitter in the ultrasonic transmitting array to emit focused ultrasonic waves toward the hand target area according to the three-dimensional spatial position and the tactile feedback control commands, so as to generate tactile feedback corresponding to the gesture semantics.

2. The interactive device according to claim 1, characterized in that, The gesture semantics include the feature changes of the gesture operation; the tactile feedback corresponding to the gesture semantics includes fixed-point tactile feedback and / or sliding tactile feedback. The fixed-point tactile feedback is used to instruct the ultrasonic emission array to form a sound pressure focus at a specific location in the target area of ​​the hand, and to dynamically adjust the intensity of the sound pressure focus according to the characteristic change. The sliding tactile feedback is used to instruct the ultrasonic emitting array to control the sound pressure focus formed by the ultrasonic emitting array to move along a specified direction or path on the target area of ​​the hand according to the characteristic change.

3. The interactive device according to claim 2, characterized in that, The fixed-point tactile feedback is generated in the following way: The controller is configured to: determine that the gesture semantics are a first type of gesture semantics for continuous intensity adjustment, and determine the target intensity value of the sound pressure focus based on the feature change amount; Based on the three-dimensional spatial position of the target area of ​​the hand, determine the required phase control parameters for each ultrasonic transmitter in the ultrasonic transmitting array; Based on the target intensity value, determine the amplitude control parameters for driving the ultrasonic transmitting array; Based on the phase control parameters and amplitude control parameters, the ultrasonic transmitting array is driven to emit ultrasonic waves to form a sound pressure focus in the target area of ​​the hand, and the intensity of the sound pressure focus corresponds to the target intensity value.

4. The interactive device according to claim 2, characterized in that, The haptic feedback is generated in the following way: The controller is configured to: determine that the gesture semantics are a second type of gesture semantics used to indicate a specified direction or path, and determine the movement vector information of the sound pressure focus based on the feature change amount; The three-dimensional spatial position of the hand target area is determined based on the image data acquired in real time by the depth camera; Based on the three-dimensional spatial position of the hand target area and the movement vector information, the three-dimensional spatial position of the sound pressure focus is determined; Based on the three-dimensional spatial position of the sound pressure focus, the required phase control parameters for each ultrasonic transmitter in the ultrasonic transmitting array are determined; Based on the phase control parameters, the ultrasonic transmitting array is driven to emit ultrasonic waves.

5. The interactive device according to claim 1, characterized in that, The controller is connected to multiple execution devices, and the controller is further configured to: Based on the gesture semantics, a device control command is generated, and the device control command is sent to the execution device corresponding to the gesture semantics, so that the execution device executes the device control command.

6. The interactive device according to claim 1, characterized in that, The depth camera and the ultrasonic transmitting array are integrated into a housing according to a specified positional relationship.

7. The interactive device according to claim 6, characterized in that, The housing is cylindrical in shape, and the depth camera is located in the central region of one end face of the cylindrical shape; the ultrasonic transmitting array includes multiple ultrasonic transmitters that are circumferentially distributed around the depth camera.

8. An interaction method, characterized in that, A controller applied in an interactive device, the method includes: Receive image data captured by a depth camera; Based on the image data, determine the semantics of the gesture and the three-dimensional spatial position of the target area of ​​the hand; Based on the gesture semantics, haptic feedback control commands are generated; Based on the three-dimensional spatial position and the tactile feedback control command, the ultrasonic wave emitting array is driven to emit focused ultrasonic waves toward the target area of ​​the hand to generate tactile feedback corresponding to the gesture semantics.

9. A vehicle, characterized in that, include: The vehicle body and the interactive device according to any one of claims 1-8.

10. The vehicle according to claim 9, characterized in that, The interactive device is located on the rear armrest of the vehicle body.