Remote control parking control method, system and device based on gestures and storage medium
By using Bluetooth positioning and on-board camera to identify driver gestures during remote parking, generating inquiry statements and re-planning the route after confirming, the problem of external gesture interference in remote parking is solved, and the accuracy and safety of control is improved.
Patent Information
- Application Number
- CN202510475301.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
AI Technical Summary
During remote parking, the gestures of other people around the vehicle can easily interfere with the driver's control command recognition, resulting in a decrease in the accuracy and safety of remote parking control.
The target vehicle's Bluetooth receiver triangularly determines the terminal position, uses the on-board camera to identify the driver's gestures, generate inquiry statements and broadcast reports, and re-plan the parking route after obtaining confirmation instructions to avoid external interference.
Improve the accuracy and safety of remote parking control, reducing external gesture interference and inaccurate identification problems.
Smart Images

Figure CN120335361A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle control, and in particular to a gesture-based remote parking control method, system, device and storage medium. Background Art
[0002] Remote parking is a technology that remotely controls a vehicle to complete the parking process outside the vehicle through a remote control device (such as a key, a mobile phone APP, etc.), and is applicable to parking scenarios that are narrow or inconvenient for manual operation. Within a certain range from the vehicle (usually about 10 meters), the driver can remotely start the vehicle through the remote control device and control the forward and backward movement of the vehicle and the left and right rotation of the steering wheel in real time to achieve parking.
[0003] During the remote parking process, the driver often needs to manually intervene in the forward, backward and steering of the vehicle in real time. To improve the operation convenience, some solutions support the driver to issue control commands by gestures. However, the gesture actions of other people around the vehicle may interfere with the vehicle's recognition of the driver's control commands, thus affecting the accuracy of remote parking control and the safety of remote parking; in addition, compared with directly obtaining control commands from a key or a mobile phone APP, the control commands based on gesture recognition still have the problem of insufficient accuracy, which further affects the accuracy of remote parking control and the safety of remote parking. Summary of the Invention
[0004] The purpose of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.
[0005] For this reason, one object of the embodiments of the present invention is to provide a gesture-based remote parking control method, which improves the accuracy of remote parking control and the safety of remote parking.
[0006] Another object of the embodiments of the present invention is to provide a gesture-based remote parking control system.
[0007] In order to achieve the above technical object, the technical solutions adopted by the embodiments of the present invention include:
[0008] In the first aspect, the embodiments of the present invention provide a gesture-based remote parking control method, including the following steps:
[0009] In response to the remote parking start instruction of the driver on the target terminal, the target vehicle is used to plan a parking route to obtain an initial parking route, and the target vehicle is controlled for parking according to the initial parking route;
[0010] The target terminal is triangulated through a plurality of Bluetooth receivers of the target vehicle to obtain the current position of the target terminal;
[0011] Obtain the first image information of the current position through the on-vehicle camera of the target vehicle, identify the target gesture action of the driver according to the first image information, and determine the corresponding target intervention intention according to the target gesture action;
[0012] Generate a corresponding inquiry statement according to the target intervention intention, and push the inquiry statement to the target terminal through the target vehicle for broadcasting;
[0013] In response to the confirmation instruction of the driver, re-plan the parking route according to the target intervention intention to obtain the target parking route, and perform parking control on the target vehicle according to the target parking route.
[0014] Further, in an embodiment of the present invention, the response to the remote control parking start instruction of the driver on the target terminal, perform parking route planning through the target vehicle to obtain the initial parking route, and perform parking control on the target vehicle according to the initial parking route, which specifically includes:
[0015] Establish a Bluetooth communication connection between the target terminal and the target vehicle, and perform operation permission authentication on the target terminal;
[0016] When receiving the remote control parking start instruction issued by the target terminal and the target terminal has the parking operation permission, obtain the surrounding environment information through the target vehicle, and perform parking route rules according to the surrounding environment information to obtain the initial parking route;
[0017] Drive the target vehicle to park according to the initial parking route.
[0018] Further, in an embodiment of the present invention, the triangulation of the target terminal by multiple Bluetooth receivers of the target vehicle to obtain the current position of the target terminal, which specifically includes:
[0019] Send a heartbeat signal from the target terminal to the target vehicle, and receive the heartbeat signal through the Bluetooth receiver;
[0020] Determine the distance between the target terminal and each Bluetooth receiver according to the heartbeat signal, and then obtain the current position through triangulation.
[0021] Further, in an embodiment of the present invention, the obtaining the first image information of the current position through the on-vehicle camera of the target vehicle, identifying the target gesture action of the driver according to the first image information, and determining the corresponding target intervention intention according to the target gesture action, which specifically includes:
[0022] Determine the shooting orientation and shooting focal length according to the current position, and control the vehicle-mounted camera to shoot the current position according to the shooting orientation and the shooting focal length to obtain the first image information;
[0023] Extract the foreground image from the first image information to obtain the human body image information of the driver;
[0024] Perform human key point detection on the human body image information to obtain a hand region image;
[0025] Determine the hand image time series data according to the hand region images corresponding to multiple consecutive frames of the first image information;
[0026] Input the hand image time series data into a pre-trained gesture recognition model to obtain the target gesture action;
[0027] Match the target gesture action according to a preset gesture action database to obtain the target intervention intention corresponding to the target gesture action.
[0028] Further, in an embodiment of the present invention, the gesture recognition model is trained through the following steps:
[0029] Obtain hand image time series samples of multiple testers, and determine gesture action labels corresponding to each hand image time series sample through manual annotation;
[0030] Input the hand image time series samples into a pre-constructed CNN-LSTM hybrid neural network to obtain a gesture action recognition result;
[0031] Determine a loss value according to the gesture action recognition result and the gesture action label;
[0032] Update the parameters of the CNN-LSTM hybrid neural network according to the loss value through the backpropagation algorithm to obtain the trained gesture recognition model;
[0033] Wherein, the CNN-LSTM hybrid neural network includes an input layer, a CNN convolutional layer, a feature fusion layer, an LSTM layer, an attention layer and an output layer. The input layer is used to input the hand image time series samples. The CNN convolutional layer is used to extract features from the hand image time series samples to obtain local time series features. The feature fusion layer is used to fuse the local time series features to obtain fused time series features. The LSTM layer is used to generate a hidden state sequence according to the fused time series features. The attention layer is used to perform dynamic weight allocation on each dimension of the hidden state sequence based on the multi-head self-attention mechanism. The output layer is used to map the hidden state sequence after dynamic weight allocation to the gesture action recognition result.
[0034] Further, in an embodiment of the present invention, generating a corresponding inquiry statement according to the target intervention intention and pushing the inquiry statement to the target terminal through the target vehicle for broadcasting specifically includes:
[0035] Obtain the current parking state of the target vehicle, perform intervention simulation according to the initial parking route, the current parking state, and the target intervention intention to obtain a target intervention result;
[0036] Generate the inquiry statement according to the target intervention result and a preset inquiry template, and push the inquiry statement to the target terminal through the target vehicle;
[0037] Display the inquiry statement through the display screen of the target terminal, and / or play the inquiry statement through the voice playback device of the target terminal.
[0038] Further, in an embodiment of the present invention, in response to the confirmation instruction of the driver, re-planning the parking route according to the target intervention intention to obtain a target parking route, and performing parking control on the target vehicle according to the target parking route specifically includes:
[0039] Obtain the confirmation instruction of the driver through voice recognition, gesture recognition, or key operation of the target terminal;
[0040] Obtain the current environment information and the current parking state through the target vehicle, and perform parking route planning according to the current environment information, the current parking state, and the target intervention intention to obtain the target parking route;
[0041] Drive the target vehicle to park according to the target parking route.
[0042] In a second aspect, an embodiment of the present invention provides a gesture-based remote control parking control system, including:
[0043] A parking start module, configured to respond to a remote control parking start instruction of a driver on a target terminal, perform parking route planning through a target vehicle to obtain an initial parking route, and perform parking control on the target vehicle according to the initial parking route;
[0044] A terminal positioning module, configured to perform triangulation positioning on the target terminal through a plurality of Bluetooth receivers of the target vehicle to obtain the current position of the target terminal;
[0045] An intervention intention recognition module, configured to obtain first image information of the current position through an in-vehicle camera of the target vehicle, recognize a target gesture action of the driver according to the first image information, and determine a corresponding target intervention intention according to the target gesture action;
[0046] An inquiry module, configured to generate a corresponding inquiry statement according to the target intervention intention, and push the inquiry statement to the target terminal through the target vehicle for broadcasting;
[0047] A parking route adjustment module, configured to, in response to the confirmation instruction of the driver, re-plan a parking route according to the target intervention intention to obtain a target parking route, and perform parking control on the target vehicle according to the target parking route.
[0048] In a third aspect, an embodiment of the present invention provides a gesture-based remote parking control device, including:
[0049] At least one processor;
[0050] At least one memory, configured to store at least one program;
[0051] When the at least one program is executed by the at least one processor, the at least one processor is caused to implement the above-mentioned gesture-based remote parking control method.
[0052] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to execute the above-mentioned gesture-based remote parking control method when executed by the processor.
[0053] The advantages and beneficial effects of the present invention will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present invention:
[0054] In an embodiment of the present invention, in response to a remote parking start instruction of a driver on a target terminal, an initial parking route is obtained by performing parking route planning on a target vehicle, and parking control is performed on the target vehicle according to the initial parking route. The current position of the target terminal is obtained by performing triangulation positioning on the target terminal through multiple Bluetooth receivers of the target vehicle. First image information of the current position is obtained through an in-vehicle camera of the target vehicle. A target gesture action of the driver is recognized according to the first image information, and a corresponding target intervention intention is determined according to the target gesture action. A corresponding inquiry statement is generated according to the target intervention intention, and the inquiry statement is pushed to the target terminal by the target vehicle for broadcasting. In response to the confirmation instruction of the driver, the parking route is re-planned according to the target intervention intention to obtain a target parking route, and parking control is performed on the target vehicle according to the target parking route. In the embodiment of the present invention, the start of remote parking is controlled through a user terminal. During remote parking, the vehicle can determine the real-time position of the user terminal through triangulation positioning and obtain the image information of the driver at the real-time position. Based on the recognized gesture action, the intervention intention of the driver is determined, a corresponding inquiry statement is generated and broadcast through the target terminal. After the driver issues a confirmation instruction, the parking route is re-planned according to the intervention intention, avoiding the interference of gesture actions of other people around the vehicle and the problem of inaccurate recognition of control instructions, and improving the accuracy of remote parking control and the safety of remote parking. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduces the drawings required to be used in the embodiments of the present invention. It should be understood that the drawings introduced below are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0056] Figure 1 It is a flowchart of the steps of a gesture-based remote parking control method provided by an embodiment of the present invention;
[0057] Figure 2 It is a schematic structural diagram of a CNN-LSTM hybrid neural network provided by an embodiment of the present invention;
[0058] Figure 3 It is a structural block diagram of a gesture-based remote parking control system provided by an embodiment of the present invention;
[0059] Figure 4 It is a structural block diagram of a gesture-based remote parking control device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0061] In the description of the present invention, the meaning of "a plurality of" is two or more. If the first and second are described, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art of this technical field.
[0062] Referring to Figure 1 , an embodiment of the present invention provides a gesture-based remote parking control method, which specifically includes the following steps:
[0063] S101. In response to a remote parking start instruction of a driver on a target terminal, perform parking route planning through a target vehicle to obtain an initial parking route, and perform parking control on the target vehicle according to the initial parking route;
[0064] S102. Perform triangulation positioning on the target terminal through a plurality of Bluetooth receivers of the target vehicle to obtain the current position of the target terminal;
[0065] S103. Obtain first image information of the current position through an in-vehicle camera of the target vehicle, identify the target gesture action of the driver according to the first image information, and determine the corresponding target intervention intention according to the target gesture action;
[0066] S104. Generate a corresponding inquiry statement according to the target intervention intention, and push the inquiry statement to the target terminal through the target vehicle for broadcast;
[0067] S105. In response to the confirmation instruction of the driver, re-perform parking route planning according to the target intervention intention to obtain a target parking route, and perform parking control on the target vehicle according to the target parking route.
[0068] In an embodiment of the present invention, the start of remote parking is controlled by a user terminal. During the remote parking process, the vehicle can determine the real-time position of the user terminal through trilateration and obtain the image information of the driver at the real-time position. Based on the recognized gesture actions, the intervention intention of the driver is determined, a corresponding inquiry statement is generated and broadcast through the target terminal. After the driver issues a confirmation instruction, the parking route is replanned according to the intervention intention, avoiding the interference of gesture actions of other people around the vehicle and the inaccurate recognition of control instructions, and improving the accuracy of remote parking control and the safety of remote parking.
[0069] Further as an optional implementation manner, in response to the remote parking start instruction of the driver on the target terminal, the target vehicle performs parking route planning to obtain an initial parking route, and performs parking control on the target vehicle according to the initial parking route, which specifically includes:
[0070] S1011. Establish a Bluetooth communication connection between the target terminal and the target vehicle, and perform operation permission authentication on the target terminal;
[0071] S1012. When receiving the remote parking start instruction issued by the target terminal and the target terminal has the parking operation permission, obtain the surrounding environment information through the target vehicle, and perform parking route rules according to the surrounding environment information to obtain an initial parking route;
[0072] S1013. Drive the target vehicle to park according to the initial parking route.
[0073] Specifically, in an embodiment of the present invention, the start of remote parking is controlled by a user terminal (key, mobile phone), and the specific process is as follows:
[0074] 1) Device discovery and pairing
[0075] Bluetooth broadcast and scanning: The target vehicle continuously broadcasts Bluetooth signals (such as vehicle control service identification) containing specific service UUIDs, and the target terminal (such as a mobile phone) discovers the vehicle through BLE scanning.
[0076] Secure pairing: The terminal and the vehicle complete pairing through ECDH key exchange or a preset PIN code (such as 1234) to generate an encrypted link key (Link Key) to ensure communication security.
[0077] Service channel establishment: The terminal queries the vehicle Bluetooth service through SDP, obtains the RFCOMM channel number, and establishes an L2CAP / RFCOMM data channel.
[0078] 2) Operation permission authentication
[0079] Digital certificate verification: The terminal needs to send a digitally signed certificate (such as TLS mutual authentication) to the vehicle, and the vehicle decrypts and verifies the validity of the certificate.
[0080] Dynamic token authorization: If the user has authorized the parking permission through the vehicle networking platform, the vehicle obtains a dynamic token from the cloud and matches it with the token submitted by the terminal.
[0081] 3) Instruction triggering and permission verification: The user triggers the "remote parking" instruction through the terminal App on the key, and the terminal sends the instruction and timestamp through the encrypted Bluetooth channel; the vehicle verifies the validity of the instruction signature and checks whether the terminal is in the authorization list (such as through the VIN code binding relationship).
[0082] 4) Environmental data collection and fusion
[0083] Multi-sensor collaboration: The vehicle activates ultrasonic radars (detecting close-range obstacles), surround-view cameras (constructing a 3D parking space model), and millimeter-wave radars (monitoring dynamic objects).
[0084] SLAM real-time mapping: Generate a high-precision semantic map around the vehicle through visual SLAM algorithms, and identify key information such as parking lines and obstacle boundaries.
[0085] 5) Path planning algorithm
[0086] Initial route generation: Based on the RRT algorithm, combined with the parking space size and vehicle kinematic model, calculate a collision-free path and optimize the number of steering and reversing times.
[0087] Dynamic obstacle avoidance adjustment: If pedestrians or moving obstacles are detected, use the DWA (Dynamic Window Approach) to adjust the path in real time to ensure a safe distance.
[0088] 6) Vehicle motion control
[0089] Wired control instruction issuance: Send steering angle and vehicle speed instructions to EPS (Electric Power Steering), ESP (Electronic Stability Program), and EPB (Electronic Parking Brake) through the CAN bus.
[0090] Closed-loop feedback control: Use data from the IMU (Inertial Measurement Unit) and wheel speed sensors to correct the path tracking error in real time (such as PID control).
[0091] Furthermore, as an optional implementation method, triangulate the target terminal through multiple Bluetooth receivers of the target vehicle to obtain the current position of the target terminal, which specifically includes:
[0092] S1021. Send a heartbeat signal from the target terminal to the target vehicle and receive the heartbeat signal through the Bluetooth receiver;
[0093] S1022. Determine the distances between the target terminal and each Bluetooth receiver based on the heartbeat signal, and then obtain the current position through triangulation.
[0094] Specifically, in the embodiments of the present invention, multiple Bluetooth receivers are arranged at different positions inside the target vehicle, and the position of the user terminal is obtained by triangulating the user terminal through multiple Bluetooth receivers, that is, the position of the driver who issues the remote parking start instruction. The specific process is as follows:
[0095] 1) Heartbeat signal design and transmission
[0096] Signal content and format: The target terminal (such as a mobile phone or in-vehicle device) periodically sends a Bluetooth broadcast packet containing a unique identifier (UUID, Major, Minor) as a heartbeat signal; the data packet needs to conform to the Bluetooth Low Energy (BLE) protocol, and the transmission interval is optimized (such as 1 second) to balance power consumption and positioning accuracy.
[0097] Signal sending mechanism: Use a timer or an independent thread to control the sending period to avoid signal interruption caused by blocking of the main program; realize non-connected signal transmission through the broadcast mode of the Bluetooth chip to reduce power consumption.
[0098] 2) Bluetooth receiver deployment and signal reception
[0099] Receiver layout principle: Deploy at least 3 Bluetooth receivers (Beacon nodes) inside the target vehicle, follow a triangular grid layout, control the spacing within 4 - 8 meters, avoid metal shielding areas (such as car doors, engine compartments), and the height is recommended to be 2.5 - 3 meters (such as the roof or rearview mirror position) to reduce multipath interference.
[0100] Signal reception and parsing: The receiver captures the heartbeat signal through the scanning mode, extracts the signal strength (RSSI) and terminal identification information; uses a filtering algorithm (such as Kalman filtering) to eliminate the influence of environmental noise on RSSI.
[0101] 3) Distance calculation and error correction: Use the RSSI ranging model to calculate the distances between the target terminal and each Bluetooth receiver, and perform error correction.
[0102] The RSSI ranging model is as follows:
[0103] d = 10 (∣RSSI∣-A) / (10×n)
[0104] where A is the reference RSSI value at 1 meter, and n is the environmental attenuation factor (usually taken as 2 - 4). For the complex electromagnetic environment inside the vehicle, the attenuation factor n and the reference value A are dynamically adjusted through measured data.
[0105] 4) Triangulation implementation
[0106] Coordinate positioning calculation: Given the coordinates of each receiver (which need to be calibrated in advance), draw a circle with each receiver as the center and the calculated distance as the radius. The intersection of the three circles is the position of the target terminal; the least squares method or geometric optimization algorithm is used to handle the situation where the multiple circles do not strictly intersect at one point.
[0107] Doppler effect compensation: If the vehicle is in a moving state, signal distortion in the dynamic scenario needs to be eliminated through phase correction or frequency shift compensation.
[0108] Furthermore, as an optional implementation manner, first image information of the current position is obtained through an in-vehicle camera of the target vehicle, a target gesture action of the driver is recognized according to the first image information, and a corresponding target intervention intention is determined according to the target gesture action, which specifically includes:
[0109] S1031. Determine the shooting direction and shooting focal length according to the current position, and control the in-vehicle camera to shoot the current position according to the shooting direction and shooting focal length to obtain the first image information;
[0110] S1032. Extract the foreground image from the first image information to obtain the human body image information of the driver;
[0111] S1033. Detect the key points of the human body in the human body image information to obtain the hand area image;
[0112] S1034. Determine the hand image time series data according to the hand area images corresponding to multiple consecutive frames of the first image information;
[0113] S1035. Input the hand image time series data into a pre-trained gesture recognition model to obtain the target gesture action;
[0114] S1036. Match the target gesture action according to a preset gesture action database to obtain the target intervention intention corresponding to the target gesture action.
[0115] Specifically, in the embodiment of the present invention, an image of the current position of the target terminal is captured by the in-vehicle camera of the target vehicle, the hand area image is extracted through image processing, the time series data is formed, and the target gesture action is recognized through a pre-trained gesture recognition model, and then the corresponding target intervention intention is obtained by matching with a pre-constructed gesture action database. The specific process is as follows:
[0116] 1) Image acquisition
[0117] Determine the shooting direction and focal length: Based on the current position of the target terminal obtained in the previous step, determine the direction and distance of the target terminal relative to the in-vehicle camera, so as to adjust the shooting direction and determine the appropriate shooting focal length to ensure that the hand movements of the driver can be clearly captured.
[0118] Controlling camera shooting: Sending the determined shooting direction and focal length parameters to the vehicle-mounted camera. The camera adjusts its own angle and focal length according to these parameters, shoots the current position of the driver, and obtains the first image information.
[0119] 2) Foreground image extraction
[0120] Image preprocessing: Preprocess the first image information, including grayscale, filtering and other operations. Grayscale can convert a color image into a grayscale image to reduce the amount of data; filtering operations (such as Gaussian filtering) can remove noise in the image and improve image quality.
[0121] Foreground extraction: Use background subtraction algorithms, such as Gaussian mixture models (GMM), to separate the driver's human image from the background. The algorithm performs statistical modeling on multiple frames of images, learns the distribution characteristics of the background, and then compares the current frame with the background model, extracting the part with larger differences as the foreground (i.e., the driver's human image) to obtain human image information.
[0122] 3) Human key point detection
[0123] Select detection algorithm: Use mature human key point detection algorithms, such as OpenPose or Mediapipe. These algorithms analyze human image information through convolutional neural networks (CNN) and can accurately detect key points of the human body, such as wrists and finger joints.
[0124] Extract the hand area image: According to the detected key points of the human body, locate the key points of the hand, and then based on these key points, intercept the image area containing the hand to obtain the hand area image.
[0125] 4) Determine the hand image time series data
[0126] Multi-frame image acquisition: continuously acquire multiple frames of first image information within a period of time, and extract the hand area image corresponding to each frame of the image according to the method in the above steps.
[0127] Constructing time series data: Arrange multiple consecutive frames of hand area images in time sequence to form hand image time series data. This data can reflect the changing process of hand movements.
[0128] 5) Gesture Recognition
[0129] Model selection and training: Select a suitable gesture recognition model, such as a model based on a recurrent neural network (RNN) or a long short-term memory network (LSTM). Use a large number of hand image time series data samples to train the model and adjust the model parameters so that it can accurately recognize different gestures.
[0130] Gesture prediction: Input the obtained time-series data of hand images into a pre-trained gesture recognition model. The model makes predictions based on the input data and outputs the target gesture actions.
[0131] 6) Intention matching
[0132] Establish a gesture action database: Collect various common parking gesture actions in advance, and define corresponding intervention intentions for each gesture action, such as "pause", "forward", "backward", "left turn", "right turn", etc., to construct a gesture action database.
[0133] Match the target intervention intention: Match the target gesture action with the gesture actions in the gesture action database, find the most similar gesture action, and obtain the target intervention intention corresponding to this gesture action.
[0134] Furthermore, as an optional implementation manner, the gesture recognition model is trained through the following steps:
[0135] S201. Obtain time-series samples of hand images of multiple testers, and determine the gesture action labels corresponding to each time-series sample of hand images through manual annotation;
[0136] S202. Input the time-series samples of hand images into a pre-constructed CNN-LSTM hybrid neural network to obtain gesture action recognition results;
[0137] S203. Determine the loss value according to the gesture action recognition results and the gesture action labels;
[0138] S204. Update the parameters of the CNN-LSTM hybrid neural network through the backpropagation algorithm according to the loss value to obtain a trained gesture recognition model;
[0139] Among them, the CNN-LSTM hybrid neural network includes an input layer, a CNN convolutional layer, a feature fusion layer, an LSTM layer, an attention layer, and an output layer. The input layer is used to input time-series samples of hand images. The CNN convolutional layer is used to extract local time-series features from the time-series samples of hand images. The feature fusion layer is used to fuse the local time-series features to obtain fused time-series features. The LSTM layer is used to generate a hidden state sequence according to the fused time-series features. The attention layer is used to dynamically allocate weights to each dimension of the hidden state sequence based on the multi-head self-attention mechanism. The output layer is used to map the hidden state sequence after dynamic weight allocation to gesture action recognition results.
[0140] Specifically, obtain the time-series samples of hand images of multiple testers during gesture parking control, and determine the gesture action labels corresponding to each time-series sample of hand images through manual annotation; input the time-series samples of hand images into the CNN-LSTM hybrid neural network to obtain the gesture action recognition result; determine the loss value according to the gesture action recognition result and the gesture action label; update the parameters of the CNN-LSTM hybrid neural network through the backpropagation algorithm based on the loss value, and then a trained gesture recognition model can be obtained.
[0141] The CNN-LSTM hybrid neural network combines the advantages of the convolutional neural network (CNN) in spatial feature extraction and the long short-term memory network (LSTM) in temporal dependence modeling, and is widely used in scenarios such as time series prediction and video analysis. The following introduces its structure and training process.
[0142] As Figure 2 shown in the structural schematic diagram of the CNN-LSTM hybrid neural network provided by an embodiment of the present invention, the input data is standardized through the input layer and then input into the CNN convolutional layer. The one-dimensional convolutional kernel is used to extract local temporal features. The extracted local temporal features undergo the feature fusion operation of the feature fusion layer to generate fused temporal features. Then, the fused temporal features are used as the feature input of the LSTM layer to capture long-term dependencies and generate a hidden state sequence. Then, based on the multi-head self-attention mechanism, the attention weights (such as the action weights of the wrist joint and finger joints) are calculated through the SoftMax function, and dynamic weight allocation is performed on each dimension of the hidden state sequence output by the LSTM layer. The hidden state sequence after dynamic weight allocation is mapped to the gesture action recognition result through the output layer. Then, the loss value is determined in combination with the gesture action label, and the loss value is backpropagated using the Adam algorithm to gradually update the model parameters layer by layer. The loss function can use binary cross entropy or a weighted loss function (to address the problem of data imbalance).
[0143] The above describes the training process of the gesture recognition model. Input the time-series data of hand images obtained in the previous steps into the trained gesture recognition model, and the target gesture action of the driver inferred by the model can be obtained.
[0144] Further, as an optional implementation, generate a corresponding inquiry statement according to the target intervention intention, and push the inquiry statement to the target terminal through the target vehicle for broadcast, which specifically includes:
[0145] S1041. Obtain the current parking state of the target vehicle, perform intervention simulation according to the initial parking route, the current parking state, and the target intervention intention to obtain the target intervention result;
[0146] S1042. Generate an inquiry statement based on the target intervention result and a preset inquiry template, and push the inquiry statement to the target terminal through the target vehicle;
[0147] S1043. Display the inquiry statement through the display screen of the target terminal, and / or play the inquiry statement through the voice playback device of the target terminal.
[0148] Specifically, after determining the target intervention intention of the driver, the embodiment of the present invention does not directly execute it. Instead, it first performs an intervention simulation based on the target intervention intention of the driver to obtain the target intervention result, generates an inquiry statement based on the target intervention result and plays it through the target terminal to confirm whether to execute the intervention, so as to avoid the interference caused by the gestures of other people around the vehicle and the potential hazards brought by gesture recognition errors. The specific process is as follows:
[0149] 1) Parking state acquisition and intervention simulation
[0150] Real-time data collection: Obtain parameters such as the current position of the vehicle, the distance to obstacles, and the heading angle through in-vehicle sensors (ultrasonic radar, surround view camera); combine the initial parking route to match the current parking progress and determine whether it deviates from the expected trajectory.
[0151] Intention simulation: According to the identified target intervention intention (such as parking on the far right), generate a corrected trajectory by combining model predictive control (MPC); detect dynamic obstacles (such as pedestrians passing by) or static obstacles (such as parking posts), and evaluate the collision risk after the intervention through multi-modal data fusion (camera image + ultrasonic signal) to obtain the target intervention result.
[0152] 2) Inquiry statement generation and push
[0153] Template-based interaction design: Preset multiple inquiry templates. For example, for the case where there is a collision risk after the intervention, the preset inquiry template is "A collision risk is detected after the intervention. Do you still want to execute the intervention command to park on the far right?", and for the case where there is no collision risk after the intervention, the preset inquiry template is "The vehicle is safe after the intervention. Are you sure you want to execute the intervention command to park on the far right?"; Combine natural language processing (NLP) to dynamically fill in parameters (such as obstacle type, distance, etc.) to generate a complete statement.
[0154] Information push: The vehicle end pushes the inquiry statement to the user terminal (key, mobile phone) through Bluetooth communication. The user terminal can display the inquiry statement through the display screen and can also trigger voice broadcast to remind the driver to confirm.
[0155] Further as an optional implementation manner, in response to the driver's confirmation command, re-plan the parking route according to the target intervention intention to obtain the target parking route, and perform parking control on the target vehicle according to the target parking route, which specifically includes:
[0156] S1051. Obtain the confirmation instruction of the driver through voice recognition, gesture recognition or key operation of the target terminal;
[0157] S1052. Obtain the current environmental information and current parking state through the target vehicle, plan the parking route according to the current environmental information, current parking state and target intervention intention, and obtain the target parking route;
[0158] S1053. Drive the target vehicle to park according to the target parking route.
[0159] Specifically, when a voice confirmation instruction (such as "yes", "confirm execution"), a gesture confirmation instruction (a pre-agreed confirmation gesture) issued by the driver is detected, or the confirmation button of the target terminal is triggered, the vehicle obtains the confirmation instruction of the driver, and then confirms to execute the target operation intention, that is, re-plan the parking route according to the current environmental information, current parking state and target intervention intention, and obtain the target parking route. The specific process is as follows:
[0160] 1) Real-time acquisition of environmental information and parking state
[0161] Multi-sensor data fusion: Obtain environmental information such as the distribution of obstacles around the vehicle, the position of the parking space line, and the size of the parking space through devices such as ultrasonic radars, cameras, and lidar.
[0162] Dynamic evaluation of parking state: The system judges the current parking state, that is, the progress of the execution of the initial parking route, based on parameters such as the current vehicle position (such as the center coordinate of the rear axle), steering angle, and driving speed, in combination with the initial parking route.
[0163] 2) Execution of the dynamic path re-planning algorithm
[0164] Path feasibility verification: Combine the real-time environmental information and the driver's intention, and verify whether the new path meets conditions such as the vehicle's steering ability and the minimum parking space size through a kinematic model and a collision detection algorithm.
[0165] Multi-objective optimization path generation:
[0166] (1) Trajectory correction based on real-time positioning: Generate a local trajectory using the current vehicle data and intervention intention, and ensure smoothness through iterative optimization (such as Bezier curves or spline interpolation);
[0167] (2) Global path reconstruction: When the initial parking route is completely infeasible (such as the intervention intention is to change the parking space), use the RRT algorithm to re-plan the global path.
[0168] 3) Path parameter output
[0169] Generate a target parking route that includes a steering angle sequence, a speed curve, and a heading angle change, and convert it into control instructions recognizable by the vehicle execution layer.
[0170] 4) Complete the gesture control of remote parking through the control instructions generated by the target vehicle execution.
[0171] The above describes the method steps of the embodiments of the present invention. It can be understood that in the embodiments of the present invention, the start of remote parking is controlled through the user terminal. During the remote parking process, the vehicle can determine the real-time position of the user terminal through triangulation and obtain the image information of the driver at the real-time position. Based on the recognized gesture actions, the intervention intention of the driver is determined, corresponding inquiry statements are generated and broadcast through the target terminal. After the driver issues a confirmation instruction, the parking route is re-planned according to the intervention intention, avoiding the interference of the gesture actions of other people around the vehicle and the inaccurate recognition of control instructions, and improving the accuracy of remote parking control and the safety of remote parking.
[0172] Refer to Figure 3 , the embodiments of the present invention provide a gesture-based remote parking control system, including:
[0173] A parking start module, configured to respond to a remote parking start instruction of a driver on a target terminal, perform parking route planning through a target vehicle to obtain an initial parking route, and perform parking control on the target vehicle according to the initial parking route;
[0174] A terminal positioning module, configured to perform triangulation on the target terminal through multiple Bluetooth receivers of the target vehicle to obtain the current position of the target terminal;
[0175] An intervention intention recognition module, configured to obtain first image information of the current position through an in-vehicle camera of the target vehicle, recognize the target gesture actions of the driver according to the first image information, and determine the corresponding target intervention intention according to the target gesture actions;
[0176] An inquiry module, configured to generate corresponding inquiry statements according to the target intervention intention, and push the inquiry statements to the target terminal through the target vehicle for broadcasting;
[0177] A parking route adjustment module, configured to respond to a confirmation instruction of the driver, re-plan the parking route according to the target intervention intention to obtain a target parking route, and perform parking control on the target vehicle according to the target parking route.
[0178] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0179] Reference Figure 4 , an embodiment of the present invention provides a gesture-based remote parking control device, including:
[0180] At least one processor;
[0181] At least one memory for storing at least one program;
[0182] When the above at least one program is executed by the above at least one processor, the above at least one processor implements the above gesture-based remote parking control method.
[0183] The content in the above method embodiment is applicable to the device embodiment of the present invention. The functions specifically implemented by the device embodiment of the present invention are the same as those in the above method embodiment, and the beneficial effects achieved are also the same as those in the above method embodiment.
[0184] An embodiment of the present invention also provides a computer-readable storage medium, in which a program executable by a processor is stored. The program executable by the processor is used to execute the above gesture-based remote parking control method when executed by the processor.
[0185] A computer-readable storage medium according to an embodiment of the present invention can execute a gesture-based remote parking control method provided by an embodiment of the method of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.
[0186] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.
[0187] In some alternative embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two consecutive blocks shown can actually be executed substantially simultaneously or the above blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are foreseeable, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.
[0188] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above-described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0189] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0190] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a predefined sequence of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0191] More specific examples (nonexhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the above programs can be printed, because the above programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0192] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0193] In the above description of this specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0194] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
[0195] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A gesture-based remote parking control method, characterized in that, It includes the following steps: In response to the remote parking start instruction of the driver on the target terminal, the target vehicle is used to plan a parking route to obtain an initial parking route, and the target vehicle is controlled for parking according to the initial parking route; The target terminal is triangulated by multiple Bluetooth receivers of the target vehicle to obtain the current position of the target terminal; The on-vehicle camera of the target vehicle is used to obtain the first image information of the current position, the target gesture action of the driver is identified according to the first image information, and the corresponding target intervention intention is determined according to the target gesture action; An inquiry statement corresponding to the target intervention intention is generated, and the inquiry statement is pushed by the target vehicle to the target terminal for broadcasting; In response to the confirmation instruction of the driver, the parking route is re-planned according to the target intervention intention to obtain a target parking route, and the target vehicle is controlled for parking according to the target parking route.
2. The method for gesture-based remote parking control according to claim 1, wherein The step of, in response to the remote parking start instruction of the driver on the target terminal, using the target vehicle to plan a parking route to obtain an initial parking route, and controlling the target vehicle for parking according to the initial parking route specifically includes: Establish a Bluetooth communication connection between the target terminal and the target vehicle, and perform an operation permission authentication on the target terminal; When the remote parking start instruction sent by the target terminal is received and the target terminal has the parking operation permission, the target vehicle is used to obtain the surrounding environment information, and the parking route is planned according to the surrounding environment information to obtain the initial parking route; Drive the target vehicle for parking according to the initial parking route.
3. The method for controlling gesture-based remote parking according to claim 1, characterized in that, The step of triangulating the target terminal by multiple Bluetooth receivers of the target vehicle to obtain the current position of the target terminal specifically includes: The target terminal sends a heartbeat signal to the target vehicle, and the heartbeat signal is received by the Bluetooth receiver; The distance between the target terminal and each Bluetooth receiver is determined according to the heartbeat signal, and then the current position is obtained through triangulation.
4. A gesture-based remote parking control method according to claim 1, characterized in that, The step of using the on-vehicle camera of the target vehicle to obtain the first image information of the current position, identifying the target gesture action of the driver according to the first image information, and determining the corresponding target intervention intention according to the target gesture action specifically includes: The shooting direction and shooting focal length are determined according to the current position, and the on-vehicle camera is controlled to shoot the current position according to the shooting direction and the shooting focal length to obtain the first image information; Foreground image extraction is performed on the first image information to obtain the human body image information of the driver; Human body key point detection is performed on the human body image information to obtain a hand region image; The hand image time series data is determined according to the hand region images corresponding to multiple consecutive frames of the first image information; The hand image time series data is input into a pre-trained gesture recognition model to obtain the target gesture action; Match the target gesture action according to a preset gesture action database to obtain the target intervention intention corresponding to the target gesture action.
5. A gesture-based remote parking control method according to claim 1, characterized in that, The gesture recognition model is trained through the following steps: Obtain time-sequence samples of hand images of multiple testers, and determine gesture action labels corresponding to each of the time-sequence samples of hand images through manual annotation; Input the time-sequence samples of hand images into a pre-constructed CNN-LSTM hybrid neural network to obtain gesture action recognition results; Determine a loss value according to the gesture action recognition results and the gesture action labels; Update the parameters of the CNN-LSTM hybrid neural network according to the loss value through the backpropagation algorithm to obtain the trained gesture recognition model; Among them, the CNN-LSTM hybrid neural network includes an input layer, a CNN convolutional layer, a feature fusion layer, an LSTM layer, an attention layer, and an output layer. The input layer is used to input the time-sequence samples of hand images. The CNN convolutional layer is used to extract local time-sequence features from the time-sequence samples of hand images. The feature fusion layer is used to fuse the local time-sequence features to obtain fused time-sequence features. The LSTM layer is used to generate a hidden state sequence according to the fused time-sequence features. The attention layer is used to perform dynamic weight allocation on each dimension of the hidden state sequence based on the multi-head self-attention mechanism. The output layer is used to map the hidden state sequence after dynamic weight allocation to the gesture action recognition results.
6. The method for controlling gesture-based remote parking according to claim 1, wherein, Generating a corresponding inquiry statement according to the target intervention intention, and pushing the inquiry statement to the target terminal for broadcasting through the target vehicle, specifically including: Obtain the current parking state of the target vehicle, perform intervention simulation according to the initial parking route, the current parking state, and the target intervention intention to obtain a target intervention result; Generate the inquiry statement according to the target intervention result and a preset inquiry template, and push the inquiry statement to the target terminal through the target vehicle; Display the inquiry statement on the display screen of the target terminal, and / or play the inquiry statement through the voice playback device of the target terminal.
7. A gesture-based remote parking control method according to any one of claims 1 to 6, characterized in that Responding to the confirmation instruction of the driver, re-planning the parking route according to the target intervention intention to obtain a target parking route, and performing parking control on the target vehicle according to the target parking route, specifically including: Obtain the confirmation instruction of the driver through voice recognition, gesture recognition, or key operation of the target terminal; Obtain the current environmental information and the current parking state through the target vehicle, and perform parking route planning according to the current environmental information, the current parking state, and the target intervention intention to obtain the target parking route; Drive the target vehicle to park according to the target parking route.
8. A gesture-based remote parking control system, characterized in that, Including: A parking start module, configured to, in response to a remote parking start instruction of a driver on a target terminal, perform parking route planning through a target vehicle to obtain an initial parking route, and perform parking control on the target vehicle according to the initial parking route; A terminal positioning module, configured to perform trilateration positioning on the target terminal through a plurality of Bluetooth receivers of the target vehicle to obtain the current position of the target terminal; An intervention intention recognition module, configured to obtain first image information of the current position through an in-vehicle camera of the target vehicle, recognize a target gesture action of the driver according to the first image information, and determine a corresponding target intervention intention according to the target gesture action; An inquiry module, configured to generate a corresponding inquiry statement according to the target intervention intention, and push the inquiry statement to the target terminal through the target vehicle for broadcasting; A parking route adjustment module, configured to, in response to a confirmation instruction of the driver, re-perform parking route planning according to the target intervention intention to obtain a target parking route, and perform parking control on the target vehicle according to the target parking route.
9. A gesture-based remote parking control device, characterized in that, Including: At least one processor; At least one memory, configured to store at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a gesture-based remote parking control method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to execute a gesture-based remote parking control method according to any one of claims 1 to 7.