Parking method and apparatus, and intelligent driving device
By acquiring voice and environmental information to control the speed and posture of intelligent driving devices, and using large language models (LLM) to generate control information, the problem of insufficient sensor coverage and inflexible driver control in existing automatic parking systems is solved, achieving greater parking convenience and safety.
Patent Information
- Application Number
- PCT/CN2025/106643
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-12
- Filing Date
- 2025-07-02
- Publication Date
- 2026-02-19
AI Technical Summary
Existing automatic parking systems do not allow drivers to flexibly adjust the parking position, trajectory, and speed when controlling the vehicle to park, resulting in a poor user experience. Furthermore, missing or inaccurate sensor signals may lead to a risk of vehicle collision.
By acquiring voice and environmental perception information, the speed and posture of intelligent driving devices are controlled, including image and radar information, enabling flexible control during parking. Large Language Model (LLM) is used to generate control information to improve the convenience and safety of the parking process.
It improves the flexibility and safety of the parking process, enhances the user's driving experience, and reduces the risk of collisions caused by insufficient obstacle recognition or sensor malfunction.
Smart Images

Figure CN2025106643_19022026_PF_FP_ABST
Abstract
Description
Parking method, device and intelligent driving equipment
[0001] The present application claims priority to the Chinese patent application No. 202411104000.6, filed on August 12, 2024, and entitled "Parking method, device and intelligent driving equipment", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of intelligent driving, and more particularly, to a parking method, device and intelligent driving equipment. BACKGROUND
[0003] Auto parking (AP) refers to automatic parking of a vehicle into a parking space, i.e., an automatic driving system can semi-automatically or fully automatically help a user to park the vehicle into a parking space. Auto parking can include auto parking assist (APA), remote parking assist (RPA) and auto valet parking (AVP), etc.
[0004] Current auto parking systems mainly rely on sensor signals to control the parking trajectory and pose when controlling the vehicle to park. During the auto parking process, the driver cannot flexibly control the parking pose, trajectory and speed, etc., which can result in poor user experience. In addition, in some parking scenarios, some sensor signals of the vehicle can be missing, or due to the limitation of sensor perception accuracy, the vehicle cannot accurately identify obstacles according to the sensor signals. The missing sensor signals and the inaccurate identification of obstacles can both result in the risk of collision of the vehicle.
[0005] In view of this, a parking solution with higher flexibility and safety is urgently needed to be developed. SUMMARY
[0006] The present application provides a parking method, device and intelligent driving equipment, which can adjust the pose and / or speed of the intelligent driving equipment during the parking process according to the voice instruction of the user, which helps to improve the flexibility and safety during the parking process.
[0007] In a first aspect, a parking method is provided, which can be executed by an intelligent driving equipment, for example, can be executed by a computing platform of the intelligent driving equipment, or can also be executed by a chip or circuit for the intelligent driving equipment.
[0008] The method comprises: obtaining voice information, the voice information being used to control a parking process of the intelligent driving device; obtaining environment perception information, the environment perception information being used to indicate an environment around the intelligent driving device; and controlling a speed and / or a pose of the intelligent driving device in a process of driving to a target parking space and / or parking into the target parking space according to the voice information and the environment perception information.
[0009] In some implementations, the voice information can be voice detected in a cabin of the intelligent driving device. The environment perception information can comprise images, or the environment perception information can also comprise information collected by a radar, such as a laser point cloud.
[0010] In the above technical solution, the speed and / or the pose of the intelligent driving device are controlled according to the voice instruction of the user in the process of driving to the target parking space or parking into the target parking space, so that the parking trajectory and the speed of the vehicle can be flexibly controlled, and the convenience and the safety of the parking process can be improved, thereby improving the driving experience of the user.
[0011] In combination with the first aspect, in some implementations of the first aspect, the voice information indicates that the parking mode is a first mode or a second mode, the first mode is used for not avoiding obstacles in the process of driving to the target parking space and / or parking into the target parking space, and the second mode is used for avoiding obstacles in the process of driving to the target parking space and / or parking into the target parking space; and the controlling of the speed and / or the pose of the intelligent driving device in the process of driving to the target parking space and / or parking into the target parking space according to the voice information and the environment perception information comprises: in the first mode, controlling the intelligent driving device to drive to the target parking space and / or park into the target parking space at a first speed; or in the second mode, controlling the intelligent driving device to drive to the target parking space and / or park into the target parking space at a second speed; and the first speed is greater than the second speed.
[0012] In some implementations, in the process of controlling the intelligent driving device to park in the first mode, when an obstacle appears around a parking trajectory or a target parking space, the second mode is switched to for parking.
[0013] In some implementations, in the second mode, when a voice instruction for controlling the intelligent driving device to accelerate is detected, the intelligent driving device is controlled to drive to the target parking space or park into the target parking space at a faster speed in the second mode. For example, when the voice information of "quick parking in the second mode" is detected, the intelligent driving device is controlled to drive to the target parking space or park into the target parking space at a faster speed in the second mode.
[0014] In some scenarios, in the first mode or the second mode, at least one of the following can also be controlled according to the voice information: a parking pose of the intelligent driving device in the target parking space, or a pose of the intelligent driving device in the process of parking into the target pose.
[0015] In the technical solution, if there is no obstacle in the parking track of the intelligent driving device and around the target parking space during parking, the speed and / or pose of the intelligent driving device in the process of driving to and / or parking into the target parking space is controlled according to the voice instruction of the user in the first mode, which helps to improve the parking efficiency.
[0016] With reference to the first aspect, in some implementations of the first aspect, the speed and / or pose of the intelligent driving device in the process of driving to and / or parking into the target parking space is controlled according to the voice information and the environment perception information, including: when the voice information indicates that there is a first obstacle in the first direction of the intelligent driving device or the target pose of the intelligent driving device in the target parking space, first control information is generated according to the position of the first obstacle or the position of the target parking space; and the speed and / or pose of the intelligent driving device is controlled according to the first control information.
[0017] In some implementations, when the voice information indicating the position of the first obstacle is obtained in the process of controlling the intelligent driving device to park in the first mode, the second mode is switched to for parking, that is, in the second mode, the first control information is generated according to the position of the first obstacle.
[0018] With reference to the first aspect, in some implementations of the first aspect, when the voice information indicates the first obstacle, the first control information is used to avoid the first obstacle and / or control the intelligent driving device to slow down; or when the voice information indicates the target pose, the first control information is used to control the intelligent driving device to park in the target parking space at the target pose.
[0019] In actual implementation, due to the processing capability of the intelligent driving device, the intelligent driving device may not be able to identify some obstacles in the parking track, in which case, by using the technical solution, the ability of the intelligent driving device to identify obstacles can be improved, so as to control the intelligent driving device to slow down and / or avoid the obstacles. In addition, by using the technical solution, the user can flexibly adjust the parking pose, which helps to improve the user driving experience.
[0020] With reference to the first aspect, in some implementations of the first aspect, the first obstacle is a negative obstacle or a scattered obstacle.
[0021] It should be noted that the negative obstacle can be an obstacle with a highest point lower than the plane on which the intelligent driving device is located, and the average distance or maximum distance between the bottom of the negative obstacle (i.e. the position near the highest point) and the plane on which the vehicle is located is referred to as the negative height. For example, the negative obstacle is a pit, and the bottom of the pit is the bottom of the negative obstacle, and the highest point of the bottom of the pit is regarded as the highest point of the negative obstacle. Taking the intelligent driving device as a vehicle as an example, the negative obstacle involved in the present application has a size (i.e. area) in the direction perpendicular to the height (i.e. parallel to the direction of the road surface on which the vehicle travels) greater than a certain threshold, for example, the size of the negative obstacle in the direction perpendicular to the height is sufficient to accommodate one or more wheels of the vehicle; in addition, the size of the negative obstacle in the direction parallel to the height (i.e. the negative height) can be a part or all of the diameter of the wheel, or greater than the diameter of the wheel. In addition, the scattered obstacle involved in the present application can be a positive obstacle with small volume or height, such as water accumulation or ice accumulation, etc.
[0022] In the above technical solution, in combination with the voice instruction of the user, the recognition ability of the intelligent driving device for the negative obstacle and / or the scattered obstacle can be improved, so that the intelligent driving device is controlled to slow down and / or avoid such obstacles during parking, thereby reducing the probability of the intelligent driving device slipping due to the existence of the scattered obstacle, and reducing the probability of the intelligent driving device scratching the negative obstacle, which helps to improve the safety during parking.
[0023] In combination with the first aspect, in some implementations of the first aspect, according to the voice information and the environment perception information, the speed and / or pose of the intelligent driving device during driving and / or parking into the target parking space is controlled, including: when the voice information indicates that there is a second obstacle in the second direction of the intelligent driving device, and the environment perception information does not include information indicating the second obstacle, generating second control information, the second control information being used to control the intelligent driving device to stop driving; controlling the speed of the intelligent driving device to be zero according to the second control information.
[0024] In some scenarios, when part of the sensors acquiring the environment perception information fails or the pose of part of the sensors acquiring the environment perception information changes, the intelligent driving device can fail to perceive the second obstacle during the parking process, and in this case, the intelligent driving device can have a risk of colliding with the second obstacle. For example, the environment perception information includes images acquired by a camera arranged at a rearview mirror, which can need to be folded when the intelligent driving device parks into a vertical narrow parking space, and in this case, the camera arranged at the rearview mirror will fail to acquire image information of the left and right sides of the intelligent driving device. Further, during the process of parking the intelligent driving device into a narrow parking space in a tail-in manner (i.e., the tail of the intelligent driving device parks into the parking space first), if the body of the intelligent driving device will rub against an obstacle, the intelligent driving device needs to drive for a distance in the head direction. If a pedestrian suddenly appears on the left side of the intelligent driving device when the intelligent driving device is preparing to drive in the head direction, and is preparing to walk to the right side via the head of the intelligent driving device, the intelligent driving device will fail to avoid the pedestrian due to the lack of perception information of the pedestrian. For the foregoing cases, through the technical solution described above, when the sensor of the intelligent driving device fails to perceive the second obstacle (such as a pedestrian), the intelligent driving device can be directly controlled to stop driving through the voice instruction of the user, so as to avoid collision with the second obstacle, which helps to improve the safety during the parking process.
[0025] With reference to the first aspect, in some implementations of the first aspect, the speed and / or the pose of the intelligent driving device during driving and / or parking into the target parking space is controlled according to the voice information and the environment perception information, including: determining a first environment feature according to the environment perception information, the first environment feature indicating a feature of a bird’s-eye-view (BEV) perspective corresponding to the intelligent driving device; determining a first text feature according to the voice information, the first text feature indicating at least one of the following: a position of the third obstacle relative to the intelligent driving device, a target pose of the intelligent driving device in the target parking space, or a target speed of the intelligent driving device; and controlling the pose and / or the speed of the intelligent driving device according to the first text feature and the first environment feature.
[0026] With reference to the first aspect, in some implementations of the first aspect, the pose and / or the speed of the intelligent driving device is controlled according to the first text feature and the first environment feature, including: extracting a second environment feature from the first environment feature according to the first text feature; wherein when the first text feature is related to the third obstacle, the second environment feature includes a feature associated with the third obstacle; when the first text feature is related to the target pose, the second environment feature includes a feature associated with the target parking space; and the pose and / or the speed of the intelligent driving device is controlled according to the first text feature and the second environment feature.
[0027] In the technical solution, the feature associated with the voice information is extracted from the environment perception information according to the voice information, and the intelligent driving device is controlled according to the feature, so that the control accuracy of the pose and / or speed is improved.
[0028] With reference to the first aspect, in some implementations of the first aspect, the pose and / or speed of the intelligent driving device is controlled according to the first text feature and the first environment feature, including: inputting the first text feature and the first environment feature into a first neural network model to obtain third control information; and controlling the pose and / or speed of the intelligent driving device according to the third control information, wherein the first neural network model is a large language model (LLM).
[0029] In the technical solution, the control information is generated by the large language model, so that the user can more conveniently and timely adjust the parking state and style through the voice instruction to meet more parking requirements of the user.
[0030] With reference to the first aspect, in some implementations of the first aspect, the method further includes: training the LLM based on control information loss and trajectory loss, the trajectory loss indicating a difference between a planned trajectory point and an actual trajectory point, wherein the planned trajectory point is generated according to the third control information, the control information loss is determined according to the third control information, the third control information indicates a planned acceleration and an actual steering angular velocity of the intelligent driving device in a future time period (such as 3 seconds or 5 seconds), the control information loss indicates a difference between an actual acceleration and the planned acceleration of the intelligent driving device, and a difference between an actual steering angular velocity and a planned steering angular velocity.
[0031] In some implementations, the LLM is trained based on the third control information and the trajectory loss, including: supervising the training process of the LLM according to the control information loss and the trajectory loss. For example, in the training process of the LLM, the trajectory loss is controlled to be as small as possible, and the control information loss is controlled to be as small as possible.
[0032] In the technical solution, the LLM is supervised according to the control information loss and the trajectory loss, which helps to improve the control accuracy of the intelligent driving device, thereby improving the safety of the intelligent driving device in the parking process.
[0033] With reference to the first aspect, in some implementations of the first aspect, the first environment feature includes a first feature indicating a position of a target parking space relative to the intelligent driving device, and a second feature indicating a position of an obstacle around the intelligent driving device; and the first environment feature is determined according to the environment perception information, including: inputting the environment perception information into a second neural network model to obtain the first environment feature, wherein the second neural network model is trained according to a parking space feature and an obstacle feature.
[0034] In some implementations, the parking space feature indicates a position of the target parking space, the obstacle feature indicates a position of an obstacle around the intelligent driving device relative to the intelligent driving device, the parking space feature can be obtained by inputting the environment perception information into the parking space segmentation neural network model, and the obstacle feature can be obtained by inputting the environment perception information into the obstacle segmentation neural network model.
[0035] In the technical solution described above, the second neural network model is trained by using the parking space feature and the obstacle feature, which helps to improve the perception ability of the second neural network model for the obstacle and the target parking space, thereby improving the accuracy of the pose control in the parking process and reducing the collision risk of the intelligent driving device in the parking process.
[0036] In a second aspect, a parking device is provided, which includes an acquisition unit and a processing unit. The acquisition unit is configured to acquire voice information, the voice information being used to control a parking process of an intelligent driving device; and the acquisition unit is further configured to acquire environment perception information, the environment perception information indicating an environment around the intelligent driving device. The processing unit is configured to control a speed and / or a pose of the intelligent driving device in a process of driving to and / or parking into a target parking space according to the voice information and the environment perception information.
[0037] In combination with the second aspect, in some implementations of the second aspect, the voice information indicates that a parking mode is a first mode or a second mode, the first mode being used for not avoiding an obstacle in the process of driving to and / or parking into the target parking space, and the second mode being used for avoiding the obstacle in the process of driving to and / or parking into the target parking space; and the processing unit is configured to control the intelligent driving device to drive to and / or park into the target parking space at a first speed in the first mode, or to control the intelligent driving device to drive to and / or park into the target parking space at a second speed in the second mode; wherein the first speed is greater than the second speed.
[0038] In combination with the second aspect, in some implementations of the second aspect, the processing unit is configured to generate first control information according to a position of a first obstacle or a position of the target parking space when the voice information indicates that the first obstacle exists in a first direction of the intelligent driving device or the target pose in the target parking space, and to control the speed and / or the pose of the intelligent driving device according to the first control information.
[0039] In combination with the second aspect, in some implementations of the second aspect, when the voice information indicates the first obstacle, the first control information is used to avoid the first obstacle and / or to control the intelligent driving device to slow down; or when the voice information indicates the target pose, the first control information is used to control the intelligent driving device to park in the target parking space at the target pose.
[0040] In combination with the second aspect, in some implementations of the second aspect, the first obstacle is a negative obstacle or a scattered obstacle.
[0041] With reference to the second aspect, in some implementations of the second aspect, the processing unit is configured to: generate the second control information for controlling the intelligent driving device to stop driving when the voice information indicates that there is a second obstacle in the second direction of the intelligent driving device, and the environment perception information does not include information indicating the second obstacle; and control the speed of the intelligent driving device to be zero according to the second control information.
[0042] With reference to the second aspect, in some implementations of the second aspect, the processing unit is configured to: determine a first environment feature according to the environment perception information, the first environment feature indicating a feature of a BEV perspective corresponding to the intelligent driving device; determine a first text feature according to the voice information, the first text feature indicating at least one of: a position of the third obstacle relative to the intelligent driving device, a target pose of the intelligent driving device in the target parking space, or a target speed of the intelligent driving device; and control the pose and / or speed of the intelligent driving device according to the first text feature and the first environment feature.
[0043] With reference to the second aspect, in some implementations of the second aspect, the control of the pose and / or speed of the intelligent driving device according to the first text feature and the first environment feature comprises: extracting a second environment feature from the first environment feature according to the first text feature; wherein when the first text feature is related to the third obstacle, the second environment feature includes a feature associated with the third obstacle; and when the first text feature is related to the target pose, the second environment feature includes a feature associated with the target parking space; and the control of the pose and / or speed of the intelligent driving device according to the first text feature and the second environment feature.
[0044] With reference to the second aspect, in some implementations of the second aspect, the processing unit is configured to: input the first text feature and the first environment feature into a first neural network model to obtain third control information; wherein the first neural network model is a large language model (LLM); and control the pose and / or speed of the intelligent driving device according to the third control information. With reference to the second aspect, in some implementations of the second aspect, the processing unit is further configured to: train the LLM based on the third control information and a trajectory loss, the trajectory loss indicating a difference between a planned trajectory point and an actual trajectory point, wherein the planned trajectory point is generated according to the third control information.
[0045] With reference to the second aspect, in some implementations of the second aspect, the first environment feature includes a first feature indicating a position of the target parking space relative to the intelligent driving device, and a second feature indicating positions of obstacles around the intelligent driving device; and the determination of the first environment feature according to the environment perception information comprises: inputting the environment perception information into a second neural network model to obtain the first environment feature; wherein the second neural network model is trained according to parking space features and obstacle features.
[0046] In a third aspect, a parking method is provided. The method comprises: obtaining voice information, the voice information being used to control a parking process of an intelligent driving device; obtaining environment perception information, the environment perception information indicating an environment around the intelligent driving device; generating first control information according to the voice information and the environment perception information, the first control information indicating a target speed and / or a target pose of the intelligent driving device for driving to and / or parking into a target parking space; and controlling the intelligent driving device according to the first control information.
[0047] With reference to the third aspect, in some implementations of the third aspect, the generating the first control information according to the voice information and the environment perception information comprises: determining a first environment feature according to the environment perception information, the first environment feature indicating a feature of a BEV perspective corresponding to the intelligent driving device; determining a first text feature according to the voice information, the first text feature indicating at least one of: a position of the third obstacle relative to the intelligent driving device, the target pose of the intelligent driving device in the target parking space, or the target speed of the intelligent driving device; and generating the first control information according to the first text feature and the first environment feature.
[0048] With reference to the third aspect, in some implementations of the third aspect, the generating the first control information according to the first text feature and the first environment feature comprises: extracting a second environment feature from the first environment feature according to the first text feature; wherein when the first text feature is related to the third obstacle, the second environment feature comprises a feature associated with the third obstacle; when the first text feature is related to the target pose, the second environment feature comprises a feature associated with the target parking space; and generating the first control information according to the first text feature and the second environment feature.
[0049] With reference to the third aspect, in some implementations of the third aspect, the generating the first control information according to the first text feature and the first environment feature comprises: inputting the first text feature and the first environment feature into a first neural network model to generate the first control information; and wherein the first neural network model is an LLM.
[0050] With reference to the third aspect, in some implementations of the third aspect, the generating the first control information according to the voice information and the environment perception information comprises: when the voice information indicates that there is a first obstacle in a first direction of the intelligent driving device, or the target pose of the intelligent driving device in the target parking space, generating the first control information according to a position of the first obstacle or a position of the target parking space indicated by the environment perception information.
[0051] With reference to the third aspect, in some implementations of the third aspect, when the voice information indicates the first obstacle, the first control information is used to avoid the first obstacle and / or control the intelligent driving device to slow down; or, when the voice information indicates the target pose, the first control information is used to control the intelligent driving device to park in the target parking space in the target pose.
[0052] With reference to the third aspect, in some implementations of the third aspect, the first obstacle is a negative obstacle or a scattered obstacle.
[0053] With reference to the third aspect, in some implementations of the third aspect, the voice information indicates that the parking mode is a first mode or a second mode, the first mode is used for not avoiding obstacles during driving to and / or parking into the target parking space, and the second mode is used for avoiding obstacles during driving to and / or parking into the target parking space; in the first mode, the first control information is used to control the intelligent driving device to drive to and / or park into the target parking space at a first speed; or, in the second mode, the first control information is used to control the intelligent driving device to drive to and / or park into the target parking space at a second speed; wherein the first speed is greater than the second speed.
[0054] With reference to the third aspect, in some implementations of the third aspect, the method further comprises: when the voice information indicates that a second obstacle exists in a second direction of the intelligent driving device, and the environment perception information does not include information indicating the second obstacle, generating second control information, the second control information being used to control the intelligent driving device to stop driving; and controlling the speed of the intelligent driving device to be zero according to the second control information.
[0055] With reference to the third aspect, in some implementations of the third aspect, the method further comprises: training the LLM based on control information loss and trajectory loss, the trajectory loss indicating a difference between a planned trajectory point and an actual trajectory point, wherein the planned trajectory point is generated according to the third control information, the control information loss is determined according to the first control information, and the control information loss indicates a difference between an actual acceleration and a planned acceleration of the intelligent driving device, and a difference between an actual steering angular velocity and a planned steering angular velocity.
[0056] With reference to the third aspect, in some implementations of the third aspect, the first environment feature includes a first feature indicating a position of the target parking space relative to the intelligent driving device, and a second feature indicating a position of an obstacle around the intelligent driving device; and the first environment feature is determined according to the environment perception information, comprising: inputting the environment perception information into a second neural network model to obtain the first environment feature; wherein the second neural network model is trained according to the parking space feature and the obstacle feature.
[0057] In a fourth aspect, a parking device is provided, which comprises a processor configured to execute a computer program stored in a memory to cause the device to perform the method in any possible implementation of the first aspect or the third aspect.
[0058] In some implementations of the fourth aspect, the device further comprises a memory.
[0059] In a fifth aspect, an intelligent driving device is provided, which comprises the device in any possible implementation of the second aspect or the fourth aspect.
[0060] In some implementations of the fifth aspect, the intelligent driving device is a vehicle.
[0061] In a sixth aspect, a computer program product is provided, which comprises computer program codes configured to cause a computer or a processor to perform the method in any possible implementation of the first aspect or the third aspect when the computer program codes are run on the computer or the processor.
[0062] It should be noted that the computer program codes can be stored in the storage medium in whole or in part, and the storage medium can be packaged together with the processor or packaged separately from the processor.
[0063] In a seventh aspect, a computer readable medium is provided, which stores instructions configured to cause a processor to implement the method in any possible implementation of the first aspect or the third aspect when the instructions are executed by the processor.
[0064] In an eighth aspect, a chip is provided, which comprises a circuit configured to perform the method in any possible implementation of the first aspect or the third aspect.
[0065] The beneficial effects not described in the second aspect to the eighth aspect can be referred to the description in the first aspect, and will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0066] FIG. 1 is a functional schematic block diagram of a vehicle according to an embodiment of the present application;
[0067] FIG. 2 is a schematic block diagram of a parking system according to an embodiment of the present application;
[0068] FIG. 3 is a schematic block diagram of a rule control signal generation module according to an embodiment of the present application;
[0069] FIG. 4 is a schematic flow chart of a parking method according to an embodiment of the present application;
[0070] FIG. 5 is a schematic diagram of an application scenario of the parking method according to an embodiment of the present application;
[0071] FIG. 6 is another schematic diagram of an application scenario of the parking method according to an embodiment of the present application;
[0072] FIG. 7 is still another schematic diagram of an application scenario of the parking method according to an embodiment of the present application;
[0073] FIG. 8 is another schematic flowchart of the parking method according to an embodiment of the present application;
[0074] FIG. 9 is a schematic block diagram of a parking device according to an embodiment of the present application;
[0075] FIG. 10 is another schematic block diagram of a parking device according to an embodiment of the present application. DETAILED DESCRIPTION
[0076] Since the embodiments of the present application relate to the application of neural networks, the following introduces the related terms and concepts of neural networks that may be involved in the embodiments of the present application for the convenience of understanding.
[0077] (1) Neural network
[0078] The neural network can be composed of neural units, and a neural unit can refer to an operation unit with x s and an intercept 1 as inputs. The output of the operation unit can be:
[0079] where s = 1, 2, … n, n is a natural number greater than 1, W s is the weight of x s , and b is the bias of the neural unit. f is an activation function of the neural unit, which is used to introduce a nonlinear characteristic into the neural network to convert the input signal in the neural unit into an output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a Relu function, a sigmoid function, etc. The neural network is a network formed by connecting multiple single neural units together, i.e., the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neural units.
[0080] (2) Deep neural network
[0081] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with multiple hidden layers. The neural networks inside a DNN can be divided into three categories according to the positions of different layers: input layer, hidden layer, and output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. The layers are fully connected, that is, any neuron in the ith layer is connected to any neuron in the (i+1)th layer.
[0082] Each layer in a DNN can be understood as a linear relationship expression as follows: wherein, is an input vector, is an output vector, is a bias vector, W is a weight matrix (also known as a coefficient), and a() is an activation function. Each layer only performs a simple operation on the input vector to obtain the output vector Due to the large number of layers in a DNN, the number of coefficients W and bias vectors is also relatively large. These parameters in a DNN are defined as follows: taking the coefficient W as an example, assuming that in a three-layer DNN, the linear coefficient from the 4th neuron of the second layer to the 2nd neuron of the third layer is defined as The superscript 3 represents the layer number of the coefficient W, and the subscripts correspond to the output third layer index 2 and the input second layer index 4.
[0083] In summary, the coefficient from the kth neuron of the (L-1)th layer to the jth neuron of the Lth layer is defined as
[0084] It should be noted that the input layer has no coefficient W. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. In theory, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is a process of learning the weight matrix, and the ultimate goal is to obtain the weight matrix of all layers of the trained deep neural network (the weight matrix formed by the coefficients W of many layers).
[0085] (3) Convolutional neural network
[0086] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. The CNN contains a feature extractor composed of convolutional layers and subsampling layers, which can be regarded as filters. A convolutional layer refers to a neuron layer in the CNN that performs convolutional processing on an input signal. In the convolutional layer of the CNN, a neuron can be connected to only part of the neurons in the adjacent layer. In a convolutional layer, there are usually several feature planes, each of which can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weights are convolutional kernels. The shared weights can be understood as being irrelevant to the way and position of extracting image information. The convolutional kernels can be initialized in the form of a matrix of random size, and the convolutional kernels can obtain reasonable weights through learning in the training process of the CNN. In addition, the shared weights directly reduce the connections between the layers of the CNN and reduce the risk of overfitting.
[0087] (4) Recurrent neural networks
[0088] A recurrent neural network (RNN) is a neural network used to process sequence data. In a traditional neural network model, the data is processed from the input layer to the hidden layer and then to the output layer, and the layers are fully connected. However, the nodes in each layer are not connected. Although this neural network has solved many problems, it is still powerless for many problems. For example, to predict the next word of a sentence, the previous words need to be used because the words in a sentence are not independent. The RNN is called a recurrent neural network because the current output of a sequence is related to the previous output. Specifically, the network memorizes the previous information and applies it to the calculation of the current output, that is, the nodes in the hidden layer are connected, and the input of the hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous time. In theory, the RNN can process sequence data of any length. The training of the RNN is the same as the training of a traditional CNN or DNN. The RNN aims to enable the machine to have the ability to remember like a human being. Therefore, the output of the RNN needs to rely on the current input information and the historical memory information.
[0089] The technical solutions in the embodiments of the present application will be described below with reference to the drawings.
[0090] FIG. 1 is a functional block diagram of a vehicle according to an embodiment of the present application. As shown in FIG. 1, the vehicle 100 can include a perception system 120, a human-machine interaction system 130, and a computing platform 150, wherein the perception system 120 can include several sensors for sensing information of an environment around the vehicle 100. For example, the perception system 120 can include a positioning system, which can be a global navigation satellite system (GNSS) such as a global positioning system (GPS), a Beidou system, etc. For another example, the perception system 120 can also include one or more of an inertial measurement unit (IMU), a laser radar, a millimeter wave radar, an ultrasonic radar, and a camera.
[0091] Exemplarily, the camera outside the cabin can include one or more of a front-view camera, a rear-view camera, an all-around-view camera, and a side-view camera. The front-view camera can be installed at a front windshield. The rear-view camera can be installed at a rear trunk. The side-view camera can be installed below a side mirror. The all-around-view camera includes four cameras installed around the vehicle, and the images obtained by the four cameras can be spliced to obtain a panoramic image of the surroundings of the vehicle. In some implementations, the all-around-view camera can coincide with the front-view camera, the rear-view camera, and the side-view camera. For example, the camera of the all-around-view camera installed on the side of the vehicle can be the side-view camera, the camera of the all-around-view camera installed in front of the vehicle can be the front-view camera, and the camera of the all-around-view camera installed behind the vehicle can be the rear-view camera. Alternatively, the all-around-view camera can also be different from the front-view camera, the rear-view camera, and the side-view camera.
[0092] The human-machine interaction system 130 can include a sound receiving device such as a microphone, etc. for receiving voice information. Alternatively, the human-machine interaction system 130 can also include a sound emitting device such as a loudspeaker, a sound box, etc. for playing voice information.
[0093] Some or all of the functionality of the vehicle 100 can be controlled by the computing platform 150. The computing platform 150 can include processors 151-15n, which are circuits that have the capability to process signals. In one implementation, the processors can be circuits that have the capability to read and execute instructions, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a type of microprocessor), a digital signal processor (DSP), or the like. In another implementation, the processors can be circuits that implement functionality through fixed or reconfigurable logic, such as an application-specific integrated circuit (ASIC) or a programmable logic device (PLD) such as a field programmable gate array (FPGA). In reconfigurable hardware circuits, the processors load configuration documents to implement the configuration of the hardware circuits, which can be understood as the processors loading instructions to implement the functionality of some or all of the units described above. Additionally, the processors can be hardware circuits designed for artificial intelligence, which can be understood as a type of ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), or the like. Additionally, the computing platform 150 can include a memory that stores instructions that can be called by some or all of the processors 151-15n to implement corresponding functionality.
[0094] The operation of the intelligent driving system can be controlled by the computing platform 150, which can include an advanced driving assistant system (ADAS) and an autonomous driving system (ADS). The intelligent driving system utilizes various sensors on the vehicle (including but not limited to: lidar, millimeter wave radar, camera, ultrasonic sensor, global positioning system, inertial measurement unit) to obtain information from the surroundings of the vehicle, and analyzes and processes the obtained information to achieve functions such as obstacle perception, target recognition, vehicle positioning, trajectory planning, driver monitoring / alerting, etc., thereby improving the safety, automation level, and comfort of vehicle driving.
[0095] At different automatic driving levels (or intelligent driving levels, a total of L0-L5 six levels), based on the information obtained by artificial intelligence algorithm and multi-sensor, the intelligent driving system can realize different levels of automatic driving assistance. The above automatic driving levels are based on the classification standard of Society of Automotive Engineers (SAE). Among them, L0 level is no automation; L1 level is driving assistance; L2 level is partial automation; L3 level is conditional automation; L4 level is high automation; L5 level is complete automation. The tasks of monitoring road conditions and responding at L1 to L3 levels are completed by the driver and the system together, and the driver needs to take over the dynamic driving task. L4 and L5 levels can make the driver completely change into the role of a passenger. At present, the functions that the intelligent driving system can realize mainly include but are not limited to: adaptive cruise assistance, automatic emergency braking, automatic parking, blind spot monitoring, front intersection traffic warning / braking, rear intersection traffic warning / braking, front vehicle collision warning, lane departure warning, lane keeping assistance, rear vehicle collision warning, traffic sign recognition, traffic congestion assistance, highway assistance, etc. It should be understood that the above various functions can have specific modes at different automatic driving levels (L0-L5), and the higher the automatic driving level, the more intelligent the corresponding mode. For example, automatic parking can include auto parking assist (APA), remote parking assist (RPA), and auto valet parking (AVP), etc. For APA, the driver does not need to manipulate the steering wheel, but still needs to manipulate the accelerator and brake on the vehicle; for RPA, the driver can use a terminal (such as a mobile phone) to remotely control the parking of the vehicle outside the vehicle; for AVP, the vehicle can complete parking without the driver. In terms of corresponding automatic driving levels, APA is approximately at the L1 level, RPA is approximately at the L2-L3 level, and AVP is approximately at the L4 level.
[0096] In this application, during parking, the computing platform 150 can control the speed and / or pose of the vehicle 100 during parking according to the voice instruction obtained by the human-computer interaction system 130 and the environment perception information obtained by the perception system 120.
[0097] FIG. 2 shows a parking system provided by an embodiment of the present application. As shown in FIG. 2, the system includes a perception module 210, a voice receiving module 220, a regulation and control signal generation module 230, and an actuator 240. The functions of each module can be as described below (I) to (IV).
[0098] The perception module 210 is configured to collect environmental perception information around the vehicle, which can include information of the target parking space and information of obstacles around the vehicle. The perception module 210 can include one or more cameras in the perception system 120 shown in FIG. 1, or can further include other sensors.
[0099] The voice receiving module 220 is configured to receive voice information of the user, which can be information for controlling the parking process of the vehicle. The voice receiving module 220 can include one or more devices in the human-computer interaction system 130 shown in FIG. 1.
[0100] The control signal generation module 230 is configured to generate vehicle control information according to the environmental perception information and the voice information, the vehicle control information being used to adjust the speed and / or pose of the vehicle during driving to the target parking space and / or parking into the target parking space. The control signal generation module 230 includes one or more processors in the computing platform 150 shown in FIG. 1. In some implementations, the control signal generation module 230 can further calculate a control quantity (such as a driving torque, a braking torque, a steering wheel angle, etc.) according to the vehicle control information, and input the control quantity to the actuator 240.
[0101] The actuator 240 is configured to execute the aforementioned control quantity, and when the aforementioned control quantity is executed, the vehicle can be controlled to drive at the speed and / or pose indicated by the vehicle control information. The actuator 240 can include a steering and braking control system in the vehicle 100.
[0102] In some implementations, the control signal generation module 230 can include more functional modules. For example, FIG. 3 shows a schematic diagram of the control signal generation module 230 and its inputs and outputs according to an embodiment of the present application. As shown in FIG. 3, the control signal generation module 230 includes an encoder and a planner. The encoder includes a camera backbone network, a depth decoder, a projector, a temporalizer, and a BEV backbone network. The planner includes an adapter 1, an adapter 2, a text processor, a fusion encoder, and an LLM.
[0103] More specifically, the encoder can obtain the BEV temporal feature through the following steps 1) to 4):
[0104] 1) input the environmental perception information into a camera backbone network for feature extraction. Illustratively, the environmental perception information can include images captured by surround-view cameras. For example, the surround-view cameras include four cameras arranged around the vehicle, the environmental perception information can include four images captured by the four cameras respectively. After inputting the four images into the camera backbone network, four two-dimensional image features can be output, i.e., visual image features of four perspectives. In addition, the four images can be input into a depth decoder to obtain depth information of the pixels in each image. Illustratively, the camera backbone network can be a CNN.
[0105] 2) input the four image features and the depth information of the pixels corresponding to each image into a projector, which can map the features in the image coordinate system to the BEV coordinate system. It can be understood that each sensor in the vehicle has its own coordinate system, and their output data or perception results will finally be aggregated to the vehicle coordinate system of the ego vehicle for processing. The mapping relationship between the vehicle coordinate system and the image coordinate system can be determined according to the extrinsic matrix and the intrinsic matrix of each camera. For the scene corresponding to the images obtained from the surround-view cameras, the purpose of the projector is to determine the rasterized representation of the scene in the BEV coordinate system. Illustratively, taking the projector as LSS (lift, splat, shoot) for example, each of the four image features can be a two-dimensional image feature, then the LSS first performs an outer product operation on the two-dimensional image features and the depth information of the pixels corresponding to the two-dimensional image features to obtain three-dimensional image features that can be processed by a CNN, then performs a sum-pooling operation (such as z accumulation and flattening) to realize dimension reduction, and then concatenates the features after dimension reduction to obtain BEV features. It should be noted that the above is described by taking the projector as LSS for example, and in actual implementation, the projector can also be other algorithms, such as Cam2BEV, PyrOccNet, etc.
[0106] 3) input the BEV features into a temporalizer to obtain temporal fusion features. More specifically, the temporalizer obtains the current BEV temporal feature (hereinafter referred to as BEV temporal feature) according to the historical BEV temporal feature (such as the BEV temporal feature of the previous moment) and the current BEV feature. For example, the historical BEV temporal feature is aligned to the current frame, and then a temporal fusion method is used for temporal fusion. It can be understood that the BEV features output by the projector are not associated with time, and the BEV temporal feature can indicate the change of the BEV feature over time. Illustratively, the temporalizer can be a long short term memory network (LSTM), or can also be other RNNs.
[0107] 4) input the BEV time sequence feature and the target parking space information into the BEV backbone network for feature extraction, so as to obtain the BEV time sequence feature fused with the position of the target parking space. It can be understood that the BEV time sequence feature fused with the position of the target parking space can indicate the change of the relative positional relationship between the target parking space and the vehicle position over time. Exemplarily, the target parking space information can be a target point corresponding to the target parking space, and the position of the target parking space in the BEV coordinate system can be determined according to the target point, and then the BEV time sequence feature containing the position of the target parking space is obtained. In some implementations, the target point corresponding to the target parking space can be a coordinate value, and the target point can be any point in the optional region in the target parking space, wherein the optional region is a region formed by the boundaries of the target parking space and a certain distance from the boundaries, and the certain distance can be 0.15 times the boundary length of the target parking space, or other multiples of the boundary length. In addition, before the target point is input into the BEV backbone network, the target point is processed using a Gaussian kernel function, and the target point processed by the Gaussian kernel function is input into the BEV backbone network.
[0108] Further, the planner generates the vehicle control information according to the BEV time sequence feature output by the encoder and the voice instruction, which can specifically include the following steps a) to e):
[0109] a) After splicing the vehicle motion state and the BEV time sequence feature, input into the adapter 1 to adjust the number of feature channels, and the number of feature channels of the BEV fusion feature output by the adapter 1 matches the number of feature channels that the fusion encoder can process. Wherein, the vehicle motion state can include the position of the ego vehicle, the speed of the ego vehicle and the like.
[0110] b) use the text processor to analyze the text feature obtained from the voice information, which can be a feature represented by numbers. For example, the text processor decomposes the text in the voice information to obtain a plurality of text units (or tokens), which are common character sequences, words and the like in the text, so the text units can be words, characters, words, numbers and the like. Further, the text processor uses a text unit index (such as token ID) to represent each text unit in the plurality of text units, which can be in the form of numbers so that the neural network can process the text unit index. Exemplarily, the text processor can be a natural language processing tool such as a text tokenizer.
[0111] c) input the text feature, the learnable query vector and the BEV fusion feature output by the adapter 1 into a fusion encoder, and the fusion encoder extracts the feature most relevant to the text feature from the BEV fusion feature according to the learnable query vector. Exemplarily, the fusion encoder can be a Transformer type neural network such as Q-Former.
[0112] d) input the feature output by the fusion encoder into the adapter 2 to adjust the number of feature channels, and the number of feature channels corresponding to the output of the adapter 2 matches the number of feature channels that can be processed by the LLM.
[0113] e) input the text feature obtained in step b) and the feature output by the adapter 2 into the LLM, and the LLM can generate vehicle control information including planned acceleration and planned steering angular velocity according to the text feature and the feature output by the adapter 2. It should be understood that the LLM can determine that the vehicle needs to control braking and / or steering according to the text feature, and since the feature output by the adapter 2 is a feature that fuses the BEV time sequence feature and the vehicle motion state, the LLM can determine the acceleration and / or steering angular velocity that the vehicle needs to perform according to the feature output by the adapter 2 and the text feature.
[0114] In some implementations, the vehicle motion state and the BEV time sequence feature can also be spliced and then directly input into the adapter 2 to adjust the number of feature channels. Then, the text feature and the feature output by the adapter 2 are input into the LLM, and the LLM can generate vehicle control information according to the text feature and the feature output by the adapter 2.
[0115] The above describes the system provided by the embodiments of the present application in combination with FIG. 2 and FIG. 3, and the foregoing system is only an exemplary description, and in actual implementation, the modules in the above system can be added or deleted according to actual needs. For example, the encoder can further include one or more backbone networks for processing signals of other sensors, and the output of the one or more backbone networks can be input into the projector to construct the BEV feature; or the output of the one or more backbone networks can also be input into the time sequencer to construct the BEV time sequence feature.
[0116] In order to facilitate understanding of the technical solutions of the present application, the parking method provided by the present application is described in detail below in combination with FIG. 4 to FIG. 7.
[0117] FIG. 4 shows a schematic flowchart of the parking method provided by the embodiments of the present application, and the method 400 can be performed by the vehicle shown in FIG. 1 or the system shown in FIG. 2, more specifically, can be performed by the rule control signal generation module 230. The method includes:
[0118] S401, obtaining environment perception information and determining a BEV time sequence feature according to the environment perception information.
[0119] In some implementations, the environment perception information can include images captured by the surround-view camera, or the environment perception information includes images captured by cameras in different orientations of the vehicle, and the method of determining the BEV time sequence feature according to the environment perception information can refer to the descriptions in the foregoing steps 1) to 4), which will not be repeated here.
[0120] In yet some implementations, in addition to the images captured by the camera, the environment perception information also includes laser point clouds captured by a laser radar, etc., and the determination of the BEV time sequence feature according to the environment perception information can be: based on a multi-modal fusion algorithm such as HDMapNET, BEVFusion, etc., to determine the BEV time sequence feature through the images and the laser point clouds.
[0121] S402, acquire voice information, and extract text features according to the voice information.
[0122] Exemplarily, the voice information can be acquired during parking, and the voice information can be acquired in the cabin of the vehicle. The specific implementation of extracting text features according to the voice information can refer to the description in the foregoing step b), which will not be repeated here.
[0123] In some implementations, the voice information is “there is a pit in front, please pay attention to avoid”, and the text features can include: a text unit index related to a direction (such as a text unit index related to “front”), a text unit index related to an obstacle (such as a text unit index related to “pit”), and a text unit index related to vehicle control (such as a text unit index related to “avoid”).
[0124] Exemplarily, the text unit related to the direction can also include left side, right side, rear (back side), left front, etc.; the text unit related to the obstacle can also include obstacle, ice, water pit, person, block, blockage, etc.; and the text unit related to vehicle control can also include acceleration, deceleration, parking, etc.
[0125] In yet some implementations, the voice information is “park on the right side”, and the text features can include: a text unit index related to a parking action (such as a text unit index related to “parking”), and a text unit index related to a parking orientation (such as a text unit index related to “on the right”).
[0126] S403, according to the text features, extract environment features associated with the text features from the BEV time sequence features.
[0127] In some implementations, the environmental features can be extracted according to all the text features in S402 in S403; or the environmental features can also be extracted according to the text unit indexes other than the text unit indexes related to vehicle control in S403.
[0128] In an example, the image features indicating the obstacles in the corresponding direction of the vehicle can be extracted according to the text unit indexes related to the direction and the text unit indexes related to the obstacles. In another example, the image features of the target parking space can also be extracted according to the text unit indexes related to the parking action and the parking orientation. For more specific implementations of extracting the environmental features associated with the text features from the BEV time sequence features, reference can be made to the description in the aforementioned step c), which will not be repeated here.
[0129] S404, generating the vehicle control information according to the text features and the environmental features.
[0130] Exemplarily, the environmental features and all the text features in S402 can be input into the LLM in FIG. 3 in S404 to make the LLM output the vehicle control information; or the environmental features and the text unit indexes related to vehicle control in the text features can be input into the LLM in FIG. 3 in S404 to make the LLM output the vehicle control information. Exemplarily, the vehicle control information can be the planned control signals for the next 30 frames, and each frame of control signal action t satisfies the following formula (1): action t t t ; (1)
[0131] wherein acc t is the acceleration of the t-th frame, and sar t is the steering angular velocity of the t-th frame. Exemplarily, each of the aforementioned frames can be 0.1 seconds, 0.15 seconds, or other time lengths.
[0132] In some implementations, after obtaining the vehicle control information, the vehicle control information and the vehicle motion state can be input into a dynamics model to obtain the planned trajectory for the next 30 frames, which can satisfy the following formulas (2) to (6):
[0133] wherein ego_state t represents the vehicle motion state of the ego vehicle at the t-th frame, x t represents the x value of the trajectory point of the ego vehicle at the t-th frame, y t represents the y value of the trajectory point of the ego vehicle at the t-th frame, and theta t a direction angle of a speed of the ego vehicle in the t-th frame (i.e., an angle between the speed of the ego vehicle and an x-axis of a whole coordinate system of the ego vehicle), v t a speed of the ego vehicle in the t-th frame, and wheelbase represents a wheelbase of the ego vehicle.
[0134] In some implementations, during the process in which the vehicle travels and / or parks into the target parking space under the control of the automatic parking function, when the voice information indicates that there is an obstacle in a certain direction of the vehicle, the obstacle-related features are extracted according to the voice information, and then the obstacle-related features control the vehicle to slow down and / or avoid the obstacle. The aforementioned “certain direction” can be the traveling direction of the vehicle.
[0135] In an example, as shown in (a) of FIG. 5, during the process in which the vehicle travels to the target parking space under the control of the automatic parking function, an obstacle 501 appears in front of the vehicle, at this time, the voice information “there is a pit in front, please pay attention to avoid” of the user in the cabin is detected (as shown in (b) of FIG. 5), and then the text features can be extracted according to the voice information. Further, according to the text features related to “front” and “pit” in the text features, the environment features related to the obstacle 501 (such as the position of the obstacle 501 relative to the ego vehicle) are extracted from the BEV time sequence features. Further, according to the text features related to “avoid” and the environment features related to the obstacle 501, the trajectory for avoiding the obstacle 501 is planned, and then the vehicle control information is generated in combination with the current motion state information of the vehicle and the planned trajectory. For example, taking the parking space 502 shown in FIG. 6 as the target parking space, the part of the original trajectory planned by the automatic parking function that passes through the obstacle 501 is shown as a dashed line in FIG. 6, that is, the vehicle will pass through the obstacle 501 during the process of traveling to the parking space 502. The adjusted trajectory according to the voice information can bypass the obstacle 501.
[0136] In another example, when there is an icy road in front of the vehicle, the voice information “the road in front is icy, please slow down” of the user in the cabin is detected, and then the text features can be extracted according to the voice information. Further, according to the text features related to “front” and “ice” in the text features, the environment features related to the icy road are extracted from the BEV time sequence features. Further, according to the text features related to “slow down” and the environment features related to the icy road, the vehicle is controlled to slow down on the icy road area.
[0137] It should be noted that the above two examples are described by taking the voice information including the text units related to the vehicle control (such as slow down and avoid) as an example, and in actual implementation, when the voice information does not include the text units related to the vehicle control, after the obstacle-related features are extracted, the vehicle can also be controlled to slow down and / or avoid the obstacle according to the text features and the obstacle-related features.
[0138] In yet some implementations, during the process that the vehicle is driving or parking into the target parking space under the control of the automatic parking function, the pose of the vehicle in the target parking space can also be adjusted in response to a voice instruction from the user to adjust the pose of the vehicle in the target parking space. For example, during the process that the vehicle is driving or parking into the target parking space, voice information “please park as close to the left as possible” of the user in the cabin is detected, and then text features can be extracted according to the voice information. Further, environment features related to the target parking space can be extracted from the BEV time sequence features according to text features related to “parking” and “left”. Further, a vehicle pose for parking on the left in the target parking space is planned according to the text features related to “left” and the environment features related to the target parking space, and then vehicle control information is generated in combination with the current motion state information of the vehicle and the vehicle pose for parking on the left. After the vehicle executes the vehicle control information, the vehicle can be parked on the left side of the target parking space.
[0139] In still some implementations, during the process that the vehicle is driving or parking into the target parking space under the control of the automatic parking function, the driving speed of the vehicle can also be adjusted in response to a voice instruction from the user to adjust only the driving speed of the vehicle.
[0140] In an example, during the process that the vehicle is driving or parking into the target parking space, voice information “please accelerate parking” or “please park faster” of the user in the cabin is detected, and then text features can be extracted. Further, environment features related to obstacles can be extracted from the BEV time sequence features according to text features related to “accelerate” or “faster”. Further, the vehicle is controlled to avoid obstacles and park into the target parking space at a faster speed according to the text features related to “accelerate” or “faster” and the environment features related to the obstacles.
[0141] In yet another example, during the process that the vehicle is driving or parking into the target parking space, voice information for controlling the speed of the vehicle such as “please accelerate parking” or “please park faster” is detected, there is no obstacle in the trajectory of the vehicle driving to the target parking space by default, and there is no obstacle around the target parking space that hinders the parking process. In this case, the vehicle can be directly controlled to park into the target parking space at a faster speed according to the voice information. It should be noted that in the foregoing scenario, the text features extracted from the voice information and the BEV time sequence features are directly input into the LLM without using the fusion encoder to extract features associated with the voice information, so that the LLM outputs vehicle control information for controlling the vehicle to park into the target parking space at a faster speed. In this case, the time length and computing power required for processing data by the fusion encoder can be saved.
[0142] In some implementations, the vehicle can include multiple parking modes, such as a regular parking mode and a fast parking mode, or a regular parking mode and an obstacle-free parking mode. Among them, in the regular parking mode, the vehicle can plan a parking trajectory according to the obstacle position, or also plan a parking trajectory according to the voice information and the obstacle position (such as the obstacle position indicated by the environmental perception information). In the fast parking mode or the obstacle-free parking mode, the vehicle can respond to the user's voice instruction to control the vehicle to park into the target parking space, and in the process of planning the parking trajectory, the obstacle-related information does not need to be considered, and the vehicle can only plan a trajectory for parking into the target parking space according to the position of the vehicle and the target parking space; or, the vehicle can also plan a trajectory for parking into the target parking space according to the position of the vehicle, the drivable area (such as the drivable area determined according to the lane line in the parking lot), and the target parking space.
[0143] It can be understood that in the fast parking mode or the obstacle-free parking mode, in the process of generating the vehicle control information through the regulation and control signal generation module 230 shown in FIG. 3, the fusion encoder does not need to be used to extract the features associated with the voice information, that is, the text features and the BEV time sequence features extracted from the voice information can be directly input into the LLM to obtain the vehicle control information. In addition, compared with the regular parking mode, the speed of parking in the fast parking mode or the obstacle-free parking mode is faster. For example, the time required for the vehicle to park into a target parking space from a certain position in the regular parking mode is longer than the time required for the vehicle to park into the same target parking space from the same position in the fast parking mode or the obstacle-free parking mode; or the average speed of the vehicle to park into a target parking space from a certain position in the regular parking mode is slower than the average speed of the vehicle to park into the same target parking space from the same position in the fast parking mode or the obstacle-free parking mode.
[0144] In some implementations, during the process of parking the vehicle into the target parking space under the control of the automatic parking function, if the parking space into which the vehicle is parked at this time is a narrow parking space, the vehicle will control the rearview mirror to be retracted (the rearview mirror 601 is in the retracted state as shown in (a) of FIG. 7) at this time to avoid scratching the vehicle with the adjacent obstacle. For example, if the vehicle needs to drive forward in the retracted state of the rearview mirror to adjust the pose during the process of parking into the target parking space, there is a pedestrian 602 moving from the left side to the right side in front of the left side of the vehicle at this time. In this case, the perception system of the vehicle can not be able to perceive the existence of the pedestrian. Then, after detecting the voice information "pedestrian on the left side, please avoid" of the user in the cabin (as shown in (b) of FIG. 7), the text features can be extracted according to the voice information. Further, the environmental features related to the pedestrian 602 can be extracted from the BEV time sequence features according to the text features related to "left side" and "pedestrian" in the text features. It can be understood that, since the rearview mirror 601 is in the retracted state, the camera device at the rearview mirror 601 can not be able to collect the image including the pedestrian 602, i.e., there is no environmental feature related to the pedestrian 602 in the BEV time sequence features, so in this case, the vehicle control information for controlling the vehicle to stop can be generated to control the vehicle to stop to avoid the collision between the vehicle and the pedestrian 602. Further, after the vehicle stops, the vehicle can continue to park into the target parking space in response to the instruction to continue parking.
[0145] The parking method provided by the embodiments of the present application can adjust the parking pose, trajectory, speed, etc. in response to the voice instruction of the user during the automatic parking process, which helps to improve the flexibility of parking. In addition, in the case that some sensor signals of the vehicle are missing or the obstacle cannot be accurately identified according to the sensor signals, the vehicle can be controlled to stop to avoid the obstacle in response to the voice instruction of the user, or the ability of the vehicle to identify the obstacle can be improved in response to the voice instruction of the user, so as to avoid the collision between the vehicle and the obstacle, which helps to improve the safety during the parking process.
[0146] The specific implementation and application scenarios of the parking method provided by the embodiments of the present application are described above in combination with FIGS. 4 to 7. In some implementations, in order to improve the control accuracy in the parking process, the encoder in FIG. 3 can be trained according to the parking space features and the obstacle features before the aforementioned parking method is executed. More specifically, the parking space features indicate the position of the target parking space, and the parking space features can be obtained by inputting the environment perception information into the parking space segmentation neural network model. The obstacle features indicate the position of the obstacles around the intelligent driving device relative to the intelligent driving device, and the obstacle features can be obtained by inputting the environment perception information into the obstacle segmentation neural network model. The encoder in FIG. 3 is trained using the parking space features and the obstacle features to adjust the weights associated with each of the depth decoder, the camera backbone network, the projector, the temporalizer, and the BEV backbone network, so that the accuracy of the encoder in identifying the obstacles and the target parking space is higher. For example, the learning rate used in the process of training the encoder can be 0.0004. In yet some implementations, in order to improve the control accuracy in the parking process, the training of the planner and the encoder can also be supervised according to the vehicle control information loss and the trajectory loss before the aforementioned parking method is executed. For example, in the process of training the LLM, the control trajectory loss is minimized, and the control vehicle control information loss is minimized. The vehicle control information loss can indicate the difference between the actual acceleration of the intelligent driving device and the planned acceleration, and the difference between the actual steering angular velocity and the planned steering angular velocity. The trajectory loss indicates the difference between the planned trajectory point and the actual trajectory point. More specifically, the vehicle control information loss can be represented by a smooth_L1 loss function, and the trajectory loss can be represented by a L1 loss function. For example, the learning rate used in the process of training the encoder and the planner according to the vehicle control information and the trajectory loss can be 0.00004.
[0147] FIG. 8 shows another exemplary flowchart of the parking method provided by the embodiments of the present application. The method 800 includes:
[0148] S810, obtaining voice information, the voice information being used to control the parking process of the intelligent driving device.
[0149] For example, the intelligent driving device can be the vehicle in the foregoing embodiments, such as the vehicle 100, or the intelligent driving device can also include other vehicles. The voice information can be the voice detected from the user in the cabin.
[0150] S820, obtaining environment perception information, the environment perception information indicating the environment around the intelligent driving device.
[0151] Exemplarily, the environment perception information can be the environment perception information in the foregoing embodiments. More specifically, the environment perception information can include images captured by the surround-view camera; or, the environment perception information can also include images captured by other cameras; or, the environment perception information can also include information captured by a radar, such as a laser point cloud, etc.
[0152] S830, controlling, according to the voice information and the environment perception information, a speed and / or a pose of the intelligent driving device in driving and / or parking into the target parking space.
[0153] In some implementations, the voice information indicates that the parking mode is a first mode or a second mode, the first mode is used for not avoiding obstacles in driving and / or parking into the target parking space, and the second mode is used for avoiding obstacles in driving and / or parking into the target parking space; S830 can be refined as: in the first mode, controlling the intelligent driving device to drive and / or park into the target parking space at a first speed; or, in the second mode, controlling the intelligent driving device to drive and / or park into the target parking space at a second speed; wherein the first speed is greater than the second speed.
[0154] In an example, the voice information indicating that the parking mode is the first mode or the second mode includes: the voice information directly indicating that the parking mode is the first mode or the second mode, for example, the voice information is “please park in the first mode”, or “please enable the first mode”, etc., and then it is determined that the parking mode is the first mode.
[0155] In another example, the voice information indicating that the parking mode is the first mode or the second mode can be: determining the parking mode as the first mode or the second mode according to semantics associated with the voice information. For example, when the semantics associated with the voice information includes related content of accelerating parking, it is determined that the parking mode is the first mode, otherwise, it is determined that the parking mode is the second mode. For example, when the voice information is “please accelerate parking”, it can be determined that the parking mode is the first mode.
[0156] Exemplarily, the first mode can be the fast parking mode or the obstacle-free parking mode in the foregoing embodiments, and the second mode can be the normal parking mode in the foregoing embodiments.
[0157] In actual implementation, the multiple parking modes of the intelligent driving device can be flexibly switched. For example, in the process of controlling the intelligent driving device to park in the first mode, when it is detected that an obstacle appears around the parking trajectory or the target parking space, the second mode is switched to park. Or, in the process of controlling the intelligent driving device to park in the first mode, when it is detected that the voice information includes a text unit related to the obstacle, or it is detected that the voice information is "switch to the second mode to park", the second mode is switched to park. For another example, in the process of controlling the intelligent driving device to park in the second mode, when it is detected that the voice information includes a text unit such as "accelerate" or "fast", or it is detected that the voice information is "switch to the first mode to park", the first mode is switched to park.
[0158] It can be understood that in the first mode or the second mode, at least one of the following can also be controlled according to the voice information: the parking pose of the intelligent driving device in the target parking space, or the pose of the intelligent driving device parking into the target pose.
[0159] In some implementations, S830 can be refined as: when the voice information indicates that there is a first obstacle in a first direction of the intelligent driving device, or a target pose of the intelligent driving device in the target parking space, generating first control information according to the position of the first obstacle or the position of the target parking space; and controlling the speed and / or pose of the intelligent driving device according to the first control information.
[0160] For example, when the intelligent driving device is a vehicle, the first control information can include the vehicle control information in the foregoing embodiments. The specific implementation of generating the first control information according to the position of the first obstacle or the position of the target parking space can refer to the description in method 400, which will not be repeated here.
[0161] When the voice information indicates the first obstacle, the first control information is used to avoid the first obstacle and / or control the intelligent driving device to slow down; or when the voice information indicates the target pose, the first control information is used to control the intelligent driving device to park in the target parking space in the target pose. The first obstacle can be a negative obstacle or a scattered obstacle.
[0162] For example, when the first obstacle is a negative obstacle, the first control information can be used to control the intelligent driving device to avoid the first obstacle; and when the first obstacle is a scattered obstacle, the first control information can be used to control the intelligent driving device to slow down.
[0163] In some embodiments, S830 can be refined as: generating the second control information for controlling the intelligent driving device to stop driving, when the voice information indicates that there is a second obstacle in the second direction of the intelligent driving device, and the environment perception information does not include information indicating the second obstacle; and controlling the speed of the intelligent driving device to be zero according to the second control information.
[0164] For example, the second obstacle can be an obstacle that is located around the intelligent driving device but cannot be perceived due to a failure of a part of sensors for acquiring the environment perception information or a change in the pose of a part of sensors for acquiring the environment perception information. For example, the second obstacle can be the pedestrian 602 in FIG. 7, and the second control information can be the vehicle control information corresponding to the part of FIG. 7. For details of generating the second control information, reference can be made to the description of the corresponding part of FIG. 7, which will not be repeated here.
[0165] In some embodiments, S830 can be refined as: determining a first environment feature according to the environment perception information, the first environment feature indicating a feature of a BEV perspective corresponding to the intelligent driving device; determining a first text feature according to the voice information, the first text feature indicating at least one of: a position of a third obstacle relative to the intelligent driving device, a target pose of the intelligent driving device in a target parking space, or a target speed of the intelligent driving device; and controlling the pose and / or speed of the intelligent driving device according to the first text feature and the first environment feature.
[0166] For example, the first environment feature can be the BEV timing feature in the foregoing embodiments, and the first text feature can be the text feature output by the text processor in the foregoing embodiments. For details of controlling the pose and / or speed of the intelligent driving device according to the first text feature and the first environment feature, reference can be made to the description of steps a) to e) in the part of FIG. 3, which will not be repeated here.
[0167] In some embodiments, controlling the pose and / or speed of the intelligent driving device according to the first text feature and the first environment feature includes: extracting a second environment feature from the first environment feature according to the first text feature; wherein when the first text feature is related to the third obstacle, the second environment feature includes a feature associated with the third obstacle; and when the first text feature is related to the target pose, the second environment feature includes a feature associated with the target parking space; and controlling the pose and / or speed of the intelligent driving device according to the first text feature and the second environment feature.
[0168] For example, the second environment feature can be the feature output by the fusion encoder in the foregoing embodiments.
[0169] In some implementations, the pose and / or speed of the intelligent driving device is controlled according to the first text feature and the first environment feature, including: inputting the first text feature and the first environment feature into a first neural network model to obtain third control information; and controlling the pose and / or speed of the intelligent driving device according to the third control information.
[0170] Exemplarily, the third control information can include the vehicle control information in the foregoing embodiments, and the third control information and the first control information or the second control information can be the same control information. The first neural network model can be the planner in FIG. 3, and more specifically, the first neural network model can be the LLM in the planner.
[0171] The parking method provided by the embodiments of the present application can control the speed and / or pose of the intelligent driving device according to the voice instruction of the user during the driving or parking of the intelligent driving device to the target parking space, can flexibly control the parking trajectory and speed of the vehicle, and can help improve the convenience and safety of the parking process, thereby improving the driving experience of the user. In addition, the parking method provided by the present application can improve the ability of the intelligent driving device to identify obstacles, and can also control the intelligent driving device to slow down and / or avoid certain obstacles that cannot be identified by the intelligent driving device. In addition, the above technical solution can enable the user to flexibly adjust the parking pose, and can help improve the driving experience of the user.
[0172] In various embodiments of the present application, the terms and / or descriptions of various embodiments are consistent and can be mutually referred to if there is no special description and logical conflict, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0173] The parking method provided by the embodiments of the present application is described in detail above in combination with FIGS. 1 to 8. The device provided by the embodiments of the present application will be described in detail below in combination with FIGS. 9 and 10. It should be understood that the description of the device embodiments corresponds to the description of the method embodiments, and therefore, the content not described in detail can be referred to the method embodiments described above, and will not be described here again for brevity.
[0174] FIG. 9 shows a schematic block diagram of a parking device 2000 provided by an embodiment of the present application, which can include units for performing the methods shown in FIGS. 4 and 8. In addition, each unit in the device 2000 is used to implement the corresponding flow of the method embodiments described above. The device 2000 includes an acquisition unit 2010, which can be used to implement the corresponding data acquisition or transceiving function. The device 2000 further includes a processing unit 2020, which can be used to implement the corresponding processing function.
[0175] Optionally, the apparatus 2000 further includes a storage unit, which can be used to store instructions and / or data. The processing unit 2020 can read the instructions and / or data in the storage unit to enable the apparatus to implement the relevant actions in the foregoing various method embodiments.
[0176] It should be understood that the specific processes by which the units perform the corresponding steps described above have been described in detail in the foregoing method embodiments, and thus will not be described here again for the sake of brevity.
[0177] It should also be understood that the apparatus 2000 here is embodied in the form of functional units. The term “module” or “unit” here can refer to an application-specific ASIC, an electronic circuit, a processor (for example, a shared processor, a dedicated processor, or a group of processors) and a memory for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components that support the described functions.
[0178] The apparatuses of the various solutions described above have the functions of implementing the corresponding steps performed by the computing platform 150 in the foregoing methods. The functions can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the functions described above; for example, the acquisition unit 2010 can be replaced by a transceiver, and other units, such as the processing unit, can be replaced by a processor, for performing the relevant processing operations in the various method embodiments.
[0179] Illustratively, the acquisition unit 2010 and the processing unit 2020 can be arranged in the vehicle 100 shown in FIG. 1, or can also be arranged in the system shown in FIG. 2. More specifically, the acquisition unit 2010 and the processing unit 2020 described above can be arranged in the regulation signal generation module 230. Illustratively, the operations performed by the acquisition unit 2010 and the processing unit 2020 described above can be performed by one processor, or can also be performed by different processors. In a specific implementation process, the one or more processors described above can be the processors arranged in the vehicle 100 shown in FIG. 1; or the apparatus 2000 described above can be a chip arranged in the vehicle 100.
[0180] In a specific implementation process, the units in the above apparatus can be integrated together, or can also be implemented independently. In one implementation, the units are integrated together to be implemented in the form of a system on a chip (SoC).
[0181] Fig. 10 is another schematic block diagram of the parking device according to an embodiment of the present application. The parking device 2100 shown in Fig. 10 can include a processor 2110, a transceiver 2120, and a memory 2130. The processor 2110, the transceiver 2120, and the memory 2130 are connected through internal connection paths. The memory 2130 is configured to store instructions, and the processor 2110 is configured to execute the instructions stored in the memory 2130 to implement the methods in the above embodiments. Alternatively, the memory 2130 can be coupled to the processor 2110 through an interface, or integrated with the processor 2110.
[0182] It should be noted that the transceiver 2120 can include, but is not limited to, a transceiving device such as an input / output interface, to enable communication between the device 2100 and other devices or communication networks.
[0183] The memory 2130 can be a volatile memory and / or a non-volatile memory. The non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM). For example, the RAM can be used as an external cache. By way of example and not limitation, the RAM includes the following varieties: static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0184] The transceiver 2120 uses a transceiving device such as, but not limited to, a transceiver, to enable communication between the device 2100 and other devices or communication networks, to receive / send data / information for implementing the methods in the above embodiments.
[0185] The embodiment of the present application further provides an intelligent driving device, which comprises the parking device 2000 or the parking device 2100 in the above embodiment.
[0186] The intelligent driving device related to the embodiment of the present application can comprise a road vehicle, a water vehicle, an air vehicle, an industrial device, an agricultural device, or an entertainment device, etc. For example, the intelligent driving device can be a vehicle, which is a general concept of a vehicle, and can be a vehicle (such as a commercial vehicle, a passenger vehicle, a motorcycle, a flying vehicle, a train, etc.), an industrial vehicle (such as a forklift, a trailer, a tractor, etc.), an engineering vehicle (such as an excavator, a bulldozer, a crane, etc.), an agricultural device (such as a mower, a harvester, etc.), an entertainment device, a toy vehicle, etc. The type of the vehicle is not limited in the embodiment of the present application.
[0187] The embodiment of the present application further provides a computer program product, which comprises computer program codes, and when the computer program codes are run on a computer, the computer program codes make the computer implement the method in the above embodiment of the present application.
[0188] The embodiment of the present application further provides a computer readable storage medium, which stores computer instructions, and when the computer instructions are run on a computer, the computer instructions make the computer implement the method in the above embodiment of the present application.
[0189] The embodiment of the present application further provides a chip, which comprises a circuit, and is used for executing the method in the above embodiment of the present application.
[0190] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the foregoing method embodiment, which will not be described herein.
[0191] In the description of the embodiment of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; "and / or" in the present application is a description of the association relationship of the associated object, which means that there can be three kinds of relationships, for example, A and / or B can represent: A alone, A and B exist at the same time, and B alone. In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, at least one of a, b, or c can represent: a, b, c, a-b, a-c, b-c, or a-b-c, wherein a, b, and c can be single or multiple.
[0192] The prefix words such as "first", "second" are used in the embodiments of the present application only for the purpose of distinguishing different described objects, and have no limitation to the position, order, priority, quantity or content of the described objects. The use of the prefix words such as ordinal numbers in the embodiments of the present application has no limitation to the described objects, and the statement of the described objects should refer to the description in the claims or embodiments, and should not be construed as superfluous limitation.
[0193] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic, and the division of the units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0194] In various embodiments of the present application, the terms and / or descriptions between various embodiments are consistent and can be mutually referred to if there is no special description and logical conflict, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0195] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.
[0196] In addition, each functional unit in various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0197] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any skilled person in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A parking method characterized by, The method comprises: obtaining voice information, the voice information being used to control a parking process of an intelligent driving device; obtaining environment perception information, the environment perception information indicating an environment around the intelligent driving device; controlling a speed and / or a pose of the intelligent driving device in a process of driving to and / or parking into a target parking space according to the voice information and the environment perception information.
2. The method of claim 1, wherein, The voice information indicates that a parking mode is a first mode or a second mode, the first mode being used for not avoiding obstacles in the process of driving to and / or parking into the target parking space, and the second mode being used for avoiding obstacles in the process of driving to and / or parking into the target parking space. The controlling the speed and / or the pose of the intelligent driving device in the process of driving to and / or parking into the target parking space according to the voice information and the environment perception information comprises: in the first mode, controlling the intelligent driving device to drive to and / or park into the target parking space at a first speed; or in the second mode, controlling the intelligent driving device to drive to and / or park into the target parking space at a second speed; wherein the first speed is greater than the second speed.
3. The method according to claim 1 or 2, characterized in that, The controlling the speed and / or the pose of the intelligent driving device in the process of driving to and / or parking into the target parking space according to the voice information and the environment perception information comprises: when the voice information indicates that there is a first obstacle in a first direction of the intelligent driving device or a target pose of the intelligent driving device in the target parking space, generating first control information according to a position of the first obstacle or a position of the target parking space; controlling the speed and / or the pose of the intelligent driving device according to the first control information.
4. The method of claim 3, wherein, When the voice information indicates the first obstacle, the first control information is used to avoid the first obstacle and / or control the intelligent driving device to decelerate; or when the voice information indicates the target pose, the first control information is used to control the intelligent driving device to park in the target pose in the target parking space.
5. The method according to claim 3 or 4, characterized in that, The first obstacle is a negative obstacle or a scattered obstacle.
6. The method according to any one of claims 1 to 5, characterized in that, The controlling the speed and / or the pose of the intelligent driving device in the process of driving to and / or parking into the target parking space according to the voice information and the environment perception information comprises: when the voice information indicates that there is a second obstacle in a second direction of the intelligent driving device and the environment perception information does not include information indicating the second obstacle, generating second control information, the second control information being used to control the intelligent driving device to stop driving; controlling the speed of the intelligent driving device to be zero according to the second control information.
7. The method according to any one of claims 1 to 6, characterized in that, The controlling the speed and / or the pose of the intelligent driving device in the process of driving to and / or parking into the target parking space according to the voice information and the environment perception information comprises: determining a first environment feature according to the environment perception information, the first environment feature indicating a feature of a bird's eye view BEV perspective corresponding to the intelligent driving device. According to the voice information, a first text feature is determined, the first text feature indicating at least one of: a position of a third obstacle relative to the intelligent driving device, a target pose of the intelligent driving device in the target parking space, or a target speed of the intelligent driving device; According to the first text feature and the first environment feature, the pose and / or speed of the intelligent driving device is controlled.
8. The method of claim 7, wherein, The control of the pose and / or speed of the intelligent driving device according to the first text feature and the first environment feature includes: According to the first text feature, a second environment feature is extracted from the first environment feature; Wherein, when the first text feature is related to the third obstacle, the second environment feature includes a feature associated with the third obstacle; when the first text feature is related to the target pose, the second environment feature includes a feature associated with the target parking space; According to the first text feature and the second environment feature, the pose and / or speed of the intelligent driving device is controlled.
9. The method of claim 8, wherein, The control of the pose and / or speed of the intelligent driving device according to the first text feature and the first environment feature includes: The first text feature and the first environment feature are input into a first neural network model to obtain third control information; Wherein, the first neural network model is a large language model LLM; According to the third control information, the pose and / or speed of the intelligent driving device is controlled.
10. A parking device, characterized in that Including: An acquisition unit is configured to acquire voice information, the voice information being used to control a parking process of an intelligent driving device; The acquisition unit is further configured to acquire environment perception information, the environment perception information indicating an environment around the intelligent driving device; A processing unit is configured to control a speed and / or pose of the intelligent driving device in a process of driving and / or parking into a target parking space according to the voice information and the environment perception information.
11. The apparatus of claim 10, wherein, The voice information indicates that a parking mode is a first mode or a second mode, the first mode being used for not avoiding obstacles in the process of driving and / or parking into the target parking space, and the second mode being used for avoiding obstacles in the process of driving and / or parking into the target parking space; The processing unit is configured to: In the first mode, control the intelligent driving device to drive and / or park into the target parking space at a first speed; or, In the second mode, control the intelligent driving device to drive and / or park into the target parking space at a second speed; Wherein, the first speed is greater than the second speed.
12. The apparatus of claim 10 or 11, wherein, The processing unit is configured to: When the voice information indicates that a first obstacle exists in a first direction of the intelligent driving device, or a target pose of the intelligent driving device in the target parking space, generate first control information according to a position of the first obstacle or a position of the target parking space; According to the first control information, the speed and / or pose of the intelligent driving device is controlled.
13. The apparatus of claim 12, wherein, When the voice information indicates the first obstacle, the first control information is used to avoid the first obstacle and / or control the intelligent driving device to slow down; or, The first control information is used to control the intelligent driving device to park in the target parking space in the target pose when the voice information indicates the target pose.
14. The apparatus of claim 12 or 13, wherein, The first obstacle is a negative obstacle or a scattered obstacle.
15. The apparatus of any one of claims 10 to 14, wherein, The processing unit is configured to: generate second control information when the voice information indicates that a second obstacle exists in a second direction of the intelligent driving device and the environment perception information does not include information indicating the second obstacle, the second control information being used to control the intelligent driving device to stop driving; control the speed of the intelligent driving device to be zero according to the second control information.
16. The apparatus of any one of claims 10 to 15, wherein, The processing unit is configured to: determine a first environment feature according to the environment perception information, the first environment feature indicating a feature of a bird's eye view BEV perspective corresponding to the intelligent driving device; determine a first text feature according to the voice information, the first text feature indicating at least one of the following: a position of a third obstacle relative to the intelligent driving device, a target pose of the intelligent driving device in the target parking space, or a target speed of the intelligent driving device; control a pose and / or speed of the intelligent driving device according to the first text feature and the first environment feature.
17. The apparatus of claim 16, wherein, The control of the pose and / or speed of the intelligent driving device according to the first text feature and the first environment feature includes: extracting a second environment feature from the first environment feature according to the first text feature; wherein, when the first text feature is related to the third obstacle, the second environment feature includes a feature associated with the third obstacle; and when the first text feature is related to the target pose, the second environment feature includes a feature associated with the target parking space; control the pose and / or speed of the intelligent driving device according to the first text feature and the second environment feature.
18. The apparatus of claim 17, wherein, The processing unit is configured to: input the first text feature and the first environment feature into a first neural network model to obtain third control information; wherein the first neural network model is a large language model LLM; control the pose and / or speed of the intelligent driving device according to the third control information.
19. A parking device, characterized in that comprises: a processor configured to execute a computer program stored in a memory to cause the apparatus to perform the method of any one of claims 1 to 9.
20. The apparatus of claim 19, wherein, The apparatus further comprises the memory. 21.An intelligent driving device, characterized in that, comprises the apparatus of any one of claims 10 to 20.
22. A computer-readable storage medium, characterized in that, instructions stored thereon, which, when executed by a processor, implement the method of any one of claims 1 to 9.
23. A computer program product, characterised in that, The computer program product comprises computer program code which, when executed by a processor, implements the method of any one of claims 1 to 9.
24. A chip, characterized by The chip comprises a circuit configured to perform the method of any one of claims 1 to 9.
Citation Information
Patent Citations
Automatic parking control system and method
CN110871792A
Parking control method and device, computer equipment and storage medium
CN111332280A
Automatic parking method, device and equipment based on voice and storage medium
CN114987448A
Method for controlling PDC through voice
CN117789712A
Method for influencing an automatic parking assistance system of a motor vehicle
DE102016209545A1