Vehicle trajectory prediction method, device and readable storage medium in parking scenarios

By using bird's-eye view image and behavior point score functions to calculate the vehicle's behavior probability in parking scenes, and combining the vehicle's trajectory historical data to generate trajectory prediction results, the shortcomings of traditional algorithms in anti-interference ability and error are solved, and the application of high accuracy and lightweight models is achieved.

CN119540842BActive Publication Date: 2025-05-23CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510107082.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

Traditional vehicle trajectory prediction algorithms have shortcomings in anti-interference ability and error, and the model is large, so they are not suitable for integration into embedded systems, which affects application.

Method used

By obtaining bird's-eye view images in the parking scene, extracting semantic information, and combining the distance and angle deviation between the driving vehicle and the parking space, the behavior point score function is used to calculate the vehicle's behavior probability, and fuse the vehicle trajectory historical data to generate trajectory prediction results.

Benefits of technology

Improve the accuracy of trajectory prediction results, and realize the lightweight design of the model, making it suitable for embedded devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119540842B_ABST
    Figure CN119540842B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of traffic control. The present invention discloses a method for predicting vehicle trajectories in a parking scenario, including: obtaining a bird's-eye view image containing all vehicles and parking spaces in the parking scenario, obtaining the distance and angle deviation between the moving vehicle and all parking spaces according to the bird's-eye view image; extracting semantic information of the bird's-eye view image; fusing the extracted semantic information, the distance and angle deviation between the moving vehicle and all parking spaces, and obtaining the behavior probability of the moving vehicle continuing to move and parking in a certain parking space through a behavior point score function; obtaining vehicle trajectory history data in the parking scenario, and inputting the behavior probability and the semantic image of the bird's-eye view image into a trajectory prediction model to generate a trajectory prediction result of the moving vehicle. The present invention captures the uncertainty in the vehicle driving process by taking the behavior probability of a certain behavior of the vehicle as the input of the trajectory prediction model, thereby improving the accuracy of the trajectory prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle trajectory prediction, and in particular to a method, device and readable storage medium for predicting vehicle trajectory in a parking scenario. Background Art

[0002] The technical background of the vehicle trajectory prediction algorithm in parking lots mainly comes from the development of intelligent transportation systems, especially the progress of vehicle trajectory prediction technology. This technology uses historical vehicle trajectory data and real-time information, through data-driven and model-based methods, combined with deep learning and collaborative control, to predict the future trajectory of vehicles. It has a wide range of applications in the fields of traffic condition prediction, traffic flow control, and autonomous driving, aiming to improve traffic efficiency and safety.

[0003] Traditional algorithms have poor anti-interference capabilities, many errors, and large models, which are not conducive to integration into embedded systems in actual engineering applications, resulting in poor applicability of the model.

[0004] Therefore, it is necessary to provide a new model algorithm to solve the above problems. Summary of the invention

[0005] The purpose of the present invention is to provide a method, device and readable storage medium for predicting vehicle trajectory in a parking scenario, which can solve at least one of the above-mentioned technical problems. The specific solution is as follows:

[0006] According to a specific embodiment disclosed in the present invention, a first aspect of the present invention provides a vehicle trajectory prediction method in a parking scenario, comprising:

[0007] Acquire a bird's-eye view image containing all vehicles and parking spaces in a parking scene, and obtain the distance and angle deviation between the moving vehicle and all parking spaces according to the bird's-eye view image;

[0008] Extracting features from the bird's-eye view image to obtain semantic information of the bird's-eye view image;

[0009] The extracted semantic information, the distance between the moving vehicle and all parking spaces, and the angle deviation are integrated, and the behavior probability of the moving vehicle continuing to move and parking in a certain parking space is obtained through the behavior point score function; wherein, the expression of the behavior point score function is:

[0010] ;

[0011] in, Mapping for long short-term memory neural networks;

[0012] For the i The vehicle is driven to jDistance to parking spaces;

[0013] For the i Angular deviation of a moving vehicle;

[0014] Indicates parking spaces with different attributes;

[0015] The vehicle trajectory history data in the parking scene is obtained, and the data is input into a trajectory prediction model together with the behavior probability and the semantic image of the bird's-eye view image to generate a trajectory prediction result of the moving vehicle.

[0016] Preferably, obtaining the distance and angle deviation between the moving vehicle and all parking spaces according to the bird's-eye view image includes:

[0017] Acquire parking space information and vehicle information in the bird's-eye view image; the parking space information includes: coordinates of each parking space, parking space boundaries, and parking space attributes; the vehicle information includes: coordinates of the center point of the current position of the moving vehicle, and the angle between the moving vehicle and the lane; obtain based on the parking space information and vehicle information:

[0018] Among the moving vehicles, the distance from the i-th moving vehicle to the j-th parking space is expressed as:

[0019] ;

[0020] The expression of the angle deviation of the i-th moving vehicle is:

[0021] ;

[0022] in, represents the coordinates of the i-th moving vehicle;

[0023] represents the coordinates of the jth parking space.

[0024] Preferably, the different attributes of the parking space include: occupied, free, reserved, and occupied.

[0025] Preferably, the step of extracting features from the bird's-eye view image to obtain semantic information of the bird's-eye view image includes:

[0026] Extracting features of the bird's-eye view image using an image feature extraction network to obtain semantic information of the bird's-eye view image;

[0027] The features of the bird's-eye view image extracted include:

[0028] The relative position of the moving vehicle and the parking space, wherein the relative position of the moving vehicle and the parking space is the distance from the center point of the moving vehicle to the center point of the parking space;

[0029] Angle, the angle being the angle between the moving vehicle and the parking space direction;

[0030] Space occupancy is used to assess whether the moving vehicle will exceed the boundary of the parking space after being parked.

[0031] Preferably, the loss function of the behavior point score function is expressed as:

[0032] ;

[0033] If the vehicle chooses the i-th parking space for parking, then , otherwise 0.

[0034] Preferably, the extraction method of the image feature extraction network includes: digitizing the image of each parking space with different attributes respectively, and transferring them into respective long-short term memory neural modules, encoding them using a dual attention mechanism module, and decoding them after global pooling to obtain semantic information.

[0035] Preferably, the trajectory prediction model includes:

[0036] Convolutional neural networks and Transformer networks;

[0037] wherein, a convolutional neural network is used to extract feature information of the semantic image of the bird's-eye view image to form a coding sequence;

[0038] The extracted feature information is combined with the historical vehicle trajectory data in the parking scenario, and the attention mechanism and position encoding technology in the Transformer network are used to capture the temporal dependency and spatial relationship of the moving vehicle under the probability of continuing to drive and parking in a certain parking space.

[0039] Preferably, the Transformer network comprises: an encoder and a decoder that introduce a causal attention mechanism;

[0040] The encoder encodes the behavior probabilities of the moving vehicle continuing to move and parking in a certain parking space into a context representation vector containing semantic information;

[0041] The decoder converts the encoded sequence into a vector through an embedding layer and then sends it to a multi-layer stacked decoder layer; each decoder layer first passes through a masked multi-head self-attention mechanism, and then uses the encoder-decoder attention mechanism to integrate the context representation vector containing semantic information and transform it through a feedforward neural network.

[0042] According to a specific embodiment disclosed in the present invention, a second aspect of the present invention discloses a vehicle trajectory prediction device in a parking scenario, comprising:

[0043] A calculation unit, used to obtain a bird's-eye view image including all vehicles and parking spaces in a parking scene, and obtain the distance and angle deviation between the moving vehicle and all parking spaces according to the bird's-eye view image;

[0044] An image processing unit, used for performing feature extraction on the bird's-eye view image to obtain semantic information of the bird's-eye view image;

[0045] The behavior probability prediction unit is used to fuse the extracted scene feature map, the distance between the moving vehicle and all parking spaces, and the angle deviation, and obtain the behavior probability of the moving vehicle continuing to move and parking in a certain parking space through a behavior point score function; wherein the expression of the behavior point score function is:

[0046] ;

[0047] in, Mapping for long short-term memory neural networks;

[0048] For the i The vehicle is driven to j Distance to parking spaces;

[0049] For the i Angular deviation of a moving vehicle;

[0050] Indicates parking spaces with different attributes;

[0051] The trajectory prediction unit is used to obtain the vehicle trajectory history data in the parking scene, input the behavior probability and the semantic image of the bird's-eye view image into a trajectory prediction model, and generate a trajectory prediction result of the moving vehicle.

[0052] According to a specific embodiment disclosed in the present invention, a third aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method for editing the content in a document as described in any one of the above items is implemented.

[0053] Compared with the prior art, the above solution disclosed in the present invention has at least the following beneficial effects:

[0054] The present invention combines the bird's-eye view image with mathematical theory to calculate the distance and angle deviation between the moving vehicle and all parking spaces; then extract features from the bird's-eye view image, fuse and further process the extracted features with the distance and angle deviation between the moving vehicle and all parking spaces, and obtain the behavior probability of the moving vehicle continuing to move and parking in a certain parking space; further combine the behavior probability with the vehicle trajectory history data and semantic image to generate corresponding multiple possible future trajectory prediction results. The present invention uses the behavior probability of a vehicle taking a certain behavior as the input of the trajectory prediction model, which can capture the uncertainty in the vehicle driving process, thereby improving the accuracy of the trajectory prediction results, while realizing the lightweight design of the model, and can be applied in embedded devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the disclosure of the present invention, and together with the specification, are used to explain the principles disclosed in the present invention. Obviously, the drawings described below are only some embodiments disclosed in the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0056] In the attached picture:

[0057] Figure 1 A flow chart of a vehicle trajectory prediction method in a parking scenario provided by the present invention;

[0058] Figure 2 A schematic diagram of an image feature extraction network according to an embodiment of the present invention;

[0059] Figure 3 It is a structural schematic diagram of a CBAM module according to an embodiment of the present invention;

[0060] Figure 4 A schematic diagram of a trajectory prediction model according to an embodiment of the present invention;

[0061] Figure 5 A schematic diagram of a vehicle trajectory prediction device in a parking scenario provided by the present invention;

[0062] Figure 6 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical scheme and advantages disclosed in the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments disclosed in the present invention, rather than all the embodiments. Based on the embodiments disclosed in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection disclosed in the present invention.

[0064] The terms used in the embodiments disclosed in the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the disclosure of the present invention. The singular forms "a", "said" and "the" used in the embodiments disclosed in the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings, and "multiple" generally includes at least two.

[0065] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0066] It should be understood that although the terms first, second, third, etc. may be used to describe in the disclosed embodiments of the present invention, these descriptions should not be limited to these terms. These terms are only used to distinguish the descriptions. For example, without departing from the scope of the disclosed embodiments of the present invention, the first may also be referred to as the second, and similarly, the second may also be referred to as the first.

[0067] It should also be noted that the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprising a ..." do not exclude the existence of other identical elements in the commodity or device including the elements.

[0068] The following is combined with Figure 1-6 Detailed description of the alternative embodiments of the present disclosure. Example

[0069] In a parking lot scenario, a vehicle's driving path often has multiple possibilities. For example, a vehicle may choose to park in a nearby parking space, choose to park in a larger parking space, or choose to continue driving to the exit to leave the parking lot.

[0070] The vehicle trajectory history data records the vehicle's driving trajectory over a period of time in the past, including location, speed, acceleration and other information.

[0071] Bird's Eye View (BEV) is a perspective of viewing an object or scene from above, similar to a bird looking down at the ground from the air. In the field of autonomous driving and robotics, BEV technology converts data obtained by sensors (such as cameras and lidar) into two-dimensional image representations to better perform tasks such as object detection and path planning.

[0072] The semantic image of the bird's-eye view is also one of the important input data, which provides detailed information about the vehicle's surroundings, such as parking space locations, lane lines, obstacles, etc. This information helps the model better understand the vehicle's driving environment and possible driving paths, which can be annotated through the Labelme software.

[0073] The first embodiment of the present invention provides a method for predicting vehicle trajectory in a parking scenario. Figure 1 As shown, the following steps are included:

[0074] Step S102: Acquire a bird's-eye view image including all vehicles and parking spaces in a parking scene, and obtain the distance and angle deviation between the moving vehicle and all parking spaces according to the bird's-eye view image.

[0075] In this step, a bird's-eye view image is obtained from the video image, and the position coordinates of each parking space in the bird's-eye view image, the parking space boundary and the parking space attributes are marked.

[0076] A rectangular coordinate system with the parking space as the horizontal axis is established, and the angular deviation between the direction of the moving vehicle and the direction of the parking space and the distance from the center point of the moving vehicle to the center point of the parking space are calculated based on the coordinates of the center points of the parking space and the moving vehicle.

[0077] Specifically, the parking space attribute refers to the use of the parking space, including: occupied, free, reserved, and occupied. In the bird's-eye view image of this embodiment, blue, green, purple, and yellow are used to represent them respectively.

[0078] Furthermore, define i The coordinates of the points corresponding to the moving vehicles are , within the planning horizon j The parking space coordinates are , where s represents a moving vehicle and p represents a parking space;

[0079] Then, i The vehicle is driven to j The distance between parking spaces is:

[0080] ,

[0081] The angular deviation is: .

[0082] Step S104: extract features from the bird's-eye view image to obtain semantic information of the bird's-eye view image.

[0083] In this embodiment, an image feature extraction network is used to extract features of the bird's-eye view image to obtain semantic information of the bird's-eye view image. The extracted features include:

[0084] The relative position of the moving vehicle and the parking space, wherein the relative position of the moving vehicle and the parking space is the distance from the center point of the moving vehicle to the center point of the parking space;

[0085] Angle, the angle being the angle between the moving vehicle and the parking space direction;

[0086] Space occupancy is used to assess whether the moving vehicle will exceed the boundary of the parking space after being parked.

[0087] Furthermore, in the image feature extraction network, the image of each parking space with different attributes is digitized separately and passed into the respective long short-term memory neural network, encoded using the dual attention mechanism module, and decoded after global pooling to obtain semantic information.

[0088] Figure 2 This is a schematic diagram of an image feature extraction network, including: an encoder consisting of a long short-term memory neural module LSTM and a CBAM (Convolutional Block Attention Module) attention mechanism module and another decoder consisting of a long short-term memory neural module LSTM.

[0089] Firstly, parking spaces with different attributes are digitized to obtain the digital feature vector sequence of parking spaces, which is then input into their respective LSTM modules to encode the parking space sequence and capture the temporal dependency in the sequence.

[0090] CBAM (Convolutional Block Attention Module) attention mechanism module, such as Figure 3As shown in the figure, it includes: channel attention mechanism and spatial attention mechanism. Among them, channel attention is used to weight each channel of LSTM output, dynamically adjust the influence of each channel of feature map on output, and highlight important feature channels; spatial attention is used to weight the spatial position of LSTM output, highlight important spatial positions, and output is the feature representation after attention weighting. The combination of the two can help the neural network pay more attention to the position information of small targets in each feature map and some detailed edge information.

[0091] Specifically, the input features , then the channel attention module one-dimensional convolution , multiply the convolution result by the original image, use the CAM output as input, and perform a two-dimensional convolution of the spatial attention module , and then multiply the output result with the original image.

[0092] The channel attention module is characterized by keeping the channel dimension unchanged and compressing the spatial dimension. This module focuses on the meaningful information in the input image. The expression is:

[0093] ;

[0094] Its working principle is: the input feature map passes through two parallel MaxPool layers and AvgPool layers, and the feature map is changed from C*H*W to C*1*1. Then, in the Share MLP module, the number of channels is first compressed to the original 1 / r, and then expanded to the original number of channels, and two activated results are obtained through the ReLU activation function. The two output results are added element by element, and then a sigmoid activation function is used to obtain the output result of Channel Attention, and then this output result is multiplied by the original image to return to the size of C*H*W.

[0095] The characteristic of the spatial attention module is that the spatial dimension remains unchanged and the channel dimension is compressed. This module focuses on the location information of the target, which is expressed as:

[0096] ;

[0097] Its working principle is: the output results of Channel Attention are subjected to maximum pooling and average pooling to obtain two 1*H*W feature maps, and then the two feature maps are spliced ​​through the Concat operation, converted into a 1-channel feature map through a 7*7 convolution, and then a sigmoid is used to obtain the feature map of Spatial Attention, and finally the output result is multiplied by the original image to return to the size of C*H*W.

[0098] Secondly, the output features of the CBAM attention mechanism processing module are globally pooled (such as global average pooling or global maximum pooling) to obtain a feature vector of fixed length.

[0099] Finally, the decoder uses LSTM to generate the output sequence.

[0100] Step S106: The extracted semantic information, the distances between the moving vehicle and all parking spaces, and the angle deviations are integrated, and the behavior probabilities of the moving vehicle continuing to move and parking in a certain parking space are obtained through a behavior point score function.

[0101] In this embodiment, the score calculation of the behavior point score includes the following aspects:

[0102] Position score: Based on the distance from the driving vehicle to the center point of the parking space, the closer the distance, the higher the score.

[0103] Direction score: Based on the angle between the moving vehicle and the parking space. The smaller the angle (the closer to parallel or perpendicular), the higher the score.

[0104] Spatial adaptability score: If the moving vehicle can be parked completely within the parking space without exceeding the boundaries, the score is high; otherwise, the score is low.

[0105] Parking space availability score: If the parking space is not occupied, the score is high; otherwise, the score is 0.

[0106] Therefore, the behavior point score function is established, and the expression is:

[0107] ;

[0108] in, Mapping for long short-term memory neural networks;

[0109] For the i The vehicle is driven to j Distance to parking spaces;

[0110] For the i Angular deviation of a moving vehicle;

[0111] Represents parking spaces with different attributes in the bird’s-eye view image, k Represents the properties of a parking space.

[0112] Furthermore, after the distances and angle deviations between the driving vehicle and all parking spaces obtained in step S104 and the semantic information obtained in step S106 are integrated, the behavior point score function is calculated to obtain the behavior probability P of the driving vehicle continuing to drive. 1and the behavior probability P of parking in a certain parking space 2 .

[0113] Among them, the expression of the loss function of the behavior point score function is:

[0114] .

[0115] Among them, if the driving vehicle selects the i If you park in a parking space, , otherwise 0.

[0116] In this embodiment, by obtaining the behavior probability of the moving vehicle as an input for subsequent trajectory prediction, interference with subsequent processing caused by deviations from a certain predicted behavior can be reduced.

[0117] Step S108: acquiring historical vehicle trajectory data in the parking scenario, and inputting the data, the behavior probability, and the semantic image of the bird's-eye view image into a trajectory prediction model to generate a trajectory prediction result of the moving vehicle.

[0118] In the trajectory prediction model, a convolutional neural network is used to extract the feature information of the semantic image of the bird's-eye view image to form a coding sequence. The extracted feature information is combined with the historical vehicle trajectory data in the parking scene, and the attention mechanism and position encoding technology in the Transformer network are used to capture the behavior probability P of the moving vehicle in continuing to move. 1 and the behavior probability P of parking in a certain parking space 2 The temporal dependencies and spatial relationships under.

[0119] Specifically, Figure 4 As shown, the Transformer network includes: an encoder and a decoder that introduce a causal attention mechanism;

[0120] The encoder is used to provide the decoder with a sequence of context representation vectors rich in global semantic information and position encoding, which captures the mutual relationship between the various semantic information in the input sequence.

[0121] Specifically, a causal attention mechanism is used to process sequence data, ensuring that the model can only focus on the current time step and the previous time step, but cannot see future information.

[0122] Add&Norm: Residual connections and layer normalization layers are used to solve some problems in the deep network training process (such as gradient disappearance, internal covariate shift, etc.), thereby improving the training efficiency and performance of the model.

[0123] In the decoder part, the semantic image is first encoded through the CNN network, and the encoding result is passed to the decoder together with the historical data of the vehicle trajectory. Inside the decoder, the input sequence is first converted into a vector through the embedding layer, and then sent to the multi-layer stacked decoder layer. Each layer of the decoder first passes through a masked multi-head self-attention mechanism to ensure that decoding only relies on the generated sequence to avoid leaking future information. Subsequently, the encoder-decoder attention mechanism is used to integrate the output information of the encoder. Finally, it is transformed through a feedforward neural network. The multi-layer decoder processes layer by layer and calculates in parallel to finally generate the path encoding of the predicted trajectory of the target sequence to achieve trajectory prediction.

[0124] In this embodiment, only the t-1 second state of the moving vehicle and the parking space is used as input to calculate the trajectory prediction, reducing excessive unnecessary inputs, achieving lightweight while ensuring accurate trajectory prediction, and the model can be fully used in embedded devices. The state of the moving vehicle is obtained through the input behavior probability and semantic image, and the t-1 second state of the parking space is obtained from the vehicle trajectory history data.

[0125] Example 2

[0126] The present invention also provides an apparatus embodiment that is consistent with the above embodiment, which is used to implement the method steps described in the above embodiment. The explanation based on the same name meaning is the same as the above embodiment, and has the same technical effect as the above embodiment, which will not be repeated here.

[0127] like Figure 5 As shown, the present invention discloses a vehicle trajectory prediction device in a parking scenario provided by the present invention, comprising:

[0128] A calculation unit 302 is used to obtain a bird's-eye view image including all vehicles and parking spaces in a parking scene, and obtain the distance and angle deviation between the moving vehicle and all parking spaces according to the bird's-eye view image;

[0129] An image processing unit 304 is used to extract features from the bird's-eye view image to obtain semantic information of the bird's-eye view image;

[0130] The behavior probability prediction unit 306 is used to fuse the extracted semantic information, the distance between the moving vehicle and all parking spaces, and the angle deviation, and obtain the behavior probability of the moving vehicle continuing to move and parking in a certain parking space through a behavior point score function; wherein the expression of the behavior point score function is:

[0131] ;

[0132] in, Mapping for long short-term memory neural networks;

[0133] For the i The vehicle is driven to j Distance to parking spaces;

[0134] For the i Angular deviation of a moving vehicle;

[0135] Indicates parking spaces with different attributes;

[0136] The trajectory prediction unit 308 is used to obtain the vehicle trajectory history data in the parking scene, input the behavior probability and the semantic image of the bird's-eye view image into a trajectory prediction model, and generate a trajectory prediction result of the moving vehicle.

[0137] Example 3

[0138] The disclosed embodiment of the present invention provides a non-volatile computer storage medium, wherein the computer storage medium stores computer executable instructions, and the computer executable instructions can execute the method steps described in the above embodiment.

[0139] Example 4

[0140] Reference below Figure 6 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the disclosed embodiment of the present invention. The terminal device in the disclosed embodiment of the present invention may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments disclosed in the present invention.

[0141] like Figure 6 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 to a random access memory (RAM) 403. In RAM 403, various programs and data required for the operation of the electronic device are also stored. The processing device 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0142] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0143] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment disclosed in the present invention are executed.

[0144] It should be noted that the computer-readable medium disclosed in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0145] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0146] Computer program code for performing the operations disclosed in the present invention may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0147] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to the various embodiments disclosed in the present invention. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0148] The units involved in the embodiments disclosed in the present invention may be implemented by software or hardware, wherein the name of a unit does not limit the unit itself in some cases.

Claims

1. A vehicle trajectory prediction method in a parking scenario, characterized in that: include: Acquire a bird's-eye view image containing all vehicles and parking spaces in a parking scene, and obtain the distance and angle deviation between the moving vehicle and all parking spaces according to the bird's-eye view image; Extracting features from the bird's-eye view image to obtain semantic information of the bird's-eye view image includes: Extracting features of the bird's-eye view image using an image feature extraction network to obtain semantic information of the bird's-eye view image; The features of the bird's-eye view image extracted include: The relative position of the moving vehicle and the parking space, wherein the relative position of the moving vehicle and the parking space is the distance from the center point of the moving vehicle to the center point of the parking space; Angle, the angle being the angle between the moving vehicle and the parking space direction; Space occupancy, used to assess whether the moving vehicle will exceed the parking space boundary after parking; The extraction method of the image feature extraction network includes: digitizing the image of each parking space with different attributes, and transferring it to the respective long-term and short-term memory neural modules, encoding it using a dual attention mechanism module, and decoding it after global pooling to obtain semantic information; The extracted semantic information, the distance between the moving vehicle and all parking spaces, and the angle deviation are integrated, and the behavior probability of the moving vehicle continuing to move and parking in a certain parking space is obtained through the behavior point score function; wherein, the expression of the behavior point score function is: Among them, Φ is the mapping of the long short-term memory neural network; ‖η i-j ‖ is the distance from the i-th moving vehicle to the j-th parking space; |Δψ [i] | is the angle deviation of the i-th moving vehicle; I [k] (0) Indicates parking spaces with different attributes, including occupied, free, reserved, and occupied; Acquire vehicle trajectory history data in the parking scenario, input the data into a trajectory prediction model together with the behavior probability and the semantic image of the bird's-eye view image, and generate a trajectory prediction result of the moving vehicle; The trajectory prediction model includes: Convolutional neural network and Transformer network; wherein the convolutional neural network is used to extract feature information of the semantic image of the bird's-eye view image to form a coding sequence; The extracted feature information and the historical data of vehicle trajectories in the parking scenario are combined to capture the temporal dependency and spatial relationship of the moving vehicle under the probability of continuing to move and parking in a certain parking space using the attention mechanism and position encoding technology in the Transformer network; The Transformer network includes: an encoder and a decoder that introduce a causal attention mechanism; The encoder encodes the behavior probabilities of the moving vehicle continuing to move and parking in a certain parking space into a context representation vector containing semantic information; The decoder converts the encoded sequence into a vector through an embedding layer and then sends it to a multi-layer stacked decoder layer; each decoder layer first passes through a masked multi-head self-attention mechanism, and then uses the encoder-decoder attention mechanism to integrate the context representation vector containing semantic information and transform it through a feedforward neural network.

2. The method according to claim 1, characterized in that The obtaining of the distance and angle deviation between the moving vehicle and all parking spaces according to the bird's-eye view image includes: Acquire parking space information and vehicle information in the bird's-eye view image; the parking space information includes: coordinates of each parking space, parking space boundaries, and parking space attributes; the vehicle information includes: coordinates of the center point of the current position of the moving vehicle, and the angle between the moving vehicle and the lane; obtain based on the parking space information and vehicle information: Among the moving vehicles, the distance from the i-th moving vehicle to the j-th parking space is expressed as: The expression of the angle deviation of the i-th moving vehicle is: in, represents the coordinates of the i-th moving vehicle; represents the coordinates of the jth parking space.

3. The method according to claim 1, characterized in that The loss function of the behavior point score function is expressed as: If the vehicle chooses the i-th parking space for parking, then Otherwise 0.

4. A vehicle trajectory prediction device in a parking scenario, characterized in that: include: A calculation unit, used to obtain a bird's-eye view image containing all vehicles and parking spaces in a parking scene, and obtain the distance and angle deviation between the moving vehicle and all parking spaces according to the bird's-eye view image; An image processing unit, used for performing feature extraction on the bird's-eye view image to obtain semantic information of the bird's-eye view image, comprises: Extracting features of the bird's-eye view image using an image feature extraction network to obtain semantic information of the bird's-eye view image; The features of the bird's-eye view image extracted include: The relative position of the moving vehicle and the parking space, wherein the relative position of the moving vehicle and the parking space is the distance from the center point of the moving vehicle to the center point of the parking space; Angle, the angle being the angle between the moving vehicle and the parking space direction; Space occupancy, used to assess whether the moving vehicle will exceed the parking space boundary after parking; The extraction method of the image feature extraction network includes: digitizing the image of each parking space with different attributes, and transferring it to the respective long-term and short-term memory neural modules, encoding it using a dual attention mechanism module, and decoding it after global pooling to obtain semantic information; The behavior probability prediction unit integrates the extracted semantic information, the distance between the moving vehicle and all parking spaces, and the angle deviation, and obtains the behavior probability of the moving vehicle continuing to move and parking in a certain parking space through a behavior point score function; wherein the expression of the behavior point score function is: Among them, Φ is the mapping of the long short-term memory neural network; ‖η i-j ‖ is the distance from the i-th moving vehicle to the j-th parking space; |Δψ [i] | is the angle deviation of the i-th moving vehicle; I [k] (0) indicates parking spaces with different attributes; A trajectory prediction unit, used to obtain the historical trajectory data of the vehicle in the parking scene, input the behavior probability and the semantic image of the bird's-eye view image into a trajectory prediction model, and generate a trajectory prediction result of the moving vehicle; The trajectory prediction model includes: Convolutional neural network and Transformer network; wherein the convolutional neural network is used to extract feature information of the semantic image of the bird's-eye view image to form a coding sequence; The extracted feature information and the historical data of vehicle trajectories in the parking scenario are combined to capture the temporal dependency and spatial relationship of the moving vehicle under the probability of continuing to move and parking in a certain parking space using the attention mechanism and position encoding technology in the Transformer network; The Transformer network includes: an encoder and a decoder that introduce a causal attention mechanism; The encoder encodes the behavior probabilities of the moving vehicle continuing to move and parking in a certain parking space into a context representation vector containing semantic information; The decoder converts the encoded sequence into a vector through an embedding layer and then sends it to a multi-layer stacked decoder layer; each decoder layer first passes through a masked multi-head self-attention mechanism, and then uses the encoder-decoder attention mechanism to integrate the context representation vector containing semantic information and transform it through a feedforward neural network.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Joint prediction method for driving intention and track of autonomous vehicle to surrounding vehicles

    CN115158364A

  • Autonomous valet parking scene-oriented vehicle motion trail prediction method

    CN117373241A