Multi-mode autonomous navigation method and system based on image and trajectory data, terminal and storage medium

By combining multimodal autonomous navigation methods with image and trajectory data, the real-time image and motion state characteristics of the smart car are extracted and fused, the problem that the existing navigation model fails to utilize the vehicle's motion state is solved, and navigation accuracy and motion prediction accuracy are improved.

CN120259988APending Publication Date: 2025-07-04SHENZHEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510301920.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing navigation models fail to effectively utilize the real-time motion state of the smart car, resulting in low navigation accuracy.

Method used

Combining image and trajectory data, the observation image group and target image group are constructed through the multimodal autonomous navigation method by obtaining real-time images and yaw angles. The multi-layer perceptron and Transformer network is used to extract features, fuse the image and trajectory features, determine the current position of the smart car and build control instructions.

Benefits of technology

It improves the accuracy of navigation success rate and subsequent action prediction, can more accurately understand the global structure and dynamic changes of the path, and enhances the stability and navigation performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259988A_ABST
    Figure CN120259988A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode autonomous navigation method, system and terminal based on image and trajectory data, and the method comprises the steps: obtaining a plurality of real-time images of a target intelligent vehicle, and constructing different image groups through the plurality of real-time images, thereby constructing different image features and trajectory data; a series of features of the image features and the trajectory data are subjected to enhancement processing and then fused into multi-modal mixed features, so that the closest target point is selected from a plurality of potential target points, and accurate navigation of the target intelligent vehicle is realized. According to the method, track data features are introduced, deep feature information of the data is extracted, and effective information in the track features is effectively extracted; meanwhile, an image track feature fusion module is designed, deep information of the image and deep information of the track are effectively fused, the model can feel important information in the image information under the guidance of the track information, the navigation success rate is increased, and the accuracy of follow-up action prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine navigation, and in particular to a multi-modal autonomous navigation method, system, terminal and computer-readable storage medium based on image and trajectory data. Background Art

[0002] With the development of computer hardware and the progress of deep learning technology in recent years, the synchronous development of both has given opportunities for robots to travel autonomously. Robots can be equipped with more powerful embedded devices as processing terminals, use cameras and computers to simulate the human visual system, and complete accurate and real-time navigation tasks by processing visual inputs through efficient navigation models.

[0003] Currently, the end-to-end visual autonomous navigation (Blind Image Quality Assessment, BIQA) method only uses visual observations during driving as model inputs, and trains the model through different scenarios of multiple datasets. The multiple datasets provide various complex scenarios, thus improving the generalization of the navigation model, and making it perform well in various scenarios both indoors and outdoors.

[0004] However, this method fails to make good use of the motion state of the intelligent vehicle itself. In the navigation task, it is very important to correctly recognize and utilize the real-time motion state of the intelligent vehicle, which can make the model pay more attention to the key information corresponding to the current motion direction of the intelligent vehicle in the image.

[0005] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0006] The main purpose of the present invention is to provide a multi-modal autonomous navigation method, system, terminal and computer-readable storage medium based on image and trajectory data, aiming to solve the problem in the existing technology that the current navigation model cannot well consider the state of the intelligent vehicle itself, resulting in low navigation accuracy of the intelligent vehicle.

[0007] To achieve the above object, the present invention provides a multi-modal autonomous navigation method based on image and trajectory data. The multi-modal autonomous navigation method based on image and trajectory data includes the following steps: Obtain multiple real-time images of a target intelligent vehicle, and construct an observation image group and a target image group according to all the real-time images to obtain multiple potential target points; Obtain the yaw angle of the target intelligent vehicle, and construct the trajectory data of the target intelligent vehicle according to all the real-time images and the yaw angle; Input the observation image group and the target image group into a multi-modal navigation model, and the image feature extraction network in the multi-modal navigation model outputs observation image features and target image features; Preprocess the trajectory data and input it into the trajectory feature extraction network in the multi-modal navigation model. The multi-modal navigation model performs multiple preliminary extractions on the preprocessed trajectory data and outputs a final feature sequence; Input the target image features and the final feature sequence into a multi-modal fusion network. After mapping the target image features and the final feature sequence, the multi-modal fusion network outputs multi-modal hybrid features; Determine the current position of the target intelligent vehicle according to the multi-modal hybrid features, select the minimum-distance target point from all the potential target points according to the current position, and construct a control instruction according to the minimum-distance target point to control the movement of the target intelligent vehicle.

[0008] Optionally, in the multi-modal autonomous navigation method based on image and trajectory data, the step of obtaining multiple real-time images of the target intelligent vehicle, constructing an observation image group and a target image group according to all the real-time images, and obtaining multiple potential target points specifically includes: Obtain multiple real-time images during the operation of the target intelligent vehicle, and count all the historical images in all the real-time images to obtain an observation image group; Obtain the starting position and the ending position of the target intelligent vehicle during the historical operation process, where the historical operation process represents the operation process according to the control instructions input by the user; Construct a sliding window according to the starting position and the ending position, and collect multiple image data at a preset frequency in the sliding window; Construct a topological map according to all the image data, and construct a target image group according to the topological map and the current image, where the current image is the last collected real-time image; Construct multiple potential target points according to the topological map, where the potential target points represent the potential positions that the target intelligent vehicle will pass through next during autonomous operation.

[0009] Optionally, in the multi-modal autonomous navigation method based on image and trajectory data, the step of obtaining the yaw angle of the target intelligent vehicle and constructing the trajectory data of the target intelligent vehicle according to all the real-time images and the yaw angle specifically includes: Obtain the yaw angle and the mileage coordinates of the target intelligent vehicle in real time; Construct a coordinate system with the mileage coordinates at a preset number of moments before the current moment as the origin, and construct the trajectory data of the target intelligent vehicle according to the yaw angle collected in real time: ; ; ; ; Among them, represents the rotation angle at the th moment, represents the yaw angle at the th moment, represents the elapsed time, and respectively represent the abscissa and ordinate at the th moment, and respectively represent the abscissa and ordinate of the origin, and respectively represent the horizontal and vertical trajectory distances of the target intelligent vehicle from the origin at the th moment, represents the trajectory data at the th moment.

[0010] Optionally, in the multi-modal autonomous navigation method based on images and trajectory data, when the observation image group and the target image group are input into the multi-modal navigation model, the image feature extraction network in the multi-modal navigation model outputs observation image features and target image features, specifically including: Input the observation image group and the target image group into the image feature extraction network of the constructed multi-modal autonomous navigation model; The image feature extraction network respectively extracts features from the observation image group and the target image group through a multi-layer perceptron, and outputs observation image features and target image features of a preset dimension: ; ; Among them, represents the observation image feature, represents the target image feature, represents the observation image group, represents the target image group, represents the multi-layer perceptron, represents forward propagation in the extraction network.

[0011] Optionally, in the multimodal autonomous navigation method based on image and trajectory data, the trajectory data is preprocessed and input into the trajectory feature extraction network in the multimodal navigation model, and the multimodal navigation model performs multiple preliminary extractions on the preprocessed trajectory data and outputs a final feature sequence, specifically including: Perform temporal sorting on all the real-time images to obtain position encodings, and add the position encodings to the trajectory data to obtain target trajectory data: ; ; Wherein, represents the position encoding, represents the time step after temporal sorting, represents the feature dimension in the real-time image, represents the feature dimension index, represents the position encoding corresponding to the even feature dimension at the th moment, represents the position encoding corresponding to the odd feature dimension at the th moment; ; Wherein, represents the target trajectory data, represents the trajectory data, represents the position encoding at the th moment; Project the target trajectory data into a preset space to obtain a first trajectory feature: ; ; ; Wherein, , and collectively represent the first trajectory feature, represents the query feature of the target trajectory data, represents the key value of the target trajectory data, represents the feature value of the target trajectory data, , and all represent learnable linear transformation weights; Perform global dependency modeling on the first trajectory feature, use the multi-head attention mechanism to perform multi-view feature extraction on the first trajectory feature, and then perform residual connection processing to obtain a second trajectory feature: ; Wherein, represents the second trajectory feature obtained after global dependency modeling, represents the residual connection, represents the first attention head, represents the second attention head, represents the th attention head, represents the weight matrix that maps the multi-head result to the output space; ; ; ; ; wherein, represents the th attention head, represents the attention mechanism, , and respectively represent the th query feature, key value and feature value, represents the length of the trajectory feature, represents the normalization function; Performing stable enhancement processing on the second trajectory feature to obtain a third trajectory feature, and using a feed-forward network to perform non-linear transformation and enhancement processing on the third trajectory feature to obtain a fourth trajectory feature: ; wherein, represents the third trajectory feature, represents layer normalization; ; wherein, represents the fourth trajectory feature, represents the activation function, and both represent learnable weight matrices, and both represent bias terms; Performing layer normalization processing on the third trajectory feature and the fourth trajectory feature, and outputting the final trajectory feature: ; wherein, represents the final trajectory feature; Repeatedly extracting the trajectory data to obtain multiple final trajectory features, and constructing a final feature sequence using a multi-layer perceptron: ; wherein, Represents the final feature sequence, Represents a multi-layer perceptron.

[0012] Optionally, in the multi-modal autonomous navigation method based on image and trajectory data, where the target image features and the final feature sequence are input into a multi-modal fusion network, and after the multi-modal fusion network maps the target image features and the final feature sequence, a multi-modal hybrid feature is output, specifically including: Input the target image features and the final feature sequence into a multi-modal fusion network, and the multi-modal fusion network projects the target image features and the final feature sequence to obtain the target query feature, target key value, target feature value of the target image features, and corresponding spatial representation: ; ; ; ; Wherein, , and respectively represent the target query feature, target key, and target feature value of the fusion feature, represents the target image feature, represents the final feature sequence, , and all represent learnable linear transformation weights, represents the spatial representation, represents the attention mechanism, represents the lengths of the target image features and the final feature sequence, represents the feature dimension in the real-time image, represents the normalization function; After performing a one-dimensional convolution operation on the spatial representation and concatenating it with the target image features, a multi-modal hybrid feature is obtained: ; Wherein, represents the multi-modal hybrid feature, represents concatenation in the channel dimension, represents performing a one-dimensional convolution operation.

[0013] Optionally, for the multi-modal autonomous navigation method based on image and trajectory data, the step of determining the current position of the target intelligent vehicle according to the multi-modal hybrid features, selecting the target point with the minimum distance among all the potential target points according to the current position, and constructing a control instruction according to the target point with the minimum distance to control the movement of the target intelligent vehicle specifically includes: Input all the potential target points into the multi-modal autonomous navigation model, and the multi-modal autonomous navigation model screens the target point with the minimum distance among all the potential target points according to the multi-modal hybrid features, where the target point with the minimum distance represents the potential target point most similar to the current image; Calculate the current position of the target intelligent vehicle in the current image, and calculate the linear velocity and angular velocity from the current position to the target point with the minimum distance; Construct a control instruction according to the linear velocity and the angular velocity to control the target intelligent vehicle to move to the target point with the minimum distance.

[0014] In addition, to achieve the above object, the present invention further provides a multi-modal autonomous navigation system based on image and trajectory data, where the multi-modal autonomous navigation system based on image and trajectory data includes: An image acquisition module, configured to acquire multiple real-time images of a target intelligent vehicle, construct an observation image group and a target image group according to all the real-time images, and obtain multiple potential target points; A trajectory acquisition module, configured to acquire the yaw angle of the target intelligent vehicle, and construct the trajectory data of the target intelligent vehicle according to all the real-time images and the yaw angle; A feature extraction module, configured to input the observation image group and the target image group into a multi-modal navigation model, and an image feature extraction network in the multi-modal navigation model outputs observation image features and target image features; A feature sorting module, configured to preprocess the trajectory data and input it into a trajectory feature extraction network in the multi-modal navigation model, and the multi-modal navigation model performs multiple preliminary extractions on the preprocessed trajectory data and outputs a final feature sequence; A feature fusion module, configured to input the target image features and the final feature sequence into a multi-modal fusion network, and the multi-modal fusion network maps the target image features and the final feature sequence and then outputs multi-modal hybrid features; A navigation module, configured to determine the current position of the target intelligent vehicle according to the multi-modal hybrid features, select the target point with the minimum distance among all the potential target points according to the current position, and construct a control instruction according to the target point with the minimum distance to control the movement of the target intelligent vehicle.

[0015] In addition, to achieve the above object, the present invention further provides a terminal, wherein the terminal includes: a memory, a processor, and a multimodal autonomous navigation program based on image and trajectory data stored on the memory and executable on the processor. When the multimodal autonomous navigation program based on image and trajectory data is executed by the processor, the steps of the multimodal autonomous navigation method based on image and trajectory data as described above are implemented.

[0016] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multimodal autonomous navigation program based on image and trajectory data. When the multimodal autonomous navigation program based on image and trajectory data is executed by a processor, the steps of the multimodal autonomous navigation method based on image and trajectory data as described above are implemented.

[0017] In the present invention, a plurality of real-time images of a target intelligent vehicle are obtained. According to all the real-time images, an observation image group and a target image group are constructed to obtain a plurality of potential target points; the yaw angle of the target intelligent vehicle is obtained, and trajectory data of the target intelligent vehicle is constructed according to all the real-time images and the yaw angle; the observation image group and the target image group are input into a multimodal navigation model, and an observation image feature and a target image feature are output by an image feature extraction network in the multimodal navigation model; the trajectory data is preprocessed and input into a trajectory feature extraction network in the multimodal navigation model, and the multimodal navigation model performs multiple preliminary extractions on the preprocessed trajectory data and outputs a final feature sequence; the target image feature and the final feature sequence are input into a multimodal fusion network, and after the multimodal fusion network maps the target image feature and the final feature sequence, a multimodal hybrid feature is output; according to the multimodal hybrid feature, the current position of the target intelligent vehicle is determined, a minimum-distance target point is selected from all the potential target points according to the current position, and a control instruction is constructed according to the minimum-distance target point to control the movement of the target intelligent vehicle; calculate the average betweenness centrality of the simulation network; construct a result analysis model, input the development data, the number of nodes of the nodes, the number of node connection relation groups, and the average betweenness centrality into the result analysis model, and output a simulation result of urban development differences. The present invention introduces trajectory data features, extracts deep feature information of the data, and effectively extracts the effective information in the trajectory features; at the same time, a fusion module of image trajectory features is designed to effectively fuse the deep information of the image and the deep information of the trajectory, enabling the model to sense the important information in the image information under the guidance of the trajectory information, improving the navigation success rate and the accuracy of subsequent action prediction. Description of the Drawings

[0018] Figure 1It is a flowchart of a preferred embodiment of the multimodal autonomous navigation method based on image and trajectory data of the present invention; Figure 2 It is a framework diagram of a preferred embodiment of the multimodal autonomous navigation method based on image and trajectory data of the present invention; Figure 3 It is a structural diagram of a feature extraction network of a preferred embodiment of the multimodal autonomous navigation method based on image and trajectory data of the present invention; Figure 4 It is a structural diagram of a trajectory feature extraction network of a preferred embodiment of the multimodal autonomous navigation method based on image and trajectory data of the present invention; Figure 5 It is a structural diagram of a multimodal fusion model of a preferred embodiment of the multimodal autonomous navigation method based on image and trajectory data of the present invention; Figure 6 It is a schematic diagram of the acquisition effect of a preferred embodiment of the multimodal autonomous navigation method based on image and trajectory data of the present invention; Figure 7 It is a visualization schematic diagram of the navigation effect of a preferred embodiment of the multimodal autonomous navigation method based on image and trajectory data of the present invention; Figure 8 It is a schematic diagram of the outdoor navigation result of a preferred embodiment of the multimodal autonomous navigation method based on image and trajectory data of the present invention; Figure 9 It is a structural diagram of a preferred embodiment of the multimodal autonomous navigation system based on image and trajectory data of the present invention; Figure 10 It is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed implementation manners

[0019] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the present invention will be further described in detail below with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0020] In the navigation task, it is very important to correctly recognize the real-time motion state of the intelligent vehicle and make use of it, which can enable the model to pay more attention to the key information corresponding to the current motion direction of the intelligent vehicle in the image.

[0021] The multimodal autonomous navigation method based on image and trajectory data according to a preferred embodiment of the present invention, as Figure 1 shown, the multimodal autonomous navigation method based on image and trajectory data includes the following steps: Step S10: Obtain multiple real-time images of the target intelligent vehicle, and construct an observation image group and a target image group according to all the real-time images to obtain multiple potential target points.

[0022] Among them, as Figure 2 shown, in this embodiment, an end-to-end multi-modal autonomous navigation model Vi-TrajNav (Vi, Visual; TrajNav, Trajectory Navigation) is constructed, which generally includes a feature extraction network, a multi-modal fusion model, and a feature decoding output network.

[0023] Specifically, multiple real-time images during the operation of the target intelligent vehicle are obtained, and all historical images in all the real-time images are counted to obtain an observation image group; the starting position and the ending position of the target intelligent vehicle during the historical operation process are obtained, where the historical operation process represents the operation process according to the control instructions input by the user; a sliding window is constructed according to the starting position and the ending position, and multiple image data are collected at a preset frequency in the sliding window; a topological map is constructed according to all the image data, and a target image group is constructed according to the topological map and the current image, where the current image is the last collected real-time image; according to the topological map, multiple potential target points are constructed, where the potential target points represent the potential positions that the target intelligent vehicle will pass through in the next step during autonomous operation.

[0024] Among them, the structure of the feature extraction network is as Figure 3 shown, specifically including an image feature extraction network and a trajectory feature extraction network, and the obtained real-time images are processed by the image feature extraction network.

[0025] Further, before inputting the real-time images into the image feature extraction network, two types of image groups need to be constructed. One part is composed of the observation image at the current moment and the images observed at the previous multiple moments (for example, if five moments are taken, then h = 1, 2, 3, 4, 5), and this group of image groups is called the observation image group ; the other part is composed of the observation image at the current moment and the images of the topological map , that is, the target image group . Among them, the topological map is the image data collected by researchers at a specific frequency during the process of manually controlling the intelligent vehicle to drive on the corresponding route, and is numbered in the order of collection time.

[0026] Among them, the real-time image is an RGB image of the scene in front of the intelligent vehicle, which is collected in real time by a front camera and is composed of the colors of the red, blue, and green channels, and rich environmental overall and detailed information of the scene in front of the intelligent vehicle can be obtained.

[0027] Step S20: Obtain the yaw angle of the target intelligent vehicle, and construct the trajectory data of the target intelligent vehicle according to all the real-time images and the yaw angle.

[0028] Among them, the trajectory data is collected by the odometer sensor of the intelligent vehicle. In the odometer sensor, the odometer coordinates of the vehicle at the th moment are recorded , as well as the yaw angle . Among them, each piece of trajectory data sent into the trajectory feature extraction network (whose structure is as Figure 4 shown) is calculated by taking the fifth moment before the current moment as the coordinate origin and combining the yaw angle of the odometer data and the waypoint.

[0029] Specifically, the yaw angle and mileage coordinates of the target intelligent vehicle are obtained in real time; a coordinate system is constructed with the mileage coordinates of a preset number of moments before the current moment as the origin, and the trajectory data of the target intelligent vehicle is constructed according to the yaw angle collected in real time: ; ; ; ; Among them, represents the rotation angle at the th moment, represents the yaw angle at the th moment, represents the elapsed time, and respectively represent the abscissa and ordinate at the th moment, and respectively represent the abscissa and ordinate of the origin, and respectively represent the horizontal and vertical trajectory distances of the target intelligent vehicle from the origin at the th moment, represents the trajectory data at the th moment.

[0030] Among them, the yaw angle of the vehicle at the th moment can be used to calculate a rotation matrix corresponding to this angle. Then, the coordinates of each subsequent moment are subtracted from the coordinates at the th moment, and then multiplied by the rotation matrix through matrix multiplication operation, so as to use the Taking a certain moment as the coordinate origin, with its facing direction as the positive x-axis direction and the left-hand side as the positive y-axis direction, a relative coordinate system is established. Calculate the relative coordinate points at the next five moments in this coordinate system, and represent the historical trajectory data of the intelligent vehicle with this set of coordinate points.

[0031] Step S30: Input the observation image group and the target image group into the multi-modal navigation model, and the image feature extraction network in the multi-modal navigation model outputs the observation image feature and the target image feature.

[0032] Specifically, input the observation image group and the target image group into the image feature extraction network of the constructed multi-modal autonomous navigation model; the image feature extraction network respectively extracts features from the observation image group and the target image group through a multi-layer perceptron, and outputs the observation image feature and the target image feature of a preset dimension: ; ; Among them, represents the observation image feature, represents the target image feature, represents the observation image group, represents the target image group, represents the multi-layer perceptron, represents the forward propagation in the extraction network.

[0033] Among them, the image feature extraction network (image feature extraction model) is mainly composed of a backbone network and a multi-layer perceptron MLP (Multi-Layer Perceptron‌). For the image branch in the network, the input consists of the observation image group and the target image group In this embodiment, EdgeVit (Edge represents the edge computing scenario, Vision Transformers represents the self-attention mechanism) is selected as the backbone network for image feature extraction, which combines the efficiency of the convolutional neural network and the powerful global modeling ability of the Transformer. While retaining the core advantages of the Transformer, it significantly reduces the computational overhead and the number of parameters, and is very suitable for application in the task of autonomous navigation. Sending the observation image group and the target image group into the EdgeVit backbone network and performing feature dimension elevation through the multi-layer perceptron MLP, the image features of the target image group and the observation image group are obtained.

[0034] Furthermore, the MLP consists of four linear layers and an activation function, and can extract the dimension of the feature information to 512.

[0035] Step S40: Preprocess the trajectory data and input it into the trajectory feature extraction network in the multimodal navigation model. The multimodal navigation model performs multiple preliminary extractions on the preprocessed trajectory data and outputs a final feature sequence.

[0036] Among them, the trajectory data is obtained through an odometer sensor. In this embodiment, real-time trajectory data of an intelligent vehicle is obtained through the odometer sensor , which reflects the position of the intelligent vehicle at each moment in a two-dimensional plane coordinate, and includes the relative coordinates of the intelligent vehicle from the th moment to the th moment, a total of 6 coordinate points, which can reflect the movement trajectory of the intelligent vehicle in real time. The trajectory data records the differential features of physical quantities such as wheel rotation angle and speed change in a short time, and these motion states reflect the current motion state of the intelligent vehicle in real time.

[0037] Specifically, time-order all the real-time images to obtain a position encoding, and add the position encoding to the trajectory data to obtain target trajectory data: ; ; Among them, represents the position encoding, represents the time step after time sorting, represents the feature dimension in the real-time image, represents the feature dimension index, represents the position encoding corresponding to the even feature dimension at the th moment, represents the position encoding corresponding to the odd feature dimension at the th moment; ; Among them, represents the target trajectory data, represents the trajectory data, represents the position encoding at the th moment; Project the target trajectory data into a preset space to obtain a first trajectory feature: ; ; ; Among them, , and together represent the first trajectory feature, represents the query feature of the target trajectory data, The key value representing the target trajectory data, The eigenvalue representing the target trajectory data, 、 and all represent learnable linear transformation weights; perform global dependence modeling on the first trajectory feature, use the multi-head attention mechanism to perform feature extraction from multiple perspectives on the first trajectory feature, and then perform residual connection processing to obtain the second trajectory feature: ; Among them, represents the second trajectory feature obtained after global dependence modeling, represents the residual connection, represents the first attention head, represents the second attention head, represents the th attention head, represents the weight matrix that maps the multi-head result to the output space; ; ; ; ; Among them, represents the th attention head, represents the attention mechanism, 、 and respectively represent the th query feature, key value and eigenvalue, represents the length of the trajectory feature, represents the normalization function; perform stable enhancement processing on the second trajectory feature to obtain the third trajectory feature, and use the feed-forward network to perform non-linear transformation and enhancement processing on the third trajectory feature to obtain the fourth trajectory feature: ; Among them, represents the third trajectory feature, represents layer normalization; ; Among them, represents the fourth trajectory feature, represents the activation function, and both represent learnable weight matrices, and Both represent the bias term; perform layer normalization on the third trajectory feature and the fourth trajectory feature, and output the final trajectory feature: ; Among them, represents the final trajectory feature; repeatedly extract the trajectory data to obtain multiple final trajectory features, and use a multi-layer perceptron to construct the final feature sequence: ; Among them, represents the final feature sequence, represents the multi-layer perceptron.

[0038] Among them, in this embodiment, by introducing a trajectory feature extraction network based on Transformer, the time series characteristics and global dependencies of the trajectory data are effectively modeled; at the same time, position encoding is introduced to capture the time order of the trajectory data, making full use of the running state of the intelligent vehicle and improving the accuracy of the trajectory data.

[0039] Furthermore, according to the multi-head self-attention mechanism, calculate the correlation matrix inside the trajectory data, which can capture the global dependencies between each time step, while paying attention to the local dynamic characteristics, capturing the global context information of the trajectory, and the residual connection and normalization processing improve the stability of the trajectory features, ensuring that the problem of gradient disappearance can be effectively alleviated during the information transmission process and enhancing the training stability of the model; after being processed by the self-attention mechanism, the trajectory features also need to be non-linearly transformed and enhanced through a feed-forward network, and its structure includes two fully connected layers and an activation function.

[0040] Stack the above-mentioned network layers to form a Transformer block to complete the preliminary extraction of the trajectory features, then pass through N such Transformer blocks for further extraction of the trajectory features, and finally obtain the final feature sequence through the MLP.

[0041] Step S50, input the target image feature and the final feature sequence into the multi-modal fusion network. After the multi-modal fusion network maps the target image feature and the final feature sequence, output the multi-modal hybrid feature.

[0042] Among them, for the two types of modal data input to the model, the image data belongs to high-dimensional grid information, while the trajectory data is a low-dimensional vector sequence. Different from the multi-modal fusion methods that match the dimensions of image-depth maps, image-point clouds, etc., when dealing with heterogeneous data with significantly different cross-domain features, directly performing feature concatenation or element addition operations in the feature extraction stage may cause serious modal interference.

[0043] Therefore, in this embodiment, a multimodal fusion network is disclosed. As Figure 5 shown, by adopting a cross-attention gating mechanism, the image features and trajectory features are effectively integrated, thereby enhancing the model's perception ability of image information.

[0044] Specifically, the target image features and the final feature sequence are input into the multimodal fusion network, and the multimodal fusion network projects the target image features and the final feature sequence to obtain the target query feature, target key-value, target feature value, and corresponding spatial representation of the target image features: ; ; ; ; Among them, , and respectively represent the target query feature, target key, and target feature value of the fusion feature, represents the target image feature, represents the final feature sequence, , and all represent learnable linear transformation weights, represents the spatial representation, represents the attention mechanism, represents the lengths of the target image feature and the final feature sequence, represents the feature dimension in the real-time image, represents the normalization function; after performing a one-dimensional convolution operation on the spatial representation, it is concatenated with the target image feature to obtain the multimodal hybrid feature: ; Among them, represents the multimodal hybrid feature, represents concatenation in the channel dimension, represents performing a one-dimensional convolution operation.

[0045] Among them, the target image group features and trajectory features with a length of are input into the multimodal fusion network, and the multimodal fusion network first maps out new Query, Key, and Value values (i.e., , and ), and then performs learnable linear transformation weights on the image features and trajectory features respectively, so as to project the input features and map them to a new representation space.

[0046] Furthermore, by calculating the similarity matrix between the Query and the Key, and then according to the formula of the spatial representation, performing a multiplication operation with the Value value, the fused attention score is obtained. This fusion method enables the trajectory features to effectively guide the attention distribution of the image features, thereby prompting the model to focus on the image regions related to the moving direction or target position of the intelligent vehicle.

[0047] Furthermore, in this embodiment, at the one-dimensional convolution operation, the convolution kernel size is set to 3 to capture the correlation and local features between adjacent elements in the sequence, thereby enhancing the expression ability of the features. At the output stage of the fusion module, the original image features and the fused features after one-dimensional convolution processing are further concatenated through residual connection to obtain the final multi-modal hybrid features.

[0048] Step S60: Determine the current position of the target intelligent vehicle according to the multi-modal hybrid features, select the minimum-distance target point from all the potential target points according to the current position, and construct a control instruction according to the minimum-distance target point to control the movement of the target intelligent vehicle.

[0049] Among them, through the processing of the feature extraction network and the fusion module, two parts of key feature information are obtained. The first part is the output from the observation image group and the target image group , and the second part is the result after fusing the target image group with the trajectory features , ( represents the real number space).

[0050] Specifically, input all the potential target points into the multi-modal autonomous navigation model. The multi-modal autonomous navigation model screens the minimum-distance target point from all the potential target points according to the multi-modal hybrid features. Among them, the minimum-distance target point represents the potential target point most similar to the current image; calculate the current position of the target intelligent vehicle in the current image, and calculate the linear velocity and angular velocity from the current position to the minimum-distance target point; construct a control instruction according to the linear velocity and the angular velocity to control the target intelligent vehicle to move to the minimum-distance target point.

[0051] Among them, based on the multi-modal hybrid features obtained above, the minimum distance between the current position of the intelligent vehicle and the potential target points can be calculated, select the potential sub-target point most similar to the current observation image as the sub-target point to be reached currently, and plan the path for the intelligent vehicle to reach this sub-target point based on the waypoints returned by the model.

[0052] Specifically, according to the formula in the following code, the corresponding linear velocity and angular velocity are calculated and fed back to the intelligent vehicle for control: " Input: The current observed image of the model and the set of the first five frame images , topological graph , real-time trajectory of the intelligent vehicle

[0053] Initialize

[0054] While( ) do if then

[0055] else

[0056] if then

[0057] else

[0058] for in do

[0059]

[0060] if then

[0061] else

[0062]

[0063]

[0064] reach goal”.

[0065] Repeat the above process for each sub-goal point until the final goal point is selected as the sub-goal point and the algorithm terminates, realizing the navigation of the complete path.

[0066] Furthermore, as Figure 6As shown in the figure, for the autonomous driving test of the simulation platform, the ROACH (Reconfigurable Open Architecture Computing Hardware) expert system was used as the control strategy to generate data for navigation data collection. Finally, a dataset covering six towns with significantly different geographical features was constructed. 100 routes were collected in each town, and the total driving duration was approximately 22 hours. After unpacking through ROS (Robot Operating System), a total of 123,135 valid training images were obtained and could be used for the training of the autonomous navigation task. Through diverse scenario and route selection, the comprehensiveness and diversity of the data were ensured, providing a multi-modal training base with spatio-temporal alignment and covering long-tail scenarios for the end-to-end autonomous driving model. Figure 6 The figure shows the acquisition effect diagram of one of the driving routes.

[0067] For the implementation experiment, three different driving datasets were selected for training in this embodiment. The three datasets cover urban, rural, indoor and outdoor environments, comprehensively demonstrating the complexity of navigation tasks in open environments and providing rich test scenarios for the development and verification of navigation algorithms. During the collection process of these datasets, multi-modal sensors such as stereo RGB cameras, thermal imaging cameras, 2D LiDAR (Light Detection and Ranging), GPS, and IMU (Inertial Measurement Unit) were used to provide rich multi-modal data for the trainer of the autonomous navigation model.

[0068] Furthermore, in order to evaluate the performance of the autonomous navigation model, the navigation success rate (SR), the number of parameters (Parameters), and the frames per second (FPS) were used as the evaluation indicators in this embodiment. Starting from the starting point, without human interference, the intelligent vehicle is considered to have successfully navigated if it reaches within 1 meter of the indoor destination and within 3 meters of the outdoor destination without collision. The navigation success rate (SR) is the number of successfully navigated routes accounting for the total number of routes and is defined as follows: .

[0069] Furthermore, in order to verify the usability of the autonomous navigation model in this embodiment first, navigation tests were carried out on the CARLA platform (CarLearning to Act, a simulation tool) first. The experimental design included 30 different routes with varying lengths. Among them, 18 routes were collected in Town 5, and 12 routes were collected in Town 7. According to the differences in the length, number of curves and intersections of the routes, these routes were divided into simple routes and difficult routes. The experiment was carried out on a computer equipped with an NVIDIA GeForce GTX 1080Ti graphics card, an Intel i7-9700K CPU and 16GB of memory. The results are shown in Table 1 below: Table 1: Platform Test Results Table

[0070] Among them, Gnm represents Global Network Management, global network management; ViNT represents Virtual InterNetwork Testbed, virtual network testbed; FPS represents Frames Per Second, frames per second. Due to the limitations of the shallow architecture, the Gnm navigation model cannot complete the navigation task in most scenarios. The performance of ViNT is medium. It can run normally in the navigation tasks of some short-distance and simple routes, but in long-distance routes with multiple curves, the navigation performance drops significantly. After replacing the backbone network of ViNT with EdgeViT, under the backbone network supported by the ViT architecture, the navigation performance is significantly improved. In contrast, the navigation success rate of the multimodal method proposed in this embodiment reaches 0.83. This shows that by introducing trajectory features and performing multimodal fusion, the model can more accurately understand the global structure and dynamic changes of the path.

[0071] For indoor scenarios, in this embodiment, office scenarios with rich obstacles and corridors with pedestrians were selected. The longest indoor route reached more than 100m. For outdoor scenarios, asphalt roads, stone brick roads and other scenarios were included, and there were also rich dynamic obstacles, such as pedestrians, bicycles in motion, cars in motion, etc. The longest outdoor route reached 273m. A total of 7 routes were selected for both indoor and outdoor, and a total of 14 routes were used as test routes for the navigation task. The conclusions shown in Table 2 and Figure 7 are as follows: Table 2: Real World Navigation Test Results Table

[0072] The above process is repeated at each sub-goal point until the final goal point is selected as the sub-goal point and the algorithm terminates, realizing the navigation of the complete path.

[0073] Different from virtual scenarios, real-world scenarios are easily affected by dynamic environments, which pose higher requirements for the navigation performance of the model. Figure 7 The navigation experiment results of the multi-modal navigation model in the real scenario of this embodiment are shown. The first row shows the first perspective of the intelligent vehicle's camera and the corresponding action outputs, and the second row shows the navigation process recorded by the researcher from the third perspective. The experiment shows that in complex road conditions such as curves, the proposed model can correctly execute steering decisions and stably drive along the predetermined route to the target position, verifying its effectiveness and robustness in the real environment.

[0074] Furthermore, as Figure 8 shown, the specific performance of the longest driving route tested in the outdoor environment on two models is presented. The experimental results show that the ViNT model has a problem of steering failure at the first intersection, while the method in this embodiment can correctly execute operations such as turning and going straight. The driving trajectory maintains a high consistency with the predetermined route and successfully completes collision-free navigation throughout the whole process. Finally, the length of the tested driving route reaches 271 meters, verifying the navigation execution ability of the algorithm in complex outdoor environments.

[0075] For the existing technical solutions, the current methods only rely on image data. After feature extraction through a single-branch network, the time distance and actions are predicted. Although this method has advantages in terms of the number of parameters and computational speed, since the time distance and action prediction are two tasks with different natures, it is difficult for a single decoder to achieve ideal effects in both tasks simultaneously.

[0076] Therefore, this embodiment proposes a two-branch architecture for independent prediction of time distance and actions respectively. The core advantage of this architecture lies in task decoupling. By configuring independent decoders for each task, more accurate feature modeling can be achieved. The two decoders are independent of each other and both adopt the structure of the Transformer decoder. By using multiple attention heads to parallelly model the interaction relationships between different time steps, the prediction accuracy is improved. Then, the MLP is used to reduce the high-dimensional feature information to the values required for our tasks. Although the two-branch structure will increase the number of parameters of the model, the impact on the computational complexity is negligible and still meets the requirements of real-time operation.

[0077] The present invention introduces trajectory data features, extracts the deep feature information of the data, and effectively extracts the effective information in the trajectory features. At the same time, a fusion module for image trajectory features is designed to effectively fuse the deep information of the image and the deep information of the trajectory, enabling the model to perceive the important information in the image information under the guidance of the trajectory information, improving the navigation success rate and the accuracy of subsequent action prediction.

[0078] Further, as Figure 9 shown, based on the above multi-modal autonomous navigation method based on image and trajectory data, the present invention also correspondingly provides a multi-modal autonomous navigation system based on image and trajectory data, wherein the multi-modal autonomous navigation system based on image and trajectory data includes: An image acquisition module 51, configured to acquire a plurality of real-time images of the target intelligent vehicle, and construct an observation image group and a target image group according to all the real-time images to obtain a plurality of potential target points; A trajectory acquisition module 52, configured to acquire the yaw angle of the target intelligent vehicle, and construct the trajectory data of the target intelligent vehicle according to all the real-time images and the yaw angle; A feature extraction module 53, configured to input the observation image group and the target image group into a multi-modal navigation model, and an image feature extraction network in the multi-modal navigation model outputs observation image features and target image features; A feature sorting module 54, configured to preprocess the trajectory data and input it into a trajectory feature extraction network in the multi-modal navigation model, and the multi-modal navigation model performs multiple preliminary extractions on the preprocessed trajectory data and outputs a final feature sequence; A feature fusion module 55, configured to input the target image features and the final feature sequence into a multi-modal fusion network, and the multi-modal fusion network maps the target image features and the final feature sequence and outputs multi-modal hybrid features; A navigation module 56, configured to determine the current position of the target intelligent vehicle according to the multi-modal hybrid features, select a minimum distance target point from all the potential target points according to the current position, and construct a control instruction according to the minimum distance target point to control the movement of the target intelligent vehicle.

[0079] Further, as Figure 10 shown, based on the above multi-modal autonomous navigation method and system based on image and trajectory data, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20 and a display 30. Figure 10 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0080] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. In some other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the terminal. Further, the memory 20 may also include both the internal storage unit of the terminal and the external storage device. The memory 20 is used to store application software installed on the terminal and various types of data, such as the program code of the installed terminal, etc. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a multi-modal autonomous navigation program 40 based on image and trajectory data is stored on the memory 20, and the multi-modal autonomous navigation program 40 based on image and trajectory data can be executed by the processor 10, so as to implement the multi-modal autonomous navigation method based on image and trajectory data in the present application.

[0081] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips, and is used to run the program code stored in the memory 20 or process data, such as executing the multi-modal autonomous navigation method based on image and trajectory data, etc.

[0082] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. The display 30 is used to display information on the terminal and to display a visual user interface. Components of the terminal communicate with each other through a system bus.

[0083] In one embodiment, when the processor 10 executes the multi-modal autonomous navigation program 40 based on image and trajectory data in the memory 20, the steps of the multi-modal autonomous navigation method based on image and trajectory data as described above are implemented.

[0084] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multi-modal autonomous navigation program based on image and trajectory data, and when the multi-modal autonomous navigation program based on image and trajectory data is executed by a processor, the steps of the multi-modal autonomous navigation method based on image and trajectory data as described above are implemented.

[0085] In summary, the present invention provides a multimodal autonomous navigation method and related devices based on image and trajectory data. The method includes: obtaining multiple real-time images of a target intelligent vehicle, constructing an observation image group and a target image group based on all the real-time images to obtain multiple potential target points; obtaining the yaw angle of the target intelligent vehicle, and constructing the trajectory data of the target intelligent vehicle based on all the real-time images and the yaw angle; inputting the observation image group and the target image group into a multimodal navigation model, and the image feature extraction network in the multimodal navigation model outputs observation image features and target image features; preprocessing the trajectory data and inputting it into the trajectory feature extraction network in the multimodal navigation model, and the multimodal navigation model performs multiple preliminary extractions on the preprocessed trajectory data and outputs a final feature sequence; inputting the target image features and the final feature sequence into a multimodal fusion network, and the multimodal fusion network maps the target image features and the final feature sequence and then outputs multimodal hybrid features; determining the current position of the target intelligent vehicle according to the multimodal hybrid features, selecting a minimum-distance target point from all the potential target points according to the current position, and constructing a control instruction according to the minimum-distance target point to control the movement of the target intelligent vehicle; calculating the average betweenness centrality of the simulation network; constructing a result analysis model, inputting the development data, the number of nodes of the nodes, the number of node connection relationship groups, and the average betweenness centrality into the result analysis model, and outputting a simulation result of urban development differences. The present invention introduces trajectory data features, extracts deep feature information of the data, and effectively extracts the effective information in the trajectory features; at the same time, a fusion module for image trajectory features is designed to effectively fuse the deep information of the image and the deep information of the trajectory, enabling the model to feel the important information in the image information under the guidance of the trajectory information, improving the navigation success rate and the accuracy of subsequent action prediction.

[0086] It should be noted that in this article, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal including that element.

[0087] Of course, those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0088] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A multi-modal autonomous navigation method based on image and trajectory data, characterized in that, The multi-modal autonomous navigation method based on image and trajectory data includes: Obtain multiple real-time images of the target intelligent vehicle, and based on all the real-time images, construct an observation image group and a target image group to obtain multiple potential target points; Obtain the yaw angle of the target intelligent vehicle, and construct the trajectory data of the target intelligent vehicle based on all the real-time images and the yaw angle; Input the observation image group and the target image group into a multi-modal navigation model, and the image feature extraction network in the multi-modal navigation model outputs observation image features and target image features; Preprocess the trajectory data and input it into the trajectory feature extraction network in the multi-modal navigation model. The multi-modal navigation model performs multiple preliminary extractions on the preprocessed trajectory data and outputs a final feature sequence; Input the target image features and the final feature sequence into a multi-modal fusion network. After mapping the target image features and the final feature sequence, the multi-modal fusion network outputs multi-modal hybrid features; Determine the current position of the target intelligent vehicle according to the multi-modal hybrid features, select the target point with the minimum distance from all the potential target points according to the current position, and construct a control instruction according to the target point with the minimum distance to control the movement of the target intelligent vehicle.

2. The multimodal autonomous navigation method based on image and trajectory data according to claim 1, wherein The step of obtaining multiple real-time images of the target intelligent vehicle, and based on all the real-time images, constructing an observation image group and a target image group to obtain multiple potential target points specifically includes: Obtain multiple real-time images during the operation of the target intelligent vehicle, and count all the historical images in all the real-time images to obtain an observation image group; Obtain the starting position and the ending position of the target intelligent vehicle during the historical operation process, where the historical operation process represents the operation process according to the control instructions input by the user; Construct a sliding window according to the starting position and the ending position, and collect multiple image data at a preset frequency in the sliding window; Construct a topological map according to all the image data, and construct a target image group according to the topological map and the current image, where the current image is the last collected real-time image; Construct multiple potential target points according to the topological map, where the potential target points represent the potential positions that the target intelligent vehicle will pass through next during autonomous operation.

3. The multimodal autonomous navigation method based on image and trajectory data according to claim 1, wherein The step of obtaining the yaw angle of the target intelligent vehicle, and constructing the trajectory data of the target intelligent vehicle based on all the real-time images and the yaw angle specifically includes: Obtain the yaw angle and the mileage coordinates of the target intelligent vehicle in real time; Construct a coordinate system with the mileage coordinates of a preset number of previous moments at the current moment as the origin, and construct the trajectory data of the target intelligent vehicle according to the real-time collected yaw angle: ; ; ; ; Among them, represents the rotation angle at the moment, represents the yaw angle at the moment, represents the elapsed time, and respectively represent the abscissa and ordinate at the moment, and respectively represent the abscissa and ordinate of the origin, and respectively represent the abscissa trajectory distance and ordinate trajectory distance of the target intelligent vehicle from the origin at the moment, represents the trajectory data at the moment.

4. The multimodal autonomous navigation method based on image and trajectory data according to claim 1, characterized in that The step of inputting the observation image group and the target image group into a multi-modal navigation model, and the image feature extraction network in the multi-modal navigation model outputs observation image features and target image features specifically includes: Input the observation image group and the target image group into the image feature extraction network of the constructed multi-modal autonomous navigation model; The image feature extraction network extracts features from the observation image group and the target image group respectively through a multi-layer perceptron, and outputs observation image features and target image features of a preset dimension: ; ; Among them, represents the observed image features, represents the target image features, represents the observed image group, represents the target image group, represents a multi-layer perceptron, represents forward propagation in the extraction network.

5. The multimodal autonomous navigation method based on image and trajectory data according to claim 1, wherein Preprocess the trajectory data and input it into the trajectory feature extraction network in the multi-modal navigation model. The multi-modal navigation model performs multiple preliminary extractions on the preprocessed trajectory data and outputs a final feature sequence, specifically including: Sort all the real-time images in time to obtain a position encoding, and add the position encoding to the trajectory data to obtain target trajectory data: ; ; Among them, represents the position encoding, represents the time step after time sorting, represents the feature dimension in the real-time image, represents the feature dimension index, represents the position encoding corresponding to the even feature dimension at the represents the position encoding corresponding to the odd feature dimension at the ; Among them, represents the target trajectory data, represents the trajectory data, represents the position encoding at the moment; Project the target trajectory data into a preset space to obtain a first trajectory feature: ; ; ; Among them, , and collectively represent the first trajectory feature, represents the query feature of the target trajectory data, represents the key value of the target trajectory data, represents the feature value of the target trajectory data, , and all represent learnable linear transformation weights; Perform global dependence modeling on the first trajectory feature, use the multi-head attention mechanism to perform feature extraction from multiple perspectives on the first trajectory feature, and then perform residual connection processing to obtain a second trajectory feature: ; Among them, represents the second trajectory feature obtained after global dependence modeling, represents residual connection, represents the first attention head, represents the second attention head, represents the th attention head, represents the weight matrix that maps the multi-head result to the output space; ; ; ; ; Among them, represents the th attention head, represents the attention mechanism, , and respectively represent the th query feature, key-value, and feature value, represents the length of the trajectory feature sequence, represents the normalization function; Perform stable enhancement processing on the second trajectory feature to obtain a third trajectory feature, and use a feed-forward network to perform non-linear transformation and enhancement processing on the third trajectory feature to obtain a fourth trajectory feature: ; Among them, represents the third trajectory feature, represents layer normalization; ; Among them, represents the fourth trajectory feature, represents the activation function, and both represent learnable weight matrices, and both represent bias terms; Perform layer normalization processing on the third trajectory feature and the fourth trajectory feature, and output a final trajectory feature: ; Among them, represents the final trajectory feature; Repeatedly extract the trajectory data to obtain multiple final trajectory features, and use a multi-layer perceptron to construct a final feature sequence: ; Among them, represents the final feature sequence, represents a multi-layer perceptron.

6. The multimodal autonomous navigation method based on image and trajectory data according to claim 1, characterized in that, Input the target image features and the final feature sequence into the multi-modal fusion network. After the multi-modal fusion network maps the target image features and the final feature sequence, it outputs multi-modal hybrid features, specifically including: Input the target image features and the final feature sequence into the multi-modal fusion network. The multi-modal fusion network projects the target image features and the final feature sequence to obtain the target query feature, target key value, target feature value and corresponding spatial representation of the target image features: ; ; ; ; Among them, , and respectively represent the target query feature, target key, and target feature value of the fusion feature, represents the target image feature, represents the final feature sequence, , and all represent learnable linear transformation weights, represents the spatial representation, represents the attention mechanism, represents the lengths of the target image feature and the final feature sequence, represents the feature dimension in the real-time image, represents the normalization function; After performing a one-dimensional convolution operation on the spatial representation, splice it with the target image features to obtain multi-modal hybrid features: ; Among them, represents the multi-modal hybrid feature, represents concatenation in the channel dimension, represents performing a one-dimensional convolution operation.

7. The multimodal autonomous navigation method based on image and trajectory data according to claim 2, wherein Determine the current position of the target intelligent vehicle according to the multi-modal hybrid features, select the minimum-distance target point from all the potential target points according to the current position, and construct a control instruction according to the minimum-distance target point to control the movement of the target intelligent vehicle, specifically including: Input all the potential target points into the multi-modal autonomous navigation model. The multi-modal autonomous navigation model filters the minimum-distance target point from all the potential target points according to the multi-modal hybrid features, where the minimum-distance target point represents the potential target point most similar to the current image; Calculate the current position of the target intelligent vehicle in the current image, and calculate the linear velocity and angular velocity from the current position to the minimum-distance target point; Construct a control instruction according to the linear velocity and the angular velocity to control the target intelligent vehicle to move to the minimum-distance target point.

8. A multimodal autonomous navigation system based on image and trajectory data, characterized in that, The multi-modal autonomous navigation system based on image and trajectory data includes: An image acquisition module, configured to acquire multiple real-time images of a target intelligent vehicle, construct an observation image group and a target image group based on all the real-time images, and obtain multiple potential target points; A trajectory acquisition module, configured to acquire the yaw angle of the target intelligent vehicle, and construct the trajectory data of the target intelligent vehicle based on all the real-time images and the yaw angle; A feature extraction module, configured to input the observation image group and the target image group into a multi-modal navigation model, and an image feature extraction network in the multi-modal navigation model outputs observation image features and target image features; A feature sorting module, configured to preprocess the trajectory data and input it into a trajectory feature extraction network in the multi-modal navigation model, and the multi-modal navigation model performs multiple preliminary extractions on the preprocessed trajectory data and outputs a final feature sequence; A feature fusion module, configured to input the target image features and the final feature sequence into a multi-modal fusion network, and the multi-modal fusion network maps the target image features and the final feature sequence and outputs multi-modal hybrid features; A navigation module, configured to determine the current position of the target intelligent vehicle according to the multi-modal hybrid features, select a minimum-distance target point from all the potential target points according to the current position, and construct a control instruction according to the minimum-distance target point to control the movement of the target intelligent vehicle.

9. A terminal, characterized in that, The terminal includes: a memory, a processor, and a multi-modal autonomous navigation program based on image and trajectory data stored on the memory and executable on the processor. When the multi-modal autonomous navigation program based on image and trajectory data is executed by the processor, the steps of the multi-modal autonomous navigation method according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a multi-modal autonomous navigation program based on image and trajectory data. When the multi-modal autonomous navigation program based on image and trajectory data is executed by a processor, the steps of the multi-modal autonomous navigation method according to any one of claims 1-7 are implemented.

Citation Information

Cited By

  • Flow dividing and converging lane detection method and device

    CN121708339A