Road topological structure sensing method and network
By fusing trajectory distribution features and trajectory vector features with road image data, fusion features are generated to predict road topology, which solves the problem of poor prior information fusion effect in the prior art and improves the accuracy of road topology perception.
Patent Information
- Application Number
- CN202510119431.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to effectively integrate prior information into road image data, resulting in inaccurate results of road topology perception.
By obtaining the road image data and prior data of the target road section, including trajectory distribution features and trajectory vector features, feature extraction and fusion features are generated to predict the road topology.
It improves the accuracy of road topology perception, and provides more accurate road topology recognition by integrating rich prior information and image features.
Smart Images

Figure CN120047922A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of environment construction, and particularly relates to a method and network for perceiving road topology structure. Background Art
[0002] Road topology structure perception plays an important role in the perception system of autonomous driving, and it can help autonomous driving vehicles better understand the road scene. Previous road topology structure perception was mainly achieved based on road image data. In order to improve the perception effect of road topology structure, related technologies have proposed a method of fusing prior information into road image data for road topology structure perception. However, the current fusion method is difficult to effectively fuse prior information into road image data, resulting in an unsatisfactory fusion effect and inaccurate road topology structure perception results. Summary of the Invention
[0003] In a first aspect, an embodiment of this application provides a method for perceiving road topology structure, and the method includes:
[0004] Obtain road image data of a target road section and prior data of the target road section; the prior data includes the trajectory distribution feature corresponding to the target road section and the trajectory vector feature of the target road section; the trajectory distribution feature is used to characterize the distribution of trajectory points of historical trajectories passing through the target road section, and the trajectory vector feature is obtained by extracting features from the vectorized historical trajectories;
[0005] Extract features from the road image data to obtain image features;
[0006] Fuse the image features, the trajectory distribution feature, and the trajectory vector feature to obtain fused features;
[0007] Predict the road topology structure of the target road section based on the fused features.
[0008] In a second aspect, an embodiment of this application provides a road topology structure perception network, including:
[0009] An encoder, configured to obtain road image data of a target road section and prior data of the target road section; the prior data includes the trajectory distribution feature corresponding to the target road section and the trajectory vector feature of the target road section; the trajectory distribution feature is used to characterize the distribution of trajectory points of historical trajectories passing through the target road section, and the trajectory vector feature is obtained by extracting features from the vectorized historical trajectories; extract features from the road image data to obtain image features; fuse the image features, the trajectory distribution feature, and the trajectory vector feature to obtain fused features; and
[0010] A decoder for predicting the road topology of the target road section based on the fused features.
[0011] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any embodiment of the present application is implemented.
[0012] In a fourth aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described in any embodiment of the present application is implemented.
[0013] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any embodiment of the present application is implemented.
[0014] In the embodiments of the present application, historical trajectories are used as prior information. On the one hand, trajectory distribution features are obtained according to the distribution of trajectory points of the historical trajectories of the target road section. Since historical trajectories are constrained by the road layout, the trajectory distribution features can reflect the road layout and help to more accurately identify the actual topological structure of the road. On the other hand, trajectory vector features are obtained by feature extraction of the vectorized historical trajectories. The vectorized historical trajectories provide specific path information of the vehicle driving on the road, including information such as the driving trajectory and direction of the vehicle, which can help to more accurately identify and model the topological structure of the road, especially in areas such as complex intersections, road connections, and lane changes. Therefore, fusing the trajectory distribution features and trajectory vector features into the image features extracted from road image data can provide richer information for road topology perception, thereby improving the accuracy of the perception results.
[0015] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings here are incorporated into the specification and form a part of this application. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solutions of the present application.
[0017] Figure 1 is a flowchart of the road topology perception method according to an embodiment of the present application.
[0018] Figure 2 is a schematic structural diagram of the road topology perception network according to an embodiment of the present application.
[0019] Figure 3 is a schematic diagram of the alignment module according to an embodiment of the present application.
[0020] Figure 4 It is a block diagram of the road topology structure perception device according to an embodiment of the present application.
[0021] Figure 5 It is a schematic diagram of the computer device according to an embodiment of the present application. Detailed implementation manners
[0022] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0023] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items. In addition, the term "at least one" as used herein represents any one of a plurality or any combination of at least two of a plurality.
[0024] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0025] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application and make the above-mentioned objects, features and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the drawings.
[0026] The road topology refers to the connection relationships and layout patterns among roads in a road network. The road topology includes, but is not limited to, the network structure of roads (such as the connection mode of roads, the layout of intersections, the topological relationships between different roads, the length, width, and number of lanes of road segments) and / or the morphological characteristics of roads (such as the curvature, slope, traffic signs, markings, signal lights, and other traffic facilities) of roads. By perceiving the road topology, it is possible to understand and identify the layout, morphology of the road network, and the connection relationships between roads, thereby providing basic support for applications such as autonomous driving, intelligent transportation systems, and vehicle-to-everything (V2X). To improve the accuracy of road topology perception, related technologies have proposed a method of fusing prior information into road image data for road topology perception. However, the current fusion methods are difficult to effectively fuse prior information into road image data, resulting in an unsatisfactory fusion effect and inaccurate road topology perception results.
[0027] Based on this, the present application proposes a method for road topology perception, see Figure 1 , the method includes:
[0028] Step S11: Obtain road image data of a target road segment and prior data of the target road segment; the prior data includes the trajectory distribution feature corresponding to the target road segment and the trajectory vector feature of the target road segment; the trajectory distribution feature is used to characterize the distribution of trajectory points of historical trajectories passing through the target road segment, and the trajectory vector feature is obtained by performing feature extraction on the vectorized historical trajectories;
[0029] Step S12: Perform feature extraction on the road image data to obtain image features, and fuse the image features, the trajectory distribution feature, and the trajectory vector feature to obtain fused features;
[0030] Step S13: Predict the road topology of the target road segment based on the fused features.
[0031] In the embodiments of the present application, historical trajectories are used as prior information. On the one hand, the trajectory distribution feature obtained according to the distribution of trajectory points of the historical trajectories of the target road segment can reflect the road layout and help to more accurately identify the actual topological structure of the road; on the other hand, the trajectory vector feature obtained by performing feature extraction on the vectorized historical trajectories can help to more accurately identify and model the topological structure of the road. Therefore, fusing the trajectory distribution feature and the trajectory vector feature into the image features extracted from road image data can provide richer information for road topology perception, thereby improving the accuracy of perception results. The following is an example of the specific implementation manner of the embodiments of the present application. The method of the embodiments of the present application can be implemented through the Figure 2 shown road topology perception network.
[0032] In step S11, the road image data of the target road section can be collected by sensors (such as cameras, lidar, infrared sensors, etc.) mounted on mobile platforms such as vehicles or drones. Taking the collection of road image data by a camera mounted on a vehicle as an example, multiple cameras can be deployed on the vehicle to cover different directions around the vehicle. The above-mentioned multiple cameras may include, but are not limited to:
[0033] Front-view camera: Installed at the front of the vehicle, providing image data of the road ahead;
[0034] Rear-view camera: Installed at the rear of the vehicle, providing image data of the road behind, commonly used for reversing and rear monitoring;
[0035] Side-view camera: Installed on both sides of the vehicle, covering the left and right fields of view respectively;
[0036] Top-view camera: Used to capture the panoramic image above the vehicle, assisting in achieving a more comprehensive perception.
[0037] To achieve a more intuitive perception of the surrounding environment, especially for path planning and obstacle detection, the images collected by the camera can be processed to be transformed into a bird's-eye view.
[0038] In the embodiments of the present application, the directly collected images by the camera can be used as the road image data of the target road section, or the transformed bird's-eye view can be used as the road image data of the target road section.
[0039] In addition to obtaining the road image data of the target road section, prior data of the target road section can also be obtained. Among them, the prior data of the target road section may include the trajectory distribution characteristics corresponding to the target road section and the trajectory vector characteristics of the target road section. The prior data can be obtained offline based on historical trajectories. The process of obtaining prior data offline is illustrated by way of example below.
[0040] First, historical trajectories can be obtained from at least one data source, and the above-mentioned at least one data source includes, but is not limited to, at least one of the following: GPS, mobile communication base stations, traffic monitoring cameras, road sensors, Internet data sources, drone data sources, Internet of Things (IoT) devices, geographic information systems (GIS), track data, map data, etc.
[0041] Since the initially obtained historical trajectories are disorderly, the initially obtained historical trajectories can be screened according to the length of the historical trajectories to ensure that only the most relevant data segments are retained for further analysis. Specifically, historical trajectories with a length greater than a preset length can be screened out. Then, the screened historical trajectories can also be filtered (such as average filtering) to reduce random noise in the data, so as to more clearly reveal the potential trend of the historical trajectories.
[0042] Then, based on the filtered historical trajectories, the trajectory distribution features corresponding to the target road segment and the trajectory vector features of the target road segment can be obtained.
[0043] The trajectory distribution features are used to characterize the distribution of trajectory points of the historical trajectories passing through the target road segment. In some embodiments, the trajectory distribution features include the position distribution features of the historical trajectory points passing through the target road segment and / or the direction distribution features of the historical trajectory points passing through the target road segment. The position distribution features are used to characterize the distribution of the spatial positions of the historical trajectory points on the target road segment, and specifically include at least one of the following information:
[0044] The spatial density of the trajectory points, that is, on the target road segment, the degree of concentration of the trajectory points in space. For example, on some road segments, vehicles may frequently pass through certain areas, resulting in a relatively dense distribution of trajectory points in these areas. By calculating the density of the trajectory points, the hot spots of object travel can be reflected.
[0045] The distribution range, which can reflect the coverage range of the trajectory points in space. By statistically analyzing the coordinate range of the trajectory points, the activity range of the object on the target road segment can be understood. For example, some objects may be concentrated in only a small part of the road segment, while others may cover a wider area.
[0046] Clustering or dispersion: By analyzing the spatial distribution pattern of the trajectory points, it is possible to determine whether there are certain specific route preferences or clustering areas. If the trajectory points show a certain concentrated distribution, it indicates that the object may tend to travel in these areas or along these paths.
[0047] The direction distribution features refer to the distribution of the driving directions of the objects when the historical trajectory points pass through the target road segment, and specifically include at least one of the following information:
[0048] The angular distribution of the driving directions: Each trajectory point has a driving direction (that is, the direction from the previous point to the current point), and these directions may show certain patterns on the target road segment. For example, on some roads, vehicles mostly travel in the same direction most of the time, forming a single direction distribution; while near curves or intersections, vehicles may have various direction changes.
[0049] The turning frequency and angle: If an object frequently turns on a certain road segment, the angular distribution of the turns can also be an important aspect of the direction distribution features. For example, vehicles may frequently turn and travel along a certain angle, while on straight road segments, they may maintain a consistent driving direction.
[0050] Directional Clustering: If, on a certain section of road, the direction changes of the trajectory points are concentrated in a certain direction, it indicates that the object is traveling in this direction most of the time. By analyzing the degree of concentration of these directions, important information about traffic flow, driving routes, etc. can be obtained.
[0051] In summary, the position distribution characteristics and direction distribution characteristics of the trajectory points can accumulate information about trajectory density and direction, thereby providing the model with a comprehensive understanding of the trajectory data distribution and its flow trend.
[0052] In some embodiments, the trajectory data can be converted into a rasterized form. Specifically, the target road section can be gridified to obtain multiple grids, and multiple historical trajectories of the target road section are obtained. Each historical trajectory includes multiple historical trajectory points. Determine the number of historical trajectory points passing through each grid and the direction of each historical trajectory point when passing through the corresponding grid. Based on the number of historical trajectory points passing through each grid, obtain the position distribution characteristics of the historical trajectory points passing through the target road section, and based on the direction of each historical trajectory point when passing through the corresponding grid, obtain the direction distribution characteristics of the historical trajectory points passing through the target road section.
[0053] Let the trajectory point set be defined as:
[0054] trajectory={P (1) ,…,P (m)}
[0055] where p i =(x i ,y i ) represents the i-th trajectory point in the historical trajectory, and x and y represent the abscissa and ordinate of the trajectory point respectively.
[0056] Define a grid that divides the space into M×M grid cells, and the side length of each grid cell is Δx. For each pair of consecutive points P (i) and P (i+1) in the trajectory, calculate their positions in the grid coordinate system:
[0057] grid_x = |(x - x_min) / Δx|, grid_y = |(y - y_min) / Δx|
[0058] where x min and y min are the minimum x and minimum y values of the defined grid area respectively. For each pair of consecutive trajectory points, the angle θ between them can be calculated to represent the direction of the trajectory in this section:
[0059] θ i,i+1 = arctan2(y i+1 -yi , x i+1 -x i )
[0060] where θ i,i+1 represents the direction of the line segment formed by the i-th trajectory point and the (i + 1)-th trajectory point.
[0061] For each grid, the number of trajectory points passing through the grid (density) and direction information (angle) can be accumulated. After obtaining the density and direction information of each grid point, rasterization processing can be performed to generate the corresponding density heat map and direction heat map. Among them, the density heat map is the above-mentioned position distribution feature, and the direction density heat map is the above-mentioned direction distribution feature.
[0062] The trajectory vector feature can be obtained by extracting features from the vectorized historical trajectory. The trajectory vector feature helps to describe the movement law of an object (such as a vehicle, a pedestrian, etc.) in space and time, and provides a basis for subsequent tasks such as trajectory analysis, behavior prediction, and path planning. The vectorization process can convert the trajectory data into a more structured and easy-to-analyze form. In some embodiments, in the face of a large amount of historical trajectories, the present application takes the obtained multiple historical trajectories as candidate historical trajectories, screens out several historical trajectories from the multiple candidate historical trajectories, and performs vectorization processing on the screened several historical trajectories. The screened historical trajectories can be considered as representative trajectory samples, and the neural network can effectively learn its potential patterns from these historical trajectories. In order to select the most representative trajectory samples, the present application adopts two strategies respectively: the clustering algorithm (such as K-means clustering) and the farthest point sampling (FPS) method.
[0063] In the example of selecting the most representative trajectory samples by using the clustering algorithm, the multiple historical trajectories can be clustered to obtain multiple clusters, and the historical trajectories corresponding to the cluster centers of each cluster can be screened out. The cluster operation can be performed through several iterations. When the preset maximum number of iterations is reached, or the relative tolerance of the change in the cluster center between two consecutive iterations is less than 0.0001, the iteration stops. Through this process, the K-means algorithm can effectively cluster similar trajectories together, and the cluster centers represent the common trends or patterns of these trajectories. These cluster centers can be used for further analysis, such as trajectory visualization or pattern recognition.
[0064] In the example of selecting the most representative trajectory samples using the FPS algorithm, several historical trajectories can be iteratively filtered from the multiple historical trajectories. Among them, in the first iterative filtering, a historical trajectory is randomly selected from the multiple historical trajectories. The process of the i-th iterative filtering is as follows: Obtain the set of historical trajectories filtered in the (i - 1)-th iteration, select the historical trajectory with the farthest distance from the set of historical trajectories obtained in the i-th iteration from the unfiltered historical trajectories, add the historical trajectory with the farthest distance to the set of historical trajectories. If the total number of historical trajectories in the set of historical trajectories reaches the preset quantity threshold, stop the iteration; otherwise, continue the next iteration.
[0065] For example, assume the total number of historical trajectories is N. In the first iteration, a historical trajectory can be randomly selected as a sampling point and added to the set of historical trajectories. In the second iteration, the historical trajectory with the farthest distance from the historical trajectories in the set of historical trajectories can be selected from the remaining N - 1 historical trajectories and also added to the set of historical trajectories. Among them, the above distance can be, for example, the Frechet distance. In the third iteration, the historical trajectory with the farthest distance from the two historical trajectories in the set of historical trajectories can be selected from the remaining N - 2 historical trajectories. And so on, until the total number of historical trajectories in the set of historical trajectories reaches the preset quantity threshold.
[0066] The FPS algorithm realizes uniform sampling of trajectory data by iteratively selecting the points with the farthest distance from the selected point set. This method ensures the uniform distribution of sampling points in the entire data space, which helps to improve the generalization ability of the model. Through FPS, we uniformly select key samples from the entire dataset, thereby increasing the diversity of the training dataset.
[0067] In step S12, the road image data can be first subjected to feature extraction to obtain image features, and then the image features, the trajectory distribution features, and the trajectory vector features can be fused to obtain fused features. The above process can be implemented by the encoder in the road topology structure perception network. Through feature fusion, the information contained in the image features can be enhanced. Since the information dimension of the vectorized trajectory information is smaller and the calculation of the cross-attention mechanism is faster, the image features and the trajectory vector features can be subjected to attention processing to obtain intermediate features, and the trajectory distribution features and the intermediate features can be fused to obtain fused features. Specifically, when performing attention processing, the image features can be used as the query (Query), and the trajectory vector features can be used as the key (Key) and value (Value).
[0068] The integration of prior information can effectively improve the perception ability of the vehicle. However, due to the inaccuracy of positioning, there may be an offset between the prior information and the online observations (i.e., road image data). Therefore, before fusing the trajectory distribution feature and the intermediate feature, the offset between the trajectory distribution feature and the intermediate feature can be determined, and the trajectory distribution feature and the intermediate feature can be aligned based on the offset. The above alignment operation can be implemented by an alignment module in the road topology perception network. The principle of the alignment module is as Figure 3 shown.
[0069] Specifically, the trajectory distribution feature and the intermediate feature can be concatenated to obtain a concatenated feature, and the concatenated feature can be convolved to obtain the offset between the trajectory distribution feature and the intermediate feature. After determining the offset between the trajectory distribution feature and the intermediate feature, the intermediate feature can be interpolated by bilinear interpolation, so as to achieve the alignment between the trajectory distribution feature and the intermediate feature.
[0070] The integration of prior information can effectively improve the perception ability of the vehicle. However, if the prior information is incorrect, it may mislead the result of road topology perception. Although the historical trajectory has been preprocessed, it may contain some incorrect data, such as the driver taking the wrong road, resulting in incorrect prior information. Therefore, the model needs to distinguish the accuracy of the prior information. To address this issue, the network needs to learn in what situations to trust which part of the information.
[0071] Specifically, the confidence of the trajectory distribution feature and the confidence of the intermediate feature can be determined. The trajectory distribution feature is weighted based on the confidence of the trajectory distribution feature to obtain a weighted trajectory distribution feature, and the intermediate feature is weighted based on the confidence of the intermediate feature to obtain a weighted intermediate feature. The weighted trajectory distribution feature and the weighted intermediate feature are fused. The fused feature can be denoted as:
[0072]
[0073] where, represents the intermediate feature at position (i, j), represents the trajectory distribution feature at position (i, j), represents the confidence of the intermediate feature at position (i, j) in the l-th channel, represents the confidence of the trajectory distribution feature at position (i, j) in the l-th channel, represents the fused feature at position (i, j) in the l-th channel. The above confidences and can be obtained by the network's adaptive learning and satisfy the following conditions:
[0074]
[0075] Regarding the acquisition of the normalized confidence, a 1x1 convolutional kernel can be used to act on the original feature map (intermediate feature or trajectory distribution feature) to respectively obtain Then, by comparing with the scalar The softmax function is used to calculate the final confidence matrix:
[0076]
[0077] The alignment of the prior data and the road image data can be achieved through alignment and confidence weighting. However, supervision needs to be introduced for the aligned quantities to ensure correct learning. Since the information of the bird's-eye view segmentation is learned from these two types of features, this application introduces map segmentation as a supervision signal to enhance the alignment effect. Specifically, the prior data can also include a map segmentation mask. After passing through the encoder, image features can be obtained, which are features similar to the picture type and are consistent with the expression way of the extracted map segmentation mask. Therefore, the two types of features can be made complementary through an addition operation to improve the richness and comprehensiveness of the feature expression. At the same time, this method is simple and effective, can maintain the original information of the two types of features without introducing complex calculations, and reduces the computational amount. When fusing the image features, the trajectory distribution features, and the trajectory vector features, the image features fused with the map segmentation mask can be fused with the trajectory distribution features and the trajectory vector features. The specific fusion method is as described in the foregoing embodiments and will not be elaborated here.
[0078] After obtaining the fused features, in step S13, the decoder in the road topology perception network can decode the fused features to obtain the road topology of the target road segment. The decoder side can use a deformable attention mechanism similar to decode the fused features. The core idea of the deformable attention mechanism is to adaptively select important positions for calculation instead of performing global interaction on all positions. This means that it makes the calculation more efficient by dynamically adjusting the attention range of the query, key, and value, and only focuses on the local regions relevant to the current task. In the deformable attention mechanism of the related art, multiple initial queries (also called initial reference points) are usually randomly selected. However, this application considers that if the positions of the initial queries can be set near the feature positions of the prior data, it can provide a better initial solution for the selection of the sampled features. Therefore, at least some of the multiple initial queries can be obtained based on the trajectory vector features. Specifically, multiple queries can be obtained; the multiple queries include queries obtained based on the trajectory vector features and randomly generated queries, the correlations between the multiple queries and multiple feature blocks in the fused features are respectively determined, the corresponding feature blocks are weighted based on the correlations to obtain weighted features, and the road topology of the target road segment is predicted based on the weighted features. In some embodiments, the multiple obtained queries are initial queries, and before respectively determining the correlations between the multiple queries and multiple feature blocks in the fused features, the multiple obtained queries can be optimized.
[0079] For example, for the trajectory vector feature T ∈ N T × C T , where N T represents the number of trajectory vector features, and C T represents the dimension after flattening the trajectory vector features. For example, when there are 10 trajectory point coordinates in the trajectory vector feature, since each coordinate includes two coordinate values of x and y, the dimension C T after flattening is 20. To align with the image features, the operation of supplementing 0 can be performed on the flattened trajectory vector features. The position of the initial reference point can be directly initialized to the randomly obtained coordinate values. Assume that the total number of the multiple obtained queries is N, then N T queries can be obtained based on the trajectory vector features, and the remaining N - N T queries are randomly obtained. Then, for the above N queries, iterative optimization is first performed, and then the correlations between the multiple optimized queries and multiple feature blocks in the fused features are respectively determined, the corresponding feature blocks are weighted based on the correlations to obtain weighted features, and the road topology of the target road segment is predicted based on the weighted features.
[0080] Such as Figure 4As shown, the present application also provides a road topology structure perception device. Refer to Figure 4 , the device includes:
[0081] An acquisition module 101, configured to acquire road image data of a target road section and prior data of the target road section; the prior data includes a trajectory distribution feature corresponding to the target road section and a trajectory vector feature of the target road section; the trajectory distribution feature is used to characterize the distribution of trajectory points of historical trajectories passing through the target road section, and the trajectory vector feature is obtained by extracting features from the vectorized historical trajectories;
[0082] A feature extraction module 102, configured to extract features from the road image data to obtain image features, and fuse the image features, the trajectory distribution feature, and the trajectory vector feature to obtain a fused feature;
[0083] A prediction module 103, configured to predict the road topology structure of the target road section based on the fused feature.
[0084] An embodiment of the present application also provides a computer device, which at least includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the program, it implements the method described in any one of the foregoing embodiments.
[0085] Figure 5 FIG. shows a more specific schematic diagram of the hardware structure of a computer device provided by an embodiment of the present application. The device may include: a processor 201, a memory 202, an input / output interface 203, a communication interface 204, and a bus 205. Wherein, the processor 201, the memory 202, the input / output interface 203, and the communication interface 204 are communicatively connected to each other inside the device through the bus 205.
[0086] The processor 201 may be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present application. The processor 201 may further include a graphics card, and the graphics card may be an Nvidia titan X graphics card or a 1080Ti graphics card, etc.
[0087] The memory 202 may be implemented in the form of a Read Only Memory (ROM), a Random Access Memory (RAM), a static storage device, a dynamic storage device, etc. The memory 202 may store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of the present application through tools or firmware, the relevant program codes are stored in the memory 202 and called and executed by the processor 201.
[0088] The input / output interface 203 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input devices may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output devices may include a display, a speaker, a vibrator, an indicator light, etc.
[0089] The communication interface 204 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. The communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0090] The bus 205 includes a path for transmitting information between various components of the device (such as the processor 201, the memory 202, the input / output interface 203, and the communication interface 204).
[0091] It should be noted that although only the processor 201, the memory 202, the input / output interface 203, the communication interface 204, and the bus 205 are shown in the above device, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary for implementing the solutions of the embodiments of the present application and does not necessarily include all the components shown in the figure.
[0092] The embodiments of the present application provide a computer program product, including a computer program, which when executed by a processor implements the method described in any embodiment of the present application.
[0093] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in any of the foregoing embodiments.
[0094] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computer device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0095] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiments. The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated. When implementing the solution of the embodiments of this application, the functions of the modules can be implemented in the same or multiple tools and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0096] The above is only the specific implementation manner of the embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the embodiments of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of this application.
Claims
1. A road topology perception method, characterized in that: The method comprises: Acquire road image data of a target road section and prior data of the target road section; the prior data includes a trajectory distribution feature corresponding to the target road section and a trajectory vector feature of the target road section; the trajectory distribution feature is used to characterize the trajectory point distribution of a historical trajectory passing through the target road section, and the trajectory vector feature is obtained by extracting features from the vectorized historical trajectory; Extracting features from the road image data to obtain image features; fusing the image feature, the trajectory distribution feature and the trajectory vector feature to obtain a fused feature; The road topology of the target road segment is predicted based on the fused features.
2. The method according to claim 1, characterized in that: The trajectory distribution features include the position distribution features and direction distribution features of the historical trajectory points passing through the target road section; the trajectory distribution features corresponding to the target road section are obtained based on the following method: Performing grid processing on the target road section to obtain a plurality of grids; Acquire multiple historical tracks of the target road section, each historical track including multiple historical track points; Determine the number of historical trajectory points passing through each grid and the direction of each historical trajectory point when passing through the corresponding grid; The position distribution characteristics of the historical trajectory points passing through the target section are obtained based on the number of historical trajectory points passing through each grid, and the direction distribution characteristics of the historical trajectory points passing through the target section are obtained based on the direction of each historical trajectory point when passing through the corresponding grid.
3. The method according to claim 1, characterized in that Before extracting features from the vectorized historical trajectory to obtain the trajectory vector features, the method further includes: Obtain multiple candidate historical trajectories; Selecting a plurality of historical tracks from the plurality of candidate historical tracks; Vectorization is performed on several selected historical trajectories.
4. The method according to claim 3, characterized in that The selecting a plurality of historical tracks from the plurality of candidate historical tracks comprises: Clustering the multiple historical trajectories to obtain multiple clusters; Filter out the historical trajectories corresponding to the cluster centers of each cluster; or Iteratively select a plurality of historical tracks from the plurality of historical tracks, wherein, in the first iterative selection, a historical track is randomly selected from the plurality of historical tracks, and the i-th iterative selection process is as follows: Get the historical trajectory set filtered out for the i-1th time; Filter out the historical trajectory farthest from the historical trajectory set obtained in the (i-1)th iteration from the historical trajectories that have not been filtered, and add the historical trajectory farthest to the historical trajectory set; If the total number of historical trajectories in the historical trajectory set reaches a preset number threshold, the iteration is stopped; otherwise, the next iteration is continued.
5. The method according to claim 1, characterized in that The fusing the image feature, the trajectory distribution feature and the trajectory vector feature to obtain a fused feature includes: Performing attention processing on the image features and the trajectory vector features to obtain intermediate features; The trajectory distribution feature and the intermediate feature are fused to obtain a fused feature.
6. The method according to claim 5, characterized in that Before fusing the trajectory distribution feature and the intermediate feature, the method further includes: Splicing the trajectory distribution feature and the intermediate feature to obtain a spliced feature; Performing convolution processing on the splicing feature to obtain an offset between the trajectory distribution feature and the intermediate feature; The trajectory distribution feature and the intermediate feature are aligned based on the offset.
7. The method according to claim 5, characterized in that The fusing the trajectory distribution feature and the intermediate feature includes: Determining the confidence of the trajectory distribution feature and the confidence of the intermediate feature; weighting the trajectory distribution feature based on the confidence of the trajectory distribution feature to obtain a weighted trajectory distribution feature; weighting the intermediate features based on the confidence of the intermediate features to obtain weighted intermediate features; The weighted trajectory distribution feature and the weighted intermediate feature are fused.
8. The method according to claim 1, characterized in that The predicting the road topology structure of the target road section based on the fusion feature includes: Acquire multiple queries; the multiple queries include queries acquired based on the trajectory vector features and randomly generated queries; respectively determining correlations between the plurality of queries and the plurality of feature blocks in the fused features; Weighting the corresponding feature blocks based on the correlation to obtain weighted features; The road topology of the target road segment is predicted based on the weighted features.
9. A road topology perception network, characterized in that: include: An encoder is used to obtain road image data of a target road section and prior data of the target road section; the prior data includes a trajectory distribution feature corresponding to the target road section and a trajectory vector feature of the target road section; the trajectory distribution feature is used to characterize the trajectory point distribution of a historical trajectory passing through the target road section, and the trajectory vector feature is obtained by performing feature extraction on the vectorized historical trajectory; and feature extraction is performed on the road image data to obtain image features; fusing the image feature, the trajectory distribution feature and the trajectory vector feature to obtain a fused feature; as well as A decoder is used to predict the road topology of the target road segment based on the fused features.
10. A computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method according to any one of claims 1 to 8.
Citation Information
Cited By
Multi-source domain fused spatial-temporal trajectory basic model construction method and application thereof
CN120893163A