Combined vehicle testing methods, electronic equipment, storage media and program products
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-08-14
AI Technical Summary
与普通乘用车相比,超长车具有尺度大、结构复杂、形态多变(如挂车、半挂车、多节车)等特点,在自动驾驶感知系统中存在以下突出挑战:目标尺度跨度大,检测不完整问题突出
[0009]依据本申请实施例,首先获取自动驾驶车辆视野范围内的各交通参与者的鸟瞰图特征,鸟瞰图特征根据多组时序相邻的目标帧融合得到,从而可以保留更完整的组合车辆特征,目标帧基于多模态特征融合得到,从而可以综合各种模态特征的优点,保证更好的环境适应能力,保证组合车辆特征的检测精度,对鸟瞰图特征分别进行不同采样倍数的下采样,得到多组下采样结果可以分别用于不同尺寸的交通参与者的检测,将多组采样结果中下采样倍数最高的目标采样结果,用于检测组合车辆,由于下采样倍数高的目标采样结果可以提供更大的感受野,从而可以保证识别出更大目标,即组合车辆的全貌,避免组合车辆被截断,保证车辆检测结果中组合车辆的完整性。
Smart Images

Figure CN122574809A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and in particular to a combined vehicle detection method, electronic device, storage medium, and program product. Background Technology
[0002] With the continuous development of autonomous driving technology, vehicles are facing higher demands on their perception capabilities of various traffic participants in the surrounding environment. Especially in long-haul logistics, port transportation, and highway scenarios, extra-long vehicles (≥25m in length) are important traffic participants, and their accurate detection and state estimation are crucial for driving safety. Compared with ordinary passenger cars, extra-long vehicles are characterized by their large scale, complex structure, and varied forms (such as trailers, semi-trailers, and multi-section vehicles), posing the following prominent challenges to autonomous driving perception systems: The target scale spans a large range, leading to significant incomplete detection issues. Extra-long vehicles often exhibit a cross-regional distribution in the perception field of view; single-viewpoint or local perception easily results in truncated detection boxes or identification of only local structures (such as the front or rear of the vehicle), making it difficult to form a complete target representation. Currently, no better solution has been proposed to address this issue. Summary of the Invention
[0003] This application provides a combined vehicle detection method, electronic device, storage medium, and program product to solve one or more of the above-mentioned technical problems.
[0004] In a first aspect, embodiments of this application provide a combined vehicle detection method, comprising: acquiring bird's-eye view features within the field of vision of an autonomous vehicle; the bird's-eye view features are obtained by fusing multiple sets of temporally adjacent target frames; the target frames are obtained by fusing multimodal features; downsampling the bird's-eye view features by different sampling multiples to obtain multiple sets of downsampling results; detecting combined vehicles based on target sampling results in the multiple sets of downsampling results to obtain combined vehicle detection results; wherein the sampling multiple of the target sampling results is higher than the sampling multiple of the sampling results other than the target sampling results in the multiple sets of downsampling results.
[0005] Secondly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements the method described in any of the above-mentioned embodiments.
[0006] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method described in any of the above-mentioned embodiments.
[0007] Fourthly, embodiments of this application provide a computer program product, wherein the computer program product includes a computer program, which, when executed by a processor, implements the method described in any of the above-mentioned embodiments.
[0008] Compared with related technologies, this application has the following advantages:
[0009] According to the embodiments of this application, firstly, bird's-eye view features of each traffic participant within the field of vision of the autonomous vehicle are obtained. The bird's-eye view features are obtained by fusing multiple sets of temporally adjacent target frames, thereby preserving more complete combined vehicle features. The target frames are obtained by fusing multimodal features, thereby combining the advantages of various modal features, ensuring better environmental adaptability, and ensuring the detection accuracy of combined vehicle features. The bird's-eye view features are downsampled by different sampling factors to obtain multiple sets of downsampled results, which can be used to detect traffic participants of different sizes. The target sampling result with the highest downsampling factor among the multiple sampling results is used to detect combined vehicles. Since the target sampling result with a high downsampling factor can provide a larger receptive field, it can ensure the identification of larger targets, i.e., the full picture of combined vehicles, avoiding the combined vehicles being truncated and ensuring the integrity of combined vehicles in the vehicle detection results.
[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application, it can be implemented according to the contents of the specification. In order to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description
[0011] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments according to this application and should not be construed as limiting the scope of this application.
[0012] Figure 1 A flowchart of a combined vehicle detection method provided in an embodiment of this application is shown;
[0013] Figure 2 A schematic diagram of the components of a combined vehicle provided in an embodiment of this application is shown;
[0014] Figure 3 A schematic diagram of a visible frame corresponding to a combined vehicle provided in an embodiment of this application is shown;
[0015] Figure 4 A schematic diagram of a vehicle and a transported object provided in an embodiment of this application is shown;
[0016] Figure 5 A visible frame diagram of a transport object provided in an embodiment of this application is shown;
[0017] Figure 6 This illustration shows a schematic diagram of a cross-modal fusion module structure provided in an embodiment of this application;
[0018] Figure 7 This illustration shows a schematic diagram of the implementation steps of a combined vehicle detection method provided in an embodiment of this application;
[0019] Figure 8 This paper shows a structural block diagram of a combined vehicle detection device provided in an embodiment of this application;
[0020] Figure 9 A block diagram of an electronic device used to implement embodiments of this application is shown. Detailed Implementation
[0021] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the concept or scope of this application. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.
[0022] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies of the embodiments of this application are described below. The following relevant technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and all of them fall within the protection scope of the embodiments of this application.
[0023] In the perception phase of assisted driving and autonomous driving systems for passenger cars and commercial vehicles (trucks), existing vehicle detection solutions cannot guarantee that the detection results of combined vehicles will show structural integrity, accurate dimensions and positions, and consistent motion of all components of the combined vehicle.
[0024] The first existing detection scheme utilizes millimeter-wave radar. Millimeter-wave radar estimates the distance, angle, and velocity of targets by emitting millimeter-wave electromagnetic signals and receiving their echoes. With the development of 4D millimeter-wave radar, it can provide a certain degree of point cloud information in the spatial dimension and has strong environmental adaptability. However, it still has significant shortcomings in the detection of ultra-long vehicles (length ≥ 25m). Limited resolution and insufficient structural representation: Although 4D millimeter-wave radar can provide sparse point cloud information, the point cloud density is low, making it difficult to depict the complete structure of ultra-long vehicles, especially the connection between the tractor and trailer. Insufficient target modeling capability: Existing radar schemes typically use a holistic target modeling approach, making it difficult to distinguish the multi-segment structure of ultra-long vehicles (such as the tractor unit and trailer), and failing to accurately describe their true physical form. Limited detection accuracy and recall: In long-range or complex multi-target scenarios, radar is prone to target aliasing or missed detections, affecting overall detection performance.
[0025] The existing second approach utilizes LiDAR and multi-sensor fusion for detection. LiDAR acquires high-precision 3D point cloud information through high-frequency laser scanning. Combined with multi-sensor fusion technology using cameras, radar, and other sensors, it can significantly improve target detection accuracy and is one of the mainstream solutions for current autonomous driving. However, this approach still has the following problems in detecting ultra-long vehicles: Insufficient structural modeling: Traditional 3D detection methods usually treat targets as rigid bodies (single 3D bounding boxes). For ultra-long vehicles, which have "multi-segment structures and hinged relationships," accurate modeling is difficult, easily leading to detection box deviation or instability. High computational complexity: Multi-sensor fusion involves processing a large amount of point cloud and image data, especially when performing high-resolution modeling in BEV space, resulting in high computational overhead and difficulty meeting real-time requirements. Insufficient multi-scale adaptability: Existing methods perform detection on BEV features at a uniform scale, lacking targeted modeling for targets of different sizes (small targets and ultra-large targets), leading to decreased detection performance for extremely large targets such as ultra-long vehicles. High cost: High-beam LiDAR and its supporting computing platform are expensive, which is not conducive to large-scale application.
[0026] The third existing solution employs a pure vision-based approach for detection. This approach relies on cameras to acquire environmental images and uses deep learning models (such as the YOLO series and Transformer) for target detection and recognition. However, this solution has the following limitations in detecting extra-long vehicles: Limited 3D perception capability: Monocular or multi-view vision suffers from depth estimation errors, especially in long-distance scenes, making it difficult to accurately determine the true size and location of the extra-long vehicle. Sensitivity to environment: In complex environments such as low light, rain, and fog, visual features degrade significantly, affecting detection stability. Insufficient spatial structure modeling: Visual methods typically focus on 2D or pseudo-3D representations, which can easily lead to incomplete detection or inaccurate boundary conditions for large-scale targets like extra-long vehicles that span multiple regions.
[0027] Furthermore, autonomous driving systems are extremely sensitive to the latency of the perception module, demanding high real-time performance and computational efficiency, especially in high-speed scenarios where the detection latency of extremely long vehicles directly impacts decision-making and control safety. Therefore, while improving detection accuracy, ensuring the model's efficient inference capabilities is also a critical issue.
[0028] Therefore, a high-precision perception method for ultra-long vehicles is urgently needed. This application provides a combined vehicle detection method, electronic device, storage medium, and program product. This method utilizes multi-source information from cameras, LiDAR, and 4D millimeter-wave radar, incorporating precise structural modeling of the vehicle's front end / trailer and multi-scale perception model design to achieve complete modeling and precise localization of ultra-long vehicles. The method aims to: improve the detection integrity of the overall structure of ultra-long vehicles, enhance the modeling capability for complex postures and multi-section structures, improve detection robustness in long-distance, occluded, and harsh environments, and achieve high-precision 3D target detection while ensuring real-time performance. This application overcomes the inaccuracy of traditional single-frame modeling by decomposing the various parts of the combined vehicle, thereby achieving multi-target joint detection. It effectively overcomes the limitations of existing technologies in ultra-long vehicle perception, improving the safety and reliability of autonomous driving systems in complex traffic scenarios.
[0029] The technical solution of this application and how it solves the aforementioned technical problems are described in detail below with specific embodiments. The listed specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0030] This application provides a combined vehicle detection method. For example... Figure 1 The diagram shown is a flowchart of a combined vehicle detection method according to an embodiment of this application. The method may include:
[0031] Step S101: Obtain bird's-eye view features within the field of vision of the autonomous vehicle; the bird's-eye view features are obtained by fusing multiple sets of temporally adjacent target frames; the target frames are obtained by fusing multimodal features.
[0032] In this possible implementation, the combined vehicle is characterized by its large size, complex structure, and varied form. In one possible implementation, the combined vehicle is a transport vehicle where the space occupied by the transported object exceeds the space occupied by the vehicle itself, resulting in a large overall space occupation for the combined vehicle. In another possible implementation, the combined vehicle comprises multiple structural sections articulated together; for example, it can be, but is not limited to, a trailer, semi-trailer, or multi-section vehicle. The embodiments of this application can be applied to autonomous vehicles. Autonomous vehicles and combined vehicles are traffic participants in the same time and space. Combined vehicles are traffic participants within the field of vision of autonomous vehicles.
[0033] Bird's-eye view features refer to the spatial transformation of environmental information perceived by autonomous vehicles to generate a two-dimensional planar feature map resembling an aerial view. This feature comprehensively reflects the layout of the vehicle's surrounding environment, aiding in path planning and obstacle detection. The generation process of bird's-eye view features is based on the fusion of multiple sets of temporally adjacent target frames. Temporally adjacent means consecutive frames that are close in time; these frames are correlated, and fusion can compensate for missing information from a single frame's perspective, improving the stability and accuracy of spatial structure recognition.
[0034] Considering that in real-world road environments, extra-long vehicles are often at a distance or partially obscured, relying on a single sensor (such as a camera or LiDAR) is prone to missed or false detections. Figure 4 The vehicle body is obscured by the fan blades. In this embodiment, the target frame is obtained based on multimodal feature fusion. Multimodal features refer to data from different sensors (such as cameras, radar, and lidar). Fusion of these features can combine visual texture information with depth and distance information, so that the generated target frame has both high-resolution image details and accurate distance and shape parameters, and can also avoid missed detections.
[0035] In this possible implementation, bird's-eye view features are used to present the vehicle's surrounding environment completely and intuitively in two dimensions, enabling more efficient path planning and risk assessment. Temporally adjacent frame fusion solves the problem of insufficient information in a single frame, enhancing the continuity and stability of environmental perception in dynamic scenes. Multimodal feature fusion improves the perception accuracy of the target frame, taking into account both visual details and spatial depth information, adapting to different lighting, weather, and complex traffic conditions.
[0036] Step S102: The bird's-eye view features are downsampled by different sampling multiples to obtain multiple sets of downsampling results.
[0037] In this possible implementation, the bird's-eye view features are downsampled at different sampling factors to obtain multiple sets of downsampled results, thus constructing a multi-scale feature map. Downsampling refers to reducing the spatial resolution of the feature map by a certain proportion, thereby providing a larger receptive field for object detection. By setting different sampling factors, features of different scales can be generated. Multi-scale features help the system simultaneously focus on close-range details and long-range overall layout, optimizing the receptive field required for combined vehicle recognition and ensuring the integrity of combined vehicle recognition.
[0038] Step S103: Detect the combined vehicle based on the target sampling result in the multiple sets of downsampling results, and obtain the combined vehicle detection result; the sampling multiple of the target sampling result is higher than the sampling multiple of the other sampling results in the multiple sets of downsampling results.
[0039] In this possible implementation, a target sampling result is selected from multiple sets of downsampling results based on the sampling factor used in the downsampling operation. This target sampling result has a higher downsampling factor than the other downsampling results in the multiple sets of downsampling results; that is, the target sampling result with the largest receptive field is used for the detection of the combined vehicle. In this embodiment, the target sampling result is input into a pre-trained large-scale object detection model, which is then used to generate vehicle detection results. Because the target sampling result provides a larger receptive field, it can be ensured that the combined vehicle detection results can include the complete picture of the combined vehicle without truncation.
[0040] This application provides a method for detecting combined vehicles. According to this embodiment, firstly, bird's-eye view features of each traffic participant within the field of vision of an autonomous vehicle are obtained. These bird's-eye view features are obtained by fusing multiple sets of temporally adjacent target frames, thus preserving more complete combined vehicle features. The target frames are obtained by fusing multimodal features, thus combining the advantages of various modal features, ensuring better environmental adaptability, and guaranteeing the detection accuracy of combined vehicle features. The bird's-eye view features are downsampled at different sampling factors to obtain multiple sets of downsampled results, which can be used to detect traffic participants of different sizes. The target sampling result with the highest downsampling factor among the multiple sampling results is used to detect combined vehicles. Since the target sampling result with a high downsampling factor can provide a larger receptive field, it can ensure the identification of larger targets, i.e., the complete picture of the combined vehicle, avoiding truncation of the combined vehicle and ensuring the integrity of the combined vehicle in the vehicle detection result.
[0041] In one possible implementation, the bird's-eye view features are downsampled by different sampling factors to obtain multiple sets of downsampling results. This can be achieved as follows: the bird's-eye view features are downsampled by two times to obtain a two-times downsampling result; the bird's-eye view features are downsampled by four times to obtain a four-times downsampling result; and the bird's-eye view features are downsampled by eight times to obtain an eight-times downsampling result.
[0042] In this possible implementation, the bird's-eye view features are downsampled by a factor of two to obtain a downsampled result. This feature layer has high resolution and is suitable for small target detection, such as pedestrians and non-motorized vehicles, where more detailed texture information needs to be preserved. The bird's-eye view features are downsampled by a factor of four to obtain a downsampled result. This feature layer balances resolution and computational efficiency and is suitable for conventional vehicle detection. The bird's-eye view features are downsampled by a factor of eight to obtain an downsampled result. This feature layer has the lowest resolution and is suitable for detecting very large targets (such as very long vehicles or combined vehicles). Low-resolution modeling reduces feature sparsity and bounding box drift problems caused by excessive scale at high resolution.
[0043] In this embodiment, it should be noted that downsampling, an operation that reduces the spatial resolution of feature maps, can be achieved through pooling or convolution. Although the amount of data is reduced, the actual scene range covered by a single feature unit is increased. The receptive field is the input spatial range corresponding to a given output feature point in the network. An increased receptive field means that the feature point can perceive a larger range of scene information. The kernel size is the window size used for convolution operations; here, a 3×3 kernel is used to capture a mixture of local and global features on low-resolution feature maps. The 3D box, or three-dimensional detection box, is used to locate and enclose the volume range of a target in space.
[0044] In this possible implementation, to adapt to the detection requirements of ultra-large targets (such as extra-long vehicles or combined vehicles) in BEV features, this embodiment performs a second downsampling on top of the original 4x downsampling features, resulting in 8x downsampling features. Although the downsampling operation reduces the spatial resolution of the feature map, it significantly expands the model's receptive field. That is, the area of the real scene covered by a single feature unit on the feature map doubles, thus enabling the aggregation of more global target information in a single feature unit. Based on the 8x downsampling feature layer, several convolutional modules (with a kernel size of 3×3) are introduced for feature extraction. This convolutional operation not only further expands the model's receptive field but also extracts the high-level semantic features required for ultra-large targets, enabling the model to describe the overall shape and structure of the target. Through the above design, the model can obtain the complete picture of ultra-large targets on the low-resolution feature layer, allowing the regressed 3D detection boxes to accurately cover the entire target and reducing the detection box truncation problem caused by the target span exceeding the high-resolution feature field of view.
[0045] In one possible implementation, detecting a combined vehicle based on the target sampling result among the multiple sets of downsampling results can be performed by the following steps: detecting a combined vehicle based on the eight-fold downsampling result; the length of the combined vehicle is greater than a specified vehicle length threshold.
[0046] In this possible implementation, the length of the combined vehicle is greater than a specified length threshold, which is 25 meters. For combined vehicles, such as ultra-large targets including but not limited to extra-long vehicles, tractors, and trailers, an eight-fold downsampling feature model is used to achieve a single feature unit covering a larger physical range on a low-resolution downsampling layer. This helps to capture the overall information of ultra-large targets and improves detection stability and accuracy.
[0047] In one possible implementation, the following steps may also be performed: detecting non-motorized vehicles based on the double downsampling results, and detecting non-combined vehicles based on the quadruple downsampling results.
[0048] In this possible implementation, multiple sets of downsampling results are input into corresponding target detection models. Each target detection model uses features at different scales to match targets of different sizes, thereby achieving differentiated detection. Non-motorized vehicles may include, but are not limited to, pedestrians, bicycles, traffic cones, traffic signs, and other small traffic participants. Non-combined vehicles may include, but are not limited to, conventional vehicles such as cars, vans, trucks, and buses.
[0049] To address the bird's-eye view features at different scales, this application embodiment sets up multiple detection heads to achieve differentiated detection task allocation. For example, for the BEV feature layers of small, medium, and very large targets, independent detection heads (output modules in the neural network responsible for specific detection tasks; different detection heads can be trained independently for different scales or tasks) are used for prediction to ensure that features at different scales can match the corresponding target types. Each detection head performs 3D target detection on the BEV features at its corresponding scale, and the output detection results (the detection results corresponding to the eight-fold downsampling results are the combined vehicle detection results) include the following information: target category: including vehicle category and component identification, such as the front of the vehicle and the trailer; 3D position coordinates (x, y, z): used to determine the precise position of the target in the vehicle coordinate system; size parameters (l, w, h): representing the length, width, and height of the target, used to describe the target's shape; orientation (yaw): the rotation angle of the target relative to the vehicle coordinate system; speed: the target's moving speed, suitable for dynamic scene tracking and prediction. Through the above detection output, the system can finally generate a complete three-dimensional detection result (3D Bounding Box). This result not only includes the target category, but also integrates position, shape, orientation and velocity information, providing a unified data interface for subsequent path planning, motion prediction and combined vehicle recognition.
[0050] In this embodiment, a multi-scale feature layer is constructed based on the fused BEV features, which can adapt to the detection tasks of targets of different sizes, including small, medium, and large. Independent detection heads are set for features of different scales to improve the detection accuracy of small, conventional, and ultra-large targets.
[0051] Considering the significant non-rigid characteristics of extra-long vehicles (such as trailer connections), their various parts may exhibit non-uniform motion during turning or lane changes, making it difficult for traditional detection models based on rigid body assumptions to accurately model them. Therefore, in one possible implementation, the combined vehicle includes multiple sets of local structures; identifying the combined vehicle based on the target sampling results in the multiple sets of downsampling results can be performed according to the following steps: based on the target sampling results in the multiple sets of downsampling results, identify the multiple sets of local structures respectively, obtaining multiple visible boxes; based on the target sampling results in the multiple sets of downsampling results, identify the spatial relationship between any two sets of adjacent local structures; the spatial relationship is used to characterize that the two sets of adjacent local structures correspond to the same vehicle; the combined vehicle is determined based on the multiple visible boxes and the spatial relationship.
[0052] In this possible implementation, the combined vehicle is an integral unit composed of multiple physically connected vehicle units. These vehicle units are considered local structures. It should be noted that a local structure can be the structure of the vehicle itself or the transported object carried by the vehicle. For example, a local structure may include, but is not limited to: the tractor unit, different sections of the trailer body, the rear unit, and onboard objects, etc. Visible boxes are used to represent the local detection results of the combined vehicle and are areas that can be displayed on the interface with a specified shape and color. Spatial relationships include, but are not limited to: the relative position, spacing, direction, and connection relationship between the components, used to determine whether these local structures belong to the same vehicle in terms of physical connection.
[0053] Based on multiple visible bounding boxes and spatial relationships, the arrangement and connection of local structures are comprehensively analyzed to determine the overall detection result of the combined vehicle. For example, when the spatial relationships between multiple visible bounding boxes satisfy connection constraints, they are merged into the detection box of the same combined vehicle.
[0054] In this possible implementation, the combined vehicle is divided into multiple identifiable units, improving the detectability of ultra-large targets in complex scenes, reducing missed detections due to occlusion or cross-field of view, enhancing the structural integrity of target detection results, and avoiding the problems of truncation and instability caused by ultra-long vehicles. Based on spatial relationship recognition operations, the connection relationship between local structures can be established, accurately determining whether they constitute the same vehicle as a whole, avoiding the mistaken splitting of the combined vehicle into multiple independent targets. By integrating local detection and spatial relationships, the accuracy of detection and positioning is improved, achieving robust identification of ultra-large targets. In the case of multiple occlusions and long vehicles occupying the field of view of multiple cameras, the embodiments of this application can still reconstruct the entire vehicle through spatial relationships, improving system robustness.
[0055] In one possible implementation, the multiple sets of local structures include a tractor and at least one trailer; the spatial relationship is a connection relationship; based on the target sampling results in the multiple sets of downsampling results, the multiple sets of local structures are identified respectively to obtain multiple visible boxes, which can be performed according to the following steps: based on the target sampling results in the multiple sets of downsampling results, the tractor and the trailer are identified respectively to obtain a first visible box of the tractor and a second visible box of the trailer.
[0056] In this possible implementation, the tractor unit can be the cab of a truck. The connection between the tractor unit and the trailer is articulated. When the partial structure includes multiple trailers, the connection between the trailers is articulated. Figure 2 A schematic diagram of the components of a modular vehicle is shown, illustrating a vehicle comprising a cab and a trailer. Figure 2The demonstrated combined vehicle identifies multiple sets of local structures, resulting in two visible frames. The visible frame of the tractor unit can be used as the first visible frame, and the visible frame of the trailer unit as the second visible frame. See also... Figure 3 The diagram shows a visible frame corresponding to a combined vehicle. The first visible frame for the front of the vehicle is a red rectangular area, and the second visible frame for the trailer is a yellow rectangular area.
[0057] In this possible implementation, spatial constraints (such as relative position or connection relationship) are established between the predicted 3D frames of the vehicle front and the trailer. This more closely resembles the real physical structure and improves detection stability and accuracy.
[0058] In one possible implementation, the multiple sets of local structures include a vehicle body and a transport object; the spatial relationship is a relative positional relationship; based on the target sampling results in the multiple sets of downsampling results, the multiple sets of local structures are identified respectively to obtain multiple visible boxes, which can be performed according to the following steps: based on the target sampling results in the multiple sets of downsampling results, the vehicle body and the transport object are identified respectively to obtain a third visible box of the vehicle body and a fourth visible box of the transport object.
[0059] In this possible implementation, the transport object is loaded in the vehicle body, and the transport object includes an area extending outside the vehicle body. Figure 4 A schematic diagram of a vehicle and a transported object is shown. The combined vehicle depicted is a vehicle comprising two partial structures: a vehicle body and a transported object. The transported object is a fan blade that occupies more space than the vehicle body itself. It should be noted that this application does not specifically limit the type of transported object; it can be determined based on the actual type of cargo being loaded. Figure 4 The illustrated combined vehicle identifies multiple sets of local structures, resulting in two visible bounding boxes. The visible bounding box of the vehicle body can be used as the third visible bounding box, and the visible bounding box of the transported objects can be used as the fourth visible bounding box. See also Figure 5 The diagram shows a visible frame of a transported object. The fourth visible frame of the transported object is a yellow rectangular area. The third visible frame of the vehicle body is not shown in the diagram.
[0060] In one possible implementation, the multimodal features include at least one or more of the following data: first point cloud data acquired by millimeter-wave radar, second point cloud data acquired by lidar, and image data acquired by camera.
[0061] In this possible implementation, the first point cloud data from the millimeter-wave radar, including point data and velocity data, can be used to characterize information such as distance, angle, altitude, and velocity. The second point cloud data from the lidar includes point data and reflectivity data. A camera acquires multi-view camera data, and after preprocessing this data such as correction and normalization, image data acquired by the camera is obtained.
[0062] In this possible implementation, the image data acquired by the camera can provide rich semantic information, the LiDAR can provide high-precision spatial structure, and the 4D millimeter-wave radar can provide stable distance and velocity information. Through multimodal fusion, a complementary fusion strategy of different sensor information in the BEV space fully leverages the complementary advantages of each sensor in ultra-long vehicle detection, effectively improving the system's performance in the following scenarios: low light, backlight, and other visually degraded environments; adverse weather conditions such as rain and fog; and long-distance and partially occluded scenarios. The target detection method based on three-modal collaborative fusion in this application embodiment can improve robustness in complex environments.
[0063] In one possible implementation, obtaining bird's-eye view features within the field of vision of an autonomous vehicle can be performed according to the following steps: feature extraction is performed based on the multimodal features to obtain multiple sets of feature extraction results; the multiple sets of feature extraction results are transformed by perspective to obtain multiple sets of data to be fused in the autonomous vehicle's own coordinate system; the weights of the multiple sets of data to be fused are determined respectively, and feature fusion is performed on the multiple sets of data to be fused according to the weights to obtain a target frame; the bird's-eye view features are obtained by fusing multiple sets of temporally adjacent target frames.
[0064] In this possible implementation, the multimodal sensors of the autonomous driving system include LiDAR, millimeter-wave radar, and cameras, whose output data have different spatial resolutions, perception dimensions, and information types. The following section combines... Figure 7 The steps for extracting features based on the multimodal features to obtain multiple sets of feature extraction results are described. See [link to documentation]. Figure 7 The diagram illustrates the implementation steps of a combined vehicle detection method. For both LiDAR point clouds and 4D millimeter-wave radar, the original 3D point cloud data is divided into pillars (pillar voxels, which are 3D data pillars in the BEV grid used to encode point cloud information). Specifically, a grid is created on the BEV plane, and point clouds in the vertical direction are aggregated within each grid to form pillar data units. Using a PointPillars-based encoding method (a neural network method that encodes point cloud pillars into high-dimensional vectors and directly generates BEV features), the point cloud within each pillar is converted into a high-dimensional vector representation and projected onto the BEV feature map. The following feature extraction results are obtained: LiDAR BEV features: possessing high-precision distance and shape information, suitable for detecting targets with clear structures; 4D millimeter-wave radar BEV features: containing point cloud coordinates along with velocity information (Doppler velocity measurement), suitable for identifying moving and distant targets.
[0065] For visual feature extraction and BEV mapping, convolutional neural networks (e.g., YOLOv8 Backbone + FPN) are used to extract multi-scale image features, obtaining two-dimensional multi-scale feature maps in perspective space. The ViewTransform method based on FastBEV (a perspective transformation method that maps perspective camera view features to the BEV spatial coordinate system) maps image features from the perspective coordinate system to the BEV plane coordinate system, achieving spatial consistency of visual information. The resulting feature extraction is visual BEV features, containing rich texture and color information.
[0066] Next, a perspective transformation is performed, that is, the multiple sets of extracted feature results are transformed to obtain multiple sets of data to be fused in the autonomous vehicle's own coordinate system. Specifically, this can be achieved as follows: Based on extrinsic parameter calibration, the two types of point cloud data can be uniformly transformed to the autonomous vehicle's own coordinate system (ego coordinate system, a fixed coordinate system with the vehicle center as the origin, the forward direction as the X-axis, and the left direction as the Y-axis), ensuring spatial alignment consistency and providing a foundation for subsequent BEV (Bird's EyeView) modeling; based on camera extrinsic and intrinsic parameters, a mapping relationship is established between the image coordinate system of the image data acquired by the camera and the autonomous vehicle's own coordinate system. Furthermore, the above three types of BEV features are respectively transformed to the autonomous vehicle's own coordinate system (a fixed coordinate system with the vehicle center as the origin), ensuring that the data of different modalities are aligned at the same spatial scale. Three sets of data to be fused are obtained: the first set: LiDAR BEV feature data; the second set: 4D millimeter-wave radar BEV feature data; the third set: visual BEV feature data. In this possible implementation, different modalities are uniformly mapped to the BEV space, avoiding the complex calculations of traditional "pixel-level alignment" or "point-level projection".
[0067] Next, multimodal BEV feature fusion is performed, which involves determining the weights of the multiple sets of data to be fused, and then fusing the features of these multiple sets of data according to the weights to obtain the target frame. Specifically, this can be achieved through the following steps: determining the weights of the multiple sets of data to be fused, which can be dynamically adjusted based on factors such as sensor signal-to-noise ratio, distance range, occlusion, weather conditions, and detection task type. In one possible implementation, a convolutional multi-scale fusion network can be used to fusing features to obtain the target frame: three types of BEV features (LiDAR / millimeter-wave radar / vision) are input into the fusion module, and channel alignment is performed on the different modal features (ensuring consistent feature dimensions); layer-by-layer feature fusion is performed, such as a combination of concat and convolution, or weighted fusion based on an attention-like mechanism. The resulting fused target frame is a comprehensive perception result of a frame in the time series. In this possible implementation, fusion is performed in the BEV space, resulting in higher computational efficiency and a clearer structure. Figure 6 This diagram illustrates the cross-modal fusion module structure in the aforementioned convolutional-based multi-scale fusion network. Figure 6 In this system, multimodal data includes data acquired by LiDAR and data acquired by cameras. The cross-modal fusion module structure may include a LiDAR feature input unit 101, a visual feature input unit 102, a stitching and fusion unit 103, a spatial fusion unit 104, an attention enhancement unit 105, and a fusion feature output unit 106.
[0068] The system comprises several modules: a LiDAR feature input module, which receives LiDAR BEV features (batch_lidar_backbone_features) extracted by the LiDAR backbone network. The features have dimensions of B×128×H×W, where B is the batch size, H and W are the height and width of the feature map, and 128 is the number of channels for that modality. A visual feature input module receives Vision BEV features (batch_vision_backbone_features) extracted by the visual backbone network. These features also have dimensions of B×128×H×W and are spatially aligned with the features output by the LiDAR feature input unit, but originate from different sensing modalities. A stitching and fusion unit, connected to both the LiDAR and visual feature input units, stitches the two input features along the channel dimension (Concat, dim=1), achieving preliminary fusion of the two modalities without altering the spatial resolution. The output feature after stitching has dimensions of B×256×H×W. The spatial fusion unit, connected to the concatenation fusion unit, performs spatial convolutional fusion processing on the concatenated features. Specifically, it includes a 2D convolutional layer (Conv2d) with a 3×3 kernel and 256→64 input / output channels, followed by a batch normalization layer (BatchNorm2d) and a ReLU activation layer, used to extract local spatial information of the fused features while reducing dimensionality. The attention enhancement unit, also connected to the spatial fusion unit, performs global context modeling on the spatially fused features. It employs a Global Context Block structure, combining channel additive enhancement with a Squeeze-and-Excitation (SE) mechanism to adaptively weight the feature responses of different channels, thereby enhancing the expressive power of key feature channels. This unit outputs 64 feature channels. The fusion feature output unit, connected to the attention enhancement unit, is used to output cross-modality fusion features obtained after splicing fusion, spatial fusion and attention enhancement. The size of the output feature is B×64×H×W, which can be further used by subsequent detection or recognition networks.
[0069] Next, temporal fusion is performed, which can be implemented based on an attention mechanism. That is, multiple sets of temporally adjacent target frames are fused to obtain the bird's-eye view features. Specifically, this can be achieved through the following steps: Multiple sets of adjacent target frames are obtained from the historical time series, and temporal alignment is performed based on the vehicle's pose information (position + orientation) to eliminate feature offsets caused by the vehicle's own motion. An attention-based temporal fusion module is used to weight the features of different time frames: higher stability weights are given to distant targets; for occluded or temporarily missing targets, historical frame information is used to complete the detection results. This process outputs the enhanced temporal BEV features (i.e., the final bird's-eye view features).
[0070] In this possible implementation, multimodal feature extraction utilizes the precise distance sensing of LiDAR, the long-range speed sensing of millimeter-wave radar, and the texture and color information of vision to achieve complementary perception and improve stability under different lighting and weather conditions. Unified viewpoint transformation (vehicle coordinate system alignment) eliminates parallax and scale differences between different sensors, providing a consistent spatial reference for subsequent fusion. Weighted fusion dynamically adjusts the importance of different modalities according to the environment, enhancing the system's adaptability. Multi-frame temporal fusion uses historical frame information to compensate for occluded or temporarily lost targets, enhancing detection continuity and robustness. The final BEV output integrates multi-source information from a unified top-down viewpoint and can be directly used for advanced tasks such as path planning, target detection, and combined vehicle recognition.
[0071] This application provides a combined vehicle detection method, electronic device, storage medium, and program product. The method comprehensively utilizes data from cameras, LiDAR, and 4D millimeter-wave radar to perform unified modeling in the BEV space. Through a multi-scale and temporal fusion mechanism, it achieves high-precision detection of extra-long vehicles (length ≥ 25m). The method divides extra-long vehicles into two sub-target categories: the tractor unit (tractor) and the trailer (trailer). Figure 2 Based on the actual structure of the vehicle, the front end and trailer are represented using two categories (TRUCK_HEAD, TRUCK_TRAILER), establishing the spatial constraint relationship between the front end and trailer, characterizing their hinged connection structure, and jointly modeling multiple substructures during the detection phase to improve the overall detection integrity and stability. This application provides a real-time multimodal BEV detection framework, combining structural modeling and multi-scale strategy computational optimization methods, and an efficient multimodal fusion network structure design. It achieves a balance between performance and efficiency through the following methods: unified modeling in the BEV space to reduce cross-modal computational complexity; utilizing structured feature representation and multi-scale strategies to reduce redundant computation; and combining lightweight convolution and attention mechanisms to achieve efficient inference.
[0072] Corresponding to the application scenarios and methods provided in the embodiments of this application, the embodiments of this application also provide a combined vehicle detection device. For example... Figure 8 The diagram shown is a structural block diagram of a combined vehicle detection device according to an embodiment of this application. The device may include:
[0073] The acquisition module 801 acquires bird's-eye view features within the field of vision of the autonomous vehicle; the bird's-eye view features are obtained by fusing multiple sets of temporally adjacent target frames; the target frames are obtained by fusing multimodal features; the downsampling module is used to downsample the bird's-eye view features by different sampling multiples to obtain multiple sets of downsampling results; the detection module is used to detect combined vehicles based on the target sampling results in the multiple sets of downsampling results to obtain combined vehicle detection results; the sampling multiple of the target sampling results is higher than the sampling multiple of the sampling results other than the target sampling results in the multiple sets of downsampling results.
[0074] According to the embodiments of this application, firstly, bird's-eye view features of each traffic participant within the field of vision of the autonomous vehicle are obtained. The bird's-eye view features are obtained by fusing multiple sets of temporally adjacent target frames, thereby preserving more complete combined vehicle features. The target frames are obtained by fusing multimodal features, thereby combining the advantages of various modal features, ensuring better environmental adaptability, and ensuring the detection accuracy of combined vehicle features. The bird's-eye view features are downsampled by different sampling factors to obtain multiple sets of downsampled results, which can be used to detect traffic participants of different sizes. The target sampling result with the highest downsampling factor among the multiple sampling results is used to detect combined vehicles. Since the target sampling result with a high downsampling factor can provide a larger receptive field, it can ensure the identification of larger targets, i.e., the full picture of combined vehicles, avoiding the combined vehicles being truncated and ensuring the integrity of combined vehicles in the vehicle detection results.
[0075] In one possible implementation, the bird's-eye view features are downsampled by different sampling factors to obtain multiple sets of downsampling results, including: downsampling the bird's-eye view features by 2 times to obtain a 2-fold downsampling result, downsampling the bird's-eye view features by 4 times to obtain a 4-fold downsampling result, and downsampling the bird's-eye view features by 8 times to obtain an 8-fold downsampling result.
[0076] In one possible implementation, detecting a combined vehicle based on the target sampling result among the multiple sets of downsampling results includes: detecting a combined vehicle based on the eight-fold downsampling result; wherein the length of the combined vehicle is greater than a specified vehicle length threshold.
[0077] In one possible implementation, the detection module is further configured to: detect non-motorized vehicles based on the double downsampling result, and detect non-combined vehicles based on the quadruple downsampling result.
[0078] In one possible implementation, the combined vehicle includes multiple sets of local structures; identifying the combined vehicle based on the target sampling results in the multiple sets of downsampling results includes: identifying the multiple sets of local structures respectively based on the target sampling results in the multiple sets of downsampling results to obtain multiple visible boxes; identifying the spatial relationship between any two sets of adjacent local structures based on the target sampling results in the multiple sets of downsampling results; the spatial relationship is used to characterize that the two sets of adjacent local structures correspond to the same vehicle; and determining the combined vehicle based on the multiple visible boxes and the spatial relationship.
[0079] In one possible implementation, the multiple sets of local structures include a tractor and at least one trailer; the spatial relationship is a connection relationship; based on the target sampling results in the multiple sets of downsampling results, the multiple sets of local structures are identified respectively to obtain multiple visible boxes, including: based on the target sampling results in the multiple sets of downsampling results, the tractor and the trailer are identified respectively to obtain a first visible box of the tractor and a second visible box of the trailer.
[0080] In one possible implementation, the multiple sets of local structures include a vehicle body and a transport object; the spatial relationship is a relative positional relationship; based on the target sampling results in the multiple sets of downsampling results, the multiple sets of local structures are identified respectively to obtain multiple visible boxes, including: based on the target sampling results in the multiple sets of downsampling results, the vehicle body and the transport object are identified respectively to obtain a third visible box of the vehicle body and a fourth visible box of the transport object.
[0081] In one possible implementation, the multimodal features include at least one or more of the following data: first point cloud data acquired by millimeter-wave radar, second point cloud data acquired by lidar, and image data acquired by camera.
[0082] In one possible implementation, obtaining bird's-eye view features within the field of vision of an autonomous vehicle includes: performing feature extraction based on the multimodal features to obtain multiple sets of feature extraction results; performing viewpoint transformation on the multiple sets of feature extraction results to obtain multiple sets of data to be fused in the autonomous vehicle's own coordinate system; determining the weights of the multiple sets of data to be fused, and performing feature fusion on the multiple sets of data to be fused according to the weights to obtain a target frame; and fusing multiple sets of temporally adjacent target frames to obtain the bird's-eye view features.
[0083] The functions of each module in each device in the embodiments of this application can be found in the corresponding description in the above method, and they have corresponding beneficial effects, which will not be repeated here.
[0084] Figure 9 This is a block diagram of an electronic device used to implement embodiments of this application. For example... Figure 9As shown, the electronic device includes a memory 901 and a processor 902. The memory 901 stores a computer program that can run on the processor 902. When the processor 902 executes the computer program, it implements the method described in the above embodiments. The number of memories 901 and processors 902 can be one or more.
[0085] The electronic device also includes:
[0086] The communication interface 903 is used to communicate with external devices and exchange and transmit data.
[0087] If the memory 901, processor 902, and communication interface 903 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0088] Optionally, in a specific implementation, if the memory 901, processor 902, and communication interface 903 are integrated on a single chip, then the memory 901, processor 902, and communication interface 903 can communicate with each other through an internal interface.
[0089] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method provided in this application.
[0090] This application provides a computer program product, wherein the computer program product includes a computer program, which, when executed by a processor, implements the method provided in this application embodiment.
[0091] This application also provides a chip including a processor for calling and executing instructions stored in a memory, causing a communication device with the chip installed to perform the method provided in this application.
[0092] This application also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in the application embodiment.
[0093] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.
[0094] Further, optionally, the aforementioned memory may include read-only memory and random access memory. The memory may be volatile memory or non-volatile memory, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Sync Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0095] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.
[0096] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0097] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
[0098] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functionality involved.
[0099] The logic and / or steps described in the flowchart or otherwise herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).
[0100] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiments.
[0101] It should be noted that the data involved in this application (including but not limited to data used for analysis, data stored, data displayed, etc.) are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0102] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.
[0103] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope described in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A combined vehicle detection method, characterized in that, include: Obtain bird's-eye view features within the field of vision of autonomous vehicles; The bird's-eye view features are obtained by fusing multiple sets of temporally adjacent target frames; The target frame is obtained based on multimodal feature fusion; The bird's-eye view features were downsampled by different sampling factors to obtain multiple sets of downsampling results; Based on the target sampling result in the multiple sets of downsampling results, the combined vehicle is detected to obtain the combined vehicle detection result; the sampling multiple of the target sampling result is higher than the sampling multiple of the other sampling results in the multiple sets of downsampling results. The combined vehicle includes multiple sets of local structures; Identifying combined vehicles based on target sampling results from the multiple sets of downsampling results includes: Based on the target sampling results in the multiple sets of downsampling results, the multiple sets of local structures are identified respectively to obtain multiple visible boxes; Based on the target sampling results in the multiple sets of downsampling results, the spatial relationship between any two sets of adjacent local structures is identified; the spatial relationship is used to characterize that the two sets of adjacent local structures correspond to the same vehicle. The combined vehicle is determined based on the multiple visible frames and the spatial relationships.
2. The method according to claim 1, characterized in that, The bird's-eye view features were downsampled by different sampling factors to obtain multiple sets of downsampling results, including: The bird's-eye view features are downsampled by a factor of 2 to obtain a downsampled result, the bird's-eye view features are downsampled by a factor of 4 to obtain a downsampled result, and the bird's-eye view features are downsampled by a factor of 8 to obtain a downsampled result.
3. The method according to claim 2, characterized in that, Detecting combined vehicles based on target sampling results from the multiple sets of downsampling results includes: The combined vehicle is detected based on the eight-fold downsampling result; the length of the combined vehicle is greater than a specified vehicle length threshold.
4. The method according to claim 2, characterized in that, Also includes: Non-motorized vehicles are detected based on the double downsampling results, and non-combined vehicles are detected based on the quadruple downsampling results.
5. The method according to claim 1, characterized in that, The multiple sets of local structures include a tractor and at least one vehicle trailer; The spatial relationship is a connection relationship; Based on the target sampling results in the multiple sets of downsampling results, the multiple sets of local structures are identified respectively, resulting in multiple visible bounding boxes, including: Based on the target sampling results in the multiple sets of downsampling results, the tractor and the trailer are identified respectively, and the first visible frame of the tractor and the second visible frame of the trailer are obtained.
6. The method according to claim 1, characterized in that, The multiple sets of local structures include the vehicle body and the transported object; the spatial relationship is a relative positional relationship. Based on the target sampling results in the multiple sets of downsampling results, the multiple sets of local structures are identified respectively, resulting in multiple visible bounding boxes, including: Based on the target sampling results in the multiple sets of downsampling results, the vehicle body and the transported object are identified respectively, and the third visible box of the vehicle body and the fourth visible box of the transported object are obtained.
7. The method according to claim 1, characterized in that, The multimodal features include at least one or more of the following data: first point cloud data acquired by millimeter-wave radar, second point cloud data acquired by lidar, and image data acquired by camera.
8. The method according to claim 1, characterized in that, Obtain bird's-eye view features within the field of view of the autonomous vehicle, including: Based on the multimodal features, feature extraction is performed separately to obtain multiple sets of feature extraction results; The multiple sets of feature extraction results are transformed by perspective to obtain multiple sets of data to be fused in the vehicle coordinate system of the autonomous vehicle. The weights of the multiple sets of data to be fused are determined respectively, and feature fusion is performed on the multiple sets of data to be fused according to the weights to obtain the target frame; The bird's-eye view features are obtained by fusing multiple sets of temporally adjacent target frames.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements the method of any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 8.
11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 8.