Target detection method based on top view angle, storage medium and electronic equipment
By establishing a coordinate mapping relationship between lane elements and target elements from a top-down perspective, image distortion and size changes are corrected, solving the problem of low target detection accuracy from a top-down perspective and achieving precise positioning and detection of target elements.
Patent Information
- Application Number
- CN202411119532.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2026-03-03
AI Technical Summary
The accuracy of target detection based on a top-down view is relatively low, mainly because the large distance between the image acquisition device and the target element leads to changes in size and sharpness, which affects the detection accuracy.
By acquiring lane images from a top-down view using an image acquisition device, the coordinate positions of lane elements and target elements in different coordinate systems are determined, a precise coordinate mapping relationship is established, and the position information of the image acquisition device in the actual coordinate system is used to correct image distortion and size changes, thereby achieving accurate positioning of the target elements.
It improves the accuracy of target detection based on a top-down view, realizes the precise position calculation of target elements in lane scenes, and enhances the accuracy of detection results.
Smart Images

Figure CN121600481A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and more specifically, to a target detection method, apparatus, storage medium, and electronic device based on a top-down view. Background Technology
[0002] In top-down view-based target detection scenarios, image acquisition devices can capture images of the lane scene from a top-down perspective, and then perform target detection on the acquired images. However, due to the considerable distance between the image acquisition device and the target element, the size and sharpness of the target element in the image can change, affecting the accuracy of target detection and resulting in low accuracy for top-down view-based target detection. Therefore, the problem of low accuracy in top-down view-based target detection exists.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This application provides a target detection method, apparatus, storage medium, and electronic device based on a top-down view, to at least solve the technical problem of low accuracy in target detection based on a top-down view.
[0005] According to one aspect of the embodiments of this application, a target detection method based on a top-down view is provided, comprising: acquiring a lane image obtained by an image acquisition device from a top-down view of a lane scene, wherein the lane image contains at least one lane element and a target element, the lane element being an element carrying attribute information in the lane scene, and the target element being an element to be detected in the lane image; acquiring a first coordinate position of the at least one lane element in a first coordinate system, acquiring a second coordinate position of the at least one lane element in a second coordinate system, and acquiring a third coordinate position of the target element in the second coordinate system, wherein the attribute information includes the first coordinate position, the first coordinate system being the actual coordinate system corresponding to the lane scene, and the second coordinate system being the pixel coordinate system corresponding to the lane image; acquiring a fourth coordinate position of the target element in the first coordinate system based on a first positional relationship between the second coordinate position and the third coordinate position, and the first coordinate position; and, if a fifth coordinate position of the image acquisition device in the first coordinate system is acquired, acquiring a detection result of the target element in the lane scene based on a second positional relationship between the fifth coordinate position and the fourth coordinate position.
[0006] According to another aspect of the embodiments of this application, a target detection device based on a top-down view is also provided, comprising: a first acquisition unit, configured to acquire a lane image obtained by an image acquisition device from a top-down view of a lane scene, wherein the lane image includes at least one lane element and a target element, the lane element being an element carrying attribute information in the lane scene, and the target element being an element to be detected in the lane image; and a second acquisition unit, configured to acquire a first coordinate position of the at least one lane element in a first coordinate system, acquire a second coordinate position of the at least one lane element in a second coordinate system, and acquire a third coordinate position of the target element in the second coordinate system. The coordinate position, wherein the aforementioned attribute information includes the aforementioned first coordinate position, the aforementioned first coordinate system is the actual coordinate system corresponding to the aforementioned lane scene, and the aforementioned second coordinate system is the pixel coordinate system corresponding to the aforementioned lane image; the third acquisition unit is used to acquire the fourth coordinate position of the aforementioned target element in the aforementioned first coordinate system based on the first positional relationship between the aforementioned second coordinate position and the aforementioned third coordinate position, and the aforementioned first coordinate position; the fourth acquisition unit is used to acquire the detection result of the aforementioned target element in the aforementioned lane scene based on the second positional relationship between the aforementioned fifth coordinate position and the aforementioned fourth coordinate position when the fifth coordinate position of the aforementioned image acquisition device in the aforementioned first coordinate system is acquired.
[0007] As an optional solution, the third acquisition unit includes: a determining module, configured to determine at least one first lane element and at least one second lane element from the at least one lane element, wherein the first lane element is in the second coordinate system and the first positional relationship satisfies that the lane element is adjacent to the target element in a first direction, and the second lane element is in the second coordinate system and the first positional relationship satisfies that the lane element is adjacent to the target element in a second direction; and a first acquisition module, configured to acquire, based on the at least one first lane element and the at least one second lane element, a first coordinate of the target element in the first coordinate system located in the first direction, and a second coordinate of the target element in the first coordinate system located in the second direction, wherein the fourth coordinate position includes the first coordinate and the second coordinate.
[0008] As an optional solution, the first acquisition module includes: a first acquisition submodule, configured to acquire the first pixel coordinates of the target element in the second coordinate system and located in the first direction, the first actual coordinates of the first element among the at least one first lane element in the first coordinate system and located in the first direction, the second actual coordinates of the second element among the at least one first lane element in the first coordinate system and located in the first direction, the second pixel coordinates of the first element in the second coordinate system and located in the first direction, and the second pixel coordinates of the second element in the second coordinate system and located in the first direction. The third pixel coordinate located in the first direction; the second acquisition submodule is used to acquire a first difference between the first pixel coordinate and the first actual coordinate, a second difference between the second actual coordinate and the first actual coordinate, and a third difference between the third pixel coordinate and the second pixel coordinate; the third acquisition submodule is used to acquire a first product of the first difference and the second difference; the fourth acquisition submodule is used to acquire a first ratio of the first product to the third difference; the fifth acquisition submodule is used to acquire a first sum of the first actual coordinate and the first ratio, and set the first sum... The first coordinate is determined as described above; and the sixth acquisition submodule is used to acquire the fourth pixel coordinate of the target element in the second coordinate system and located in the second direction, the third actual coordinate of the third element of the at least one second lane element in the first coordinate system and located in the second direction, the fourth actual coordinate of the fourth element of the at least one second lane element in the first coordinate system and located in the second direction, the fifth pixel coordinate of the third element in the second coordinate system and located in the second direction, and the fifth actual coordinate of the fourth element in the second coordinate system and located in the second direction. The sixth pixel coordinate; the seventh acquisition submodule, used to acquire the fourth difference value between the fourth pixel coordinate and the third actual coordinate, the fifth difference value between the fourth actual coordinate and the third actual coordinate, and the sixth difference value between the third pixel coordinate and the second pixel coordinate; the eighth acquisition submodule, used to acquire the second product value of the product of the fourth difference value and the fifth difference value; the ninth acquisition submodule, used to acquire the second ratio value of the ratio of the second product value and the sixth difference value; the tenth acquisition submodule, used to acquire the second sum value of the sum of the third actual coordinate and the second ratio value, and determine the second sum value as the second coordinate.
[0009] As an optional solution, the above-mentioned device further includes: a storage module, configured to, after obtaining the first coordinate of the target element in the first coordinate system located in the first direction and the second coordinate of the target element in the first coordinate system located in the second direction based on the at least one first lane element and the at least one second lane element, store the projection information between the first coordinate and the second coordinate in a pixel projection mapping table, wherein the projection information between the first coordinate and the second coordinate is used to indicate whether the target element is projected from the first coordinate system to the second coordinate system or from the second coordinate system to the first coordinate system, and the pixel projection mapping table is used to store and query the projection information corresponding to each element.
[0010] As an optional solution, the above-mentioned device further includes: a second acquisition module, configured to acquire a projection query request triggered by the target range of the lane image after storing the projection information between the first coordinate and the second coordinate in a pixel projection mapping table, wherein the projection query request is used to request a query for the target position of a target element in the lane image located within the target range on the first coordinate system; a third acquisition module, configured to acquire multiple target elements in the lane image located within the target range after storing the projection information between the first coordinate and the second coordinate in a pixel projection mapping table; a query module, configured to query the pixel projection mapping table for multiple candidate coordinate positions of the multiple target elements in the first coordinate system based on the multiple coordinate positions of the multiple candidate elements in the second coordinate system after storing the projection information between the first coordinate and the second coordinate in a pixel projection mapping table; and a filtering module, configured to filter the coordinate positions located outside a predetermined detection range among the multiple candidate coordinate positions after storing the projection information between the first coordinate and the second coordinate in a pixel projection mapping table to obtain at least one target position.
[0011] As an optional solution, the above-mentioned device further includes: an input unit, configured to input the lane image obtained by the image acquisition device from a top-down perspective into a target detection model, wherein the target detection model is a neural network model trained by multiple image samples for detecting the target element; and a fifth acquisition unit, configured to acquire the detection result output by the target detection model after the image acquisition device obtains the lane image from a top-down perspective.
[0012] As an optional solution, the image acquisition device includes N acquisition sub-devices, characterized in that the device further includes: an integration unit, used to integrate the N lane images acquired by the N acquisition sub-devices into a multi-dimensional tensor lane image after the lane image is input into the target detection model, wherein the multi-dimensional tensor includes a tensor corresponding to the input size dimension received by the image feature encoding layer of the target detection model, and a tensor corresponding to the number N dimensions, where N is an integer greater than 1; an extraction unit, used to extract the image features corresponding to the lane image of the multi-dimensional tensor through the image feature encoding layer of the target detection model after the lane image is input into the target detection model, to obtain a first feature map and a second feature map at different scales; and an upsampling unit, used to... After inputting the lane image into the target detection model, the first or second feature map of the first and second feature maps at different scales is upsampled to obtain the first and second feature maps of the same scale; the fusion unit is used to perform cumulative fusion processing on the first and second feature maps of the same scale after inputting the lane image into the target detection model to obtain the first target feature map; the downsampling unit is used to perform downsampling processing on the first target feature map after inputting the lane image into the target detection model to obtain the second target feature map; the fifth acquisition unit includes: a processing module, used to process the second target feature map through the subsequent model structure of the target detection model to obtain the detection result.
[0013] As an optional solution, the above processing module includes: a transformation submodule, used to perform frustum transformation processing on the second target feature map through the element frustum transformation layer of the above target detection model to obtain a top-view feature map, wherein the top-view feature map is used to represent the first positional relationship, and wherein the above subsequent model structure includes the element frustum transformation layer; and a processing submodule, used to perform subsequent processing on the top-view feature map through at least one target model structure to obtain the detection result, wherein the above subsequent model structure includes at least one target model structure.
[0014] As an optional solution, the above processing submodule includes: a summation subunit, used to sum the features falling at the same pixel position in the above top-view feature map through the view feature pooling layer of the above target detection model to obtain a pooled top-view feature map, wherein the above at least one target model structure includes the above view feature pooling layer; and a first processing subunit, used to perform subsequent processing on the above pooled top-view feature map through at least one first model structure to obtain the above detection result, wherein the above at least one target model structure includes the above at least one first model structure.
[0015] As an optional solution, the above processing submodule includes: an alignment subunit, used to perform temporal alignment processing on the top-view feature map at the current time and the top-view feature map at at least one historical time through the temporal feature fusion layer of the above target detection model, to obtain aligned top-view feature maps at different times, wherein the above at least one target model structure includes the above temporal feature fusion layer; a splicing subunit, used to splice the above top-view feature maps at different times through the above temporal feature fusion layer, to obtain a top-view feature map with fused temporal information; and a second processing subunit, used to perform subsequent processing on the above top-view feature map with fused temporal information through at least one second model structure to obtain the above detection result, wherein the above at least one target model structure includes the above at least one second model structure.
[0016] As an optional solution, the above-mentioned processing submodule includes: a third processing subunit, used to perform subsequent processing on the above-view feature map through the detection result output layer of the above-mentioned target detection model, to obtain the three-dimensional position, element size, element orientation angle, and element velocity of the above-mentioned target element in the above-mentioned first coordinate system, wherein the above-mentioned detection result includes the above-mentioned three-dimensional position, the above-mentioned element size, the above-mentioned element orientation angle, and the above-mentioned element velocity.
[0017] As an optional solution, the second acquisition unit includes: a projection module, used to project the at least one lane element from the first coordinate system to the second coordinate system using a projection matrix to obtain the projected position of the at least one lane element, wherein the second coordinate position includes the projected position, and the projection matrix represents the mapping relationship between the first coordinate system and the second coordinate system.
[0018] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the target detection method based on a top-down view as described above.
[0019] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described target detection method based on a top-down view through the computer program.
[0020] In this embodiment, a lane image is obtained by an image acquisition device capturing images of a lane scene from a top-down perspective. The lane image includes at least one lane element and a target element. The lane element is an element carrying attribute information in the lane scene, and the target element is an element to be detected in the lane image. The process involves acquiring the first coordinate position of the at least one lane element in a first coordinate system, the second coordinate position of the at least one lane element in a second coordinate system, and the third coordinate position of the target element in the second coordinate system. The attribute information includes the first coordinate position. The first coordinate system is the actual coordinate system corresponding to the lane scene, and the second coordinate system is the pixel coordinate system corresponding to the lane image. Based on a first positional relationship between the second and third coordinate positions, and the first coordinate position, a fourth coordinate position of the target element in the first coordinate system is obtained. If a fifth coordinate position of the image acquisition device in the first coordinate system is obtained, the detection result of the target element in the lane scene is obtained based on a second positional relationship between the fifth and fourth coordinate positions.
[0021] This embodiment establishes a precise coordinate mapping relationship by acquiring the coordinate positions of lane elements and target elements in different coordinate systems in a lane image. Furthermore, utilizing this coordinate transformation model, this embodiment can accurately convert the position of the target element in the pixel coordinate system (third coordinate position) to its position in the actual coordinate system (fourth coordinate position), effectively correcting image distortion and size changes caused by distance variations, resulting in a more accurate position of the target element.
[0022] Furthermore, this embodiment also considers the position of the image acquisition device in the actual coordinate system (the fifth coordinate position). By combining the position information of the image acquisition device and the position of the target element in the actual coordinate system, the actual position of the target element in the lane scene can be calculated more accurately, thereby achieving the goal of obtaining more accurate detection results. This improves the accuracy of target detection based on a top-down view and solves the technical problem of low accuracy in target detection based on a top-down view. Attached Figure Description
[0023] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0024] Figure 1 This is a schematic diagram of an application environment for an optional target detection method based on a top-down view according to an embodiment of this application;
[0025] Figure 2 This is a schematic diagram of the flow of an optional target detection method based on a top-down view according to an embodiment of this application;
[0026] Figure 3 This is a schematic diagram of an optional target detection method based on a top-down view according to an embodiment of this application;
[0027] Figure 4 This is a schematic diagram of another optional target detection method based on a top-down view according to an embodiment of this application;
[0028] Figure 5 This is a schematic diagram of another optional target detection method based on a top-down view according to an embodiment of this application;
[0029] Figure 6 This is a schematic diagram of another optional target detection method based on a top-down view according to an embodiment of this application;
[0030] Figure 7 This is a schematic diagram of another optional target detection method based on a top-down view according to an embodiment of this application;
[0031] Figure 8 This is a schematic diagram of an optional target detection device based on a top-down view according to an embodiment of this application;
[0032] Figure 9 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application. Detailed Implementation
[0033] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0035] According to one aspect of the embodiments of this application, a target detection method based on a top-down view is provided. Optionally, as an optional implementation, the above-described target detection method based on a top-down view can be applied to, but is not limited to, [examples of other methods]. Figure 1 The environment shown may include, but is not limited to, user equipment 102 and server 112. User equipment 102 may include, but is not limited to, a display 104, a processor 106 and a memory 108. Server 112 includes a database 114 and a processing engine 116.
[0036] The specific process can be summarized in the following steps:
[0037] Step S102: User equipment 102 acquires a lane image obtained by the image acquisition device from a top-down perspective after acquiring the lane scene.
[0038] Step S104: Send the lane image to server 112 via network 110;
[0039] In steps S106-S110, server 112 obtains the first coordinate position of at least one lane element in the first coordinate system, the second coordinate position of at least one lane element in the second coordinate system, and the third coordinate position of the target element in the second coordinate system through processing engine 116. Further, based on the first positional relationship between the second and third coordinate positions and the first coordinate position, server 112 obtains the fourth coordinate position of the target element in the first coordinate system. If the fifth coordinate position of the image acquisition device in the first coordinate system is obtained, the detection result of the target element in the lane scene is obtained based on the second positional relationship between the fifth and fourth coordinate positions.
[0040] In step S112, the detection results are sent to user equipment 102 via network 110. User equipment 102 displays the detection results on display 104 via processor 106 and stores the detection results in memory 108.
[0041] remove Figure 1 Beyond the examples shown, the terminal devices described above can be terminal devices configured with a target client, including but not limited to at least one of the following: mobile phones (such as Android phones, iOS phones, etc.), laptops, tablets, PDAs, MIDs (Mobile Internet Devices), PADs, desktop computers, smart TVs, etc. The target client can be a video client, instant messaging client, browser client, educational client, etc. The networks described above can include, but are not limited to, wired networks and wireless networks. The wired networks include local area networks (LANs), metropolitan area networks (MANs), and wide area networks (WANs). The wireless networks include Bluetooth, Wi-Fi, and other networks that enable wireless communication. The server described above can be a single server, a server cluster consisting of multiple servers, or a cloud server. The above is merely an example, and no limitations are imposed in this embodiment.
[0042] Alternatively, as an alternative implementation method, such as Figure 2 As shown, the target detection method based on a top-down view can be executed by an electronic device, such as... Figure 1 The user equipment or server shown includes the following specific steps:
[0043] S202, acquire a lane image obtained by the image acquisition device from a top-down perspective after acquiring the lane scene. The lane image contains at least one lane element and a target element. The lane element is an element in the lane scene that carries attribute information, and the target element is an element in the lane image that is expected to be detected.
[0044] S204, obtain the first coordinate position of at least one lane element in the first coordinate system, obtain the second coordinate position of at least one lane element in the second coordinate system, and obtain the third coordinate position of the target element in the second coordinate system, wherein the attribute information includes the first coordinate position, the first coordinate system is the actual coordinate system corresponding to the lane scene, and the second coordinate system is the pixel coordinate system corresponding to the lane image;
[0045] S206, Based on the first positional relationship between the second and third coordinate positions, and the first coordinate position, obtain the fourth coordinate position of the target element in the first coordinate system;
[0046] S208, after obtaining the fifth coordinate position of the image acquisition device in the first coordinate system, the detection result of the target element in the lane scene is obtained according to the second positional relationship between the fifth coordinate position and the fourth coordinate position.
[0047] Optionally, in this embodiment, the above-described target detection method based on a top-down perspective can be applied to multiple scenarios, such as in an urban intelligent traffic detection system, where image acquisition devices (such as high-definition cameras) are installed at high locations at traffic intersections or important road sections to capture lane scenes from a top-down perspective. The system first acquires lane images containing lane elements (such as lane lines, traffic signs, etc.) and target elements (such as vehicles, pedestrians, etc.).
[0048] By identifying and locating lane elements in both the physical coordinate system and the pixel coordinate system, the system establishes a coordinate mapping relationship. When a target element (such as a vehicle violating traffic rules) appears, the system can quickly calculate its position in the pixel coordinate system and, using the previously established coordinate mapping relationship, convert it into its position in the physical coordinate system.
[0049] By combining the actual installation location of the image acquisition equipment, the system can accurately determine the specific location of the target element in the lane scene, thereby triggering corresponding traffic management measures, such as sending alarms and recording violations.
[0050] For example, in autonomous vehicles, onboard cameras capture the surrounding lane scene from a top-down perspective. By processing these images, the vehicle needs to accurately perceive its surroundings, including lane lines, traffic signs, obstacles, and other vehicles.
[0051] Using the aforementioned target detection method based on a top-down perspective, the perception system of an autonomous vehicle can first identify the positions of lane elements (such as lane lines) in the pixel coordinate system and the actual coordinate system, establishing an accurate coordinate mapping model. When a target element (such as a vehicle or pedestrian in front) appears, the system can quickly calculate its position in the pixel coordinate system.
[0052] Through the coordinate mapping model, the perception system can convert the position of the target element in the pixel coordinate system into the position in the actual coordinate system, thereby accurately determining the relative distance and orientation between the target element and the autonomous vehicle.
[0053] For example, in a drone patrol system, the drone's high-definition camera captures images of the ground from a top-down perspective. The system needs to accurately identify and locate target elements on the ground, such as buildings, vehicles, and people.
[0054] By acquiring the positional information of lane elements (such as ground markers, road boundaries, etc.) in both the pixel coordinate system and the actual coordinate system, the system can establish a precise coordinate mapping relationship. When the UAV detects a target element during its cruise, the system can quickly calculate its position in the pixel coordinate system.
[0055] By using coordinate mapping, the system can convert the position of the target element in the pixel coordinate system to its position in the actual coordinate system, thereby accurately determining the specific location of the target element on the ground.
[0056] Optionally, in this embodiment, the image acquisition device refers to a device used to capture and record images, such as a camera or video camera. These devices are capable of converting optical information in a scene into digital images.
[0057] Optionally, in this embodiment, a top-down view can refer to a perspective looking down from a high place, like a person standing at a high place looking down at a scene below. In fields such as traffic detection, drone photography, or autonomous driving, a top-down view can provide a wider field of view and more comprehensive scene information.
[0058] Optionally, in this embodiment, the lane scene may refer to the environment or background related to road traffic, including the overall picture composed of elements such as lanes, traffic signs, surrounding buildings, and other vehicles.
[0059] Optionally, in this embodiment, lane images may refer to images captured and recorded from a lane scene by an image acquisition device. These images show the actual situation of the lane and its surroundings.
[0060] Optionally, in this embodiment, lane elements can refer to those inherent elements in the lane scene that have certain quantifiable information (such as position, size, etc.). For example, lane lines, traffic signs, and curbs are all lane elements.
[0061] Optionally, in this embodiment, the target element can be an element that is of particular interest and that we want to detect in the lane image, such as a moving vehicle or a pedestrian. These elements can be dynamic or static.
[0062] Optionally, in this embodiment, attribute information can refer to various types of data describing the specific characteristics and properties of lane elements. This data includes, but is not limited to, the shape, size, color, position, and type of the lane element. For example, the attribute information of a lane line may include whether it is a solid or dashed line, its color (such as white or yellow), and its specific position in the image.
[0063] To further illustrate, consider an example at a city intersection where a high-definition camera captures an image of the entire traffic lane from above. This image shows white lane lines, a red stop sign, and a pedestrian crossing the road. In this embodiment, the white lane lines and the red stop sign are lane elements, carrying attribute information such as color, shape, and position, while the target element can be the pedestrian.
[0064] Optionally, in this embodiment, the first coordinate system may refer to the actual coordinate system corresponding to the lane scene, also known as the world coordinate system or physical coordinate system. In this coordinate system, the position information corresponds to the position in the actual physical space, and the unit may be a meter or other unit of length.
[0065] Optionally, in this embodiment, the first coordinate position may refer to the position of the lane element in the first coordinate system (the actual coordinate system). This position reflects the precise location of the lane element in the real world.
[0066] Optionally, in this embodiment, the second coordinate system may refer to the pixel coordinate system corresponding to the lane image. In this coordinate system, the position information may be defined based on the pixel grid of the image, and the unit may be pixels.
[0067] Optionally, in this embodiment, the second coordinate position may refer to the position of the lane element in the second coordinate system (pixel coordinate system). This position reflects the pixel coordinates of the lane element in the image.
[0068] Optionally, in this embodiment, the third coordinate position may refer to the position of the target element in the second coordinate system (pixel coordinate system). Similar to the second coordinate position, it may also be defined based on the pixel grid of the image, but it applies to the target element rather than the lane element.
[0069] To further illustrate, consider an optional assumption: a lane image captured by a traffic detection camera. In this image, lane lines (lane elements) and a car (target element) can be identified. Using image processing techniques, the pixel positions of these elements (second and third coordinate positions) can be determined in a pixel coordinate system (second coordinate system). Furthermore, given the camera's parameters and installation location, these pixel positions can be further converted to their positions in a real-world coordinate system (first coordinate system), thus revealing the exact location of these elements in the real world.
[0070] Optionally, in this embodiment, the first positional relationship may refer to the relative positional relationship between the lane element (represented by the second coordinate position) and the target element (represented by the third coordinate position) in the second coordinate system (pixel coordinate system). This relationship may include direction, distance, or other spatial relativity descriptions, reflecting the positioning of the target element relative to the lane element in the image.
[0071] Optionally, in this embodiment, the fourth coordinate position may refer to the position of the target element in the first coordinate system (actual coordinate system). This position is derived by combining the first positional relationship between the second and third coordinate positions, and the first coordinate position of the lane element in the first coordinate system. In other words, it converts the position (pixel coordinates) of the target element in the image into its position (actual coordinates) in the real world.
[0072] To further illustrate, consider an optional assumption: in a traffic detection image, this embodiment identifies a lane line (lane element) and a car (target element). The position of the lane line in the second coordinate system is known (second coordinate position), and the position of the car is also known (third coordinate position). By analyzing the image, the positional relationship of the car relative to the lane line is determined (i.e., the first positional relationship), for example, an image representation of the car being 5 meters north of the lane line. Then, using the position of the lane line in the real coordinate system (first coordinate position) and this relative positional relationship, this embodiment can calculate the exact position of the car in the real world, i.e., the fourth coordinate position.
[0073] Optionally, in this embodiment, the fifth coordinate position may refer to the position of the image acquisition device (such as a camera) in the first coordinate system (the actual coordinate system). For example, this position can be used to describe the installation location of the camera in the real world and is a reference point for determining the actual position of elements in the image.
[0074] Optionally, in this embodiment, the second positional relationship may refer to the relative positional relationship between the image acquisition device (represented by the fifth coordinate position) and the target element (represented by the fourth coordinate position) in the first coordinate system. This relationship helps to understand and analyze the specific position and state of the target element in the lane scene.
[0075] Optionally, in this embodiment, the detection result may refer to the specific location, state, or other relevant information of the target element in the lane scene, obtained by analyzing and processing the lane image captured by the image acquisition device. This information may be derived based on image processing and coordinate transformation techniques and is used to describe and explain the situation of the target element in the real world.
[0076] To further illustrate, consider an optional assumption: a traffic detection system has a camera mounted above an intersection, and the camera's actual installation location (the fifth coordinate position) is known. The system processes the images captured by the camera, identifies a car running a red light (the target element), and calculates the car's position in the actual coordinate system (the fourth coordinate position). By analyzing the relative positional relationship between the camera and the car (the second positional relationship), the system can accurately detect the car's specific location and behavior within the lane, such as whether it violated traffic rules. This detection result can be used for traffic management, law enforcement, or data analysis purposes.
[0077] It should be noted that this embodiment describes an image processing and coordinate transformation process, which aims to accurately detect and locate target elements from a top-down view of a lane image.
[0078] To further illustrate, optional examples include... Figure 3 As shown, the image acquisition device 302 acquires images of the lane scene 304 from a top-down perspective to obtain a lane image 306. The lane image 306 contains at least one lane element (such as lane element A1, lane element A2, lane element A3) and a target element (such as target element B1, target element B2). The lane element is then used as the basis for detecting the target element to obtain a detection result 308.
[0079] Specifically, the system first obtains the first coordinate positions of lane elements A1, A2, and A3 in the first coordinate system; then obtains their second coordinate positions in the second coordinate system; and finally obtains the third coordinate positions of target elements B1 and B2 in the second coordinate system. The attribute information includes the first coordinate position, where the first coordinate system is the actual coordinate system corresponding to lane scene 304, and the second coordinate system is the pixel coordinate system corresponding to lane image 306. Next, based on the first positional relationship between the second and third coordinate positions, and the first coordinate position, the system obtains the fourth coordinate position of target elements B1 and B2 in the first coordinate system. Further, after obtaining the fifth coordinate position of the image acquisition device 302 in the first coordinate system, the system obtains the detection result 308 of target elements B1 and B2 in the lane scene based on the second positional relationship between the fifth and fourth coordinate positions.
[0080] The embodiments provided in this application obtain the coordinate positions of lane elements and target elements in lane images under different coordinate systems, establishing a precise coordinate mapping relationship. Furthermore, utilizing this coordinate transformation model, this embodiment can accurately convert the position of the target element in the pixel coordinate system (third coordinate position) to its position in the actual coordinate system (fourth coordinate position), effectively correcting image distortion and size changes caused by distance variations, resulting in a more accurate position of the target element.
[0081] Furthermore, this embodiment also considers the position of the image acquisition device in the actual coordinate system (the fifth coordinate position). By combining the position information of the image acquisition device and the position of the target element in the actual coordinate system, the actual position of the target element in the lane scene can be calculated more accurately, thereby achieving the goal of obtaining more accurate detection results and thus realizing the technical effect of improving the accuracy of target detection based on the top-down view.
[0082] As an optional approach, based on the first positional relationship between the second and third coordinate positions, and the first coordinate position, the fourth coordinate position of the target element in the first coordinate system is obtained, including:
[0083] S1-1, determine at least one first lane element and at least one second lane element from at least one lane element, wherein the first lane element is in the second coordinate system and the first positional relationship satisfies the requirement that the lane element is adjacent to the target element in the first direction, and the second lane element is in the second coordinate system and the first positional relationship satisfies the requirement that the lane element is adjacent to the target element in the second direction.
[0084] S1-2, based on at least one first lane element and at least one second lane element, obtain the first coordinate of the target element in the first coordinate system and located in the first direction, and the second coordinate of the target element in the first coordinate system and located in the second direction, wherein the fourth coordinate position includes the first coordinate and the second coordinate.
[0085] Optionally, in the second coordinate system (pixel coordinate system), there exists at least one lane element (referred to as the first lane element) that is adjacent to the target element in a specific direction (referred to as the first direction, such as the horizontal or vertical direction). This adjacency relationship is determined by analyzing the positional data in the second coordinate system. The first positional relationship in the first direction refers to the relative positional relationship between the lane element and the target element in the pixel coordinate system, and this relationship indicates that they are closely adjacent in a specific direction (such as the horizontal direction).
[0086] In the second coordinate system, at least one lane element (referred to as the second lane element) is adjacent to the target element in another specific direction (referred to as the second direction; if the first direction is horizontal, the second direction can be vertical). The first positional relationship in the second direction describes the relative positional relationship between the lane element and the target element in another specific direction in the pixel coordinate system, indicating that they are adjacent in this direction.
[0087] Optionally, the target element's position coordinates in the actual coordinate system (first coordinate system) along a first direction (e.g., horizontal). These coordinates are obtained by analyzing the positional relationship between the target element and adjacent lane elements in the second coordinate system, combined with the position of the lane elements in the first coordinate system, and then performing coordinate transformation. The first coordinate is the coordinate value of the target element in the actual coordinate system along the first direction (e.g., horizontal or lateral).
[0088] The target element's position coordinates in the actual coordinate system are along the second direction (such as the vertical direction). These coordinates are also determined by analyzing the positional relationship in the second coordinate system and performing coordinate transformation. The second coordinate is the coordinate value of the target element in the actual coordinate system along the second direction (such as vertical or longitudinal).
[0089] It should be noted that this embodiment describes a process in which the position of the target element in the first coordinate system is calculated by analyzing and utilizing the relative positional relationship between the lane element and the target element in the second coordinate system (pixel coordinate system) and the position of the lane element in the first coordinate system (actual coordinate system). This process involves identifying lane elements adjacent to the target element in a specific direction from multiple lane elements and using the positional information of these adjacent elements to calculate the coordinates of the target element in the actual scene.
[0090] To illustrate further, consider the optional assumption that a target vehicle has been identified in the pixel coordinate system of a lane image, and that the vehicle's relative position to the left and right lane lines has been obtained. By finding the positions of these lane lines in the real coordinate system, these relative positions can be used to estimate the approximate position of the target vehicle in the real coordinate system.
[0091] The embodiments provided in this application enable more precise location of target elements in real-world scenarios, thereby providing support for accurate navigation, obstacle avoidance, and decision-making required in scenarios such as intelligent transportation systems and autonomous driving technologies.
[0092] As an optional approach, based on at least one first lane element and at least one second lane element, obtaining the first coordinates of the target element in a first coordinate system and located in a first direction, and the second coordinates of the target element in a first coordinate system and located in a second direction, includes:
[0093] S2-1, obtain the first pixel coordinate of the target element in the second coordinate system and located in the first direction, the first actual coordinate of the first element in the first coordinate system and located in the first direction of at least one first lane element, the second actual coordinate of the second element in the first coordinate system and located in the first direction of at least one first lane element, the second pixel coordinate of the first element in the second coordinate system and located in the first direction of the first direction, and the third pixel coordinate of the second element in the second coordinate system and located in the first direction of the second element.
[0094] S2-2, obtain the first difference between the first pixel coordinate and the first actual coordinate, the second difference between the second actual coordinate and the first actual coordinate, and the third difference between the third pixel coordinate and the second pixel coordinate;
[0095] S2-3, obtain the first product value of the product of the first difference and the second difference;
[0096] S2-4, obtain the first ratio of the first product value to the third difference value;
[0097] S2-5, obtain the first sum of the first actual coordinate and the first ratio, and determine the first sum as the first coordinate; and,
[0098] S2-6, obtain the fourth pixel coordinate of the target element in the second coordinate system and located in the second direction, the third actual coordinate of the third element in the first coordinate system and located in the second direction of at least one second lane element, the fourth actual coordinate of the fourth element in the first coordinate system and located in the second direction of at least one second lane element, the fifth pixel coordinate of the third element in the second coordinate system and located in the second direction of the second direction, and the sixth pixel coordinate of the fourth element in the second coordinate system and located in the second direction of the second direction;
[0099] S2-7, obtain the fourth difference between the fourth pixel coordinate and the third actual coordinate, the fifth difference between the fourth actual coordinate and the third actual coordinate, and the sixth difference between the third pixel coordinate and the second pixel coordinate.
[0100] S2-8, obtain the second product value of the product of the fourth difference and the fifth difference;
[0101] S2-9, obtain the second ratio of the second product value to the sixth difference value;
[0102] S2-10, obtain the second sum of the third actual coordinate and the second ratio, and determine the second sum as the second coordinate.
[0103] It should be noted that this embodiment describes how to determine the position (first coordinate and second coordinate) of the target element in the actual coordinate system by using the position information of lane elements (first lane element and second lane element) in the pixel coordinate system (second coordinate system) and the actual coordinate system (first coordinate system) through a series of calculations and transformations. This process involves calculating the difference, product, ratio, and sum between the pixel coordinates and the actual coordinates.
[0104] To further illustrate, let's assume the target element (e.g., a car) has its first pixel coordinate in the first direction (e.g., horizontal) of the pixel coordinate system as Px, while the first and second actual coordinates of two adjacent lane elements (the first lane element) in the actual coordinate system in the corresponding first direction are Rx1 and Rx2, respectively. By calculating the differences between Px and Rx1, Rx2 and Rx1, and their corresponding values in the pixel coordinate system, a ratio can be obtained. This ratio, along with Rx1 (or Rx2), can then be used to estimate the first coordinate of the target element in the actual coordinate system.
[0105] The embodiments provided in this application enable more accurate determination of the position of target elements in real-world scenarios, which not only improves the accuracy of target detection but also provides reliable position information for subsequent decision-making and control.
[0106] As an optional approach, after obtaining the first coordinates of the target element in a first coordinate system and located in a first direction, and the second coordinates of the target element in a first coordinate system and located in a second direction, based on at least one first lane element and at least one second lane element, the method further includes:
[0107] The projection information between the first coordinate and the second coordinate is stored in the pixel projection mapping table. The projection information between the first coordinate and the second coordinate is used to indicate whether to project the target element from the first coordinate system to the second coordinate system or from the second coordinate system to the first coordinate system. The pixel projection mapping table is used to store and query the projection information corresponding to each element.
[0108] Optionally, in this embodiment, projection information can refer to data or parameters describing the transformation relationship of an element (such as a target element) from one coordinate system (such as the first coordinate system, i.e., the actual coordinate system) to another coordinate system (such as the second coordinate system, i.e., the pixel coordinate system). In short, it provides a method for mapping points in one coordinate system to another. This projection information may include coordinate transformation formulas, transformation matrices, scaling factors, offsets, etc., all of which are crucial data necessary for coordinate transformation.
[0109] Optionally, in this embodiment, the pixel projection mapping table can be a data structure used to store and query the projection information of various elements (such as lane elements and target elements) between different coordinate systems. This mapping table can be viewed as a database or lookup table, recording the correspondence between elements in the actual coordinate system and the pixel coordinate system. By querying this mapping table, the transformation relationship of an element between the two coordinate systems can be obtained quickly and accurately, thereby achieving rapid coordinate system transformation and accurate element positioning.
[0110] It should be noted that images of the lane scene are acquired using an image acquisition device, and lane elements and target elements are identified from them. The positional relationships of these elements in different coordinate systems (actual coordinate system and pixel coordinate system) are used to determine the position of the target element in the actual coordinate system. After determining the first coordinate (located in the first direction) and the second coordinate (located in the second direction) of the target element in the actual coordinate system (first coordinate system), this embodiment also includes storing the projection information between these coordinates in a pixel projection mapping table. This mapping table is essentially a database or data structure used to store and query the projection correspondence of each element between the two coordinate systems.
[0111] The pixel projection map can be used not only for dynamic targets such as vehicles, but also for static lane elements, such as lane lines and traffic signs. By continuously updating this map, the system can detect and transform the position of all elements in the image in real time. Furthermore, this map can be used to calibrate the viewing angle and position of the image acquisition device, thereby improving the accuracy of lane and vehicle detection.
[0112] To further illustrate, consider an optional assumption of a lane image containing a vehicle (target element) and lane lines (lane elements). After determining the vehicle's position in the actual coordinate system (i.e., the first and second coordinates), the projection relationship between these coordinates and the vehicle's position in the pixel coordinate system can be stored in a pixel projection mapping table. Thus, if this embodiment needs to transform the vehicle's position between the two coordinate systems, this mapping table can be directly consulted without recalculation.
[0113] Through the embodiments provided in this application, since the projection relationships have been pre-calculated and stored, they can be directly queried when needed without the need for complex coordinate transformation calculations. By storing accurate projection relationships, errors that may occur during coordinate system transformation can be reduced. Furthermore, the mapping table can be dynamically updated to adapt to changes in the position or viewing angle of the image acquisition device, thereby ensuring the stability and accuracy of the system.
[0114] As an optional approach, after storing the projection information between the first and second coordinates in a pixel projection mapping table, the method further includes:
[0115] S3-1, Obtain a projection query request triggered by the target range of the lane image, wherein the projection query request is used to request the target position of the target element in the lane image located within the target range in the first coordinate system.
[0116] S3-2, Obtain multiple target elements located within the target range in the lane image;
[0117] S3-3, using the multiple coordinate positions of multiple candidate elements in the second coordinate system, query the pixel projection mapping table to find the multiple candidate coordinate positions of multiple target elements in the first coordinate system;
[0118] S3-4: Filter the coordinate positions that are outside the predetermined detection range from the multiple candidate coordinate positions to obtain at least one target position.
[0119] Optionally, in this embodiment, the target range may refer to a specific area in the lane image that the user or system is interested in. This area can be defined based on actual needs or points of interest; for example, it could be a traffic congestion area, an accident scene, or a road segment that requires special detection in the lane image. Target elements (such as vehicles, pedestrians, etc.) within this range will receive special attention and processing.
[0120] Optionally, in this embodiment, the predetermined detection range can be a valid or acceptable detection area defined in the actual coordinate system (first coordinate system). This range can be set based on the actual application scenario and safety considerations to ensure that the detected target element position is valid and reliable. For example, in an autonomous driving scenario, the predetermined detection range can be a safe and drivable area on the vehicle's driving road to exclude target elements from roadside obstacles or non-driving areas.
[0121] It should be noted that this embodiment describes a method for acquiring lane images through an image acquisition device, determining the positional relationship between lane elements and target elements from the images, and then determining the position of the target elements in the actual coordinate system by querying a pixel projection mapping table.
[0122] To further illustrate, consider a possible scenario in a highway detection system where a "target range" is defined, covering the area near an exit on the highway. The system will specifically detect all vehicles (target elements) within this area. Simultaneously, this embodiment defines a "predetermined detection range," including only the highway lanes to exclude roadside guardrails, signs, etc. When the system detects vehicles within the target range, it determines their positions in the actual coordinate system by querying a pixel projection mapping table and filters out position information outside the predetermined detection range.
[0123] The embodiments provided in this application allow for focusing on a specific area (target range), enabling concentrated resource processing of critical information and reducing unnecessary computation and data processing. By setting an effective detection range (predetermined detection range), invalid or interfering information can be eliminated, improving the accuracy of target element location detection. Furthermore, the target range can be dynamically adjusted according to actual needs, making the system more flexible and scalable, adapting to different scenarios and application requirements.
[0124] As an optional approach, after acquiring the lane image obtained by the image acquisition device from a top-down perspective of the lane scene, the method further includes:
[0125] S4-1, Input the lane image into the target detection model, wherein the target detection model is a neural network model trained from multiple image samples for detecting target elements;
[0126] S4-2, Obtain the detection results output by the target detection model.
[0127] Optionally, in this embodiment, the object detection model is a neural network model trained using machine learning or deep learning techniques, used to automatically identify and locate specific target elements in images or videos. This model is typically able to identify the category of the target and provide its specific location in the image, for example, represented by a bounding box.
[0128] It should be noted that after acquiring the lane image obtained by the image acquisition device from the top-down view of the lane scene, this embodiment further includes two key steps: first, the lane image is input into a neural network model called the "object detection model", and then the detection result output by this model is obtained.
[0129] To further illustrate, for example, suppose the object detection model in this embodiment is trained to detect vehicles in a lane image. After acquiring the lane image, this embodiment inputs the image into the model. The model analyzes the image content and outputs a series of vehicle detection results, including the position of each vehicle in the image (usually represented by a rectangle) and the corresponding confidence score.
[0130] The embodiments provided in this application introduce a target detection model that can automatically and accurately identify and locate target elements in lane images. This not only greatly reduces the workload of manual analysis but also improves the accuracy and efficiency of detection.
[0131] As an optional solution, the image acquisition device includes N acquisition sub-devices, characterized in that, after inputting the lane image into the target detection model, the method further includes:
[0132] S5-1 integrates the N lane images collected by the N acquisition sub-devices into a multidimensional tensor lane image. The multidimensional tensor includes a tensor corresponding to the dimension of the input size accepted by the image feature encoding layer of the target detection model, and a tensor corresponding to the dimension of the quantity N, where N is an integer greater than 1.
[0133] S5-2, through the image feature encoding layer of the target detection model, extracts the image features corresponding to the lane image of the multidimensional tensor to obtain the first feature map and the second feature map at different scales;
[0134] S5-3, upsampling is performed on the first or second feature maps of the first and second feature maps at different scales to obtain the first and second feature maps of the same scale.
[0135] S5-4, perform cumulative fusion processing on the first feature map and the second feature map of the same scale to obtain the first target feature map;
[0136] S5-5, perform downsampling on the first target feature map to obtain the second target feature map;
[0137] Obtaining the detection results output by the target detection model includes: processing the second target feature map through the subsequent model structure of the target detection model to obtain the detection results.
[0138] Optionally, in this embodiment, N acquisition sub-devices can refer to an image acquisition device composed of multiple (N) independent acquisition units or cameras. These acquisition sub-devices can capture images of the lane scene from different angles or positions, thereby providing more comprehensive and multi-angle visual information. For example, in a traffic detection system, multiple cameras can be used to simultaneously detect different road segments or angles.
[0139] Optionally, the multidimensional tensor can be a high-dimensional array capable of storing complex data structures. In this embodiment, the multidimensional tensor is used to integrate N lane images captured by N acquisition sub-devices. This tensor includes not only image data (pixel values), but also a dimension corresponding to the input size of the object detection model, and a dimension (N) representing the number of images. In other words, it is a high-dimensional data structure capable of storing multiple image data simultaneously.
[0140] Optionally, in this embodiment, feature maps of different scales can refer to image features extracted through operations such as convolution in a neural network, which are represented in different resolutions or sizes. These feature maps capture details and information at different levels in the image. "First feature map" and "second feature map" can refer to feature maps extracted from different network layers or stages, which have different sizes and receptive fields, thus enabling them to capture different features in the image.
[0141] Optionally, in this embodiment, upsampling can be an image processing technique used to increase the resolution of the feature map. In neural networks, upsampling can be achieved through interpolation (such as bilinear interpolation, nearest neighbor interpolation, etc.) or transposed convolution (also known as deconvolution). In this context, upsampling is used to resize the first and second feature maps of different scales to the same size for subsequent feature fusion.
[0142] Optionally, in this embodiment, after upsampling, the first and second feature maps, which were originally at different scales, are adjusted to the same resolution or size. This allows them to be spatially aligned, facilitating subsequent feature fusion operations.
[0143] Optionally, in this embodiment, the cumulative fusion process can refer to the element-wise addition of the first and second feature maps of the same scale. Through cumulative fusion, complementary information from different feature maps can be integrated to obtain a richer and more comprehensive feature representation. This fusion method helps improve the accuracy and robustness of object detection.
[0144] Optionally, downsampling can be a process of reducing the resolution of the feature map to reduce the dimensionality of the data and computational complexity. In neural networks, downsampling can be achieved through pooling (such as max pooling, average pooling, etc.) or convolution operations (using convolutional kernels with a stride greater than 1). In this embodiment, downsampling is used to reduce the size of the first target feature map and generate a second target feature map to further reduce the computational burden and improve processing efficiency.
[0145] It should be noted that, in the description of this embodiment, the image acquisition device is designed to include N acquisition sub-devices, which means that images of the lane scene can be captured simultaneously from multiple different perspectives or positions. These acquired images are then integrated into a multidimensional tensor lane image. A multidimensional tensor here refers to a data structure capable of accommodating high-dimensional data, including a tensor corresponding to the dimension of the input size accepted by the image feature encoding layer of the object detection model, and a tensor corresponding to the number N dimensions.
[0146] To further illustrate, consider an optional assumption: an advanced traffic detection system equipped with four high-definition cameras as acquisition sub-devices (N=4). These cameras are mounted at different locations on the road to capture a more comprehensive view of the lanes. The system first integrates the images captured by these four cameras into a multidimensional tensor of the lane image. This multidimensional tensor contains the image data captured by each camera and is formatted according to the input size accepted by the model.
[0147] The multidimensional tensor lane image is then processed through the image feature encoding layer of the object detection model. This layer extracts features from the image, generating first and second feature maps at different scales. These feature maps capture different levels of detail and information in the image. For further processing and analysis, the system upsamples these feature maps to ensure they are at the same scale. Then, by accumulating and fusing feature maps of the same scale, a more comprehensive and richer first target feature map is obtained. To reduce data dimensionality and improve processing efficiency, the first target feature map is downsampled to generate the second target feature map.
[0148] The embodiments provided in this application enable the effective utilization of image data captured by multiple acquisition sub-devices to extract more comprehensive and accurate image features. These features not only contain information from a single image but also fuse data from multiple perspectives, thereby improving the accuracy and robustness of target detection.
[0149] As an optional approach, the second target feature map is processed through the subsequent model structure of the target detection model to obtain the detection result, including:
[0150] S6-1, The second target feature map is processed by the element-view frustum transformation layer of the target detection model to obtain the top view feature map, wherein the top view feature map is used to represent the first positional relationship, and the subsequent model structure includes the element-view frustum transformation layer;
[0151] S6-2, using at least one target model structure, performs subsequent processing on the top-view feature map to obtain the detection result, wherein the subsequent model structure includes at least one target model structure.
[0152] Optionally, in this embodiment, the top-view feature map can be a feature map obtained after processing by the element view frustum transformation layer, which displays the lane scene from a top-view perspective and is used to understand and analyze the positional relationships of elements in the lane scene.
[0153] It should be noted that, in the description of this embodiment, it is mentioned that the second target feature map is processed through the subsequent model structure of the target detection model to obtain the detection result. This process specifically includes two main steps: First, the second target feature map is subjected to view frustum transformation processing through the element view frustum transformation layer of the target detection model to generate a top-view feature map; second, the top-view feature map is further processed through at least one target model structure to finally obtain the detection result.
[0154] To further illustrate, consider an advanced traffic detection system that uses a target detection model comprising an element-view frustum transformation layer and a target model structure. After capturing a lane image, the system first processes the data to obtain a second target feature map. This feature map is then passed to the element-view frustum transformation layer and transformed into a top-down feature map. This top-down feature map acts like a "map" overlooking the entire lane scene from above, allowing the system to more clearly see the positional relationships between vehicles, pedestrians, and other elements. Finally, the target model structure further analyzes this "map" to accurately identify and locate target elements, such as illegally parked vehicles or pedestrians running red lights.
[0155] The embodiments provided in this application enable the object detection model to utilize information from lane images more effectively. The use of an element-based view frustum transformation layer allows the model to understand the scene from different perspectives, which helps improve the accuracy and robustness of detection. Furthermore, by performing subsequent processing on the top-view feature map using at least one object model structure, the detection results can be further refined and optimized.
[0156] As an optional approach, the top-view feature map is processed using at least one target model structure to obtain detection results, including:
[0157] S7-1, through the view feature pooling layer of the target detection model, the features falling at the same pixel position in the top view feature map are summed to obtain the pooled top view feature map, wherein at least one target model structure includes a view feature pooling layer.
[0158] S7-2, using at least one first model structure, performs subsequent processing on the pooled top-view feature map to obtain the detection result, wherein at least one target model structure includes at least one first model structure.
[0159] Optionally, in this embodiment, the pooled top-view feature map is a top-view feature map processed by the view feature pooling layer, and its features are aggregated at each pixel position, which helps the model to better understand and identify target elements in the image.
[0160] It should be noted that this embodiment illustrates how to process the top-view feature map through the subsequent model structure of the object detection model to obtain the final detection result. This involves two key steps: first, pooling the top-view feature map through the view feature pooling layer of the object detection model; second, performing subsequent processing on the pooled top-view feature map through at least one first model structure.
[0161] To further illustrate, consider the optional assumption that the object detection model in this embodiment is used to identify vehicles and pedestrians on the road. During vehicle movement, the image acquisition device continuously captures lane images and inputs these images into the object detection model. The model first converts the feature map into a top-view feature map using an element-based view frustum transformation layer to better understand the positional relationships between road elements. Next, a view feature pooling layer pools the top-view feature map, summing features falling at the same pixel location to simplify the feature map's complexity. Finally, the first model structure performs in-depth analysis of the pooled top-view feature map to accurately identify vehicles and pedestrians on the road.
[0162] Through the embodiments provided in this application, the object detection model can more effectively process and analyze information in lane images. The application of the viewpoint feature pooling layer not only reduces the dimensionality of the feature map but also improves the computational efficiency of the model. Furthermore, by performing subsequent processing on the pooled top-view feature map using the first model structure, the detection results can be further refined and optimized.
[0163] As an optional approach, the top-view feature map is processed using at least one target model structure to obtain detection results, including:
[0164] S8-1, through the temporal feature fusion layer of the target detection model, the top view feature map at the current time is temporally aligned with the top view feature map at least one historical time to obtain aligned top view feature maps at different times, wherein at least one target model structure includes a temporal feature fusion layer.
[0165] S8-2, through the temporal feature fusion layer, stitches together the top-view feature maps at different times to obtain a top-view feature map that incorporates temporal information;
[0166] S8-3, by using at least one second model structure, the top-view feature map with fused time-series information is further processed to obtain the detection result, wherein at least one target model structure includes at least one second model structure.
[0167] Optionally, in this embodiment, the aligned top-view feature maps at different times can refer to top-view feature maps from different time points after time-series alignment processing. These feature maps are consistent in the time dimension, which facilitates subsequent time-series information fusion.
[0168] Optionally, in this embodiment, the top-view feature map that incorporates time-series information can be a feature map obtained by stitching together top-view feature maps from different times. It incorporates information from multiple time points, which helps the model understand dynamically changing lane scenarios.
[0169] It should be noted that this embodiment illustrates how to further process the top-view feature map to obtain the detection result through the subsequent model structure of the target detection model, especially the temporal feature fusion layer. This process involves temporally aligning the top-view feature map at the current moment with the top-view feature map at at least one historical moment, then concatenating them to incorporate temporal information, and finally performing subsequent processing through at least one second model structure.
[0170] To further illustrate, consider an optional assumption: an intelligent traffic detection system that needs to detect and monitor vehicles in lanes in real time. The system uses an object detection model to process continuously captured lane images. At the current moment, the model generates a top-down feature map. To more accurately identify vehicles and predict their trajectories, the model also utilizes a temporal feature fusion layer to temporally align and stitch the current top-down feature map with top-down feature maps from the previous moment or several moments ago. In this way, the model can more accurately predict the future position of a vehicle based on its position changes over a past period.
[0171] By introducing a temporal feature fusion layer, the object detection model can handle dynamically changing lane scenes, improving detection accuracy and robustness. Especially when dealing with moving targets (such as vehicles and pedestrians), incorporating temporal information helps the model better understand the target's motion patterns and trajectories.
[0172] As an optional approach, the top-view feature map is processed using at least one target model structure to obtain detection results, including:
[0173] The detection result output layer of the target detection model is used to perform subsequent processing on the top-view feature map to obtain the three-dimensional position, element size, element orientation angle, and element velocity of the target element in the first coordinate system. The detection results include the three-dimensional position, element size, element orientation angle, and element velocity.
[0174] Optionally, in this embodiment, the three-dimensional position may refer to the coordinate position of the target element in the actual three-dimensional space, which helps to determine the precise position of the element in the lane scene.
[0175] Optionally, in this embodiment, the element size can represent the actual size of the target element, such as the length, width, and height of the vehicle.
[0176] Optionally, in this embodiment, the element orientation angle can describe the angle of the target element relative to a certain reference direction (such as due north), reflecting the orientation of the element.
[0177] Optionally, in this embodiment, element velocity can represent the movement speed of the target element in the lane scene, including the magnitude and direction of the velocity.
[0178] It should be noted that the "detection result output layer" mentioned in this embodiment is an important component of the target detection model. Its main function is to further analyze and process the top-view feature map after processing by the previous layers, so as to output detailed information such as the specific position, size, orientation angle, and velocity of the target element in three-dimensional space. This information is crucial for understanding and analyzing dynamic elements in the lane scene.
[0179] To further illustrate, an optional hypothetical object detection model is applied to an autonomous driving system to detect other vehicles in the lane. The detection output layer can analyze the top-down feature map and output the 3D position, size, orientation angle, and speed of each vehicle. For example, a moving car might be detected as being located at a specific latitude and longitude coordinate, 4.5 meters long, 1.8 meters wide, currently facing due east, and traveling at a speed of 60 kilometers per hour.
[0180] By conducting in-depth analysis of the top-down feature map through the detection result output layer, the target detection model can provide comprehensive information about target elements, rather than just simple location markings, enabling it to better adapt to complex and ever-changing lane scenarios.
[0181] As an optional approach, obtaining the second coordinate position of at least one lane element in the second coordinate system includes:
[0182] Using a projection matrix, at least one lane element is projected from a first coordinate system to a second coordinate system to obtain the projected position of at least one lane element. The second coordinate position includes the projected position. The projection matrix represents the mapping relationship between the first coordinate system and the second coordinate system.
[0183] Optionally, the projection matrix can be a mathematical matrix used to transform a point in one coordinate system to another. In this embodiment, the projection matrix can be used to transform a three-dimensional real coordinate system to a two-dimensional pixel coordinate system.
[0184] It should be noted that the description of this embodiment mentions a key process: projecting lane elements from a first coordinate system (actual coordinate system) to a second coordinate system (pixel coordinate system) using a projection matrix. Here, the "projection matrix" is a mathematical tool used to achieve transformations between different coordinate systems. In this way, this embodiment can map points or objects in three-dimensional space onto a two-dimensional image plane.
[0185] To further illustrate, consider an optional assumption: this embodiment presents a lane scene, including multiple lane elements (such as traffic signs, road markings, etc.). These elements have specific locations in actual three-dimensional space. To accurately represent the positions of these elements on a two-dimensional image, this embodiment uses a projection matrix to convert these three-dimensional positions into two-dimensional pixel coordinates. Thus, when viewing the lane image, the exact locations of these lane elements on the image can be determined.
[0186] The embodiments provided in this application utilize a projection matrix for coordinate transformation, enabling precise mapping of elements in a lane scene onto a lane image. This makes subsequent image processing and analysis (such as object detection, object identification, etc.) more accurate and efficient.
[0187] As an alternative approach, for ease of understanding, the aforementioned target detection method based on a top-down perspective is applied to a smart highway scenario to perform end-to-end Bird's Eye View (BEV) target detection. In the construction of smart highways, multiple camera sensors installed on the road gantries play a crucial role, capturing traffic targets from different directions and providing data support for comprehensive detection. Simultaneously, high-precision map information allows for detailed modeling of the 3D traffic environment, providing a solid foundation for subsequent target detection.
[0188] The sensor layout in this embodiment is as follows: Figure 4 As shown, the gantry is equipped with five cameras. The fisheye camera in the center effectively compensates for the blind spot directly below the gantry, while the four cameras on the left and right sides detect traffic conditions in two different directions. Among these four cameras, there are two short-focus and two long-focus cameras, and the introduction of the long-focus cameras significantly enhances the ability to observe distant targets.
[0189] Building upon this foundation, this embodiment proposes an innovative end-to-end visual BEV target detection method, particularly suitable for high-speed scenarios. The method first utilizes an image feature encoding layer to extract deep features from images from multiple cameras, ensuring comprehensive information capture. Subsequently, through a high-precision map view frustum transformation layer, this embodiment accurately projects feature maps from a conventional viewpoint to a bird's-eye view; this step is crucial for subsequent target detection. Next, a BEV encoding layer performs secondary encoding on the bird's-eye view features, further optimizing feature representation. Finally, through a BEV detection head network, this embodiment achieves efficient and accurate target detection.
[0190] It should be noted that in autonomous driving scenarios, BEV target detection typically relies on surround-view multi-camera systems, whose detection range is generally limited to within 60 meters and does not involve the fusion of high-precision map information. However, in the context of smart highway applications, the demand for target detection has expanded to a greater distance (within 500 meters). At the same time, due to the availability of high-precision maps, how to effectively fuse this information in an end-to-end manner has become a core research issue.
[0191] To address the aforementioned core issues, this embodiment proposes an end-to-end target detection method for BEVs (Browser-Electric Vehicles) specifically designed for intelligent highway scenarios. This method first extracts features from image data from multiple cameras using an image feature encoding layer to ensure the capture of rich visual information. Next, through a high-precision map view frustum transformation layer, this embodiment accurately transforms feature maps from conventional viewpoints to a bird's-eye view; this step is crucial for integrating information from different perspectives. Then, a BEV encoding layer further encodes and enhances the bird's-eye view features to improve their representativeness and discriminative power. Finally, through a BEV detection head network, this embodiment achieves efficient and accurate target detection.
[0192] This embodiment not only fills the research gap in end-to-end BEV target detection methods in high-speed scenarios, but also solves the key problem of how to effectively integrate and utilize high-precision map information for BEV target detection within an end-to-end framework.
[0193] To further illustrate, the optional intelligent highway scenario's BEV end-to-end target detection process is as follows: Figure 5 As shown, the specific steps are as follows:
[0194] S502, multi-camera image acquisition:
[0195] Five cameras (including a fisheye camera, a short-focus camera, and a long-focus camera) mounted on a gantry are used to capture traffic targets (target elements) in high-speed scenes.
[0196] S504, Image Feature Coding Layer:
[0197] Deep feature extraction is performed on images from multiple cameras to ensure comprehensive capture of key information in the traffic environment.
[0198] Using Ncams cameras as the data source, the input image size is set to H*W*3, where b represents the number of images input each time. To enable the image feature encoding layer to efficiently process image features from multiple cameras simultaneously, the input data needs to be specifically reassembled.
[0199] Specific examples Figure 6As shown, the standard input format for the convolution module is b*C*H*W, where C represents the number of channels. To meet this requirement, this embodiment reorganizes the input data into the format [b*Ncams, C, H, W]. This reorganization strategy essentially treats the image from each camera as part of an independent sample, thereby enabling the simultaneous extraction and processing of image features from multiple batches and multiple cameras during a single forward propagation.
[0200] After a series of convolutional processes, the network outputs two feature maps of different scales, denoted as f0 and f1, respectively. Their dimensionality expressions are shown in Equations (1) and (2) below:
[0201] The dimension of f0 is: b*Ncams*1024*(H / 16)*(W / 16)(1)
[0202] The dimension of f1 is: b*Ncams*2048*(H / 16)*(W / 16)(2)
[0203] To achieve effective feature fusion, this embodiment uses bilinear interpolation to upsample f0 so that its dimension matches that of f1, and then performs cumulative fusion using the following formula (3). This fusion strategy can fully utilize the information in feature maps at different scales, thereby enhancing the richness and robustness of the features.
[0204] F=Conv3×3(Up(Conv1×1(L1))+Conv1×1(L0)) (3)
[0205] Finally, the fused feature map is downsampled using a convolution operation to obtain a new feature map F. For ease of subsequent processing, this embodiment reorganizes the dimensions of F as [b, Ncams, 512, H / 16, W / 16]. This step not only makes the structure of the feature map clearer but also facilitates subsequent object detection tasks.
[0206] S506, High-precision map view frustum conversion layer (element view frustum conversion layer):
[0207] By leveraging high-precision map information, feature maps from conventional perspectives are accurately projected onto bird's-eye view (BEV) perspectives, integrating information from different perspectives to provide an accurate data foundation for subsequent detection.
[0208] In the high-definition map frustum transformation layer, the primary task is to complete the joint calibration of the high-definition map and the camera. This step is crucial for subsequent 3D information acquisition and projection.
[0209] The high-precision maps collected by the intelligent highway system contain key information such as static lane lines on the road, which are presented as strings of white dots in the high-precision map. In order to combine this map information with the image data captured by the camera, this embodiment requires a precise calibration algorithm to accurately calibrate the elements of the high-precision map onto the image, thereby obtaining the three-dimensional information of the pixels in the image.
[0210] The calibration process mainly consists of two steps:
[0211] First, within the camera's field of view in the physical world, this embodiment uses a GPS marker to manually record the marker positions and their corresponding latitude and longitude coordinates. To ensure calibration accuracy, data from at least 10 points is typically recorded. Simultaneously, this embodiment also records the pixel coordinates of these points on the camera image. Since the camera position is fixed and its sensor field of view remains constant, ignoring minor sensor jitter, it can be reasonably assumed that the depth value of each pixel on the camera image relative to the camera remains essentially constant.
[0212] Next, this embodiment utilizes the one-to-one correspondence between recorded pixels and GPS points to calculate the homography matrix. By solving this matrix, this embodiment can obtain a projection matrix R from the high-precision map to the image. Using this projection matrix R, this embodiment can accurately project the acquired lane line information from the high-precision map onto the camera image, such as... Figure 7 The dashed lane lines are shown in the image. This process ensures accurate alignment between high-precision map information and image data, providing a solid foundation for subsequent target detection and scene understanding.
[0213] By jointly calibrating a high-precision map and a camera, the precise matching relationship between the 3D coordinates of all lane lines in the image and their corresponding pixel coordinates is obtained. However, further processing is required to obtain the 3D coordinates of each pixel in the image. This embodiment uses an interpolation method to calculate the 3D coordinates of all pixels using the known 3D coordinates of the lane lines, thereby constructing a complete pixel projection mapping table.
[0214] To further illustrate, based on Figure 7 The scene shown is used for calculation. Figure 7 Taking the 3D coordinates of pixel number 1 as an example:
[0215] First, in this embodiment, all lane line points in the image are searched, and the nearest neighbor points of the lanes to the left and right of the point numbered 1 are determined. Figure 7Points numbered 3 and 4 are used. Next, this embodiment uses a linear interpolation algorithm to calculate the lateral 3D coordinates (X3d) of point number 1. Specifically, this embodiment uses the following formula (4) to perform interpolation calculation based on the lateral coordinates (Ix) of the pixel point and the lateral coordinates (Px2 and Px1) of the actual 3D position of the searched lane line points and the corresponding lateral coordinates (Ipx1 and Ipx2):
[0216] X3d=Px1+Ipx2-Ipx1(Ix-Ipx1)×(Px2-Px1)(4)
[0217] Subsequently, this embodiment searches all lane line points in the image again, finding the nearest neighbor and adjacent points of the pixel numbered 1. Figure 7 Points numbered 3 and 2 are used. Using the same interpolation algorithm as in the first step, this embodiment calculates the vertical 3D position coordinates of point number 1.
[0218] In summary, this embodiment can determine the corresponding 3D position coordinates of each pixel in the image plane. Finally, a complete pixel projection mapping table, HDSet, is constructed. This mapping table provides a convenient way to subsequently query the bird's-eye view (BEV) position information corresponding to any pixel (e.g., img_x, img_y), that is, to achieve fast retrieval through HDSet[img_x, img_y].
[0219] Furthermore, in autonomous vehicle scenarios, generating a bird's-eye view (BEV) feature map typically requires predicting a depth distribution map for each pixel in the image, and then projecting features at different depths onto the BEV feature map using the camera's intrinsic and extrinsic parameters. However, this approach significantly reduces the algorithm's efficiency and performance due to the need to predict the depth distribution, thus affecting the algorithm's real-time performance. In the intelligent highway scenario, this embodiment possesses the advantage of high-precision maps, allowing for the direct acquisition of the relatively accurate depth value of each pixel through a lookup table, thereby significantly improving efficiency.
[0220] Specifically, after obtaining the HDSet (pixel projection map), this embodiment needs to project the image features [b, Ncams, 512, H / 16, W / 16] obtained in the first step onto the BEV viewpoint. Assume the goal of this embodiment is to detect BEV targets within a 500-meter radius in a highway scene, and set the horizontal and vertical perception resolution to 1 meter. Therefore, this embodiment first defines a BEV feature map of size [500-(-500) / 1.0, 500-(-500) / 1.0, 512] and initializes it to 0. Next, this embodiment uses the HDSet to query the BEV position (X, Y) of the corresponding image pixels [b, Ncams, 512, H / 16, W / 16]. Then, the features are projected (or placed) onto the corresponding positions of the BEV feature map using the following formulas (5) and (6), where Xgrid and Ygrid represent the pixel coordinates on the BEV feature map after projection:
[0221] Xgrid = 1.0X - (-500) (5)
[0222] Ygrid = 1.0Y - (-500) (6)
[0223] Finally, this embodiment filters out all view frustums outside the horizontal and vertical ranges of the bird's-eye view, as well as features not on the ground (i.e., Z is not equal to 0). This step further improves the accuracy and effectiveness of the BEV feature map. The entire process fully utilizes the precise depth information provided by high-precision maps, significantly improving the efficiency and performance of BEV feature map generation from a technical perspective.
[0224] S508, BEV feature pooling layer (viewpoint feature pooling layer):
[0225] After processing by the view frustum transformation layer of a high-precision map, sometimes a pixel location on the same BEV feature map may be occupied by multiple feature projections. This phenomenon is mainly caused by three factors: First, the resolution limitation of the BEV feature map may cause multiple features to be projected onto the same pixel location; second, pixels on targets at the same height may overlap when projected onto the BEV feature map; and finally, overlapping fields of view between different cameras may also cause target features in overlapping fields of view to be projected onto the same location on the BEV feature map.
[0226] To address this issue, this embodiment specifically introduces a BEV feature pooling layer. The core function of this layer is to perform a summation operation on all features projected to the same pixel location, thereby obtaining the final feature vector for that location. Specifically, this embodiment adds all features falling at the same BEV pixel location to obtain the final feature representation for that location. The mathematical expression is shown in the following formula (7):
[0227] f = sum(f1+f2+…+fn) (7)
[0228] Here, f represents the final feature vector, while f1, f2, ..., fn are all features projected onto the same BEV pixel location.
[0229] This step not only effectively solves the problem of feature overlap, but also ensures from a technical perspective that the feature representation at each pixel location on the BEV feature map is accurate and comprehensive. Through the processing of the BEV feature pooling layer, this embodiment can better utilize the projected features for subsequent target detection and scene understanding tasks, improving the performance and accuracy of the entire system.
[0230] S510, Temporal Feature Fusion Layer:
[0231] To improve the model's ability to detect BEV (bird's-eye view) targets, especially considering that in real-world scenarios, when a target is occluded at certain times, the lack of temporal correlation between preceding and following frames may lead to missed detections, this embodiment specifically designs a temporal feature fusion layer.
[0232] Specifically, assume the current BEV feature map is Ft, and the feature maps for historical times are Ft-1, Ft-2, ..., Ft-n. Each time-series BEV feature map has the same size, i.e., [B, Ci, H, W]. To effectively fuse these temporal features, this embodiment uses a concat operation on the channel dimension Ci to connect the feature maps from all times. Thus, the size of the concatenated feature map will be the sum of the number of channels in all time-series feature maps, while other dimensions remain unchanged.
[0233] Finally, to further refine and integrate these temporal features, this embodiment processes the stitched feature map through a series of 1×1 and 3×3 convolutional layers. These convolutional layers are designed to effectively capture and fuse temporal information without changing the size of the feature map, thereby enhancing the model's ability to detect targets in dynamic scenes. The entire process not only considers the utilization of temporal information but also ensures the effective fusion and refinement of features from a technical perspective, providing richer and more accurate feature representations for subsequent target detection tasks.
[0234] S512, BEV detection head network (detection result output layer):
[0235] By leveraging enhanced bird's-eye view features, a dedicated detection network enables efficient and accurate target detection, identifying and locating traffic targets in high-speed scenes.
[0236] After the aforementioned processing and feature fusion steps, this embodiment finally generates a BEV feature map of size [B, 64*10, 500, 500]. Here, 64 represents the feature length of each BEV feature map, and multiplying by 10 is because this embodiment effectively stitches together historical feature maps from the past 10 frames, thereby enriching the temporal information of the features.
[0237] To derive specific BEV target detection results from this feature map, a BEV detection head network was designed. This network structure consists of a series of consecutive 1×1 convolutional layers, specifically designed to predict various attributes of the BEV target. Specifically, it can predict the target's 3D spatial position (x, y, z), size (l, w, h), orientation angle (yaw), and velocity (v). Through this design, this embodiment achieves comprehensive and accurate detection of BEV targets.
[0238] The following is an application example of this embodiment in a real high-speed scenario: This scenario is equipped with 5 cameras, including 2 telephoto cameras, 2 short-focus cameras, and 1 fisheye camera, to ensure all-around detection without blind spots. Simultaneously, combined with high-precision map information, the system of this embodiment can generate corresponding BEV target detection results.
[0239] It should be noted that this embodiment is not only applicable to BEV target detection in intelligent highway scenarios, but can also be widely applied to any task that requires multi-camera fusion for BEV target detection and is equipped with high-precision map information. It not only fills the gap in the industry for end-to-end BEV target detection methods in highway scenarios, but more importantly, it solves the technical challenge of how to effectively fuse and utilize high-precision map information for BEV target detection under end-to-end conditions.
[0240] The embodiments provided in this application improve the accuracy and efficiency of BEV target detection, promoting the application of BEV detection algorithms in high-speed scenarios and a wider range of fields. From a technical perspective, these embodiments represent a new breakthrough in the field of BEV target detection and provide valuable reference and inspiration for future related research.
[0241] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0242] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0243] According to another aspect of the embodiments of this application, a target detection device based on a top-down view is also provided for implementing the above-described target detection method based on a top-down view. For example... Figure 8 As shown, the device includes:
[0244] The first acquisition unit 802 is used to acquire a lane image obtained by the image acquisition device from a top-down perspective after acquiring the lane scene. The lane image contains at least one lane element and a target element. The lane element is an element in the lane scene that carries attribute information, and the target element is an element in the lane image that is expected to be detected.
[0245] The second acquisition unit 804 is used to acquire the first coordinate position of at least one lane element in the first coordinate system, acquire the second coordinate position of at least one lane element in the second coordinate system, and acquire the third coordinate position of the target element in the second coordinate system. The attribute information includes the first coordinate position, the first coordinate system is the actual coordinate system corresponding to the lane scene, and the second coordinate system is the pixel coordinate system corresponding to the lane image.
[0246] The third acquisition unit 806 is used to acquire the fourth coordinate position of the target element in the first coordinate system based on the first positional relationship between the second coordinate position and the third coordinate position, and the first coordinate position.
[0247] The fourth acquisition unit 808 is used to acquire the detection result of the target element in the lane scene based on the second positional relationship between the fifth coordinate position and the fourth coordinate position when the fifth coordinate position of the image acquisition device in the first coordinate system is acquired.
[0248] For specific implementation examples, please refer to the examples shown in the above-described target detection method based on a top-down view. These examples will not be repeated here.
[0249] As an optional solution, the third acquisition unit 806 includes:
[0250] A determining module is used to determine at least one first lane element and at least one second lane element from at least one lane element, wherein the first lane element is in a second coordinate system and the first positional relationship satisfies that the lane element adjacent to the target element in a first direction is a lane element, and the second lane element is in a second coordinate system and the first positional relationship satisfies that the lane element adjacent to the target element in a second direction is a lane element.
[0251] The first acquisition module is used to acquire, based on at least one first lane element and at least one second lane element, a first coordinate of a target element in a first coordinate system and located in a first direction, and a second coordinate of the target element in a first coordinate system and located in a second direction, wherein the fourth coordinate position includes the first coordinate and the second coordinate.
[0252] For specific implementation examples, please refer to the examples shown in the above-described target detection method based on a top-down view. These examples will not be repeated here.
[0253] As an optional solution, the first acquisition module includes:
[0254] The first acquisition submodule is used to acquire the first pixel coordinate of the target element in the second coordinate system and located in the first direction, the first actual coordinate of the first element in the first coordinate system and located in the first direction of at least one first lane element, the second actual coordinate of the second element in the first coordinate system and located in the first direction of at least one first lane element, the second pixel coordinate of the first element in the second coordinate system and located in the first direction of the first direction, and the third pixel coordinate of the second element in the second coordinate system and located in the first direction of the second element.
[0255] The second acquisition submodule is used to acquire the first difference between the first pixel coordinate and the first actual coordinate, the second difference between the second actual coordinate and the first actual coordinate, and the third difference between the third pixel coordinate and the second pixel coordinate.
[0256] The third acquisition submodule is used to acquire the first product value of the product of the first difference and the second difference;
[0257] The fourth acquisition submodule is used to obtain the first ratio of the first product value to the third difference value;
[0258] The fifth acquisition submodule is used to acquire the first sum of the first actual coordinates and the first ratio, and to determine the first sum as the first coordinate; and,
[0259] The sixth acquisition submodule is used to acquire the fourth pixel coordinate of the target element in the second coordinate system and located in the second direction, the third actual coordinate of the third element in the first coordinate system and located in the second direction of at least one second lane element, the fourth actual coordinate of the fourth element in the first coordinate system and located in the second direction of at least one second lane element, the fifth pixel coordinate of the third element in the second coordinate system and located in the second direction of the second direction, and the sixth pixel coordinate of the fourth element in the second coordinate system and located in the second direction of the second direction.
[0260] The seventh acquisition submodule is used to acquire the fourth difference value between the fourth pixel coordinate and the third actual coordinate, the fifth difference value between the fourth actual coordinate and the third actual coordinate, and the sixth difference value between the third pixel coordinate and the second pixel coordinate.
[0261] The eighth submodule is used to obtain the second product value of the product of the fourth difference and the fifth difference;
[0262] The ninth submodule is used to obtain the second ratio of the second product value to the sixth difference value;
[0263] The tenth acquisition submodule is used to obtain the second sum of the third actual coordinate and the second ratio, and to determine the second sum as the second coordinate.
[0264] For specific implementation examples, please refer to the examples shown in the above-described target detection method based on a top-down view. These examples will not be repeated here.
[0265] As an optional solution, the device also includes:
[0266] The storage module is used to store the projection information between the first coordinate and the second coordinate in a pixel projection mapping table after obtaining the first coordinate of the target element in a first coordinate system and the second coordinate of the target element in a second direction in a first coordinate system based on at least one first lane element and at least one second lane element. The projection information between the first coordinate and the second coordinate is used to indicate whether the target element is projected from the first coordinate system to the second coordinate system or from the second coordinate system to the first coordinate system. The pixel projection mapping table is used to store and query the projection information corresponding to each element.
[0267] For specific implementation examples, please refer to the examples shown in the above-described target detection method based on a top-down view. These examples will not be repeated here.
[0268] As an optional solution, the device also includes:
[0269] The second acquisition module is used to acquire a projection query request triggered by the target range of the lane image after storing the projection information between the first coordinate and the second coordinate into a pixel projection mapping table. The projection query request is used to request the target position of the target element in the lane image located within the target range in the first coordinate system.
[0270] The third acquisition module is used to acquire multiple target elements located within the target range in the lane image after storing the projection information between the first coordinate and the second coordinate into a pixel projection mapping table.
[0271] The query module is used to query the pixel projection mapping table after storing the projection information between the first coordinate and the second coordinate in the pixel projection mapping table, and to query the multiple candidate coordinate positions of multiple candidate elements in the first coordinate system by multiple coordinate positions of multiple candidate elements in the second coordinate system.
[0272] The filtering module is used to filter coordinate positions that are outside a predetermined detection range from multiple candidate coordinate positions after storing the projection information between the first coordinate and the second coordinate in a pixel projection mapping table, so as to obtain at least one target position.
[0273] For specific implementation examples, please refer to the examples shown in the above-described target detection method based on a top-down view. These examples will not be repeated here.
[0274] As an optional solution, the device also includes:
[0275] The input unit is used to input the lane image obtained by the image acquisition device from the top view of the lane scene into the target detection model after acquiring the lane image. The target detection model is a neural network model for detecting target elements, which is trained by multiple image samples.
[0276] The fifth acquisition unit is used to acquire the detection results output by the target detection model after acquiring the lane image obtained by the image acquisition device from a top-down perspective of the lane scene.
[0277] For specific implementation examples, please refer to the examples shown in the above-described target detection method based on a top-down view. These examples will not be repeated here.
[0278] As an optional solution, the image acquisition device includes N acquisition sub-devices, characterized in that the device further includes:
[0279] The integration unit is used to integrate the N lane images acquired by the N acquisition sub-devices into a multidimensional tensor lane image after the lane image is input into the target detection model. The multidimensional tensor includes a tensor corresponding to the input size of the image feature encoding layer of the target detection model and a tensor corresponding to the number N, where N is an integer greater than 1.
[0280] The extraction unit is used to extract the image features corresponding to the multidimensional tensor lane image through the image feature encoding layer of the target detection model after the lane image is input into the target detection model, so as to obtain the first feature map and the second feature map at different scales.
[0281] The upsampling unit is used to upsample the first or second feature map of the first feature map and the second feature map of different scales after the lane image is input into the target detection model, so as to obtain the first feature map and the second feature map of the same scale.
[0282] The fusion unit is used to perform cumulative fusion processing on the first feature map and the second feature map of the same scale after the lane image is input into the target detection model to obtain the first target feature map.
[0283] The downsampling unit is used to downsample the first target feature map after the lane image is input into the target detection model to obtain the second target feature map;
[0284] The fifth acquisition unit includes a processing module, which processes the second target feature map through the subsequent model structure of the target detection model to obtain the detection result.
[0285] For specific implementation examples, please refer to the examples shown in the above-described target detection method based on a top-down view. These examples will not be repeated here.
[0286] As an optional solution, the processing module includes:
[0287] The transformation submodule is used to perform frustum transformation on the second target feature map through the element frustum transformation layer of the target detection model to obtain a top view feature map, wherein the top view feature map is used to represent the first positional relationship, and the subsequent model structure includes the element frustum transformation layer;
[0288] The processing submodule is used to perform subsequent processing on the top-view feature map through at least one target model structure to obtain the detection result, wherein the subsequent model structure includes at least one target model structure.
[0289] For specific implementation examples, please refer to the examples shown in the above-described target detection method based on a top-down view. These examples will not be repeated here.
[0290] As an optional solution, the processing submodule includes:
[0291] The summation subunit is used to sum the features falling at the same pixel position in the top view feature map through the view feature pooling layer of the target detection model to obtain the pooled top view feature map. At least one target model structure includes a view feature pooling layer.
[0292] The first processing subunit is used to perform subsequent processing on the pooled top-view feature map through at least one first model structure to obtain the detection result, wherein at least one target model structure includes at least one first model structure.
[0293] For specific implementation examples, please refer to the examples shown in the above-described target detection method based on a top-down view. These examples will not be repeated here.
[0294] As an optional solution, the processing submodule includes:
[0295] The alignment subunit is used to perform temporal alignment processing between the top view feature map at the current moment and the top view feature map at at least one historical moment through the temporal feature fusion layer of the target detection model, so as to obtain the aligned top view feature maps at different moments, wherein at least one target model structure includes a temporal feature fusion layer.
[0296] The splicing subunit is used to splice the top-view feature maps at different times through the temporal feature fusion layer to obtain a top-view feature map that incorporates temporal information.
[0297] The second processing subunit is used to perform subsequent processing on the top-view feature map that incorporates time-series information through at least one second model structure to obtain the detection result, wherein at least one target model structure includes at least one second model structure.
[0298] For specific implementation examples, please refer to the examples shown in the above-described target detection method based on a top-down view. These examples will not be repeated here.
[0299] As an optional solution, the processing submodule includes:
[0300] The third processing subunit is used to perform subsequent processing on the top-view feature map through the detection result output layer of the target detection model to obtain the three-dimensional position, element size, element orientation angle, and element velocity of the target element in the first coordinate system. The detection results include the three-dimensional position, element size, element orientation angle, and element velocity.
[0301] For specific implementation examples, please refer to the examples shown in the above-described target detection method based on a top-down view. These examples will not be repeated here.
[0302] As an optional solution, the second acquisition unit 804 includes:
[0303] The projection module is used to project at least one lane element from a first coordinate system to a second coordinate system using a projection matrix, thereby obtaining the projected position of at least one lane element. The second coordinate position includes the projected position, and the projection matrix represents the mapping relationship between the first coordinate system and the second coordinate system.
[0304] For specific implementation examples, please refer to the examples shown in the above-described target detection method based on a top-down view. These examples will not be repeated here.
[0305] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described target detection method based on a top-down view is also provided. This electronic device may, but is not limited to, […]. Figure 1 The user equipment 102 or server 112 shown in the figure, in this embodiment, is taken as an example of an electronic device, namely user equipment 102. Further, as shown in the figure... Figure 9 As shown, the electronic device includes a memory 902 and a processor 904. The memory 902 stores a computer program, and the processor 904 is configured to execute the steps of any of the above method embodiments through the computer program.
[0306] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0307] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0308] S1, acquire a lane image obtained by the image acquisition device from a top-down perspective after acquiring the lane scene. The lane image contains at least one lane element and a target element. The lane element is an element in the lane scene that carries attribute information, and the target element is an element in the lane image that is expected to be detected.
[0309] S2, obtain the first coordinate position of at least one lane element in the first coordinate system, obtain the second coordinate position of at least one lane element in the second coordinate system, and obtain the third coordinate position of the target element in the second coordinate system, wherein the attribute information includes the first coordinate position, the first coordinate system is the actual coordinate system corresponding to the lane scene, and the second coordinate system is the pixel coordinate system corresponding to the lane image;
[0310] S3, based on the first positional relationship between the second and third coordinate positions, and the first coordinate position, obtain the fourth coordinate position of the target element in the first coordinate system;
[0311] S4. After obtaining the fifth coordinate position of the image acquisition device in the first coordinate system, the detection result of the target element in the lane scene is obtained according to the second positional relationship between the fifth coordinate position and the fourth coordinate position.
[0312] Alternatively, as those skilled in the art will understand, Figure 9 The structure shown is for illustrative purposes only. Figure 9 This does not limit the structure of the aforementioned electronic devices. For example, the electronic device may also include components that are more... Figure 9 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 9 The different configurations shown.
[0313] The memory 902 can be used to store software programs and modules, such as the program instructions / modules corresponding to the target detection method and device based on the top-down view in this embodiment. The processor 904 executes various functional applications and data processing by running the software programs and modules stored in the memory 902, thereby realizing the aforementioned target detection method based on the top-down view. The memory 902 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 902 may further include memory remotely located relative to the processor 904, and these remote memories can be connected to electronic devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 902 may be used, but is not limited to, to store information such as lane images, lane elements, target elements, and detection results. As an example, such as... Figure 9 As shown, the memory 902 may include, but is not limited to, the first acquisition unit 802, the second acquisition unit 804, the third acquisition unit 806, and the fourth acquisition unit 808 in the target detection device based on the top-down view. Furthermore, it may include, but is not limited to, other module units in the target detection device based on the top-down view, which will not be elaborated upon in this example.
[0314] Optionally, the transmission device 906 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 906 includes a Network Interface Controller (NIC), which can be connected to other network devices and routers via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 906 is a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0315] In addition, the aforementioned electronic device also includes: a display 908 for displaying information such as the lane image, lane elements, target elements, and detection results; and a connection bus 910 for connecting the various module components in the aforementioned electronic device.
[0316] In other embodiments, the aforementioned user equipment or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer network, and any form of computing device, such as a server, user equipment, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.
[0317] According to one aspect of this application, a computer program product is provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions provided in embodiments of this application.
[0318] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0319] It should be noted that the computer system of the electronic device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0320] A computer system includes a Central Processing Unit (CPU), which performs various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) or loaded from RAM. ROM also stores various programs and data required for system operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output interfaces (I / O interfaces) are also connected to the bus.
[0321] The following components are connected to the input / output interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard drives; and communication sections including network interface cards such as LAN cards and modems. The communication section performs communication processing via a network such as the Internet. Drives are also connected to the input / output interface as needed. Removable media, such as disks, optical discs, magneto-optical discs, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as required.
[0322] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions defined in the system of this application.
[0323] According to one aspect of this application, a computer-readable storage medium is provided, wherein a processor of a computer device reads computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations described above.
[0324] Optionally, in this embodiment, the computer-readable storage medium described above may be configured to store a computer program for performing the following steps:
[0325] S1, acquire a lane image obtained by the image acquisition device from a top-down perspective after acquiring the lane scene. The lane image contains at least one lane element and a target element. The lane element is an element in the lane scene that carries attribute information, and the target element is an element in the lane image that is expected to be detected.
[0326] S2, obtain the first coordinate position of at least one lane element in the first coordinate system, obtain the second coordinate position of at least one lane element in the second coordinate system, and obtain the third coordinate position of the target element in the second coordinate system, wherein the attribute information includes the first coordinate position, the first coordinate system is the actual coordinate system corresponding to the lane scene, and the second coordinate system is the pixel coordinate system corresponding to the lane image;
[0327] S3, based on the first positional relationship between the second and third coordinate positions, and the first coordinate position, obtain the fourth coordinate position of the target element in the first coordinate system;
[0328] S4. After obtaining the fifth coordinate position of the image acquisition device in the first coordinate system, the detection result of the target element in the lane scene is obtained according to the second positional relationship between the fifth coordinate position and the fourth coordinate position.
[0329] Optionally, in embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0330] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware of an electronic device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0331] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0332] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods of the various embodiments of this application.
[0333] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0334] In the several embodiments provided in this application, it should be understood that the disclosed user equipment can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.
[0335] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0336] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0337] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A target detection method based on a top-down view, characterized in that, include: The image acquisition device acquires a lane image from a top-down perspective of the lane scene. The lane image contains at least one lane element and a target element. The lane element is an element in the lane scene that carries attribute information, and the target element is an element in the lane image that is expected to be detected. The method involves obtaining the first coordinate position of the at least one lane element in a first coordinate system, obtaining the second coordinate position of the at least one lane element in a second coordinate system, and obtaining the third coordinate position of the target element in the second coordinate system. The attribute information includes the first coordinate position, the first coordinate system is the actual coordinate system corresponding to the lane scene, and the second coordinate system is the pixel coordinate system corresponding to the lane image. Based on the first positional relationship between the second coordinate position and the third coordinate position, and the first coordinate position, the fourth coordinate position of the target element in the first coordinate system is obtained; Having obtained the fifth coordinate position of the image acquisition device in the first coordinate system, the detection result of the target element in the lane scene is obtained according to the second positional relationship between the fifth coordinate position and the fourth coordinate position.
2. The method according to claim 1, characterized in that, The step of obtaining the fourth coordinate position of the target element in the first coordinate system based on the first positional relationship between the second coordinate position and the third coordinate position, and the first coordinate position, includes: At least one first lane element and at least one second lane element are determined from the at least one lane element, wherein the first lane element is in the second coordinate system and the first positional relationship satisfies the requirement that the lane element is adjacent to the target element in a first direction, and the second lane element is in the second coordinate system and the first positional relationship satisfies the requirement that the lane element is adjacent to the target element in a second direction. Based on the at least one first lane element and the at least one second lane element, obtain the first coordinate of the target element in the first coordinate system and located in the first direction, and the second coordinate of the target element in the first coordinate system and located in the second direction, wherein the fourth coordinate position includes the first coordinate and the second coordinate.
3. The method according to claim 2, characterized in that, The step of obtaining the first coordinate of the target element in the first coordinate system and located in the first direction, and the second coordinate of the target element in the first coordinate system and located in the second direction, based on the at least one first lane element and the at least one second lane element, includes: Obtain the first pixel coordinate of the target element in the second coordinate system and located in the first direction, the first actual coordinate of the first element in the first coordinate system and located in the first direction, the second actual coordinate of the second element in the first coordinate system and located in the first direction, the second pixel coordinate of the first element in the second coordinate system and located in the first direction, and the third pixel coordinate of the second element in the second coordinate system and located in the first direction; Obtain a first difference between the first pixel coordinate and the first actual coordinate, a second difference between the second actual coordinate and the first actual coordinate, and a third difference between the third pixel coordinate and the second pixel coordinate; Obtain the first product value of the product of the first difference and the second difference; Obtain a first ratio value of the ratio of the first product value to the third difference value; Obtain the first sum of the first actual coordinates and the first ratio, and determine the first sum as the first coordinates; and, Obtain the fourth pixel coordinate of the target element in the second coordinate system and located in the second direction, the third actual coordinate of the third element of the at least one second lane element in the first coordinate system and located in the second direction, the fourth actual coordinate of the fourth element of the at least one second lane element in the first coordinate system and located in the second direction, the fifth pixel coordinate of the third element in the second coordinate system and located in the second direction, and the sixth pixel coordinate of the fourth element in the second coordinate system and located in the second direction; Obtain the fourth difference value between the fourth pixel coordinate and the third actual coordinate, the fifth difference value between the fourth actual coordinate and the third actual coordinate, and the sixth difference value between the third pixel coordinate and the second pixel coordinate; Obtain the second product value of the product of the fourth difference and the fifth difference; Obtain a second ratio value of the ratio of the second product value to the sixth difference value; Obtain the second sum of the third actual coordinate and the second ratio, and determine the second sum as the second coordinate.
4. The method according to claim 2, characterized in that, After obtaining the first coordinates of the target element in the first coordinate system, located in the first direction, and the second coordinates of the target element in the first coordinate system, located in the second direction, based on the at least one first lane element and the at least one second lane element, the method further includes: The projection information between the first coordinate and the second coordinate is stored in a pixel projection mapping table. The projection information between the first coordinate and the second coordinate is used to indicate whether the target element is projected from the first coordinate system to the second coordinate system or from the second coordinate system to the first coordinate system. The pixel projection mapping table is used to store and query the projection information corresponding to each element.
5. The method according to claim 4, characterized in that, After storing the projection information between the first coordinate and the second coordinate into a pixel projection mapping table, the method further includes: Obtain a projection query request triggered by the target range of the lane image, wherein the projection query request is used to request the target position of the target element in the lane image located within the target range on the first coordinate system; Obtain multiple target elements located within the target range in the lane image; By querying the pixel projection mapping table using the multiple coordinate positions of the multiple candidate elements in the second coordinate system, the multiple candidate coordinate positions of the multiple target elements in the first coordinate system are obtained. The coordinate positions outside the predetermined detection range among the multiple candidate coordinate positions are filtered to obtain at least one target position.
6. The method according to claim 1, characterized in that, After acquiring the lane image obtained by the image acquisition device from a top-down perspective of the lane scene, the method further includes: The lane image is input into the target detection model, wherein the target detection model is a neural network model trained from multiple image samples for detecting the target element; Obtain the detection results output by the target detection model.
7. The method according to claim 6, The image acquisition device comprises N acquisition sub-devices, characterized in that, After inputting the lane image into the target detection model, the method further includes: The N lane images acquired by the N acquisition sub-devices are integrated into a multidimensional tensor lane image, wherein the multidimensional tensor includes a tensor corresponding to the dimension of the input size received by the image feature encoding layer of the target detection model, and a tensor corresponding to the dimension of the quantity N, where N is an integer greater than 1; The image features corresponding to the lane image of the multidimensional tensor are extracted through the image feature encoding layer of the target detection model to obtain first feature maps and second feature maps at different scales. Upsampling is performed on the first or second feature map in the first and second feature maps of different scales to obtain the first and second feature maps of the same scale. The first and second feature maps of the same scale are accumulated and fused to obtain the first target feature map; The first target feature map is downsampled to obtain the second target feature map; The step of obtaining the detection result output by the target detection model includes: processing the second target feature map through the subsequent model structure of the target detection model to obtain the detection result.
8. The method according to claim 7, characterized in that, The step of processing the second target feature map through the subsequent model structure of the target detection model to obtain the detection result includes: The second target feature map is processed by the element-view frustum transformation layer of the target detection model to obtain a top-view feature map, wherein the top-view feature map is used to represent the first positional relationship, and the subsequent model structure includes the element-view frustum transformation layer; The detection result is obtained by further processing the top-view feature map using at least one target model structure, wherein the subsequent model structure includes the at least one target model structure.
9. The method according to claim 8, characterized in that, The step of processing the top-view feature map using at least one target model structure to obtain the detection result includes: The view feature pooling layer of the target detection model is used to sum the features that fall at the same pixel position in the top view feature map to obtain the pooled top view feature map. The at least one target model structure includes the view feature pooling layer. The pooled top-view feature map is processed using at least one first model structure to obtain the detection result, wherein the at least one target model structure includes the at least one first model structure.
10. The method according to claim 8, characterized in that, The step of processing the top-view feature map using at least one target model structure to obtain the detection result includes: The temporal feature fusion layer of the target detection model is used to temporally align the top-view feature map at the current moment with the top-view feature map at least one historical moment to obtain aligned top-view feature maps at different moments. The at least one target model structure includes the temporal feature fusion layer. The temporal feature fusion layer stitches together the top-view feature maps from different times to obtain a top-view feature map that incorporates temporal information. The detection result is obtained by further processing the top-view feature map that incorporates time-series information through at least one second model structure, wherein the at least one target model structure includes the at least one second model structure.
11. The method according to claim 8, characterized in that, The step of processing the top-view feature map using at least one target model structure to obtain the detection result includes: The target detection model output layer is used to process the top-view feature map to obtain the three-dimensional position, element size, element orientation angle, and element velocity of the target element in the first coordinate system. The detection results include the three-dimensional position, element size, element orientation angle, and element velocity.
12. The method according to any one of claims 1 to 11, characterized in that, Obtaining the second coordinate position of the at least one lane element in the second coordinate system includes: Using a projection matrix, the at least one lane element is projected from the first coordinate system to the second coordinate system to obtain the projected position of the at least one lane element, wherein the second coordinate position includes the projected position, and the projection matrix represents the mapping relationship between the first coordinate system and the second coordinate system.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program is executed by an electronic device to perform the method according to any one of claims 1 to 12.
14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 12.
15. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 12 through the computer program.