A road-end BEV perception method, system, computer device, and storage medium
By extracting and converting the road-end camera image, using a multi-layer perceptron and deformable transformer to generate BEV foreground segmentation images, the problems of diverse installation positions and environmental changes of the road-end camera are solved, and high-precision BEV perception is achieved.
Patent Information
- Application Number
- CN202410227393.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-02-29
AI Technical Summary
The existing BEV perception algorithm relies on camera calibration parameters, making it difficult to adapt to the problems of diverse position positions and large changes in view angles and environment of the road-end camera installation, resulting in low detection accuracy.
By extracting the feature of the road-end camera image, using a multi-layer perceptron and deformable transformer for image conversion and fusion, BEV foreground segmentation images are generated, and BEV perception does not depend on camera calibration parameters are realized.
It improves the flexibility and detection accuracy of BEV perception on the roadside, and can accurately identify target objects in different installation positions and environments.
Smart Images

Figure CN118230312B_ABST
Abstract
Description
Background Art
[0002] Currently, most BEV perception algorithms are designed for vehicle terminals. Due to differences in perspectives, poses, etc. between roadside cameras and vehicle-mounted cameras, the performance of directly applying these algorithms to the roadside is poor.
[0003] A key problem in BEV perception is to construct BEV features from images. Existing methods usually construct BEV features through camera calibration parameters, and the performance of the algorithms is greatly affected by the accuracy of the camera calibration parameters. However, the installation poses of roadside cameras are diverse, and the perspectives, heights, and environments vary greatly, making it difficult to obtain accurate calibration parameters.
[0004] Therefore, there is an urgent need to provide a technical solution to solve the above problems. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a roadside BEV perception method, system, computer device, and storage medium.
[0006] In a first aspect, the present invention provides a roadside BEV perception method, and the technical solution of this method is as follows:
[0007] Extract features from the to-be-detected image collected by the roadside camera to obtain a feature image and convert it into a BEV feature image;
[0008] When the camera calibration parameters of the roadside camera are not provided, according to the pixel coordinates projected by the bottom border of the 3D target bounding box extracted from the feature image in the feature image, obtain a foreground segmentation perspective view and convert it into a BEV foreground segmentation image;
[0009] Decode and perform 3D detection on the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image respectively to obtain a BEV perception result.
[0010] The beneficial effects of a roadside BEV perception method of the present invention are as follows:
[0011] The method of the present invention can perform roadside BEV perception without relying on camera calibration parameters, has a more flexible feature, and higher detection accuracy.
[0012] On the basis of the above solution, a roadside BEV perception method of the present invention can also be improved as follows.
[0013] In an optional manner, the step of extracting features from the to-be-detected image collected by the roadside camera to obtain a feature image includes:
[0014] Use the ResNet-50 network to extract features from the to-be-detected image to obtain the feature image.
[0015] In an alternative manner, the step of converting the feature image into the BEV feature image includes:
[0016] Using a multi-layer perceptron to convert the feature image into the BEV feature image.
[0017] In an alternative manner, the BEV perception result includes: a BEV foreground segmentation prediction image, 3D prediction information of the target object, pixel coordinates of the 3D bounding box projected onto the image, 3D coordinates of the center of each side of the 3D bounding box, and the object type;
[0018] The step of performing decoding and 3D detection on the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain the BEV perception result includes:
[0019] Using a deformable transformer to decode the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain a BEV foreground segmentation prediction image, and performing 3D detection on the fused image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain the 3D prediction information of the target object, pixel coordinates of the 3D bounding box projected onto the image, 3D coordinates of the center of each side of the 3D bounding box, and the object type.
[0020] In an alternative manner, it further includes:
[0021] When the camera calibration parameters of the roadside camera are provided, according to the pixel coordinates of the bottom side of the 3D target bounding box extracted from the feature image projected onto the feature image, a foreground segmentation perspective view is obtained, and according to the pixel coordinates of the bottom side of the 3D target bounding box projected onto the feature image, the height of the roadside camera from the ground, and the camera calibration parameters, a transformation matrix between the foreground segmentation perspective view and the BEV foreground segmentation image is calculated, and according to the transformation matrix, the foreground segmentation perspective view is converted into the BEV foreground segmentation image.
[0022] In the above alternative manner, when providing the camera calibration parameters, the inverse perspective transformation method is adopted to more accurately obtain the BEV foreground segmentation image, and the detection process is more flexible and has higher accuracy.
[0023] In a second aspect, the present invention provides a roadside BEV perception system, and the technical solution of this system is as follows:
[0024] It includes: a processing module, a conversion module, and a perception module;
[0025] The processing module is configured to: extract features from the to-be-tested image collected by the roadside camera to obtain a feature image and convert it into a BEV feature image;
[0026] The conversion module is configured to: when the camera calibration parameters of the roadside camera are not provided, obtain a foreground segmentation perspective view based on the pixel coordinates projected by the bottom border of the 3D target bounding box extracted from the feature image in the feature image and convert it into a BEV foreground segmentation image;
[0027] The perception module is configured to: respectively decode and perform 3D detection on the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain a BEV perception result.
[0028] The beneficial effects of a roadside BEV perception system of the present invention are as follows:
[0029] The system of the present invention can perform roadside BEV perception without relying on camera calibration parameters, has a more flexible feature and higher detection accuracy.
[0030] On the basis of the above solution, a roadside BEV perception system of the present invention can also be improved as follows.
[0031] In an optional manner, the step of extracting features from the to-be-tested image collected by the roadside camera in the processing module to obtain a feature image includes:
[0032] Using a ResNet-50 network to extract features from the to-be-tested image to obtain the feature image.
[0033] In an optional manner, the step of converting the feature image into the BEV feature image in the processing module includes:
[0034] Using a multi-layer perceptron to convert the feature image into the BEV feature image.
[0035] In a third aspect, the technical solution of a computer device of the present invention is as follows:
[0036] It includes a memory, a processor, and a program stored on the memory and running on the processor. When the processor executes the program, it implements the steps of the roadside BEV perception method of the present invention.
[0037] In a fourth aspect, the technical solution of a computer-readable storage medium provided by the present invention is as follows:
[0038] Instructions are stored in the computer-readable storage medium. When the computer-readable storage medium reads the instructions, it causes the computer-readable storage medium to execute the steps of the roadside BEV perception method of the present invention.
[0039] The above description is only an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention, it can be implemented according to the content of the specification. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings are only used to illustrate the embodiments and are not considered as limiting the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0041] Figure 1 is a schematic flow chart of an embodiment of a roadside BEV perception method of the present invention;
[0042] Figure 2 is a schematic structural diagram of an embodiment of a roadside BEV perception system of the present invention;
[0043] Figure 3 is a schematic structural diagram of an embodiment of a computer device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The exemplary embodiments of the present invention will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein.
[0045] Figure 1 shows a schematic flow chart of an embodiment of a roadside BEV perception method provided by the present invention. As Figure 1 shown, the following steps are included:
[0046] S1. Extract features from the to-be-detected image collected by the roadside camera to obtain a feature image and convert it into a BEV feature image.
[0047] Among them, the roadside camera is a camera installed at any position on the roadside. The to-be-detected image is the image captured by the roadside camera, and the image may contain target objects (people, vehicles, buildings, etc.). The to-be-detected image is the image after feature extraction, and the BEV feature image is the feature image in the BEV view.
[0048] S2. When the camera calibration parameters of the roadside camera are not provided, obtain a foreground segmentation perspective view according to the pixel coordinates projected by the bottom border of the 3D target bounding box extracted from the feature image in the feature image, and convert it into a BEV foreground segmentation image.
[0049] Before executing S2, it is necessary to determine whether the roadside camera provides camera calibration parameters. If not, execute S2. In S2, an image detection algorithm is used to extract the pixel coordinates of the bottom border of the 3D object bounding box in the feature image, projected into the feature image. Then, a multi-layer perceptron (MLP) is used to convert the foreground segmentation perspective image into a BEV foreground segmentation image.
[0050] It should be noted that compared with the technical solution of directly obtaining the BEV foreground segmentation image from image features, the technical solution in S2 first predicts the pixel coordinates of the bottom border of the 3D target bounding box projected in the feature image and then converts it into the BEV foreground segmentation image. This method can explicitly supervise the process of foreground segmentation perspective map generation, better locate the position of the target in the foreground segmentation perspective map, and thus obtain a more accurate BEV foreground segmentation image.
[0051] S3. Decode and 3D detect the image features obtained after the BEV feature image and the BEV foreground segmentation image are fused to obtain a BEV perception result.
[0052] Among them, the BEV perception results include but are not limited to: BEV foreground segmentation prediction image, 3D prediction information of the target object, pixel coordinates of the 3D bounding box projected to the image, the center 3D coordinates of each frame of the 3D bounding box, and the object type.
[0053] In an optional manner, the step of extracting features from an image to be tested captured by a roadside camera to obtain a feature image includes:
[0054] The ResNet-50 network is used to extract features of the image to be tested to obtain the feature image.
[0055] Among them, the ResNet-50 network is a pre-trained ResNet-50 network that uses a deep convolutional architecture to effectively extract high-level features of images.
[0056] In an optional manner, the step of converting the characteristic image into the BEV characteristic image includes:
[0057] The feature image is converted into the BEV feature image using a multi-layer perceptron.
[0058] A multi-layer perceptron (MLP) is used to convert the feature image into a BEV feature image. It should be noted that the process of converting extracted image features (feature images) into BEV view features (BEV feature images) using a multi-layer perceptron is well known in the art, and its principles and process will not be elaborated on here.
[0059] In an optional manner, S3 includes:
[0060] Using a deformable Transformer, decode the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain a BEV foreground segmentation prediction image, and perform 3D detection on the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain 3D prediction information of the target object, the pixel coordinates of the 3D bounding box projected onto this image, the central 3D coordinates of each side of the 3D bounding box, and the object type.
[0061] Among them, fusing the BEV feature image and the BEV foreground segmentation image can improve the accuracy of the final BEV perception result. The deformable Transformer is a Transformer in deep learning, which is used to decode an image and output a predicted BEV foreground segmentation (BEV foreground segmentation prediction image). The target object is the object detected in the image to be measured in this embodiment (such as people, vehicles, buildings, trees, etc.). The 3D prediction information includes: the size, position, orientation, etc. of the target object. The object types include: large vehicles, medium-sized vehicles, small vehicles, pedestrians, etc., and the basis for object type division is determined after performing a clustering algorithm calculation on big data.
[0062] It should be noted that the prediction of the pixel coordinates of the 3D bounding box of the target object projected onto this image can improve the detection ability of the algorithm for the target object. The calculation method of the 3D central coordinates of the target object can be appropriately selected according to the distance of the target object from the roadside camera. Since the features at the top of the object are more obvious, the prediction of the top center of the 3D bounding box of the target object is relatively more accurate. Therefore, compared with directly predicting the 3D central coordinates of the object, indirect calculation is more accurate. Therefore, when the main part of the target object in the image is the top of the object, the 3D center of the target object can be indirectly obtained by subtracting the predicted height of the target object from the top of the 3D bounding box of the target object. The detection ability of the object size (large vehicles, medium-sized vehicles, small vehicles, pedestrians, etc.) has a great impact on the BEV perception result.
[0063] In an alternative manner, it further includes:
[0064] When the camera calibration parameters of the roadside camera are provided, according to the pixel coordinates of the bottom border of the 3D target bounding box extracted from the feature image projected onto the feature image, obtain a foreground segmentation perspective view, and according to the pixel coordinates of the bottom border of the 3D target bounding box projected onto the feature image, the height of the roadside camera from the ground, and the camera calibration parameters, calculate the transformation matrix between the foreground segmentation perspective view and the BEV foreground segmentation image, and according to the transformation matrix, convert the foreground segmentation perspective view into the BEV foreground segmentation image.
[0065] Among them, in the inference stage, if the road-side camera has provided the camera calibration parameters, the foreground segmentation under the perspective view is accurately calculated according to the camera calibration parameters to obtain the BEV foreground segmentation. Specifically, for the image captured by the road-side camera, it can be assumed that the road surface has the same height H (the height of the road-side camera from the ground). Through the road surface height H, pixel coordinates (u, v), and camera calibration parameters, the conversion matrix between the foreground segmentation perspective view and the foreground segmentation under the BEV view is calculated, and according to the conversion matrix, the foreground segmentation under the foreground segmentation perspective view is converted into the foreground segmentation under the BEV view (BEV foreground segmentation image). Compared with using MLP conversion, using IPM (inverse perspective transform) to convert the BEV view introduces geometric prior information and improves the accuracy of BEV foreground segmentation.
[0066] The technical solution of this embodiment can perform road-side BEV perception without relying on camera calibration parameters, and is more flexible and has higher detection accuracy.
[0067] Figure 2 FIG. 2 shows a schematic structural diagram of an embodiment of a roadside BEV perception system 200 provided by the present invention. Figure 2 As shown, the system 200 includes: a processing module 210, a conversion module 220 and a perception module 230;
[0068] The processing module 210 is used to: extract features from the image to be tested collected by the roadside camera, obtain a feature image and convert it into a BEV feature image;
[0069] The conversion module 220 is configured to: when the camera calibration parameters of the roadside camera are not provided, obtain a foreground segmentation perspective image based on the pixel coordinates of the bottom border of the 3D object bounding box extracted from the feature image projected in the feature image and convert the image into a BEV foreground segmentation image;
[0070] The perception module 230 is used to respectively decode and perform 3D detection on the image features obtained after the fusion of the BEV feature image and the BEV foreground segmentation image to obtain a BEV perception result.
[0071] In an optional manner, the step of extracting features from the image to be tested captured by the roadside camera to obtain a feature image in the processing module 210 includes:
[0072] The ResNet-50 network is used to extract features of the image to be tested to obtain the feature image.
[0073] In an optional manner, the step of converting the characteristic image into the BEV characteristic image in the processing module 210 includes:
[0074] Using a multi-layer perceptron, convert the feature image into the BEV feature image.
[0075] In an alternative embodiment, the BEV perception result includes: a BEV foreground segmentation prediction image, 3D prediction information of the target object, pixel coordinates of the 3D bounding box projected onto the image, the central 3D coordinates of each side of the 3D bounding box, and the object type; specifically, the perception module 230 is configured to:
[0076] Using a deformable transformer, decode the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain a BEV foreground segmentation prediction image, and perform 3D detection on the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain the 3D prediction information of the target object, the pixel coordinates of the 3D bounding box projected onto the image, the central 3D coordinates of each side of the 3D bounding box, and the object type.
[0077] In an alternative embodiment, it further includes: a second conversion module; the second conversion module is configured to:
[0078] When the camera calibration parameters of the roadside camera are provided, according to the pixel coordinates of the bottom border of the 3D target bounding box extracted from the feature image projected onto the feature image, obtain a foreground segmentation perspective view, and according to the pixel coordinates of the bottom border of the 3D target bounding box projected onto the feature image, the height of the roadside camera from the ground, and the camera calibration parameters, calculate the conversion matrix between the foreground segmentation perspective view and the BEV foreground segmentation image, and according to the conversion matrix, convert the foreground segmentation perspective view into the BEV foreground segmentation image.
[0079] The technical solution of this embodiment can perform roadside BEV perception without relying on camera calibration parameters, has the characteristics of greater flexibility and higher detection accuracy.
[0080] For the parameters and steps of each module in the above-mentioned roadside BEV perception system 200 of this embodiment to implement corresponding functions, reference can be made to the parameters and steps in the embodiment of the roadside BEV perception method in the above text, which will not be elaborated here.
[0081] As Figure 3 shown, a computer device 300 according to an embodiment of the present invention, the computer device 300 includes a processor 320, the processor 320 is coupled to a memory 310, and at least one computer program 330 is stored in the memory 310. The at least one computer program 330 is loaded and executed by the processor 320 to enable the computer device 300 to implement any one of the above-mentioned roadside BEV perception methods. Specifically:
[0082] The computer device 300 can vary significantly due to differences in configuration or performance. It may include one or more processors 320 (Central Processing Units, CPUs) and one or more memories 310. Among them, at least one computer program 330 is stored in the one or more memories 310. The at least one computer program 330 is loaded and executed by the one or more processors 320, enabling the computer device 300 to implement any of the roadside BEV perception methods provided in the above embodiments. Of course, the computer device 300 may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The computer device 300 may further include other components for implementing device functions, which will not be elaborated here.
[0083] A computer-readable storage medium according to an embodiment of the present invention stores at least one computer program, and the at least one computer program is loaded and executed by a processor to enable a computer to implement any of the above roadside BEV perception methods.
[0084] Optionally, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0085] In an exemplary embodiment, a computer program product or computer program is further provided. The computer program product or computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute any of the above roadside BEV perception methods.
[0086] It should be noted that the terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, rather than to limit a specific order or sequence. Under appropriate circumstances, the order of use of similar objects may be interchanged so that the embodiments of the present application described here can be implemented in an order other than the illustrated or described order.
[0087] Those skilled in the art of the present technology know that the present invention can be implemented as a system, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, which is generally referred to as "circuit", "module", or "system" herein. In addition, in some embodiments, the present invention can also be implemented in the form of a computer program product in one or more computer-readable media, which contain computer-readable program code.
[0088] Any combination of one or more computer-readable media can be adopted. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in combination with an instruction execution system, apparatus, or device.
[0089] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limitations on the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A road-end BEV perception method, characterized in that, Including: Using the ResNet-50 network to extract features from the to-be-tested image collected by the roadside camera to obtain a feature image, and using a multi-layer perceptron to convert the feature image into a BEV feature image; When the camera calibration parameters of the roadside camera are not provided, according to the pixel coordinates projected by the bottom border of the 3D target bounding box extracted from the feature image in the feature image, obtain a foreground segmentation perspective view and convert it into a BEV foreground segmentation image; wherein, using an image detection algorithm, extract the pixel coordinates projected by the bottom border of the 3D target bounding box in the feature image in the feature image, and use a multi-layer perceptron to convert the foreground segmentation perspective view into the BEV foreground segmentation image; Perform decoding and 3D detection on the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image respectively to obtain a BEV perception result; The BEV perception result includes: a BEV foreground segmentation prediction image, 3D prediction information of the target object, pixel coordinates of the 3D bounding box projected onto this image, center 3D coordinates of each border of the 3D bounding box, and the object type; The step of performing decoding and 3D detection on the image obtained after fusing the BEV feature image and the BEV foreground segmentation image respectively to obtain a BEV perception result includes: Using a deformable transformer to decode the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain a BEV foreground segmentation prediction image, and performing 3D detection on the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain the 3D prediction information of the target object, pixel coordinates of the 3D bounding box projected onto this image, center 3D coordinates of each border of the 3D bounding box, and the object type; wherein, the deformable transformer is a transformer in deep learning, used to decode the image and output the predicted BEV foreground segmentation prediction image; the target object is the object detected in the to-be-tested image; the 3D prediction information includes: the size, position, and orientation of the target object; the object type includes: large vehicle, medium vehicle, small vehicle, pedestrian, and the object type is determined based on the calculation of the clustering algorithm for big data; It also includes: when the camera calibration parameters of the roadside camera are provided, according to the pixel coordinates projected by the bottom border of the 3D target bounding box extracted from the feature image in the feature image, obtain a foreground segmentation perspective view, and according to the pixel coordinates projected by the bottom border of the 3D target bounding box in the feature image, the height of the roadside camera from the ground, and the camera calibration parameters, calculate the transformation matrix between the foreground segmentation perspective view and the BEV foreground segmentation image, and according to the transformation matrix, convert the foreground segmentation perspective view into the BEV foreground segmentation image; wherein, the BEV foreground segmentation image is obtained by using inverse perspective transformation.
2. A road-end BEV perception system, characterized in that, Including: A processing module, a conversion module, and a perception module; The processing module is used to: extract features from the to-be-tested image collected by the roadside camera using the ResNet-50 network to obtain a feature image, and use a multi-layer perceptron to convert the feature image into a BEV feature image; The conversion module is used to: when the camera calibration parameters of the roadside camera are not provided, obtain a perspective view of foreground segmentation and convert it into a BEV foreground segmentation image according to the pixel coordinates projected in the feature image by the bottom border of the 3D target bounding box extracted from the feature image; wherein, use an image detection algorithm to extract the pixel coordinates projected in the feature image by the bottom border of the 3D target bounding box in the feature image, and use a multi-layer perceptron to convert the perspective view of foreground segmentation into the BEV foreground segmentation image; The perception module is used to: respectively decode and perform 3D detection on the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain a BEV perception result; The BEV perception result includes: a BEV foreground segmentation prediction image, 3D prediction information of the target object, pixel coordinates of the 3D bounding box projected onto this image, central 3D coordinates of each border of the 3D bounding box, and the object type; specifically, the perception module is used to: Use a deformable transformer to decode the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain a BEV foreground segmentation prediction image, and perform 3D detection on the image features obtained after fusing the BEV feature image and the BEV foreground segmentation image to obtain the 3D prediction information of the target object, pixel coordinates of the 3D bounding box projected onto this image, central 3D coordinates of each border of the 3D bounding box, and the object type; wherein, the deformable transformer is a transformer in deep learning, used to decode an image and output a predicted BEV foreground segmentation prediction image; the target object is an object detected in the to-be-tested image; the 3D prediction information includes: the size, position, and orientation of the target object; the object type includes: large vehicle, medium vehicle, small vehicle, pedestrian, and the object type is determined based on the calculation of a clustering algorithm for big data; It further includes: a second conversion module; the second conversion module is used to: When the camera calibration parameters of the roadside camera are provided, obtain a perspective view of foreground segmentation according to the pixel coordinates projected in the feature image by the bottom border of the 3D target bounding box, and calculate a conversion matrix between the perspective view of foreground segmentation and the BEV foreground segmentation image according to the pixel coordinates projected in the feature image by the bottom border of the 3D target bounding box, the height of the roadside camera from the ground, and the camera calibration parameters, and convert the perspective view of foreground segmentation into the BEV foreground segmentation image according to the conversion matrix; wherein, the BEV foreground segmentation image is obtained by using inverse perspective transformation.
3. A computer device, characterized in that, The computer device includes a processor, the processor is coupled to a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by the processor so that the computer device implements the roadside BEV perception method as described in claim 1.
4. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium. The at least one computer program is loaded and executed by a processor so that the computer-readable storage medium implements the roadside BEV perception method as described in claim 1.