Image generation method, electronic device, storage medium, program product, and vehicle

By correcting the vehicle image and building a BEV data set, training an adversarial network to generate BEV images, solving the problem of stretching deformation at the field of view junction in the prior art, and achieving more realistic and highly applicable BEV image generation.

CN120163887APending Publication Date: 2025-06-17BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510143418.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The BEV images generated by the prior art have tensile deformation at the junction of the field of view, which affects their application.

Method used

By performing correction processing on the first image set of the vehicle, the first target image is determined, and the BEV data set is constructed in combination with the second image acquired by the fisheye camera, and the adversarial network is trained to generate a BEV image not affected by the fisheye image correction.

Benefits of technology

The generated BEV image does not have tensile deformation at the junction of the field of view, and is closer to the real top view, which improves the applicability of the BEV image and can generate a rear top view image blocked by the object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163887A_ABST
    Figure CN120163887A_ABST
Patent Text Reader

Abstract

The invention provides a BEV image generation method, electronic equipment, a storage medium, a program product and a vehicle. The method comprises the following steps: determining a first target image according to a first image set of the vehicle; determining a BEV data set according to the first target image and a collected second image; determining an adversarial network according to the BEV data set; and determining a BEV image according to the adversarial network and the second image. The top view in the first image set of the vehicle is corrected to obtain the first target image, the first target image and the second image collected by the fisheye camera on the vehicle serve as the BEV data set, the BEV data set is used for training the adversarial network, and the BEV image is obtained in combination with the second image. Therefore, the generated BEV image is not affected by edge stretching of the fisheye image, tensile deformation does not exist at the FOV junction, the BEV image is closer to a real top view image, and the applicability of the BEV image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent driving, and in particular, to a method for generating a BEV image, an electronic device, a computer-readable storage medium, a computer program product, and a vehicle. Background Art

[0002] BEV (Bird's Eye View) images are very important for intelligent driving. In related technologies, through correction steps such as camera calibration, image transformation, and image stitching, multiple images are converted into BEV images, or correction operations such as establishing a mapping relationship between the camera source image and the BEV image and projecting the camera image into the BEV space can generate seamless 360-degree BEV images.

[0003] However, the BEV images generated by the above methods are affected by the characteristics of the images captured by the fish-eye camera, and there is a large stretching deformation at the junction of the field of view (FOV), which is quite different from the real BEV image, reducing the applicability of the BEV image. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art.

[0005] To this end, an object of the present invention is to propose a method for generating a BEV image, which avoids the influence of edge stretching caused by the correction of fish-eye camera images. The generated BEV image has no stretching deformation at the junction of the FOV and is more realistic, and can generate a rear overhead image blocked by an object, improving the applicability of the BEV image.

[0006] To this end, a second object of the present invention is to propose an electronic device.

[0007] To this end, a third object of the present invention is to propose a computer-readable storage medium.

[0008] To this end, a fourth object of the present invention is to propose a computer program product.

[0009] To this end, a fifth object of the present invention is to propose a vehicle.

[0010] To achieve the above object, an embodiment of the first aspect of the present invention proposes a method for generating a BEV image, the method comprising: determining a first target image according to a first image set of a vehicle; determining a BEV data set according to the first target image and a second image collected; determining an adversarial network according to the BEV data set; and determining a BEV image according to the adversarial network and the second image.

[0011] The BEV image generation method according to an embodiment of the present invention corrects the top view in the first image set of the vehicle to obtain a first target image with the vehicle angle, position, and scale normalized. The first target image and the second image collected by the fisheye camera on the vehicle are used as the BEV data set, which enriches the training sample features of the adversarial network. The adversarial network is trained using the BEV data set. Based on the trained adversarial network, the second image can be used to generate a BEV image, so that the generated BEV image is not affected by the edge stretching caused by fisheye image correction, there is no stretching deformation at the FOV junction, and it is closer to the real top view image, improving the applicability of the BEV image. Moreover, a rear top view image blocked by an object can be generated. The adversarial network introduces an STN structure, enabling the network to autonomously learn the conversion process from the fisheye view to the top view space, saving the cost of feature detection and calculation of the transformation model compared to traditional methods.

[0012] In some embodiments, determining the first target image according to the first image set of the vehicle includes: obtaining the first image in the first image set and an initial correction parameter; correcting the first image and the remaining images in the first image set according to the initial correction parameter to obtain a first corrected image and the remaining preliminarily corrected images; determining the first target image according to the remaining preliminarily corrected images and the first corrected image.

[0013] In some embodiments, determining the first target image according to the remaining preliminarily corrected images and the first corrected image includes: obtaining detection feature points of an interested region in a target preliminarily corrected image among the remaining preliminarily corrected images; determining a correction parameter according to the detection feature points and a preset correction parameter formula; correcting the target preliminarily corrected image according to the correction parameter; storing the corrected target preliminarily corrected image and the first corrected image to determine the first target image.

[0014] In some embodiments, after storing the corrected target preliminarily corrected image and the first corrected image to determine the first target image, it further includes: if the target preliminarily corrected image is not a preset image, updating the interested region; correcting the remaining target preliminarily corrected images in the remaining preliminarily corrected images according to the updated interested region.

[0015] In some embodiments, obtaining the initial correction parameter includes: obtaining vehicle feature points; obtaining the initial correction parameter according to the vehicle feature points.

[0016] In some embodiments, when obtaining the initial correction parameter according to the vehicle feature points, substituting the vehicle feature points into the following formula: …(1) …(2) …(3) Among them, the is the vehicle deflection angle, the and are the vehicle feature points, the is the vehicle center coordinate, the is the marking length, the and are the distances of the vehicle feature points in the X upward direction, and the and are the distances of the vehicle feature points in the Y direction.

[0017] In some embodiments, determining the BEV dataset according to the first target image and the acquired second image includes: acquiring the second image when the vehicle is at a preset form speed; using the first target image and the second image as the BEV dataset.

[0018] In some embodiments, when determining the adversarial network according to the BEV dataset, the following formula is introduced:

[0019] Among them, the is the total loss function of the adversarial network, the and are empirical input values, the is the adversarial loss function, the is the perceptual loss function, the is the feature matching loss function, the is the generator, and the is the discriminator.

[0020] In some embodiments, when determining the adversarial loss function, the following formula is introduced:

[0021] Among them, the is the adversarial loss function, the is the undistorted image of the second image, the is the first target image, the is that the input first target image and the undistorted image of the second image conform to the training data distribution, the is the discriminator output result, is the generator output result.

[0022] In some embodiments, when determining the perceptual loss, the following formula is introduced:

[0023] Among them, the is the perceptual loss function, the is the second image undistorted image, the is the first target image, the is that the input first target image and the second image undistorted image conform to the training data distribution, the is a variable weight used to measure the importance of each layer i, the is the feature extracted by the first target image passing through the layer network, the is the feature extracted by the output result of the generator passing through the layer network.

[0024] In some embodiments, when determining the feature matching loss function, the following formula is introduced:

[0025] Among them, the is the feature matching loss function, is a variable weight used to measure the importance of each layer i, the is the feature extracted by the first target image passing through the discriminator at the scale, the layer discriminator, the is the feature extracted by the output result of the generator passing through the discriminator at the scale, the layer discriminator.

[0026] In some embodiments, determining the BEV image according to the adversarial network and the second image includes: performing undistortion processing on the second image to obtain a second target image; determining the BEV image according to the second target image and the adversarial network.

[0027] To achieve the above object, an embodiment of the second aspect of the present invention proposes an electronic device, the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a BEV image generation program executable by the at least one processor, and when the BEV image generation program is executed by the at least one processor, the at least one processor executes the BEV image generation method as described in the above embodiment.

[0028] An electronic device according to an embodiment of the present invention uses the BEV image generation method described in the above embodiment. By performing a correction process on the top view in the first image set of the vehicle, a first target image with the vehicle angle, position, and scale normalized is obtained. The first target image and the second image collected by the fisheye camera on the vehicle are used as the BEV data set, enriching the training sample features of the adversarial network. Using the BEV data set to train the adversarial network, based on the trained adversarial network, using the second image, the generation of the BEV image can be achieved. The generated BEV image is not affected by the edge stretching caused by the fisheye image correction, there is no stretching deformation at the FOV junction, it is closer to the real top view image, improving the applicability of the BEV image, and the rear top view image blocked by objects can be generated. Moreover, the adversarial network introduces the STN structure, enabling the network to autonomously learn the conversion process from the fisheye view to the top view space, saving the cost of feature detection and calculating the transformation model compared with the traditional method.

[0029] To achieve the above object, an embodiment of the third aspect of the present invention proposes a computer-readable storage medium, on which a BEV image generation program is stored. When the BEV image generation program is executed by a processor, the device installed with the BEV image generation program realizes the BEV image generation method described in the above embodiment.

[0030] To achieve the above object, an embodiment of the fourth aspect of the present invention proposes a computer program product, which includes a computer program. When the computer program is executed by a processor, the BEV image generation method described in the above embodiment is realized.

[0031] A computer program product according to an embodiment of the present invention obtains a first target image with the vehicle angle, position, and scale normalized by performing a correction process on the top view in the first image set of the vehicle. The first target image and the second image collected by the fisheye camera on the vehicle are used as the BEV data set, enriching the training sample features of the adversarial network. Using the BEV data set to train the adversarial network, based on the trained adversarial network, using the second image, the generation of the BEV image can be achieved. The generated BEV image is not affected by the edge stretching caused by the fisheye image correction, there is no stretching deformation at the FOV junction, it is closer to the real top view image, improving the applicability of the BEV image, and the rear top view image blocked by objects can be generated. Moreover, the adversarial network introduces the STN structure, enabling the network to autonomously learn the conversion process from the fisheye view to the top view space, saving the cost of feature detection and calculating the transformation model compared with the traditional method.

[0032] To achieve the above object, an embodiment of the fifth aspect of the present invention proposes a vehicle, which includes: the electronic device described in the above embodiment.

[0033] According to the vehicle of the embodiment of the present invention, by performing correction processing on the top view in the first image set of the vehicle, a first target image with normalized vehicle angle, position, and scale is obtained. The first target image and the second image collected by the fisheye camera on the vehicle are used as the BEV data set, enriching the training sample features of the adversarial network. Using the BEV data set to train the adversarial network, based on the trained adversarial network, using the second image, the generation of the BEV image can be achieved, so that the generated BEV image is not affected by the edge stretching caused by fisheye image correction, there is no stretching deformation at the FOV junction, and it is closer to the real top view image, improving the applicability of the BEV image. Moreover, the rear top view image blocked by an object can be generated. And the adversarial network introduces the STN structure, enabling the network to autonomously learn the conversion process from the fisheye view to the top view space, saving the cost of feature detection and calculating the transformation model compared with the traditional method.

[0034] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Brief Description of the Drawings

[0035] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where: Figure 1 is a flowchart of a BEV image generation method according to an embodiment of the present invention; Figure 2 is a schematic diagram of vehicle correction parameters according to an embodiment of the present invention; Figure 3 is a flowchart of constructing a BEV data set according to an embodiment of the present invention; Figure 4 is a flowchart of generating a BEV image from a fisheye image according to an embodiment of the present invention; Figure 5 is a flowchart of image processing for images collected by a drone according to an embodiment of the present invention; Figure 6 is a flowchart of a BEV image generation method according to another embodiment of the present invention; Figure 7 is a block diagram of an electronic device according to an embodiment of the present invention; Figure 8 is a block diagram of a vehicle according to an embodiment of the present invention.

[0036] Reference Signs: Processor 101; Memory 102; Electronic device 100; Vehicle 99. Detailed Description of the Embodiments

[0037] The embodiments described with reference to the accompanying drawings are exemplary, and the embodiments of the present invention will be described in detail below.

[0038] A BEV image is an image that presents a scene in a bird's-eye view. BEV images are usually obtained by cameras or sensors to acquire surrounding environmental information and present it from an overhead perspective. They are widely used in the field of intelligent driving technology and are often applied to tasks such as intelligent parking assistance, environmental perception, and path planning. BEV images can improve the driver's field of vision and perception ability, enhance the vehicle's environmental perception ability, and provide important data support for the intelligent driving system, thereby improving driving safety and driving efficiency, and playing a very important role in the field of intelligent driving.

[0039] Among them, the camera is generally a fish-eye camera. Through an extremely wide-angle lens and a special projection method, objects within a field of view of 180 degrees or more can be captured in the image, and a circular field-of-view effect is presented in the image.

[0040] In related technologies, for example, through correction steps such as camera calibration, image transformation, and image stitching, multiple images are converted into BEV images, or operations such as establishing a mapping relationship between the camera source image and the BEV image and projecting the camera image into the BEV space can generate a seamless 360-degree BEV image.

[0041] However, when generating BEV images using the above methods, due to the large radial distortion of the images captured by the fish-eye camera, the objects captured may have significant deformation. After correcting the images captured by the fish-eye camera, the closer to the edge, the more obvious the stretching deformation after correction. Therefore, for the BEV image generated based on the corrected image of the fish-eye camera, there is a large stretching deformation at the junction of the FOV, which is quite different from the real top view, reducing the applicability of the BEV image.

[0042] Accordingly, by adopting the BEV image generation method of the embodiment of the present invention, a first target image with normalized vehicle angle, position, and scale is obtained by correcting the top view in the first image set of the vehicle. The first target image and the second image collected by the fisheye camera on the vehicle are used as the BEV data set, enriching the training sample features of the adversarial network. Using the BEV data set to train the adversarial network, and based on the trained adversarial network, using the second image, the generation of the BEV image can be achieved. The generated BEV image is not affected by the edge stretching caused by fisheye image correction, there is no stretching deformation at the FOV junction, and it is closer to the real top view image, improving the applicability of the BEV image. Moreover, the rear top view image blocked by objects can be generated. And the adversarial network introduces the STN structure, enabling the network to autonomously learn the conversion process from the fisheye view to the top view space, saving the cost of feature detection and calculation of the transformation model compared with the traditional method.

[0043] The following combines Figures 1-6 to describe the BEV image generation method of the embodiment of the present invention.

[0044] As Figure 1 shown, it is a flowchart of the BEV image generation method of an embodiment of the present invention. The BEV image generation method of the embodiment of the present invention at least includes steps S1 - step S4.

[0045] Step S1, determine the first target image according to the first image set of the vehicle.

[0046] In the embodiment, the first image set is the top view of the vehicle collected during the data collection stage; the first target image is a more standard top view obtained by correcting the top view in the data collection stage during the image processing stage, and it is an image set. Determining the first target image according to the first image set of the vehicle. For example, using a drone to fly above the vehicle and following the vehicle, the camera shoots the top view of the vehicle vertically downward to obtain the first image set, and a series of correction processes are performed on the top view in the first image set to obtain the first target image, which involves steps such as inclination correction, feature point detection, cropping, region of interest division, and calculation of correction parameters to achieve the normalization of the vehicle angle, position, and scale in the top view.

[0047] Step S2, determine the BEV data set according to the first target image and the collected second image.

[0048] In an embodiment, the second image is an image collected by a camera on a vehicle. For example, a fisheye image is collected by four fisheye cameras on the vehicle, including images from four perspectives, namely, the front, rear, left, and right of the vehicle, that is, the second image. The first target image and the second image are used as the BEV dataset. Since the first target image is a set of real top-view images, training based on the first target image can ensure that the generated BEV image has no stretching distortion at the FOV junction and is closer to the real top-view image. Since the BEV dataset includes top-view images and fisheye images, the features of the dataset are more complete, providing data support for the training of the adversarial network and making the generated BEV image unaffected by the edge stretching caused by fisheye image correction.

[0049] Step S3: Determine the adversarial network according to the BEV dataset.

[0050] In an embodiment, the adversarial network includes a generator and a discriminator. The generator is mainly composed of convolutional layers, an STN (Spatial Transformer Networks), and a ResNet (Residual Network). The input is divided into four paths and can simultaneously receive the input of four undistorted fisheye images from the front, rear, left, and right. After feature extraction and spatial transformation, the generated BEV image is output. The discriminator then determines whether the BEV image generated by the generator is real, and the two are adversarially trained to improve the quality of the BEV image generated by the generator.

[0051] Perform undistortion processing on the second image in the BEV dataset. According to the undistorted second image and the first target image, combine with the loss function of the adversarial network to obtain the adversarial network. By using the BEV dataset to train the generative adversarial network, a non-deformed BEV image directly generated from the fisheye image can be obtained, making it closer to the real top view and solving the problem of BEV image deformation caused by fisheye images.

[0052] Step S4: Determine the BEV image according to the adversarial network and the second image.

[0053] In an embodiment, based on the trained adversarial network and STN, the BEV image can be generated using the second image, and a rear top-view image blocked by an object can also be generated. For example, input the four undistorted fisheye images (the second image) into the generator of the adversarial network. At this time, the generator has completed training and can generate the BEV image according to the undistorted fisheye image. The generator projects the fisheye image onto the top-view perspective to generate a BEV image without edge deformation. Moreover, the adversarial network introduces the STN structure, enabling the network to autonomously learn the conversion process from the fisheye view to the top-view space, saving the cost of feature detection and calculation of the transformation model compared with traditional methods.

[0054] According to the BEV image generation method of an embodiment of the present invention, a top view in a first image set of a vehicle is corrected to obtain a first target image with the vehicle angle, position, and scale normalized. The first target image and a second image collected by a fisheye camera on the vehicle are used as a BEV data set, enriching the training sample features of the adversarial network. The adversarial network is trained using the BEV data set. Based on the trained adversarial network, using the second image, a BEV image can be generated, such that the generated BEV image is not affected by the edge stretching caused by fisheye image correction, there is no stretching deformation at the FOV junction, it is closer to the real top view image, improving the applicability of the BEV image, and a rear top view image blocked by an object can be generated. Moreover, the adversarial network introduces an STN structure, enabling the network to autonomously learn the conversion process from the fisheye view to the top view space, saving the cost of feature detection and calculating the transformation model compared with traditional methods.

[0055] In some embodiments, determining the first target image according to the first image set of the vehicle includes: obtaining the first image in the first image set and an initial correction parameter; correcting the first image and the remaining images in the first image set according to the initial correction parameter to obtain a first corrected image and the remaining preliminarily corrected images; determining the first target image according to the remaining preliminarily corrected images and the first corrected image.

[0056] In an embodiment, as Figure 3 shown, it is a flowchart of constructing a BEV data set according to an embodiment of the present invention. The flowchart of constructing a BEV data set according to an embodiment of the present invention includes at least step S40 - step S42.

[0057] Step S40, prepare a test vehicle, a fisheye camera, a drone, and a camera.

[0058] Step S41, the test vehicle drives slowly, and the drone follows and shoots above the vehicle.

[0059] Step S42, perform image processing to normalize the vehicle angle, position, and scale in the top view.

[0060] Specifically, in step S40, in the preparation stage, prepare a test vehicle, a fisheye camera, a drone, and an ordinary camera; equip 4 fisheye cameras on the test vehicle, and the installation positions are respectively on the front car logo, the rear fender, and the left and right rearview mirrors. Each camera points in a different direction, enabling 360-degree coverage of the viewing angle; the test vehicle needs to deploy drone tracking markers at both the front and rear ends to assist the drone in tracking the test vehicle, and at the same time, it can also be used as feature points to provide a reference for subsequent image processing; the drone needs to be equipped with an ordinary camera, and the camera shoots horizontally downward; calibrate the camera to obtain the internal parameters, external parameters, and distortion parameters of the fisheye camera.

[0061] Step S41, in the data collection stage, the test vehicle travels at a relatively low speed within the parking lot, with the traveling speed being within 20 km / h; the drone follows the vehicle and flies above the vehicle, and the camera takes a top view of the vehicle perpendicular to the horizontal plane downward.

[0062] Step S42, in the image processing stage, the normalization of the vehicle angle, position, and scale in the top view is realized. Since it is impossible for the drone to remain relatively stationary with respect to the test vehicle during the shooting process, the direction, position, and scale of the test vehicle in the image will all change to a certain extent. In order for the top view to be used for the training of the network, these parameters need to be normalized, so the drone images must be corrected.

[0063] The first image set is the top view of the vehicle collected in the data collection stage; the initial correction parameters are the parameters calculated according to the feature points of the vehicle for correcting the image; the first target image is a more standard top view obtained after the top view in the data collection stage is corrected through the image processing stage, which is an image set; during the correction of the drone images, the first image in the first image set is obtained. For example, the first top view is obtained. The data set captured by the drone is a sequence of images with a time relationship. First, the first top view is taken for processing; the initial correction parameters are obtained. For example, according to the feature points that can represent the vehicle angle and position, the initial correction parameters are calculated to provide a data basis for the preliminary correction of the positive image; the first top view is rotated, scaled, and cropped according to the initial correction parameters to obtain the corrected top view, that is, the first corrected image. The remaining top views are rotated and scaled according to the initial correction parameters to obtain the preliminarily corrected top views, that is, the other preliminarily corrected images, so as to realize the preliminary correction of the top view of the vehicle collected in the data collection stage; the other preliminarily corrected images and the first corrected image are further corrected to obtain the first target image, so as to make the obtained first target image more standard, that is, the vehicle angle, position, and scale in the top view are all normalized.

[0064] In some embodiments, determining the first target image according to the other preliminarily corrected images and the first corrected image includes: obtaining the detected feature points of the region of interest in the target preliminarily corrected image among the other preliminarily corrected images; determining the correction parameters according to the detected feature points and the preset correction parameter formula; correcting the target preliminarily corrected image according to the correction parameters; storing the corrected target preliminarily corrected image and the first corrected image to determine the first target image.

[0065] In an embodiment, the target preliminary correction image is processed to narrow the range of detected feature points for the remaining preliminary correction images; an area of interest is selected on the top view to narrow the range of detected feature points, facilitating the calculation of correction parameters. Feature points are detected within the area of interest. For example, the SIFT (Scale-Invariant Feature Transform) corner detection algorithm is used to detect feature points to more accurately extract the feature points in the remaining preliminary correction images; the detected feature points and a preset correction parameter formula are used to determine the correction parameters. The preset correction parameter formula is the same as the calculation formula for the initial correction parameters. For example, assuming the detected feature point coordinates are P4( , ) and the corresponding coordinates of P5 are ( , ), and the correction parameters include: the deflection angle θ, the vehicle center coordinate P3, and the marker length l. Then the preset correction parameter formula is

[0066]

[0067] ; The target preliminary correction image is corrected according to the correction parameters, and the target preliminary correction image is rotated, scaled, and cropped to complete the correction, so as to obtain a more accurate top view; the corrected target preliminary correction image and the first corrected image are stored to obtain the first target image, so as to use the first target image as the true reference image of the top view.

[0068] In some embodiments, after storing the corrected target preliminary correction image and the first corrected image to determine the first target image, it further includes: if the target preliminary correction image is not the preset image, update the area of interest; correct the remaining target preliminary images in the remaining preliminary correction images according to the updated area of interest.

[0069] In an embodiment, the preset image is the target preliminary correction image corresponding to the last image in the first image set. It is determined whether the currently processed target preliminary correction image is the last one. If not, the next top view image in the first image set is continued to be corrected. Specifically, according to the feature point coordinates detected from the current target preliminary correction image, the region of interest is updated to narrow the range for the feature point detection of the next top view image. The feature points of the next top view image are detected, and the correction parameters are determined according to the detected feature points and the preset correction parameters. The target preliminary correction image is corrected according to the correction parameters, and the corrected target preliminary correction image is stored. Then, it is continuously determined whether the target preliminary correction image is the last one until the target preliminary correction image corresponding to the last image in the first image set is corrected, so as to complete the correction of all top view images in the first image set in the image processing stage and realize the normalization of the angle, position, and scale of the vehicle in the top view image.

[0070] For example, as Figure 5 shown, it is a flowchart of image processing for the images collected by a drone according to an embodiment of the present invention. Taking the image processing of the images collected by a drone as an example, the process of image processing for the images collected by a drone in the embodiment of the present invention at least includes step S50-step S62.

[0071] Step S50, obtain the first top view image.

[0072] Step S51, select feature points.

[0073] Step S52, calculate the initial correction parameters. Execute step S53 and step S55.

[0074] Step S53, correct the first top view image.

[0075] Step S54, save the correction result.

[0076] Step S55, perform pre-correction.

[0077] Step S56, select the region of interest.

[0078] Step S57, detect feature points.

[0079] Step S58, calculate the current correction parameters.

[0080] Step S59, correct the current top view image. Execute step S54 and step S60.

[0081] Step S60, determine whether it is the last top view image. If so, execute step S61; otherwise, execute step S62.

[0082] Step S61, complete the correction.

[0083] Step S62, update the region of interest. Execute step S57.

[0084] In some embodiments, obtaining the initial calibration parameters includes: obtaining vehicle feature points; obtaining the initial calibration parameters according to the vehicle feature points.

[0085] In an embodiment, as Figure 2 shown, it is a schematic diagram of vehicle calibration parameters according to an embodiment of the present invention. During the calibration of the UAV image, vehicle feature points are obtained. For example, two feature points P1 and P2 are manually selected on the test vehicle. The feature points should be able to represent the angle and position of the vehicle. The UAV tracking marker can be selected as the feature point to prepare for obtaining the initial calibration parameters; the initial calibration parameters are obtained according to the vehicle feature points. For example, the initial calibration parameters are calculated. According to the corresponding coordinates ( , ) of the selected reference point P1 and the corresponding coordinates ( , ) of P2, the initial calibration parameters can be calculated to prepare for correcting the image according to the initial calibration parameters.

[0086] In some embodiments, when obtaining the initial calibration parameters according to the vehicle feature points, the vehicle feature points are substituted into the following formula: …(1) …(2) …(3) Where is the vehicle deflection angle, and are the vehicle feature points, is the vehicle center coordinate, is the marker length, and are the distances of the vehicle feature points in the X direction, and are the distances of the vehicle feature points in the Y direction.

[0087] In an embodiment, as Figure 2 shown, according to the corresponding coordinates ( , ) of the selected reference point P1 and the corresponding coordinates ( , ) of P2, assuming the deflection angle θ, the vehicle center coordinate P3, and the marker length l, the initial calibration parameters can be calculated to realize calculating the initial correction parameters according to the vehicle feature points and correcting the image. The calculation formula is as follows:

[0088]

[0089]

[0090] In some embodiments, determining a BEV dataset based on a first target image and a captured second image includes: obtaining the second image when the vehicle is at a preset form speed; using the first target image and the second image as the BEV dataset.

[0091] In an embodiment, the second image is an image captured by a camera on the vehicle; the preset form speed is to make the vehicle travel in different speed scenarios according to requirements; the second image is obtained when the vehicle is at the preset form speed. For example, in the data collection stage, the test vehicle travels at a relatively slow speed within the parking lot, with a traveling speed of less than 20 kilometers per hour, and a fisheye image is captured by a fisheye camera, including images of 4 perspectives of the front, rear, left, and right of the vehicle, that is, the second image; the second image and the first target image are used as the BEV dataset, making the BEV dataset more real and comprehensive, providing a guarantee for the training of the GAN (Generative Adversarial Networks), and making the BEV image generated by the network after training closer to the real top view.

[0092] In some embodiments, when determining an adversarial network based on the BEV dataset, the following formula is introduced:

[0093] where, is the total loss function of the adversarial network, and are empirical input values, is the adversarial loss function, is the perceptual loss function, is the feature matching loss function, is the generator, is the discriminator.

[0094] In an embodiment, the adversarial network mainly includes two parts: a generator and a discriminator. A generative adversarial network is constructed and trained using the constructed BEV dataset to enable the generator and the discriminator to continuously optimize their respective capabilities through mutual confrontation, and finally generate realistic data. The generator adopts a multi-input, single-output structure, and inputs the second image after distortion removal. For example, 4 undistorted fisheye images are input, and a BEV image is output. The discriminator adopts a multi-scale discriminator, which consists of 3 discriminators with the same structure, and each discriminator is trained for different scales; the input of the discriminator is a 15-channel feature map, which is composed of connecting 4 undistorted fisheye images with the first target image or the generated image.

[0095] The generator and discriminator in the adversarial network each have a loss function, which are the keys to the training and optimization of the adversarial network. The total loss function consists of three parts: adversarial loss, feature matching loss, and perceptual loss. The total loss function is as follows:

[0096] Among them, and are empirical input values. For example, they are taken as 5 and 2 respectively. is the total loss function of the adversarial network, is the adversarial loss function, is the perceptual loss function, is the feature matching loss function, is the generator, is the discriminator. In some embodiments, when determining the adversarial loss, the following formula is brought in:

[0097] Among them, is the adversarial loss function, is the second image de-distorted image, is the first target image, is that the input first target image and the second image de-distorted image conform to the training data distribution, is the discriminator output result, is the generator output result.

[0098] In the embodiment, the definition of the adversarial loss is as follows:

[0099] The adversarial loss function is a loss function used to train the generator and discriminator in the generative adversarial network. The goal of the generator is to generate forged data as close as possible to the real data, while the goal of the discriminator is to accurately distinguish whether the input data is real or forged data generated by the generator. To achieve this goal, the GAN introduces an adversarial loss function to measure the difference between the forged data generated by the generator and the real data (for the generator) or the ability of the discriminator to distinguish between real data and forged data (for the discriminator).

[0100] In some embodiments, when determining the perceptual loss function, the following formula is brought in:

[0101] Among them, is the perceptual loss function, is the second image de-distorted image, is the first target image, The first target image of the input and the undistorted second image conform to the training data distribution. is a variable weight for measuring the importance of each layer i. For the first target image passing through layer The features extracted by the network. The output result of the generator passes through layer The features extracted by the network.

[0102] In the embodiment, the perceptual loss uses a VGG (Visual Geometry Group Network) pre-trained model as a reference. Feature maps of the generated image and the real image are extracted from the intermediate layer of the network, and the L1 distance is used to calculate the difference between them. To achieve that the generated image and the real image extract similar low-level and high-level feature differences from the loss network. The definition of the perceptual loss is as follows:

[0103] where x is the undistorted second image, and y is the first target image. is a variable weight for measuring the importance of each layer i. , n is the number of intermediate layers used, for example, n = 4 is taken.

[0104] In some embodiments, when determining the feature matching loss function, the following formula is brought in:

[0105] where is the feature matching loss function. is a variable weight for measuring the importance of each layer i. For the first target image at the scale passing through layer discriminator extracts features, and the is the features extracted by the generator output result at the scale passing through layer discriminator extracts features.

[0106] In the embodiment, the feature matching loss function directly extracts the feature map from the intermediate layer of the discriminator, and its definition is as follows:

[0107] In the feature matching loss, features are usually extracted from different scales of the true and false images and "matched" to measure the difference between the generated image and the real image in the feature space.

[0108] In some embodiments, determining the BEV image based on the adversarial network and the second image includes: performing undistortion processing on the second image to obtain a second target image; and determining the BEV image based on the second target image and the adversarial network.

[0109] In an embodiment, a camera on the vehicle, such as a fisheye camera, is used to collect an image to obtain a second image, such as a fisheye image. Undistortion processing is performed according to pre-calculated distortion parameters to obtain a second target image, so as to ensure that the generated BEV image is not affected by distortion. The four-view undistorted fisheye images (second images) are input into the generator of the adversarial network. At this time, the generator has completed training and can generate a BEV image according to the undistorted fisheye image. The generator projects the fisheye image onto the top-down view to generate a BEV image without edge distortion. Among them, the generator adopts an architecture with multiple inputs and one output, provides a separate input path for each input image, and performs downsampling operations on each path. The structure includes a convolutional layer, an STN (Spatial Transformer Networks), and a ResNet (Residual Network). The STN is used to ensure the spatial consistency between the input image and the BEV image, and the ResNet is used to restore the slight blur caused by the spatial transformation. Then, the features of the four input paths are connected to a single feature map, and convolution is performed through a modified ResNet block to reduce the number of channels of the output feature map. The feature map passes through a series of upsampling layers and finally outputs a BEV image without edge distortion, so as to generate a BEV image without edge distortion on the premise of using fisheye camera images, and the STN structure is introduced, enabling the network to autonomously learn the conversion process from the fisheye view to the top-down space, saving the cost of feature detection and calculation of the transformation model compared with traditional methods.

[0110] For example, as Figure 4 shown, it is a flowchart of generating a BEV image from a fisheye image according to an embodiment of the present invention. Taking a fisheye camera as an example, the flowchart of generating a BEV image from a fisheye image according to an embodiment of the present invention includes at least steps S30 - S35.

[0111] Step S30: Calibrate the camera to obtain the internal parameters, external parameters, and distortion parameters of the fisheye camera.

[0112] Step S31: Construct a BEV dataset, including real top-down views and fisheye images of four views.

[0113] Step S32: Construct an adversarial network and train it using the BEV dataset.

[0114] Step S33: The fisheye camera collects an image and performs undistortion processing according to the distortion parameters.

[0115] Step S34: Input the undistorted fisheye images from four perspectives into the generator.

[0116] Step S35: The generator projects the fisheye images onto the top-down perspective to generate a BEV image without edge distortion.

[0117] The following specifically describes the BEV image generation method according to an embodiment of the present invention with reference to Figure 6 the accompanying drawings.

[0118] As Figure 6 shown in the figure, it is a flowchart of the BEV image generation method according to another embodiment of the present invention. The BEV image generation method according to the embodiment of the present invention at least includes steps S10 - S25.

[0119] Step S10: Obtain vehicle feature points.

[0120] Step S11: Obtain initial calibration parameters based on the vehicle feature points.

[0121] Step S12: Obtain the first image in the first image set.

[0122] Step S13: Calibrate the first image and the remaining images in the first image set according to the initial calibration parameters to obtain the first calibrated image and the remaining preliminarily calibrated images.

[0123] Step S14: Obtain the region of interest in the target preliminarily calibrated image among the remaining preliminarily calibrated images.

[0124] Step S15: Detect feature points.

[0125] Step S16: Determine calibration parameters according to the detected feature points and the preset calibration parameter formula.

[0126] Step S17: Calibrate the target preliminarily calibrated image according to the calibration parameters.

[0127] Step S18: Store the calibrated target preliminarily calibrated image and the first calibrated image to determine the first target image.

[0128] Step S19: Determine whether the target preliminarily calibrated image is a preset image. If so, execute step S21; otherwise, execute step S20.

[0129] Step S20: Update the region of interest. Execute step S15.

[0130] Step S21: Obtain a second image when the vehicle is at a preset driving speed.

[0131] Step S22: Use the first target image and the second image as the BEV dataset.

[0132] Step S23: Determine the adversarial network according to the BEV dataset.

[0133] Step S24: Perform undistortion processing on the second image to obtain the second target image.

[0134] Step S25: Determine the BEV image according to the second target image and the adversarial network.

[0135] According to the BEV image generation method of the embodiment of the present invention, the top view in the first image set of the vehicle is corrected to obtain the first target image with the vehicle angle, position, and scale normalized. The first target image and the second image collected by the fisheye camera on the vehicle are used as the BEV dataset, which enriches the training sample features of the adversarial network. The adversarial network is trained using the BEV dataset. Based on the trained adversarial network, the BEV image can be generated using the second image, so that the generated BEV image is not affected by the edge stretching caused by fisheye image correction, there is no stretching deformation at the FOV junction, and it is closer to the real top view image, improving the applicability of the BEV image. Moreover, the rear top view image blocked by objects can be generated. And the adversarial network introduces the STN structure, enabling the network to autonomously learn the conversion process from the fisheye view to the top view space, saving the cost of feature detection and calculating the transformation model compared with the traditional method.

[0136] Next, refer to Figure 7 Describe the electronic device of the embodiment of the present invention.

[0137] As Figure 7 shown, it is a block diagram of an electronic device according to an embodiment of the present invention. The electronic device 100 according to the embodiment of the present invention includes: at least one processor 101; and a memory 102 communicatively connected to the at least one processor 101; wherein, the memory 102 stores a BEV image generation program executable by the at least one processor 101. When the BEV image generation program is executed by the at least one processor 101, the at least one processor 101 is caused to execute the BEV image generation method as described in the above embodiment.

[0138] An electronic device according to an embodiment of the present invention uses the BEV image generation method of the above embodiment. By performing a correction process on the top view in the first image set of the vehicle, a first target image with normalized vehicle angle, position, and scale is obtained. The first target image and the second image collected by the fisheye camera on the vehicle are used as the BEV data set, enriching the training sample features of the adversarial network. Using the BEV data set to train the adversarial network, based on the trained adversarial network, using the second image, a BEV image can be generated, making the generated BEV image not affected by the edge stretching caused by fisheye image correction, without stretching deformation at the FOV junction, being closer to the real top view image, improving the applicability of the BEV image, and being able to generate the rear top view image blocked by objects. Moreover, the adversarial network introduces an STN structure, enabling the network to autonomously learn the conversion process from the fisheye view to the top view space, saving the cost of feature detection and calculating the transformation model compared to traditional methods.

[0139] In some embodiments, the processor 101 determines the first target image according to the first image set of the vehicle, including: obtaining the first image in the first image set and the initial correction parameters; correcting the first image and the remaining images in the first image set according to the initial correction parameters to obtain the first corrected image and the remaining preliminarily corrected images; determining the first target image according to the remaining preliminarily corrected images and the first corrected image.

[0140] In some embodiments, the processor 101 determines the first target image according to the remaining preliminarily corrected images and the first corrected image, including: obtaining the detection feature points of the region of interest in the target preliminarily corrected image among the remaining preliminarily corrected images; determining the correction parameters according to the detection feature points and the preset correction parameter formula; correcting the target preliminarily corrected image according to the correction parameters; storing the corrected target preliminarily corrected image and the first corrected image to determine the first target image.

[0141] In some embodiments, after the processor 101 stores the corrected target preliminarily corrected image and the first corrected image to determine the first target image, it further includes: if the target preliminarily corrected image is not the preset image, updating the region of interest; correcting the remaining target preliminarily images in the remaining preliminarily corrected images according to the updated region of interest.

[0142] In some embodiments, the processor 101 obtains the initial correction parameters, including: obtaining the vehicle feature points; obtaining the initial correction parameters according to the vehicle feature points.

[0143] In some embodiments, when the processor 101 obtains the initial correction parameters according to the vehicle feature points, the vehicle feature points are substituted into the following formula: …(1) …(2) …(3) Among them, is the vehicle deflection angle, and are vehicle feature points, is the vehicle center coordinate, is the marked length, and are the distances of the vehicle feature points in the X upward direction, and are the distances of the vehicle feature points in the Y direction.

[0144] In some embodiments, the processor 101 determines the BEV dataset according to the first target image and the acquired second image, including: obtaining the second image when the vehicle is at a preset form speed; using the first target image and the second image as the BEV dataset.

[0145] In some embodiments, when the processor 101 determines the adversarial network according to the BEV dataset, the following formula is brought in:

[0146] Among them, is the total loss function of the adversarial network, and are empirical input values, is the adversarial loss function, is the perceptual loss function, is the feature matching loss function, is the generator, is the discriminator.

[0147] In some embodiments, when the processor 101 determines the adversarial loss, the following formula is brought in:

[0148] Among them, is the adversarial loss function, is the de-distorted image of the second image, is the first target image, is that the de-distorted images of the input first target image and the second image conform to the training data distribution, is the discriminator output result, is the generator output result.

[0149] In some embodiments, when the processor 101 determines the perceptual loss, the following formula is brought in:

[0150] Among them, is the perceptual loss function, is the undistorted image of the second image, is the first target image, The undistorted images of the input first target image and the second image conform to the training data distribution, is a variable weight used to measure the importance of each layer i, After the first target image passes through layer The features extracted by the network, The output result of the generator passes through layer The features extracted by the network.

[0151] In some embodiments, when the processor 101 determines the feature matching loss function, the following formula is brought in:

[0152] Wherein, is the feature matching loss function, is a variable weight used to measure the importance of each layer i, is the feature extracted by the discriminator for the first target image at the scale after passing through layer discriminator, is the feature extracted by the discriminator for the output result of the generator at the scale after passing through layer discriminator.

[0153] In some embodiments, the processor 101 determines the BEV image according to the adversarial network and the second image, including: performing undistortion processing on the second image to obtain a second target image; determining the BEV image according to the second target image and the adversarial network.

[0154] According to the electronic device of the embodiment of the present invention, using the BEV image generation method of the above embodiment, by performing correction processing on the top view in the first image set of the vehicle, the first target image with the vehicle angle, position and scale normalized is obtained, and the first target image and the second image collected by the fisheye camera on the vehicle are used as the BEV data set, enriching the training sample features of the adversarial network. Using the BEV data set to train the adversarial network, based on the trained adversarial network, using the second image, the BEV image can be generated, so that the generated BEV image is not affected by the edge stretching caused by the fisheye image correction, there is no stretching deformation at the FOV junction, it is closer to the real top view image, improving the applicability of the BEV image, and the rear top view image blocked by objects can be generated. Moreover, the adversarial network introduces the STN structure, enabling the network to autonomously learn the conversion process from the fisheye view to the top view space, saving the cost of feature detection and calculating the transformation model compared with the traditional method.

[0155] The computer-readable storage medium of the embodiments of the present invention is described below.

[0156] A BEV image generation program is stored on the computer-readable storage medium. When the BEV image generation program is executed by a processor, the device installed with the BEV image generation program implements the BEV image generation method as described in the above embodiments.

[0157] The computer program product of the embodiments of the present invention is described below.

[0158] The computer program product includes a computer program. When the computer program is executed by a processor, the BEV image generation method as described in the above embodiments is implemented.

[0159] According to the computer program product of the embodiments of the present invention, a first target image with normalized vehicle angle, position, and scale is obtained by rectifying the top view in the first image set of the vehicle. The first target image and the second image collected by the fisheye camera on the vehicle are used as the BEV data set, enriching the training sample features of the adversarial network. The adversarial network is trained using the BEV data set. Based on the trained adversarial network, the second image can be used to generate a BEV image, such that the generated BEV image is not affected by the edge stretching caused by fisheye image correction, there is no stretching deformation at the FOV junction, it is closer to the real top view image, improving the applicability of the BEV image, and a rear top view image blocked by an object can be generated. Moreover, the adversarial network introduces an STN structure, enabling the network to autonomously learn the conversion process from the fisheye view to the top view space, saving the cost of feature detection and calculating the transformation model compared to traditional methods.

[0160] The following refers to Figure 8 the vehicle of the embodiments of the present invention is described.

[0161] As Figure 8 shown, it is a block diagram of a vehicle according to an embodiment of the present invention. The vehicle 99 includes: the electronic device 100 as described in the above embodiments.

[0162] A vehicle 99 according to an embodiment of the present invention obtains a first target image with normalized vehicle angle, position, and scale by rectifying the top view in the first image set of the vehicle. The first target image and the second image collected by the fisheye camera on the vehicle are used as the BEV data set, enriching the training sample features of the adversarial network. Using the BEV data set to train the adversarial network, based on the trained adversarial network, using the second image, a BEV image can be generated, such that the generated BEV image is not affected by the edge stretching caused by fisheye image correction, there is no stretching deformation at the FOV junction, is closer to the real top view image, improves the applicability of the BEV image, and can generate the rear top view image blocked by objects. Moreover, the adversarial network introduces the STN structure, enabling the network to autonomously learn the conversion process from the fisheye view to the top view space, saving the cost of feature detection and calculating the transformation model compared to traditional methods.

[0163] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example.

[0164] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A BEV image generation method, characterized in that: include: Determining a first target image based on a first set of images of the vehicle; Determine a BEV data set according to the first target image and the acquired second image; Determine an adversarial network based on the BEV data set; A BEV image is determined based on the adversarial network and the second image.

2. The BEV image generation method according to claim 1, characterized in that: Determining a first target image according to a first image set of the vehicle includes: Acquire the first image and initial correction parameters in the first image set; Correcting the first image and the remaining images in the first image set according to the initial correction parameters to obtain a first corrected image and the remaining preliminary corrected images; The first target image is determined according to the remaining preliminary corrected images and the first corrected image.

3. The BEV image generation method according to claim 2, characterized in that: Determining the first target image according to the remaining preliminary corrected images and the first corrected image includes: Acquire detection feature points of an area of ​​interest in a target preliminary corrected image in the remaining preliminary corrected images; Determining correction parameters according to the detected characteristic points and a preset correction parameter formula; Correcting the target preliminary correction image according to the correction parameters; The corrected target preliminary corrected image and the first corrected image are stored to determine the first target image.

4. The BEV image generation method according to claim 3, characterized in that: After storing the corrected target preliminary corrected image and the first corrected image to determine the first target image, the method further includes: If the target preliminary corrected image is not a preset image, updating the region of interest; The remaining target preliminary images in the remaining preliminary corrected images are corrected according to the updated region of interest.

5. The BEV image generation method according to claim 2, characterized in that: Acquiring the initial calibration parameters includes: Obtain vehicle feature points; The initial correction parameters are obtained according to the vehicle feature points.

6. The BEV image generation method according to claim 5, characterized in that: When obtaining the initial correction parameters according to the vehicle feature points, the vehicle feature points are substituted into the following formula: …(1) …(2) …(3) Among them, the is the vehicle deflection angle, and is the vehicle feature point, is the center coordinate of the vehicle, is the mark length, and is the distance of the vehicle feature point in the X direction, and is the distance of the vehicle feature point in the Y direction.

7. The BEV image generation method according to claim 1, characterized in that: Determining a BEV data set according to the first target image and the acquired second image includes: Acquiring the second image when the vehicle is at a preset speed; The first target image and the second image are used as the BEV dataset.

8. The BEV image generation method according to claim 1, characterized in that: When determining the adversarial network based on the BEV dataset, the following formula is used: Among them, the is the total loss function of the adversarial network, and is the empirical input value, To combat the loss function, is the perceptual loss function, is the feature matching loss function, For the generator, the For the discriminator.

9. The BEV image generation method according to claim 8, characterized in that: When determining the adversarial loss function, the following formula is used: Among them, the To combat the loss function, is the dedistorted image of the second image, is the first target image, the The dedistorted images of the input first target image and the second image conform to the distribution of the training data. is the discriminator output result, Output the result for the generator.

10. The BEV image generation method according to claim 8, characterized in that: When determining the perceptual loss function, the following formula is used: Among them, the is the perceptual loss function, is the dedistorted image of the second image, is the first target image, the The dedistorted images of the input first target image and the second image conform to the distribution of the training data, is a variable weight used to measure the importance of each layer i. For the first target image layer The features extracted by the network are The output of the generator is layer Features extracted by the network.

11. The BEV image generation method according to claim 8, characterized in that: When determining the feature matching loss function, the following formula is used: Among them, the is the feature matching loss function, is a variable weight used to measure the importance of each layer i. The first target image is Passing by scale The features extracted by the layer discriminator are Output the result for the generator in Passing by scale Features extracted by the layer discriminator.

12. The BEV image generation method according to claim 1, characterized in that: Determining a BEV image according to the adversarial network and the second image includes: Performing a dedistortion process on the second image to obtain a second target image; A BEV image is determined based on the second target image and the adversarial network.

13. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein, the memory stores a BEV image generation program that can be executed by the at least one processor, and when the BEV image generation program is executed by the at least one processor, the at least one processor executes the BEV image generation method as described in any one of claims 1-12.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a BEV image generation program, and when the BEV image generation program is executed by a processor, a device installed with the BEV image generation program implements the BEV image generation method according to any one of claims 1 to 12.

15. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the BEV image generation method according to any one of claims 1 to 12 is implemented.

16. A vehicle comprising: The electronic device as claimed in claim 13.