Vision-based upright column detection method and device, storage medium and electronic equipment

By using fisheye cameras and deep learning detection models during vehicle column identification, combined with vehicle motion information, the cost of equipment, blind spots and weather impact of column detection in underground parking lots is solved, and high-precision and stable column identification and distance measurement are achieved.

CN120472430APending Publication Date: 2025-08-12ENBOTAI TIANJIN TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510707553.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art has problems such as expensive equipment, blind spots for detection, weather-affected measurement effects, poor target recognition effect, easy to miss detection and insufficient distance measurement accuracy when detecting columns in underground parking lots, especially at close distance, low recognition rate and poor distance measurement stability.

Method used

During the process of identifying the column, the fisheye camera is used to collect images and input a pre-trained deep learning detection model, calculate the position and size information of the column, combine the vehicle motion information and the parameters of the camera calibration, missed detection and secondary perception are performed, and the position and size of the column are calculated.

Benefits of technology

It realizes accurate identification and ranging during the vehicle identification of columns, avoids missed inspection, improves the accuracy and stability of ranging, and ensures the accuracy and reliability of column inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472430A_ABST
    Figure CN120472430A_ABST
Patent Text Reader

Abstract

The invention provides a vision-based upright post detection method and device, a storage medium and electronic equipment, and the method comprises the steps: at any moment in the process of identifying an upright post by a vehicle, if the upright post cannot be detected in each camera; if yes, second position information at the moment is calculated according to the motion information of the vehicle at the moment and first position information recorded at the previous moment; after the vehicle starts to be parked, for a first stand column involved in a parking path, the first stand column is cut out from an image collected by a camera and is sent into a pre-trained perception model, so that distance measurement key point information, output by the perception model, of the first stand column is obtained; and calculating third position information and size information of the first stand column according to the distance measurement key point information. According to the invention, the missed upright post can be accurately identified and the position information of the upright post at the missed time can be calculated, and the upright post can be dug out in the parking process and sent into the sensing model for secondary sensing and calculation, so that more accurate position information and size information can be obtained, and missed detection can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine vision, and in particular to a vision-based column detection method, device, storage medium and electronic equipment. Background Art

[0002] Underground parking column pillars are the most common large obstacles in automated parking scenarios. They are large, frequently appear, and are extremely close to the target parking location. Therefore, they are highly susceptible to collisions or scratches during parking, posing a high risk and making them one of the most critical obstacles to identify in automated parking scenarios.

[0003] Currently, the primary method for detecting pillars in underground parking lots is through lidar or ultrasonic radar. However, these two solutions suffer from high equipment costs, blind spots due to installation location constraints, weather susceptibility, and an inability to pinpoint the specific type and size of obstacles. Fisheye camera-based visual solutions are increasingly popular with automakers due to their low cost, minimal weather impact, and ability to pinpoint target type and size. However, current solutions suffer from poor target recognition, a high risk of missed detections, insufficient ranging accuracy, and poor stability. Summary of the Invention

[0004] In view of the above problems, the present invention provides a vision-based column detection method, device, storage medium and electronic device that overcome the above problems or at least partially solve the above problems.

[0005] In a first aspect, a vision-based column detection method includes:

[0006] At any moment during the vehicle recognition process, if none of the cameras can detect the pillar, second position information at that moment is calculated based on the vehicle's motion information at that moment and the first position information recorded at the previous moment, where the first position information is the position information of the pillar in the world coordinate system, and the second position information is the position information of the pillar in the vehicle coordinate system;

[0007] After the vehicle starts parking, for a first pillar involved in the parking path, the first pillar is extracted from an image captured by the camera and fed into a pre-trained perception model to obtain ranging key point information of the first pillar output by the perception model;

[0008] The third position information and size information of the first pillar are calculated based on the ranging key point information, wherein the third position information is the position information of the first pillar in the vehicle coordinate system.

[0009] Optionally, in certain optional embodiments, at any time during the process of the vehicle identifying the pillar, if none of the cameras can detect the pillar, before calculating the second position information at that time based on the motion information of the vehicle at that time and the first position information recorded at a previous time, the method further includes:

[0010] During the process of identifying the pillar by the vehicle, fourth position information of the identified pillar is obtained, wherein the fourth position information is position information of the identified pillar in the vehicle coordinate system;

[0011] Converting the fourth position information into fifth position information according to the motion information of the vehicle itself, wherein the fifth position information is position information in a world coordinate system;

[0012] Record the fifth position information corresponding to each moment.

[0013] Optionally, in certain optional embodiments, obtaining fourth position information of the identified pillar during the process of the vehicle identifying the pillar includes:

[0014] During the process of vehicle recognition of the pillar, the image of the pillar is acquired through camera capture;

[0015] Inputting the column image into a pre-trained deep learning detection model to obtain information of a plurality of key points output by the deep learning detection model, wherein the key points are points on the bottom edge of the identified column;

[0016] According to the pre-calibrated parameters of the camera and the information of each key point, the fourth position information of the corresponding column is calculated.

[0017] Optionally, in certain optional embodiments, during the process of the vehicle identifying the pillar, acquiring the pillar image by a camera includes:

[0018] During the process of vehicle recognition of the pillar, multiple scene images are acquired through multiple cameras, wherein one camera corresponds to one scene image at the same time;

[0019] For any scene image, format conversion and image preprocessing are performed on the scene image to obtain a corresponding pillar image.

[0020] Optionally, in certain optional embodiments, inputting the pillar image into a pre-trained deep learning detection model to obtain information of multiple key points output by the deep learning detection model includes:

[0021] Inputting each of the pillar images into a pre-trained deep learning detection model to obtain information of multiple key points output by the deep learning detection model, wherein each pillar image outputs corresponding information of multiple key points;

[0022] The fourth position information of the corresponding column is calculated based on the pre-calibrated parameters of the camera and the information of each key point, including:

[0023] For information of multiple key points of any pillar image, fourth position information of the pillar corresponding to the pillar image is calculated based on pre-calibrated parameters of a camera used to capture the pillar image and information of each key point;

[0024] pairing each of the fourth position information with each other;

[0025] For any pair of fourth position information, calculating a matching factor between the corresponding two pieces of fourth position information;

[0026] For any pair of fourth position information, if the corresponding matching factor is greater than a preset factor threshold, the corresponding pillars are determined to be different pillars, and the corresponding two fourth position information are output respectively;

[0027] For any pair of fourth position information, if the corresponding matching factor is less than the preset factor threshold, the corresponding pillars are determined to be the same pillar, and the corresponding two fourth position information are combined into the same fourth position information for output.

[0028] Optionally, in certain optional embodiments, at any moment during the process of the vehicle identifying the pillar, if none of the cameras can detect the pillar, calculating the second position information at that moment based on the motion information of the vehicle at that moment and the first position information recorded at the previous moment includes:

[0029] At any time during the vehicle's pillar recognition process, if none of the cameras can detect the pillar, it is determined that the pillar is missed at that moment;

[0030] When determining that a pillar has been missed, obtain the movement information of the vehicle at that moment;

[0031] Obtain the first position information of the previous moment from the record;

[0032] The second position information at the moment is calculated based on the motion information of the vehicle at the moment and the first position information.

[0033] Optionally, in certain optional implementations, the calculating and obtaining the third position information and size information of the first column according to the ranging key point information includes:

[0034] Calculating fifth position information and size information of the first column based on the ranging key point information, wherein the fifth position information is position information of the first column in a world coordinate system;

[0035] If the first column is outside the preset area, performing weighted calculation based on the fifth position information of the first column at the current moment and the fifth position information of the first column at the historical moment to obtain corresponding weighted position information;

[0036] After smoothing the weighted position information using a filtering algorithm, third position information of the first column is calculated;

[0037] If the first column is within the preset area, the third position information of the first column is calculated based on the fifth position information of the first column at the current moment.

[0038] In a second aspect, a vision-based column detection device includes: a missed detection calculation unit, a secondary perception unit, and an accurate calculation unit;

[0039] The missed detection calculation unit is configured to, at any moment during the process of the vehicle identifying the pillar, if none of the cameras can detect the pillar, calculate second position information at that moment based on the vehicle's motion information at that moment and the first position information recorded at a previous moment, wherein the first position information is the position information of the pillar in the world coordinate system, and the second position information is the position information of the pillar in the vehicle coordinate system;

[0040] The secondary perception unit is configured to, after the vehicle begins parking, extract a first pillar involved in a parking path from an image captured by the camera and feed the image into a pre-trained perception model to obtain ranging key point information of the first pillar output by the perception model;

[0041] The accurate calculation unit is used to calculate the third position information and size information of the first pillar based on the ranging key point information, wherein the third position information is the position information of the first pillar in the vehicle coordinate system.

[0042] In a third aspect, a computer-readable storage medium stores a program, which, when executed by a processor, implements any of the above-mentioned vision-based column detection methods.

[0043] In a fourth aspect, an electronic device comprises at least one processor, and at least one memory and a bus connected to the processor; wherein the processor and the memory communicate with each other through the bus; and the processor is used to call program instructions in the memory to execute any of the above-mentioned vision-based column detection methods.

[0044] By means of the above technical solution, the present invention provides a vision-based pillar detection method, device, storage medium, and electronic device. At any time during the process of vehicle identification of a pillar, if none of the cameras can detect the pillar, second position information at that moment is calculated based on the vehicle's motion information at that moment and the first position information recorded at the previous moment, wherein the first position information is the position information of the pillar in the world coordinate system, and the second position information is the position information of the pillar in the vehicle coordinate system. After the vehicle starts parking, for the first pillar involved in the parking path, the first pillar is cut out from the image captured by the camera and fed into a pre-trained perception model to obtain ranging key point information of the first pillar output by the perception model. Third position information and size information of the first pillar are calculated based on the ranging key point information, wherein the third position information is the position information of the first pillar in the vehicle coordinate system. It can be seen from this that the present invention can accurately identify missed pillars and calculate the position information of the pillars at the time of missed detection. Moreover, during the parking process, the pillars can be picked out and sent to the perception model for secondary perception and calculation, thereby obtaining more accurate position information and size information. The recognition effect is better, missed detection is avoided, and the ranging accuracy and stability are good.

[0045] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0047] Figure 1 The flowchart of the first vision-based column detection method provided by the present invention is shown;

[0048] Figure 2 A flow chart of a second vision-based column detection method provided by the present invention is shown;

[0049] Figure 3 A flow chart of a third vision-based column detection method provided by the present invention is shown;

[0050] Figure 4 A flowchart of a fourth vision-based column detection method provided by the present invention is shown;

[0051] Figure 5 A schematic structural diagram of a vision-based column detection device provided by the present invention is shown;

[0052] Figure 6 A schematic structural diagram of an electronic device provided by the present invention is shown. DETAILED DESCRIPTION

[0053] Underground parking column pillars are the most common large obstacles in automated parking scenarios. They are large, frequently appear, and are extremely close to the target parking location. Therefore, they are highly susceptible to collisions or scratches during parking, posing a high risk and making them one of the most critical obstacles to identify in automated parking scenarios.

[0054] Currently, the main method for detecting pillars in underground parking lots is through lidar or ultrasonic radar. However, these two solutions have shortcomings such as expensive equipment, blind spots due to installation location restrictions, measurement results are easily affected by weather conditions, and the inability to determine the specific type and size of obstacles.

[0055] The inventors of this solution have discovered that current fisheye camera-based vision solutions are increasingly popular with automakers due to their low price, minimal weather impact, and ability to provide information such as target type and size. However, these solutions have the following issues:

[0056] (1) The recognition rate is low when the vehicle is close to the pillar, especially when the image of the pillar fills the entire camera, which makes it very easy to miss the detection.

[0057] (2) Insufficient column distance measurement accuracy. When the distance to the column is far, other obstacles such as vehicles blocking the bottom of the column, unclear contact surface between the bottom of the column and the ground, the column is too close to the camera so that the bottom of the column cannot be seen, etc., may cause inaccurate regression of the column distance key points, thus affecting the column distance measurement accuracy.

[0058] (3) Poor ranging stability. When the target appears in multiple camera viewpoints and switches back and forth between multiple cameras, the ranging jump will be large and the ranging stability cannot be guaranteed.

[0059] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0060] like Figure 1 As shown, the present invention provides a vision-based column detection method, including: S100, S200 and S300;

[0061] S100. At any time during the process of the vehicle identifying a pillar, if none of the cameras can detect the pillar, calculating second position information at that time based on the vehicle's motion information at that time and first position information recorded at a previous time, wherein the first position information is position information of the pillar in a world coordinate system, and the second position information is position information of the pillar in a vehicle coordinate system;

[0062] Alternatively, the vehicle identifying the pillars mentioned in the present invention may occur while the vehicle is driving in a parking lot or while the vehicle is parking in a corresponding parking space for the first time. The present invention does not limit this. The vehicle may continuously detect and collect pillar data using a fisheye camera while driving in a parking lot or parking, and calculate and record corresponding position information.

[0063] Optionally, a fisheye camera refers to a camera with a fisheye lens, which is a lens with an extremely short focal length and a viewing angle close to or equal to 180°, such as a lens with a focal length of 16 mm or shorter. It should be noted that the present invention is not limited to fisheye cameras, and any feasible camera falls within the scope of protection of the present invention.

[0064] Optionally, the present invention involves two coordinate systems, one is the vehicle coordinate system of the vehicle itself, that is, the position of each object outside the vehicle is identified based on the coordinate system calibrated on the vehicle; the other is the world coordinate system, which is the absolute coordinate system of the system. Before the user coordinate system is established, the coordinates of all points on the screen are determined by the origin of the coordinate system.

[0065] Optionally, during the movement of the vehicle, the present invention can collect the movement information of the vehicle in real time based on the sensors set on the vehicle, including the displacement information of the vehicle at four angles of front, back, left and right, and the up and down angles of the vehicle in the up and down directions. The present invention does not impose any restrictions on this.

[0066] Optionally, the present invention can record the movement information of the vehicle at each moment and collect and calculate the position information of the pillar in the world coordinate system in real time, so that the corresponding information can be directly read from the record when needed for identifying the pillar. The present invention does not impose any restrictions on this.

[0067] For example, Figure 2 As shown, in some optional embodiments, before S100, the method further includes: S70, S80 and S90;

[0068] S70. During the process of identifying a pillar by the vehicle, obtaining fourth position information of the identified pillar, wherein the fourth position information is position information of the identified pillar in the vehicle coordinate system;

[0069] Optionally, as previously mentioned, the camera on the vehicle of the present invention can continuously collect and identify the position information of the pillar. Since the camera is mounted on the vehicle, to ensure accuracy, the coordinate system used by the camera is the vehicle coordinate system. That is, the coordinate system is calibrated with the vehicle as the center. Therefore, the position information collected and identified by the camera is position information in the vehicle coordinate system, which is not limited by the present invention.

[0070] For example, in some optional embodiments, the S70 includes: step 1.1, step 1.2 and step 1.3;

[0071] Step 1.1: During the process of vehicle recognition of the pillar, an image of the pillar is acquired by a camera;

[0072] Optionally, as previously mentioned, the camera described in the present invention may be a fisheye camera. Using a fisheye camera, the present invention can continuously capture images of pillars while the vehicle is in motion. For example, in an underground parking lot, to avoid collisions with pillars, the present invention can continuously capture images of the pillars and determine distances, although this is not a limitation of the present invention.

[0073] For example, in some optional embodiments, the step 1.1 includes: step 2.1 and step 2.2;

[0074] Step 2.1, during the process of the vehicle identifying the pillar, multiple scene images are acquired by using multiple cameras, wherein one camera corresponds to one scene image at the same time;

[0075] Step 2.2: For any scene image, perform format conversion and image preprocessing on the scene image to obtain a corresponding pillar image.

[0076] Optionally, the present invention can install one fisheye camera at each of four different locations on the vehicle, namely the front, rear, left, and right, for a total of four fisheye cameras. The four fisheye cameras can be used to obtain image data from four different perspectives. Of course, the present invention is not limited to four fisheye cameras, and more fisheye cameras can be installed as needed to ensure that there are no blind spots. The present invention does not impose any restrictions on this.

[0077] Optionally, any camera may continuously capture pillar images at regular intervals during the process of vehicle pillar recognition. That is, any camera may capture a sequence of pillar images during the process of vehicle pillar recognition, and the pillar images may be arranged in chronological order, although this is not a limitation of the present invention.

[0078] Optionally, there may be some problems with the photos taken directly by the fisheye camera. Therefore, in order to improve the accuracy of the present invention, the present invention can process the images obtained by the fisheye camera to complete image format conversion and image preprocessing (image dedistortion, image cropping, scaling and shuffle, etc.) operations, and the present invention does not impose any restrictions on this.

[0079] Step 1.2: Input the column image into a pre-trained deep learning detection model to obtain information on a plurality of key points output by the deep learning detection model, wherein the key points are points on the bottom edge of the identified column;

[0080] For example, in some optional embodiments, the step 1.2 includes: step 3.1;

[0081] Step 3.1: Input each of the pillar images into a pre-trained deep learning detection model to obtain information on a plurality of key points output by the deep learning detection model, wherein each pillar image outputs corresponding information on a plurality of key points;

[0082] Optionally, the present invention can pre-train a corresponding deep learning detection model based on actual needs, and then detect images based on the deep learning detection model to obtain key point information. It should be noted that the present invention can calibrate training samples (training set and test set) to calibrate key point information in the samples, and then use the samples to train and verify the deep learning detection model until the accuracy of the deep learning detection model meets actual requirements. This is not a limitation of the present invention.

[0083] Optionally, the present invention does not impose any specific restrictions on the information and number of key points extracted from each pillar image; any feasible method falls within the scope of protection of the present invention. For example, the present invention may use a deep learning detection model to perform inference on each pillar image, regress four points on the bottom edge of the pillar as key points, and obtain corresponding information, and the present invention does not impose any restrictions on this.

[0084] Step 1.3: Calculate the fourth position information of the corresponding column based on the pre-calibrated parameters of the camera and the information of each key point.

[0085] For example, in some optional embodiments, the step 1.3 includes: step 4.1, step 4.2, step 4.3, step 4.4 and step 4.5;

[0086] Step 4.1: For information on multiple key points of any pillar image, calculate fourth position information of the pillar corresponding to the pillar image based on pre-calibrated parameters of a camera used to capture the pillar image and information on each key point;

[0087] Optionally, given the different installation locations of each camera, performance parameters may also vary. Therefore, the present invention can calibrate corresponding parameters (including internal camera parameters and external installation parameters) for each camera based on actual needs. These parameters are used to determine the position of the pillar image captured by the camera, and the present invention is not limited to this.

[0088] Optionally, the fourth position information mentioned in the present invention can be understood as the position information of the key point in the vehicle coordinate system, and the present invention is not limited to this.

[0089] Step 4.2, pairing the fourth position information in pairs;

[0090] Step 4.3: For any pair of fourth position information, calculate the matching factor between the corresponding two pieces of fourth position information;

[0091] Step 4.4: For any pair of fourth position information, if the corresponding matching factor is greater than a preset factor threshold, the corresponding pillars are determined to be different pillars, and the corresponding two fourth position information are output respectively;

[0092] Step 4.5: For any pair of fourth position information, if the corresponding matching factor is less than the preset factor threshold, the corresponding pillars are determined to be the same pillar, and the corresponding two fourth position information are combined into the same fourth position information for output.

[0093] Optionally, considering that in the four fisheye camera views of the surround view, the same pillar may appear in a single fisheye camera view or in the intersection of two fisheye camera views as the vehicle moves. Therefore, in order to solve the problem of the same pillar appearing in the intersection of two fisheye camera views, leading to misidentification as two different pillars. The present invention first calculates the matching factor between the fourth position information of each pair of pillars based on the fourth position information of the pillars in the images of each fisheye camera (for example, if the fourth position information of the pillar in the image of fisheye camera 1 is D1, and the fourth position information of the pillar in the image of fisheye camera 2 is D2, then the matching factor between D1 and D2 is calculated). If the value of the matching factor is less than a set threshold, the pillars are considered to be the same target (i.e., the same pillar). Based on the time sequence of the pillars' appearance in different fisheye cameras and their relative distances from the fisheye cameras, the fourth position information of the pillars in the images of each fisheye camera is fused, stored as the fourth position information of a single pillar, and output. If the value of the matching factor is greater than the set threshold, the pillars are considered to be different targets and are stored and output separately, but the present invention does not impose any limitation on this.

[0094] S80. Convert the fourth position information into fifth position information based on the motion information of the vehicle itself, wherein the fifth position information is position information in a world coordinate system;

[0095] S90: Record the fifth position information corresponding to each moment.

[0096] Optionally, as previously mentioned, the fourth position information is positional information in the vehicle coordinate system, while the pillars identified by the present invention are physical objects. Therefore, to improve the accuracy of the present invention, the present invention can combine the vehicle's own motion information to convert the fourth position information into positional information in the world coordinate system and record and store it. This prevents the failure to identify pillars due to missed detection by the model or vehicle movement.

[0097] Optionally, the present invention incorporates the vehicle's global pose information (the position and orientation of an object or system in a world coordinate system) to record the vehicle's lateral and longitudinal displacements. The fourth position information of the pillar and the global pose information are then used to calculate the pillar's coordinates in the world coordinate system, i.e., the fifth position information. This is not a limitation of the present invention.

[0098] Optional, such as Figure 3 As shown, in some optional embodiments, the S100 includes: S110, S120, S130 and S140;

[0099] S110: If, at any time during the vehicle identification process, no camera can detect the pillar, it is determined that the pillar is missed at that time.

[0100] S120: When it is determined that a pillar has been missed, obtaining movement information of the vehicle at that moment;

[0101] S130, obtaining first position information at the last moment from the record;

[0102] S140: Calculate the second position information at the moment based on the motion information of the vehicle at the moment and the first position information.

[0103] Optionally, during the vehicle's pillar recognition process, if a pillar was recognized by a camera at one moment but cannot be recognized by all cameras at the next moment, this indicates a possible missed detection. Therefore, to ensure data continuity and pillar detection accuracy, the present invention can determine a missed detection by calculating the pillar's position information at the next moment (the missed detection moment) (i.e., second position information) based on the pillar's position information monitored at the previous moment (i.e., first position information). The present invention does not impose any limitations on this.

[0104] That is, if at a certain moment, no pillar matching the one being recorded (continuously recorded) can be detected from each camera, the target is considered to have been missed at that moment. The global pose information at the moment of loss and the first position information of the target at the last moment before it was lost are read, and the position information corresponding to the target after the vehicle's movement is calculated (the calculated position information at this time is the position information in the world coordinate system). This position information is then converted to the position information in the vehicle coordinate system (i.e., the second position information) and output as the target position information at the moment of loss, thereby ensuring the continuity of the target output. The present invention is not limited to this.

[0105] S200: After the vehicle starts parking, for a first pillar involved in the parking path, extract the first pillar from an image captured by a camera and feed the image into a pre-trained perception model to obtain ranging key point information of the first pillar output by the perception model;

[0106] S300: Calculate third position information and size information of the first pillar based on the ranging key point information, wherein the third position information is position information of the first pillar in the vehicle coordinate system.

[0107] Optionally, after parking begins, for the target (the first pillar) on the planned parking path, the present invention can extract the pillar from the original image based on the geometric shape obtained by the previous detection (the rectangular frame used in computer vision to locate objects in the image when detecting the pillar) in combination with the size information of the pillar, and feed the model into a high-precision perception model for secondary perception, thereby obtaining more accurate ranging key point information. The ranging key point information can then be used to calculate a more accurate pillar position and regress a more accurate pillar size. The present invention is not limited to this.

[0108] For example, Figure 4 As shown, in some optional embodiments, the S300 includes: S310, S320, S330 and S340;

[0109] S310: Calculate fifth position information and size information of the first column based on the ranging key point information, wherein the fifth position information is position information of the first column in a world coordinate system;

[0110] S320: If the first pillar is outside the preset area, perform weighted calculation based on the fifth position information of the first pillar at the current moment and the fifth position information of the first pillar at previous moments to obtain corresponding weighted position information;

[0111] S330: Smoothing the weighted position information using a filtering algorithm to obtain third position information of the first column;

[0112] S340: If the first column is within a preset area, calculate the third position information of the first column according to the fifth position information of the first column at the current moment.

[0113] Optionally, due to changes in the fisheye camera's viewing angle, unstable factors in model detection, and errors in long-range camera calibration, the relative positions of the pillars detected at different times may fluctuate. Therefore, in order to obtain a relatively stable and accurate distance output, the present invention can smooth the output results.

[0114] For example, based on actual conditions, the present invention sets a high-quality area. The output of the columns outside this area is weighted using the fifth position information of the current frame and the fifth position information of the historical frame. The result is smoothed by the Kalman filtering algorithm, and then the result is converted into the final output third position information according to the global pose information. The present invention does not impose any restrictions on this.

[0115] Although the operations are depicted in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in a sequential order.Multitasking and parallel processing may be advantageous under certain circumstances.

[0116] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0117] like Figure 5 As shown, the present invention provides a vision-based pillar detection device, comprising: a missed detection calculation unit 100, a secondary perception unit 200 and an accurate calculation unit 300;

[0118] The missed detection calculation unit 100 is configured to, at any moment during the process of the vehicle identifying a pillar, if none of the cameras can detect the pillar, calculate second position information at that moment based on the vehicle's motion information at that moment and the first position information recorded at a previous moment, wherein the first position information is the position information of the pillar in the world coordinate system, and the second position information is the position information of the pillar in the vehicle coordinate system;

[0119] The secondary perception unit 200 is configured to, after the vehicle begins parking, extract a first pillar involved in the parking path from an image captured by the camera and feed the image into a pre-trained perception model to obtain ranging key point information of the first pillar output by the perception model;

[0120] The accurate calculation unit 300 is configured to calculate third position information and size information of the first pillar based on the ranging key point information, wherein the third position information is position information of the first pillar in the vehicle coordinate system.

[0121] Optionally, in certain optional embodiments, the apparatus further comprises: a fourth information obtaining subunit, a fifth information obtaining subunit, and a fifth information recording subunit;

[0122] The fourth information obtaining subunit is configured to, if at any moment during the process of the vehicle identifying the pillar, none of the cameras can detect the pillar, obtain fourth position information of the identified pillar during the process of the vehicle identifying the pillar, before calculating the second position information at that moment based on the motion information of the vehicle at that moment and the first position information recorded at a previous moment, wherein the fourth position information is position information of the identified pillar in the vehicle coordinate system;

[0123] The fifth information obtaining subunit is configured to convert the fourth position information into fifth position information according to the motion information of the vehicle itself, wherein the fifth position information is position information in a world coordinate system;

[0124] The fifth information recording subunit is used to record the fifth position information corresponding to each moment.

[0125] Optionally, in certain optional embodiments, the fourth information obtaining subunit includes: a pillar image obtaining subunit, a key point information obtaining subunit, and a fourth information calculating subunit;

[0126] The pillar image acquisition subunit is used to acquire a pillar image through camera capture during the process of vehicle recognition of the pillar;

[0127] The key point information obtaining subunit is used to input the column image into a pre-trained deep learning detection model to obtain information of a plurality of key points output by the deep learning detection model, wherein the key points are points on the bottom edge of the identified column;

[0128] The fourth information calculation subunit is used to calculate the fourth position information of the corresponding column based on the pre-calibrated parameters of the camera and the information of each key point.

[0129] Optionally, in certain optional embodiments, the pillar image acquisition subunit includes: a scene image acquisition subunit and an image processing subunit;

[0130] The scene image acquisition subunit is used to acquire multiple scene images through multiple cameras during the process of vehicle recognition of the pillar, wherein one camera corresponds to one scene image at the same time;

[0131] The image processing subunit is used to perform format conversion and image preprocessing on any scene image to obtain a corresponding pillar image.

[0132] Optionally, in certain optional implementations, the key point information obtaining subunit includes: a model detection subunit;

[0133] The model detection subunit is configured to input each of the pillar images into a pre-trained deep learning detection model to obtain information of a plurality of key points output by the deep learning detection model, wherein each pillar image outputs corresponding information of a plurality of key points;

[0134] The fourth information calculation subunit includes: an image information calculation subunit, an information pairing subunit, a matching factor calculation subunit, a first output subunit and a second output subunit;

[0135] The image information calculation subunit is configured to calculate, based on information of multiple key points of any pillar image, fourth position information of the pillar corresponding to the pillar image according to pre-calibrated parameters of a camera used to capture the pillar image and information of each key point;

[0136] The information pairing subunit is used to pair each of the fourth position information with each other;

[0137] The matching factor calculation subunit is configured to calculate, for any pair of fourth position information, a matching factor between two corresponding pieces of fourth position information;

[0138] The first output sub-unit is configured to, for any pair of fourth position information, determine that the corresponding pillars are different pillars if the corresponding matching factor is greater than a preset factor threshold, and output the corresponding two pieces of fourth position information respectively;

[0139] The second output subunit is configured to determine, for any pair of fourth position information, if the corresponding matching factor is less than a preset factor threshold, that the corresponding pillars are the same pillar, and combine the corresponding two fourth position information into the same fourth position information for output.

[0140] Optionally, in some optional embodiments, the missed detection calculation unit 100 includes: a missed detection determination subunit, a motion information acquisition subunit, a first information acquisition subunit, and a second information calculation subunit;

[0141] The missed detection determination subunit is configured to determine that a pillar is missed at any moment during the process of identifying a pillar by the vehicle if none of the cameras can detect the pillar at that moment;

[0142] The motion information obtaining subunit is used to obtain the motion information of the vehicle at that moment when it is determined that the pillar has been missed;

[0143] The first information obtaining subunit is used to obtain the first position information of the last moment from the record;

[0144] The second information calculation subunit is configured to calculate the second position information at the moment based on the motion information of the vehicle at the moment and the first position information.

[0145] Optionally, in some optional implementations, the accurate calculation unit 300 includes: a fifth information calculation subunit, a position information weighting subunit, a smoothing calculation subunit, and a direct calculation subunit;

[0146] The fifth information calculation subunit is configured to calculate fifth position information and size information of the first column based on the ranging key point information, wherein the fifth position information is position information of the first column in a world coordinate system;

[0147] The position information weighting subunit is configured to perform weighted calculation based on the fifth position information of the first column at the current moment and the fifth position information of the first column at the historical moment to obtain corresponding weighted position information if the first column is outside the preset area;

[0148] The smoothing calculation subunit is configured to smooth the weighted position information using a filtering algorithm to calculate and obtain the third position information of the first column;

[0149] The direct calculation subunit is configured to calculate the third position information of the first column based on the fifth position information of the first column at a current moment if the first column is within a preset area.

[0150] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0151] The vision-based column detection device includes a processor and a memory. The missed detection calculation unit 100, the secondary perception unit 200 and the accurate calculation unit 300 are all stored in the memory as program units, and the processor executes the program units stored in the memory to realize corresponding functions.

[0152] The processor contains a kernel, which retrieves the corresponding program unit from memory. One or more kernels can be configured. By adjusting kernel parameters, the kernel can accurately identify missed pillars and calculate their positions at the moment of missed detection. Furthermore, during parking, pillars can be extracted and fed into the perception model for secondary perception and calculation, yielding even more accurate position and size information. This improves recognition, avoids missed detections, and provides high ranging accuracy and stability.

[0153] An embodiment of the present invention provides a computer-readable storage medium having a program stored thereon, which implements the vision-based column detection method when executed by a processor.

[0154] An embodiment of the present invention provides a processor, which is used to run a program, wherein the vision-based column detection method is executed when the program is run.

[0155] like Figure 6 As shown, an embodiment of the present invention provides an electronic device 700, which includes at least one processor 701, at least one memory 702 connected to the processor 701, and a bus 703. The processor 701 and the memory 702 communicate with each other via the bus 703. The processor 701 is configured to invoke program instructions stored in the memory 702 to execute the aforementioned vision-based pillar detection method. The electronic device herein may be a server, a PC, a PAD, a mobile phone, or the like.

[0156] The present invention also provides a computer program product, which, when executed on an electronic device, is adapted to execute a program for initializing the following method steps:

[0157] A vision-based column detection method, comprising:

[0158] At any moment during the vehicle recognition process, if none of the cameras can detect the pillar, second position information at that moment is calculated based on the vehicle's motion information at that moment and the first position information recorded at the previous moment, where the first position information is the position information of the pillar in the world coordinate system, and the second position information is the position information of the pillar in the vehicle coordinate system;

[0159] After the vehicle starts parking, for a first pillar involved in the parking path, the first pillar is extracted from an image captured by the camera and fed into a pre-trained perception model to obtain ranging key point information of the first pillar output by the perception model;

[0160] The third position information and size information of the first pillar are calculated based on the ranging key point information, wherein the third position information is the position information of the first pillar in the vehicle coordinate system.

[0161] Optionally, in certain optional embodiments, at any time during the process of the vehicle identifying the pillar, if none of the cameras can detect the pillar, before calculating the second position information at that time based on the motion information of the vehicle at that time and the first position information recorded at a previous time, the method further includes:

[0162] During the process of identifying the pillar by the vehicle, fourth position information of the identified pillar is obtained, wherein the fourth position information is position information of the identified pillar in the vehicle coordinate system;

[0163] Converting the fourth position information into fifth position information according to the motion information of the vehicle itself, wherein the fifth position information is position information in a world coordinate system;

[0164] Record the fifth position information corresponding to each moment.

[0165] Optionally, in certain optional embodiments, obtaining fourth position information of the identified pillar during the process of the vehicle identifying the pillar includes:

[0166] During the process of vehicle recognition of the pillar, the image of the pillar is acquired through camera capture;

[0167] Inputting the column image into a pre-trained deep learning detection model to obtain information of a plurality of key points output by the deep learning detection model, wherein the key points are points on the bottom edge of the identified column;

[0168] According to the pre-calibrated parameters of the camera and the information of each key point, the fourth position information of the corresponding column is calculated.

[0169] Optionally, in certain optional embodiments, during the process of the vehicle identifying the pillar, acquiring the pillar image by a camera includes:

[0170] During the process of vehicle recognition of the pillar, multiple scene images are acquired through multiple cameras, wherein one camera corresponds to one scene image at the same time;

[0171] For any scene image, format conversion and image preprocessing are performed on the scene image to obtain a corresponding pillar image.

[0172] Optionally, in certain optional embodiments, inputting the pillar image into a pre-trained deep learning detection model to obtain information of multiple key points output by the deep learning detection model includes:

[0173] Inputting each of the pillar images into a pre-trained deep learning detection model to obtain information of multiple key points output by the deep learning detection model, wherein each pillar image outputs corresponding information of multiple key points;

[0174] The fourth position information of the corresponding column is calculated based on the pre-calibrated parameters of the camera and the information of each key point, including:

[0175] For information of multiple key points of any pillar image, fourth position information of the pillar corresponding to the pillar image is calculated based on pre-calibrated parameters of a camera used to capture the pillar image and information of each key point;

[0176] pairing each of the fourth position information with each other;

[0177] For any pair of fourth position information, calculating a matching factor between the corresponding two pieces of fourth position information;

[0178] For any pair of fourth position information, if the corresponding matching factor is greater than a preset factor threshold, the corresponding pillars are determined to be different pillars, and the corresponding two fourth position information are output respectively;

[0179] For any pair of fourth position information, if the corresponding matching factor is less than the preset factor threshold, the corresponding pillars are determined to be the same pillar, and the corresponding two fourth position information are combined into the same fourth position information for output.

[0180] Optionally, in certain optional embodiments, at any moment during the process of the vehicle identifying the pillar, if none of the cameras can detect the pillar, calculating the second position information at that moment based on the motion information of the vehicle at that moment and the first position information recorded at the previous moment includes:

[0181] At any time during the vehicle's pillar recognition process, if none of the cameras can detect the pillar, it is determined that the pillar is missed at that moment;

[0182] When determining that a pillar has been missed, obtain the movement information of the vehicle at that moment;

[0183] Obtain the first position information of the previous moment from the record;

[0184] The second position information at the moment is calculated based on the motion information of the vehicle at the moment and the first position information.

[0185] Optionally, in certain optional implementations, the calculating and obtaining the third position information and size information of the first column according to the ranging key point information includes:

[0186] Calculating fifth position information and size information of the first column based on the ranging key point information, wherein the fifth position information is position information of the first column in a world coordinate system;

[0187] If the first column is outside the preset area, performing weighted calculation based on the fifth position information of the first column at the current moment and the fifth position information of the first column at the historical moment to obtain corresponding weighted position information;

[0188] After smoothing the weighted position information using a filtering algorithm, third position information of the first column is calculated;

[0189] If the first column is within the preset area, the third position information of the first column is calculated based on the fifth position information of the first column at the current moment.

[0190] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatuses, electronic devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable device to produce a machine, so that the instructions executed by the processor of the computer or other programmable device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0191] In a typical configuration, an electronic device includes one or more processors (CPUs), a memory, and a bus. The electronic device may also include an input / output interface, a network interface, and the like.

[0192] Memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip. Memory is an example of a computer-readable medium.

[0193] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0194] In the description of the present invention, it should be understood that if the terms "up", "down", "front", "back", "left" and "right" are used to indicate directions or positional relationships, they are based on the directions or positional relationships shown in the accompanying drawings. They are only used to facilitate the description of the present invention and simplify the description, and do not indicate or imply that the positions or elements referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, they should not be understood as limitations of the present invention.

[0195] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. It should also be noted that the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, commodity, or device comprising the element.

[0196] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0197] The above are merely embodiments of the present invention and are not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A vision-based column detection method, characterized in that: include: At any moment during the vehicle recognition process, if none of the cameras can detect the pillar, second position information at that moment is calculated based on the vehicle's motion information at that moment and the first position information recorded at the previous moment, where the first position information is the position information of the pillar in the world coordinate system, and the second position information is the position information of the pillar in the vehicle coordinate system; After the vehicle starts parking, for a first pillar involved in the parking path, the first pillar is extracted from an image captured by the camera and fed into a pre-trained perception model to obtain ranging key point information of the first pillar output by the perception model; The third position information and size information of the first pillar are calculated based on the ranging key point information, wherein the third position information is the position information of the first pillar in the vehicle coordinate system.

2. The method according to claim 1, characterized in that At any time during the process of the vehicle identifying the pillar, if none of the cameras can detect the pillar, before calculating the second position information at that time based on the motion information of the vehicle at that time and the first position information recorded at the previous time, the method further includes: During the process of identifying the pillar by the vehicle, fourth position information of the identified pillar is obtained, wherein the fourth position information is position information of the identified pillar in the vehicle coordinate system; Converting the fourth position information into fifth position information according to the motion information of the vehicle itself, wherein the fifth position information is position information in a world coordinate system; Record the fifth position information corresponding to each moment.

3. The method according to claim 2, characterized in that The step of obtaining fourth position information of the identified pillar during the process of identifying the pillar by the vehicle includes: During the process of vehicle recognition of the pillar, the image of the pillar is acquired through camera capture; Inputting the column image into a pre-trained deep learning detection model to obtain information of a plurality of key points output by the deep learning detection model, wherein the key points are points on the bottom edge of the identified column; According to the pre-calibrated parameters of the camera and the information of each key point, the fourth position information of the corresponding column is calculated.

4. The method according to claim 3, characterized in that In the process of identifying the pillar by the vehicle, acquiring the pillar image by the camera includes: During the process of vehicle recognition of the pillar, multiple scene images are acquired through multiple cameras, wherein one camera corresponds to one scene image at the same time; For any scene image, format conversion and image preprocessing are performed on the scene image to obtain a corresponding pillar image.

5. The method according to claim 4, characterized in that Inputting the pillar image into a pre-trained deep learning detection model to obtain information of multiple key points output by the deep learning detection model includes: Inputting each of the pillar images into a pre-trained deep learning detection model to obtain information of multiple key points output by the deep learning detection model, wherein each pillar image outputs corresponding information of multiple key points; The fourth position information of the corresponding column is calculated based on the pre-calibrated parameters of the camera and the information of each key point, including: For information of multiple key points of any pillar image, fourth position information of the pillar corresponding to the pillar image is calculated based on pre-calibrated parameters of a camera used to capture the pillar image and information of each key point; pairing each of the fourth position information with each other; For any pair of fourth position information, calculating a matching factor between the corresponding two pieces of fourth position information; For any pair of fourth position information, if the corresponding matching factor is greater than a preset factor threshold, the corresponding pillars are determined to be different pillars, and the corresponding two fourth position information are output respectively; For any pair of fourth position information, if the corresponding matching factor is less than the preset factor threshold, the corresponding pillars are determined to be the same pillar, and the corresponding two fourth position information are combined into the same fourth position information for output.

6. The method according to claim 1, wherein At any time during the vehicle recognition process, if the pillar cannot be detected by any camera, the second position information at that time is calculated based on the vehicle's motion information at that time and the first position information recorded at the previous time, including: At any time during the vehicle's pillar recognition process, if none of the cameras can detect the pillar, it is determined that the pillar is missed at that moment; When determining that a pillar has been missed, obtain the movement information of the vehicle at that moment; Obtain the first position information of the previous moment from the record; The second position information at the moment is calculated based on the motion information of the vehicle at the moment and the first position information.

7. The method according to claim 1, characterized in that The calculating of the third position information and size information of the first column according to the ranging key point information includes: Calculating fifth position information and size information of the first column based on the ranging key point information, wherein the fifth position information is position information of the first column in a world coordinate system; If the first column is outside the preset area, performing weighted calculation based on the fifth position information of the first column at the current moment and the fifth position information of the first column at the historical moment to obtain corresponding weighted position information; After smoothing the weighted position information using a filtering algorithm, third position information of the first column is calculated; If the first column is within the preset area, the third position information of the first column is calculated based on the fifth position information of the first column at the current moment.

8. A vision-based column detection device, characterized in that: include: Missed detection calculation unit, secondary perception unit and accurate calculation unit; The missed detection calculation unit is configured to, at any moment during the process of the vehicle identifying the pillar, if none of the cameras can detect the pillar, calculate second position information at that moment based on the vehicle's motion information at that moment and the first position information recorded at a previous moment, wherein the first position information is the position information of the pillar in the world coordinate system, and the second position information is the position information of the pillar in the vehicle coordinate system; The secondary perception unit is configured to, after the vehicle begins parking, extract a first pillar involved in a parking path from an image captured by the camera and feed the image into a pre-trained perception model to obtain ranging key point information of the first pillar output by the perception model; The accurate calculation unit is used to calculate the third position information and size information of the first pillar based on the ranging key point information, wherein the third position information is the position information of the first pillar in the vehicle coordinate system.

9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the vision-based pillar detection method according to any one of claims 1 to 7 is implemented.

10. An electronic device comprising at least one processor, and at least one memory and a bus connected to the processor; wherein: The processor and the memory communicate with each other via the bus; The processor is configured to call program instructions in the memory to execute the vision-based pillar detection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image processing method and device, storage medium and electronic equipment

    CN112637482A

  • Multi-camera target matching and tracking method and device for automobile

    CN113077511A

  • Target track generation method and device, electronic equipment and medium

    CN114066974A

  • Method for detecting traffic lights and electronic equipment

    CN114529883A

  • Obstacle detection method and training method and device of image detection model

    CN114612882A