Traffic information collection method and device applied to road section scenario and electronic equipment
By integrating and correcting the visual and radar sensors in the radar-visual integrated machine, the optical distortion problem in traffic information collection has been solved, enabling accurate collection of traffic information and accurate reflection of movement trajectories, thus improving the effectiveness of traffic control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2024-02-20
- Publication Date
- 2026-07-21
AI Technical Summary
Existing visual and radar sensors suffer from optical distortion and aberrations in traffic information collection in road scenarios, making it impossible to accurately obtain the fused trajectory and traffic information of various objects.
The system employs a radar-integrated machine that combines visual and radar sensors. By correcting the lane inclination and distortion-reducing mapping relationship of the visual sensor and combining it with the global projection relationship, it corrects and fuses position information to obtain an accurate fused trajectory.
It enables precise collection of traffic information, better reflects the actual movement trajectory of objects, and improves the accuracy of traffic control.
Smart Images

Figure CN117953689B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of traffic management, and in particular to methods, devices and electronic equipment for collecting traffic information in road segment scenarios. Background Technology
[0002] Intelligent traffic management relies on accurately collecting traffic information for road segments, such as traffic flow, vehicle spacing, time required for vehicles or pedestrians to pass through the segment, and the number and length of vehicles queuing during congestion. Currently, in intelligent traffic management applications, the location information of the same object collected by visual sensors and radar sensors at the same time point is typically fused in the same coordinate system to obtain the fused trajectory of the target object, thereby enabling traffic information collection based on the fused trajectories of various objects on the road segment.
[0003] However, on the one hand, visual sensors suffer from optical distortion when collecting location information, resulting in inaccurate location information. On the other hand, the location information collected by visual sensors is also distorted during the conversion process. This makes it impossible for the fused trajectories of the objects to accurately reflect their actual movement trajectories, thus making it impossible to accurately obtain traffic information for that road segment. Summary of the Invention
[0004] In view of this, embodiments of this application provide a method, apparatus and electronic device for collecting traffic information in road segment scenarios, so as to correct the above-mentioned distortions and accurately obtain the fusion trajectory of each target, thereby accurately obtaining traffic information of the road segment scenario.
[0005] This application provides a method for collecting traffic information in a road segment scenario. The method is applied to a radar-visual integrated machine, which integrates at least two visual sensors and at least two radar sensors. The method includes:
[0006] For any visual sensor, obtain the first position information of the first object in the first lane at the first time point; correct the first position information according to the lane inverse slope k in the image coordinate system corresponding to the visual sensor to obtain the second position information; the lane inverse slope k is determined based on the image of the first lane collected by the visual sensor.
[0007] Check whether the second position information is within the target distortion region corresponding to the vision sensor. If so, based on the obtained distortion correction mapping relationship and the obtained global projection relationship corresponding to the vision sensor, map the second position information to the specified global coordinate system to obtain the first global position information of the first object at the first time point under the vision sensor. If not, map the second position information to the global coordinate system according to the global projection relationship to obtain the first global position information of the first object at the first time point under the vision sensor. The distortion correction mapping relationship is used to compensate for the deviation of the second position information mapped to the global coordinate system. The target distortion region is determined based on the sampling points where distortion occurs on each lane line.
[0008] The first global position information of the first object at each time point under each visual sensor is fused with the second global position information of the first object at each time point under each radar sensor to obtain the fused trajectory of the first object. The fused trajectory is used for traffic information collection.
[0009] This application embodiment also provides a traffic information collection device applied in a road segment scenario. This device is applied to a radar-visual integrated machine, which integrates at least two visual sensors and at least two radar sensors; the device includes:
[0010] The correction module is used to obtain, for any visual sensor, the first position information of a first object in a first lane at a first time point; and to correct the first position information according to the lane inverse slope k in the image coordinate system corresponding to the visual sensor to obtain second position information; wherein the lane inverse slope k is determined based on the image of the first lane acquired by the visual sensor.
[0011] The mapping module is used to check whether the second position information is within the target distortion region corresponding to the visual sensor. If so, based on the obtained distortion correction mapping relationship and the obtained global projection relationship, the second position information is mapped to a specified global coordinate system to obtain the first global position information of the first object at the first time point under the visual sensor. If not, the second position information is mapped to the global coordinate system based on the global projection relationship to obtain the first global position information of the first object at the first time point under the visual sensor. The distortion correction mapping relationship is used to compensate for the deviation of the second position information mapped to the global coordinate system. The target distortion region is determined based on the sampling points where distortion occurs on each lane line.
[0012] The fusion module is used to fuse the first global position information of the first object at each time point under each visual sensor with the second global position information of the first object at each time point under each radar sensor to obtain the fused trajectory of the first object, which is used for traffic information collection.
[0013] This application also provides an electronic device, including: a processor and a memory for storing computer program instructions, which, when executed by the processor, cause the processor to perform the steps of the method described above.
[0014] This application also provides a machine-readable storage medium storing computer program instructions that, when executed, enable the implementation of the steps described above.
[0015] As can be seen from the above technical solution, in this embodiment, for any visual sensor, after obtaining the first position information of the first object in the first lane at the first time point collected by the visual sensor, the first position information is corrected using the lane inverse slope k in the image coordinate system corresponding to the visual sensor. This method of simultaneously correcting the distortion of the first position information of the first object in the first lane based on correcting the linear distortion of the first lane effectively corrects linear distortion and obtains accurate second position information, providing accurate second position information for subsequent fusion. This ensures that the fused trajectories of each object obtained subsequently accurately reflect the actual movement trajectories of each object, resulting in more accurate traffic information and better traffic management.
[0016] Furthermore, in this embodiment, when the obtained second position information is detected to be within the target distortion region corresponding to the visual sensor, the second position information is mapped to a specified global coordinate system based on the obtained distortion correction mapping relationship corresponding to the visual sensor and the obtained global projection relationship. This compensation for the deviation in mapping the second position information to the global coordinate system through the distortion correction mapping relationship effectively corrects the nonlinear distortion generated during the mapping process, thereby obtaining more accurate first global position information. This further enables the fused trajectory obtained by fusing the first global position information of the first object at each time point under each visual sensor with the second global position information of the first object at each time point under each radar sensor to accurately reflect the actual movement trajectory of each object, precisely collect traffic information, and achieve intelligent traffic control. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the method flow provided in the embodiments of this application.
[0018] Figure 2 This is a flowchart illustrating the process of obtaining the target distortion region corresponding to the visual sensor, as provided in an embodiment of this application.
[0019] Figure 3 This is a flowchart illustrating the process of obtaining the target distortion region and distortion correction mapping matrix corresponding to the visual sensor, as provided in an embodiment of this application.
[0020] Figure 4 A flowchart for obtaining the global projection relationship provided in an embodiment of this application.
[0021] Figure 5 This is a schematic diagram of the device structure provided in the embodiments of this application.
[0022] Figure 6 This is a schematic diagram of the electronic device structure provided in an embodiment of this application. Detailed Implementation
[0023] Exemplary embodiments will now be described in detail, examples of which are identified in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings identify the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0024] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0025] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0026] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, and to make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, the technical solutions in the embodiments of this application will be further described in detail below with reference to the accompanying drawings.
[0027] This embodiment provides a method for collecting traffic information in road segment scenarios. Here, the road segment scenario can be an intersection, urban road, highway, tunnel, etc., and this embodiment is not specifically limited to these scenarios.
[0028] The method provided in the embodiments of this application is described below:
[0029] See Figure 1 , Figure 1 This is a schematic diagram of the method flow provided in an embodiment of this application. The process is applied to a radar-vision integrated machine, which integrates at least two visual sensors and at least two radar sensors. In this embodiment, the integrated setup of the radar-vision integrated machine, compared to setting up multiple sensors separately, not only effectively ensures the time synchronization between the sensors but also facilitates installation and debugging, effectively improving the adaptability of the radar-vision integrated machine. Optionally, the visual sensor can be a monocular fisheye camera, and the radar sensor can be a millimeter-wave radar; this embodiment is not specifically limited.
[0030] like Figure 1 As shown, the process may include the following steps:
[0031] S101, for any vision sensor, obtain the first position information of the first object in the first lane at the first time point collected by the vision sensor; correct the first position information according to the lane inverse slope k in the image coordinate system corresponding to the vision sensor to obtain the second position information; the lane inverse slope k is determined based on the image of the first lane collected by the vision sensor.
[0032] In this embodiment, for any visual sensor, target recognition is performed on the image acquired by the visual sensor at the first time point to obtain the first position information of the first object in the first lane at the first time point.
[0033] When the image is acquired by the vision sensor, the first lane is distorted due to optical distortion. This distortion can be corrected using the lane anti-slope k. Since both the first lane and the first object in the first lane are distorted due to optical distortion, correcting the distortion of the first lane is equivalent to correcting the first position information. Therefore, in this embodiment, the first position information is corrected using the lane anti-slope k to obtain the corrected second position information. Optionally, the lane anti-slope k can be obtained by identifying the position information of the two endpoints of the first lane in the image acquired by the vision sensor and calculating k based on the obtained position information of the two endpoints.
[0034] In this embodiment, step S101, correcting the first position information based on the lane anti-slope k in the image coordinate system corresponding to the visual sensor, can be implemented in many ways. For example, obtaining the width value w along the width direction and the height value h along the height direction of the first object detection region in the first image, and adjusting the values along the width direction and the height direction of the first position information based on the width value w, the height value h, and the lane anti-slope k. Here, the first image is the image of the first object in the first lane at the first time point acquired by the visual sensor, and the first object detection region refers to the region in the first image that contains at least the first object.
[0035] It should be noted that the first position information and the second position information are position information in the image coordinate system corresponding to the vision sensor, such as the coordinates in the image coordinate system.
[0036] For example, the first object detection area mentioned above is the target detection box that encloses the first object. The first position information is the coordinates of the center point of the bottom edge of the target detection box, denoted as (x, y). Then the second position information is (x+kw, y-kh), where w represents the width of the target detection box and h represents the height of the target detection box.
[0037] S102, check whether the second position information is in the target distortion region corresponding to the vision sensor. If so, based on the obtained distortion correction mapping relationship and the obtained global projection relationship corresponding to the vision sensor, map the second position information to the specified global coordinate system to obtain the first global position information of the first object at the first time point under the vision sensor. If not, map the second position information to the global coordinate system according to the global projection relationship to obtain the first global position information of the first object at the first time point under the vision sensor. The distortion correction mapping relationship is used to compensate for the deviation of the second position information mapped to the global coordinate system. The target distortion region is determined based on the sampling points where distortion occurs on each lane line.
[0038] In this embodiment, the global coordinate system can be the radar coordinate system corresponding to any radar sensor, or a designated radar coordinate system other than the radar coordinate system corresponding to all radar sensors.
[0039] For any given vision sensor, both the testing and operational phases involve collecting location information from the same road segment and projecting it onto the same global coordinate system. Therefore, the distortion caused by the mapping process and the methods for compensating for this distortion are similar in both phases. It is evident that the target distortion region of the vision sensor obtained during the testing phase can be used to characterize areas prone to distortion in subsequent operational phases. The distortion correction mapping relationship of the vision sensor obtained during the testing phase can also compensate for the deviation in mapping the second location information to the global coordinate system during subsequent operational phases. Therefore, in this embodiment, step S102 corrects the distortion during the mapping of the second location information to the global coordinate system, thereby obtaining the corrected first global location information.
[0040] In this embodiment, the target distortion region and distortion correction mapping relationship corresponding to any visual sensor can be obtained during the testing process of this Ravage all-in-one machine. Examples will be provided below, and details will not be elaborated here. The global projection relationship can also be obtained during the testing process of this Ravage all-in-one machine, which will also be described with examples below, and details will not be elaborated here.
[0041] S103, the first global position information of the first object at each time point under each visual sensor is fused with the second global position information of the first object at each time point under each radar sensor to obtain the fused trajectory of the first object, and the fused trajectory is used for traffic information collection.
[0042] In this embodiment, the second global position information of any object at any time point under any radar sensor refers to the position information of the mapped point obtained by mapping the position information of the object collected by the radar sensor at that time point from the current radar coordinate system to the global coordinate system.
[0043] In this embodiment, under the global coordinate system, the first global location information and the second global location information that meet the set similarity requirement at the same time point are determined to belong to the same object. For example, the second global location information that meets the set similarity requirement with the first global location information of the first object at time point t1 is determined to be the second global location information of the first object at time point t1. The set similarity requirement can be set according to the specific application scenario, and this embodiment is not specifically limited.
[0044] In this embodiment, step S103 involves fusing the first global position information of the first object at each time point under each visual sensor with the second global position information of the first object at each time point under each radar sensor to obtain the fused trajectory of the first object. This process can be implemented in many ways. For example, as one embodiment, for each time point, the first and second global position information of the first object at that time point are processed in a specific way, such as by averaging or weighted averaging, to obtain the fused trajectory points of the first object at that time point. Furthermore, target tracking is performed on the first object to establish the association between the fused trajectory points of the first object at each time point, thereby obtaining the fused trajectory of the first object.
[0045] It should be noted that the field of view collected by each visual sensor overlaps. When any object is within the overlapping area of the field of view, multiple visual sensors will collect the object's position information at the same time. If the object is within the overlapping area of the field of view, only one visual sensor will collect the object's position information at the same time. When multiple visual sensors collect the object's position information at the same time, there will be multiple first global position information for the object at each time point in the global coordinate system. Therefore, a deduplication operation is performed on the multiple first global position information belonging to the object, retaining only one first global position information.
[0046] It should also be noted that due to limitations imposed by the installation angle of the radar-visual integrated machine and the specific structure of the radar sensors within it, blind zones exist for each radar sensor. Therefore, when any object is in a blind zone, the radar sensor cannot obtain the object's position information. At that point in time, there is no second global position information for the object in the global coordinate system. In this case, there is no need for fusion; the first global position information of the object is directly used as the position information of the fused trajectory point.
[0047] This concludes the process. Figure 1 The process is shown below.
[0048] pass Figure 1 As shown in the flowchart, for any visual sensor, after obtaining the first position information of the first object in the first lane at the first time point, the first position information is corrected using the lane inverse slope k in the image coordinate system corresponding to the visual sensor. This method of simultaneously correcting the distortion of the first position information of the first object in the first lane based on correcting the linear distortion of the first lane effectively corrects linear distortion and obtains accurate second position information, providing accurate second position information for subsequent fusion. This ensures that the fused trajectories of each object obtained subsequently accurately reflect their actual movement trajectories, resulting in more accurate traffic information and better traffic management.
[0049] Furthermore, in this embodiment, when the obtained second position information is detected to be within the target distortion region corresponding to the visual sensor, the second position information is mapped to a specified global coordinate system based on the obtained distortion correction mapping relationship corresponding to the visual sensor and the obtained global projection relationship. This compensation for the deviation in mapping the second position information to the global coordinate system through the distortion correction mapping relationship effectively corrects the nonlinear distortion generated during the mapping process, thereby obtaining more accurate first global position information. This further enables the fused trajectory obtained by fusing the first global position information of the first object at each time point under each visual sensor with the second global position information of the first object at each time point under each radar sensor to accurately reflect the actual movement trajectory of each object, precisely collect traffic information, and achieve intelligent traffic control.
[0050] The following describes the target distortion region and distortion correction mapping relationship obtained for any visual sensor during the testing phase:
[0051] See Figure 2 , Figure 2 This is a flowchart illustrating the process of obtaining the target distortion region corresponding to the visual sensor, as provided in an embodiment of this application. Figure 2 As shown, the process may include the following steps:
[0052] S201, according to the stitching parameters configured for the vision sensor, convert each sampling point on at least one lane line already obtained by the vision sensor to obtain the conversion point corresponding to each sampling point.
[0053] In this embodiment, the stitching parameter m of any visual sensor is pre-configured and pre-calculated using conventional stitching parameter calculation methods.
[0054] In this embodiment, at least one lane line is identified from images of a road segment scene acquired by the vision sensor at any historical time point during the testing phase. Multiple sampling points are selected on each lane line according to a set sampling interval. That is, the sampling points on at least one lane line already acquired by the vision sensor are obtained. It should be noted that the aforementioned set sampling interval can be set according to the actual application scenario, and this embodiment is not specifically limited to it.
[0055] S202, performing specified processing on each sampling point on at least one lane line already acquired by the vision sensor to select the distorted points that have been distorted from the sampling points; the specified processing includes at least rotation and translation.
[0056] In this embodiment, step S202 involves specifying the processing of each sampling point on at least one lane line acquired by the visual sensor to select the distorted points. This process can be implemented in many ways. For example, as one embodiment, specifying the processing of each sampling point on at least one lane line acquired by the visual sensor yields a target processing point corresponding to each sampling point. For each sampling point, the positional deviation between the corresponding conversion point and the target processing point is checked to see if it meets the distortion requirements. If it does, the sampling point is determined to be a distorted point; otherwise, it is determined not to be a distorted point.
[0057] Optionally, the specific implementation of the above-mentioned specified processing of each sampling point on at least one lane line acquired by the vision sensor to obtain the target processing point corresponding to each sampling point can be as follows: According to the current processing parameters, specified processing is performed on each sampling point on at least one lane line acquired by the vision sensor to obtain candidate processing points corresponding to each sampling point. If the deviation between the conversion point corresponding to each sampling point and the candidate processing point meets the set requirements, then the candidate processing point corresponding to each sampling point is determined as the target processing point corresponding to each sampling point; otherwise, the current processing parameters are adjusted, and the process returns to the step of specified processing of each sampling point on at least one lane line acquired by the vision sensor according to the current processing parameters.
[0058] Optionally, the specific implementation of checking whether the positional deviation between the conversion point and the target processing point corresponding to the sampling point meets the distortion requirements can be as follows: Determine a deviation threshold based on the positional deviation between the conversion point and the target processing point corresponding to each sampling point. For example, the average of the obtained positional deviations can be used as the aforementioned deviation threshold. For any sampling point, if the positional deviation between the conversion point and the target processing point corresponding to that sampling point is greater than or equal to the deviation threshold, then that sampling point is determined to be a distorted point; otherwise, it is determined that the sampling point is not a distorted point.
[0059] In this embodiment, since each visual sensor can only capture a portion of the road scene, the process of mapping the second location information to the global coordinate system encompasses the stitching process. The stitching process involves mapping using stitching parameters, which involves linear transformations (rotation and translation) and nonlinear distortions. Linear transformations are simulated by applying specified processing, including rotation and translation, to the sampling points. Nonlinear distortions are represented by the deviation between the target processing point and the transformation point corresponding to each sampling point. The greater the deviation between these two, the deeper the nonlinear distortion. Therefore, sampling points whose positional deviation between the corresponding transformation point and the target processing point is greater than or equal to a deviation threshold are identified as distortion points.
[0060] S203, determine the target distortion region based on each distortion point.
[0061] In this embodiment, the determination of the target distortion region based on each distortion point in step S203 can be implemented in many ways. For example, the region enclosed by the connected domain formed by each distortion point can be determined as the target distortion region.
[0062] As an example, after obtaining the target distortion region corresponding to any vision sensor, for any vision sensor, based on the positional deviation between the conversion point and the target processing point corresponding to each distortion point in the target distortion region corresponding to that vision sensor, a distortion correction vector corresponding to that distortion point is determined; based on the distortion correction vectors corresponding to each distortion point, a distortion correction mapping relationship is determined. Here, the distortion correction vector is used to adjust the conversion point corresponding to the distortion point to be closer to the target processing point. For example, if the conversion point corresponding to the distortion point is offset 0.5 meters to the upper right relative to the target processing point, then the distortion correction vector will move the distortion point 0.5 meters to the lower left.
[0063] It should be noted that, in another embodiment, after obtaining the distortion points of each lane line, the distortion correction mapping matrix corresponding to the visual sensor can also be obtained by solving the functional relationship between the conversion point corresponding to each distortion point and the target processing point corresponding to each distortion point.
[0064] To illustrate this in more detail, the following examples illustrate the acquisition of the target distortion region and distortion correction mapping relationship corresponding to the visual sensor.
[0065] Example 1:
[0066] In this embodiment 1, refer to Figure 3 Obtaining the target distortion region corresponding to visual sensor A includes the following steps:
[0067] S301. From the images of the road segment scene collected by the vision sensor at any historical time point during the testing phase, obtain each sampling point on each lane line. Here, the sequence formed by each sampling point is denoted as (A, B).
[0068] S302. Using the stitching parameters m1 configured on the visual sensor A, the sampling point sequence (A, B) is transformed to obtain the transformed point sequence (A', B').
[0069] S303. The current processing parameters are represented by the current rotation and translation matrix D. According to the current D, the visual sensor sampling point sequence (A, B) is processed to obtain the candidate processing point sequence (HA, HB”) corresponding to each sampling point.
[0070] S304. Determine whether the deviation between the candidate processing point sequence (HA”, HB”) and the transition point sequence (A’, B’) meets the set requirements.
[0071] If the result of step S304 is yes, then proceed to step S305; otherwise, adjust D and return to step S303.
[0072] S305, then the candidate processing point sequence (HA”, HB”) is determined as the target processing point sequence (A”, B”) corresponding to each sampling point.
[0073] S306. Obtain the positional deviation E between the conversion point and the target processing point corresponding to each sampling point, and obtain the mean value SE of each positional deviation E.
[0074] S307. For any sampling point, if the positional deviation between the conversion point and the target processing point corresponding to the sampling point is greater than or equal to SE, then the sampling point is determined to be a distorted point where distortion has occurred; otherwise, the sampling point is determined not to be a distorted point where distortion has occurred.
[0075] S308. After obtaining all the distortion points on each lane line in the above image, the region enclosed by the connected domain formed by each distortion point is determined as the target distortion region.
[0076] S309. For each distortion point in the target distortion region obtained above, based on the positional deviation between the conversion point and the target processing point corresponding to the distortion point, determine the distortion correction vector corresponding to the distortion point, and concatenate the distortion correction vectors corresponding to each distortion point to obtain the distortion correction mapping relationship corresponding to the video sensor.
[0077] The following describes the global projection relationship obtained during the testing phase:
[0078] See Figure 4 , Figure 4 A flowchart illustrating the process of obtaining global projection relationships, provided for embodiments of this application. For example... Figure 4 As shown, the process may include the following steps:
[0079] S401. Obtain the overlapping area in the images acquired by each vision sensor at the same time point.
[0080] Given that the field of view of each visual sensor overlaps, there will also be overlapping areas in the images acquired by each visual sensor at the same time point. Optionally, as an example, the images acquired by each visual sensor are stitched together using the stitching parameters configured for each visual sensor, and the overlapping areas are obtained in the stitched image through feature matching.
[0081] S402. Using the current calibration mapping matrix of each vision sensor, the overlapping area is mapped to the radar coordinate system corresponding to the radar sensor paired with the vision sensor to obtain the mapped area; if the distance between at least one pair of target points in the mapped area does not meet the requirements, the current calibration mapping matrix of the vision sensor is adjusted; if the distance between all pairs of target points in the mapped area meets the requirements, the current calibration mapping matrix of the vision sensor is determined as the target calibration mapping matrix of the vision sensor; wherein, any pair of target points in the mapped area refers to the mapping points obtained by mapping two points on adjacent lane lines in the overlapping area to the radar coordinate system.
[0082] In this embodiment, the paired visual sensors and radar sensors are installed in positions that meet the requirements, and visual sensors and radar sensors located on the same side are paired with each other. The image coordinate system corresponding to the paired visual sensor and the radar coordinate system corresponding to the radar sensor have a calibration relationship. Through a conventional radar-visual calibration algorithm, the calibration mapping matrix between the above-mentioned image coordinate system and radar coordinate system can be obtained. At this time, the calibration mapping matrix configured for the visual sensor is obtained. Then, in step S402, the above-mentioned configured calibration mapping matrix is used as the initial value of the current calibration mapping matrix.
[0083] S403. Determine the global projection relationship based on the target calibration mapping matrix of each visual sensor.
[0084] In this embodiment, after obtaining the target calibration mapping matrix of the visual sensor, a specified operation is performed on the obtained target calibration mapping matrix to obtain the global projection relationship.
[0085] The specific implementation methods of the above steps S402 and S403 will be illustrated with examples later, and will not be repeated here.
[0086] To illustrate this in more detail, the following examples demonstrate how the global projection relationship can be obtained.
[0087] Example 2:
[0088] In this embodiment 2, for ease of explanation, two visual sensors are used as an example. In the radar-visual integrated machine, the two sensors are arranged back-to-back, and the two radar sensors are arranged back-to-back. Obtaining the global projection relationship includes the following steps:
[0089] 1. Obtain Image 1 and Image 2, acquired by visual sensor 1 and visual sensor 2 at any historical time point during the testing phase, and obtain the overlapping region in these two images. The overlapping region in Image 1 is denoted as v1, and the overlapping region in Image 2 is denoted as v2. Here, v1 is the sequence of pixels in the overlapping region of Image 1, and v2 is the sequence of pixels in the overlapping region of Image 2.
[0090] 2. Obtain the calibration mapping matrix c1 configured for vision sensor 1, and obtain the calibration mapping matrix c2 configured for vision sensor 2.
[0091] 3. In the overlapping area V1, select two points on adjacent lane lines and use c1 to map them to radar coordinate system 1 to obtain two target points. Calculate the distance between the two target points. If the obtained distance is less than or equal to the set value, then the current c1 is determined as the target calibration mapping matrix c1' of the vision sensor. Otherwise, adjust the c1 parameter until the distance between the two target points is less than or equal to the set value, until the target calibration mapping matrix c1' is obtained.
[0092] Here, the above setting value can be the standard road width value of 3.75m.
[0093] 4. In the overlapping area V1, select two points on adjacent lane lines and use c2 to map them to radar coordinate system 2 to obtain two target points. Calculate the distance between the two target points. If the obtained distance is less than or equal to the set value, then the current c2 is determined as the target calibration mapping matrix c2' of the vision sensor. Otherwise, adjust the c2 parameter until the distance between the two target points is less than or equal to the set value, until the target calibration mapping matrix c2' is obtained.
[0094] 5. After obtaining the target calibration mapping matrix c1' corresponding to visual sensor 1 and the target calibration mapping matrix c2' corresponding to visual sensor 2, the transformation relationship between radar coordinate system 1 corresponding to radar sensor 1 and radar coordinate system 2 corresponding to radar sensor 2 is derived in the following way.
[0095] v1*c1'=r1, v2*c2'=r2, v1=r1*c1'^-1, v2=r2*c2'^-1, v1 and v2 belong to the overlapping region. In the visual dimension, v1=v2, then r1*c1'^-1=r2*c2'^-1, and therefore r1*c1'^-1*c2'=r2.
[0096] It can be seen that the transformation relationship between radar coordinate system 1 and radar coordinate system 2 is c1'^-1*c2'.
[0097] 6. The global projection relationship can be represented by the following expression: Q = c1 * c1'^-1 * c2', or c2 * c1'^-1 * c2'.
[0098] The methods provided in the embodiments of this application have been described above. The apparatus provided in the embodiments of this application is described below:
[0099] See Figure 5 , Figure 5 This is a structural diagram of the device provided in an embodiment of this application. The device is applied to a radar-visual integrated machine, which integrates at least two visual sensors and at least two radar sensors, such as... Figure 5 As shown, the device may include: a correction module 501, a mapping module 502, and a fusion module 503.
[0100] Correction module 501 is used to obtain, for any visual sensor, the first position information of the first object in the first lane at a first time point collected by the visual sensor; and to correct the first position information according to the lane anti-slope k in the image coordinate system corresponding to the visual sensor to obtain second position information; wherein the lane anti-slope k is determined based on the image of the first lane collected by the visual sensor.
[0101] The mapping module 502 is used to check whether the second position information is located in the target distortion region corresponding to the vision sensor. If so, based on the obtained distortion correction mapping relationship and the obtained global projection relationship corresponding to the vision sensor, the second position information is mapped to the specified global coordinate system to obtain the first global position information of the first object under the vision sensor at the first time point. If not, based on the global projection relationship, the second position information is mapped to the global coordinate system to obtain the first global position information of the first object under the vision sensor at the first time point. The distortion correction mapping relationship is used to compensate for the deviation of the second position information mapped to the global coordinate system. The target distortion region is determined based on the sampling points where distortion occurs on each lane line.
[0102] The fusion module 503 is used to fuse the first global position information of the first object at each time point under each visual sensor with the second global position information of the first object at each time point under each radar sensor to obtain the fused trajectory of the first object, and the fused trajectory is used for traffic information collection.
[0103] As one embodiment, correcting the first position information based on the lane inverse slope k in the image coordinate system corresponding to the visual sensor includes:
[0104] The width value w along the width direction and the height value h along the height direction of the first object detection region in the first image are obtained; the first image is an image of the first object in the first lane at a first time point acquired by the vision sensor; the first object detection region refers to the region in the first image that contains at least the first object.
[0105] Based on the width value w, the height value h, and the lane inclination k, adjust the values along the width direction and the height direction in the first position information.
[0106] As an example,
[0107] The target distortion region corresponding to any visual sensor is determined through the following steps:
[0108] According to the stitching parameters configured for the vision sensor, each sampling point on at least one lane line already obtained by the vision sensor is converted to obtain the conversion point corresponding to each sampling point.
[0109] The visual sensor performs specified processing on each sampling point on at least one lane line to select distorted points from the sampling points; the specified processing includes at least rotation and translation.
[0110] The target distortion region is determined based on each distortion point.
[0111] As an example,
[0112] The specified processing of each sampling point on at least one lane line acquired by the vision sensor to select the distorted points from the sampling points includes:
[0113] The visual sensor has already acquired at least one sampling point on a lane line. The sampling points are then processed in a specified manner to obtain the target processing points corresponding to each sampling point.
[0114] For each sampling point, check whether the positional deviation between the corresponding conversion point and the target processing point meets the distortion requirements. If so, determine that the sampling point is a distortion point where distortion has occurred; otherwise, determine that the sampling point is not a distortion point where distortion has occurred.
[0115] As one embodiment, the step of performing specified processing on each sampling point on at least one lane line already obtained by the vision sensor to obtain the target processing point corresponding to each sampling point includes:
[0116] Based on the current processing parameters, each sampling point on at least one lane line already obtained by the vision sensor is processed in a specified manner to obtain the candidate processing point corresponding to each sampling point.
[0117] If the deviation between the conversion point and the candidate processing point corresponding to each sampling point meets the set requirements, then the candidate processing point corresponding to each sampling point is determined as the target processing point corresponding to each sampling point; otherwise, the current processing parameters are adjusted, and the process of performing specified processing on each sampling point on at least one lane line obtained by the vision sensor according to the current processing parameters is returned.
[0118] As an example, the distortion correction mapping relationship corresponding to any visual sensor is determined through the following steps:
[0119] For any vision sensor, based on the positional deviation between the conversion point and the target processing point corresponding to each distortion point in the target distortion region corresponding to the vision sensor, a distortion correction vector corresponding to the distortion point is determined; the distortion correction vector is used to adjust the conversion point corresponding to the distortion point to be closer to the target processing point.
[0120] The distortion correction mapping relationship is determined based on the distortion correction vector corresponding to each distortion point.
[0121] As an example, the global projection relationship is determined through the following steps:
[0122] Obtain the overlapping region in images acquired by each visual sensor at the same time point;
[0123] Using the current calibration mapping matrix of each vision sensor, the overlapping region is mapped to the radar coordinate system corresponding to the radar sensor paired with that vision sensor to obtain the mapped region. If the distance between at least one pair of target points in the mapped region does not meet the requirements, the current calibration mapping matrix of the vision sensor is adjusted. If the distance between all pairs of target points in the mapped region meets the requirements, the current calibration mapping matrix of the vision sensor is determined as the target calibration mapping matrix of the vision sensor. Here, any pair of target points in the mapped region refers to the mapping points obtained by mapping two points on adjacent lane lines in the overlapping region to the radar coordinate system.
[0124] The global projection relationship is determined based on the target calibration mapping matrix of each visual sensor.
[0125] This concludes the process. Figure 5 Structural description of the device shown.
[0126] Please see Figure 6This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device may include a processor 601, a communication interface 602, a memory 603, and a communication bus 60. The processor 601, communication interface 602, and memory 603 communicate with each other via the communication bus 60. The memory 603 stores a computer program; the processor 601 can execute the steps of the method described in the above embodiments by executing the program stored in the memory 603. Depending on the actual function of the electronic device, other hardware may also be included, which will not be elaborated further.
[0127] The embodiments of the subject matter and functional operation described in this specification can be implemented in the following ways: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, the program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information and transmit it to a suitable receiving device for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof.
[0128] The processing and logic flow described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform corresponding functions by operating on input data and generating output. The processing and logic flow can also be executed by dedicated logic circuitry—such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the device can also be implemented as dedicated logic circuitry.
[0129] Suitable computers for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or the computer will be operatively coupled to such mass storage devices to receive data from or transfer data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0130] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROMs and DVD-ROMs. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.
[0131] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.
[0132] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0133] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0134] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for collecting traffic information in road segment scenarios, characterized in that, This method is applied to a radar-visual integrated machine, which integrates at least two visual sensors and at least two radar sensors; the method includes: For any visual sensor, obtain the first position information of the first object in the first lane at the first time point; correct the first position information according to the lane inverse slope k in the image coordinate system corresponding to the visual sensor to obtain the second position information; the lane inverse slope k is determined based on the image of the first lane collected by the visual sensor. Check whether the second position information is within the target distortion region corresponding to the vision sensor. If so, based on the obtained distortion correction mapping relationship and the obtained global projection relationship corresponding to the vision sensor, map the second position information to the specified global coordinate system to obtain the first global position information of the first object at the first time point under the vision sensor. If not, map the second position information to the global coordinate system based on the global projection relationship to obtain the first global position information of the first object at the first time point under the vision sensor. The distortion correction mapping relationship is used to compensate for the deviation of the second position information mapped to the global coordinate system. The target distortion region is determined based on the sampling points where distortion occurs on each lane line. The first global position information of the first object at each time point under each visual sensor is fused with the second global position information of the first object at each time point under each radar sensor to obtain the fused trajectory of the first object. The fused trajectory is used for traffic information collection.
2. The method according to claim 1, characterized in that, The step of correcting the first position information based on the lane inverse slope k in the image coordinate system corresponding to the visual sensor includes: The width value w along the width direction and the height value h along the height direction of the first object detection region in the first image are obtained; the first image is an image of the first object in the first lane at a first time point acquired by the vision sensor; the first object detection region refers to the region in the first image that contains at least the first object. Based on the width value w, the height value h, and the lane inclination k, adjust the values along the width direction and the height direction in the first position information.
3. The method according to claim 1, characterized in that, The target distortion region corresponding to any visual sensor is determined through the following steps: According to the stitching parameters configured for the vision sensor, each sampling point on at least one lane line already obtained by the vision sensor is converted to obtain the conversion point corresponding to each sampling point. The visual sensor performs specified processing on each sampling point on at least one lane line to select distorted points from the sampling points; the specified processing includes at least rotation and translation. The target distortion region is determined based on each distortion point.
4. The method according to claim 3, characterized in that, The specified processing of each sampling point on at least one lane line acquired by the vision sensor to select the distorted points from the sampling points includes: The visual sensor has already acquired at least one sampling point on a lane line. The sampling points are then processed in a specified manner to obtain the target processing points corresponding to each sampling point. For each sampling point, check whether the positional deviation between the corresponding conversion point and the target processing point meets the distortion requirements. If so, determine that the sampling point is a distortion point where distortion has occurred; otherwise, determine that the sampling point is not a distortion point where distortion has occurred.
5. The method according to claim 4, characterized in that, The step of performing specified processing on each sampling point on at least one lane line obtained by the visual sensor to obtain the target processing point corresponding to each sampling point includes: Based on the current processing parameters, each sampling point on at least one lane line already obtained by the vision sensor is processed in a specified manner to obtain the candidate processing point corresponding to each sampling point. If the deviation between the conversion point and the candidate processing point corresponding to each sampling point meets the set requirements, then the candidate processing point corresponding to each sampling point is determined as the target processing point corresponding to each sampling point; otherwise, the current processing parameters are adjusted, and the process of performing specified processing on each sampling point on at least one lane line obtained by the vision sensor according to the current processing parameters is returned.
6. The method according to claim 1, characterized in that, The distortion correction mapping relationship for any visual sensor is determined through the following steps: For any vision sensor, based on the positional deviation between the conversion point and the target processing point corresponding to each distortion point in the target distortion region corresponding to the vision sensor, a distortion correction vector corresponding to the distortion point is determined; the distortion correction vector is used to adjust the conversion point corresponding to the distortion point to be closer to the target processing point. The distortion correction mapping relationship is determined based on the distortion correction vector corresponding to each distortion point.
7. The method according to claim 1, characterized in that, The global projection relationship is determined through the following steps: Obtain the overlapping region in images acquired by each visual sensor at the same time point; Using the current calibration mapping matrix of each visual sensor, the overlapping region is mapped to the radar coordinate system corresponding to the radar sensor paired with that visual sensor to obtain the mapped region. If the distance between at least one pair of target points in the mapping area does not meet the requirements, the current calibration mapping matrix of the vision sensor is adjusted. If the distance between all pairs of target points in the mapping area meets the requirements, the current calibration mapping matrix of the vision sensor is determined as the target calibration mapping matrix of the vision sensor. Here, any pair of target points in the mapping area refers to the mapping points obtained by mapping two points on adjacent lane lines in the overlapping area to the radar coordinate system. The global projection relationship is determined based on the target calibration mapping matrix of each visual sensor.
8. A traffic information collection device applied in road segment scenarios, characterized in that, This device is used in a radar-view integrated machine, which integrates at least two visual sensors and at least two radar sensors; the device includes: The correction module is used to obtain, for any visual sensor, the first position information of a first object in a first lane at a first time point; and to correct the first position information according to the lane inverse slope k in the image coordinate system corresponding to the visual sensor to obtain second position information; wherein the lane inverse slope k is determined based on the image of the first lane acquired by the visual sensor. The mapping module is used to check whether the second position information is within the target distortion region corresponding to the visual sensor. If so, based on the obtained distortion correction mapping relationship and the obtained global projection relationship, the second position information is mapped to a specified global coordinate system to obtain the first global position information of the first object at the first time point under the visual sensor. If not, the second position information is mapped to the global coordinate system based on the global projection relationship to obtain the first global position information of the first object at the first time point under the visual sensor. The distortion correction mapping relationship is used to compensate for the deviation of the second position information mapped to the global coordinate system. The target distortion region is determined based on the sampling points where distortion occurs on each lane line. The fusion module is used to fuse the first global position information of the first object at each time point under each visual sensor with the second global position information of the first object at each time point under each radar sensor to obtain the fused trajectory of the first object, which is used for traffic information collection.
9. The apparatus according to claim 8, characterized in that, The step of correcting the first position information based on the lane inverse slope k in the image coordinate system corresponding to the visual sensor includes: The width value w along the width direction and the height value h along the height direction of the first object detection region in the first image are obtained; the first image is an image of the first object in the first lane at a first time point acquired by the vision sensor; the first object detection region refers to the region in the first image that contains at least the first object. Based on the width value w, the height value h, and the lane inclination k, adjust the values along the width direction and the height direction in the first position information; And / or, The target distortion region corresponding to any visual sensor is determined through the following steps: According to the stitching parameters configured for the vision sensor, each sampling point on at least one lane line already obtained by the vision sensor is converted to obtain the conversion point corresponding to each sampling point. The visual sensor performs specified processing on each sampling point on at least one lane line to select distorted points from the sampling points; the specified processing includes at least rotation and translation. The target distortion region is determined based on each distortion point; And / or, The specified processing of each sampling point on at least one lane line acquired by the vision sensor to select the distorted points from the sampling points includes: The visual sensor has already acquired at least one sampling point on a lane line. The sampling points are then processed in a specified manner to obtain the target processing points corresponding to each sampling point. For each sampling point, check whether the positional deviation between the corresponding conversion point and the target processing point meets the distortion requirements. If so, determine that the sampling point is a distortion point where distortion occurs; otherwise, determine that the sampling point is not a distortion point where distortion occurs. And / or, The step of performing specified processing on each sampling point on at least one lane line obtained by the visual sensor to obtain the target processing point corresponding to each sampling point includes: Based on the current processing parameters, each sampling point on at least one lane line already obtained by the vision sensor is processed in a specified manner to obtain the candidate processing point corresponding to each sampling point. If the deviation between the conversion point corresponding to each sampling point and the candidate processing point meets the set requirements, then the candidate processing point corresponding to each sampling point is determined as the target processing point corresponding to each sampling point; otherwise, the current processing parameters are adjusted, and the process of performing specified processing on each sampling point on at least one lane line obtained by the vision sensor according to the current processing parameters is returned. And / or, The distortion correction mapping relationship for any visual sensor is determined through the following steps: For any vision sensor, based on the positional deviation between the conversion point and the target processing point corresponding to each distortion point in the target distortion region corresponding to the vision sensor, a distortion correction vector corresponding to the distortion point is determined; the distortion correction vector is used to adjust the conversion point corresponding to the distortion point to be closer to the target processing point. The distortion correction mapping relationship is determined based on the distortion correction vector corresponding to each distortion point; And / or, The global projection relationship is determined through the following steps: Obtain the overlapping region in images acquired by each visual sensor at the same time point; Using the current calibration mapping matrix of each vision sensor, the overlapping region is mapped to the radar coordinate system corresponding to the radar sensor paired with that vision sensor to obtain the mapped region. If the distance between at least one pair of target points in the mapped region does not meet the requirements, the current calibration mapping matrix of the vision sensor is adjusted. If the distance between all pairs of target points in the mapped region meets the requirements, the current calibration mapping matrix of the vision sensor is determined as the target calibration mapping matrix of the vision sensor. Here, any pair of target points in the mapped region refers to the mapping points obtained by mapping two points on adjacent lane lines in the overlapping region to the radar coordinate system. The global projection relationship is determined based on the target calibration mapping matrix of each visual sensor.
10. An electronic device, characterized in that, include: processor; as well as A memory storing computer program instructions that, when executed by the processor, cause the processor to perform the steps of the method as described in any one of claims 1 to 7.