Vision sensor, information acquisition system and roadside base station
Through visual sensors, the mapping model of image pixel points and point cloud depth information is established, which solves the problem of lack of depth information in radar failure and realizes traffic safety guarantee in the case of radar failure.
Patent Information
- Application Number
- CN202010837537.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-19
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-08-19
AI Technical Summary
In the event of radar failure, the prior art is difficult to provide in-depth information to the traffic information center, resulting in a decrease in traffic safety.
The visual sensor is used to map the pixel point position of the image with the point cloud depth information of the radar sensor through a mapping model. The image is acquired by the camera and the depth information is calculated through the processor, including interpolation and fitting curve processing, and the mapping relationship between the pixel point and point cloud depth information is established.
In the event of a radar failure, the vision sensor can independently provide depth information, ensuring that the traffic information center obtains image pixel location and depth information, improving traffic safety.
Smart Images

Figure CN114078148B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of sensors, and particularly to a vision sensor, a multi-sensor information acquisition system, a roadside base station, and a vision information acquisition system. Background Art
[0002] An intelligent transportation system refers to a system in which traffic participants provide real-time traffic information from various locations to a traffic information center through sensors and transmission devices installed on roads, vehicles, etc. The traffic information center processes this information and can provide traffic participants with road traffic information and other travel-related information. Travelers can determine their travel modes and select routes based on this information, thereby ensuring traffic safety.
[0003] In related technologies, when providing real-time traffic information from various locations to a traffic information center, image data in a road scene is usually collected by a monocular camera, and depth information such as distance in the road scene is collected by a radar. Then, the camera and the radar transmit the information they collect respectively to the traffic information center for processing.
[0004] However, when the radar fails in the above technology, it is difficult to provide depth information to the traffic information center. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a vision sensor, a multi-sensor information acquisition system, a roadside base station, and a vision information acquisition system that can still provide depth information when the radar fails.
[0006] A vision sensor includes a camera, a memory, and a processor. The memory stores a mapping model.
[0007] The camera is used to obtain an image within a sensing range; the image includes a plurality of pixel points.
[0008] The processor is used to call the mapping model in the memory to process the image and obtain depth information; wherein, the mapping model includes a mapping relationship between the positions of the pixel points of the image and the depth information of the point cloud of a radar sensor, and the radar sensor is used to collect the point cloud within the sensing range.
[0009] In one embodiment, the processor is specifically configured to determine whether there is depth information corresponding to the position of the pixel point in the mapping model according to the position of the pixel point on the image; if there is depth information corresponding to the position of the pixel point, the depth information is obtained.
[0010] In one embodiment, the above-mentioned processor is specifically configured to, if there is no depth information corresponding to the position of the above-mentioned pixel point, obtain a plurality of target pixel points around the position of the above-mentioned pixel point; the above-mentioned target pixel points are pixel points with depth information; and use a preset interpolation algorithm to perform interpolation processing on the depth information corresponding to the positions of the above-mentioned target pixel points to obtain the depth information corresponding to the position of the above-mentioned pixel point.
[0011] In one embodiment, the above-mentioned processor is specifically configured to obtain the distances between each of the above-mentioned target pixel points and the above-mentioned pixel point according to the positions of each of the above-mentioned target pixel points and the position of the above-mentioned pixel point; and perform interpolation processing on the depth information corresponding to the positions of each of the above-mentioned target pixel points according to the distances between each of the above-mentioned target pixel points and the above-mentioned pixel point to obtain the depth information corresponding to the position of the above-mentioned pixel point.
[0012] In one embodiment, the above-mentioned processor is specifically configured to determine a plurality of fitting curves formed by each of the above-mentioned depth information in the above-mentioned image according to each of the above-mentioned depth information in the above-mentioned mapping model; based on the two-dimensional coordinate axes in the above-mentioned image, determine an extension line with one of the position coordinates of the above-mentioned pixel point as the center point and along the coordinate axis direction of the other position coordinate of the above-mentioned pixel point, and obtain the intersection points of the extension line and the above-mentioned plurality of fitting curves; select the two intersection points closest to the above-mentioned pixel point from the intersection points of the above-mentioned plurality of fitting curves as the above-mentioned plurality of target pixel points.
[0013] In one embodiment, the above-mentioned processor is specifically configured to obtain first historical data collected by the above-mentioned camera and second historical data collected by the above-mentioned radar sensor in the same time period and the same scene; perform spatio-temporal synchronization processing on the above-mentioned first historical data and the second historical data to obtain the spatio-temporal mapping relationship between the above-mentioned first historical data and the second historical data; extract feature information from the above-mentioned first historical data to obtain first feature information, and extract feature information from the above-mentioned second historical data to obtain second feature information; and associate the above-mentioned first feature information and the second feature information based on the spatio-temporal mapping relationship between the above-mentioned first historical data and the second historical data to establish the above-mentioned mapping model.
[0014] In one embodiment, the above-mentioned mapping model is a fitting mapping model, the relative positions of the above-mentioned camera and the above-mentioned radar sensor are fixed, the above-mentioned first historical data is a two-dimensional image of a historical object in a target scene, the above-mentioned second historical data is point cloud data of the historical object in the above-mentioned target scene, the above-mentioned first feature information is the pixel position of the historical object in the image coordinate system of the above-mentioned two-dimensional image, the above-mentioned second feature information is the depth information of the historical object, and the depth information is used to characterize the distance between the historical object and the above-mentioned radar sensor;
[0015] The above-mentioned processor is specifically configured to map the above-mentioned depth information to the corresponding pixel positions in the above-mentioned two-dimensional image to obtain a depth image, so that some pixel positions on the above-mentioned depth image have depth information; perform curve fitting connection on the above-mentioned partial pixel positions in the above-mentioned depth image to obtain multiple depth fitting curves, and determine the above-mentioned multiple depth fitting curves as the above-mentioned fitting mapping model.
[0016] In one embodiment, the above-mentioned mapping model is a deep learning model, the above-mentioned first historical data is a two-dimensional image of a historical object in a target scene, the above-mentioned second historical data is the point cloud data of the historical object in the above-mentioned target scene, the above-mentioned first feature information is the pixel position of the historical object in the image coordinate system of the above-mentioned two-dimensional image, the above-mentioned second feature information is the depth information of the historical object, and the above-mentioned depth information is used to characterize the distance between the historical object and the above-mentioned radar sensor;
[0017] The above-mentioned processor is specifically configured to use the first feature information corresponding to the above-mentioned first historical data as the training input sample of the above-mentioned deep learning model, use the second feature information corresponding to the above-mentioned second historical data as the sample label of the training input sample of the above-mentioned deep learning model to obtain a training data set; use the pixel position of the historical object in the image coordinate system of the above-mentioned two-dimensional image in the above-mentioned training data set as the input of the above-mentioned initial deep learning model, use the corresponding depth information of the historical object as the reference output of the above-mentioned initial deep learning model, and train the above-mentioned initial deep learning model to obtain the above-mentioned deep learning model.
[0018] In one embodiment, the above-mentioned processor is specifically configured to perform time synchronization on the above-mentioned first historical data and the above-mentioned second historical data to obtain multiple time-synchronized data frame pairs; each data frame pair includes a first data frame and a second data frame with synchronized sampling times; perform coordinate system conversion on the first data frame and the second data frame in each data frame pair to obtain a spatially synchronized data pair, and each data pair includes the first data in the above-mentioned first data frame and the second data in the above-mentioned second data frame that are spatially synchronized.
[0019] A multi-sensor information acquisition system includes a radar sensor and the above-mentioned vision sensor.
[0020] A roadside base station includes the above-mentioned multi-sensor information acquisition system and a roadside unit, and the vision sensing of the above-mentioned multi-sensor information acquisition system and the radar sensor of the above-mentioned multi-sensor information acquisition system are communicatively connected to the above-mentioned roadside unit,
[0021] The above-mentioned vision sensor is configured to acquire an image within a sensing range and transmit the above-mentioned image to the roadside unit;
[0022] The above-mentioned radar sensor is used to obtain the point cloud within the above-mentioned sensing range and transmit the above-mentioned point cloud to the above-mentioned roadside unit;
[0023] The above-mentioned roadside unit is used to process the received above-mentioned image and point cloud to obtain the sensing information of the roadside base station.
[0024] A visual information acquisition system includes a camera and an edge device,
[0025] The above-mentioned camera is used to obtain an image within the sensing range; the above-mentioned image includes a plurality of pixel points;
[0026] The above-mentioned edge device is used to process the above-mentioned image by using a preset mapping model to obtain depth information; wherein, the above-mentioned mapping model includes the mapping relationship between the positions of the pixel points of the image and the depth information of the point cloud of the radar sensor, and the above-mentioned radar sensor is used to collect the point cloud within the above-mentioned sensing range.
[0027] The above-mentioned visual sensor, multi-sensing information acquisition system, roadside base station and visual information acquisition system, the visual sensor includes a camera, a memory and a processor, the memory stores a mapping model, the above-mentioned camera is used to obtain an image within the sensing range; the above-mentioned image includes a plurality of pixel points; the above-mentioned processor is used to call the above-mentioned mapping model in the above-mentioned memory to process the above-mentioned image to obtain depth information; wherein, the above-mentioned mapping model includes the mapping relationship between the positions of the pixel points of the image and the depth information of the point cloud of the radar sensor, and the above-mentioned radar sensor is used to collect the point cloud within the above-mentioned sensing range. Through this visual sensor, a mapping model including the mapping relationship between the positions of the pixel points of the image and the depth information of the point cloud can be established. In this way, when the radar cannot provide depth information, the visual sensor itself can also obtain the depth information at the positions of the pixel points through the pre-established mapping model, so as to provide the image pixel point positions and depth information for the intelligent transportation information center, so that the intelligent transportation information center can process the provided image pixel point positions and depth information to ensure the safe driving of each object in the target scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is the internal structure diagram of the visual sensor in one embodiment;
[0029] Figure 2 It is the flow schematic diagram of obtaining depth information by using the interpolation method in another embodiment;
[0030] Figure 3 It is the example diagram of obtaining depth information by using the interpolation method in another embodiment;
[0031] Figure 4 It is the example diagram of multiple depth fitting curves fitted in another embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0032] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0033] In one embodiment, a vision sensor is provided. Referring to Figure 1 as shown, the vision sensor includes a camera, a memory, and a processor. The memory stores a mapping model. The camera is used to acquire an image within a sensing range. The image includes a plurality of pixel points. The processor is used to call the mapping model in the memory to process the image to obtain depth information. Among them, the mapping model includes a mapping relationship between the positions of the pixel points of the image and the depth information of the point cloud of the radar sensor. The radar sensor is used to collect point cloud within the sensing range.
[0034] Among them, the processor and the camera can be an integrated device or a split device. When the processor and the camera are split devices, the processor can be an edge device or a cloud server.
[0035] The sensing range refers to the range within which data can be detected. The sensing ranges of the camera and the radar sensor can be different or the same. When performing depth information mapping in this embodiment, it is mainly for the pixel points on the image within the sensing range of the radar sensor collected by the camera.
[0036] The image within the sensing range can be an image of a target scene. The target scene can be a road scene, and of course it can also be other scenes. The road scene can be an outdoor road scene or an indoor amusement road scene, etc. The target scene may or may not include a target object. When including a target object, the target object can be a vehicle, a pedestrian, etc. in the road scene. In addition, the plurality of pixel points included in the image can be a plurality of ground pixel points on the target object (the ground pixel points refer to the pixel points on the target object close to the ground).
[0037] The vision sensor can be a monocular camera, which can usually be set on a side pole in the road (it can be a vertical pole, a horizontal pole, etc.). In this way, the vision sensor can collect images of vehicles or pedestrians, etc. in the road scene. The camera can be a camera on a camera, a camera on a video camera, etc. Of course, the vision sensor can also be set on a target object in the road scene. For example, it can be set on a vehicle in the road to capture images of other vehicles or pedestrians in the road scene.
[0038] Before the above-mentioned processor processes the above-mentioned image using the above-mentioned mapping model in the above-mentioned memory to obtain depth information, the processor can also determine whether the roadside only includes visual sensors. If so, the processor can directly use the mapping model to process the above-mentioned image. If not, that is, the roadside includes visual sensors and radar sensors, etc., then it can first detect whether the radar sensor fails. If the radar sensor fails, the processor can use the mapping model to process the above-mentioned image. The failure of the radar sensor here refers to situations such as the radar sensor being damaged, some of the data collected by the radar sensor being damaged, the data collected by the radar being distorted, lost, or unavailable due to objective reasons such as weather conditions. The radar sensor here can be a lidar, a millimeter-wave radar, etc. The lidar can include 8-line, 16-line, 24-line, 128-line lidars, and the millimeter-wave radar can be a 24G, 77G radar, etc.
[0039] In addition, the mapping model here can be a fitting mapping model or a deep learning model. Then, before obtaining the corresponding depth information based on the position of the pixel points on the image, a fitting mapping model or a deep learning model between the position of the pixel points on the image and the depth information of the point cloud can also be obtained first.
[0040] When obtaining the fitting mapping model and the deep learning model, historical image data and historical point cloud data at various moments in the same scene can be collected. Among them, the historical image data includes the positions of the pixel points of the historical object, and the historical image data can be measured by a visual sensor. The historical point cloud data includes the depth information of the historical object at this pixel point, and this depth information can represent the distance between the historical object and the acquisition device, as well as the x, y coordinates and related angles in the physical coordinate system. The acquisition device here refers to the radar, and the historical object can be a vehicle, a pedestrian, etc. in the scene; then, the positions of the pixel points in the historical image data and the depth information in the point cloud data at the same moment are associated to obtain the mapping relationship between the two, that is, the fitting mapping model is obtained. Similarly, the positions of the pixel points in the historical image data at the same moment can be used as the input of the initial deep learning model, and the depth information in the historical point cloud data at this moment can be used as the label to train the initial deep learning model to obtain the deep learning model.
[0041] It should be noted that in terms of the roadside, usually the visual sensor and the radar sensor are installed at the same position on the roadside. Therefore, the depth information in the above-mentioned historical point cloud data can represent the distance between the historical object and the radar sensor, which is essentially the distance between the historical object and the visual sensor.
[0042] After establishing the fitting mapping model or the deep learning model, the positions of the pixel points on the image can be input into the fitting mapping model or the deep learning model to obtain the depth information of the pixel points.
[0043] Through the vision sensor in this embodiment, both the pixel positions of the pixel points on the image and the depth information of the pixel points can be obtained. It can be seen that this embodiment can obtain the pixel positions and the depth information through one vision sensor. In this way, in the actual application process, it is not necessary to install too many sensors on the roadside or the vehicle. Only by installing one vision sensor of this embodiment can the functions of both the camera and the radar be realized, thereby reducing the overall cost of the system; at the same time, only this vision sensor needs to be maintained during maintenance, which is relatively easy, thus reducing the maintenance cost; further, it can also save system space and reduce the number of sensors in the system.
[0044] It should be noted that Figure 1 The structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the vision sensor to which the solution of this application is applied. The specific vision sensor may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In addition, any reference to a memory, storage, database or other medium used in the embodiments provided in this application may include at least one of non-volatile and volatile memories. The non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. The volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0045] The above visual sensor includes a camera, a memory, and a processor. The memory stores a mapping model. The above camera is used to acquire an image within a sensing range. The above image includes a plurality of pixel points. The above processor is used to call the above mapping model in the above memory to process the above image to obtain depth information. Among them, the above mapping model includes a mapping relationship between the positions of the pixel points of the image and the depth information of the point cloud of the radar sensor. The above radar sensor is used to collect point cloud within the above sensing range. Through this visual sensor, a mapping model including the mapping relationship between the positions of the pixel points of the image and the depth information of the point cloud can be established. In this way, when the radar cannot provide depth information, the visual sensor itself can also obtain the depth information at the position of the pixel point through the pre-established mapping model, so as to provide the image pixel point position and depth information for the intelligent transportation information center, so that the intelligent transportation information center can process the provided image pixel point position and depth information to ensure the safe driving of each object in the target scene.
[0046] In another embodiment, another visual sensor is provided. Specifically, the above processor is used to determine whether there is depth information corresponding to the position of the pixel point on the above image in the above mapping model. If there is depth information corresponding to the position of the pixel point, then obtain the above depth information.
[0047] In this embodiment, the mapping model is mainly taken as a fitting mapping model for illustration. If the mapping model is a deep learning model, then when applied, by inputting the position of a pixel point, the depth information at the position of the pixel point can be obtained. However, for the fitting mapping model, since not all pixel points may be covered when establishing the mapping relationship, after obtaining the positions of the pixel points on the image, taking a pixel point as an example, it can first be determined whether there is exactly corresponding depth information for the position of the pixel point in the mapping relationship established by the fitting mapping model. If so, then the found depth information can be used as the depth information at the position of the pixel point. Of course, there may also be cases where there is no corresponding depth information for a pixel point in the mapping relationship. Then, optionally, the above processor is specifically used to, if there is no depth information corresponding to the position of the pixel point, obtain a plurality of target pixel points around the position of the pixel point. The above target pixel points are pixel points with depth information. Using a preset interpolation algorithm, interpolate the depth information corresponding to the positions of the above target pixel points to obtain the depth information corresponding to the position of the pixel point.
[0048] Among them, the number of the above target pixel points can be two or more. In this embodiment, mainly two target pixel points are taken as an example for illustration.
[0049] When obtaining the target pixel points, optionally, the above-mentioned processor is specifically configured to determine a plurality of fitting curves formed by the above-mentioned depth information on the above-mentioned image according to the respective depth information in the above-mentioned mapping model; based on the two-dimensional coordinate axes on the above-mentioned image, use one of the position coordinates of the above-mentioned pixel point as the center point and determine an extension line along the coordinate axis direction of the other position coordinate of the above-mentioned pixel point, and obtain the intersection points of the above-mentioned extension line and the above-mentioned plurality of fitting curves; select the two intersection points closest to the above-mentioned pixel point from the intersection points of the above-mentioned plurality of fitting curves as the above-mentioned plurality of target pixel points. For example, if the pixel point coordinates are (x_0, y_0), then an extension line (x = x_0) can be determined along the y-axis direction with x_0 as the center point to obtain the intersection points of x = x_0 and the plurality of fitting curves, or an extension line (y = y_0) can be determined along the x-axis direction with y_0 as the center point to obtain the intersection points of y = y_0 and the plurality of fitting curves.
[0050] When forming the fitting curve, it can be obtained by establishing the historical point cloud data of the mapping model. The specific method can be: the point cloud data can include the depth information of the historical object. This depth information can be the three-dimensional coordinates of the historical object in the world coordinate system or the polar coordinates in the polar coordinate system. Here, the three-dimensional coordinates are used for illustration. After installing the visual sensor and the radar, taking the visual sensor as the camera, the internal parameter matrix and external parameter matrix of the camera can be calculated according to the relative position between the two. Then, the three-dimensional coordinates in the historical point cloud data are transformed into the three-dimensional coordinates in the camera coordinate system through the external parameter matrix, so that the points outside the camera's field of view and the points that are not received successfully can be filtered out; then, a series of top-down points (X, Y, 0) of a series of three-dimensional points (X, Y, Z) after filtering are taken. On the Z = 0 plane, a series of points (X, Y) are transformed into the polar coordinate system to obtain the polar angles of each point in the historical point cloud data. Because the lidar has different beam numbers, at this time, two points with adjacent beam numbers and equal polar angles can be filled with points by equally dividing the polar length (i.e., performing densification processing) to obtain a plurality of filled three-dimensional coordinates in the camera coordinate system. Then, the plurality of filled three-dimensional coordinates in the camera coordinate system can be transformed into the pixel coordinate system through the internal parameter matrix, so that they can correspond to each pixel point in the historical image data in the pixel coordinate system, forming a one-to-one correspondence between the three-dimensional coordinates (i.e., depth information) of a plurality of pixel points and the positions of the pixel points. Then, curve fitting can be performed on the plurality of three-dimensional coordinates in the pixel coordinate system (for example, the three-dimensional coordinate points can be directly connected), and a plurality of fitting curves can be obtained. The fitting curve here can be a fitting circle, and of course, it can also be other curves.
[0051] Then, in actual application, after obtaining the positions of the pixel points on the image, taking a pixel point as an example, the position of this pixel point is generally a two-dimensional coordinate, assumed to be (x_0, y_0). Then, a straight line can be drawn along the positive and negative directions of the y-axis with x = x_0, or a straight line can be drawn along the positive and negative directions of the x-axis with y = y_0. Optionally, it can also be other forms of lines. In this embodiment, the description is mainly made with straight lines. This straight line will have intersections with the multiple fitting curves established above, that is, multiple intersections will be obtained. Since these intersections are all points on the fitting curves, these intersections also have corresponding depth information. Then, two intersections adjacent to the position of this pixel point, that is, the two intersections closest to this pixel point, can be found among these multiple intersections, and these two intersections are used as target pixel points, that is, two target pixel points are obtained.
[0052] Further, the interpolation algorithm here can be a spline interpolation algorithm, etc. When interpolating the depth information corresponding to the positions of the target pixel points, optionally, the above-mentioned processor is specifically configured to obtain the distances between the target pixel points and the pixel point according to the positions of the target pixel points and the position of the pixel point; according to the distances between the target pixel points and the pixel point, perform interpolation processing on the depth information corresponding to the positions of the target pixel points to obtain the depth information corresponding to the position of the pixel point.
[0053] Exemplarily, the specific process can be referred to Figure 2 as shown. The specific example diagram can be referred to Figure 2 as shown. Continuing with the position of the above-mentioned pixel point being (x_0, y_0), which is the Figure 3 No. 1 point in, and taking the obtained fitting curve as a fitting circle as an example, then two fitting circles c_1 and c_2 clamping this pixel point can be obtained according to the multiple fitting circles obtained. Selecting x_0 as the center point and drawing a straight line with x = x_0, the intersections of this straight line and these two fitting circles can be obtained, that is, two target pixel points are obtained. Assuming that the coordinates of the two obtained target pixel points are (x_0, y_1) and (x_0, y_2) respectively, which are respectively Figure 3Points 2 and 3 in it. Assume that the depth information corresponding to these two target pixel points on the fitted curve is (X_1, Y_1, Z_1) and (X_2, Y_2, Z_2) respectively. Then, through the distance ratio from (x_0, y_0) to (x_0, y_1) and (x_0, y_2), proportional interpolation calculation can be performed on (X_1, Y_1, Z_1) and (X_2, Y_2, Z_2) to obtain the three-dimensional coordinates corresponding to (x_0, y_0). Of course, it is also possible to set a weight for (X_1, Y_1, Z_1) and (X_2, Y_2, Z_2) respectively, and then perform interpolation calculation on (X_1, Y_1, Z_1) and (X_2, Y_2, Z_2) according to the set weights to obtain the three-dimensional coordinates corresponding to (x_0, y_0).
[0054] The vision sensor provided in this embodiment can determine whether there is depth information corresponding to the position of a pixel point in the mapping model through the position of the pixel point. When there is corresponding depth information, the depth information can be directly obtained. When there is no depth information, the depth information of the pixel point can be obtained by performing interpolation calculation on the depth information of multiple target pixel points with depth information, thereby expanding the applicable range and calculating the depth information corresponding to the positions of all pixel points.
[0055] In another embodiment, another vision sensor is provided. The above-mentioned processor is specifically used to obtain the first historical data collected by the above-mentioned camera and the second historical data collected by the above-mentioned radar sensor in the same time period and the same scene; perform spatio-temporal synchronization processing on the first historical data and the second historical data to obtain the spatio-temporal mapping relationship between the first historical data and the second historical data; extract feature information from the first historical data to obtain first feature information, and extract feature information from the second historical data to obtain second feature information; based on the spatio-temporal mapping relationship between the first historical data and the second historical data, associate the first feature information and the second feature information to establish the above-mentioned mapping model.
[0056] In this embodiment, the first historical data collected by the camera can be historical images, which can be images of historical objects in the target scene; the second historical data collected by the radar sensor can be historical point cloud data, which can be point cloud data of historical objects in the target scene. The first historical data and the second historical data are data collected for the historical objects in the same target scene in the same time period at a historical moment.
[0057] Before establishing the mapping model based on the first historical data and the second historical data, the vision sensor can perform spatio-temporal synchronization on the first historical data and the second historical data to correspond the first historical data and the second historical data collected in the same time period and the same scene.
[0058] Optionally, the above-mentioned processor is specifically configured to perform time synchronization on the above-mentioned first historical data and the above-mentioned second historical data to obtain multiple pairs of time-synchronized data frames; each pair of data frames includes a first data frame and a second data frame with synchronized sampling times; perform coordinate transformation on the first data frame and the second data frame in each pair of data frames to obtain spatially synchronized data pairs, and each data pair includes the first data in the above-mentioned first data frame and the second data in the above-mentioned second data frame that are spatially synchronized.
[0059] The above-mentioned first historical data may include multiple first data frames, and the second historical data may include multiple second data frames. The processor may determine whether the first data frame and the second data frame are time-synchronized according to the acquisition times of the first data frame and the second data frame, and then determine the first data frame and the second data frame with synchronized sampling times as a data pair.
[0060] When specifically performing time synchronization according to the sampling time, the processor may obtain the sampling times of each first data frame and each second data frame, and then compare the obtained sampling times to determine the synchronized data frames; in addition, the processor may also first determine the second data frame that is time-synchronized with the first first data frame, and then calculate the positions of the second data frames that are time-synchronized with other first data frames in turn according to the sampling frequencies of the first historical data and the second historical data; the above time synchronization method is not limited herein.
[0061] The sampling frequencies of the above-mentioned first historical data and the second historical data may be the same or different. When they are different, the sampling frequency of the second historical data may be a multiple of the sampling frequency of the first historical data, and the multiple may be 1, 2, 3, etc. For example, the sampling frequency of the second historical data is 20KHz, and the sampling frequency of the first historical data is 10KHz. After the processor determines the data pair with the first synchronized sampling time, it may determine the data frames after the first data frame in the first historical data that are in this data pair and the data frames after the second data frame in the second historical data that are in this data pair as the second data pair, and complete the time synchronization of each data frame in turn, which can improve the time synchronization efficiency of the first historical data and the second historical data.
[0062] When specifically performing spatial synchronization on the time-synchronized data frame pairs, each first data frame may include multiple first data, and each second data frame may include multiple second data. After the processor obtains multiple pairs of data with synchronized sampling times, it may correspond the first data and the second data in each data pair to obtain the first data and the second data obtained by the vision sensor and the radar sensor for collecting the same historical object.
[0063] Specifically, the processor can perform spatial synchronization through coordinate transformation, enabling the data collected by the visual sensor and the radar sensor to be labeled in the same coordinate system, thereby obtaining the correspondence between the first data and the second data. The processor can transform the first data to the coordinate system where the second data is located, or transform the second data to the coordinate system where the first data is located. Additionally, it can also transform both the first data and the second data to other coordinate systems, such as the Earth coordinate system, the world coordinate system, etc. The above spatial synchronization methods are not limited here. In short, through spatial synchronization, multiple spatially synchronized data pairs can be obtained. After performing spatio-temporal synchronization here, data pairs after spatio-temporal synchronization in the first historical data and the second historical data can be obtained, and thus the spatio-temporal mapping relationship between the first historical data and the second historical data can be obtained.
[0064] Furthermore, the first feature information can be the position of a pixel point, and the second feature information can be the depth information of the pixel point. Then, based on the spatio-temporal mapping relationship between the first historical data and the second historical data, the first feature information and the second feature information can be associated to establish a mapping model.
[0065] Here, optionally, the mapping model can be a fitting mapping model or a deep learning model. The following will explain these two cases separately.
[0066] The first case is to explain that the mapping model is a fitting mapping model:
[0067] When establishing the fitting mapping model, optionally, the relative position of the above camera and the above radar sensor is fixed. The above first historical data is a two-dimensional image of a historical object in the target scene, the above second historical data is the point cloud data of the historical object in the above target scene, the above first feature information is the pixel position of the historical object in the image coordinate system of the above two-dimensional image, the above second feature information is the depth information of the historical object, and the above depth information is used to represent the distance between the historical object and the above radar sensor. Specifically, the processor is used to map the above depth information to the corresponding pixel position in the above two-dimensional image to obtain a depth image, so that some pixel positions on the above depth image have depth information. Curve fitting connections are performed on the above partial pixel positions in the above depth image to obtain multiple depth fitting curves, and the above multiple depth fitting curves are determined as the above fitting mapping model.
[0068] The relative positions of the camera and the radar sensor are fixed here, which means that before establishing the fitting mapping model, when collecting historical data using the camera and the radar sensor, the relative positions between the camera and the radar sensor should be fixed. Only then can the historical data collected by them be synchronized in time and space, and the data after spatio-temporal synchronization can be used to establish the fitting mapping model. Then, when actually using the established fitting mapping model, the camera also needs to be in the position when the fitting mapping model was established, so that the depth information mapping can be performed on the data collected by the camera.
[0069] Specifically, when establishing the fitting mapping model, the depth information can be corresponded to the pixel positions in the two-dimensional image. In this way, depth information is available at some pixel positions in the two-dimensional image. After that, curve connection or curve fitting can be performed on these pixel positions with depth information, and multiple curves fitted from these pixel positions with depth information can be obtained. The fitted curves are shown in Figure 4 As shown, there are pixel positions and corresponding depth information on these fitted curves, that is, the corresponding relationship between the positions and depth information of the pixel points is established. Then, these multiple depth fitting curves here can be directly used as the fitting mapping model.
[0070] The second case: An explanation of the mapping model being a deep learning model:
[0071] When establishing the deep learning model, optionally, the above first historical data is the two-dimensional image of the historical object in the target scene, the above second historical data is the point cloud data of the historical object in the above target scene, the above first feature information is the pixel position of the historical object in the image coordinate system of the above two-dimensional image, the above second feature information is the depth information of the historical object, and the depth information is used to represent the distance between the historical object and the above radar sensor; the above processor is specifically used to use the first feature information corresponding to the above first historical data as the training input sample of the above deep learning model, use the second feature information corresponding to the above second historical data as the sample label of the training input sample of the above deep learning model, and obtain a training data set; use the pixel position of the historical object in the image coordinate system of the above two-dimensional image in the above training data set as the input of the above initial deep learning model, use the corresponding depth information of the historical object as the reference output of the above initial deep learning model, and train the above initial deep learning model to obtain the above deep learning model.
[0072] That is to say, the position of the pixel point and the corresponding depth information can be used as a data-label pair, that is, a training data pair is obtained. Then, the position of the pixel point is used as the input of the initial deep learning model, and the depth information corresponding to the position of the pixel point is used as the label of the initial deep learning model, and the initial deep learning model is trained to obtain a deep learning model. The above training data pairs may include positive samples or negative samples. When obtaining the training data pairs, training samples can be obtained based on the positions of the pixel points in all the first data and the depth information in the second data with a corresponding relationship; or training data pairs can be obtained according to the positions of the pixel points in some of the first data and the corresponding depth information in the second data, which is not limited here. In addition, the training weights of the positions and depth information of the pixel points in the training data pairs can be set according to the importance of the first data and the second data in the data pairs.
[0073] The vision sensor provided in this embodiment can perform spatio-temporal synchronization on the first historical data and the second historical data to obtain the spatio-temporal mapping relationship between the first historical data and the second historical data, and correlate the first feature information and the second feature information based on this spatio-temporal mapping relationship to establish a mapping model. By performing spatio-temporal synchronization on the first historical data and the second historical data, the data source used when establishing the mapping model can be made more accurate, and thus the mapping model finally obtained according to the more accurate data source will be more accurate.
[0074] In another embodiment, another vision sensor is provided. When specifically performing time synchronization, the processor is specifically configured to convert the first historical data and the second historical data to the same time axis; under this time axis, obtain the first sampling moments of each first data frame in the first historical data and the second sampling moments of each second data frame in the second historical data; determine the first data frame and the second data frame corresponding to the same sampling moment as a data pair according to the first sampling moment and the second sampling moment.
[0075] Among them, the first historical data is collected by the vision sensor for historical objects in the target scene, and the second historical data is collected by the radar sensor for historical objects in the target scene. The vision sensor and the radar sensor are two different devices, so the time axes of the data obtained by the two may be different. For example, the time axis of the vision sensor is the Global Positioning System (GPS) time axis, while the time axis of the radar is determined by the device itself, with a certain time axis difference. The vision sensor can convert the first historical data and the second historical data to the same time axis so that the system can obtain the first data frame and the second data frame with synchronized sampling moments.
[0076] Specifically, the processor can convert the first historical data to the time axis of the second historical data, or convert the second historical data to the time axis of the first historical data, or convert both the first historical data and the second historical data to another time axis, such as the GPS time axis. The above conversion methods are not limited herein.
[0077] In addition, the processor can obtain the first sampling moments of the first data frames in the first historical data and the second sampling moments of the second data frames in the second historical data. The above first sampling moment can be the timestamp marked on the first data frame when the visual sensor collects the first historical data, or the first sampling moment obtained according to the order of each first data frame and the start sampling moment. The obtaining method of the first sampling moment is not limited herein. The obtaining method of the second sampling moment is similar to that of the first sampling moment and will not be elaborated herein.
[0078] Further, the processor can determine the first data frame and the second data frame with the same first sampling moment and second sampling moment as a data pair; or, when the absolute value of the difference between the first sampling moment and the second sampling moment is within a certain range, the first data frame and the second data frame at this time can be determined as a data pair.
[0079] Optionally, the processor can calculate the absolute value of the difference between the first sampling moment and the second sampling moment; if the absolute value of the difference is less than a preset threshold, it is determined that the first data frame corresponding to the first sampling moment and the second data frame corresponding to the second sampling moment are a data pair. The preset threshold here can be determined according to the actual situation, such as 5ms, 10ms, 15ms, etc. Of course, if the absolute value of the difference is not less than the preset threshold, that is, greater than or equal to the preset threshold, then the data at the next sampling moment is searched according to a certain sampling frequency for time synchronization. The sampling frequency here can be 10Hz, 15Hz, etc. This can avoid the time synchronization failure caused by the incomplete coincidence of the first sampling moment and the second sampling moment due to differences in sampling frequency, etc., and improve the stability of the image depth information acquisition process.
[0080] The visual sensor in this embodiment can convert the first historical data and the second historical data to the same time axis. On this time axis, according to the first sampling moments of the first data frames in the first historical data and the second sampling moments of the second data frames in the second historical data, a data pair composed of the first data frame and the second data frame at the same sampling moment is determined. In this embodiment, since the first historical data and the second historical data can be converted to the same time axis for time synchronization, the time synchronization process can be relatively simple and reliable, and at the same time, the timeliness of the overall time synchronization can be improved.
[0081] In another embodiment, another vision sensor is provided. When specifically performing spatial synchronization, the processor is specifically configured to convert the second data in the data pair into the pixel coordinate system according to a preset transformation matrix to obtain the point coordinates of each second data in the pixel coordinate system; in the data pair, obtain the first data corresponding to the point coordinates; and associate the second data corresponding to each point coordinate with the first data corresponding to the point coordinate to obtain a corresponding relationship.
[0082] Here, the transformation matrix can be composed of an extrinsic matrix and an intrinsic matrix. The extrinsic matrix refers to the matrix that converts the data collected by the radar sensor into the vision sensor coordinate system, which can be determined according to the relative pose between the vision sensor and the radar sensor. The above relative pose can include the translation amount and the rotation angle between the vision sensor and the radar sensor; the intrinsic matrix can be the transformation matrix inside the vision sensor, which characterizes the pose of the vision sensor, etc., and can be the transformation matrix from the vision sensor coordinate system to the pixel coordinate system. The extrinsic matrix and the intrinsic matrix here can be obtained by manual measurement by the staff and then input into the vision sensor; or they can be automatically calibrated by the vision sensor according to the current relative pose relationship between the vision sensor and the radar sensor and the internal pose of the vision sensor, etc., which is not limited here; in addition, the extrinsic matrix and the intrinsic matrix here can include 6 degrees of freedom such as translation and rotation.
[0083] The processor can first use the extrinsic matrix to convert the second data in the data pair into the vision sensor coordinate system according to the preset transformation matrix, and then use the intrinsic matrix to convert the data converted by the extrinsic matrix into the pixel coordinate system, so that the coordinates of each second data in the pixel coordinate system can be obtained.
[0084] Further, the above second data can be the three-dimensional coordinates in the radar coordinate system, and of course, it can also be the three-dimensional coordinates in the world coordinate system. The processor can first convert the above three-dimensional coordinates to the vision sensor coordinate system through the calculated extrinsic matrix, and then convert the three-dimensional coordinates in the vision sensor coordinate system to the three-dimensional coordinates in the pixel coordinate system through the intrinsic matrix. At this time, the three-dimensional coordinates of the second data in the pixel coordinate system correspond one-to-one with the first data at the pixel coordinate, that is, there is a corresponding coordinate of the second data at the position of one first data. In this way, the first data and the second data in the same data pair can be spatially synchronized.
[0085] The visual sensor in this embodiment can convert the second data in each data pair into the pixel coordinate system according to a preset conversion matrix, obtain the coordinates of each second data in the pixel coordinate system, and associate them with the first data at this coordinate to obtain the above corresponding relationship. Through the data space conversion in this embodiment, the first historical data and the second historical data can be synchronized in space, thereby providing a data basis for the establishment of the mapping model and also making the depth information obtained from the established data pairs more accurate subsequently.
[0086] In one embodiment, a multi-sensor information acquisition system is provided, including the above visual sensor and radar sensor. Of course, the multi-sensor information acquisition system may also include other components.
[0087] Through the multi-sensor new acquisition system of this embodiment, since the visual sensor it includes can establish a mapping model including the position of the pixel points of the image and the depth information of the point cloud, when the radar sensor cannot provide depth information, the roadside base station itself can also obtain the depth information at the position of the pixel points through the mapping model pre-established by its visual sensor, so as to provide the image pixel point position and depth information for the intelligent transportation information center, so that the intelligent transportation information center can process the provided image pixel point position and depth information to ensure the safe driving of each object in the target scene.
[0088] In one embodiment, a roadside base station is provided, including the above multi-sensor information acquisition system and a roadside unit. The visual sensor of the multi-sensor information acquisition system and the radar sensor of the multi-sensor information acquisition system are communicatively connected to the roadside unit. The visual sensor is used to acquire an image within the sensing range and transmit the image to the roadside unit; the radar sensor is used to acquire the point cloud within the sensing range and transmit the point cloud to the roadside unit; the roadside unit is used to process the received image and point cloud to obtain the sensing information of the roadside base station.
[0089] Among them, the above visual sensor and radar sensor are both connected to the roadside unit and can transmit the data collected by each to the roadside unit. The connection between the roadside unit and the visual sensor and the radar sensor can be wired or wireless. The roadside unit can be a computer device, such as a terminal, a server, an edge device, etc. The sensing information of the roadside base station can be the depth information of each pixel point on the image, the speed, acceleration, trajectory, angular velocity, heading angle, etc. of the target object on the image.
[0090] With the roadside base station of this embodiment, since the included vision sensor can establish a mapping model including the mapping relationship between the positions of the pixel points of the image and the depth information of the point cloud, when the radar sensor cannot provide depth information, the roadside base station itself can also obtain the depth information at the positions of the pixel points through the mapping model pre-established by its vision sensor, so as to provide the intelligent transportation information center with the position and depth information of the image pixel points, so that the intelligent transportation information center can process the provided position and depth information of the image pixel points to ensure the safe driving of each object in the target scene.
[0091] In one embodiment, a vision information acquisition system is provided, including a camera and an edge device. The camera is used to acquire an image within the sensing range; the image includes a plurality of pixel points; the edge device is used to process the image by using a preset mapping model to obtain depth information; wherein, the mapping model includes the mapping relationship between the positions of the pixel points of the image and the depth information of the point cloud of the radar sensor, and the radar sensor is used to acquire the point cloud within the sensing range.
[0092] Here, the process of the edge device specifically establishing the mapping model and using the established mapping model can be the same as the steps executed by the processor in the above vision sensor, and will not be elaborated here.
[0093] With the vision information acquisition system of this embodiment, a mapping model including the mapping relationship between the positions of the pixel points of the image and the depth information of the point cloud can be established. When the radar sensor cannot provide depth information, the vision information acquisition system itself can also obtain the depth information at the positions of the pixel points through the pre-established mapping model, so as to provide the intelligent transportation information center with the position and depth information of the image pixel points, so that the intelligent transportation information center can process the provided position and depth information of the image pixel points to ensure the safe driving of each object in the target scene.
[0094] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should be considered as the scope described in this specification.
[0095] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A visual sensor, comprising a camera, a memory, and a processor, wherein the memory stores a mapping model, and is characterized in that the camera is configured to acquire an image within a sensing range; the image includes a plurality of pixel points; the processor is configured to call the mapping model in the memory to process the image to obtain depth information; wherein, the mapping model includes a mapping relationship between the positions of the pixel points of the image and the depth information of the point cloud of a radar sensor, and the radar sensor is configured to collect point cloud within the sensing range; specifically, the processor is configured to determine whether there is depth information corresponding to the position of the pixel point in the mapping model according to the position of the pixel point on the image; specifically, the processor is configured to obtain the depth information if there is depth information corresponding to the position of the pixel point; specifically, the processor is configured to obtain a plurality of target pixel points around the position of the pixel point if there is no depth information corresponding to the position of the pixel point; the target pixel points are pixel points with depth information; and an interpolation algorithm is used to perform interpolation processing on the depth information corresponding to the positions of the target pixel points to obtain the depth information corresponding to the position of the pixel point; specifically, the processor is configured to obtain the distances between each target pixel point and the pixel point according to the positions of each target pixel point and the position of the pixel point; and interpolation processing is performed on the depth information corresponding to the positions of each target pixel point according to the distances between each target pixel point and the pixel point to obtain the depth information corresponding to the position of the pixel point; specifically, the processor is configured to determine a plurality of fitting curves formed by the respective depth information in the mapping model on the image; based on the two-dimensional coordinate axes on the image, an extension line is determined with one of the position coordinates of the pixel point as the center point and along the axis direction of the other position coordinate of the pixel point, and the intersections of the extension line and the plurality of fitting curves are obtained; and two intersections closest to the pixel point are selected from the intersections with the plurality of fitting curves as the plurality of target pixel points.
2. The visual sensor according to claim 1, wherein specifically, the processor is configured to obtain first historical data collected by the camera and second historical data collected by the radar sensor in the same time period and the same scene; perform spatio-temporal synchronization processing on the first historical data and the second historical data to obtain a spatio-temporal mapping relationship between the first historical data and the second historical data; extract first feature information from the first historical data, extract second feature information from the second historical data; and associate the first feature information and the second feature information based on the spatio-temporal mapping relationship between the first historical data and the second historical data to establish the mapping model.
3. The visual sensor according to claim 2, characterized in that, The mapping model is a fitting mapping model. The relative positions of the camera and the radar sensor are fixed. The first historical data is a two-dimensional image of historical objects in the target scene. The second historical data is the point cloud data of the historical objects in the target scene. The first feature information is the pixel position of the historical object in the image coordinate system of the two-dimensional image. The second feature information is the depth information of the historical object, and the depth information is used to represent the distance between the historical object and the radar sensor; Specifically, the processor is configured to map the depth information to the corresponding pixel positions in the two-dimensional image to obtain a depth image, so that some pixel positions on the depth image have depth information; perform curve fitting connection on the partial pixel positions in the depth image to obtain a plurality of depth fitting curves, and determine the plurality of depth fitting curves as the fitting mapping model.
4. The visual sensor according to claim 2, wherein The mapping model is a deep learning model. The first historical data is a two-dimensional image of historical objects in the target scene. The second historical data is the point cloud data of the historical objects in the target scene. The first feature information is the pixel position of the historical object in the image coordinate system of the two-dimensional image. The second feature information is the depth information of the historical object, and the depth information is used to represent the distance between the historical object and the radar sensor; Specifically, the processor is configured to use the first feature information corresponding to the first historical data as the training input sample of the deep learning model, use the second feature information corresponding to the second historical data as the sample label of the training input sample of the deep learning model to obtain a training data set; use the pixel position of the historical object in the image coordinate system of the two-dimensional image in the training data set as the input of the initial deep learning model, and use the corresponding depth information of the historical object as the reference output of the initial deep learning model to train the initial deep learning model to obtain the deep learning model.
5. The vision sensor according to claim 2, wherein Specifically, the processor is configured to perform time synchronization on the first historical data and the second historical data to obtain a plurality of pairs of time-synchronized data frames; each pair of data frames includes a first data frame and a second data frame with synchronized sampling times; Perform coordinate transformation on the first data frame and the second data frame in each pair of data frames to obtain a spatially synchronized data pair, and each data pair includes the first data in the first data frame and the second data in the second data frame that are spatially synchronized.
6. A multi-sensor information acquisition system, characterized in that, It includes a radar sensor and the vision sensor according to any one of claims 1-5.
7. A roadside base station, characterized in that, It includes the multi-sensor information acquisition system according to claim 6 and a roadside unit. The vision sensing of the multi-sensor information acquisition system and the radar sensor of the multi-sensor information acquisition system are communicatively connected to the roadside unit, The vision sensor is configured to acquire an image within the sensing range and transmit the image to the roadside unit; The radar sensor is used to obtain the point cloud within the sensing range and transmit the point cloud to the roadside unit; The roadside unit is used to process the received image and point cloud to obtain the sensing information of the roadside base station.
8. A visual information acquisition system, comprising a camera and an edge device, characterized in that The camera is used to obtain an image within the sensing range; the image includes a plurality of pixel points; The edge device is used to process the image by using a preset mapping model to obtain depth information; wherein, the mapping model includes the mapping relationship between the positions of the pixel points of the image and the depth information of the point cloud of the radar sensor, and the radar sensor is used to collect the point cloud within the sensing range; The edge device is specifically used to determine whether there is depth information corresponding to the position of the pixel point in the mapping model according to the position of the pixel point on the image; The edge device is specifically used to obtain the depth information if there is depth information corresponding to the position of the pixel point; The edge device is specifically used to obtain a plurality of target pixel points around the position of the pixel point if there is no depth information corresponding to the position of the pixel point; the target pixel points are pixel points with depth information; an interpolation algorithm is used to perform interpolation processing on the depth information corresponding to the positions of the target pixel points to obtain the depth information corresponding to the position of the pixel point; The edge device is specifically used to obtain the distance between each target pixel point and the pixel point according to the positions of each target pixel point and the position of the pixel point; and perform interpolation processing on the depth information corresponding to the positions of each target pixel point according to the distance between each target pixel point and the pixel point to obtain the depth information corresponding to the position of the pixel point; The edge device is specifically used to determine a plurality of fitting curves formed by each depth information in the mapping model on the image; based on the two-dimensional coordinate axis on the image, determine an extension line with the position coordinate of one of the pixel points as the center point and along the axis direction of the other position coordinate of the pixel point, and obtain the intersection points of the extension line and the plurality of fitting curves; select the two intersection points closest to the pixel point from the intersection points of the plurality of fitting curves as the plurality of target pixel points.
Citation Information
Patent Citations
Method for collecting and storing multidimensional information data of an environment and a target
CN109522951A