Data processing methods and equipment

CN116762094BActive Publication Date: 2026-09-01SZ ZHUOYU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180079742.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-17
Publication Date
2026-09-01
Estimated Expiration
2041-03-17

AI Technical Summary

Technical Problem

[0004]然而,智能汽车对高精地图的需求是不同的,现有的高精地图供给方式无法适应智能汽车的个性化的地图需求

Benefits of technology

[0017]综上所述,本申请实施例提供数据处理方法和设备,可移动平台在移动时采集空间场景的图像数据,并基于图像数据生成地图数据,再使用所采集地图数据控制其运动,可满足可移动平台个性化地图数据需求。基于图像数据生成地图数据,可使用可移动平台上已有图象传感器,无需再配置激光雷达等高成本传感器采集点云,降低地图建构成本。并且,在判断地图元数据满足建图质量要求后再根据地图元数据生成地图数据,保证可移动平台所生成地图数据的准确度。此外,本申请中基于图像数据生成地图数据,地图数据存储更轻量,对于地图的实时更新、维护极具便利。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116762094B_ABST
    Figure CN116762094B_ABST
Patent Text Reader

Abstract

A data processing method and apparatus include: controlling an image sensor located on a mobile platform to acquire multiple frames of image data of the spatial scene when the platform moves within a spatial scene; processing the multiple frames of image data to obtain map metadata; wherein the map metadata includes any one or more combinations of 3D feature points, texture data, and semantic information; determining whether the map metadata meets mapping quality requirements; if the map metadata meets the mapping quality requirements, generating map data based on the map metadata; wherein the map data is used to control the movement of the mobile platform within the spatial scene. This solution can meet the personalized map data needs of the mobile platform and also ensure the accuracy of the generated map data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, and in particular to a data processing method and device. Background Technology

[0002] High-precision map data is an important foundation for autonomous driving and is of great significance to the development of intelligent vehicles.

[0003] High-definition map data is typically provided by map providers. Map providers usually only offer high-definition map data with high usage, and not high-definition map data with low usage.

[0004] However, intelligent vehicles have different requirements for high-precision maps, and the existing high-precision map supply methods cannot meet the personalized map needs of intelligent vehicles. Summary of the Invention

[0005] This application provides a data processing method and apparatus, aiming to provide a solution that can adapt to the personalized map needs of different mobile platforms.

[0006] Firstly, this application provides a data processing method, including:

[0007] When the mobile platform moves within the spatial scene, the image sensor located on the mobile platform is controlled to acquire multi-frame image data of the spatial scene;

[0008] Map metadata is obtained by processing multiple frames of image data; the map metadata includes any one or more combinations of three-dimensional feature points, texture data, and semantic information.

[0009] Determine whether the map metadata meets the mapping quality requirements;

[0010] If the map metadata meets the mapping quality requirements, map data is generated based on the map metadata; the map data is used to control the movement of the mobile platform within the spatial scene.

[0011] Secondly, this application provides a control device, comprising: a memory for storing instructions and a processor for executing the instructions stored in the memory, wherein the processor is used to specifically execute:

[0012] When the mobile platform moves within the spatial scene, the image sensor located on the mobile platform is controlled to acquire multi-frame image data of the spatial scene;

[0013] Map metadata is obtained by processing multiple frames of image data; the map metadata includes any one or more combinations of three-dimensional feature points, texture data, and semantic information.

[0014] Determine whether the map metadata meets the mapping quality requirements;

[0015] If the map metadata meets the mapping quality requirements, map data is generated based on the map metadata; wherein, the map data is used to control the movement of the mobile platform within the spatial scene.

[0016] Thirdly, this application provides a mobile platform, including an image sensor and the data processing method involved in the second aspect.

[0017] In summary, the embodiments of this application provide a data processing method and apparatus. A mobile platform collects image data of a spatial scene while moving, generates map data based on the image data, and then uses the collected map data to control its movement, thus meeting the personalized map data needs of the mobile platform. Generating map data based on image data allows the use of existing image sensors on the mobile platform, eliminating the need for high-cost sensors such as LiDAR to collect point clouds, thereby reducing map construction costs. Furthermore, generating map data based on map metadata only after determining that the map metadata meets the mapping quality requirements ensures the accuracy of the map data generated by the mobile platform. In addition, the map data generation based on image data in this application results in lighter map data storage, greatly facilitating real-time map updates and maintenance. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the structure of a movable platform provided in an embodiment of this application;

[0020] Figure 2 A schematic flowchart illustrating a data processing method provided in another embodiment of this application;

[0021] Figure 3 A schematic flowchart illustrating a data processing method provided in another embodiment of this application;

[0022] Figure 4 A schematic flowchart illustrating a data processing method provided in another embodiment of this application;

[0023] Figure 5 A schematic flowchart illustrating a data processing method provided in another embodiment of this application;

[0024] Figure 6 This is a schematic diagram of the structure of a control device provided in another embodiment of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] It should be noted that when a component is said to be "fixed" to another component, it can be directly on the other component or it can be in a middle component. When a component is said to be "connected" to another component, it can be directly connected to the other component or it may be in a middle component.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0028] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0029] Existing high-precision map supply methods cannot meet the personalized map requirements of intelligent vehicles. To address the aforementioned technical problems, this application provides a data processing method and apparatus. The technical concept of this application is as follows: a mobile platform collects image data of a spatial scene while moving within it, and uses the collected data to generate map data, which can adapt to the personalized map data requirements of mobile platforms. Before generating map data, it is determined whether the metadata used to generate the map data meets the mapping quality requirements to ensure that the map data generated by the mobile platform can accurately reflect the spatial scene. Furthermore, generating map data based on image data collected by image sensors eliminates the need for high-cost image sensors, reducing map construction costs.

[0030] like Figure 1As shown, one embodiment of this application provides a mobile platform 100, which includes an image sensor 101, a driving sensor (not shown), and a control device (not shown). The image sensor 101 is used to collect image data of the scene surrounding the mobile platform 100, the driving sensor is used to collect driving data of the mobile platform, and the control device is used to execute the data processing method described below. For details, please refer to the following description, which will not be repeated here.

[0031] This application can be used to solve the problem of automatic parking in autonomous driving functions, and can be used for map building during short-distance automatic parking, such as within 300 meters. The map is mainly used to record various landmarks within the parking lot, including parking spaces, traffic signs, road lane lines, and landmark buildings. After map building, it can assist in realizing functions such as parking lot location recognition, automatic parking space search within the map area, and locating vehicles at any position within the map area during automatic parking.

[0032] like Figure 2 As shown, this application provides a data processing method, in which a control device is the executing entity, and the method specifically includes the following steps:

[0033] S201. When the mobile platform moves within the spatial scene, the control device controls the image sensor located on the mobile platform to collect multiple frames of image data of the spatial scene.

[0034] When the mobile platform enters a certain spatial scene, the image sensor on the mobile platform is controlled to work. The image sensor collects image data of the spatial scene and transmits the image data to the control device.

[0035] Preferably, the driving sensor collects the position information of the mobile platform and transmits the collected position information to the control device. The control device determines whether it has entered a certain spatial scene based on the position information. When it is determined that it has entered the designated spatial scene, it controls the image sensor on the mobile platform to work. The image sensor transmits the collected multi-frame image data to the control device.

[0036] S202. The control equipment processes multi-frame image data to obtain map metadata.

[0037] Map metadata is used to generate map data, and it includes any one or more combinations of 3D feature points, texture data, and semantic information.

[0038] Three-dimensional feature points are used to reflect the position and shape of objects in the scene space, texture data is used to reflect the surface information of objects in the scene space, and semantic information is used to reflect the category of objects represented by texture data and three-dimensional feature points.

[0039] The aforementioned map metadata is obtained by performing processes such as extracting two-dimensional feature data from image data, matching two-dimensional feature data from multiple frames of image data, and semantic recognition.

[0040] Furthermore, to determine whether the map metadata meets the mapping quality requirements, step S203 lists a specific implementation method:

[0041] S203. If the map metadata meets the mapping quality requirements, the control device generates map data based on the map metadata.

[0042] Among them, the mapping quality requirements are used to determine whether the above map metadata is rich enough, that is, whether the amount of map metadata is sufficient and whether the types of map metadata are sufficient.

[0043] If the map metadata is rich enough, the quality of the map data built using that metadata will be higher, meaning the map data will more accurately describe the spatial scene. If the map metadata is small in volume and limited in type, the quality of the map data built using that metadata will be lower, meaning the map data will not accurately describe the spatial scene.

[0044] When the map metadata meets the mapping quality requirements, the map metadata is processed to obtain map data, for example, the map metadata is processed into images to obtain each layer.

[0045] Map data is used to control the movement of the mobile platform within the spatial scene. The control device can control the mobile platform's movement within the spatial scene at the current moment based on the map data generated in the previous moment. The control device can also control the mobile platform's movement within the spatial scene based on the generated map data the next time it re-enters the spatial scene.

[0046] In the above technical solution, the mobile platform collects image data during its movement, generates map data based on the image data, and then uses the collected map data to control its movement. This satisfies the mobile platform's personalized map data requirements and directly utilizes the existing image sensors of the mobile platform to collect data, eliminating the need for high-cost sensors such as LiDAR. Furthermore, after determining that the map metadata meets the mapping quality requirements, map data is generated based on the map metadata, ensuring the accuracy of the generated map data.

[0047] like Figure 3 As shown, another embodiment of this application provides a data processing method, in which the execution subject is a control device, and the method specifically includes the following steps:

[0048] S301. When the mobile platform moves within the spatial scene, the control device controls the image sensor located on the mobile platform to collect multiple frames of image data of the spatial scene.

[0049] The mobile platform is equipped with multiple image sensors located around it. These sensors are used to collect image data within the spatial scene where the mobile platform is located. When the control device acquires image data, it also acquires the image data from the image sensors that are collecting the image data.

[0050] S302, The control device processes multi-frame image data to obtain map metadata.

[0051] This step has been described in detail in the above embodiments and will not be repeated here.

[0052] S303. The control device uses the identifier of the image sensor to mark the source of the map metadata and obtain the marked map metadata.

[0053] Specifically, after processing multiple frames of image data to obtain map metadata, the obtained map metadata is tagged with the identifier of the image sensor that acquired the image data. In other words, the map metadata is tagged with its data source.

[0054] S304. If the marked map metadata meets the mapping quality requirements, the control device generates marked map data based on the marked map metadata.

[0055] Among them, the mapping quality requirements are used to determine whether the labeled map data is rich enough. Map metadata includes any one or more combinations of 3D feature points, texture data, and semantic information.

[0056] The quality requirements for map construction include at least one of the following: the total number of three-dimensional feature points reaches the first threshold; there are at least two three-dimensional feature points with different components on all three coordinate axes; the total number of texture data reaches the second threshold; the number of texture data types reaches the third threshold; the total number of semantic information reaches the fourth threshold; and the number of semantic information types reaches the fifth threshold.

[0057] The abundance of 3D feature points is determined by checking if the total number of 3D feature points reaches a first threshold. The abundance of 3D feature points in terms of type is determined by checking if there are at least two 3D feature points with distinct components on all three coordinate axes. If all 3D feature points lie on the same plane, meaning all 3D feature points have the same component on one of the coordinate axes (e.g., all 3D feature points have the same z-axis component), then the 3D feature points can only represent a single plane and cannot represent a rich 3D spatial scene.

[0058] By determining whether the total number of texture data reaches a second threshold, we can determine whether the texture data is sufficiently abundant in quantity. By determining whether the number of texture data types reaches a third threshold, we can determine whether the texture data is sufficiently abundant in type.

[0059] By determining whether the total amount of semantic information reaches the fourth threshold, it is determined whether the texture data is sufficiently rich in quantity. By determining whether the number of semantic information categories reaches the fifth threshold, it is determined whether the semantic information is sufficiently rich in type.

[0060] Based on the richness of the map metadata in terms of quantity and type, it is determined whether the map metadata meets the requirements for mapping quality. The map data generated based on the map metadata can accurately reflect the spatial scene.

[0061] After obtaining map metadata that meets the mapping quality requirements, map data is generated based on the labeled map metadata. The process of generating map data specifically includes at least one of the following:

[0062] A labeled feature data layer is generated based on the labeled 3D feature points and labeled texture data; and a labeled semantic information layer is generated based on the labeled semantic information.

[0063] Preferably, the labeled 3D features and labeled texture data are image-processed to obtain a labeled feature data layer. The labeled semantic information is then image-processed to obtain a labeled semantic information layer.

[0064] In other embodiments, a parking space layer can be generated based on the marked semantic information, specifically including: extracting semantic information representing parking spaces from the marked semantic information, and performing image processing based on the semantic information representing parking spaces to generate a marked parking space layer.

[0065] The marked map data includes marker information indicating the data source. When using map data, it can be filtered based on the marker information and direction of movement. The filtered map data can then be used to control the movement of the mobile platform, reducing the amount of data processing involved in using map data and enabling the mobile platform to generate control commands more quickly based on the map data.

[0066] In the above technical solution, the richness of the map metadata in terms of quantity and type is combined to determine whether the map metadata meets the mapping quality requirements. Accurate map data reflecting the spatial scene can be obtained based on this map metadata. Furthermore, source tagging processing of the map metadata ensures that the obtained map data also reflects its source. When using map data, data can be filtered based on its source, reducing data processing volume. This allows for the rapid generation of control commands based on the map data, enabling precise movement of the mobile platform.

[0067] like Figure 4 As shown, another embodiment of this application provides a data processing method, in which the execution subject is a control device, and the method specifically includes the following steps:

[0068] S401. When the mobile platform moves within the spatial scene, the control device controls the image sensor located on the mobile platform to collect multi-frame image data of the spatial scene.

[0069] S402. The control device processes multi-frame image data to obtain map metadata.

[0070] S403, The control device uses the identifier of the image sensor to mark the source of the map metadata, and obtains the marked map metadata.

[0071] S404. If the marked map metadata meets the mapping quality requirements, the control device generates marked map data based on the marked map metadata.

[0072] S401 to S404 have been described in detail in the above embodiments and will not be repeated here.

[0073] S405. The control device acquires the moving direction of the mobile platform and real-time data collected by the image sensor.

[0074] After generating marked map data, the control device can control the movable platform to move within the spatial scene at the current moment based on the map data generated at the previous moment. It can also control the movable platform to move within the spatial scene again upon re-entering it, based on the generated map data.

[0075] As the mobile platform moves within the spatial scene, the direction of movement is collected by a driving sensor, and real-time data of the spatial scene is collected by an image sensor. The control equipment then controls the mobile platform's movement within the spatial scene in real time based on its direction of movement, real-time data, and map data.

[0076] S406. The control device obtains target data matching the direction of movement from the marked map data based on the marking information.

[0077] The map data's marker information reflects its data source, identifying the image sensor that acquired the map data. Since the image sensor's installation location on the mobile platform is fixed, the marker information can be used to determine the sensor's position. This, combined with the mobile platform's direction of movement and the aforementioned position information, allows for the selection of target data from the map data.

[0078] More specifically, if the movable vehicle is moving forward, the target data is obtained from the map data sourced from the image sensor installed in front of the movable platform. If the movable vehicle is moving backward, the target data is obtained from the map data sourced from the image sensor installed behind the movable platform.

[0079] S407. The control equipment generates control commands based on real-time data and target data.

[0080] The real-time data collected by the image sensor is also image data. The control device performs feature extraction and other processing on the image data, and determines the location information of the mobile platform based on the processed real-time data and map data. Then, based on the location information and map data, it generates control commands to control the movement of the mobile platform, so that the mobile platform can move within the spatial scene under the control of the control commands.

[0081] When determining the location information of a mobile platform based on processed real-time data and map data, the processed real-time data and map data are matched to obtain a matching result, and the location information of the mobile platform is determined based on the successfully matched map data.

[0082] In another embodiment, the control device matches real-time data with target data to obtain a matching result, and then sets a reliable value for the target data based on the matching result.

[0083] More specifically, if the matching result is successful, the reliability value of the target data is set to the first reliability value; if the matching result is unsuccessful, the reliability value of the target data is set to the second reliability value. The first reliability value is greater than the second reliability value.

[0084] After obtaining the reliable value of the target data, the control equipment statistically analyzes the reliable value of the target data to obtain the reliability statistics. When the reliability statistics meet the low reliability condition, the target data is deleted to optimize the map data.

[0085] When statistically analyzing the reliability values ​​of target data, if the reliability statistical result is the mean of the reliability values, the low reliability condition is that the mean of the reliability values ​​is less than the preset mean.

[0086] In the above technical solution, after the control device generates map data with source markers, it selects target data from the map data based on the marker information. Then, it controls the movement of the mobile platform based on the target data and real-time data collected by the image sensor. By filtering to obtain the target data, the amount of data processing during map data usage is reduced, allowing the control device to generate control commands more quickly and enabling the mobile device to move reliably within the spatial scene. Furthermore, when the control device matches the real-time data and the target data, it marks the target data with the matching result, thus optimizing the map data.

[0087] The following uses a mobile platform, specifically a smart car, as an example to illustrate the data processing method provided in this application. The execution subject of this method is the control device within the smart car, such as the vehicle's computer. The method specifically includes the following steps:

[0088] S501. When the mobile platform moves within the spatial scene, the control device controls the image sensor located on the mobile platform to collect multiple frames of image data of the spatial scene.

[0089] The intelligent vehicle is equipped with monocular cameras, such as dashcams, and fisheye cameras installed around the perimeter of the vehicle. These cameras are used to collect image data of a specific spatial scene, such as an underground parking lot. The intelligent vehicle also has driving sensors, such as low-precision inertial navigation units, odometers, and GPS.

[0090] The data processing method provided in this application does not require the addition of new sensors to intelligent vehicles; the aforementioned sensors can be used to generate map data and control the driving of intelligent vehicles.

[0091] S502, The control device processes multi-frame image data to obtain map metadata.

[0092] The control device, after acquiring multiple frames of image data from the aforementioned camera and driving data from the aforementioned driving sensors, processes the multiple frames of image data to obtain map metadata. Specifically, this includes the following steps:

[0093] S5001. Calculate the inter-frame pose between the two frames of image data using the above driving data.

[0094] VIO and VO algorithms can be used to process image data and data acquired by driving sensors to estimate the inter-frame pose between two image frames. Inter-frame pose serves as the foundation for image data processing. Alternatively, driving sensors can be used to estimate inter-frame pose; for example, integrating data acquired by the odometer and inertial measurement unit (IMU) can yield the inter-frame pose.

[0095] S5002, Extract two-dimensional feature data from multi-frame image data.

[0096] This process involves feature extraction from image data acquired by a monocular camera and an image data acquired by a fisheye camera. Preferably, geometric features are extracted from the image data, such as object edges, corners, planes, salient points, and special textures. Texture data, gradient data, and pixel color data are also extracted from the image data. These features are time-stable, angle-stable, and scale-stable, ensuring stable and consistent observation across different angles, distances, and time periods.

[0097] When extracting features from image data to obtain two-dimensional feature data, the extracted texture data, gradient data, pixel color, etc. are also used to encode the two-dimensional feature data in order to perform feature matching and build a map dictionary.

[0098] During feature extraction, the image data acquired by the fisheye camera needs to be corrected. For example, the image data under the fisheye camera model can be converted into the image data under the pinhole camera model. The two-dimensional feature data in the pinhole image data and the two-dimensional feature data in the converted fisheye image data can be fused, which is beneficial for feature matching.

[0099] S5003, Perform feature matching on inter-frame pose and two-dimensional feature data.

[0100] The main focus is on temporal correlation of the extracted features. Common temporal correlation methods include inter-frame correlation, window correlation, and loop closure correlation algorithms.

[0101] Inter-frame association mainly involves two adjacent images, such as images acquired at a time interval of 50 milliseconds, or images displayed at a distance of 20 centimeters. Inter-frame matching usually involves a lot of feature association.

[0102] Window association mainly refers to associating all features within a time period or distance. By statistically analyzing the number of associated features, we can measure performance indicators such as feature stability and consistency.

[0103] For example, when a two-dimensional feature data can be associated with a large number of images within a window, such as 30 frames of images, the two-dimensional feature data is high-quality two-dimensional feature data and has good robustness to temporal and spatial variations.

[0104] S5004. Calculate the three-dimensional feature points based on the feature matching results.

[0105] The calculation of 3D feature points is based on the assumption of three points being coplanar. It involves observing the same object from two different locations and calculating the object's 3D coordinates. In this application, the triangulation of image feature data is achieved using the inter-frame pose and matching results of two frames of image data to obtain 3D feature points.

[0106] S5005. Extract semantic information from multiple frames of image data.

[0107] Semantic information extraction primarily involves extracting information about objects with clearly defined categories within a spatial scene, such as ground lane lines, parking spaces, and directional arrows, as well as overhead safety barriers, hanging signs, large walls, and pillars. Semantic information is typically a relatively stable element, usually only becoming ineffective when the environment changes, such as during parking lot repairs or reconstruction, thus accurately reflecting the spatial scene.

[0108] S503. Use the image sensor identifier to mark the source of the map metadata and obtain the marked map metadata.

[0109] To achieve one-time map data construction, images are collected using cameras located at different angles, such as monocular and fisheye cameras. During image data collection, the images from different cameras are labeled based on their characteristics. When using the map data, different map metadata is selected and matched based on the intelligent vehicle's driving direction.

[0110] S504. If the marked map metadata meets the mapping quality requirements, the control device generates marked map data based on the marked map metadata.

[0111] Specifically, the richness of spatial 3D feature points, semantic information, and texture data is used to determine whether the obtained map metadata meets the mapping quality requirements. If it does not meet the requirements, the control equipment issues a warning message to prompt the selection of a more suitable spatial scene and a more appropriate time period for map construction.

[0112] If the obtained map metadata meets the mapping quality requirements, labeled map data is generated. More specifically, the labeled 3D features and texture data are visualized to obtain a labeled feature data layer. The labeled semantic information is also visualized to obtain a labeled semantic information layer.

[0113] After generating feature data layers and semantic information layers, corresponding layers can be generated based on personalized needs to support autonomous driving in smart cars. For example, building a keyframe dictionary.

[0114] The following example illustrates the process of generating map data for a parking lot: When an intelligent vehicle travels along the path in the map, a fisheye camera and a forward-facing camera collect image data within the parking lot. The image data collected by the fisheye camera is stitched together to form a panoramic top-down view. Then, the panoramic top-down view and the image collected by the monocular camera are stitched together to form a ground image. Finally, deep learning methods are used to identify and extract parking spaces, mainly including lane lines, parking spaces, and ground directional arrows, which serve as an important component of the semantic map.

[0115] By recognizing parking spaces, information such as valid parking spaces, invalid parking spaces, exclusive parking spaces, and parking space numbers can be identified and stored on the map for customers to select during automatic parking.

[0116] The detection results of parking spaces may contain noise and false detections. It is necessary to perform fusion filtering on the location, type, size and other information of parking spaces observed multiple times to obtain a semantic layer. Then, combined with the feature data layer of the parking lot, the final map data of the parking lot is formed.

[0117] After generating map data, it can be optimized. This includes optimizing the positions of 3D points, camera poses, and the quality of the semantic map.

[0118] S505, The control device acquires the moving direction of the mobile platform and real-time data collected by the image sensor.

[0119] The control equipment acquires the driving direction of the intelligent vehicle and obtains image data from fisheye cameras and monocular cameras.

[0120] S506. Obtain target data matching the direction of movement from the marked map data according to the marking information.

[0121] Specifically, when the intelligent vehicle is traveling forward, it uses a monocular camera to collect image data and corresponding map data for positioning. When the intelligent vehicle is traveling backward, it uses a fisheye camera to collect image data and corresponding map data for positioning, so that the intelligent vehicle has a faster and more powerful positioning capability.

[0122] S507. The control device generates control commands based on the real-time data and the target data.

[0123] The control device matches real-time data with map data to obtain a matching result, determines the location information of the intelligent vehicle based on the successfully matched map data, and then controls the intelligent vehicle to drive based on the location information of the intelligent vehicle.

[0124] In the above technical solution, after the control device generates map data with source markers, target data is selected from the map data according to the marker information, and the vehicle driving is controlled according to the target data and real-time data collected by the image sensor. This reduces the amount of data processing during the use of map data, so as to achieve rapid positioning of the intelligent vehicle and thus more reliably control the driving of the intelligent vehicle.

[0125] like Figure 5 As shown, another embodiment of this application provides a data processing method. The execution subject of this method is a control device in a smart car, such as a vehicle computer. The method specifically includes the following steps:

[0126] S601. Obtain inter-frame pose based on multi-frame data collected by sensors on the mobile platform.

[0127] The mobile platform is a smart car, with a dashcam installed at the front, using a pinhole camera as an image sensor. Fisheye cameras are installed around the smart car, also serving as image sensors. Both the pinhole and fisheye cameras are used to collect image data of the spatial scene.

[0128] Intelligent vehicles are also equipped with odometers, inertial measurement units (IMUs), and GPS. The IMU, GPS, and odometer are used to collect driving data of intelligent vehicles, such as acceleration, speed, mileage, and driving position.

[0129] After acquiring multiple frames of image data using a pinhole camera and acquiring driving data of the intelligent vehicle using an odometer and an IMU, the inter-frame pose is estimated based on the multiple frames of image data, the odometer data, and the driving data acquired by the IMU.

[0130] When estimating inter-frame pose, the visual inertial system (VINS) algorithm, the visual inertial odometer (VIO) algorithm, or the visual odometer (VO) algorithm can be used. Alternatively, the inter-frame pose can be estimated based on the integration of wheel speed and IMU output driving data.

[0131] Estimating inter-frame pose is a crucial system foundation for the entire data processing method, and its quality directly affects the computation time of subsequent map data optimization steps.

[0132] S602. Extract two-dimensional feature data from multi-frame image data.

[0133] This step primarily involves extracting features from the image data acquired by the pinhole camera and the image data acquired by the fisheye camera. This feature data is two-dimensional.

[0134] Two-dimensional feature data includes geometric features such as object edges, corners, planes, salient points, and special textures. These features are time-stable, angle-stable, and scale-stable, remaining relatively stable across different angles, distances, and time periods while maintaining consistent observability. Furthermore, effective representation of these two-dimensional feature data involves encoding them using texture data, gradient data, and pixel color data. The encoded two-dimensional feature data is then used for feature matching and dictionary data.

[0135] During feature extraction, the image data acquired by the fisheye camera is corrected to transform the image data under the fisheye camera model into the pinhole camera model. After the transformation, the two-dimensional feature data of the image data acquired by the pinhole camera and the two-dimensional feature data of the image data acquired by the fisheye camera can be fused to achieve feature matching with higher stability and consistency.

[0136] S603. Extract semantic information from the image data acquired by the pinhole camera.

[0137] The semantic information processing section primarily involves classifying and identifying objects with clear meanings in the spatial scene image data, such as ground objects like lane lines, parking spaces, and directional arrows, and spatial objects like crash barriers, hanging signs, large walls, and pillars, to obtain semantic information. Semantic information is typically a relatively stable element, and it usually only becomes unusable during large-scale environmental changes, such as parking lot repairs or reconstruction.

[0138] S604. Stitch together the images captured by the fisheye camera to form a panoramic top view.

[0139] The specific stitching method uses existing technology and will not be elaborated here. After obtaining the panoramic top view, semantic information is extracted from the panoramic top view, mainly including semantic information such as lane lines, parking spaces, and ground directional arrows, which can be used as an important component of the semantic information layer.

[0140] By identifying parking spaces in the panoramic view, information such as valid parking spaces, invalid parking spaces, exclusive parking spaces, and parking space numbers can be identified and stored on the map for customers to select during automatic parking.

[0141] S605. Match the two-dimensional feature data extracted in S602 based on the inter-frame pose estimated in S601.

[0142] In the two-dimensional model, feature data matching mainly involves temporal correlation of the extracted two-dimensional feature data. Algorithms such as inter-frame correlation, window correlation, and loop closure correlation can be used for matching.

[0143] Inter-frame association mainly involves associating two adjacent images, for example, with an interval of 50 milliseconds or 20 centimeters. Inter-frame matching usually involves a lot of feature association.

[0144] Window association primarily involves associating all features within a fixed or non-fixed time or distance range. Performance metrics such as stability and consistency of a feature are obtained by statistically analyzing the number of associated features. For example, if a two-dimensional feature can be associated with a large number of images within a window, such as 30 frames, it indicates that the two-dimensional feature data is of very high quality and has good robustness to temporal and spatial variations.

[0145] Loop closure matching refers to the phenomenon where data may be collected multiple times for the same spatial scene. In this case, by associating two-dimensional feature data with image data from different time periods, it is possible to identify whether a map has been built at the current location, and also to effectively fuse and update image data of a spatial scene observed multiple times.

[0146] S606. Based on S602, extract two-dimensional feature data to construct a keyframe dictionary in the map data.

[0147] Constructing a keyframe dictionary involves clustering the features of image data of a spatial scene and using a combination of multiple two-dimensional feature data to represent the current scene. The keyframe dictionary can be constructed using not only two-dimensional feature data but also semantic information, deep learning descriptors, and more.

[0148] The functions of the keyframe dictionary include: First, expressing the richness of a scene. A rich keyframe dictionary in the map data of a spatial scene indicates that the map data possesses abundant texture data, geometric features, and semantic information. Conversely, a low keyframe dictionary richness indicates poor map data quality, making it unsuitable for user control of intelligent vehicles, such as automatic parking, thus hindering user expectation management. Second, the keyframe dictionary can be used for location identification during parking relocation. During relocation initialization, the vehicle needs to find its current position on the map. By matching keyframe dictionary data, the approximate position of the vehicle on the map can be quickly found. Then, by matching semantic information and two-dimensional feature data, accurate position estimation can be achieved.

[0149] S607. Extract parking space information from the image data collected by the panoramic top view and pinhole camera.

[0150] The panoramic top-down view, stitched together from images captured by a fisheye camera, and the pinhole camera image data together provide a complete image of the ground at a distance. Parking space identification and extraction are achieved using deep learning methods, which will not be elaborated upon here.

[0151] The parking space recognition results may contain noise and misidentification issues. It is necessary to perform fusion filtering on the multiple collected image data to obtain information such as the location, type, and size of the parking spaces, so as to obtain multiple continuous, stable, and high-quality parking space layers.

[0152] S608. Obtain three-dimensional feature points based on feature matching results.

[0153] Among them, obtaining three-dimensional feature points based on feature matching results is to realize the triangulation of feature data, that is, to convert two-dimensional feature data into three-dimensional feature data.

[0154] The triangulation step is mainly based on the assumption that three points are coplanar. It uses images of the same object observed from two different positions to calculate the three-dimensional coordinates of the object. Traditional SFM technology requires simultaneous estimation of the pose between cameras and the three-dimensional coordinates of the point cloud. In this application, after estimating the inter-frame pose, the two-dimensional feature data is matched based on the inter-frame pose, and then the feature data is quickly triangulated based on the matching result.

[0155] S609. Determine whether the map metadata meets the mapping quality requirements.

[0156] Map metadata refers to the 3D feature points, semantic information, and texture data obtained in the above steps. By comprehensively evaluating information such as the number of 3D feature points, the spatial richness of these feature points, the richness of semantic features, the richness of scene textures, and ambient illumination, the quality of the map data constructed based on the currently obtained map metadata is determined to meet the mapping quality requirements. Specific mapping quality requirements have been detailed in the above embodiments and will not be repeated here.

[0157] If the mapping quality does not meet the requirements, the control equipment will issue a warning message to alert the user to change the spatial scene, better weather, and a better time to build the map.

[0158] S610. Generate map data based on map metadata and optimize the map data.

[0159] Specifically, if the map data constructed based on map metadata meets the requirements for automatic parking control while satisfying the mapping quality requirements, then the map data will undergo unified quality optimization. Optimization includes the position of 3D feature points, camera pose, and the quality of semantic layers. For example, optimizations will be made to the position, angle, size, and category of semantic elements.

[0160] S611. Store map data according to different layers.

[0161] Once the map data is constructed, it is stored in different layers, such as feature data layers, semantic information layers, navigation data layers, and keyframe dictionary layers. These different layers can be used to provide data for the automatic parking positioning function.

[0162] Furthermore, to achieve one-time map data construction, map metadata for image data collected by different cameras is encoded, and the map metadata from image data collected by different cameras is used for matching at different positioning stages. For example, when the vehicle is moving forward, map metadata from image data collected by the forward-looking camera is used for matching and positioning, while when the vehicle is reversing, map metadata from image data collected by the rear-looking fisheye camera is used for matching and positioning, achieving a faster and more powerful map building capability.

[0163] When using map data for automatic parking, the control device increases the weight of matching map metadata and decreases the weight of unmatched map metadata. It also adds a certain amount of high-quality map metadata to the map to achieve timely updates of map metadata and ensure the long-term timeliness and high quality of the map.

[0164] The above technical solution constructs lightweight, purely visual map data, avoiding the complex and manually-intervention-required construction of high-precision maps. The constructed map data has extremely low hardware cost requirements and can quickly achieve high-quality map data construction. Extensive testing has shown that the constructed map data can be effectively used for positioning problems in the automatic parking process. Through extensive learning, it enables functions such as dedicated parking space selection, parking in any selected space, and parking lot cruising. It effectively assists human-computer interaction and provides map navigation and real-time route planning for functions such as valet parking and memory parking.

[0165] like Figure 6 As shown, this application provides a control device 700, which specifically includes: a memory 701 for storing instructions and a processor 702 for executing the instructions stored in the memory. The processor 702 is specifically used to execute:

[0166] When the mobile platform moves within the spatial scene, the image sensor located on the mobile platform is controlled to acquire multi-frame image data of the spatial scene;

[0167] Map metadata is obtained by processing multiple frames of image data; the map metadata includes any one or more combinations of three-dimensional feature points, texture data, and semantic information.

[0168] Determine whether the map metadata meets the mapping quality requirements;

[0169] If the map metadata meets the mapping quality requirements, map data is generated based on the map metadata; the map data is used to control the movement of the mobile platform within the spatial scene.

[0170] Optionally, the mapping quality requirements include at least one of the following:

[0171] The total number of three-dimensional feature points has reached the first threshold.

[0172] There exist at least two three-dimensional feature points whose components on all three coordinate axes are different;

[0173] The total amount of texture data has reached the second threshold.

[0174] The number of texture data types has reached the third threshold.

[0175] The total amount of semantic information has reached the fourth threshold.

[0176] The number of semantic information categories has reached the fifth threshold.

[0177] Optionally, processor 702 is used to specifically perform:

[0178] The map metadata is source-tagged using the identifiers of the image sensor to obtain the tagged map metadata;

[0179] Map data is generated based on map metadata, specifically including:

[0180] Generate labeled map data based on the labeled map metadata;

[0181] The marked map data includes labeling information to indicate the data source.

[0182] Optionally, the processor 702 is used to specifically perform at least one of the following:

[0183] Based on the labeled 3D feature points and labeled texture data, a labeled feature data layer is generated;

[0184] Generate a labeled semantic information layer based on the labeled semantic information.

[0185] Optionally, processor 702 is used to specifically perform:

[0186] Acquire the movement direction of the mobile platform and real-time data collected by the image sensor;

[0187] Target data matching the direction of movement is obtained from the marked map data based on the marking information;

[0188] Control commands are generated based on real-time data and target data; these commands are used to control the movement of the mobile platform.

[0189] Optionally, processor 702 is used to specifically perform:

[0190] If the direction of movement is forward, the target data is obtained from the map data sourced from the image sensor installed in front of the movable platform.

[0191] If the direction of movement is backward, the target data is obtained from the map data sourced from the image sensor installed behind the mobile platform.

[0192] Optionally, processor 702 is used to specifically perform:

[0193] Matching real-time data with target data yields matching results;

[0194] Obtain the location information of the mobile platform based on the successfully matched target data;

[0195] Control commands are generated based on the location information of the mobile platform.

[0196] Optionally, processor 702 is used to specifically perform:

[0197] Set the reliability value of the target data based on the matching results.

[0198] Optionally, processor 702 is used to specifically perform:

[0199] If the matching result is successful, set the reliability value of the target data to the first reliability value;

[0200] If the matching result is a failure, set the reliability value of the target data to the second reliability value;

[0201] The first reliability value is greater than the second reliability value.

[0202] Optionally, processor 702 is used to specifically perform:

[0203] Statistically analyze the reliability values ​​of the target data to obtain reliability statistics.

[0204] Target data is deleted when reliability statistics meet the low reliability criteria.

[0205] Optionally, the image sensor includes a monocular camera and / or a fisheye camera.

[0206] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0207] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A data processing method, characterized by, include: When the mobile platform moves within the spatial scene, the image sensor located on the mobile platform is controlled to acquire multiple frames of image data of the spatial scene; The multi-frame image data is processed to obtain map metadata; The map metadata includes three-dimensional feature points, texture data, and semantic information; the three-dimensional feature points are obtained based on the triangulation of two-dimensional feature points in the multi-frame image data. The semantic information is used to characterize the object category under the texture data and the three-dimensional feature points; Determine whether the quantity and type of the three-dimensional feature points, texture data, and semantic information meet the mapping quality requirements; If the mapping quality requirements are met, map data is generated based on the map metadata; wherein, the map data is used to control the movement of the mobile platform within the spatial scene.

2. The method of claim 1, wherein, The mapping quality requirements include: The total number of the three-dimensional feature points reaches a first quantity threshold; There exist at least two three-dimensional feature points whose components on all three coordinate axes are different; The total amount of texture data reaches the second quantity threshold; The number of texture data types reaches the third threshold. The total amount of semantic information reaches the fourth quantity threshold; The number of semantic information types reaches the fifth threshold.

3. The method according to claim 1 or 2, characterized in that, After processing the multi-frame image data to obtain map metadata, the method further includes: The map metadata is source-marked using the identifier of the image sensor to obtain the marked map metadata; Generating map data based on the map metadata specifically includes: Generate labeled map data based on the labeled map metadata; The marked map data includes marker information to indicate the source of the data.

4. The method according to claim 3, characterized in that, The map data is generated based on the marked map metadata, including at least one of the following: Based on the labeled 3D feature points and labeled texture data, a labeled feature data layer is generated; Generate a labeled semantic information layer based on the labeled semantic information.

5. The method according to claim 3 or 4, characterized in that, After generating labeled map data based on the labeled map metadata, the method further includes: The moving direction of the mobile platform and the real-time data collected by the image sensor are obtained; Target data matching the direction of movement is obtained from the marked map data based on the marking information; Control commands are generated based on the real-time data and the target data; wherein the control commands are used to control the movement of the mobile platform.

6. The method according to claim 5, characterized in that, Based on the marking information, target data matching the direction of movement is obtained from the marked map data, specifically including: If the direction of movement is forward, map data from the image sensor installed in front of the mobile platform is obtained from the map data as the target data; If the direction of movement is backward, the map data obtained from the image sensor installed behind the mobile platform is used as the target data.

7. The method according to claim 5, characterized in that, Control commands are generated based on the real-time data and the target data, specifically including: The real-time data and the target data are matched to obtain a matching result; The location information of the mobile platform is obtained based on the successfully matched target data; The control commands are generated based on the location information of the mobile platform.

8. The method according to claim 7, characterized in that, After matching the real-time data and the target data to obtain a matching result, the method further includes: Set the reliability value of the target data based on the matching result.

9. The method according to claim 8, characterized in that, Setting a reliable value for the target data based on the matching result specifically includes: If the matching result is a successful match, the reliability value of the target data is set to the first reliability value; If the matching result is a failure, the reliability value of the target data is set to the second reliability value; Wherein, the first reliability value is greater than the second reliability value.

10. The method according to claim 8, characterized in that, After setting the reliability value of the target data based on the matching result, the method further includes: Statistically analyze the reliability values ​​of the target data to obtain reliability statistics. When the reliability statistics result meets the low reliability condition, the target data is deleted.

11. The method according to any one of claims 1 to 10, characterized in that, The image sensor includes a monocular camera and / or a fisheye camera.

12. A control device, characterized in that, include: A memory for storing instructions and a processor for executing the instructions stored in the memory, the processor being used to specifically execute: When the mobile platform moves within the spatial scene, the image sensor located on the mobile platform is controlled to acquire multiple frames of image data of the spatial scene; The multi-frame image data is processed to obtain map metadata; The map metadata includes three-dimensional feature points, texture data, and semantic information; the three-dimensional feature points are obtained based on the triangulation of two-dimensional feature points in the multi-frame image data. The semantic information is used to characterize the object category under the texture data and the three-dimensional feature points; Determine whether the quantity and type of the three-dimensional feature points, texture data, and semantic information meet the mapping quality requirements; If the mapping quality requirements are met, map data is generated based on the map metadata; wherein, the map data is used to control the movement of the mobile platform within the spatial scene.

13. The control device according to claim 12, characterized in that, The mapping quality requirements include: The total number of the three-dimensional feature points reaches a first quantity threshold; There exist at least two three-dimensional feature points whose components on all three coordinate axes are different; The total amount of texture data reaches the second quantity threshold; The number of texture data types reaches the third threshold. The total amount of semantic information reaches the fourth quantity threshold; The number of semantic information types reaches the fifth threshold.

14. The control device according to claim 12 or 13, characterized in that, The processor is used to specifically execute: The map metadata is source-marked using the identifier of the image sensor to obtain the marked map metadata; Generating map data based on the map metadata specifically includes: Generate labeled map data based on the labeled map metadata; The marked map data includes marker information to indicate the source of the data.

15. The control device according to claim 14, characterized in that, The processor is used to specifically perform at least one of the following: Based on the labeled 3D feature points and labeled texture data, a labeled feature data layer is generated; Generate a labeled semantic information layer based on the labeled semantic information.

16. The control device according to claim 14 or 15, characterized in that, The processor is used to specifically execute: The moving direction of the mobile platform and the real-time data collected by the image sensor are obtained; Target data matching the direction of movement is obtained from the marked map data based on the marking information; Control commands are generated based on the real-time data and the target data; wherein the control commands are used to control the movement of the mobile platform.

17. The control device according to claim 16, characterized in that, The processor is used to specifically execute: If the direction of movement is forward, map data from the image sensor installed in front of the mobile platform is obtained from the map data as the target data; If the direction of movement is backward, the map data obtained from the image sensor installed behind the mobile platform is used as the target data.

18. The control device according to claim 16, characterized in that, The processor is used to specifically execute: The real-time data and the target data are matched to obtain a matching result; The location information of the mobile platform is obtained based on the successfully matched target data; Control commands are generated based on the location information of the mobile platform.

19. The control device according to claim 18, characterized in that, The processor is used to specifically execute: Set the reliability value of the target data based on the matching result.

20. The control device according to claim 19, characterized in that, The processor is used to specifically execute: If the matching result is a successful match, the reliability value of the target data is set to the first reliability value; If the matching result is a failure, the reliability value of the target data is set to the second reliability value; Wherein, the first reliability value is greater than the second reliability value.

21. The control device according to claim 19, characterized in that, The processor is used to specifically execute: Statistically analyze the reliability values ​​of the target data to obtain reliability statistics. When the reliability statistics result meets the low reliability condition, the target data is deleted.

22. The control device according to any one of claims 12 to 21, characterized in that, The image sensor includes a monocular camera and / or a fisheye camera.

23. A mobile platform, characterized in that, It includes an image sensor and a control device as described in any one of claims 12 to 22.

Citation Information

Patent Citations

  • Method, device, equipment and system for map building

    CN107145578A

  • Map constructing method and system and storage medium

    CN109785731A