A data processing method and device, computer equipment and readable storage medium

By optimizing image matching and reprojection error functions, the problem of sensor data errors in autonomous vehicles was solved, and the optimization of camera extrinsic parameters and object 3D spatial coordinates was achieved, ensuring the accuracy of autonomous driving.

CN122289368APending Publication Date: 2026-06-26MOMENTA (SUZHOU) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MOMENTA (SUZHOU) TECHNOLOGY CO LTD
Filing Date
2024-12-25
Publication Date
2026-06-26

Smart Images

  • Figure CN122289368A_ABST
    Figure CN122289368A_ABST
Patent Text Reader

Abstract

This application provides a data processing method, apparatus, computer device, and readable storage medium, relating to the field of autonomous driving. The method includes: acquiring an image data packet of an autonomous vehicle's estimated initial trajectory and driving environment, wherein the image data packet includes multiple image data points acquired by the autonomous vehicle's camera at different positions along the estimated initial trajectory; performing image matching on the multiple image data points, determining the coordinates of different image feature points in 3D space based on the image matching results, obtaining multiple 3D space points; constructing a reprojection error function based on the camera's extrinsic parameters and the 3D space coordinates of the 3D space points; and optimizing the reprojection error function to obtain target camera extrinsic parameters and target 3D space coordinates that minimize the reprojection error function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving, and in particular to a data processing method, apparatus, computer device, and readable storage medium. Background Technology

[0002] Truth data on the driving environment is crucial for autonomous driving control, and it plays the following roles: (1) Training and validating models: Truth data is the foundation for training and validating autonomous driving system algorithms. With a large amount of truth data, more accurate and reliable autonomous driving models can be trained. (2) Evaluating system performance: Truth data can also be used to evaluate the performance of autonomous driving systems. By comparing with truth data, the accuracy and reliability of autonomous driving systems in recognizing and handling various scenarios can be quantified. (3) Improving safety and reliability: Accurate truth data helps improve the safety and reliability of autonomous driving systems. By continuously optimizing and adjusting the algorithm model, it can be made more adaptable to the complex scenarios and changes in the real world, thereby reducing the risk of accidents.

[0003] In existing technologies, true data about the driving environment is collected by various sensors on autonomous vehicles. However, due to sensor accuracy issues, the collected data may have errors compared to the actual data, leading to inaccurate autonomous driving control. Summary of the Invention

[0004] In view of this, this application provides a data processing method, apparatus, computer equipment, and readable storage medium, which solves the problem of inaccurate data in the related art.

[0005] In a first aspect, embodiments of this application provide a data processing method, including:

[0006] The autonomous vehicle acquires an image data packet of its estimated initial trajectory and driving environment. The image data packet includes multiple image data, which are images captured by the autonomous vehicle's camera at different positions along the estimated initial trajectory as the autonomous vehicle travels along it.

[0007] Image matching is performed on multiple image data sets, and the coordinates of different image feature points in the image data in 3D space are determined based on the image matching results to obtain multiple 3D space points;

[0008] Based on the camera extrinsic parameters of the camera and the 3D spatial coordinates of the 3D spatial points, a reprojection error function is constructed;

[0009] The reprojection error function is optimized to obtain the target camera extrinsic parameters and target 3D spatial coordinates that minimize the reprojection error function.

[0010] Secondly, embodiments of this application provide a data processing apparatus, including:

[0011] An acquisition module is used to acquire an image data packet of the estimated initial trajectory of the autonomous vehicle and the driving environment. The image data packet includes multiple image data, which are images captured by the autonomous vehicle's camera at different positions on the estimated initial trajectory when the autonomous vehicle is traveling along the estimated initial trajectory.

[0012] The spatial point determination module is used to perform image matching on multiple image data, and determine the coordinates of different image feature points in the image data in 3D space based on the image matching results, thereby obtaining multiple 3D spatial points;

[0013] The optimization module is used to construct a reprojection error function based on the camera extrinsic parameters of the camera and the 3D spatial coordinates of the 3D spatial points, and to optimize the reprojection error function to obtain the target camera extrinsic parameters and target 3D spatial coordinates that minimize the reprojection error function.

[0014] Thirdly, embodiments of this application provide a computer device including a processor and a memory, the memory storing programs or instructions that can run on the processor, the programs or instructions implementing the steps of the method as described in the first aspect when executed by the processor.

[0015] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method as described in the first aspect.

[0016] In this embodiment, image data packets of the estimated initial trajectory of the autonomous vehicle and the driving environment are acquired. Image matching is performed on image data collected by the autonomous vehicle at different positions on the same estimated initial trajectory to obtain image matching results. Then, based on the image matching results, the coordinates of each feature point in the image in 3D space are identified to obtain the corresponding 3D spatial points. A reprojection error function is constructed, and by solving the reprojection error function, the camera extrinsic parameters and the 3D spatial coordinates of the 3D spatial points are optimized.

[0017] In this embodiment, the collected image data is optimized to obtain optimized camera extrinsic parameters and the 3D spatial coordinates of the object. This optimizes both the camera extrinsic parameters and the 3D spatial coordinates of the object in the driving environment. This ensures that the position of the same object captured by the vehicle at different driving positions on the same trajectory is consistent in 3D space, independent of the vehicle's current position or viewpoint. This allows for the determination of the object's unique position, reduces data errors, and ensures the accuracy of autonomous driving control.

[0018] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0020] Figure 1 One of the flowcharts of the data processing method according to an embodiment of this application is shown;

[0021] Figure 2 A second schematic flowchart of a data processing method according to an embodiment of this application is shown;

[0022] Figure 3 A schematic diagram illustrating an autonomous vehicle traveling along a predicted initial trajectory according to an embodiment of this application is shown;

[0023] Figure 4 A structural block diagram of a data processing apparatus according to an embodiment of this application is shown;

[0024] Figure 5 A structural block diagram of a computer device according to an embodiment of this application is shown. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0027] The data processing method, apparatus, computer equipment, and readable storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0028] This application provides a data processing method that optimizes the collected image data to obtain optimized camera extrinsic parameters and the 3D spatial coordinates of objects. This optimizes both the camera extrinsic parameters and the 3D spatial coordinates of objects in the driving environment, i.e., it achieves GPO (General Pose Optimization). This ensures that the position of the same object captured by the vehicle at different driving positions is consistent in 3D space, independent of the vehicle's current position or viewing angle, and can determine the unique position of the object.

[0029] like Figure 1 and Figure 2 As shown in the embodiment of this application, a data processing method includes:

[0030] Step 101: Obtain image data packets of the estimated initial trajectory of the autonomous vehicle and the driving environment. The image data packets include multiple image data, which are images captured by the autonomous vehicle's camera at different positions on the estimated initial trajectory when the autonomous vehicle is driving along the estimated initial trajectory.

[0031] In this step, the estimated initial trajectory of the autonomous vehicle, also known as the egopose, and the image data packet of the driving environment corresponding to the estimated initial trajectory are obtained. This image data packet can be called single packet data, which refers to the images collected by the camera at different positions of the estimated initial trajectory when the autonomous vehicle is driving along a predicted initial trajectory.

[0032] Autonomous vehicles are equipped with multiple cameras that can capture images of the driving environment from different angles while the vehicle is in motion. For example, an autonomous vehicle might have six cameras: one at the front, one at the rear, two on the left, and two on the right. These multiple cameras allow for the collection of image data from various angles. Figure 3 As shown, at any position of the predicted initial trajectory, multiple cameras are used to acquire images from different angles.

[0033] Step 102: Perform image matching on multiple image data, and determine the coordinates of different image feature points in the image data in 3D space based on the image matching results to obtain multiple 3D space points.

[0034] In this step, image matching is performed on image data collected by the autonomous vehicle at different positions along the same estimated initial trajectory to obtain image matching results. Then, based on the image matching results, the coordinates of each feature point in the image in 3D space are identified to obtain the corresponding 3D space points.

[0035] In one embodiment of this application, image matching is performed on multiple image data, and the coordinates of different image feature points in the image data in 3D space are determined based on the image matching results to obtain multiple 3D space points, including:

[0036] Image matching is performed on multiple image data to obtain multiple image matching pairs;

[0037] For each image matching pair, feature point matching is performed to obtain the feature point matching pair corresponding to the image matching pair;

[0038] Based on feature point matching pairs, the coordinates of different image feature points in the image data in 3D space are obtained, resulting in multiple 3D space points.

[0039] In this embodiment, image matching pairs are established based on image data collected by the autonomous vehicle at different positions in the same estimated initial trajectory. Each image matching pair includes two image data collected at different positions. These two image data can be collected by the same camera or by different cameras.

[0040] After establishing image matching pairs, an image matching algorithm is used to perform feature point matching on each image matching pair to obtain the corresponding feature point matching pairs, that is, to establish image feature point matching relationships. The specific process of the image matching algorithm may include: extracting feature points from two images, including corner points, edges, etc. of objects in the images, matching feature points in different images to obtain image feature point matching pairs.

[0041] In one embodiment, the SuperPoint algorithm and the LightGlue algorithm can be used for feature extraction and image matching. The SuperPoint algorithm is a feature point detector and descriptor based on a deep neural network, capable of stably extracting key points in various complex environments and generating robust feature descriptors. These feature descriptors are used for accurate matching between images, which is crucial for tasks such as image stitching and image registration. The SuperPoint algorithm is trained through self-supervised learning, enabling it to accurately detect key feature points in various images. The LightGlue algorithm is a flexible image registration library focused on optimizing image geometric transformations and stitching. It can calculate the optimal image fusion strategy based on provided feature point information (such as feature points generated by SuperPoint), thereby achieving smooth transitions and eliminating discontinuities in overlapping areas. LightGlue is more efficient in terms of memory and computation while maintaining high accuracy.

[0042] By combining SuperPoint with LightGlue, a powerful feature extraction and image matching solution is formed. SuperPoint is responsible for extracting key feature points in the image and generating descriptors, while LightGlue uses this feature point information to optimize the matching and registration between images. It has the following advantages: (1) Stable feature points, with relatively consistent feature points extracted in different seasons and times, which is conducive to multiple matching runs; (2) Strong matching ability under different lighting scenarios, which is conducive to multiple matching runs; (3) Free from the dependence on semantic features, it obtains features from the original image, making it more versatile.

[0043] Furthermore, based on feature point matching pairs, the coordinates of different image feature points in the image data in 3D space are obtained, resulting in multiple 3D spatial points. Specifically, triangulation is performed based on the estimated initial trajectory and image matching results. That is, by using image feature point matching pairs in images taken from different perspectives, combined with the estimated initial trajectory (i.e., vehicle posture information), the positions of these feature points in three-dimensional space are calculated, resulting in 3D spatial points.

[0044] The embodiments of this application can achieve accurate feature extraction and image matching to obtain precise 3D spatial points.

[0045] In one embodiment of this application, image matching is performed on multiple image data to obtain multiple image matching pairs, including: taking two image data acquired by two cameras that are within a first preset distance and have a co-view relationship as image matching pairs.

[0046] In this embodiment, image matching pairs are established based on image data collected by the autonomous vehicle at different positions in the same estimated initial trajectory. The image matching pair consists of two image data collected by cameras at different positions but within a first preset distance and having a common viewing relationship. Here, having a common viewing relationship means that the orientation of the camera is within 150 degrees and the first preset distance can be within 50m.

[0047] By using the above methods, effective image matching pairs can be selected, thereby improving the accuracy of the 3D spatial coordinates of objects.

[0048] In one embodiment of this application, image matching is performed on multiple image data to obtain multiple image matching pairs, including:

[0049] According to the second preset distance of the estimated initial trajectory, perform image frame extraction processing on multiple image data once;

[0050] Image matching is performed on the image data after image frame extraction to obtain multiple image matching pairs.

[0051] In this embodiment, image matching is a relatively time-consuming module. For example, if a single matching takes 200ms, and an image data packet contains 30s and 2100 images, brute-force matching would result in 2.2 million matching pairs, which would be extremely time-consuming. To address this, image frame extraction is used to improve timeliness. Image frame extraction is performed on multiple image data sets at every second preset distance from the estimated initial trajectory. For example, image frame extraction is performed every 5 meters. Then, image matching is performed on the image data after frame extraction to obtain multiple image matching pairs.

[0052] In this embodiment of the application, image frame extraction reduces the amount of data, saves matching time, and improves data processing efficiency.

[0053] Step 103: Construct a reprojection error function based on the camera's extrinsic parameters and the 3D spatial coordinates of the 3D spatial points.

[0054] In this step, a reprojection error function is constructed. Subsequently, by solving the reprojection error function, the camera extrinsic parameters and the 3D spatial coordinates of 3D spatial points are optimized.

[0055] In one embodiment of this application, the reprojection error function is constructed as follows:

[0056] (1) Calculate the position of a 3D point in the world coordinate system relative to the camera coordinate system:

[0057] P rig= R rw ·P 3D +t rw

[0058] P cam= R cr ·P rig +t cr

[0059] Among them, P 3D P represents the position of a point in 3D space in the world coordinate system. rig For point P in 3D space 3D At the position in the reference coordinate system, P cam For point P in 3D space 3D Position R in the camera coordinate system cr t represents the rotation matrix of the reference coordinate system relative to the camera coordinate system. cr R represents the translation vector of the reference coordinate system relative to the camera coordinate system. rw The rotation matrix t represents the rotation of the world coordinate system relative to the reference coordinate system. rw This represents the translation vector of the world coordinate system relative to the reference coordinate system.

[0060] (2) P cam Projecting onto a two-dimensional image plane, we obtain the theoretical projection point (u', v'):

[0061]

[0062] Where (X',Y',Z') is P cam In the coordinate components of the x-axis, y-axis, and z-axis, f x The focal length, expressed in pixels, represents the camera's magnification capability in the x-direction. y The focal length, expressed in pixels, represents the y-axis magnification capability of a camera lens in the y-direction. Since most cameras use equidistant lenses, f... x and f y They are often very close or identical. x This represents the x-coordinate of the principal point, which is the position of the center point on the horizontal axis of the image in the pixel coordinate system. y This represents the y-coordinate of the principal point, which is the position of the center point on the vertical axis of the image in the pixel coordinate system.

[0063] (3) Calculate the feature points (x) obs ,y obs The residual between the projection point (u', v') and the projection point (u', v'):

[0064] r x =u'-x obs

[0065] ry =v'-y obs

[0066] Therefore, the reprojection error function, or loss function L1, can be expressed as:

[0067]

[0068] L1 represents the Euclidean distance between the feature point and the predicted projection point on the two-dimensional image plane. The goal of the optimization process is to adjust the parameter R. cr t cr R rw t rw P 3D Thus minimizing L1,r x r represents the residual in the x-direction between the feature point and the projection point. y This represents the residual in the y-direction between the feature point and the projection point.

[0069] If the obtained 3D spatial points are used directly to construct reprojection constraints, it will be subject to the prior estimation of the initial trajectory. Because if there are any deviations in the estimated initial trajectory or extrinsic parameters, the 3D spatial points cannot be generated, resulting in unsatisfactory optimization results. Therefore, all 3D spatial points are determined based on the image matching results. After initially filtering out mismatches based on the reprojection error, the remaining 3D spatial points will participate in the reprojection optimization.

[0070] Since some 3D spatial points with large reprojection errors exhibit mismatches, which can interfere with the overall optimization, these 3D spatial points are assigned lower confidence levels, while those with smaller reprojection errors are assigned higher confidence levels. Introducing confidence levels as a weighting factor into the reprojection error function reduces the impact of 3D spatial points with large reprojection errors during optimization. Therefore, in one embodiment of this application, before constructing the reprojection error function based on camera extrinsic parameters and the 3D spatial coordinates of the 3D spatial points, the method further includes:

[0071] Calculate the initial value of the reprojection error for each 3D spatial point, and determine the confidence level of the 3D spatial point based on the initial value of the reprojection error. The confidence level is inversely proportional to the initial value of the reprojection error.

[0072] Based on the camera's extrinsic parameters and the 3D spatial coordinates of 3D spatial points, a reprojection error function is constructed, including:

[0073] A reprojection error function is constructed based on confidence level, camera extrinsic parameters, and 3D spatial coordinates of 3D spatial points.

[0074] In this embodiment, an initial value for the reprojection error is first calculated, and the confidence level of the 3D spatial points is set based on this initial value. A larger initial value results in a smaller confidence level. The magnitude of the initial reprojection error is obtained by comparing the projected points of the 3D spatial points on the two-dimensional image plane with the actual 2D pixels observed on the two-dimensional image plane under the initial parameter values. Figure 2 As shown, 3D spatial points with an initial reprojection error value less than a preset threshold can be set to high confidence, while 3D spatial points with an initial reprojection error value greater than or equal to the preset threshold can be set to low confidence.

[0075] Then, based on the confidence level, the camera's extrinsic parameters, and the 3D spatial coordinates of the 3D spatial points, a reprojection error function is constructed. The reprojection error function is as follows:

[0076]

[0077] Where L1 represents the reprojection error function, R cr t represents the rotation matrix of the reference coordinate system relative to the camera coordinate system. cr R represents the translation vector of the reference coordinate system relative to the camera coordinate system. rw The rotation matrix t represents the rotation of the world coordinate system relative to the reference coordinate system. rw P represents the translation vector of the world coordinate system relative to the reference coordinate system. 3D p represents the 3D spatial coordinates of a point in the world coordinate system. 2D θ represents the 2D spatial coordinates of a 2D pixel on a 2D image plane. c This represents the camera's intrinsic parameters, where N represents the number of 3D spatial points, and ω... i r represents the confidence level of the i-th 3D spatial point. x,i r represents the residual in the x-direction between the projection of the i-th 3D spatial point onto the 2D image plane and the actual observed 2D pixel on the 2D image plane. y,i This represents the residual in the y-direction between the projection of the i-th 3D spatial point onto the 2D image plane and the actual observed 2D pixel on the 2D image plane.

[0078] By introducing such a weighting mechanism, we can better handle mismatches and outlier issues, thereby improving the robustness of the system.

[0079] Step 104: Optimize the reprojection error function to obtain the target camera extrinsic parameters and target 3D spatial coordinates that minimize the reprojection error function.

[0080] In this embodiment, the optimization process seeks camera extrinsic parameters and 3D spatial coordinates that minimize the reprojection error function.

[0081] Optimization algorithms, such as gradient descent and Newton's method, can be used to iteratively update parameter values. In each iteration, the reprojection error function is recalculated and updated. Through multiple iterations, the value of the reprojection error function can be gradually reduced until a certain convergence condition is met, such as the error change being less than a certain threshold or the number of iterations reaching an upper limit.

[0082] In this embodiment, the collected image data is optimized to obtain optimized camera extrinsic parameters and the 3D spatial coordinates of the object. This optimizes both the camera extrinsic parameters and the 3D spatial coordinates of the object in the driving environment. This ensures that the position of the same object captured by the vehicle at different driving positions on the same trajectory is consistent in 3D space, independent of the vehicle's current position or viewpoint. This allows for the determination of the object's unique position, reduces data errors, and ensures the accuracy of autonomous driving control.

[0083] In one embodiment of this application, the method further includes:

[0084] Acquire the absolute position prior data and relative position prior data of the autonomous vehicle, as well as the extrinsic parameter prior data between multiple cameras of the autonomous vehicle;

[0085] Prior residual functions are established for absolute position prior data, relative position prior data, and extrinsic parameter prior data, respectively;

[0086] Each prior residual function is optimized to obtain prior data of the target's absolute position, target's relative position, and target's extrinsic parameters that minimize each prior residual function.

[0087] In this embodiment, prior residuals are obtained and optimized. These prior residuals include absolute position prior data for the global egopose, relative position prior data for the local egopose, and residuals of extrinsic parameter prior data between multiple cameras. The absolute position prior data for the global egopose refers to the absolute position information of the autonomous vehicle in the global coordinate system, which typically originates from GPS, map matching, or other external positioning systems. Introducing this prior information can significantly improve the accuracy and robustness of the positioning system. The absolute position prior of the global egopose can be used as part of the optimization objective to reduce accumulated errors and improve the accuracy of map construction.

[0088] Local egopose relative position prior data refers to the position information of an autonomous vehicle relative to a local reference point or a previous moment. This information typically comes from the vehicle's odometer, wheel speed sensors, or other internal positioning systems. In the absence of global positioning information, local egopose relative position priors can help the system maintain a relatively accurate position estimate. By continuously updating the relative position information, the system can track the vehicle's trajectory and avoid drift. In visual odometry or inertial navigation systems, local egopose relative position priors can be used to constrain and optimize the vehicle's attitude estimation.

[0089] Extrinsic prior data between multiple cameras refers to the relative position and pose information between these cameras. This information typically originates from the camera calibration process or a predefined camera configuration. The relative position and pose information between multiple cameras is crucial for achieving accurate image matching, 3D reconstruction, and depth estimation. Introducing this prior information can significantly improve system performance and accuracy. In multi-camera systems, extrinsic priors can be used to constrain and optimize camera pose estimation and image matching processes. For example, in multi-camera systems for autonomous vehicles, introducing extrinsic prior information can enable more accurate vehicle detection and tracking.

[0090] For the absolute position prior data, the relative position prior data, and the extrinsic parameter prior data, respectively, a prior residual function is established. That is, a prior residual function for the absolute position prior data, the relative position prior data, and the extrinsic parameter prior data is established. Any of these a prior residual functions is as follows:

[0091]

[0092] Where L(q,t) represents the prior residual function, and S is a scaling matrix, equivalent to the inverse of the square root of the information matrix, i.e., the inverse of the covariance matrix, used to weight different errors. true Let q represent the prior rotation component of the current pose, and t represent the rotation component of the current pose. true Let t represent the prior position component of the current pose, and t represent the position component of the current pose.

[0093] Furthermore, each prior residual function is optimized to obtain target absolute position prior data that minimizes the prior residual function of the absolute position prior data, target relative position prior data that minimizes the prior residual function of the relative position prior data, and target extrinsic parameter prior data that minimizes the prior residual function of the extrinsic parameter prior data.

[0094] This application embodiment optimizes the prior data of absolute position, prior data of relative position, and prior data of extrinsic parameters.

[0095] As a specific implementation of the above data processing method, this application provides a data processing apparatus. For example... Figure 4 As shown, the data processing device 400 includes: an acquisition module 401, a spatial point determination module 402, and an optimization module 403.

[0096] The acquisition module 401 is used to acquire the estimated initial trajectory of the autonomous vehicle and the image data packet of the driving environment. The image data packet includes multiple image data, which are images acquired by the autonomous vehicle's camera at different positions on the estimated initial trajectory when the autonomous vehicle is driving along the estimated initial trajectory.

[0097] The spatial point determination module 402 is used to perform image matching on multiple image data, and determine the coordinates of different image feature points in the image data in 3D space based on the image matching results, thereby obtaining multiple 3D spatial points;

[0098] The optimization module 403 is used to construct a reprojection error function based on the camera extrinsic parameters and the 3D spatial coordinates of the 3D spatial points, and to optimize the reprojection error function to obtain the target camera extrinsic parameters and target 3D spatial coordinates that minimize the reprojection error function.

[0099] Furthermore, the device also includes: an initial value calculation module, used to: calculate the initial value of the reprojection error of each 3D spatial point, and determine the confidence level of the 3D spatial point based on the initial value of the reprojection error of the 3D spatial point, wherein the confidence level is inversely proportional to the initial value of the reprojection error;

[0100] The optimization module 403 is specifically used to construct a reprojection error function based on confidence level, camera extrinsic parameters, and 3D spatial coordinates of 3D spatial points.

[0101] Furthermore, the reprojection error function is:

[0102]

[0103] Where L1 represents the reprojection error function, R cr t represents the rotation matrix of the reference coordinate system relative to the camera coordinate system. cr R represents the translation vector of the reference coordinate system relative to the camera coordinate system. rw The rotation matrix t represents the rotation of the world coordinate system relative to the reference coordinate system. rw P represents the translation vector of the world coordinate system relative to the reference coordinate system. 3D p represents the 3D spatial coordinates of a point in 3D space.2D θ represents the 2D spatial coordinates of a 2D pixel on a 2D image plane. c This represents the camera's intrinsic parameters, where N represents the number of 3D spatial points, and ω... i r represents the confidence level of the i-th 3D spatial point. x,i r represents the residual in the x-direction between the projection of the i-th 3D spatial point onto the 2D image plane and the 2D pixel on the 2D image plane. y,i This represents the residual in the y-direction between the projection of the i-th 3D spatial point onto the 2D image plane and the 2D pixel point on the 2D image plane.

[0104] Furthermore, the spatial point determination module 402 is specifically used for:

[0105] Image matching is performed on multiple image data to obtain multiple image matching pairs;

[0106] For each image matching pair, feature point matching is performed to obtain the feature point matching pair corresponding to the image matching pair;

[0107] Based on feature point matching pairs, the coordinates of different image feature points in the image data in 3D space are obtained, resulting in multiple 3D space points.

[0108] Furthermore, the spatial point determination module 402 is specifically used to: take two image data collected by two cameras that are within a first preset distance and have a co-view relationship as an image matching pair.

[0109] Furthermore, the spatial point determination module 402 is specifically used for:

[0110] According to the second preset distance of the estimated initial trajectory, perform image frame extraction processing on multiple image data once;

[0111] Image matching is performed on the image data after image frame extraction to obtain multiple image matching pairs.

[0112] Furthermore, the optimization module 403 is also used for:

[0113] Acquire the absolute position prior data and relative position prior data of the autonomous vehicle, as well as the extrinsic parameter prior data between multiple cameras of the autonomous vehicle;

[0114] Prior residual functions are established for absolute position prior data, relative position prior data, and extrinsic parameter prior data, respectively;

[0115] Each prior residual function is optimized to obtain prior data of the target's absolute position, target's relative position, and target's extrinsic parameters that minimize each prior residual function.

[0116] The data processing device 400 in this embodiment can be a computer device or a component within the computer device, such as an integrated circuit or a chip. The computer device can be a terminal or other devices besides a terminal. For example, the computer device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle computer device, etc., and can also be a server, network attached storage (NAS), personal computer (PC), etc. This embodiment does not specifically limit the specific type of device.

[0117] The data processing device 400 provided in this embodiment can achieve... Figure 1 and Figure 2 The various processes implemented in the data processing method embodiment will not be described again here to avoid repetition.

[0118] This application also provides a computer device, such as... Figure 5 As shown, the computer device 500 includes a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. When the program or instructions are executed by the processor 501, they implement the various steps of the above-described data processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0119] The memory 502 can be used to store software programs and various data. The memory 502 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 502 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 502 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0120] Processor 501 may include one or more processing units; optionally, processor 501 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 501.

[0121] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described data processing method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0122] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0123] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A data processing method, characterized in that, include: The autonomous vehicle acquires an image data packet of its estimated initial trajectory and driving environment. The image data packet includes multiple image data, which are images captured by the autonomous vehicle's camera at different positions along the estimated initial trajectory as the autonomous vehicle travels along it. Image matching is performed on multiple image data sets, and the coordinates of different image feature points in the image data in 3D space are determined based on the image matching results to obtain multiple 3D space points; Based on the camera extrinsic parameters of the camera and the 3D spatial coordinates of the 3D spatial points, a reprojection error function is constructed; The reprojection error function is optimized to obtain the target camera extrinsic parameters and target 3D spatial coordinates that minimize the reprojection error function.

2. The method according to claim 1, characterized in that, Before constructing the reprojection error function based on the camera's extrinsic parameters and the 3D spatial coordinates of the 3D spatial points, the method further includes: Calculate the initial value of the reprojection error for each of the 3D spatial points, and determine the confidence level of the 3D spatial points based on the initial value of the reprojection error. The confidence level is inversely proportional to the initial value of the reprojection error. The reprojection error function is constructed based on the camera's extrinsic parameters and the 3D spatial coordinates of the 3D spatial points, including: Based on the confidence level, the camera extrinsic parameters, and the 3D spatial coordinates of the 3D spatial points, a reprojection error function is constructed.

3. The method according to claim 2, characterized in that, The reprojection error function is: Where L1 represents the reprojection error function, R cr t represents the rotation matrix of the reference coordinate system relative to the camera coordinate system. cr R represents the translation vector of the reference coordinate system relative to the camera coordinate system. rw The rotation matrix t represents the rotation of the world coordinate system relative to the reference coordinate system. rw P represents the translation vector of the world coordinate system relative to the reference coordinate system. 3D p represents the 3D spatial coordinates of the 3D spatial point. 2D θ represents the 2D spatial coordinates of a 2D pixel on a 2D image plane. c This represents the camera's intrinsic parameters, N represents the number of 3D spatial points, and ω represents the number of points in the 3D space. i r represents the confidence level of the i-th 3D spatial point. x,i r represents the residual in the x-direction between the projection of the i-th 3D spatial point onto the 2D image plane and the 2D pixel on the 2D image plane. y,i This represents the residual in the y-direction between the projection of the i-th 3D spatial point onto the 2D image plane and the 2D pixel point on the 2D image plane.

4. The method according to claim 1, characterized in that, The step involves performing image matching on multiple image data sets, determining the coordinates of different image feature points in the image data in 3D space based on the image matching results, and obtaining multiple 3D space points, including: Image matching is performed on multiple image data sets to obtain multiple image matching pairs; For each image matching pair, feature point matching is performed to obtain the feature point matching pair corresponding to the image matching pair; Based on the feature point matching pairs, the coordinates of different image feature points in the image data in 3D space are obtained, resulting in multiple 3D space points.

5. The method according to claim 4, characterized in that, The step of performing image matching on multiple image data to obtain multiple image matching pairs includes: The two image data collected by the two cameras that are within a first preset distance and have a co-view relationship are used as the image matching pair.

6. The method according to claim 4, characterized in that, The step of performing image matching on multiple image data to obtain multiple image matching pairs includes: According to each second preset distance of the estimated initial trajectory, perform image frame extraction processing on multiple image data once; Image matching is performed on the image data after image frame extraction to obtain multiple image matching pairs.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Acquire the absolute position prior data and relative position prior data of the autonomous vehicle, as well as the extrinsic parameter prior data between multiple cameras of the autonomous vehicle; A priori residual function is established for the absolute position prior data, the relative position prior data, and the extrinsic parameter prior data, respectively; The prior residual functions are optimized to obtain the target absolute position prior data, target relative position prior data, and target extrinsic parameter prior data that minimize the respective prior residual functions.

8. A data processing apparatus, characterized in that, include: An acquisition module is used to acquire an image data packet of the estimated initial trajectory of the autonomous vehicle and the driving environment. The image data packet includes multiple image data, which are images captured by the autonomous vehicle's camera at different positions on the estimated initial trajectory when the autonomous vehicle is traveling along the estimated initial trajectory. The spatial point determination module is used to perform image matching on multiple image data, and determine the coordinates of different image feature points in the image data in 3D space based on the image matching results, thereby obtaining multiple 3D spatial points; The optimization module is used to construct a reprojection error function based on the camera extrinsic parameters of the camera and the 3D spatial coordinates of the 3D spatial points, and to optimize the reprojection error function to obtain the target camera extrinsic parameters and target 3D spatial coordinates that minimize the reprojection error function.

9. A computer device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that run on the processor, the program or instructions being executed by the processor to implement the steps of the data processing method as described in any one of claims 1 to 7.

10. A readable storage medium having a program or instructions stored thereon, characterized in that, When the program or instructions are executed by the processor, they implement the steps of the data processing method as described in any one of claims 1 to 7.