Real-time visual positioning correction method and device, electronic equipment and storage medium
By characterizing the initial and real-time images collected by the camera in the traffic scene and generating a calibration matrix, the camera jitter problem in the traffic scene is solved, and real-time positioning is achieved through deep learning without adding additional equipment.
Patent Information
- Application Number
- CN202311456288.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-06
AI Technical Summary
In traffic scenarios, roadside cameras have jitter problems. The existing multimodal data fusion method is difficult to widely use in traditional scenarios due to the high equipment cost and limited detection range.
By feature registration of the initial image and real-time image collected by the image acquisition device, a calibration matrix is generated, and the target identified in the real-time image is calibrated based on the matrix, the pixel coordinates after calibration of the target are determined, and coordinate conversion is performed through the calibration information of the image acquisition device to determine the real-time positioning information of the target.
It realizes that without adding additional equipment, the target information is perceived in real time through deep learning, and the target is converted to world coordinates based on the calibration data obtained in the initialization stage and the image pixel information obtained in the target detection, real-time positioning is achieved, and the camera jitter problem is solved.
Smart Images

Figure CN119941839A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence / IT (Information Technology) applications, and in particular to a real-time visual positioning correction method, device, electronic device and storage medium. Background Art
[0002] In traffic scenarios, roadside cameras have jitter problems in positioning. To solve this problem, the existing solution is to use multimodal data fusion methods, that is, to use perception devices including cameras, lidar, millimeter-wave radar, etc. at the same time, take the advantages of each sensor, and fuse the perception results of each device through post-processing to obtain a more comprehensive perception result.
[0003] However, due to the high cost and limited detection range of lidar, and the expensive millimeter-wave radar devices and their detection limitations on stationary and small objects, such multi-sensor solutions are difficult to apply in some traditional scenarios. Summary of the invention
[0004] The present disclosure is proposed in view of the above problems. The present disclosure provides a real-time visual positioning correction method, device, electronic device and storage medium.
[0005] According to one aspect of the present disclosure, a real-time visual positioning correction method is provided, the method comprising: performing feature registration on initial eigenvalues of an initial image acquired by an image acquisition device and eigenvalues of a real-time image to generate a calibration matrix; calibrating current pixel coordinates of a target identified in the real-time image based on the calibration matrix to determine the calibrated pixel coordinates of the target; and performing coordinate conversion on the calibrated pixel coordinates based on calibration information of the image acquisition device to determine real-time positioning information of the target.
[0006] In addition, according to one aspect of the present disclosure, a real-time visual positioning correction method is provided, in which the initial eigenvalues of the initial image captured by the image acquisition device and the eigenvalues of the real-time image are feature aligned to generate a calibration matrix, including: mask setting and feature extraction of the initial image captured by the image acquisition device to generate initial eigenvalues; the initial eigenvalues include: initial key points and initial descriptors; frame sampling and acquisition processing is performed on the real-time video to obtain a real-time image; feature extraction is performed on the real-time image within the mask setting range to generate current eigenvalues; the current eigenvalues include: current key points and current descriptors; feature alignment is performed on the current eigenvalues and the initial eigenvalues to generate a calibration matrix.
[0007] In addition, according to an aspect of the present disclosure, a real-time visual positioning correction method is provided, wherein, within the mask setting range, based on the current descriptor and the initial descriptor, the current key point and the initial key point are similarly matched to generate a matching point pair, including: within the mask setting range, by comparing the difference between a certain pixel and other pixels in the first preset area, the current key point is determined, and the descriptor is determined based on the grayscale value of the pixel in the second preset area around the current key point; the descriptor is optimized to generate a current descriptor with higher discrimination and stability; the distance metric is calculated for the current descriptor and the initial descriptor, and similarity matching is performed based on the distance metric to obtain the current key point that matches the initial key point, and generate a matching point pair.
[0008] In addition, according to one aspect of the present disclosure, a real-time visual positioning correction method is provided, in which matching point pairs are screened based on a preset threshold to obtain reasonable matching point pairs, including: setting a threshold based on the maximum distance in the distance metric, the threshold being N times the maximum distance, wherein 0<N<1; screening matching point pairs based on the threshold, when the distance metric corresponding to the matching point pairs is less than the threshold, it is considered reasonable, and a reasonable matching point pair is obtained; when the distance metric corresponding to the matching point pairs is greater than or equal to the threshold, it is considered unreasonable, and the unreasonable point pairs are eliminated.
[0009] In addition, according to one aspect of the present disclosure, a real-time visual positioning correction method is provided, in which a calibration matrix is generated based on reasonable matching point pairs, including: generating a matrix equation group based on reasonable matching point pairs; solving the matrix equation group based on the least squares method to generate an optimal solution as the calibration matrix.
[0010] In addition, according to an aspect of the present disclosure, a real-time visual positioning correction method is provided, in which feature alignment is performed on the current eigenvalues and the initial eigenvalues to generate a calibration matrix, including: within the mask setting range, based on the current descriptor and the initial descriptor, similarity matching is performed on the current key points and the initial key points to generate matching point pairs; based on a preset threshold, the matching point pairs are screened to obtain reasonable matching point pairs; based on the reasonable matching point pairs, a calibration matrix is generated.
[0011] In addition, according to one aspect of the present disclosure, a real-time visual positioning correction method is provided, in which the current pixel coordinates of a target identified in a real-time image are calibrated based on a calibration matrix, and determining the calibrated pixel coordinates of the target includes: performing target recognition on the real-time image to obtain the current pixel coordinates of the target; and calibrating the current pixel coordinates based on the calibration matrix to generate the calibrated pixel coordinates of the target.
[0012] In addition, according to a real-time visual positioning correction method of one aspect of the present disclosure, the calibrated pixel coordinates are converted into coordinates based on the calibration information of the image acquisition device, and the real-time positioning information of the target is determined, which includes: acquiring the calibration information of the image acquisition device; the calibration information includes: intrinsic parameters and extrinsic parameters; based on the intrinsic parameters and the calibrated pixel coordinates, the camera coordinates of the target are determined; based on the extrinsic parameters and the camera coordinates, the world coordinates of the target are determined as the real-time positioning information.
[0013] In addition, according to a real-time visual positioning correction method according to one aspect of the present disclosure, the calibrated pixel coordinates are transformed based on the calibration information of the image acquisition device, and determining the real-time positioning information of the target also includes: obtaining the distortion coefficient; and normalizing the coordinates of the initial image based on the distortion coefficient.
[0014] According to another aspect of the present disclosure, a real-time visual positioning correction device is provided, and the device includes: an initial module, which is used to obtain calibration information of an image acquisition device and initial eigenvalues of an initial image acquired by the image acquisition device; a real-time module, which includes: a frame extraction unit, which is used to perform frame extraction and acquisition processing on the real-time video acquired by the image acquisition device to acquire a real-time image; an extraction unit, which is used to perform feature extraction on the real-time image within a mask setting range to generate a current eigenvalue; an identification unit, which is used to perform target identification on the real-time image to obtain the current pixel coordinates of the target; a calculation module, which is used to perform feature registration on the current eigenvalue and the initial eigenvalue to generate a calibration matrix; a calibration module, which is used to calibrate the current pixel coordinates based on the calibration matrix to generate the calibrated pixel coordinates of the target; and a conversion module, which is used to perform coordinate conversion on the calibrated pixel coordinates based on the calibration information to determine the real-time positioning information of the target.
[0015] In addition, according to a real-time visual positioning correction device according to one aspect of the present disclosure, the calculation module includes: a matching unit, which is used to perform similarity matching on the current key point and the initial key point within the mask setting range based on the current descriptor and the initial descriptor to generate matching point pairs; a screening unit, which is used to screen the matching point pairs based on a preset threshold to obtain reasonable matching point pairs; and a generation unit, which is used to generate a calibration matrix based on the reasonable matching point pairs.
[0016] According to another aspect of the present disclosure, an electronic device is provided, including: a memory for storing computer-readable instructions; and a processor for executing the computer-readable instructions so that the electronic device performs the real-time visual positioning correction method as described above.
[0017] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium is provided for storing computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, the processor executes the real-time visual positioning correction method as described above.
[0018] As will be described in detail below, according to the real-time visual positioning correction method of the embodiment of the present disclosure, only an image acquisition device is used without adding additional equipment (gyroscope, IMU (Inertial Measurement Unit, inertial sensor), etc.), and the target information is perceived in real time through deep learning. According to the calibration data obtained in the initialization phase and the image pixel information obtained by target detection, the target is converted to world coordinates to achieve real-time positioning.
[0019] It is to be understood that both the foregoing general description and the following detailed description are exemplary, and are intended to provide further explanation of the technology as claimed. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other purposes, features and advantages of the present disclosure will become more apparent by describing the embodiments of the present disclosure in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0021] Figure 1 is a diagram illustrating an application scenario of a real-time visual positioning correction method according to an embodiment of the present disclosure;
[0022] Figure 2 is a schematic diagram illustrating the preparation work of the image acquisition device in the initialization stage according to an embodiment of the present disclosure;
[0023] Figure 3 is a flowchart illustrating a real-time visual positioning correction method according to an embodiment of the present disclosure;
[0024] Figure 4 is a flow chart further illustrating a method for generating a real-time visual positioning calibration matrix according to an embodiment of the present disclosure;
[0025] Figure 5 is a schematic diagram further illustrating real-time visual positioning feature registration according to an embodiment of the present disclosure;
[0026] Figure 6 is a schematic diagram further illustrating real-time visual positioning target coordinate conversion according to an embodiment of the present disclosure;
[0027] Figure 7 is a diagram illustrating a real-time visual positioning correction device according to an embodiment of the present disclosure;
[0028] Figure 8 is a hardware block diagram illustrating an electronic device according to an embodiment of the present disclosure; and
[0029] Fig. 9 is a schematic diagram illustrating a computer-readable storage medium according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solution and advantages of the present disclosure more obvious, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described here.
[0031] First, refer to Figure 1 The application scenarios according to the embodiments of the present disclosure are summarized.
[0032] Figure 1 2 is a diagram illustrating an application scenario of the real-time visual positioning correction method according to an embodiment of the present disclosure. Figure 1 As shown, the application scenario may include at least: a server 101, an image acquisition device 102, and a scene captured by the image acquisition device 102. The scene captured by the image acquisition device 102 may include at least: a traffic scene 105 (e.g. Figure 1 traffic lights and zebra crossings), traffic participants 103 and 104.
[0033] Specifically, server 101 can be a server device for real-time processing and correction of visual positioning data. Considering the real-time requirements of the application scenario, this server needs to have high computing power and response speed in order to process a large amount of data in a timely manner and output the corrected positioning results.
[0034] The image acquisition device 102 may be a road test sensing device in the traffic scene 105, for example Figure 1 The road test camera, monocular camera, etc. shown in the figure.
[0035] It is easy to understand that Figure 1 101-105 are exemplary only and do not constitute any limitation on the above application scenarios, types and numbers of traffic participants, and models and numbers of servers.
[0036] As mentioned above, the image acquisition device 102 needs to perform two preparatory tasks in the initialization phase, that is, before formally performing image processing and computer vision tasks: pre-calibration and initial image feature value extraction. Figure 2 As shown, Figure 2 1 is a schematic diagram illustrating the preparation work of the image acquisition device in the initialization stage according to an embodiment of the present disclosure.
[0037] Preparation 1: Pre-calibration. The purpose of pre-calibration is to obtain accurate internal and external parameters of the camera to obtain the conversion relationship from the camera coordinate system to the world coordinate system (where the camera coordinate system can be understood as the coordinate system of the camera itself, which is an important reference system for describing the internal and external parameters of the camera). Specifically, the calibration data of pre-calibration can include at least:
[0038] Intrinsic Parameters: Also known as camera internal parameters, they describe the optical characteristics of the camera itself. Common intrinsic parameters include focal length, optical center position, etc. Among them, the focal length represents the focusing ability of the camera, and the optical center position represents the location of the origin of the camera coordinate system.
[0039] Extrinsic Parameters: Also known as camera external parameters, they describe the position and posture (also called pose) of the camera in the world coordinate system. Common extrinsic parameters include the position and rotation of the camera.
[0040] In one embodiment of the present disclosure, the internal parameter may be in the form of a 3*3 matrix:
[0041]
[0042] Where (f x ,f y ) is the focal length of the camera, (c x ,c y ) is the origin of the camera coordinate system, that is, the location of the optical center.
[0043] The external parameter can be a matrix composed of a rotation matrix R (3*3) and a translation vector t (3*1), which can be written as a 4*4 augmented matrix:
[0044]
[0045] Based on the above parameters, the conversion relationship from the camera coordinate system to the world coordinate system can be obtained. The intrinsic parameter acts on the conversion process from the camera coordinate system to the pixel coordinate system; the extrinsic parameter acts on the conversion process from the world coordinate system to the camera coordinate system, which is a rigid body transformation that reflects the relative position of the camera in the world coordinate system. This conversion relationship will be used in step S303 below.
[0046] Optionally, the calibration data of the pre-calibration as described above may also include a distortion coefficient. Distortion refers to image distortion caused by factors such as lens manufacturing and installation accuracy, that is, the lines at the edge of the image will be more curved than the image in the center of the image, which is caused by the inherent characteristics of the camera. The distortion coefficient is a parameter that describes the distortion.
[0047] Distortion can mainly include radial distortion and tangential distortion. Among them, radial distortion is caused by the nonlinear propagation of light at the camera lens, which will cause the scale of the image center and the surrounding area to be inconsistent. Tangential distortion is caused by the camera lens not being installed parallel to the image plane, which will cause the straight lines in the image to appear not to be completely straight. Therefore, it is necessary to dedistort the image through the distortion coefficient so that the straight lines in the image remain straight and the shape and size of the object are accurately measured.
[0048] In one embodiment of the present disclosure, the distortion coefficient may be a 5*1 vector, which may be used to convert the original image coordinates obtained from the lens into ideal normalized image coordinates. The dedistorted image may be closer to the actual situation, and the accuracy of target positioning may be better than directly using the original distorted image.
[0049] Preparation 2: Extraction of initial image feature values. The purpose of extracting initial image feature values is mainly to serve as a reference standard for subsequent calculation of the calibration matrix. Specifically, the initial image refers to the picture collected during the initialization phase for calibrating the initial external parameters of the image acquisition device 102. The initial image feature value extraction may at least include: extracting information such as initial key points and initial descriptors of the initial image. Among them, key points and descriptors are important concepts used to represent local features in an image. Specifically,
[0050] Key Points: represent locations with significant properties or specific structures in an image. Usually selected are locations with significant features such as texture changes, edges, corners, etc. Key points are designed to extract local features that are unique and robust so that they can be used in subsequent tasks such as image matching and target tracking.
[0051] Descriptors: Descriptors are descriptions of the local area around a key point. A descriptor is a vector or feature description that is used to characterize the image information around a key point, such as capturing the texture, color, gradient and other features around the key point. The construction of descriptors usually takes into account factors such as rotation invariance, scale invariance and illumination invariance to improve their robustness and reliability.
[0052] All of the above information will be stored in the memory of the server 101. In order to reduce the amount of information stored in the memory, facilitate the calculation of subsequent steps, and reduce the interference of matching noise, before performing this step (ie, extracting the initial image feature value), the original image can be masked.
[0053] Masking is a commonly used technique in computer vision and image processing. Its essence is the bit operation of the image. It can be used to select areas of interest or filter certain parts of the image (i.e., cut out a specified area on the image).
[0054] In one embodiment of the present disclosure, mask setting can be divided into two methods: manual setting and automatic setting.
[0055] Manual setting: You can manually select a specified area in the original image. The area shape can include polygons, circles, irregular shapes, etc. You can customize the focus area according to the actual situation;
[0056] Automatic setting: The default value of the mask setting. The system will take a rectangular area with a certain height (this value can be set as a configuration, the default value is 3 / 5 of the image height) and a certain aspect ratio (the default value is proportional to the image resolution) as the mask area according to the position of the image optical center.
[0057] At this point, the two preparations in the initialization phase have been completed. The following will introduce a method for real-time visual positioning correction in the process of performing image processing and computer vision tasks after the image acquisition device 102 is officially put into use. The specific real-time visual positioning correction method will refer to the following Figure 3-Figure 6 Describe in more detail.
[0058] Figure 3 FIG. 1 is a flowchart illustrating a real-time visual positioning correction method according to an embodiment of the present disclosure. Figure 3 As shown, the correction method according to the embodiment of the present disclosure includes the following steps.
[0059] In step S301, the initial eigenvalues of the initial image and the eigenvalues of the real-time image acquired by the image acquisition device are feature-aligned to generate a calibration matrix. As mentioned above, the initial key points and initial descriptors of the initial image have been acquired in the preparation 2 of the initialization stage. These initial eigenvalues can be used as a reference to calibrate the deviation of the eigenvalues of the real-time image, thereby calculating a calibration matrix that can be used for calibration. The method for generating the calibration matrix is as follows: Figure 4 shown.
[0060] Figure 4 is a flow chart further illustrating a method for generating a real-time visual positioning calibration matrix according to an embodiment of the present disclosure. Figure 4 As shown, the method for generating a calibration matrix may at least include the following steps.
[0061] In step S401, the real-time video is subjected to frame sampling and acquisition processing to obtain a real-time image. As described above, at this stage, the image acquisition device 102 has been officially put into use, and real-time video will be generated. Frame sampling and acquisition of the real-time video will obtain real-time image data, and multiple real-time images will construct a data set.
[0062] In step S402, feature extraction is performed on the real-time image within the mask setting range to generate the current feature value; the current feature value includes: the current key point and the current descriptor. As mentioned above, for the subsequent feature registration, it is necessary to identify the real-time image obtained in step S401, identify the target in the image and annotate it. Among them, the target in the traffic scene can be a traffic participant. Common traffic participants can be as follows Figure 1 As shown in 103 and 104, for example, motor vehicles, non-motor vehicles, pedestrians, etc., of course, it can also be ships, trains, etc., which are not limited here.
[0063] Since the initial key points and initial descriptors are extracted in the mask area in the preparation work 2 of the initialization stage, for the convenience of subsequent comparison, this step also needs to be performed in the mask area when extracting the current key points and current descriptors in the real-time image. Image recognition can first use a 2D target detection algorithm to train the data set constructed in step S401 to generate a visual target detection model, and then use the model for image recognition.
[0064] In one embodiment of the present disclosure, YOLOv6 2D target detection can be used to train the constructed data set to ensure the accuracy of the visual detection model. YOLOv6 is a 2D target detection model based on deep learning and is the latest version of the YOLO (You Only Look Once) series. Unlike traditional methods based on region extraction and classifiers, YOLOv6 uses a single neural network to directly predict object categories and bounding boxes in images, with the advantages of faster reasoning speed, higher accuracy, strong versatility and adaptability.
[0065] In step S403, feature registration is performed on the current eigenvalues and the initial eigenvalues to generate a calibration matrix. As described above, feature registration is performed on the initial key points and initial descriptors in the initial image obtained in preparation 2 and the current key points and current descriptors in the implementation image obtained in step S402, thereby obtaining a homography matrix between the two images as a calibration matrix.
[0066] Among them, the homography matrix is also called the projection mapping matrix. It represents the transformation relationship between some points on the common plane between two images and is independent of the scale (that is, aH and H have the same effect). Therefore, its degree of freedom is 8, and a unique solution can be obtained through at least four matching point pairs.
[0067] For detailed feature registration, please refer to Figure 5 To understand, Figure 5 It is a schematic diagram further illustrating real-time visual positioning feature registration according to an embodiment of the present disclosure. Figure 5It is the image comparison in real scenes. Figure 5 The picture on the left is the original image. Figure 5 The picture on the right side of the middle is a real-time image. The feature point examples are marked with dots, and the feature point pairing examples are marked with straight lines (with dots at both ends of the straight line), representing the pairing relationship.
[0068] Furthermore, considering the application scenarios' requirements for real-time performance and feature matching accuracy, the ORB feature extraction + BEBLID descriptor algorithm can be used for feature registration to generate the optimal homography matrix, which can ensure the matching accuracy of feature points while ensuring operational efficiency.
[0069] Among them, ORB (Oriented FAST and Rotated BRIEF) is a feature extraction and descriptor algorithm commonly used in computer vision, which is used to detect and describe key points in images. BEBLID (Boosted Efficient Binary Local Image Descriptor) is an improved descriptor used in the ORB algorithm, which can enhance the efficient local feature descriptor. The key to its effectiveness is to select a set of image features in a discriminative way. It has low computational requirements and strong real-time performance.
[0070] The ORB feature extraction algorithm is based on the combination of the FAST corner detector and the BRIEF binary descriptor. It first uses the FAST corner detector to find some key points in the image, and then calculates the corresponding BRIEF descriptors based on the surrounding pixels of these key points.
[0071] The BEBLID descriptor has made some improvements on the ORB descriptor. It improves the discriminative ability and robustness of the ORB descriptor by introducing the Boosting algorithm. The BEBLID descriptor converts the binary features of the ORB descriptor into feature vectors with continuous values to better represent the local structural information of the image. The generation of this continuous value feature is based on the Boosting mechanism, which improves the performance of the descriptor by learning the difference between the target and background contrast samples.
[0072] The combination of ORB feature extraction and BEBLID descriptor makes the algorithm have a good balance in computational efficiency and description ability. Compared with traditional SIFT and SURF algorithms, ORB+BEBLID has faster speed and better robustness, and has high applicability in real-time applications and embedded devices. The specific steps may include at least:
[0073] S1: Get feature point information in the mask area through ORB feature extraction algorithm, and get descriptor information through BEBLID;
[0074] S2: Use BFMatcher (Brute-Force Matcher) in the feature point matching stage. It is a feature matching algorithm provided in the OpenCV library and is suitable for ORB. BFMatcher performs matching by calculating the distance between two sets of feature descriptors. For each feature point in the query image, BFMatcher will find the feature point that best matches it in the target image. You can use Euclidean distance, Hamming distance, etc. as similarity metrics, combined with descriptor information to make a rough match between the two images;
[0075] S3: Since feature point mismatch may occur in the rough matching, it is necessary to further filter the matching points. The specific filtering method can be: remove the point pairs whose distance exceeds the given threshold, and the retained point pairs are reasonable matching point pairs. When matching feature points, the similarity between the descriptors of the two feature points is determined by comparing their distances. The threshold can be set with reference to the maximum distance, which is the maximum distance between the feature point descriptors in the S2 rough matching. If the distance between the descriptors of the two feature points is less than a certain multiple of the maximum distance, they are considered to be matched, that is, a reasonable matching point pair; otherwise, they are considered to be mismatched and should be eliminated.
[0076] In one embodiment of the present disclosure, the threshold may be set to 0.9*maximum distance, that is, 0.9 times the maximum distance.
[0077] S4: Use the model to find the optimal solution for the filtered reasonable matching point pairs. The model may include: minimum mean square error, RANSAC, etc., so as to obtain the optimal homography matrix between the two images.
[0078] As mentioned above, the homography matrix is a matrix that describes the transformation relationship between two images, which can be used for tasks such as image registration, stitching, and even 3D reconstruction. However, due to factors such as noise, distortion, and occlusion, the match between the two images is not perfect. Therefore, it is necessary to use methods such as minimum mean square error or RANSAC to solve the optimal homography matrix between the two images.
[0079] Among them, the RANSAC method can be understood as first randomly selecting a group of feature points from the reasonable matching point pairs obtained in S3, generating an initial homography matrix estimate, and then using this estimate to verify whether other feature points in the reasonable matching point pairs conform to the model. If not, the initial homography matrix estimate is optimized. This process is repeated until the optimal homography matrix is obtained.
[0080] In one embodiment of the present disclosure, the point p is represented as in the camera system A. A (u A ,v A,1) and the camera system B is represented as point p B (u B ,v B ,1), the relationship between the two camera systems can be obtained through the homography matrix:
[0081]
[0082] At this point, the calibration matrix is obtained, so the generation method is introduced. Then return Figure 3 Continue to introduce the application of the calibration matrix after obtaining the calibration matrix.
[0083] In step S302, the current pixel coordinates of the target identified in the real-time image are calibrated based on the calibration matrix to determine the calibrated pixel coordinates of the target.
[0084] As described above, when the real-time image is recognized in step S402, the current pixel coordinates of the target are obtained while the current feature values are obtained. The obtained calibration matrix is applied to the current pixel coordinates to calibrate them and obtain the calibrated current pixel coordinates.
[0085] In one embodiment of the present disclosure, the homography matrix has been obtained as Based on this homography matrix, the current pixel coordinate p C (u C ,v C ,1) Perform calibration to generate the calibrated current pixel coordinate p′ C =(u′ C ,v′ C ,1).
[0086] In step S303, coordinate conversion is performed on the calibrated pixel coordinates based on the calibration information of the image acquisition device to determine the real-time positioning information of the target. As described above, the calibrated current pixel coordinates obtained in step S302 are projected into the world coordinate system to obtain the corresponding position of the traffic participant in the visual detection range in the actual world coordinate system. The specific conversion steps may at least include:
[0087] Step 1: Based on the internal parameters obtained in the initialization stage preparation work 1 and the calibrated current pixel coordinates obtained in step S402, the camera coordinates of the traffic participant in the camera coordinate system are determined;
[0088] Step 2: Based on the external parameters obtained in the initialization phase preparation work 1 and the camera coordinates obtained in step 1, determine the world coordinates of the traffic participant in the world coordinate system.
[0089] Figure 6is a schematic diagram further illustrating the real-time visual positioning target coordinate conversion according to an embodiment of the present disclosure. Figure 6 The specific process is introduced with a specific embodiment. Figure 6 As shown in Figure 2, in computer vision and image processing, conversion between different coordinate systems is often required.
[0090] Pixel Coordinate System: The pixel coordinate system is the most basic coordinate system. It takes the upper left corner of the image as the origin, the horizontal and vertical resolution of the image as units, and uses integer values to represent the position of the pixel. For example, the pixel coordinates (u, v) of an image can represent the row and column where the pixel is located.
[0091] Image Coordinate System: The image coordinate system is a coordinate system with the center of the image as the origin and the width and height of the image as units. The horizontal axis of the image coordinate system extends to the right, and the vertical axis extends downward. The image coordinate system is usually the same as the pixel coordinate system, except that the coordinate values are floating point numbers.
[0092] Camera Coordinate System: The camera coordinate system is the coordinate system inside the camera, with the camera optical center as the origin and aligned with the camera's optical axis (usually a vertical line pointing to the image plane). The units of the camera coordinate system are usually millimeters or meters.
[0093] World Coordinate System: The world coordinate system is a coordinate system that describes the position of objects in the real world. It is defined by specifying reference points and scales. The world coordinate system can be two-dimensional or three-dimensional, depending on the requirements of the scene. The world coordinate system is usually represented by real numbers and is used to describe the real position of objects in the real world.
[0094] To convert pixel coordinate system, image coordinate system, camera coordinate system and world coordinate system, it is usually necessary to use the camera's intrinsic and extrinsic parameters and appropriate numerical calculation methods. The conversion process is as follows:
[0095] Conversion from pixel coordinate system to image coordinate system: Since the two are basically the same, no conversion is required. You only need to convert the integer value of the pixel coordinate to a floating-point value.
[0096] Conversion from image coordinate system to camera coordinate system: Through inverse projection transformation, the image coordinates (u, v) are mapped to the point (X C ,Y C ,Z C ).
[0097] Conversion from camera coordinate system to world coordinate system: Through the external parameter matrix, the camera coordinate (X C ,Y C ,Z C ) is mapped to a point in the world coordinate system (X W ,Y W ,Z W )
[0098] In one embodiment of the present disclosure, the conversion formula from the pixel coordinate system to the camera coordinate system is:
[0099]
[0100] The conversion formula from the camera coordinate system to the world coordinate system is:
[0101]
[0102] Furthermore, for real-time considerations, TensorRT can be used to accelerate the visual target detection model. At the same time, the visual target detection and the calculation of the image feature homography matrix are performed asynchronously, and different computing units are allocated accordingly. The respective calculation speeds are in the millisecond level, which can fully guarantee the real-time performance of the system when processing real-time video.
[0103] At this point, the real-time visual positioning correction method has been introduced, and the real-time visual positioning correction device will be introduced below.
[0104] Figure 7 2 is a diagram illustrating a real-time visual positioning correction device according to an embodiment of the present disclosure. Figure 7 As shown, the real-time visual positioning correction device 700 may at least include:
[0105] Initial module 701, used to obtain calibration information and initial eigenvalues of the image acquisition device;
[0106] The real-time module 702 includes:
[0107] The frame extraction unit 7021 is used to perform frame extraction and acquisition processing on the real-time video of the image acquisition device to obtain a real-time image;
[0108] An extraction unit 7022, configured to extract features of the real-time image within the mask setting range to generate a current feature value;
[0109] The recognition unit 7023 is used to perform target recognition on the real-time image and obtain the current pixel coordinates of the target;
[0110] The calculation module 703 is used to perform feature registration on the current eigenvalue and the initial eigenvalue to generate a calibration matrix, which includes:
[0111] A matching unit 7031 is used to perform similarity matching on the current key point and the initial key point within the mask setting range based on the current descriptor and the initial descriptor to generate a matching point pair;
[0112] A screening unit 7032 is used to screen the matching point pairs based on a preset threshold to obtain reasonable matching point pairs;
[0113] A generating unit 7033, configured to generate the calibration matrix based on the reasonable matching point pairs;
[0114] A calibration module 704, configured to calibrate the current pixel coordinates based on the calibration matrix to generate calibrated pixel coordinates of the target;
[0115] The recognition unit 7041 is used to perform target recognition on the real-time image and obtain the current pixel coordinates of the target;
[0116] The calibration unit 7042 is used to calibrate the current pixel coordinates based on the calibration matrix to generate calibrated pixel coordinates of the target.
[0117] The conversion module 705 is used to perform coordinate conversion on the calibrated pixel coordinates based on the calibration information to determine the real-time positioning information of the target.
[0118] The camera conversion unit 7051 is used to determine the camera coordinates of the target based on the intrinsic parameters in the calibration information and the calibrated pixel coordinates;
[0119] The world conversion unit 7052 is used to determine the world coordinates of the target as real-time positioning information based on the external parameters and camera coordinates in the calibration information.
[0120] Figure 8 1 is a hardware block diagram of an electronic device 800 according to an embodiment of the present disclosure. The electronic device according to an embodiment of the present disclosure includes at least a processor; and a memory for storing computer-readable instructions. When the computer-readable instructions are loaded and executed by the processor, the processor executes the real-time visual positioning correction method as described above.
[0121] Figure 8 The electronic device 800 shown specifically includes: a central processing unit (CPU) 801, a graphics processing unit (GPU) 802 and a main memory 803. These units are connected to each other via a bus 804. The central processing unit (CPU) 801 and / or the graphics processing unit (GPU) 802 can be used as the above-mentioned processor, and the memory 803 can be used as the above-mentioned memory for storing computer-readable instructions. In addition, the electronic device 800 may also include a communication unit 805, a storage unit 806, an output unit 807, an input unit 808 and an external device 809, which are also connected to the bus 804.
[0122] Fig. 9 is a schematic diagram illustrating a computer-readable storage medium according to an embodiment of the present disclosure. Fig. 9 As shown, a computer-readable storage medium 900 according to an embodiment of the present disclosure has computer-readable instructions 901 stored thereon. When the computer-readable instructions 901 are executed by a processor, the real-time visual positioning correction method according to the embodiment of the present disclosure described with reference to the above figures is executed. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, optical disk, magnetic disk, etc.
[0123] The real-time visual positioning correction method, device and electronic device according to the embodiments of the present disclosure are described above with reference to the accompanying drawings. According to the real-time visual positioning correction method of the embodiments of the present disclosure, only an image acquisition device is used without adding additional equipment (gyroscope, IMU, etc.). The target information is perceived in real time through deep learning, and the target is converted to world coordinates according to the calibration data obtained in the initialization stage and the image pixel information obtained by target detection, so as to achieve real-time positioning.
[0124] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this disclosure.
[0125] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limitations. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.
[0126] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," and the like are open words, referring to "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or," and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0127] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.
[0128] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0129] Various changes, substitutions, and modifications of the techniques described herein may be made without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of the present disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and actions described above. Currently existing or later to be developed processes, machines, manufactures, compositions of events, means, methods, or actions that perform substantially the same functions or achieve substantially the same results as the corresponding aspects described herein may be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or actions within their scope.
[0130] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0131] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A real-time visual positioning correction method, characterized in that: The method comprises: Performing feature registration on the initial eigenvalues of the initial image collected by the image acquisition device and the eigenvalues of the real-time image to generate a calibration matrix; Calibrate the current pixel coordinates of the target identified in the real-time image based on the calibration matrix to determine the calibrated pixel coordinates of the target; The calibrated pixel coordinates are transformed based on the calibration information of the image acquisition device to determine the real-time positioning information of the target.
2. The calibration method according to claim 1, characterized in that: The step of performing feature registration on the initial eigenvalues of the initial image and the eigenvalues of the real-time image collected by the image acquisition device to generate a calibration matrix includes: Performing mask setting and feature extraction on the initial image collected by the image acquisition device to generate the initial feature value; the initial feature value includes: initial key points and initial descriptors; Performing frame extraction and acquisition processing on the real-time video to obtain the real-time image; Extracting features of the real-time image within the mask setting range to generate the current feature value; the current feature value includes: a current key point and a current descriptor; Perform feature registration on the current eigenvalue and the initial eigenvalue to generate the calibration matrix.
3. The calibration method according to claim 2, characterized in that: The performing feature registration on the current eigenvalue and the initial eigenvalue to generate the calibration matrix includes: Within the mask setting range, based on the current descriptor and the initial descriptor, similarity matching is performed on the current key point and the initial key point to generate a matching point pair; Screening the matching point pairs based on a preset threshold to obtain reasonable matching point pairs; The calibration matrix is generated based on the reasonable matching point pairs.
4. The calibration method according to claim 3, characterized in that: The step of performing similarity matching on the current key point and the initial key point based on the current descriptor and the initial descriptor within the mask setting range to generate a matching point pair includes: Within the mask setting range, a current key point is determined by comparing a certain pixel with other pixels in a first preset area, and a descriptor is determined based on the grayscale values of pixels in a second preset area around the current key point; Optimizing the descriptor to generate a more discriminative and stable current descriptor; A distance metric is calculated for the current descriptor and the initial descriptor, and similarity matching is performed based on the distance metric to obtain the current key point that matches the initial key point, thereby generating the matching point pair.
5. The calibration method according to claim 3 or 4, characterized in that: The step of screening the matching point pairs based on a preset threshold to obtain reasonable matching point pairs includes: The threshold is set based on the maximum distance in the distance metric, and the threshold is N times the maximum distance, where 0<N<1; The matching point pairs are screened based on the threshold. When the distance metric corresponding to the matching point pairs is less than the threshold, they are considered reasonable and the reasonable matching point pairs are obtained. When the distance metric corresponding to the matching point pairs is greater than or equal to the threshold, they are considered unreasonable and the unreasonable point pairs are eliminated.
6. The calibration method according to claim 5, characterized in that: The step of generating the calibration matrix based on the reasonable matching point pairs includes: generating an initial calibration matrix based on a plurality of reasonable matching point pairs among the reasonable matching point pairs; The initial calibration matrix is optimized based on the other reasonable matching point pairs to generate an optimal solution as the calibration matrix.
7. The calibration method according to claim 1 or 2, characterized in that: The step of calibrating the current pixel coordinates of the target identified in the real-time image based on the calibration matrix to determine the calibrated pixel coordinates of the target includes: Performing target recognition on the real-time image to obtain current pixel coordinates of the target; The current pixel coordinates are calibrated based on the calibration matrix to generate calibrated pixel coordinates of the target.
8. The calibration method according to claim 1, wherein: The step of performing coordinate conversion on the calibrated pixel coordinates based on the calibration information of the image acquisition device to determine the real-time positioning information of the target includes: Acquire calibration information of the image acquisition device; the calibration information includes: internal parameters and external parameters; Determining the camera coordinates of the target based on the intrinsic parameters and the calibrated pixel coordinates; Based on the external parameters and the camera coordinates, the world coordinates of the target are determined as real-time positioning information.
9. The calibration method according to claim 1, characterized in that: The step of performing coordinate conversion on the calibrated pixel coordinates based on the calibration information of the image acquisition device to determine the real-time positioning information of the target further includes: Get the distortion coefficient; The coordinates of the initial image are normalized based on the distortion coefficient.
10. A real-time visual positioning correction device, characterized in that: The device comprises: An initial module, used to obtain calibration information of an image acquisition device and initial feature values of an initial image acquired by the image acquisition device; Real-time modules, including: A frame extraction unit, used to perform frame extraction and acquisition processing on the real-time video acquired by the image acquisition device to acquire a real-time image; An extraction unit is used to extract features of the real-time image within a mask setting range, Generate current eigenvalue; An identification unit, used to perform target identification on the real-time image and obtain the current pixel coordinates of the target; A calculation module, used for performing feature registration on the current eigenvalue and the initial eigenvalue to generate a calibration matrix; A calibration module, configured to calibrate the current pixel coordinates based on the calibration matrix to generate calibrated pixel coordinates of the target; A conversion module is used to perform coordinate conversion on the calibrated pixel coordinates based on the calibration information to determine the real-time positioning information of the target.
11. The calibration device according to claim 10, characterized in that: The computing module comprises: A matching unit, configured to perform similarity matching on the current key point and the initial key point within the mask setting range based on the current descriptor and the initial descriptor to generate a matching point pair; A screening unit, used to screen the matching point pairs based on a preset threshold to obtain reasonable matching point pairs; A generating unit is used to generate the calibration matrix based on the reasonable matching point pairs.
12. An electronic device, characterized in that: include: a memory for storing computer readable instructions; as well as A processor is used to run the computer-readable instructions so that the electronic device performs the real-time visual positioning correction method as described in any one of claims 1 to 9.
13. A non-transitory computer-readable storage medium for storing computer-readable instructions, characterized in that: When the computer-readable instructions are executed by a processor, the processor is caused to perform the real-time visual positioning correction method according to any one of claims 1 to 9.
Citation Information
Cited By
Printed circuit board (PCB) positioning and multifunctional testing integrated device based on visual identification
CN120510108A
Real-time shelf stockout detection method and device based on inspection robot and storage medium
CN120808299A