An image correction method, device and storage medium for a multi-view camera

By processing video data from multiple cameras and matching features, a correction scheme is generated, which solves the problems of misalignment in multi-camera image stitching and unstable correction accuracy, and achieves high-precision image correction and improved stability.

CN122391034APending Publication Date: 2026-07-14SHENZHEN ZHUOCHENG ELECTRONICS CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN ZHUOCHENG ELECTRONICS CO LTD
Filing Date
2026-04-09
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

When stitching images together, multi-camera systems can cause image misalignment and geometric distortion due to non-parallel optical axes, positional deviations, or time synchronization errors. Existing correction methods are complex, costly, or have unstable accuracy.

Method used

By acquiring video data from multiple cameras, frame extraction and feature matching are performed to generate a correction scheme. The temporal information of the video data is used to correct the time synchronization problem, and high-precision spatial correction is achieved through feature point matching and offset calculation.

Benefits of technology

It achieves high-precision image correction, eliminates time asynchrony and spatial deviation, improves the stability and robustness of correction, and solves the problems of misalignment in multi-camera image stitching and unstable correction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391034A_ABST
    Figure CN122391034A_ABST
Patent Text Reader

Abstract

The application discloses an image correction method and device for a multi-view camera and a storage medium, relates to the technical field of image processing, and solves the technical problems of the existing multi-view camera image splicing misplacement and unstable correction accuracy. The method comprises the following steps: acquiring video data of the multi-view camera; performing frame extraction based on each independent video data in the video data, and obtaining a plurality of image groups through feature matching; performing feature point extraction on each frame image in the image group to obtain a plurality of feature points; matching the plurality of feature points to obtain a plurality of groups of feature point pairs; generating a correction scheme of different frame images based on the plurality of feature point pairs; acquiring a plurality of correction schemes corresponding to the same independent video data, and correcting the camera corresponding to the independent video data based on the correction scheme; and improving the correction accuracy when the multi-view camera is spliced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing technology, specifically an image correction method, device, and storage medium for multi-view cameras. Background Technology

[0002] Currently, with the development of video surveillance and panoramic imaging technologies, multi-view cameras are being used more and more widely. Multi-view cameras typically consist of multiple independent cameras, and a wide-angle or panoramic image is obtained by stitching together the images captured by each camera.

[0003] However, in practical applications, due to manufacturing processes, installation errors, or mechanical displacement after prolonged use, problems such as non-parallel optical axes, positional deviations, or time synchronization discrepancies often exist between individual cameras. These deviations can lead to misalignment, ghosting, or geometric distortion in the stitched image, severely affecting image quality. Existing correction methods typically rely on high-precision hardware calibration equipment, which is complex and costly; or image feature-based correction methods require extremely high accuracy in feature point selection and are easily affected by factors such as changes in lighting and sparse scene textures, resulting in unstable correction accuracy. Therefore, there is an urgent need for a method that can efficiently and accurately correct images from multiple cameras. Summary of the Invention

[0004] This application provides an image correction method, apparatus, and storage medium for multi-view cameras, which solves the technical problems of image stitching misalignment and unstable correction accuracy in existing technologies.

[0005] To achieve the above objectives, this application adopts the following technical solution: Firstly, an image correction method for multi-view cameras is provided, comprising: Acquire video data from a multi-camera system, the video data including independent video data collected by each camera of the multi-camera system, the independent video data being the video data collected by a single camera in the multi-camera system; perform frame extraction and feature matching based on each independent video data to obtain several image groups, the image groups containing one frame image from each independent video data; Feature points are extracted from each frame image in the image group to obtain several feature points; several feature points are matched to obtain several sets of feature point pairs; each set of corresponding feature points includes two feature points, and the two feature points come from different frame images; a correction scheme for different frame images is generated based on several feature point pairs. Several correction schemes corresponding to the same independent video data are obtained, and the camera corresponding to the independent video data is corrected based on the correction schemes.

[0006] Based on the above technical solution, the image correction method, apparatus, and storage medium for multi-cameras provided in this application acquire video data from the multi-cameras, perform frame extraction and feature matching to obtain image groups, then extract feature points and match them to obtain feature point pairs, and finally generate a correction scheme to correct the cameras. This method utilizes the temporal information of the video data, eliminates the time asynchrony problem between different cameras through time deviation correction, and achieves high-precision spatial correction through feature point matching and offset calculation. Simultaneously, the final correction value is determined through statistical methods, improving the stability and robustness of the correction.

[0007] In conjunction with the first aspect above, in one possible implementation, the step of extracting frames from individual video data and performing feature matching to obtain several image groups includes: Select any independent video data as reference video data, obtain the first frame step length, extract frames from the reference video data based on the first frame step length to obtain several frame images, and record the frame images as reference images. Obtain the set time neighborhood radius, and construct the matching time period of the corresponding reference image with the frame time corresponding to each reference image as the center and the time neighborhood radius as the radius. Select any other independent video data as the deviation video data, and extract video segment data within the matching time period from the deviation video data; obtain the second frame step length, extract frames from the video segment data based on the second frame step length to obtain several frame images, and record the frame images as the matching images; A number of matching images corresponding to the reference image are integrated into a time difference determination group corresponding to the reference image; a time deviation between the reference video data and the deviation video data is generated based on the reference image and its corresponding time difference determination group; the time axis of the deviation video data is corrected based on the time deviation to obtain the corrected video data; Other independent video data are selected sequentially as the deviation video data, and the time axis of the deviation video data is corrected to obtain the corresponding corrected video data; Image groups are obtained by extracting frames based on reference video data and corrected video data; it is understood that each frame image in the image group needs to be preprocessed here, including distortion correction, etc.

[0008] In conjunction with the first aspect above, in one possible implementation, generating the time deviation between the reference video data and the deviation video data based on the reference image and its corresponding time difference determination group includes: Calculate the similarity between the reference image and each matching image in the time difference determination group; it is understood that the calculation of image similarity is a relatively existing technology, and will not be elaborated on here. It is worth noting that the matching image corresponding to the largest similarity value is recorded as the deviation image of the reference image, and the difference between the frame time corresponding to the reference image and the frame time corresponding to the deviation image is recorded as the time deviation corresponding to the reference image. The time deviations corresponding to each reference image are obtained sequentially; the average value of the time deviations corresponding to each reference image is used as the final time deviation; it is understandable that the mode can also be used as the final time deviation.

[0009] In conjunction with the first aspect above, in one possible implementation, the step of extracting image groups based on reference video data and corrected video data includes: The overlapping portions of the reference video data and each correction video data on the time axis are cropped, and frames are extracted from the overlapping portions to obtain several frame images; the frame images of the reference video data and each correction video data with the same frame time are grouped into an image group.

[0010] In conjunction with the first aspect above, in one possible implementation, the matching of several feature points to obtain several sets of feature point pairs includes: Select any frame image from the extracted image group as the reference image, and record the other frame images as images to be corrected; Acquire several feature points in the reference image, and acquire several feature points in any image to be corrected; match the feature points in the reference image and the image to be corrected to obtain the matching status between the corresponding feature points. Two feature points that are successfully matched are recorded as a pair of feature points. Each pair of feature points includes a feature point from the reference image and a feature point from the image to be corrected. Several sets of feature point pairs are generated sequentially between each image to be corrected and the reference image; it can be understood that there are generally several sets of feature point pairs between an image to be corrected and the reference image.

[0011] In conjunction with the first aspect mentioned above, in one possible implementation, feature points in the reference image and the image to be corrected are matched to obtain the matching state between corresponding feature points, including: Obtain the feature descriptor vectors of several feature points corresponding to the reference image, and the feature descriptor vectors of several feature points corresponding to the image to be corrected; Using any feature point in the reference image as the reference point, calculate the Euclidean distance between the feature descriptor vector corresponding to the reference point and the feature descriptor vector corresponding to any feature point in the image to be corrected. When the Euclidean distance is less than a set similarity distance threshold, set the matching status between the feature point and the feature point corresponding to the reference point as a preliminary successful match; otherwise, set the matching status between the feature point and the feature point corresponding to the reference point as a failed match. It is understood that the specific value of the similarity distance threshold is set by expert experience and is used to distinguish whether feature descriptors are similar. Traverse all feature points in the image to be corrected to obtain the matching status between the feature point corresponding to the reference point and each feature point in the image to be corrected; select the feature point with the smallest Euclidean distance from the feature points with the initial matching status, and record the matching status between the feature point and the reference point as a successful match; The feature points in the reference image are traversed to obtain the matching status between each feature point in the reference image and each feature point in the image to be corrected. It can be understood that whenever a reference point has a corresponding successfully matched feature point, the successfully matched feature point will be deleted when matching other reference points in the future. That is, when matching other reference points in the future, the previously successfully matched feature points will not be considered.

[0012] In conjunction with the first aspect above, in one possible implementation, the correction scheme for generating different frame images based on several feature point pairs includes: Obtain several sets of feature point pairs from the reference image and the image to be corrected, and select the feature point pair with the smallest Euclidean distance as the reference feature point pair; The coordinates of feature points belonging to the image to be corrected in the reference feature point pair are offset to obtain offset coordinate points. The offset coordinate points are located in the neighborhood corresponding to the set neighborhood radius of the feature points. The same offset is performed on the coordinates of feature points belonging to the image to be corrected in other feature point pairs to obtain the corresponding offset coordinate points. Obtain the set feature neighborhood radius, construct the feature neighborhood corresponding to the offset coordinate point with the feature neighborhood radius centered on the offset coordinate point, obtain the feature descriptor vector corresponding to the feature neighborhood, and mark it as the offset feature descriptor vector; Calculate the Euclidean distance between the offset feature descriptor vector and the feature descriptor vector corresponding to another feature point in its corresponding feature point pair, and denot it as the offset distance; calculate the average value of the offset distances corresponding to each feature point pair under the same offset, and select the offset with the smallest average value as the correction scheme; the offset includes a horizontal offset and a vertical offset, and the correction scheme includes a horizontal offset and a vertical offset.

[0013] In conjunction with the first aspect above, in one possible implementation, the camera corresponding to the independent video data is calibrated based on the calibration scheme, including: Extract the longitudinal and lateral offsets from several correction schemes corresponding to independent video data; Calculate the variance of each longitudinal offset. When the variance is less than a set variance threshold, obtain the mode of the longitudinal offset as the longitudinal correction value; otherwise, obtain the average value of the longitudinal offset as the longitudinal correction value. Calculate the variance of each lateral offset. When the variance is less than a set variance threshold, obtain the mode of the lateral offset as the lateral correction value; otherwise, obtain the average value of the lateral offset as the lateral correction value. The original longitudinal and lateral dimensions of the camera corresponding to the independent video data are corrected based on the longitudinal and lateral correction values. Specifically, a translation transformation matrix is ​​set according to the longitudinal and lateral correction values, and the image data obtained by the corresponding camera is translated based on the translation transformation matrix.

[0014] In a second aspect, an image correction device for a multi-view camera is provided, comprising: a data acquisition module, a video data processing module, an image data processing module, and a correction module; The data acquisition module is used to acquire video data from the multi-camera system, and the video data includes video data from each camera in the multi-camera system. The video data processing module is used to extract frames and match features to obtain several image groups based on each independent video data in the video data. The image data processing module includes a feature point matching unit and a correction scheme generation unit; The feature point matching unit is used to extract feature points from each frame image in the image group to obtain a number of feature points; to match the number of feature points to obtain a number of feature point pairs; a set of corresponding feature points includes two feature points, and the two feature points come from different frame images; The correction scheme generation unit is used to generate correction schemes for different frame images based on several feature point pairs; obtain several correction schemes corresponding to the same independent video data, and generate horizontal correction values ​​and vertical correction values ​​for the independent video data based on the correction schemes; The correction module corrects the original longitudinal and lateral dimensions of the camera corresponding to the independent video data based on the longitudinal and lateral correction values.

[0015] Thirdly, this application provides a computer-readable storage medium storing instructions that, when executed on an image correction apparatus for a multi-view camera, cause the image correction apparatus for a multi-view camera to perform the methods described in the first aspect and any possible implementation thereof.

[0016] Fourthly, this application provides a computer program product containing instructions that, when run on an image correction apparatus for a multi-view camera, cause the image correction apparatus for the multi-view camera to perform the methods described in the first aspect and any possible implementation thereof.

[0017] This application provides an image correction method for multi-camera systems. It acquires video data from multiple cameras, performs frame extraction and feature matching to obtain image groups, extracts feature points and matches them to obtain feature point pairs, and finally generates a correction scheme to correct the cameras. This method utilizes the temporal information of the video data, eliminating time asynchrony issues between different cameras through time deviation correction, and achieving high-precision spatial correction through feature point matching and offset calculation. Simultaneously, it determines the final correction value using statistical methods, effectively eliminating interference from abnormal data, improving the stability and robustness of the correction, and solving the problems of image stitching misalignment and unstable correction accuracy in existing multi-camera technologies.

[0018] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram illustrating the steps of the image correction method for multi-view cameras in this application; Figure 2 This is a schematic diagram of the module connections for the image correction system used in the multi-view camera in this application. Detailed Implementation

[0021] The technical solutions of this application will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0022] Please see Figure 1 The first aspect of this application provides an image correction method for a multi-view camera, comprising: Acquire video data from the multi-camera system. This video data includes individual video data captured by each camera within the multi-camera system, with each individual camera capturing video data from a single camera within the system. Specifically, a multi-camera system typically contains two or more independent camera modules used to capture images from different perspectives to stitch together a panoramic image. The video data refers to the continuous frame sequence data captured by these independent camera modules. In practical applications, the video data can be a real-time video stream or historical video files stored on a storage medium. This video data may have initial deviations in the temporal and spatial dimensions, such as inconsistent startup time differences or slight offsets in the optical axis angle. This step acquires this raw data through data interfaces such as USB, network interface, and HDMI to provide a data source for subsequent processing.

[0023] Based on the extraction of frames from individual video data and feature matching, several image groups are obtained. Each image group contains one frame from each individual video data. Specifically, due to the massive amount of video data, directly processing all frames is inefficient and contains a lot of redundant information. This step uses frame extraction technology to extract keyframes from the continuous video stream according to preset rules, such as fixed time intervals, fixed frame intervals, or scene change detection. Subsequently, feature matching technology is used to analyze the correlation between the frame images extracted from different independent video data. An image group refers to a collection of several frame images captured by different cameras at the same time or in the same scene. Through feature matching, the content captured by different cameras can be logically linked to ensure that the subsequent processing objects are observations of the same scene from different perspectives, thereby eliminating the frame misalignment caused by time asynchrony.

[0024] Feature points are extracted from each frame image in the image group to obtain several feature points; these feature points are then matched to obtain several sets of feature point pairs; each set of corresponding feature points includes two feature points, and these two feature points come from different frame images; specifically, feature points are local regions in the image that have significant identifiability, such as corners, spots, edges, etc. This step can use algorithms such as SIFT (Scale Invariant Feature Transform), SURF (Speed-Up Robust Feature Transform), or ORB (Oriented Fast and Rotated BRIEF) to extract feature points and their corresponding feature descriptors; each set of corresponding feature points includes two feature points, and these two feature points come from different frame images, which means that they are the imaging projections of the same physical point under different camera perspectives; through feature point matching, a pixel-level correspondence is established between different images, which is the basis for calculating the geometric transformation relationship between images; it should be understood that the choice of feature point extraction algorithm is not limited to the above-listed examples.

[0025] A correction scheme is generated for different frame images based on several feature point pairs. Specifically, the correction scheme refers to a set of parameters used to correct the geometric position of the image. Due to position and angle deviations between cameras, the coordinates of the same physical point differ in different images. This step uses the matched feature point pairs to analyze their coordinate transformation rules and calculates the transformation parameters that can eliminate this difference, i.e., the correction scheme. The correction scheme may include horizontal offset, vertical offset, etc. By calculating the statistical rules of multiple sets of feature point pairs, the influence of mismatched points can be effectively eliminated, generating a high-precision correction scheme.

[0026] This method acquires several correction schemes corresponding to the same independent video data, and corrects the camera corresponding to the independent video data based on these correction schemes. Specifically, to improve the robustness of the correction, this method does not rely solely on the matching results of a single frame image, but rather integrates several correction schemes generated from multiple frames. The final correction value is determined through statistical analysis, such as calculating the average, mode, or weighted average. If the hardware supports this, the final correction operation can be performed at the hardware level by adjusting the physical position of the camera; alternatively, it can be performed at the software level by using image processing algorithms such as affine transformation and perspective transformation to perform pixel-level translation or transformation on the acquired image data, thereby outputting the corrected video stream. Through the above scheme, this embodiment achieves full automation from the acquisition of original video data to the final correction execution, effectively solving the problem of misalignment in multi-camera stitching and improving the panoramic imaging quality.

[0027] In one possible implementation, several image groups are obtained by extracting frames from individual video data and performing feature matching, including: Select any independent video data as reference video data, obtain the first frame step size, extract frames from the reference video data based on the first frame step size to obtain several frame images, and record these frame images as reference images. Obtain the set temporal neighborhood radius, and construct the matching time period of the corresponding reference image with the frame time corresponding to each reference image as the center and the temporal neighborhood radius as the radius. In a multi-camera system, although the cameras theoretically start synchronously, there are often millisecond-level startup delays or transmission jitters in actual operation. In order to determine a unified benchmark, this step randomly selects the video stream of one of the cameras as reference video data. The setting of the first frame step size aims to balance processing efficiency and data coverage. For example, if the video frame rate is 30fps, the first frame step size can be set to 100 frames, that is, extract a reference image every approximately 3.3 seconds. The temporal neighborhood radius defines the time window for searching and matching. For example, setting it to 0.5 seconds means that the system will search for the corresponding matching image within a time range of 0.5 seconds before and after the reference image. This setting tolerates the large time deviation that may exist between cameras.

[0028] Select any other independent video data as the deviation video data, and extract video segments within the matching time period from the deviation video data; obtain the second frame step size, and extract frames from the video segment data based on the second frame step size to obtain several frame images, which are recorded as matching images. It can be understood that the first frame step size is much larger than the second frame step size, which can be 10 times, 20 times, etc. The concept of "second frame step size" is introduced here, and the second frame step size is much smaller than the first frame step size. This "coarse-fine combination" strategy is crucial: the reference image is extracted with a large step size, which reduces the amount of data processing and allows for rapid traversal of the entire video duration; while the matching image is extracted intensively with a small step size within the matching time period, which ensures that high-precision corresponding frames can be captured within the local time window, avoiding missing the best matching moment due to an excessively large step size. Thus, while ensuring computational efficiency, the accuracy of time deviation calculation is ensured.

[0029] Several matching images corresponding to the reference image are integrated into a time difference determination group corresponding to the reference image; a time deviation between the reference video data and the deviation video data is generated based on the reference image and its corresponding time difference determination group; the time axis of the deviation video data is corrected based on the time deviation to obtain the corrected video data. Specifically, the process of generating the time deviation includes: calculating the similarity between the reference image and each matching image in the time difference determination group; recording the matching image with the largest similarity value as the deviation image of the reference image, and recording the difference between the frame time corresponding to the reference image and the frame time corresponding to the deviation image as the time deviation corresponding to the reference image; the similarity calculation can be performed using various methods such as perceptual hashing algorithm, structural similarity index (SSIM), or feature point matching number. Since the reference image and the matching image are captured in the same scene, their content similarity is highest when their time points are aligned. Therefore, by traversing, calculating, and locking the maximum similarity value, the delay of the deviation video data relative to the reference video data can be accurately located.

[0030] Other independent video data are selected sequentially as deviation video data, and the time axis of the deviation video data is corrected to obtain the corresponding corrected video data. By traversing all non-reference video data, all video streams are uniformly corrected to the time reference of the reference video data, thus realizing the time synchronization of the multi-camera system.

[0031] Image groups are obtained by extracting frames from reference video data and corrected video data. It is understood that preprocessing is required for each frame image in the image group, including distortion correction, etc. Specifically, this step includes: cropping the overlapping parts of the reference video data and each corrected video data on the time axis, extracting frames from the overlapping parts to obtain several frame images; and grouping the frame images of the reference video data and each corrected video data with the same frame time into one image group.

[0032] Because the startup times or recording durations of each camera differ, the beginning and end of the video stream may not be complete. Extracting the overlapping portion of the timeline ensures that each set of images used for subsequent correction contains complete information from all cameras, avoiding correction failures due to missing data. Dividing frames from the same moment into image groups means that the frames within an image group should theoretically capture the scene at the same instant, providing a strict temporal consistency guarantee for subsequent correction based on spatial feature points. Through these steps, this embodiment successfully eliminates the frame misalignment problem caused by asynchronous acquisition from multiple cameras, laying a solid foundation for high-precision image stitching and correction.

[0033] In one possible implementation, generating the time deviation between the reference video data and the deviation video data based on the reference image and its corresponding time difference determination group includes: The similarity between the reference image and each matching image in the time difference determination group is calculated. It is understood that image similarity calculation is a relatively existing technique, and will not be elaborated upon here. It is worth noting that this embodiment uses the relationship between the various cameras of a multi-camera system to crop the reference image and matching images, obtaining the overlapping portion of the reference image and matching images, and calculating the similarity of the overlapping portion images. The matching image with the highest similarity value is recorded as the deviation image of the reference image, and the difference between the frame time corresponding to the reference image and the frame time corresponding to the deviation image is recorded as the time deviation corresponding to the reference image. Specifically, the frame time corresponding to the reference image is subtracted from the frame time of the deviation image. When the resulting time deviation is positive, it indicates that the reference image is delayed relative to the deviation image; when the time deviation is negative, it indicates that the deviation image is delayed relative to the reference image. The time deviations corresponding to each reference image are obtained sequentially; the average value of the time deviations corresponding to each reference image is used as the final time deviation; it is understandable that the mode can also be used as the final time deviation.

[0034] Furthermore, the temporal deviations corresponding to each reference image are sequentially obtained; the average of the temporal deviations corresponding to each reference image is used as the final temporal deviation. Using the average value instead of a single measurement as the final temporal deviation is to eliminate interference from random factors. In real-world scenarios, occlusion by moving objects, instantaneous changes in lighting, or frame drops during transmission can lead to misjudgments in a single similarity calculation. By calculating the average of the temporal deviations corresponding to multiple reference images, random errors can be effectively smoothed, making the final determined temporal deviation more stable and reliable, truly reflecting the inherent time difference between cameras. Based on this temporal deviation, the time axis of the deviated video data is shifted and corrected to align it temporally with the reference video data.

[0035] In one possible implementation, frame extraction is performed based on reference video data and correction video data to obtain an image group, including: cropping the overlapping parts of the reference video data and each correction video data on the time axis, extracting frames from the overlapping parts to obtain several frame images; and grouping the frame images of the reference video data and each correction video data with the same frame time into an image group.

[0036] In one possible implementation, several feature points are matched to obtain several sets of feature point pairs. This includes: selecting any one frame image from the image group as the reference image, and recording the other frame images as images to be corrected. Specifically, since the frame images within the image group have already undergone time synchronization correction, theoretically they were captured at the same time and in the same scene. Therefore, selecting which frame as the reference image will not affect the final calculation result of the relative position relationship. In practice, a frame image captured by one of the cameras can be randomly selected as the reference image, or the frame with the best quality can be selected as the reference image based on image quality evaluation indicators such as sharpness and brightness uniformity to improve the accuracy of subsequent matching. The frame images from the remaining cameras are then used as images to be corrected, waiting to be registered with the reference image.

[0037] Several feature points are obtained from a reference image and several feature points are obtained from any image to be corrected. The feature points in the reference image and the image to be corrected are matched to obtain the matching state between corresponding feature points. Feature point extraction can employ algorithms such as SIFT, SURF, or ORB. These algorithms can extract key points in the image that are rotation-invariant and scale-invariant, and generate a feature descriptor vector for each key point. The matching state is used to characterize whether two feature points are projections of the same physical point from different viewpoints. To accurately determine the matching state, this embodiment adopts a hierarchical matching mechanism based on Euclidean distance. Two feature points that are successfully matched are recorded as a pair of feature points. Each pair of feature points includes a feature point from the reference image and a feature point from the image to be corrected. Several sets of feature point pairs are generated sequentially between each image to be corrected and the reference image; it can be understood that there are generally several sets of feature point pairs between an image to be corrected and the reference image.

[0038] Through the above steps, this embodiment successfully established a precise correspondence between the reference image and each image to be corrected, providing sufficient and accurate data support for subsequent calculation of spatial offset based on feature points. This matching strategy based on Euclidean distance hierarchical screening and anti-repeated traversal effectively balances matching efficiency and accuracy, significantly improving the robustness of the image correction method.

[0039] In one possible implementation, feature points in the reference image and the image to be corrected are matched to obtain the matching state between corresponding feature points, including: The process involves obtaining feature descriptor vectors for several feature points in the reference image and the image to be corrected. Each feature descriptor vector is a multi-dimensional vector, such as a 128-dimensional SIFT descriptor, which characterizes the texture, gradient, and other local features of the neighborhood surrounding the feature point. The more similar two feature points are, the closer their descriptor vectors are in multi-dimensional space. Using any feature point in the reference image as the reference point, calculate the Euclidean distance between the feature descriptor vector corresponding to the reference point and the feature descriptor vector corresponding to any feature point in the image to be corrected. Euclidean distance is a commonly used metric for measuring the similarity between two vectors; a smaller distance indicates greater similarity in the local features of the two feature points. By calculating the Euclidean distances between the reference point and all feature points in the image to be corrected, a distance set can be constructed. Specifically, the formula for calculating the Euclidean distance is: ; Where p represents the feature descriptor vector of the reference point in the reference image, q represents the feature descriptor vector of the feature point in the image to be corrected, and n represents the dimension of the feature descriptor vector, for example, n=128 in the SIFT algorithm, p i and q i Let p and q represent the components of vectors p and q in the i-th dimension, respectively. Euclidean distance is a commonly used metric for measuring the similarity between two vectors; a smaller distance indicates greater similarity in the local features of the two feature points. The Euclidean distance between the reference point and all feature points in the image to be corrected is calculated. When the Euclidean distance is less than a set similarity distance threshold, the matching status between the feature point and the feature point corresponding to the reference point is set as a preliminary successful match; otherwise, the matching status is set as a failed match. It is understood that the specific value of the similarity distance threshold is set by expert experience to distinguish whether feature descriptors are similar. The similarity distance threshold is an empirical value used to quickly eliminate obviously irrelevant feature points. For example, if the reference point is a "corner point" in the scene, then the Euclidean distance between the "spots" in the flat area of ​​the image to be corrected and the reference point is usually large, and they will be directly judged as a failed match. This step greatly narrows the scope of subsequent filtering and reduces computational complexity.

[0040] The process iterates through all feature points in the image to be corrected, obtaining the matching status between the feature point corresponding to the reference point and each feature point in the image to be corrected. From the feature points with a preliminary matching status, the feature point with the smallest Euclidean distance is selected, and the matching status between this feature point and the reference point is recorded as a successful match. In the candidate point set after threshold filtering, the point with the smallest Euclidean distance is selected as the final matching point. This follows the "nearest neighbor" principle, ensuring the uniqueness and optimality of the matching. It should be understood that although multiple points may be smaller than the threshold, only the point with the smallest distance is the most likely correct corresponding point.

[0041] The feature points in the reference image are traversed to obtain the matching status between each feature point in the reference image and each feature point in the image to be corrected. It can be understood that whenever a reference point has a corresponding successfully matched feature point, the successfully matched feature point will be deleted when matching other reference points in the future. That is, when matching other reference points in the future, previously successfully matched feature points will not be considered. This mechanism is crucial. It prevents the "one-to-many" error situation where multiple reference points match the same feature point to be corrected, establishes a defense depth at the algorithm level, and ensures the one-to-one correspondence between feature point pairs, thereby providing a reliable data foundation for subsequent calculation of the correction scheme.

[0042] In one possible implementation, a correction scheme based on several feature point pairs to generate different frame images includes: acquiring several sets of feature point pairs of the reference image and the image to be corrected, and selecting the feature point pair with the smallest Euclidean distance as the reference feature point pair. Specifically, in the feature point matching process, the Euclidean distance characterizes the similarity between two feature points in the feature space. The smaller the Euclidean distance, the closer the local texture, gradient, and other features of the two feature points are, the higher the probability that they are projections of the same physical point, and the lower the probability of interference from factors such as noise and illumination changes. Therefore, selecting the feature point pair with the smallest Euclidean distance as the reference feature point pair is equivalent to selecting the pair with the "highest confidence" among all matching point pairs as the reference for subsequent offset traversal. This strategy constructs the "anchor point" of the algorithm, effectively preventing the risk of correction direction errors due to individual mismatched point pairs.

[0043] The coordinates of feature points belonging to the image to be corrected in the reference feature point pair are offset to obtain offset coordinate points. These offset coordinate points are located within the neighborhood corresponding to the set neighborhood radius of the reference feature point. The same offset is applied to the coordinates of feature points belonging to the image to be corrected in other feature point pairs to obtain corresponding offset coordinate points. Specifically, the essence of the correction scheme is to find an optimal translation vector, including horizontal and vertical offsets, so that the image to be corrected best fits the reference image after this translation. To find this optimal solution, this embodiment employs a traversal search strategy. A neighborhood radius is set for each feature point, for example, a radius of 2 pixels. Within the circular or rectangular area covered by this radius, the coordinates of the feature points of the image to be corrected in the reference feature point pair are offset according to a preset step size, for example, 0.1 pixels. It is worth noting that when traversing each possible offset, the coordinates of the feature points of the image to be corrected in all other feature point pairs need to undergo the same offset operation. This "overall linkage" mechanism ensures the consistency of the correction scheme and simulates the translation effect of the entire image.

[0044] A predetermined feature neighborhood radius is obtained. A feature neighborhood corresponding to the offset coordinate point is constructed with the offset coordinate point as the center and the feature neighborhood radius as the radius. The feature descriptor vector corresponding to the feature neighborhood is obtained and marked as the offset feature descriptor vector. Specifically, when the feature point coordinates undergo a slight translation, the surrounding image content will also change. To quantify the impact of this change on feature matching quality, the feature descriptor at the offset position needs to be recalculated. The feature neighborhood radius defines the range for calculating the feature descriptor, usually consistent with the neighborhood size during initial feature extraction. Centered on the offset coordinate point, image pixel information within this neighborhood is extracted, and a new feature descriptor vector, i.e., the offset feature descriptor vector, is generated using algorithms such as SIFT, SURF, or ORB. This process maps the spatial position change of pixels to a vector change in the feature space, providing a mathematical basis for subsequent quality assessment.

[0045] Calculate the Euclidean distance between the offset feature descriptor vector and the feature descriptor vector corresponding to the other feature point in its feature point pair, and denote it as the offset distance; specifically, the formula for calculating the offset distance is: ; Where Δx and Δy represent the lateral offset and vertical offset, respectively. Indicates at offset The offset feature descriptor vector calculated below, Let v_offset,i represent the feature descriptor vector corresponding to the baseline feature point, n represent the dimension of the feature descriptor vector, and v_offset,i and v_ref,i represent the vector values ​​respectively. and The component in the i-th dimension. The offset distance reflects the similarity between the image to be corrected and the reference image at the location of that feature point at the current offset. The smaller the offset distance, the closer the offset feature point is to the reference feature point in the feature space, that is, the better the image alignment effect.

[0046] For the same offset The formula for calculating the average offset distance for all m pairs of feature points is: ; Where m represents the total number of feature point pairs, Indicates the j-th pair of feature points at the offset The offset distance is calculated. Due to potential interference such as occlusion and repetitive textures in the image, the offset distance of a single feature point may be random. Therefore, this embodiment calculates the average offset distance of all feature points at the same offset. The average value strategy effectively smooths out errors caused by individual outliers and reflects the overall alignment quality of the image. After traversing all possible offset combinations, the system selects the offset with the smallest average value as the optimal correction scheme, i.e.: ; It should be understood that, since the offset step size can be set to a value much smaller than 1 pixel, such as 0.01 pixels, this method can overcome the limitation of integer pixels and achieve high-precision correction at the sub-pixel level, which is one of the core technical contributions of this application. Through the above steps, this embodiment uses the vector distance of feature descriptors as an evaluation index, and accurately determines the optimal offset of the image to be corrected in the horizontal and vertical directions by combining traversal and statistics, providing reliable single-frame data support for subsequent multi-frame statistical correction.

[0047] The average offset distance of each feature point under the same offset is calculated, and the offset with the smallest average value is selected as the correction scheme. The offset includes horizontal and vertical offsets, and the correction scheme includes both horizontal and vertical offsets. Specifically, the offset distance reflects the similarity between the image to be corrected and the reference image at the feature point position under the current offset. The smaller the offset distance, the closer the offset feature point is to the reference feature point in the feature space, i.e., the better the image alignment effect. Since there may be interference such as occlusion and repeated textures in the image, the offset distance of a single feature point may be random. Therefore, this embodiment calculates the average offset distance of all feature points under the same offset. The average value strategy can effectively smooth the error caused by individual outliers and reflect the overall alignment quality of the image. After traversing all possible offset combinations, the system selects the offset with the smallest average value as the optimal correction scheme. It should be understood that since the offset step size can be set to a value much smaller than 1 pixel, this method can break through the limitation of integer pixels and achieve sub-pixel level high-precision correction, which is one of the core technical contributions of this application. Through the above steps, this embodiment uses the vector distance of feature descriptors as an evaluation index. By combining traversal and statistics, it accurately determines the optimal offset of the image to be corrected in the horizontal and vertical directions, providing reliable single-frame data support for subsequent multi-frame statistical correction.

[0048] In conjunction with the first aspect above, in one possible implementation, the camera corresponding to the independent video data is calibrated based on the calibration scheme, including: Extract the longitudinal and lateral offsets from several correction schemes corresponding to independent video data; Calculate the variance of each vertical offset. When the variance is less than a set variance threshold, obtain the mode of the vertical offset as the vertical correction value; otherwise, obtain the average of the vertical offsets as the vertical correction value. Variance is a key indicator for measuring the dispersion of data. This step determines the central tendency of the data by calculating the variance of the vertical offsets. When the variance is less than the set variance threshold, it indicates that most vertical offset values ​​are very close, and the data distribution is concentrated. In this case, using the mode, i.e., the value with the highest frequency, as the final correction value has a significant advantage: the mode can naturally eliminate individual outliers that deviate from the mainstream value, reflecting the most common deviation. For example, if most values ​​in the dataset are 5 pixels, and only one value is 10 pixels, a mode of 5 pixels is more representative of the true deviation. Conversely, when the variance is greater than or equal to the variance threshold, it indicates that the data distribution is relatively dispersed, and there is no obvious "mainstream" value. In this case, the mode may not represent the overall situation, while the average can integrate the information of all data, smooth out positive and negative random errors, and thus obtain a compromise correction value. It should be understood that the specific value of the variance threshold can be set according to the accuracy requirements of the camera and the application scenario, for example, it can be set to the square of 1 pixel. This invention does not limit this.

[0049] The variance of each horizontal offset is calculated. When the variance is less than a set variance threshold, the mode of the horizontal offset is taken as the horizontal correction value; otherwise, the average of the horizontal offsets is taken as the horizontal correction value. The determination logic for the horizontal offset is completely consistent with that for the vertical offset, using the same variance judgment mechanism to process the horizontal dataset separately. This method of processing horizontal and vertical errors independently avoids mutual interference between them, further improving the accuracy of the correction.

[0050] The system corrects the original vertical and horizontal dimensions of the camera corresponding to the independent video data based on the vertical and horizontal correction values. Specifically, a translation transformation matrix is ​​set according to the vertical and horizontal correction values, and the image data obtained by the corresponding camera is translated based on the translation transformation matrix. Specifically, after determining the final vertical and horizontal correction values, the system generates a translation transformation matrix. This matrix contains translation parameters on the X-axis (horizontal) and Y-axis (vertical). For each subsequent frame acquired by the camera, the system applies this translation transformation matrix to perform an affine transformation, translating the image pixels as a whole according to the correction values. Through the above steps, this embodiment utilizes statistical principles to effectively eliminate the random errors that may exist in a single correction scheme, ensuring the high reliability and robustness of the final correction value, and solving the problem of unstable correction results due to poor quality of individual frames in the prior art.

[0051] Please see Figure 2Secondly, an image correction device for a multi-view camera is provided, comprising: a data acquisition module, a video data processing module, an image data processing module, and a correction module; The data acquisition module is used to acquire video data from the multi-camera system, and the video data includes video data from each camera in the multi-camera system. The video data processing module is used to extract frames and match features to obtain several image groups based on each independent video data in the video data. The image data processing module includes a feature point matching unit and a correction scheme generation unit; The feature point matching unit is used to extract feature points from each frame image in the image group to obtain a number of feature points; to match the number of feature points to obtain a number of feature point pairs; a set of corresponding feature points includes two feature points, and the two feature points come from different frame images; The correction scheme generation unit is used to generate correction schemes for different frame images based on several feature point pairs; obtain several correction schemes corresponding to the same independent video data, and generate horizontal correction values ​​and vertical correction values ​​for the independent video data based on the correction schemes; The correction module corrects the original longitudinal and lateral dimensions of the camera corresponding to the independent video data based on the longitudinal and lateral correction values.

[0052] Thirdly, this application provides a computer-readable storage medium storing instructions that, when executed on an image correction apparatus for a multi-view camera, cause the image correction apparatus for a multi-view camera to perform the methods described in the first aspect and any possible implementation thereof.

[0053] Fourthly, this application provides a computer program product containing instructions that, when run on an image correction apparatus for a multi-view camera, cause the image correction apparatus for the multi-view camera to perform the methods described in the first aspect and any possible implementation thereof.

[0054] Some of the data in the above formula are calculated by removing dimensions and taking their numerical values. The formula is the closest to the real situation obtained by software simulation of a large amount of collected data. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.

[0055] How this application works: This method acquires video data from multiple cameras, performs frame extraction and feature matching to obtain image groups, then extracts feature points and matches them to obtain feature point pairs, and finally generates a correction scheme to correct the cameras. This method utilizes the temporal information of the video data, eliminating the time asynchrony problem between different cameras through time deviation correction, and achieving high-precision spatial correction through feature point matching and offset calculation. Simultaneously, statistical methods are used to determine the final correction value, effectively eliminating interference from abnormal data, improving the stability and robustness of the correction, and solving the problems of image stitching misalignment and unstable correction accuracy in existing technologies for multi-camera systems.

[0056] The above embodiments are only used to illustrate the technical methods of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of this application without departing from the spirit and scope of the technical methods of this application.

Claims

1. An image correction method for multi-view cameras, characterized in that, include: Acquire video data from multiple cameras; Frame extraction and feature matching are performed on each independent video data in the video data to obtain several image groups; Feature points are extracted from each frame image in the image group to obtain several feature points; several feature points are matched to obtain several sets of feature point pairs; each set of corresponding feature points includes two feature points, and the two feature points come from different frame images; a correction scheme for different frame images is generated based on several feature point pairs. Several correction schemes corresponding to the same independent video data are obtained, and the camera corresponding to the independent video data is corrected based on the correction schemes.

2. The image correction method for a multi-view camera according to claim 1, characterized in that, The process of extracting frames and matching features from individual video data to obtain several image groups includes: Select any independent video data as reference video data, obtain the first frame step length, extract frames from the reference video data based on the first frame step length to obtain several frame images, and record the frame images as reference images. Obtain the set time neighborhood radius, and construct the matching time period of the corresponding reference image with the frame time corresponding to each reference image as the center and the time neighborhood radius as the radius. Select any other independent video data as the deviation video data, and extract video segments within the matching time period from the deviation video data; obtain the second frame step length, and perform frame extraction on the video segment data based on the second frame step length to obtain several frame images, and record the frame images as the matching images. A number of matching images corresponding to the reference image are integrated into a time difference determination group corresponding to the reference image; a time deviation between the reference video data and the deviation video data is generated based on the reference image and its corresponding time difference determination group; the time axis of the deviation video data is corrected based on the time deviation to obtain the corrected video data; Other independent video data are selected sequentially as the deviation video data, and the time axis of the deviation video data is corrected to obtain the corresponding corrected video data; Image groups are obtained by extracting frames based on reference video data and corrected video data.

3. The image correction method for a multi-view camera according to claim 2, characterized in that, Based on the reference image and its corresponding time difference, a time deviation between the reference video data and the deviation video data is generated, including: Calculate the similarity between the reference image and each matching image in the time difference determination group; record the matching image with the highest similarity value as the deviation image of the reference image, and record the difference between the frame time corresponding to the reference image and the frame time corresponding to the deviation image as the time deviation corresponding to the reference image; The time deviations corresponding to each reference image are obtained sequentially; the average of the time deviations corresponding to each reference image is used as the final time deviation.

4. The image correction method for a multi-view camera according to claim 2, characterized in that, The step of extracting images from reference video data and corrected video data includes: The overlapping portions of the reference video data and each correction video data on the time axis are cropped, and frames are extracted from the overlapping portions to obtain several frame images; the frame images of the reference video data and each correction video data with the same frame time are grouped into one image group.

5. The image correction method for a multi-view camera according to claim 1, characterized in that, The process of matching several feature points to obtain several pairs of feature points includes: Select any frame image from the extracted image group as the reference image, and record the other frame images as images to be corrected; Several feature points are obtained from the reference image, and several feature points are obtained from any image to be corrected; the feature points in the reference image and the image to be corrected are matched to obtain the matching status between the corresponding feature points; two feature points with a matching status of successful matching are recorded as a pair of feature points, and the pair of feature points includes a feature point from the reference image and a feature point from the image to be corrected. Several pairs of feature points are generated sequentially between each image to be corrected and the reference image.

6. The image correction method for a multi-view camera according to claim 5, characterized in that, Feature points in the reference image and the image to be corrected are matched to obtain the matching status between corresponding feature points, including: Obtain the feature descriptor vectors of several feature points corresponding to the reference image, and the feature descriptor vectors of several feature points corresponding to the image to be corrected; Using any feature point in the reference image as the reference point, calculate the Euclidean distance between the feature descriptor vector corresponding to the reference point and the feature descriptor vector corresponding to any feature point in the image to be corrected; when the Euclidean distance is less than a set similarity distance threshold, set the matching status between the feature point and the feature point corresponding to the reference point to preliminary matching success; otherwise, set the matching status between the feature point and the feature point corresponding to the reference point to matching failure. Traverse all feature points in the image to be corrected to obtain the matching status between the feature point corresponding to the reference point and each feature point in the image to be corrected; select the feature point with the smallest Euclidean distance from the feature points with the initial matching status, and record the matching status between the feature point and the reference point as a successful match; By traversing the feature points in the reference image, the matching status of each feature point in the reference image and each feature point in the image to be corrected is obtained.

7. The image correction method for a multi-view camera according to claim 1, characterized in that, The correction scheme based on several feature point pairs to generate different frame images includes: Obtain several sets of feature point pairs from the reference image and the image to be corrected, and select the feature point pair with the smallest Euclidean distance as the reference feature point pair; The coordinates of feature points belonging to the image to be corrected in the reference feature point pair are offset to obtain offset coordinate points. The offset coordinate points are located in the neighborhood corresponding to the set neighborhood radius of the feature points. The same offset is performed on the coordinates of feature points belonging to the image to be corrected in other feature point pairs to obtain the corresponding offset coordinate points. Obtain the set feature neighborhood radius, construct the feature neighborhood corresponding to the offset coordinate point with the feature neighborhood radius centered at the offset coordinate point, obtain the feature descriptor vector corresponding to the feature neighborhood, and mark it as the offset feature descriptor vector; Calculate the Euclidean distance between the offset feature descriptor vector and the feature descriptor vector corresponding to another feature point in its corresponding feature point pair, and denot it as the offset distance; calculate the average value of the offset distances corresponding to each feature point pair under the same offset, and select the offset with the smallest average value as the correction scheme; the offset includes a horizontal offset and a vertical offset, and the correction scheme includes a horizontal offset and a vertical offset.

8. The image correction method for a multi-view camera according to claim 1, characterized in that, Based on the correction scheme, the camera corresponding to the independent video data is calibrated, including: Extract the longitudinal and lateral offsets from several correction schemes corresponding to independent video data; Calculate the variance of each longitudinal offset. When the variance is less than a set variance threshold, obtain the mode of the longitudinal offset as the longitudinal correction value; otherwise, obtain the average value of the longitudinal offset as the longitudinal correction value. Calculate the variance of each lateral offset. When the variance is less than a set variance threshold, obtain the mode of the lateral offset as the lateral correction value; otherwise, obtain the average value of the lateral offset as the lateral correction value. The original vertical and horizontal dimensions of the camera corresponding to the independent video data are corrected based on the vertical and horizontal correction values.

9. An image correction apparatus for a multi-view camera, based on the application of an image correction method for a multi-view camera according to any one of claims 1 to 8, characterized in that, include: Data acquisition module, video data processing module, image data processing module, and correction module; The data acquisition module is used to acquire video data from the multi-camera system, and the video data includes video data from each camera in the multi-camera system. The video data processing module is used to extract frames and match features to obtain several image groups based on each independent video data in the video data. The image data processing module includes a feature point matching unit and a correction scheme generation unit; The feature point matching unit is used to extract feature points from each frame image in the image group to obtain a number of feature points; to match the number of feature points to obtain a number of feature point pairs; a set of corresponding feature points includes two feature points, and the two feature points come from different frame images; The correction scheme generation unit is used to generate correction schemes for different frame images based on several feature point pairs; obtain several correction schemes corresponding to the same independent video data, and generate horizontal correction values ​​and vertical correction values ​​for the independent video data based on the correction schemes; The correction module corrects the original longitudinal and lateral dimensions of the camera corresponding to the independent video data based on the longitudinal and lateral correction values.

10. A computer storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on the image correction apparatus for a multi-camera system, perform an image correction method for a multi-camera system as described in any one of claims 1 to 8.