Multi-image stitching method, system, device and medium for target tracking and recognition

By performing feature extraction and pose calculation on the images of multiple cameras, combined with image transformation and stitching preview, problems such as false detection and missed detection in multi-picture stitching are solved, efficient target recognition and tracking are achieved, and the accuracy and efficiency of multi-picture stitching are improved.

CN114581307BActive Publication Date: 2025-05-13WUXI FANTE INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210273485.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2025-05-13
Estimated Expiration
2042-03-18

AI Technical Summary

Technical Problem

The existing multi-picture splicing effect is poor, the picture synthesis efficiency is low, and the accuracy of multiple target recognition and tracking tasks is affected by factors such as targets across cameras, across regions, a wide variety of targets, lighting changes, occlusion and camera transfer, resulting in problems such as false detection, missed detection, target trajectory disconnection and series connection.

Method used

By extracting the original images collected by multiple cameras, calculating the pose information and image transformation parameters, performing image transformation and stitching previews, determining the seam lines using the maximum flow graph separator, and optimizing the gain coefficient through the error function, combining semantic segmentation and object detection model for target tracking and identification.

Benefits of technology

The multi-screen splicing effect is improved, and the error detection and missed detection problems caused by targets across cameras and regions are solved, the calculation amount and delay are reduced, and the accuracy of target recognition and tracking tasks is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114581307B_ABST
    Figure CN114581307B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-image stitching method, system, device and medium for target tracking and identification, and the method includes: extracting features from original images acquired by multiple cameras; obtaining the position information of each camera according to the feature extraction result, and then obtaining the image transformation parameters of each camera; performing image transformation on multiple original images according to the image transformation parameters to obtain multiple first-class transformation images; after reducing the resolution of multiple original images by a predetermined multiple, performing image transformation according to the image transformation parameters to obtain multiple second-class transformation images, and performing image stitching preview to obtain stitching preview parameters; performing image stitching on multiple first-class transformation images according to the stitching preview parameters to obtain a stitching map; and tracking and identifying each target on the stitching map through a pre-trained model. The amount of calculation and the time delay of multi-image stitching of the present invention are both greatly superior to the existing stitching schemes, and have certain practical promotion significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a multi-image stitching method, system, device and medium for target tracking and identification. Background Art

[0002] With the development of pattern recognition technology and video analysis and processing technology, people have an increasingly strong demand for the safety of daily activities. Intelligent video surveillance systems have been widely used in the field of security to provide protection for people's property and life safety. Object detection and recognition based on video sequences is an important part of intelligent video surveillance systems and is used in important places such as shopping malls, parking lots, banks, exhibitions, and railway stations. Object detection and recognition of its state can be used to analyze whether there are any abnormal events in the monitoring scene.

[0003] At present, the real-time target tracking and recognition task for multiple surveillance images is easily affected by many factors, such as the cross-camera and cross-region targets, the wide variety of targets, lighting changes, occlusion problems, camera transfer, etc., which may lead to false detection, missed detection, disconnection and connection of target tracks, resulting in low accuracy of the entire recognition task. Summary of the invention

[0004] 1. Technical issues to be resolved

[0005] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a multi-image stitching method, system, device and medium for target tracking and identification, which solves the technical problems of poor multi-image stitching effect, low image synthesis efficiency and low accuracy of multiple target identification and tracking tasks.

[0006] (II) Technical solution

[0007] In order to achieve the above object, the main technical solutions adopted by the present invention include:

[0008] In a first aspect, an embodiment of the present invention provides a multi-image stitching method for target tracking and recognition, comprising:

[0009] Perform feature extraction on the original images acquired by multiple cameras;

[0010] According to the feature extraction result of each original image, the position information of each camera is obtained;

[0011] Obtaining image transformation parameters of each camera according to the posture information;

[0012] Performing image transformation on the plurality of original images according to corresponding image transformation parameters to obtain a plurality of first-type transformed images;

[0013] After reducing the resolution of the plurality of original images by a predetermined multiple, performing image transformation according to corresponding image transformation parameters to obtain a plurality of second-type transformed images, and performing image splicing preview on the plurality of the second-type transformed images to obtain splicing preview parameters;

[0014] According to the splicing preview parameters, multiple first-type transformed images are spliced ​​to obtain a spliced ​​image.

[0015] Optionally, N feature sub-images are extracted from each original image, and each feature sub-image includes feature point coordinates and a feature vector.

[0016] Optionally, obtaining the pose information of each camera according to the feature extraction result of each original image includes:

[0017] By calculating the cosine distance between the feature vectors of each two original images, the matching degree of each two original images is obtained;

[0018] Determining the adjacent relationship between each camera according to the matching degree;

[0019] According to the coordinates of the feature points, the initial intrinsic parameter matrix and the initial extrinsic parameter matrix of each camera are obtained by using the internal and external parameter formulas;

[0020] The initial matrix of internal parameters and the initial matrix of external parameters are optimized by the bundle adjustment method to obtain the internal parameter matrix and the external parameter matrix of each camera;

[0021] Obtaining posture information of each camera based on the intrinsic parameter matrix and the extrinsic parameter matrix;

[0022] Obtaining the position information of each camera according to the posture information of each camera and the adjacent relationship between each camera;

[0023] Among them, for each group of adjacent cameras, the internal and external parameter formulas are:

[0024]

[0025] In formula (1), K1 and K2 are the internal parameter matrices of any group of adjacent cameras 1 and 2 to be solved, R1 and R2 are the external parameter matrices of any group of adjacent cameras 1 and 2 to be solved, and x 2n+1 and x 2n+2 is a known set of matching feature point coordinates.

[0026] Optionally, obtaining the image transformation parameters of each camera according to the posture information includes:

[0027] Solving the remapping matrix corresponding to the correction transformation of each camera based on the intrinsic parameter matrix and the extrinsic parameter matrix;

[0028] Based on the extrinsic matrix of each camera, by solving the vertex coordinates of each original image, the size of the mosaic image and the translation amount required for each original image to be translated to the corresponding position in the mosaic image are obtained;

[0029] Obtaining image transformation parameters according to the remapping matrix, the image stitching scale and the translation amount;

[0030] Wherein, the remapping matrix is:

[0031]

[0032] In formula (2), u and v are the transformed pixel coordinates, and x and y are the corresponding pixel coordinates in the original image, x = sin(π-v)·sin(u), y = cos(π-v), z = sin(π-v)·cos(u); p = sin(π-v)·sin(u), q = cos(π-v), r = sin(π-v)·cos(u).

[0033] Optionally, performing image stitching preview on a plurality of the second-type transformed images to obtain stitching preview parameters includes:

[0034] According to the position information of each camera, performing image stitching preview on the second type of transformed images to obtain a stitching preview image;

[0035] In the overlapping area of ​​each image in the splicing preview image, a seam line is determined by a maximum flow graph cut method; the seam line is a line connecting a number of pixels in the overlapping area that meet a preset similarity degree;

[0036] Taking the seam line as a reference, only pixels of the adjacent left image are retained on the left side of the seam line, and only pixels of the adjacent right image are retained on the right side of the seam line;

[0037] Amplifying the resolution of the overlapping area after the seam line optimization by the predetermined multiple;

[0038] Assigning a corresponding gain coefficient to each part of the spliced ​​preview image through an error function to make the image intensity of the overlapping area equal or similar;

[0039] Obtaining a splicing preview parameter according to the overlapped area magnified by the predetermined multiple and the gain coefficient;

[0040] Wherein, the error function is:

[0041]

[0042] In formula (3), g i and g jis the gain coefficient of image i and image j, R(i,j) represents the overlapping area of ​​image i and image j, I i (u i ) represents the average intensity I of image i in the overlapping area R(i,j) ij ;u i and u j They respectively represent the same point on the overlapping area R(i,j) corresponding to the point position in the images of image i and image j;

[0043]

[0044] In formula (4), R, G and B represent the intensity values ​​of the red, green and blue components of the color image respectively, and N ij Represents the number of pixels in the overlapping area R(i,j);

[0045] The empirical formula of the error function is:

[0046]

[0047] In formula (5), I ij represents the average intensity of image i in the overlapping area of ​​image i and image j, I ji represents the average intensity of image i in the overlapping area of ​​image i and image j; σ N and σ g denote the standard deviation of error and gain, σ N =10,σ g =0.1;

[0048] The empirical formula for g i The derivative of is:

[0049]

[0050] Make formula (6) equal to 0 and expand it into g1, g2, ..., g n The equation for the variable is:

[0051]

[0052] By taking the derivative of e with respect to all g, a system of equations including n linear equations as shown in equation (7) is established, and n gain coefficients g are obtained by solving the system of equations.

[0053] Optionally, the method further comprises:

[0054] Dividing the spliced ​​image into multiple sub-images using a pre-trained semantic segmentation model;

[0055] Performing target detection on the multiple sub-images using a pre-trained target detection model, and mapping the detected target coordinates back to the spliced ​​image;

[0056] Each target is tracked individually according to the mapped target coordinates, and the attributes of each tracked target are recognized based on deep learning.

[0057] Optionally, dividing the spliced ​​graph into a plurality of sub-graphs by using a pre-trained semantic segmentation model comprises:

[0058] Detecting the reference area in the spliced ​​image by using a pre-trained semantic segmentation model;

[0059] The reference area is drawn on a binary image of the same size, and a lateral expansion operation is performed on the binary first area so that there is only one connected domain in the binary image;

[0060] Based on the connected domain, the spliced ​​graph is segmented into three initial segmentation areas;

[0061] The highest point and the lowest point of each of the initially segmented regions are used as segmentation lines to cut the spliced ​​image into three sub-images with overlapping areas.

[0062] In a second aspect, an embodiment of the present invention provides a multi-image stitching system for target tracking and recognition, comprising:

[0063] A feature extraction module is used to extract features from the original images acquired by the multiple cameras;

[0064] A posture determination module, used to obtain the posture information of each camera based on the feature points extracted from each of the original images;

[0065] An image transformation parameter output module, used to obtain the image transformation parameters of each camera according to the posture information;

[0066] A first image transformation module, configured to perform image transformation on the original images captured by the plurality of cameras according to corresponding image transformation parameters to obtain a plurality of first-type transformed images;

[0067] A second image transformation module, configured to reduce the resolution of the original images collected by the plurality of cameras by a predetermined multiple, and then perform image transformation according to corresponding image transformation parameters to obtain a plurality of second-type transformed images;

[0068] A splicing preview module, used for performing image splicing preview on a plurality of the second-type transformed images to obtain splicing preview parameters;

[0069] An image stitching module is used to stitch a plurality of the first-type transformed images to obtain a stitching image according to the stitching preview parameters.

[0070] In a third aspect, an embodiment of the present invention provides a target tracking and recognition device based on multi-image splicing, comprising:

[0071] at least one database;

[0072] and a memory in communication with the at least one database;

[0073] The memory stores instructions that can be executed by the at least one database, and the instructions are executed by the at least one database so that the at least one database can execute the multi-image stitching method for target tracking and identification as described above.

[0074] In a fourth aspect, an embodiment of the present invention provides a computer-readable medium having computer executable instructions stored thereon, and when the executable instructions are executed by a processor, the multi-image stitching method for target tracking and recognition as described above is implemented.

[0075] (III) Beneficial effects

[0076] The beneficial effects of the present invention are as follows: the present invention obtains the image transformation parameters of each camera by judging the posture information of each camera, which not only improves the multi-screen stitching effect, but also effectively solves the problems of false detection, missed detection, target track disconnection and series connection in multi-screen stitching display due to many factors such as targets crossing cameras and regions, a wide variety of targets, lighting changes, occlusion problems, camera transfer, etc., thereby laying a solid foundation for subsequent detection and recognition tasks; at the same time, the present invention first performs a preview of screen stitching at a low resolution to obtain effective stitching preview parameters, and then in subsequent actual stitching, based on the stitching preview parameters, its computational complexity and the time delay of multi-screen stitching are greatly advantageous over the existing stitching scheme, and has certain practical promotion significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 A schematic diagram of a flow chart of a multi-image stitching method for target tracking and identification provided by an embodiment of the present invention;

[0078] Figure 2 The original images captured by multiple cameras in a multi-image stitching method for target tracking and recognition provided by an embodiment of the present invention;

[0079] Figure 3 A corrected image of a multi-image stitching method for target tracking and recognition provided by an embodiment of the present invention;

[0080] Figure 4 , Figure 5 , Figure 6 as well as Figure 7They are respectively a first image, a second image, a third image and a fourth image after image mapping in a multi-image stitching method for target tracking and recognition provided by an embodiment of the present invention;

[0081] Figure 8 , Fig. 9 , Fig.10 as well as Fig.11 A first image, a second image, a third image and a fourth image using a seam line as a reference line in a multi-image stitching method for target tracking and recognition provided by an embodiment of the present invention;

[0082] Fig.12 A mosaic graph of a multi-image mosaic method for target tracking and recognition provided by an embodiment of the present invention;

[0083] Fig.13 A schematic diagram of a reference area obtained through semantic segmentation in a multi-image stitching method for target tracking and recognition provided by an embodiment of the present invention;

[0084] Fig.14 A schematic diagram of drawing a reference area on a binary image in a multi-image stitching method for target tracking and recognition provided by an embodiment of the present invention;

[0085] Fig.15 A schematic diagram of a connected domain of a multi-image stitching method for target tracking and recognition provided by an embodiment of the present invention;

[0086] Fig.16 A schematic diagram of the initial segmentation area of ​​a multi-image stitching method for target tracking and recognition provided by an embodiment of the present invention;

[0087] Fig.17 , Fig.18 as well as Fig.19 They are respectively a first sub-image, a second sub-image and a third sub-image of a multi-image stitching method for target tracking and recognition provided by an embodiment of the present invention;

[0088] Fig. 20 A tracking and identification process of a multi-image stitching method for target tracking and identification is provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0089] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation modes in conjunction with the accompanying drawings.

[0090] like Figure 1As shown, an embodiment of the present invention proposes a multi-image stitching method for target tracking and recognition, comprising: first, extracting features from the original images acquired by multiple cameras; second, obtaining the posture information of each camera based on the feature extraction result of each original image; then, obtaining the image transformation parameters of each camera based on the posture information; further, performing image transformation on the multiple original images based on the corresponding image transformation parameters to obtain multiple first-class transformed images; then, after reducing the resolution of the multiple original images by a predetermined multiple, performing image transformation on the corresponding image transformation parameters to obtain multiple second-class transformed images, and performing image stitching preview on the multiple second-class transformed images to obtain stitching preview parameters; then, based on the stitching preview parameters, performing image stitching on the multiple first-class transformed images to obtain a stitching image.

[0091] The present invention obtains the image transformation parameters of each camera by judging the posture information of each camera, which not only improves the multi-screen stitching effect, but also effectively solves the problems of false detection, missed detection, target track disconnection and series connection in multi-screen stitching display due to many factors such as targets crossing cameras and regions, various target types, illumination changes, occlusion problems, camera transfer, etc., laying a solid foundation for subsequent detection and recognition tasks; at the same time, the present invention first performs a preview of screen stitching at a low resolution to obtain effective stitching preview parameters, and then performs the subsequent actual stitching based on the stitching preview parameters. The amount of calculation and the delay of multi-screen stitching are both much more advantageous than those of existing stitching schemes, and has certain practical promotion significance.

[0092] In order to better understand the above technical solution, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0093] Specifically, the present invention provides a multi-image stitching method for target tracking and recognition, comprising:

[0094] S1. Extract features from the original images captured by multiple cameras. The present invention uses SIFT (or SURF or ORB) to extract feature points from each image, and extracts N feature sub-images from each original image, each of which includes a feature point coordinate and a 128-dimensional feature vector.

[0095] S2. Based on the feature extraction results of each original image, the pose information of each camera is obtained.

[0096] Further, step S2 includes:

[0097] S21. Obtain the matching degree of each two original images by calculating the cosine distance between the feature vectors of each two original images.

[0098] S22. Determine the neighbor relationship between each camera according to the matching degree.

[0099] S23. According to the coordinates of the feature points, the initial internal parameter matrix and the initial external parameter matrix of each camera are obtained through the internal and external parameter formulas.

[0100] S23. Optimize the initial internal parameter matrix and the initial external parameter matrix by using the bundle adjustment method to obtain the internal parameter matrix and the external parameter matrix of each camera.

[0101] S24. Obtain the posture information of each camera based on the intrinsic parameter matrix and the extrinsic parameter matrix.

[0102] S25. Obtain the position and posture information of each camera based on the posture information of each camera and the adjacent relationship between each camera.

[0103] Among them, for each group of adjacent cameras (such as camera 1 and camera 2), the internal and external parameter formulas are:

[0104]

[0105] In formula (1), K1 and K2 are the internal parameter matrices of any group of adjacent cameras 1 and 2 to be solved, R1 and R2 are the external parameter matrices of any group of adjacent cameras 1 and 2 to be solved, and x 2n+1 and x 2n+2 is a known set of matching feature point coordinates.

[0106] In a specific embodiment, the present invention calculates the similarity between all cameras for multiple feature vectors, and the calculation is based on the cosine distance between feature vectors. If there are 4 cameras, a 4×4 similarity matrix will be obtained in the end:

[0107]

[0108] The value in the i-th row and j-th column of the matrix represents the matching degree between the i-th camera and the j-th camera. The larger the value, the greater the probability that they are adjacent. Generally, we consider the camera groups with values ​​greater than 1 to be adjacent cameras. Taking the above matrix as an example, our conclusion is that camera 1 is adjacent to camera 2, camera 2 is adjacent to camera 3, and camera 3 is adjacent to camera 4.

[0109] The camera posture is determined by the camera's intrinsic parameters and extrinsic parameters. The former includes the focal length, the position of the principal point and the pixel size, and the latter includes the world coordinates of the camera center point and the camera rotation angle. The present invention uses a 3×3 intrinsic parameter matrix K and a 3×3 extrinsic parameter matrix R to characterize them.

[0110] Preferably, the specific steps of solving the intrinsic and extrinsic parameter matrices according to the coordinates of the feature points are:

[0111] Given a set of matching points x1 and x2 of camera 1 and camera 2, they correspond to the same point X in space. According to the transformation relationship from the world coordinate system to the camera coordinate system, we have:

[0112] x1=K1·R1·X

[0113] x2=K2·R2·X

[0114] Combining the above two equations, we get

[0115] R1 -1 K1 -1 x1=R2 -1 K2 -1 x2

[0116] Among them, R1 and K1 are the intrinsic and extrinsic matrix of camera 1 to be solved, R2 and K2 are the intrinsic and extrinsic matrix of camera 2 to be solved, and x1 and x2 are the image coordinates of a known set of matched feature points. A set of matching points can get a linear equation. Take the 50+ matching points with the highest scores and establish a linear equation system to solve R1, K1, R2, K2. The specific algorithm for solving is the RANSAC (random consensus sampling) algorithm. The process is as follows:

[0117] (1) Randomly sample N points

[0118] (2) According to the N sampled points, the least squares method is used to solve the local optimal solution of the unknown parameters, thereby obtaining a temporary model

[0119] (3) Calculate the Mean Square Error of the remaining sampling points (points other than N points) based on the temporary model

[0120] (4) The points with errors greater than the threshold are recorded as outliers, and the points with errors less than the threshold are recorded as inliers.

[0121] (5) Repeat the above four steps 3 to 5 times

[0122] (6) All internal points are taken as the final sampling points, and the least squares method is used to find the optimal solution of the unknown parameters, which is recorded as the final model.

[0123] After the above steps, the intrinsic parameter matrix (K) and extrinsic parameter matrix (R) of the four cameras are obtained.

[0124]

[0125] The above is just a rough calculation of the internal and external parameters based on the feature points between adjacent cameras. Now we need to use Bundle Adjustment to globally optimize the internal and external parameters. (The bundle adjustment method itself is not introduced.) After this method, the internal and external parameter matrix will be more accurate:

[0126]

[0127] S3. Obtain image transformation parameters of each camera based on the posture information.

[0128] Further, step S3 includes:

[0129] S31. Based on the intrinsic parameter matrix and the extrinsic parameter matrix, solve the remapping matrix corresponding to the correction transformation of each camera.

[0130] S32. Based on the external parameter matrix of each camera, by solving the vertex coordinates of each original image, the size of the mosaic image and the translation amount required for each original image to be translated to the corresponding position in the mosaic image are obtained.

[0131] S33, obtaining image transformation parameters according to the remapping matrix, the image stitching scale and the translation amount.

[0132] In a specific embodiment, according to the intrinsic and extrinsic parameter matrices of each camera, we solve the remapping matrix corresponding to the correction transformation of each camera:

[0133] The remapping matrix is:

[0134]

[0135] In formula (2), u and v are the pixel coordinates after transformation, and x and y are the corresponding pixel coordinates in the original image, x = sin(π-v)·sin(u), y = cos(π-v), z = sin(π-v)·cos(u);

[0136] p=sin(π-v)·sin(u), q=cos(π-v), r=sin(π-v)·cos(u).

[0137] like Figure 3 As shown in the figure, after the above correction, the images of different cameras are guaranteed to be in the same plane, with only up, down, left, and right translation errors. The size of each correction image is: [(2206, 1512), (2204, 1515), (2216, 1521), (2192, 1505).

[0138] Then, according to the external parameter matrix of each camera, the vertex coordinates of the upper left corner of each screen are solved:

[0139] [(-3182, 3408), (-1689, 3435), (-592, 3480), (1080, 3469)]

[0140] Combining the coordinates of the upper left corner vertex and the size of each correction image, we can calculate

[0141] (1) The size of the final mosaic image

[0142] (2) The amount of translation required to move each corrected image to the appropriate position in the stitched image.

[0143] S4. Performing image transformation on the multiple original images according to corresponding image transformation parameters to obtain multiple first-type transformed images.

[0144] According to the remapping matrix and translation of the correction transformation, the final transformation matrix of each camera is calculated. With the transformation matrix, the corresponding camera image only needs to be remapped once (put the pixels in one image to the specified position in another image) to become the corresponding position and posture in the spliced ​​image. Figure 4 , 5 , 6 and 7 are diagrams of each original image after image transformation according to the transformation parameters of each camera.

[0145] S5. After reducing the resolution of the plurality of original images by a predetermined multiple, performing image transformation according to corresponding image transformation parameters to obtain a plurality of second-type transformed images, and performing image splicing preview on the plurality of second-type transformed images to obtain splicing preview parameters.

[0146] In step S5, the original image is downsampled to obtain the splicing preview parameters in order to reduce the amount of calculation and improve the efficiency of synthesis and splicing.

[0147] Furthermore, performing image splicing preview on a plurality of the second-type transformed images to obtain splicing preview parameters includes:

[0148] S51. Perform image stitching preview on the second-type transformed images according to the position information of each camera to obtain a stitching preview image.

[0149] S52, in the overlapping area of ​​each image in the splicing preview image, the seam line is determined by the maximum flow graph cut method; the seam line is the connection line of several pixels in the overlapping area that meet the preset similarity. The Boykov-Kolmogorov maximum flow segmentation method is used here, and the overlapping area of ​​adjacent images is regarded as a directed graph (V, E) in the data structure, V represents the node, and E represents the edge. The nodes are divided into general nodes (analogous to the pixels in the overlapping area) and terminal nodes (two categories, S and T, analogous to the left and right sides of the seam line). In this way, the problem is converted into finding a terminal node for each general node, which is connected by edges, and the energy value of all edges is the lowest in the end. When the specific segmentation is performed, it is iterated continuously through the "growth stage", "augmentation stage" and "adoption stage" (not specifically expanded) until the sum of the energy values ​​of all edges in the directed graph is less than the given threshold. At the end, the pixel with terminal S and the pixel with terminal T will divide the overlapping area into two blocks, and the seam between the two blocks is the seam line we are looking for.

[0150] S53, taking the seam line as a reference, retaining only pixels of the adjacent left image on the left side of the seam line, and retaining only pixels of the adjacent right image on the right side of the seam line.

[0151] S54, amplifying the resolution of the overlapping area after the seam line optimization by the predetermined multiple.

[0152] like Figure 4-Figure 7 As shown in , there are overlapping areas in adjacent images, so we need to find the best seam line for each adjacent image. The seam is the line connecting the most similar pixels in the overlapping area. After the seam is determined, in the overlapping area, only the adjacent left image is selected on the left side of the line, and only the adjacent right image is selected on the right side of the line. Finally, Figure 8-Figure 11 .

[0153] S55, assigning a corresponding gain coefficient to each part of the spliced ​​preview image through an error function to make the image intensity of the overlapping area equal or similar.

[0154] S56, obtaining splicing preview parameters according to the overlapped area magnified by the predetermined multiple and the gain coefficient.

[0155] A stitched image can be obtained by directly stacking the effective pixels (color areas) of multiple images. However, if the exposure levels of different images are different, an obvious sudden change in brightness will appear near the seam of the stitched image, making the image look very unnatural. Therefore, the present invention also performs exposure compensation on each pair of images to make all images have the same exposure level.

[0156] The following gain compensation method is used:

[0157] Gain compensation is to assign a gain coefficient to each image so that the image intensity in the overlapping area is equal or similar. It can be achieved using an error function, which is:

[0158]

[0159] In formula (3), g i and g j is the gain coefficient of image i and image j, R(i,j) represents the overlapping area of ​​image i and image j, I i (u i ) represents the average intensity I of image i in the overlapping area R(i,j) ij ;u i and u j They respectively represent the same point on the overlapping area R(i,j) corresponding to the point position in the images of image i and image j;

[0160]

[0161] In formula (4), R, G and B represent the intensity values ​​of the red, green and blue components of the color image respectively, and N ij Represents the number of pixels in the overlapping area R(i,j);

[0162] The empirical formula of the error function is:

[0163]

[0164] In formula (5), I ij represents the average intensity of image i in the overlapping area of ​​image i and image j, I ji represents the average intensity of image i in the overlapping area of ​​image i and image j, σ N and σ g denote the standard deviation of error and gain, σ N =10 (if the intensity range is 0 to 255), σ g =0.1;

[0165] Formula (5) is a quadratic objective function of the gain coefficient g, which can be solved in closed form by making its derivative equal to 0. The empirical formula (5) is i The derivative of is:

[0166]

[0167] Make formula (6) equal to 0 and expand it into g1, g2, ..., g n The equation for the variable is:

[0168]

[0169] By taking the derivative of e with respect to all g, a system of equations including n linear equations as shown in equation (7) is established, and n gain coefficients g are obtained by solving the system of equations.

[0170] S6. Based on the splicing preview parameters, multiple first-class transformed images are spliced ​​to obtain a spliced ​​image. The effective pixels of multiple first-class transformed images can be directly stacked to obtain the following: Fig.12 The mosaic shown.

[0171] S7. Track and identify each target on the spliced ​​image through the pre-trained model.

[0172] Further, step S7 includes:

[0173] S71. Divide the spliced ​​image into multiple sub-images using a pre-trained semantic segmentation model.

[0174] S72: Perform target detection on the multiple sub-images using a pre-trained target detection model, and map the detected target coordinates back to the spliced ​​image.

[0175] S73. Track each target separately according to the mapped target coordinates, and perform attribute recognition based on deep learning on each tracked target.

[0176] The semantic segmentation model is any one of DeepLabV1, DeepLabV2, DeeplabV3 and DeepLabV3+; the target detection model is any one of Faster R-CNN, SSD and YOLOS; the target tracking adopts any one of Kalman filter, CenterTrack, FairMot and ByteTrack; the attribute recognition adopts contrastive learning. The use and training methods of the above models are all realized by existing technologies.

[0177] like Fig. 20 As shown, in the detailed step of the above step S7, target detection is performed on the sub-graphs corresponding to each area, and the coordinates are mapped back to the mosaic graph, and target tracking such as aircraft tracking, vehicle tracking, and pedestrian tracking is performed on the mosaic graph, and various attributes of the tracked targets are identified. For example, the direction, take-off or landing of the aircraft; the specific type of the vehicle is a car, truck, or bus; the attribute identification of the pedestrian is a child, an adult, or an elderly person.

[0178] Furthermore, step S71 includes:

[0179] S711, using a pre-trained semantic segmentation model to detect the reference area in the spliced ​​image. Fig.12The image data shown in FIG. 1 is input into the pre-trained semantic segmentation model to detect the reference area in the picture. In this embodiment, the "house" area is selected. Fig.13 The pre-trained semantic segmentation model detects and segments the house area in the picture. Therefore, the reference area is all the house areas in the picture. The reason for selecting the "house" area here is that the "house" area is a general object that can distinguish the ground and the sky.

[0180] S712, such as Fig.14 As shown in FIG. 1 , the reference area (“house” area) is drawn on a binary image of the same size, and a horizontal expansion operation is performed on the binary first area so that there is only one connected domain in the binary image (in Fig.15 (displayed in white in the figure).

[0181] S713: Based on the connected domain, the mosaic image is segmented into three initial segmentation areas. Fig.16 As shown, based on the connected domain, the spliced ​​image is divided into three parts, the upper (grid area) is the sky, the middle (horizontal line area) is the house, and the lower (vertical line area) is the road / river.

[0182] S714, using the highest and lowest points of each of the initially segmented regions as a dividing line, and cutting the mosaic image into three sub-images with overlapping areas. Based on the three initially segmented regions, using the highest and lowest points of each region as a dividing line, and cutting the mosaic image into three sub-images with overlapping areas, respectively representing the aerial area ( Fig.17 ), Housing Area( Fig.18 ), Ground / River Area( Fig.19 ).

[0183] In addition, an embodiment of the present invention further provides a multi-image stitching system for target tracking and recognition, comprising:

[0184] The feature extraction module is used to extract features from the original images acquired by multiple cameras.

[0185] The posture determination module is used to obtain the posture information of each camera based on the feature points extracted from each original image.

[0186] The image transformation parameter output module is used to obtain the image transformation parameters of each camera according to the posture information.

[0187] The first image transformation module is used to perform image transformation on the original images collected by the multiple cameras according to corresponding image transformation parameters to obtain multiple first-type transformed images.

[0188] The second image transformation module is used to reduce the resolution of the original images collected by the multiple cameras by a predetermined multiple, and then perform image transformation according to corresponding image transformation parameters to obtain multiple second-type transformed images.

[0189] The splicing preview module is used to perform image splicing preview on a plurality of the second-type transformed images to obtain splicing preview parameters.

[0190] An image stitching module is used to stitch a plurality of the first-type transformed images to obtain a stitching image according to the stitching preview parameters.

[0191] The multi-image stitching system for target tracking and recognition also includes:

[0192] The tracking and identification module is used to track and identify each target on the spliced ​​image through a pre-trained model.

[0193] Since the system / device described in the above embodiments of the present invention is a system / device used to implement the method of the above embodiments of the present invention, a person skilled in the art can understand the specific structure and deformation of the system / device based on the method described in the above embodiments of the present invention, and thus will not be described in detail here. All systems / devices used in the method of the above embodiments of the present invention belong to the scope of protection of the present invention.

[0194] Furthermore, an embodiment of the present invention further provides a target tracking and identification device based on multi-image stitching, comprising: at least one database; and a memory communicatively connected to the at least one database; wherein the memory stores instructions executable by the at least one database, and the instructions are executed by the at least one database so that the at least one database can execute a multi-image stitching method for target tracking and identification as described above.

[0195] Furthermore, the present invention provides a computer-readable medium having computer-executable instructions stored thereon, and when the executable instructions are executed by a processor, the multi-image stitching method for target tracking and identification as described above is implemented.

[0196] In summary, the present invention provides a multi-image stitching method, system, device and medium for target tracking and identification. The present invention implements a complete set of feasible multi-screen target tracking and identification solutions by first stitching, then detecting, and then tracking by category and identifying each target. The detection and screen stitching effects implemented by the present invention are good and the detection and tracking task accuracy is high, which is a great improvement over the prior art.

[0197] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0198] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions.

[0199] It should be noted that in the claims, any reference numerals placed between brackets shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention may be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In the claims enumerating several means, several of these means may be embodied by the same hardware. The use of the words first, second, third, etc., is for convenience of expression only and does not indicate any order. These words may be understood as part of the component name.

[0200] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0201] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments after knowing the basic creative concept. Therefore, the claims should be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.

[0202] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention should also include these modifications and variations.

Claims

1. A multi-image stitching method for target tracking and recognition, characterized in that: include: Perform feature extraction on the original images acquired by multiple cameras; According to the feature extraction result of each original image, the position information of each camera is obtained; Obtaining image transformation parameters of each camera according to the posture information; Performing image transformation on the plurality of original images according to corresponding image transformation parameters to obtain a plurality of first-type transformed images; After reducing the resolution of the plurality of original images by a predetermined multiple, image transformation is performed according to the corresponding image transformation parameters to obtain a plurality of second-type transformation images, and image stitching preview is performed on the plurality of second-type transformation images to obtain stitching preview parameters; wherein, performing image stitching preview on the plurality of second-type transformation images to obtain stitching preview parameters comprises: performing image stitching preview on the second-type transformation images according to the posture information of each camera to obtain a stitching preview image; determining the seam line in the overlapping area of ​​each image in the stitching preview image by the maximum flow graph cut method; the seam line is a line connecting a number of pixels in the overlapping area that meet a preset similarity; taking the seam line as a reference, only pixels of the adjacent left image are retained on the left side of the seam line, and only pixels of the adjacent right image are retained on the right side of the seam line; amplifying the resolution of the overlapping area after the seam line optimization by the predetermined multiple; assigning a corresponding gain coefficient to each part of the stitching preview image by an error function to make the image intensity of the overlapping area equal or similar; obtaining the stitching preview parameters according to the overlapping area amplified by the predetermined multiple and the gain coefficient; According to the splicing preview parameters, multiple first-type transformed images are spliced ​​to obtain a spliced ​​image.

2. A multi-image stitching method for target tracking and recognition as claimed in claim 1, characterized in that: N feature sub-images are extracted from each of the original images, and each feature sub-image includes feature point coordinates and a feature vector.

3. A multi-image stitching method for target tracking and recognition as claimed in claim 2, characterized in that: According to the feature extraction results of each original image, the pose information of each camera is obtained, including: By calculating the cosine distance between the feature vectors of each two original images, the matching degree of each two original images is obtained; Determining the adjacent relationship between each camera according to the matching degree; According to the coordinates of the feature points, the initial intrinsic parameter matrix and the initial extrinsic parameter matrix of each camera are obtained by using the internal and external parameter formulas; The initial matrix of internal parameters and the initial matrix of external parameters are optimized by the bundle adjustment method to obtain the internal parameter matrix and the external parameter matrix of each camera; Obtaining posture information of each camera based on the intrinsic parameter matrix and the extrinsic parameter matrix; Obtaining the position information of each camera according to the posture information of each camera and the adjacent relationship between each camera; Among them, for each group of adjacent cameras, the internal and external parameter formulas are: In formula (1), K1 and K2 are the internal parameter matrices of any group of adjacent cameras 1 and 2 to be solved, R1 and R2 are the external parameter matrices of any group of adjacent cameras 1 and 2 to be solved, and x 2n+1 and x 2n+2 is a known set of matching feature point coordinates.

4. A multi-image stitching method for target tracking and recognition as claimed in claim 3, characterized in that: The image transformation parameters of each camera obtained according to the posture information include: Solving the remapping matrix corresponding to the correction transformation of each camera based on the intrinsic parameter matrix and the extrinsic parameter matrix; Based on the extrinsic matrix of each camera, by solving the vertex coordinates of each original image, the size of the mosaic image and the translation amount required for each original image to be translated to the corresponding position in the mosaic image are obtained; Obtaining image transformation parameters according to the remapping matrix, the image stitching scale and the translation amount; Wherein, the remapping matrix is: In formula (2), u and v are the transformed pixel coordinates, and x and y are the corresponding pixel coordinates in the original image, x = sin(π-v)·sin(u), y = cos(π-v), z = sin(π-v)·cos(u); p = sin(π-v)·sin(u), q = cos(π-v), r = sin(π-v)·cos(u).

5. The multi-image stitching method for target tracking and recognition according to claim 1, characterized in that: The error function is: In formula (3), g i and g j is the gain coefficient of image i and image j, R(i,j) represents the overlapping area of ​​image i and image j, I i (u i ) represents the average intensity I of image i in the overlapping area R(i,j) ij ;u i and u j They respectively represent the same point on the overlapping area R(i,j) corresponding to the point position in the images of image i and image j; In formula (4), R, G and B represent the intensity values ​​of the red, green and blue components of the color image respectively, and N ij Represents the number of pixels in the overlapping area R(i,j); The empirical formula of the error function is: In formula (5), I ij represents the average intensity of image i in the overlapping area of ​​image i and image j, I ji represents the average intensity of image j in the overlapping area of ​​image i and image j; σ N and σ g denote the standard deviation of error and gain, σ N =10,σ g =0.1; The empirical formula for g i The derivative of is: Make formula (6) equal to 0 and expand it into g1, g2, ..., g n The equation for the variable is: By taking the derivative of e with respect to all g, a system of equations including n linear equations as shown in equation (7) is established, and n gain coefficients g are obtained by solving the system of equations.

6. The multi-image stitching method for target tracking and recognition according to claim 1, characterized in that: The method further comprises: Dividing the spliced ​​image into multiple sub-images using a pre-trained semantic segmentation model; Performing target detection on the multiple sub-images using a pre-trained target detection model, and mapping the detected target coordinates back to the spliced ​​image; Each target is tracked individually according to the mapped target coordinates, and the attributes of each tracked target are recognized based on deep learning.

7. A multi-image stitching method for target tracking and recognition as claimed in claim 6, characterized in that: The spliced ​​image is divided into multiple sub-images by a pre-trained semantic segmentation model, including: Detecting the reference area in the spliced ​​image by using a pre-trained semantic segmentation model; The reference area is drawn on a binary image of the same size, and a first area of ​​the binary image is expanded horizontally so that there is only one connected domain in the binary image; Based on the connected domain, the spliced ​​graph is segmented into three initial segmentation areas; The highest point and the lowest point of each of the initially segmented regions are used as segmentation lines to cut the spliced ​​image into three sub-images with overlapping areas.

8. A multi-image stitching system for target tracking and recognition, characterized in that: include: A feature extraction module is used to extract features from the original images acquired by the multiple cameras; A posture determination module, used to obtain the posture information of each camera based on the feature points extracted from each of the original images; An image transformation parameter output module, used to obtain the image transformation parameters of each camera according to the posture information; A first image transformation module, used for performing image transformation on the original images collected by the plurality of cameras according to corresponding image transformation parameters to obtain a plurality of first-type transformed images; A second image transformation module, configured to reduce the resolution of the original images collected by the plurality of cameras by a predetermined multiple, and then perform image transformation according to corresponding image transformation parameters to obtain a plurality of second-type transformed images; A stitching preview module is used to perform image stitching preview on a plurality of the second-type transformed images to obtain stitching preview parameters; wherein, performing image stitching preview on a plurality of the second-type transformed images to obtain stitching preview parameters comprises: performing image stitching preview on the second-type transformed images according to the posture information of each camera to obtain a stitching preview map; determining a seam line in the overlapping area of ​​each image in the stitching preview map by the maximum flow graph cut method; the seam line is a line connecting a number of pixels in the overlapping area that meet a preset similarity; taking the seam line as a reference, only retaining pixels of the adjacent left image on the left side of the seam line, and only retaining pixels of the adjacent right image on the right side of the seam line; amplifying the resolution of the overlapping area after the seam line optimization by the predetermined multiple; assigning a corresponding gain coefficient to each part of the stitching preview map by an error function to make the image intensity of the overlapping area equal or similar; obtaining stitching preview parameters according to the overlapping area amplified by the predetermined multiple and the gain coefficient; An image stitching module is used to stitch a plurality of the first-type transformed images to obtain a stitching image according to the stitching preview parameters.

9. A target tracking and recognition device based on multi-image splicing, characterized in that: include: at least one database; and a memory in communication with the at least one database; Wherein, the memory stores instructions that can be executed by the at least one database, and the instructions are executed by the at least one database so that the at least one database can execute a multi-image stitching method for target tracking and identification as described in any one of claims 1-7.

10. A computer-readable medium having computer-executable instructions stored thereon, characterized in that: When the executable instructions are executed by the processor, a multi-image stitching method for target tracking and identification as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Online deep learning SLAM based image cloud computing method and system

    CN108921893A

  • Unmanned aerial vehicle positioning method and device, computer and storage medium

    CN109974693A