A method for panoramic reconstruction based on multiple perspectives in a high-occlusion multi-target environment
By using multi-view cameras and deep learning technology in a high-occlusion multi-object environment, a virtual coordinate mapping model is established and the image mapping matrix is calculated, which solves the problem of difficulty in identifying target objects under high-occlusion background, and achieves fast and accurate panoramic reconstruction and recognition effects.
Patent Information
- Application Number
- CN202111585932.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-12-20
AI Technical Summary
In a high-occlusion multi-target environment, it is difficult for a single camera to accurately detect the installation location and type of workpieces, and the panoramic artifacts of multi-view images are severe, resulting in difficulty in target recognition.
Multiple cameras that control pitch and rotation angles by the gimbal are adopted to establish a global and local virtual two-dimensional coordinate mapping model based on the physical position distribution relationship of the target object. The image mapping matrix is calculated through deep learning algorithms, SIFT, least squares method and RANSAC algorithms to achieve panoramic reconstruction of the target object.
It realizes the rapid and accurate identification of the number, type, installation position, distribution relationship and background fixed objects in the high occlusion background, reducing the labor intensity of workers.
Smart Images

Figure CN114332758B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image analysis, and particularly relates to a method for panoramic reconstruction based on multiple perspectives in a high-occlusion multi-object environment. Background Art
[0002] Currently, the detection and replacement of workpieces inside many machines rely on manual labor. Some workpieces are installed on mounting plates deep underwater, and there are a large number and various types of workpieces. Most regulatory departments rely on the labor experience of workers to monitor the workpieces regularly. However, using the above method for monitoring not only takes a long time, has low recognition efficiency, and high labor costs, but also cannot effectively ensure whether the placement or replacement of workpieces is correct due to fatigue in human eye detection. Therefore, it is particularly important and meaningful to use modern non-manual methods to quickly and correctly identify the replacement of workpieces to assist workers.
[0003] Currently, there are many pipes of different sizes (also known as background fixtures or occluders) on the mounting plates for installing workpieces, which block the workpieces (also known as target objects). Therefore, when using a camera for monitoring, the pipes will block the camera, making it difficult for a single camera to detect the installation positions and types of all workpieces. In addition, in order to better save space, some workpieces are designed in irregular shapes, making it difficult to accurately extract features. Currently, in order to accurately extract the features of target objects and identify their categories, multiple cameras are used to collect multi-perspective images of various highly occluded target objects. However, there are serious artifacts in the panoramic stitching of multi-perspective images under a highly occluded background, resulting in difficulties in identifying the target objects to be detected. Especially after the camera shooting angle changes, how to map the workpiece categories identified by each camera to the global in real time is a rather difficult problem. Summary of the Invention
[0004] The present invention discloses a multi-object panoramic reconstruction method under high occlusion and multiple perspectives, aiming to solve the technical problem of how to map the workpiece (target object) categories identified by each camera to the global in real time mentioned in the background art.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0006] A multi-object panoramic reconstruction method under high occlusion and multiple perspectives includes the following steps:
[0007] Step 1: Above the surrounding and middle positions of the mounting plate for installing the target object, a plurality of cameras that can be controlled by a pan-tilt head for pitch and rotation angles are arranged as needed;
[0008] Step 2: According to the physical position distribution relationship of the actual target objects, different labels are assigned to the positions where each target object is located, and a global virtual two-dimensional coordinate mapping model of the target objects and local virtual two-dimensional coordinate mapping models of each angle are respectively established;
[0009] Step 3: Take images at their respective front-facing angles by each camera as template images for calibration.
[0010] Step 4: Obtain the position coordinates of each target object and the position coordinates of each background fixed object in each template image through a deep learning algorithm.
[0011] Step 5: Calibrate the local virtual two-dimensional coordinate mapping models for each angle according to the topological distribution relationship between the position coordinates of the target objects and the background fixed objects in each template image, so that the local virtual two-dimensional coordinate mapping models for each angle are aligned with the center of the target object installation position in their respective corresponding template images.
[0012] Step 6: Adjust the shooting angles of each camera through the pan-tilt head, take images at different angles, and obtain target images for each angle.
[0013] Step 7: Calculate the mapping matrix between the template image and the target image through the SIFT, least squares method, and RANSAC algorithms, and map the calibrated local virtual two-dimensional coordinate mapping models to the target images at each angle through the mapping matrix; then identify the position coordinates and categories of each target object in the target image through a deep learning algorithm.
[0014] The full English name of the SIFT algorithm is: scale-invariant feature transform; the full Chinese name is: Scale Invariant Feature Transform Algorithm.
[0015] The full English name of the RANSAC algorithm is: random sample consensus; the full Chinese name is: Random Sample Consensus.
[0016] Step 8: According to the difference between the Euclidean distance of the position coordinates of the target objects in each target image and the point coordinates of the local virtual two-dimensional coordinate mapping model under each target image, exclude the points with Euclidean distance greater than the set threshold, screen out the labels and categories corresponding to the target images identified by deep learning, traverse the coordinate nodes of the target objects in the global virtual two-dimensional coordinate mapping model, and output the categories of the target objects in sequence, so that the categories of the target objects are mapped to the global virtual two-dimensional coordinate mapping model.
[0017] The present invention aims at the problem of panoramic map reconstruction for multiple objects under a highly occluded background. First, according to the distribution relationship of the installation positions corresponding to the objects, a global and local virtual two-dimensional coordinate mapping model of the objects is manually established; then, under the occluded background, multi-angle shooting is used to shoot template images, and according to the control points, rotation factors, rotation correction factors, and scale factors calculated from the template images, the local virtual two-dimensional coordinate mapping model is aligned with the center of the installation position of the object in the template and mapped to the global virtual two-dimensional coordinate mapping model; when re-identification is required, the shooting angle is changed at multiple angles to shoot local target images, and the mapping matrix between the local target images and the template images is obtained, and the local virtual two-dimensional coordinate mapping model aligned with the center of the installation position of the object in the template is mapped to the local target images; then, deep learning is combined to identify the positions and categories of the objects in each local target image; finally, an Euclidean distance judgment is made based on the position coordinates of the objects in the target images and the local virtual two-dimensional coordinate mapping model after alignment with the center of the installation position of the object under the target images, the points with Euclidean distance greater than the set threshold are excluded, the position labels and categories corresponding to the objects predicted by deep learning are selected, and the category of each object label node is output in turn to map it to the global virtual two-dimensional coordinate mapping model to achieve panoramic reconstruction of the objects. Finally, the purpose of quickly and accurately identifying the total number of all objects, the types of objects, the installation positions corresponding to the objects, the distribution relationship of the objects, and the distribution relationship of the background fixed objects is achieved.
[0018] Preferably, the global virtual two-dimensional coordinate mapping model described in step 2 is generated according to the distribution of the installation positions corresponding to the real objects; and the installation position corresponding to each object is drawn with a hexagon, and the objects are rendered by traversing multiple times, and each installation position corresponds to a label.
[0019] Preferably, in step 5, according to the deep learning algorithm, the control points, rotation factors, rotation correction factors, and scale factors are calculated based on the position coordinates of the objects in the template images at each angle, so that the local virtual two-dimensional coordinate mapping models at each angle are aligned with the centers of the installation positions of the objects in the template images.
[0020] Preferably, the determination of the control points, rotation factors, rotation correction factors, and scale factors in step 5 is as follows:
[0021] Determination of control points: According to the distribution law of the objects, the camera installation method is fixed to find the control points;
[0022] Determination of rotation factor: The rotation factor is the rotation angle between the coordinate system of each local virtual two-dimensional coordinate mapping model and the camera pixel coordinate system;
[0023] Determination of rotation correction factor: Based on the coordinate positions of the target objects recognized by the deep learning algorithm, a fitting line for multiple points is determined, the rotation angle of the line is calculated, and the rotation correction factor is determined;
[0024] Determination of scale factor: According to the position coordinates of the target objects around the control points, the average Euclidean distance between the position coordinates of the target objects is calculated, and the average value is taken as the scale factor.
[0025] Preferably, step 5 includes the following steps:
[0026] Step 5.1: Based on the position coordinates of each target object in each angle template map determined in step 4, denoted as set B; compare the y pixel coordinates of the target objects in each template map, obtain the set of target points with the top three y pixel coordinates, locate the first row of the target objects, and define the set as C; and compare the x pixels of each target point within set C, obtain the point with the largest x pixel coordinate, locate the rightmost point in the first row, and take this point as the control point, and define this control point as Y(x0, y0);
[0027] Step 5.2: The rotation factor is the rotation angle between the coordinate system of each local virtual two-dimensional coordinate mapping model and the camera pixel coordinate system, denoted as θ1 (generally, the angle of θ1 is taken as follows: if it is a local virtual two-dimensional coordinate mapping model with upper, lower, and middle angles, the angle rotation factor is taken as 60 degrees; if it is a local virtual two-dimensional coordinate mapping model with left and right angles, the angle rotation factor is taken as 30 degrees); fit a line through set C, and calculate the rotation angle of this line, and determine the rotation angle of this line as the rotation correction factor, denoted as θ;
[0028] Step 5.3: Taking the y coordinate y0 of Y as a reference, search for the set of coordinate points in set B that meet the y pixel coordinates in the range from y0 - d max to y0, denoted as set G(x k , y k ), calculate the Euclidean distance between the coordinates of adjacent two target objects through set G. If the calculated Euclidean distance is within the preset threshold d min to d max , then add this Euclidean distance to set Q, and take the average value of all Euclidean distances in set Q as the scale R1;
[0029] Step 5.4: For the template map of the middle angle, based on the position coordinate points of the background fixed objects determined in step 4, use the method of permutation and combination to randomly select four points, fit a line, and then calculate whether the distances from the four points to this line are greater than the preset threshold d e ; if there is one point greater than the preset threshold d e, then exclude this combination. If it is less than or equal to, randomly select any one of these four points and calculate the Euclidean distance between it and other points respectively;
[0030] Step 5.5: Sort the calculation results in Step 5.4 in descending order. If the first value in the descending order is within the preset threshold range and is twice the second value, then it is determined as set L. Take the point a with the smallest sum of x and y pixel coordinates within set L, and take the point b with the largest sum of x and y pixel coordinates. Using a and b as the corner points of a rectangle, obtain the locked area K of the image;
[0031] Step 5.6: Based on the position coordinates of the target object determined in Step 4, denoted as set P, search for set M within set P that belongs to area K; traverse set M and find the coordinate point whose six surrounding coordinate points belong to the coordinate points within set P, and denote this coordinate point as set O;
[0032] Step 5.7: Calculate the sum of the x and y coordinates of each coordinate point within set O, and find the coordinate point with the middle sum of x and y coordinates, and use this coordinate point as the control point; then, according to the method of determining the scale factor, rotation factor, and rotation correction factor for the upper, lower, left, and right angles, determine the scale factor, rotation factor, and rotation correction factor of the local virtual two-dimensional coordinate mapping model for the middle angle;
[0033] Step 5.8: Through the calculated control point, rotation factor, and scale factor, align the local virtual two-dimensional coordinate mapping model γ of each angle with the center of the installation position of the target object in the template map of each angle, and obtain the calibrated local virtual two-dimensional coordinate mapping model γ”(X Z , Y Z ); Denote the set of coordinate points in the calibrated local virtual two-dimensional coordinate mapping model of each angle as set U. Set U contains the topological distribution relationship of each coordinate point and the x and y pixel coordinates of each coordinate point; and the Euclidean distance between each coordinate point in the calibrated local virtual two-dimensional coordinate mapping model and the center coordinate point of the target object in the template map is not greater than the threshold d error
[0034] According to the distribution positions of the objects on the mounting plate, a local virtual two-dimensional coordinate mapping model is established at each angle. The labels and topological relationships of the local virtual two-dimensional coordinate mapping model are established based on the object mounting position labels and topological relationships in each angle template diagram. Based on the mounting method of the fixed camera, as well as the control points, rotation factors, and scales identified and analyzed in step 5 for each angle, the local virtual two-dimensional coordinate mapping model is calibrated. The coordinate point set in the calibrated local virtual two-dimensional coordinate mapping model is denoted as set U, and the Euclidean distance between each coordinate point in the calibrated local virtual two-dimensional coordinate mapping model and the central coordinate point of the object in the template diagram is not greater than the threshold d error , realizing the alignment of the local virtual two-dimensional coordinate mapping model with the centers of the objects in each template diagram.
[0035] Preferably, step 7 includes the following steps:
[0036] Step 7.1: Randomly change the camera shooting angle, shoot the target diagrams at each angle, and extract the feature point pairs in each angle template diagram and the target diagram through SIFT;
[0037] Step 7.2: Obtain the mapping matrix through the least squares method, then perform selection through the RANSAC algorithm to obtain the optimized mapping matrix, and then map set U to the target diagram through the optimized mapping matrix, denoted as set T.
[0038] Preferably, step 8 includes the following steps:
[0039] Step 8.1: Based on the object position coordinates determined in step 7, identify the central coordinates of the objects in each angle target diagram, defined as set S;
[0040] Step 8.2: By sequentially extracting the points in set T, traversing the points in set S, and making an Euclidean distance judgment based on the corresponding position coordinates under the same global label of set T and set S; if the Euclidean distance is lower than d, it is determined that the point is found and the category of the object is assigned; if they are all greater than the preset threshold, the point is skipped and it is determined that it is not recognized, obtaining the label category set.
[0041] Furthermore, it also includes color rendering for the label categories, with different labels corresponding to different colorings. By rendering the label categories and different labels corresponding to different colors, the corresponding staff can intuitively see the distribution of the label categories, that is, the distribution of the objects.
[0042] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows: The present invention captures template images at various angles, and based on the distribution relationship of the installation positions corresponding to the target objects, as well as the control points, rotation factors, rotation correction factors, and scale factors calculated from the images, establishes a global virtual two-dimensional coordinate mapping model and a local virtual two-dimensional coordinate mapping model of the target objects. When performing recognition, the shooting angle is randomly changed at various angles to capture the target images, and the mapping matrix between the target images and the template images is obtained. The local and global virtual two-dimensional coordinate mapping models of the template images are mapped to the target images, and combined with the results of deep learning recognition, finally, the total number of target objects, the types of target objects, the installation positions corresponding to the target objects, the distribution relationship of the target objects, and the distribution relationship of the background fixed objects are quickly and accurately recognized. And different label categories are rendered, with different colors corresponding to different label categories, enabling the corresponding staff to intuitively see the distribution of the label categories. Its advantages lie in that in actual production or life, it not only provides a modern scientific method for multi-target panoramic reconstruction with high occlusion and multiple perspectives, but also realizes the rapid and accurate recognition of similar working condition scenarios, achieving the purpose of reducing the labor intensity of workers. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The present invention will be described by way of examples with reference to the accompanying drawings, where:
[0044] Figure 1 is a schematic block diagram of the camera installation position of the present invention;
[0045] Figure 2 is a flow block diagram of the present invention;
[0046] Figure 3 is the category of workpieces (target objects) in the embodiment of the present invention;
[0047] Figure 4 is a diagram of the global virtual two-dimensional coordinate mapping model of the present invention;
[0048] Figure 5 are template images of the present invention at various angles;
[0049] Figure 6 are target images of the present invention at various angles;
[0050] Figure 7 is a diagram of the category recognition result of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are only some of the embodiments of this application, rather than all of them. The components of the embodiments of this application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents the selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts belong to the scope of protection of this application.
[0052] See the attached Figure 1 and Figure 2 As shown, the optimal embodiment of the present invention is described as follows:
[0053] A multi-object panoramic reconstruction method under high occlusion and multiple perspectives includes the following steps:
[0054] Step 1: A camera controlled by a pan-tilt head for pitch and rotation angles is respectively arranged above the four-week and middle positions of the mounting disk for mounting the target object; see the attached Figure 1 As shown, the present invention arranges a camera at intervals of 90 degrees in a circle above the mounting disk for mounting the target object, and arranges a camera directly above the center of the mounting disk, for a total of five cameras.
[0055] Step 2: According to the physical position distribution relationship of the actual target objects, different labels are assigned to the positions where each target object is located, and a global virtual two-dimensional coordinate mapping model of the target objects and local virtual two-dimensional coordinate mapping models of each angle are respectively established;
[0056] The global virtual two-dimensional coordinate mapping model is compiled according to the installation position distribution of the real target objects; and the installation positions corresponding to each target object are drawn in hexagons, and the target objects are rendered by traversing multiple times, and each installation position corresponds to a label.
[0057] Step 3: The images taken by each camera at its respective front view angle are used as template images;
[0058] Step 4: The position coordinates of each target object and the position coordinates of each background fixed object in each template image are obtained through a deep learning algorithm;
[0059] Step 5: According to the topological distribution relationship between the position coordinates of the target objects and the background fixed objects in each template image, the local virtual two-dimensional coordinate mapping models of each angle are calibrated so that the local virtual two-dimensional coordinate mapping models of each angle are centered on the installation positions of the target objects in their respective corresponding template images;
[0060] Calculate the control points, rotation factors, rotation correction factors, and scale factors based on the position coordinates of the objects in the template images at each angle identified by the deep learning algorithm, so that the local virtual two-dimensional coordinate mapping models at each angle are aligned with the installation position centers of the objects in the template images.
[0061] The determination of the control points, rotation factors, rotation correction factors, and scale factors is as described below:
[0062] Determination of control points: According to the distribution law of the objects, fix the camera installation method and find the control points;
[0063] Determination of rotation factors: The rotation factor is the rotation angle between the coordinate systems of the local virtual two-dimensional coordinate mapping models and the camera pixel coordinate system;
[0064] Determination of rotation correction factors: According to the coordinate positions of the objects identified by deep learning, determine the fitting lines of multiple points, calculate the rotation angle of the lines, and determine the rotation correction factors;
[0065] Determination of scale factors: According to the position coordinates of the objects around the control points, calculate the average Euclidean distance between the position coordinates of the objects, and take the average value as the scale factor.
[0066] Step 5 includes the following steps:
[0067] Step 5.1: Based on the position coordinates of each object in the template images at each angle determined in Step 4, denoted as set B; compare the y pixel coordinates of the objects in each template image, obtain the set of target points with the top three y pixel coordinates, locate the first row of the objects, define the set as C; and compare the x pixels of each target point within set C, obtain the point with the largest x pixel coordinate, locate the rightmost point in the first row, and take this point as the control point, define this control point as Y(x0, y0);
[0068] Step 5.2: The rotation factor is the rotation angle between the coordinate systems of the local virtual two-dimensional coordinate mapping models and the camera pixel coordinate system, denoted as θ1 (generally, the angle of θ1 is taken as follows: if it is the local virtual two-dimensional coordinate mapping model of the upper, lower, and middle angles, the angle rotation factor is taken as 60 degrees; if it is the local virtual two-dimensional coordinate mapping model of the left and right angles, the angle rotation factor is taken as 30 degrees); fit a line through set C and calculate the rotation angle of this line, determine the rotation angle of this line as the rotation correction factor, denoted as θ;
[0069] Step 5.3: Taking the y coordinate y0 of Y as the benchmark, search for the set of coordinate points in set B that meet the y pixel coordinate in the range from y0 - d max to y0, denoted as set G(x k , yk ) Calculate the Euclidean distance between the coordinates of two adjacent objects through set G. If the calculated Euclidean distance is within the preset threshold d min to d max , then add this Euclidean distance to set Q, and take the average value of all Euclidean distances in set Q as the scale R1;
[0070] Step 5.4: For the template image of the intermediate angle, based on the position coordinate points of the background fixed objects determined in Step 4, randomly select four points using the method of permutation and combination, fit a straight line, and then calculate whether the distances from the four points to this straight line are greater than the preset threshold d e ; If one point is greater than the preset threshold d e , then exclude this combination. If it is less than or equal, randomly select any one of the four points and calculate the Euclidean distances between it and the other points respectively;
[0071] Step 5.5: Sort the calculation results in Step 5.4 in descending order. If the first value in the descending order is within the preset threshold range and is twice or three times the second value, then it is determined as set L. Take the point a with the smallest sum of x pixel coordinates and y pixel coordinates in set L, and take the point b with the largest sum of x pixel coordinates and y pixel coordinates. Using a and b as the corner points of a rectangle, obtain the locked area K of the image;
[0072] Step 5.6: Based on the position coordinates of the objects determined in Step 4, denoted as set P, search for set M in set P that belongs to region K; traverse set M, and find the coordinate points in set M whose six surrounding coordinate points belong to the coordinate points in set P. Denote this coordinate point as set O;
[0073] Step 5.7: Calculate the sum of the x and y coordinates of each coordinate point in set O, and find the coordinate point with the middle sum of x and y coordinates. Take this coordinate point as the control point; then, according to the method of determining the scale factor, rotation factor, and rotation correction factor for the upper, lower, left, and right angles, determine the scale factor, rotation factor, and rotation correction factor of the local virtual two-dimensional coordinate mapping model for the intermediate angle;
[0074] Step 5.8: Align each angle local virtual two-dimensional coordinate mapping model γ with the installation position center of the object in each angle template image through the calculated control point, rotation factor, and scale factor, and obtain the calibrated local virtual two-dimensional coordinate mapping model γ”(X Z ,Y Z);Denote the set of coordinate points in the calibrated local virtual two-dimensional coordinate mapping model at each angle as set U. Set U contains the topological distribution relationships of each coordinate point and the x and y pixel coordinates of each coordinate point; and the distance error between each coordinate point in the calibrated local virtual two-dimensional coordinate mapping model and the central coordinate point of the target object in the template image is not greater than the threshold d error
[0075] In view of the distribution positions of the target objects on the mounting plate, the present invention establishes local virtual two-dimensional coordinate mapping models at each angle. The labels and topological relationships of the local virtual two-dimensional coordinate mapping models are established according to the installation position labels and topological relationships of the target objects in the template images at each angle; based on the installation method of the fixed camera and the control points, rotation factors and scales identified and analyzed in step 5 at each angle, calibrate the local virtual two-dimensional coordinate mapping models. Denote the set of coordinate points in the calibrated local virtual two-dimensional coordinate mapping model as set U, and the Euclidean distance between each coordinate point in the calibrated local virtual two-dimensional coordinate mapping model and the central coordinate point of the target object in the template image is not greater than the threshold d error , realizing the alignment of the local virtual two-dimensional coordinate mapping model with the central points of each target object in each template image.
[0076] Step 6: Adjust the shooting angles of each camera through the pan-tilt head, shoot images at different angles, and obtain target images at each angle;
[0077] Step 7: Calculate the mapping matrix between the template image and the target image through SIFT, the least squares method, and RANSAC, and through the mapping matrix, map the calibrated local virtual two-dimensional coordinate mapping models at each angle to the target images at each angle; then identify the position coordinates and categories of each target object in the target images through deep learning;
[0078] The said step 7 includes the following steps:
[0079] Step 7.1: Randomly change the shooting angle of the camera, shoot target images at each angle, and extract the feature point pairs in the template images and target images at each angle through the SIFT algorithm;
[0080] Step 7.2: Obtain the mapping matrix through the least squares method, then perform refinement through the RANSAC algorithm to obtain the optimized mapping matrix, and then map set U to the target image through the optimized mapping matrix, denoted as set T.
[0081] Step 8: According to the difference between the position coordinates of the target object in each target image and the Euclidean distance of each point coordinate in the local virtual two-dimensional coordinate mapping model under each target image, exclude the points whose Euclidean distance is greater than the set threshold, screen out the label and category corresponding to the target image identified by deep learning, traverse the coordinate nodes of the target object in the global virtual two-dimensional coordinate mapping model, and output the category of the target object in sequence, so that the category of the target object is mapped to the global virtual two-dimensional coordinate mapping model.
[0082] The said step 8 includes the following steps:
[0083] Step 8.1: Based on the position coordinates of the target object determined in step 7, identify the center coordinates of the target object in each target image at each angle, and define it as set S;
[0084] Step 8.2: By sequentially extracting the points in set T, traverse the points in set S, and make a Euclidean distance judgment according to the corresponding position coordinates under the same global label of set T and set S; if the Euclidean distance is lower than d, it is determined that the point is found, and the category of the target object is assigned; if they are all greater than the preset threshold, skip the point and determine that it is not recognized, and obtain the label category set.
[0085] It also includes rendering the label category, and different labels correspond to different colors. By rendering the label category and different labels corresponding to different colors, the corresponding staff can intuitively see the distribution of the label categories, that is, the distribution of the target objects.
[0086] The present invention establishes a global virtual two-dimensional coordinate mapping model and a local virtual two-dimensional coordinate mapping model of the target object by taking template images at various angles and according to the distribution relationship of the installation positions corresponding to the target object and the control points in the image; when performing recognition, randomly change the shooting angle at each angle to shoot the target image, calculate the mapping matrix between the target image and the template image, map the template local virtual two-dimensional coordinate mapping model to the target image, and combine the recognition results of deep learning to finally quickly and accurately identify the total number of target objects, the types of target objects, the installation positions corresponding to the target objects, the distribution relationship of the target objects, and the distribution relationship of the background fixed objects.
[0087] See Appendix Figure 1 to Appendix Figure 7 , and the embodiments of the present invention will be further described by way of example below:
[0088] A method for panoramic reconstruction based on multiple perspectives in a high-occlusion multi-target environment of the present invention is implemented based on the following device:
[0089] See Appendix Figure 1 As shown, it includes five cameras arranged above, below, left, right, and in the middle of the workpiece disk for taking images; Figure 1The camera shown (in the middle) is located directly above the center of the workpiece disk in a three-dimensional relationship, and the camera faces the workpiece disk directly;
[0090] Monitor: Used to display the mapping results of workpiece recognition;
[0091] Workpiece recognition system: Real-time display of the monitoring screen, adjust the camera angle to cover the workpiece within the maximum range and capture the template image, arbitrarily adjust the angles of each camera to capture the second mapping image, and generate the workpiece recognition result map;
[0092] The said workpiece is denoted as the target object;
[0093] See Appendix Figure 3 As shown, there are a total of five different workpieces, namely A, B, C, D, and E, and the shape of each workpiece is different.
[0094] The mapping principle of workpiece recognition is as follows:
[0095] Based on the physical distribution relationship of the workpiece pits (i.e., the installation positions of the workpieces on the workpiece disk, corresponding to the installation positions of the target objects) at various angles, taking the rightmost workpiece pit in the first row at the upper, lower, left, and right angles as the control points [the control points at the intermediate angles are determined according to the distribution relationship of the small pipes (i.e., the corresponding fixed background objects)], establish a virtual global and local virtual two-dimensional coordinate mapping model (see Appendix Figure 4 As shown); by capturing the template images and target images at each angle, and performing registration, calculate the mapping matrix at each angle, map the local virtual two-dimensional coordinate mapping model to the target image, and finally obtain the recognition category of the workpiece through the deep learning recognition result, and finally generate a schematic diagram of the category recognition confidence and a category recognition result map.
[0096] The specific process of workpiece recognition mapping is as follows:
[0097] Set four pan-tilt heads around the workpiece observation port, and arrange an industrial camera on each pan-tilt head. The rotation and pitch angles of the camera can be adjusted through the pan-tilt head. Arrange a camera without a pan-tilt head in the middle of the observation port, and its rotation and pitch angles are fixed (if necessary, a camera with a pan-tilt head can also be arranged in the middle of the observation port for use in special cases).
[0098] According to the physical arrangement relationship of the workpiece placement positions on the workpiece tray, a global virtual two-dimensional coordinate mapping model is established using the workpiece pit simulation diagram. Each pit is drawn as a hexagon. After traversing 313 times, the workpiece installation positions are rendered. Each pit corresponds to a label. From top to bottom, the labels of the first row to the last row are A3 - A9, B2 - B11, C1 - C13, D1 - D14, E1 - E15, F1 - F16, G1 - G17, H1 - H18, I1 - I19, J2 - J19, K2 - K20, L3 - L20, M3 - M21, N4 - N21, P5 - P21, Q6 - Q21, R7 - R21, S8 - S21, T9 - T21, U11 - U20, V13 - V19, a total of 313 pits and labels. See Table 1 for details. Then, based on the camera installation position and the global two-dimensional virtual coordinate mapping model, a local virtual two-dimensional coordinate mapping model γ is designed (in Figure 3 in, at the upper, lower, left, right, middle, and angular positions, the centers of the regions named A3, V19, M3, I19, and K11 are used as the virtual model coordinate origin (0, 0);
[0099] For the upper angle, the line connecting the centers of the regions labeled A3 and B3 is used as the x-axis, and the line connecting the centers of the regions named A3 and B5 is used as the y-axis. The same name as the global coordinate model is given to each workpiece installation position name, and the coordinates increase proportionally;
[0100] For the lower angle, the line connecting the centers of the regions labeled V19 and U19 is used as the x-axis, and the line connecting the centers of the regions named V19 and U17 is used as the y-axis. The same name as the global coordinate model is given to each workpiece installation position name, and the coordinates increase proportionally;
[0101] For the left angle, the line connecting the centers of the regions labeled M3 and N4 is used as the x-axis, and the line connecting the centers of the regions labeled M3 and L4 is used as the y-axis. The same label as the global coordinate model is given to each workpiece installation position name, and the coordinates increase proportionally;
[0102] For the right angle, the line connecting the centers of the regions labeled I19 and H18 is used as the x-axis, and the line connecting the centers of the regions labeled I19 and J18 is used as the y-axis. The same label as the global coordinate model is given to each workpiece installation position name, and the coordinates increase proportionally;
[0103] For the middle angle, the line connecting the centers of the regions labeled K11 and J11 is used as the x-axis, and the line connecting the centers of the regions labeled K11 and J9 is used as the y-axis. The same label as the global coordinate model is given to each workpiece installation position name, and the coordinates increase proportionally).
[0104] Adjust the pan-tilt head and take a template image so that the camera can capture most of the workpieces at various angles, ensuring that the workpieces in the first row are within the image at the four angles of up, down, left, and right, and the slope of the straight line formed by the center coordinates of the workpieces in the first row does not exceed 15 degrees; see the appendix Figure 5 As shown, for the image captured by the camera in the middle, it is necessary to ensure that the small tubes are within the image, and save the captured image as a template image.
[0105] See the appendix Figure 5 As shown, the images captured by the five cameras of up, down, left, right, and middle are used as template images. Based on the illustration, it can be clearly seen that there are multiple small tubes in the image; and to make the small tubes within the image captured by the middle camera, only the installation position of the middle camera needs to be adjusted accordingly during the installation and debugging phase.
[0106] Detect the workpieces in the template images at the four angles of up, down, left, and right through deep learning, and obtain their center coordinate set B(x i ,y i ) and categories. For the middle angle, obtain the center position coordinates and categories of the workpieces and small pipes; the small pipes can be seen in the appendix Figure 5 , appendix Figure 5 The image fed back by the middle camera at the bottom right corner of the appendix. The small pipes are the smaller pipes in the image.
[0107] Traverse the coordinate point set B of the workpiece centers identified in the images at the four angles of up, down, left, and right, compare the y-pixel coordinates of a single image at a single angle, obtain the target point set with the top three y-pixel coordinates, locate the first row of the target object, define the set as C, then traverse the coordinate set C, find the point with the largest x coordinate in the pixel coordinate system and denote it as the control point Y(x0,y0) (Y is the control point that the origin of the virtual two-dimensional coordinate mapping model at each angle needs to align with), and fit the coordinate points in the set C by the least squares method to obtain the straight line:
[0108] y = ax + b (Formula 1);
[0109] By:
[0110] θ = arctan(α) (Formula 2);
[0111] Obtain the straight line rotation angle θ, and θ is the rotation correction factor of the local virtual two-dimensional coordinate mapping model around.
[0112] Taking the y-pixel coordinate y0 of Y as the reference, search for the coordinate point set in the coordinate point set B that meets the y-pixel coordinate in the area from y0 - d max to y0 and denote it as G(x k ,y k) The coordinate point set G contains m points. The estimated central distances between the centers of the workpieces within the set G are obtained through the following formula:
[0113]
[0114]
[0115] where 0 < k ≤ m, d max (usually taken as 300) and d min (usually taken as 150) are the number of pixels obtained by converting the maximum and minimum Euclidean distances between the central positions of adjacent real target objects into images under the condition that the camera resolution remains unchanged and the object distance is fixed; the set Q is the set of target object coordinates whose pixel Euclidean distance between target objects conforms to the threshold range from d min to d max ; the scale factor is obtained through the following formula:
[0116]
[0117] where R1 is the scale factor (R1 is the average Euclidean distance between two adjacent target objects in the set Q).
[0118] For the intermediate angles, the categories and central coordinate sets of each workpiece are identified through deep learning, and the central coordinate set F of the small pipes in the entire image is identified through a small pipe model (a small pipe model trained by deep learning). n For the set F n of coordinate points, permutations and combinations are performed on the coordinate points, and all combination methods are obtained through the following formula:
[0119]
[0120] In the formula, n is the number of central coordinates included in F n . For each combination of coordinate point sets P j of four coordinate points (the coordinate point set P j includes ), the straight line is fitted by the least square method as follows:
[0121]
[0122] For any one of the four coordinate points taken arbitrarily, the Euclidean distance differences are made with the remaining coordinate points in turn to obtain the set D m (D m includes the Euclidean distance between a and b, the Euclidean distance between a and c, and the Euclidean distance between a and d), and then the small pipe coordinate set is screened out according to the following formula:
[0123]
[0124] Among them, the set of coordinate points P that satisfies the conditions of formula 8 j That is, the set L, d e (Generally taken as 50) is a parameter for determining whether four coordinates are roughly on a straight line and equally spaced under the condition that the camera resolution remains unchanged and the object distance is fixed.
[0125] Calculate the point with the smallest sum of the x-pixel coordinate and the y-pixel coordinate in the set L as point a, and the point with the largest sum as point b. Then, with point a as the upper left corner and point b as the lower right corner, divide the area K in the image. According to the area K, screen out the set M of the position coordinates of the centers of each workpiece recognized by deep learning within the area.
[0126] According to the set M of the area range, through deep learning, identify the set S of the center coordinates of the target objects in the target images of each angle. Search for the set M that belongs to the area K in the set S. Traverse whether there are six points belonging to the set S around each coordinate point in the set M. Finally, obtain the set O. By comparing the sum of the x and y coordinates of the coordinate points in the set O through three coordinate points, determine the point with the middle size as the control point Y(x0, y0). Then, according to the methods of determining the scale factor, rotation factor, and rotation correction factor for the upper, lower, left, and right angles, determine the scale factor, rotation factor, and rotation correction factor of the mapping model for the middle angle.
[0127] Then, perform scale transformation and translation on the local virtual two-dimensional coordinate mapping model γ of each angle through formulas 9 - 10:
[0128]
[0129] Among them: γ is the local virtual two-dimensional coordinate mapping model, x1...x u is the abscissa of the pit position numbers of each workpiece, y1...y u is the abscissa of the pit position numbers of each workpiece.
[0130]
[0131] Among them, γ' is the local virtual two-dimensional coordinate mapping model obtained after γ undergoes scale transformation and translation.
[0132] Then, rotate γ' clockwise around the control point by θ1 - θ degrees through formulas 11 - 12, and then symmetrically with respect to the x-axis:
[0133]
[0134]
[0135] Among them, (θ1 - θ) is the final rotation angle of the local virtual two-dimensional coordinate mapping model for each angle; where θ is the rotation correction factor; θ1 is the rotation factor. Generally, the angle of θ1 is taken as follows: if it is the local virtual two-dimensional coordinate mapping model for the upper, lower, and middle angles, the angle rotation factor is 60 degrees; if it is the local virtual two-dimensional coordinate mapping model for the left and right angles, the angle rotation factor is 30 degrees; u is the number of workpieces; X Z and Y Z are the x and y coordinates after rotation symmetry of the x and y coordinates of each coordinate point set in γ'.
[0136] Through formula 13 for each combination of X Z and Y Z :
[0137] γ” = (X Z , Y Z ) (Formula 13);
[0138] The local virtual two-dimensional coordinate mapping model γ”(X Z , Y Z ) after calibration for each angle is obtained. The coordinate point set in the model is denoted as set U, and the Euclidean distance between each coordinate point in the calibrated local virtual two-dimensional coordinate mapping model and the central coordinate point of the target object in the template image is not greater than the threshold d error (d error is the value converted to image pixels of the maximum Euclidean distance between each coordinate in the calibrated local virtual two-dimensional coordinate mapping model and the true central coordinate of the target object in the target image under the condition that the camera resolution remains unchanged and the object distance is fixed).
[0139] When the user terminal takes the second photo, the camera is offset by an arbitrary angle to take the target image. The coordinate pairs of the feature points of the template image and the target image are extracted by SIFT. Assuming (x a , y a ) is the feature point of the template image, and (x b , y b ) is the corresponding point on the target image, then there is:
[0140]
[0141] Therefore, to recover the 8 parameters in the transformation matrix, at least 4 pairs of matching feature points are required. The process is as follows:
[0142]
[0143] For the solution of the overdetermined equation like formula 15, first let:
[0144]
[0145] Then solve it by the least squares method. For example:
[0146]
[0147] So far, the H matrix has been obtained. Then, it is refined through Random Sample Consensus (RANSAC) to obtain the optimized H. Next, the set U is mapped to the target image through the mapping matrix H, and the calibrated local virtual two-dimensional coordinate mapping models are corrected to obtain the corrected local virtual mapping coordinate set T (the set T is the set of coordinate points generated by the local virtual two-dimensional coordinate mapping model aligned with the target image).
[0148] The central coordinate set S of the target object in the target images at each angle is identified through deep learning.
[0149] By sequentially extracting T and traversing the points in the set S, and calculating the Euclidean distance. If the Euclidean distance is lower than d (the threshold d is the number of image pixels converted from the Euclidean distance between each coordinate in the corrected local virtual two-dimensional coordinate mapping models at each angle and the position coordinates of the identified target object under the condition that the camera resolution remains unchanged and the object distance is fixed), then it is determined that the point is found and the category of the target object is assigned; if all are greater than d, then the point is skipped and it is determined as not recognized, obtaining a rough labeled category set. According to the same labels at each angle, the points that are recognized multiple times but have different categories and the points with relatively low recognition credibility are removed to obtain the final labeled category set. Finally, the labeled category set is rendered, with different labels corresponding to different colors.
[0150] Table 1 Workpiece Label Table
[0151]
[0152]
[0153] The above-described embodiments merely represent the specific implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the protection scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the technical solution of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application.
Claims
1. A method for panoramic reconstruction based on multiple perspectives in a high-occlusion multi-target environment, characterized in that It includes the following steps: Step 1: Above the perimeter and the middle position of the mounting disc for mounting the target object, arrange multiple cameras whose pitch and rotation angles are controlled by a pan-tilt head as required; Step 2: According to the physical position distribution relationship of the actual target objects, assign different labels to the positions where each target object is located, and respectively establish a global virtual two-dimensional coordinate mapping model of the target objects and local virtual two-dimensional coordinate mapping models for each angle; Step 3: Take the images captured by each camera at its respective front view angle as template images for calibration; Step 4: Use a deep learning algorithm to obtain the position coordinates of each target object and the position coordinates of each background fixed object in each template image; Step 5: According to the topological distribution relationship between the position coordinates of the target objects and the background fixed objects in each template image, calibrate the virtual two-dimensional coordinate mapping models for each angle, so that the local virtual two-dimensional coordinate mapping models for each angle are aligned with the center of the installation position of the target objects in their respective corresponding template images; Step 6: Adjust the shooting angles of each camera through the pan-tilt head, and take images at different angles to obtain target images for each angle; Step 7: Calculate the mapping matrix between the template image and the target image through the SIFT, least squares method, and RANSAC algorithms, and through the mapping matrix, map the calibrated local virtual two-dimensional coordinate mapping models to the target images for each angle; then use a deep learning algorithm to identify the position coordinates and categories of each target object in the target image; Step 8: According to the difference between the Euclidean distance of the position coordinates of the target objects in each target image and the point coordinates of the local virtual two-dimensional coordinate mapping model under each target image, exclude the points whose Euclidean distance is greater than the set threshold, screen out the labels and categories corresponding to the target images identified by deep learning, traverse the coordinate nodes of the target objects in the global virtual two-dimensional coordinate mapping model, and output the categories of the target objects in sequence, so that the categories of the target objects are mapped to the global virtual two-dimensional coordinate mapping model.
2. The method for panoramic reconstruction based on multiple perspectives in a high-occlusion multi-target environment according to claim 1, wherein The global virtual two-dimensional coordinate mapping model described in Step 2 is generated according to the installation position distribution of the actual target objects; and the installation position corresponding to each target object is drawn with a hexagon, and the target objects are rendered by traversing multiple times, and each installation position corresponds to a label.
3. A method for panoramic reconstruction based on multiple perspectives in a high-occlusion multi-target environment according to claim 1, characterized in that, In Step 5, according to the position coordinates of the target objects in the template image for each angle identified by the deep learning algorithm, calculate the control points, rotation factors, rotation correction factors, and scale factors, so that the local virtual two-dimensional coordinate mapping models for each angle are aligned with the center of the installation position of the target objects in the template image.
4. A method for panoramic reconstruction based on multiple perspectives in a high-occlusion multi-target environment according to claim 3, characterized in that, The determination of the coordinate calculation control points, rotation factors, rotation correction factors, and scale factors described in Step 5 is as follows: Determination of the coordinate calculation control points: According to the distribution law of the target objects, fix the camera installation method and find the control points; Determination of the rotation factor: The rotation factor is the rotation angle between the coordinate system of each local virtual two-dimensional coordinate mapping model and the camera pixel coordinate system; Determination of the rotation correction factor: According to the coordinate position of the target object identified by the deep learning algorithm, determine the fitting straight line of multiple points, and calculate the rotation angle of the straight line to determine the rotation correction factor; Determination of the scale factor: According to the position coordinates of the objects around the control points, calculate the average Euclidean distance between the position coordinates of the objects, and take the average value as the scale factor.
5. A method for panoramic reconstruction based on multiple perspectives in a high-occlusion multi-target environment according to claim 1, characterized in that Step 5 includes the following steps: Step 5.1: Based on the position coordinates of each object in each angle template map determined in Step 4, denoted as set B; compare the y-pixel coordinates of the objects in each template map to obtain the set of target points with the top three y-pixel coordinates, locate the first row of the objects, and define the set as C; and compare the x-pixels of each target point within set C to obtain the point with the largest x-pixel coordinate, locate the rightmost point in the first row, and take this point as the control point, denoted as ; Step 5.2: The rotation factor is the rotation angle between the coordinate systems of each local virtual two-dimensional coordinate mapping model and the camera pixel coordinate system, denoted as ; Fit a straight line through set C and calculate the rotation angle of this straight line, and determine the rotation angle of this straight line as the rotation correction factor, denoted as ; Step 5.3: Taking the y-coordinate y0 of Y as the reference, search for the set of coordinate points in set B that meet the y-pixel coordinate in the range of to, denoted as set G, calculate the Euclidean distance between the coordinates of adjacent two objects through set G, if the calculated Euclidean distance is within the preset threshold range of to, then add this Euclidean distance to set Q, and take the average value of all Euclidean distances in set Q as the scale R1; Step 5.4: For the template map of the intermediate angle, based on the position coordinate points of the background fixed objects determined in Step 4, use the method of permutation and combination to randomly select four points, fit a straight line, and then calculate whether the distances of the four points from this straight line are greater than the preset threshold de; if one point is greater than the preset threshold de, then exclude this combination, if it is less than or equal, then randomly select any one of the four points and calculate the Euclidean distance between it and other points respectively; Step 5.5: Sort the calculation results in Step 5.4 in descending order. If the first value in the descending order is within the preset threshold range and is twice the second value, then it is determined as set L. Take the point a with the smallest sum of the x-pixel coordinate and y-pixel coordinate in set L, and take the point b with the largest sum of the x-pixel coordinate and y-pixel coordinate. Using a and b as the corner points of the rectangle, obtain the locked area K of the image; Step 5.6: Based on the position coordinates of the objects determined in Step 4, denoted as set P, search for set M in set P that belongs to area K; traverse set M and find the coordinate point whose six surrounding coordinate points belong to the coordinate points in set P, and denote this coordinate point as set O; Step 5.7: Calculate the sum of the x and y coordinates of each coordinate point in set O, find the coordinate point with the middle sum of the x and y coordinates, and take this coordinate point as the control point; then according to the methods of determining the scale factor, rotation factor and rotation correction factor for the upper, lower, left and right angles, determine the scale factor, rotation factor and rotation correction factor of the local virtual two-dimensional coordinate mapping model for the intermediate angle; Step 5.8: By means of the calculated control points, rotation factors and scale factors, align the local virtual two-dimensional coordinate mapping models at each angle with the center of the installation position of the target object in the template maps at each angle, so as to obtain the calibrated local virtual two-dimensional coordinate mapping models at each angle; Denote the set of coordinate points in the calibrated local virtual two-dimensional coordinate mapping models at each angle as set U, where set U includes the topological distribution relationships of each coordinate point and the x and y pixel coordinates of each coordinate point; and the Euclidean distance between each coordinate point in the calibrated local virtual two-dimensional coordinate mapping model and the central coordinate point of the target object in the template map is not greater than the threshold value.
6. A method for panoramic reconstruction based on multiple perspectives in a high-occlusion multi-target environment according to claim 1, characterized in that, The said step 7 includes the following steps: Step 7.1: Randomly change the camera shooting angle, shoot the target maps at each angle, and extract the feature point pairs in the template maps and target maps at each angle by means of the SIFT algorithm; Step 7.2: Obtain the mapping matrix by means of the least squares method, then perform selection through the RANSAC algorithm to obtain the optimized mapping matrix, and then map set U to the target map by means of the optimized mapping matrix, denoted as set T.
7. A method for panoramic reconstruction based on multiple perspectives in a high-occlusion multi-target environment according to claim 6, characterized in that, The said step 8 includes the following steps: Step 8.1: Based on the position coordinates of the target object determined in step 7, identify the central coordinates of the target object in the target maps at each angle, defined as set S; Step 8.2: By sequentially extracting the points in set T, traverse the points in set S, and make a Euclidean distance judgment according to the position coordinates corresponding to each other under the same global label of set T and set S; If the Euclidean distance is lower than d, it is determined that the point is found and the category of the target object is assigned; If they are all greater than the preset threshold value, skip the point and determine that it is not recognized, so as to obtain the label category set.
8. A method for panoramic reconstruction based on multiple perspectives in a high-occlusion multi-target environment according to claim 1, characterized in that It also includes color rendering of the label categories, with different labels corresponding to different colorings.