A target recognition and localization method based on self-switching of multi-view vision system

CN117058236BActive Publication Date: 2026-09-01CHINA UNIV OF MINING & TECH +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202311043208.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-17
Publication Date
2026-09-01
Estimated Expiration
2043-08-17

AI Technical Summary

Technical Problem

该专利通过多组双目相机对人体进行识别定位,无法根据识别定位的结果进行自主的切换以获得更为精准距离信息

Benefits of technology

[0047]本申请的一种基于多目视觉系统自切换的目标识别定位方法,通过多对不同焦距的双目相机系统对目标物体进行识别与定位,同时构建多目视觉系统自主切换逻辑,并建立识别定位结果评价模型,根据切换逻辑与各对相机的评价结果实现多对双目相机的自切换,利用最优的双目相机以获取更为精准的目标定位信息,克服双目定位距离及视野的限制,弥补了单对双目视觉装备固定焦距的限制,克服了低焦距双目相机对远距离目标定位精度低及高焦距双目相机视野小的问题,通过多对双目相机的自主切换,同时满足对近距离目标及远距离目标的精准感知与测距,大幅提高目标识别定位的准确性与效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058236B_ABST
    Figure CN117058236B_ABST
Patent Text Reader

Abstract

This invention discloses a target recognition and localization method based on self-switching of a multi-view vision system. The method uses a target recognition algorithm to identify and perceive the target object, and acquires local depth information based on the binocular imaging principle. Based on this local depth information, three indicators are calculated to evaluate localization accuracy: single-pixel deviation influence factor σ, depth information error rate ε, and depth dispersion ξ. Two evaluation indicators, Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity-to-Simple (SSIM), are obtained based on the local image information of the target. A target localization evaluation model is established to quantify the localization accuracy. According to the autonomous switching logic, the target recognition result and the localization accuracy evaluation score are compared sequentially to achieve autonomous switching between multiple binocular cameras with different focal lengths. This satisfies the requirements for both near-range and far-range target recognition and localization, overcoming the problems of low localization accuracy of low-focal-length binocular cameras for distant targets and small field of view of high-focal-length binocular cameras.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multi-view vision system technology, and more specifically to a target recognition and localization method based on self-switching of multi-view vision systems. Background Technology

[0002] To enhance the autonomy and intelligence of robotic operations and enable tasks such as object recognition, localization, and measurement, acquiring the target's three-dimensional information is crucial. Numerous studies have been conducted on this topic, attracting significant attention from scholars both domestically and internationally. Existing target perception and localization equipment can be categorized into laser, radar, ultrasonic, infrared, and machine vision image ranging. While active ranging equipment such as lasers and radar can provide rapid and accurate non-contact target localization, their high cost and limitations in their working principles prevent them from locating and measuring objects like smoke, water surfaces, and flames. Ultrasonic waves have long propagation distances, strong penetration, and low cost, but their slow propagation speed makes it difficult to guarantee accuracy when measuring distant targets. Infrared ranging offers advantages such as fast propagation speed and rapid response, but it is expensive, has a short range, and unstable data output. Machine vision ranging, on the other hand, acquires surrounding image information using a visible light camera sensor, establishes a three-dimensional relationship model using the camera's imaging principle, and then establishes a transformation from image coordinates to spatial coordinates to obtain the object's spatial distance. Multi-view vision technology can not only autonomously identify targets, but also spatially locate the identified targets. It has a wide range, reliable accuracy, and low cost, and is widely used in target perception and localization.

[0003] In existing target perception and localization methods, acquiring target image information and performing depth calculations using fixed binocular cameras has become a mainstream approach, widely applied in transportation, industry, and emergency rescue. However, due to limitations in binocular ranging principles, low-focal-length, small-baseline binocular cameras have a large field of view. In long-distance ranging, the pixel proportion of the target decreases, increasing the difficulty of identification and significantly increasing localization errors. High-focal-length, large-baseline binocular cameras offer higher accuracy for long-distance target localization, but their smaller field of view means targets outside the field of view may remain undetected. For these reasons, a single pair of binocular cameras cannot simultaneously achieve accurate identification and localization of both near and far-distance targets, significantly limiting its application scope.

[0004] Patent application CN202110163869.8 discloses a multi-view vision measurement method and device for tunnels based on line structured light scanning. It utilizes multiple sets of stereo binocular cameras fixedly installed every 100 meters within the tunnel to track the position of a tunnel vehicle, and combines a platform binocular camera on the tunnel vehicle with a laser scanning probe to achieve three-dimensional information measurement of each tunnel surface segment. This patent uses multiple sets of fixed stereo binocular cameras, achieving sequential positioning and tracking of specific targets on the tunnel vehicle through fixed-distance installation. However, due to the limitation of the fixed binocular camera focal length, each set of cameras can only locate targets within a specific distance. Therefore, positioning moving targets can only be achieved by installing multiple sets of binocular cameras at fixed intervals, significantly increasing the cost of target positioning.

[0005] Patent application CN202210053723.2 discloses a method for robot-assisted visual recognition and localization of flames. It utilizes a calibrated camera to acquire video images, uses a background subtraction method based on a flame color model to identify potential dynamic flame targets, and inputs the fused model into a trained support vector machine (SVM). If the flame is identified, a coarse localization of the suspected flame is calculated based on monocular, binocular, or multi-view vision, guiding the robot to the designated location. This patent proposes a localization method using visible light vision, but it can only provide a rough estimate of the approximate location range, and the localization range is very limited, with accuracy not guaranteed.

[0006] Patent application CN201610255694.2 discloses a fire monitoring and positioning device and method. Two cameras with infrared lenses are placed horizontally and fixed to both ends of a metal rod. The center distance between the lenses of the two cameras is 80-120cm, forming a camera support. Flame recognition and positioning are achieved by determining the occurrence of a fire through a flame recognition algorithm and calculating the location of the ignition point based on the spatial structure of the cameras. This patent uses only a pair of binocular cameras for flame recognition and positioning. Due to the focal length limitation of binocular cameras, it cannot simultaneously satisfy both near-range and long-range recognition and positioning.

[0007] Patent application number 202010441103.7 proposes a method for measuring human body dimensions based on a multi-view vision system. It involves constructing three sets of binocular stereo vision systems to capture color images of the subject, performing semantic segmentation and stereo matching to obtain the spatial coordinates of marker points, projecting all marker points onto the XOZ plane for fitting, and obtaining the dimension of the measured area based on the length of the fitted curve. However, this patent uses multiple sets of binocular cameras for human body identification and positioning, and cannot autonomously switch between them based on the identification and positioning results to obtain more accurate distance information. Summary of the Invention

[0008] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a target recognition and localization method based on self-switching of a multi-view vision system, so as to solve the problems mentioned in the background technology.

[0009] According to one aspect of this application, a target recognition and localization method based on self-switching of a multi-view vision system includes the following steps:

[0010] Step 1: Acquire real-time images using a low-focal-length binocular camera, perform distortion correction on each frame, establish a target detection model, input the distortion-corrected images into the target detection model, and extract the target local image Mat. t arg et ;

[0011] Step 2: Extract the target local image Mat t arg et Then, using the binocular camera imaging principle and the SGBM stereo matching algorithm, the depth information of each pixel is obtained with the left camera coordinate system as the reference. The depth information of each pixel is described by the gray value of each pixel in the disparity map Dep. In the disparity map Dep, the local image Mat of the target to be located is found. t arg et The corresponding target local disparity image Dep t arg et ;

[0012] Step 3: Calculate the single-pixel deviation influence factor σ using the binocular ranging principle;

[0013] Step 4: Obtain the target local disparity image Dep t arg et Then, using the SGBM stereo matching algorithm, the target local disparity image Dep is traversed and calculated. t arg et The depth data of each pixel in the image is used to obtain the depth information error rate ε.

[0014] Step 5: The depth reliability of the target region is represented by calculating the depth dispersion ξ of the target region;

[0015] Step 6: Calculate the target local image information indicators, inputting the target local image Mat. t arg et and target local disparity image Dep t arg et The peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) of the two were obtained respectively.

[0016] Step 7: Calculate the binocular localization evaluation score S using the five indicators obtained: single-pixel deviation influence factor σ, depth information error rate ε, depth dispersion ξ, peak signal-to-noise ratio (PSNR), and structural similarity index (SSIM). i ;

[0017] Step 8: After completing all the switching and acquisition of data from the binocular cameras and obtaining the binocular localization evaluation score S... iThen, based on the visual system's self-switching logic, and according to the number of targets recognized by each pair of binocular cameras and the binocular localization evaluation score S, the system then... i Select the optimal stereo camera and switch to it for subsequent target localization processing.

[0018] Preferably, in step one, the distortion-corrected image is input into the target detection model for target perception. If a target is detected, the number of targets is saved, and the pixel areas of multiple targets are compared. The local image Mat of the target with the largest pixel area is extracted. t arg et If no target is detected, it determines whether all evaluations of the binocular cameras have been completed, and switches to the high-focal-length camera for evaluation or performs autonomous switching of the vision system based on the results.

[0019] Preferably, in step two, the target local disparity image Dep is acquired. t arg et The formula is:

[0020] Dep t arg et .x = Dep.(Mat t arg et Formula (1)

[0021] Dep t arg et .y = Dep.(Mat t arg et .y) Formula (2)

[0022] In equations (1) and (2), Dep t arg et .x, Dep t arg et .y represents Dep t arg et The x and y coordinates of Mat t arg et .x, Mat t arg et .y represents Mat t arg et The x and y coordinates.

[0023] Preferably, in step three, the formula for calculating the single-pixel deviation influence factor σ is:

[0024]

[0025] In equation (4), Z is the distance to be measured, B is the baseline length, f is the focal length, μ is the pixel size of the camera, and σ is the single pixel deviation influence factor, that is, the ranging error corresponding to one pixel.

[0026] The formula for calculating the distance Z to be measured in the formula for determining the single-pixel deviation influence factor σ is as follows:

[0027]

[0028] In equation (3), Z is the distance to be measured, B is the baseline length, f is the focal length, and X is the focal length. r -Xl For parallax.

[0029] Preferably, in step four, the target local disparity image Dep is acquired. t arg et Subsequently, using the SGBM stereo matching algorithm, all stereo matching error points were assigned a fixed value of 10000mm, meaning that the depth information error values ​​within the flame target area were all 10000mm. The local flame depth image (Dep) was then iterated through to obtain the desired depth. t arg et The depth information error rate ε is obtained by taking the depth data of each pixel in the image, and its formula is as follows:

[0030]

[0031] In equation (5), ε is the depth information error rate, and dep error Let m and n be the local disparity images of the target, respectively, where m is the pixel with a depth of 10000mm and n is the pixel with the local disparity of the target. t arg et The magnitudes of the x and y coordinates.

[0032] Preferably, in step five, the formula for calculating the depth dispersion ξ is as follows:

[0033]

[0034] In equation (6), dep(i,j) is the depth value of the pixel at (i,j) within the interval. Let m and n be the average pixel depth values ​​within the interval, and m and n be the corresponding local disparity images of the target, respectively. t arg et The magnitudes of the x and y coordinates.

[0035] Preferably, in step six, the formula for calculating the peak signal-to-noise ratio (PSNR) is as follows:

[0036]

[0037]

[0038] In equations (7) and (8), (i,j) are the pixels to be calculated, MSE is the mean square error, and PSNR describes the similarity between the local image of the target and the corresponding disparity map.

[0039] The formula for calculating structural similarity (SSIM) is as follows:

[0040]

[0041] In equation (9), x is the local image of the target, Mat. t arg et y represents the local disparity image of the target, Dep. t arg et μ x , Let μ be the mean and standard deviation of x. y , Let σ be the mean and standard deviation of y. xy Let c1 be the covariance of x and y, and c1 = (k1L). 2 c2=(k2L) 2 The constants used to maintain stability are L, which is the dynamic range of pixel values, k1 = 0.01, k2 = 0.03.

[0042] Preferably, in step seven, the binocular localization evaluation score S is calculated. i The formula is as follows:

[0043]

[0044] In equation (10), a is a constant term, b, c, d, e, and f are the weights of the five evaluation indicators, k is the number of targets identified, and S i Let be the binocular localization evaluation score for the i-th pair of binocular cameras.

[0045] Preferably, in step eight, it is first determined whether the switching of all pairs of binocular cameras has been completed. If not, the process switches to the next pair of high-focal-length binocular cameras and begins from step one. If the switching of all pairs of binocular cameras has been completed, the process is then based on the vision system's self-switching logic, according to the number of target recognitions and the binocular localization evaluation score S of each pair of binocular cameras. i Select the optimal binocular camera and switch to the optimal binocular camera for subsequent target localization processing;

[0046] The self-switching logic steps of the vision system are as follows: First, compare the number of targets recognized by all binocular cameras. If the number of targets recognized is different, switch to the binocular camera with the most targets recognized, and the self-switching ends. Second, compare the binocular positioning evaluation scores of all binocular cameras. Switch to the binocular camera with the highest binocular positioning evaluation score, and the self-switching ends. When the number of targets recognized and the binocular positioning evaluation score are the same, switch to the low-focal-length camera, and the self-switching ends.

[0047] This application presents a target recognition and localization method based on self-switching of a multi-view vision system. It uses multiple pairs of binocular cameras with different focal lengths to identify and locate target objects. Simultaneously, it constructs an autonomous switching logic for the multi-view vision system and establishes an evaluation model for the recognition and localization results. Based on the switching logic and the evaluation results of each pair of cameras, it achieves self-switching of the multiple pairs of binocular cameras, utilizing the optimal binocular camera to obtain more accurate target localization information. This overcomes the limitations of binocular localization distance and field of view, compensates for the fixed focal length limitation of a single pair of binocular vision equipment, and overcomes the problems of low accuracy in locating distant targets with low-focal-length binocular cameras and small field of view with high-focal-length binocular cameras. Through the autonomous switching of multiple pairs of binocular cameras, it simultaneously satisfies the accurate perception and ranging of both near and far targets, significantly improving the accuracy and efficiency of target recognition and localization. Attached Figure Description

[0048] Figure 1 This is an overall flowchart according to an embodiment of the present invention.

[0049] Figure 2 This is a schematic diagram of the binocular imaging principle according to an embodiment of the present invention.

[0050] Figure 3 This is a schematic diagram of the autonomous switching logic of a multi-view vision system according to an embodiment of the present invention.

[0051] Figure 4 This is a schematic diagram of a multi-pair binocular camera switching device according to an embodiment of the present invention. Detailed Implementation

[0052] To make the content of this application easier to understand, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0053] like Figure 1 As shown, a target recognition and localization method based on self-switching of a multi-view vision system includes the following steps:

[0054] Step 1: First, use a low-focal-length binocular camera to capture real-time images. Correct the distortion of each frame using pre-calibrated intrinsic and extrinsic parameters to obtain a distortion-corrected image. Establish a target detection model and input the distortion-corrected image into the model for target localization. If a target is detected, save the number of targets and compare the pixel areas of multiple targets. Extract the local image (Mat) of the target with the largest pixel area. target If no target is detected, it determines whether all evaluations of the binocular cameras have been completed, and switches to the high-focal-length camera for evaluation or performs autonomous switching of the vision system based on the results.

[0055] Step 2: Extract the target local image Mat targetThen, stereo matching is performed using the imaging principle of a binocular camera. The depth information of each pixel, based on the left camera coordinate system, is calculated using the SGBM stereo matching algorithm. The depth information of each pixel can be described and calculated using the grayscale value of each pixel in the disparity map Dep. The local image Mat of the target to be located is then found within the disparity map Dep. target The corresponding target local disparity image Dep target The formula is as follows:

[0056] Dep target .x = Dep.(Mat target Formula (1)

[0057] Dep target .y = Dep.(Mat target .y) Formula (2)

[0058] In equations (1) and (2), Dep target .x, Dep target .y represents Dep target The x and y coordinates of Mat target .x, Mat target .y represents Mat target The x and y coordinates.

[0059] Step 3: Calculate the single-pixel deviation influence factor σ using the binocular ranging principle. Figure 2 As shown in the diagram illustrating the principle of binocular imaging, the formula for calculating the distance to a point P(X,Y,Z) in space is:

[0060]

[0061] In equation (3), Z is the distance to be measured, B is the baseline length, f is the focal length, and X is the focal length. r -X l For parallax.

[0062] For binocular cameras with different baselines, focal lengths, and distances, based on the camera pixel size and the ranging calculation formula, the ranging error of the binocular camera when the image has a one-pixel error can be obtained as follows:

[0063]

[0064] In equation (4), Z is the distance to be measured, B is the baseline length, f is the focal length, μ is the pixel size of the camera, and σ is the single pixel deviation influence factor, which is the ranging error corresponding to one pixel.

[0065] Step 4: Obtain the target local disparity image Dep targetSubsequently, based on the previous SGBM stereo matching algorithm settings, all stereo matching error points were assigned a fixed value of 10000mm, meaning that the depth information error values ​​within the flame target area were all 10000mm. The local flame depth image Dep was then iterated through to obtain the desired depth. t arg et The depth information error rate ε is obtained by taking the depth data of each pixel in the image, and its formula is as follows:

[0066]

[0067] In equation (5), ε is the depth information error rate, and dep error Let m and n be the local disparity images of the target, respectively, where m is the pixel with a depth of 10000mm and n is the pixel with the local disparity of the target. t arg et The magnitudes of the x and y coordinates.

[0068] Step 5: Besides the majority of the target foreground, the target area also contains a small portion of the ground and other background. The depth information of the target foreground and the background is inconsistent. The depth data of the foreground is roughly within the same range, while the depth data of the background shows a significant deviation. Therefore, by calculating the dispersion of the target area's depth, we can represent the depth reliability of the target area. A large dispersion indicates lower reliability of the target depth information within the recognition box, and vice versa. The formula for the depth dispersion ξ of the target area is as follows:

[0069]

[0070] In equation (6), dep(i,j) is the depth value of the pixel at (i,j) within the interval. Let m and n be the average pixel depth values ​​within the interval, and m and n be the corresponding local disparity images of the target, respectively. t arg et The magnitudes of the x and y coordinates.

[0071] Step 6: Calculate the target local image information indicators, inputting the target local image Mat. t arg et and target local disparity image Dep t arg et The peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) of the two components were calculated respectively.

[0072] The formula for calculating Peak Signal-to-Noise Ratio (PSNR) is as follows:

[0073]

[0074]

[0075] In equations (7) and (8), (i,j) are the pixels to be calculated, MSE is the mean square error, and PSNR describes the similarity between the local image of the target and the corresponding disparity map. For a high-quality disparity map, the image should have less noise and clear texture edges, and should be more similar to the original image, and its PSNR should be higher. Conversely, its PSNR should be lower.

[0076] The formula for calculating structural similarity (SSIM) is as follows:

[0077]

[0078] In equation (9), x is the local image of the target, Mat. target y represents the local disparity image of the target, Dep. target μ x , Let μ be the mean and standard deviation of x. y , Let σ be the mean and standard deviation of y. xy Let c1 be the covariance of x and y, and c1 = (k1L). 2 c2=(k2L) 2 The constant used to maintain stability is L, which is the dynamic range of pixel values, k1 = 0.01, k2 = 0.03; Structural similarity (SSIM) compares the target local image with the corresponding disparity map in terms of brightness, contrast, and structure. For a high-quality disparity map, the higher its SSIM value; conversely, the lower its SSIM value.

[0079] Step 7: Based on the five indicators obtained above—single-pixel deviation influence factor σ, depth information error rate ε, depth dispersion ξ, peak signal-to-noise ratio (PSNR), and structural similarity (SSIM)—construct a binocular positioning accuracy evaluation model and comprehensively calculate the binocular positioning evaluation score S. i Each indicator was normalized, and the positioning evaluation score S for binocular cameras with different focal lengths was calculated. i The formula is as follows:

[0080]

[0081] In equation (10), a is a constant term, b, c, d, e, and f are the weights of the five evaluation indicators, k is the number of targets identified, and S i Let be the binocular localization evaluation score for the i-th pair of binocular cameras.

[0082] Step 8: First, determine if all pair of binocular cameras have been switched. If not, switch to the next pair of high-focal-length binocular cameras and start processing from step 1. If all pair of binocular cameras have been switched, then according to the vision system's self-switching logic, based on the number of target recognitions and the binocular localization evaluation score S of each pair of binocular cameras... iSelect the optimal binocular camera and switch to the optimal binocular camera for subsequent target localization processing;

[0083] like Figure 3 As shown, the self-switching logic steps of the vision system are as follows: First, compare the number of targets recognized by all binocular cameras. If the number of targets recognized is different, switch to the binocular camera with the most targets recognized, and the self-switching ends. Second, compare the binocular positioning evaluation scores of all binocular cameras. Switch to the binocular camera with the highest binocular positioning evaluation score, and the self-switching ends. When the number of targets recognized and the binocular positioning evaluation score are the same, switch to the low-focal-length camera, and the self-switching ends.

[0084] In summary, this application uses a target recognition algorithm to identify and perceive the target object, and obtains local depth information of the recognition result based on the binocular imaging principle. Based on the local depth information, three indicators for evaluating positioning accuracy are calculated, and two evaluation indicators are obtained based on the local image information of the target. A target positioning evaluation model is then established to quantify the positioning accuracy. According to the autonomous switching logic, the target recognition result and the positioning accuracy evaluation score are compared sequentially to achieve autonomous switching between multiple binocular cameras with different focal lengths, thereby meeting the recognition and positioning needs of both near-range and far-range targets.

[0085] By using a multi-pair binocular camera system with different focal lengths to identify and locate target objects, an autonomous switching logic for the multi-view vision system is constructed, and an evaluation model for the identification and positioning results is established. Based on the switching logic and the evaluation results of each pair of cameras, the system enables the autonomous switching of multiple pairs of binocular cameras. The optimal binocular camera is used to obtain more accurate target positioning information, overcoming the limitations of binocular positioning distance and field of view. This compensates for the limitations of fixed focal length in single-pair binocular vision equipment and overcomes the problems of low positioning accuracy of low focal length binocular cameras for distant targets and small field of view of high focal length binocular cameras. Through the autonomous switching of multiple pairs of binocular cameras, the system simultaneously satisfies the accurate perception and ranging of both near and far targets, significantly improving the accuracy and efficiency of target identification and positioning.

[0086] It should be noted that the appendix Figure 4 The diagram shown is a simplified schematic of a device for switching between three pairs of binocular cameras. This invention includes, but is not limited to, the autonomous switching of the three pairs of cameras shown. The number of binocular camera pairs can be freely set, and their placement can be different from that shown in the diagram, all of which are within the protection scope of this invention patent.

[0087] The above embodiments are only used to illustrate the technical solutions of the embodiments of this application, and are not intended to limit them. Although the embodiments of this application have been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features, without departing from the spirit and scope defined by the claims of this application.

Claims

1. A target recognition and localization method based on self-switching of a multi-view vision system, characterized in that, The method includes the following steps: Step 1: Acquire real-time images using a low-focal-length binocular camera, perform distortion correction on each frame, establish a target detection model, input the distortion-corrected images into the target detection model, and extract local target images. ; Step 2: Extract the target local image Then, using the imaging principle of a binocular camera and the SGBM stereo matching algorithm, the depth information of each pixel is obtained with the left camera coordinate system as the reference. The depth information of each pixel is then used to construct a disparity map. The grayscale value of each pixel is described in the disparity map. Find the local image of the target to be located The corresponding target local disparity image ; Step 3: Calculate the single-pixel deviation influence factor using the binocular ranging principle. ; Step 4: Obtain the target local disparity image Then, the target local disparity image is obtained by traversing through the SGBM stereo matching algorithm. The depth data of each pixel in the image is used to obtain the depth information error rate. ; Step 5: Calculate the depth dispersion of the target region. , to represent the depth credibility of the target region; Step 6: Calculate the target local image information indicators, and input the target local image. and target local disparity image The peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) of the two were obtained respectively. Step 7: Using the obtained single-pixel deviation influence factor Depth information error rate Depth of dispersion The binocular localization evaluation score is calculated based on five metrics: peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and peak signal-to-noise ratio (PSNR). ; Step 8: After completing all the switching and acquisition of data from the binocular cameras and obtaining the binocular localization evaluation score. Then, based on the visual system's self-switching logic, the system calculates the target recognition count and binocular localization evaluation score for each pair of binocular cameras. Select the optimal stereo camera and switch to it for subsequent target localization processing.

2. The target recognition and localization method based on self-switching of a multi-view vision system according to claim 1, characterized in that, In step one, the distortion-corrected image is input into the target detection model for target perception. If a target is detected, the number of targets is saved, and the pixel areas of multiple targets are compared. The local image of the target with the largest pixel area is extracted. If no target is detected, it determines whether all evaluations of the binocular cameras have been completed, and switches to the high-focal-length camera for evaluation or performs autonomous switching of the vision system based on the results.

3. The target recognition and localization method based on self-switching of a multi-view vision system according to claim 1, characterized in that, In step two, the target local disparity image is acquired. The formula is: Official (1) Official (2) In equations (1) and (2), , They are respectively x and y coordinates , They are respectively The x and y coordinates.

4. The target recognition and localization method based on self-switching of a multi-view vision system according to claim 1, characterized in that, In step three, the single-pixel deviation influence factor is calculated. The formula is: Official (4) In equation (4), Z is the distance to be measured, B is the baseline length, and f is the focal length. The pixel size of the camera. This is the single-pixel deviation influence factor, which is the ranging error corresponding to one pixel; Determine the single-pixel deviation influence factor The formula for calculating the distance Z to be measured in the formula is: Official (3) In equation (3), Z is the distance to be measured, B is the baseline length, and f is the focal length. For parallax.

5. The target recognition and localization method based on self-switching of a multi-view vision system according to claim 1, characterized in that, In step four, the target local disparity image is obtained. Subsequently, using the SGBM stereo matching algorithm, all stereo matching error points were assigned a fixed value of 10000mm, meaning that the depth information error value within the target interval was always 10000mm. The local depth image of the target was then iterated through to obtain the correct values. The depth data of each pixel in the image is used to obtain the depth information error rate. The formula is as follows: Official (5) In equation (5), For depth information error rate, For the incorrectly positioned pixel with depth information of 10000mm, , These are the corresponding target local disparity images. The magnitudes of the x and y coordinates.

6. The target recognition and localization method based on self-switching of a multi-view vision system according to claim 1, characterized in that, In step five, the depth dispersion of the target region is determined. The formula is as follows: Official (6) In equation (6), Within the interval The depth value of a pixel. This represents the average pixel depth value within the interval. , These are the corresponding target local disparity images. The magnitudes of the x and y coordinates.

7. The target recognition and localization method based on self-switching of a multi-view vision system according to claim 1, characterized in that, In step six, the formula for calculating the peak signal-to-noise ratio (PSNR) is as follows: Official (7) Official (8) In equations (7) and (8), For each pixel to be calculated, MSE is the mean squared error, and PSNR describes the similarity between the local image of the target and the corresponding disparity map. The formula for calculating structural similarity (SSIM) is as follows: Official (9) In equation (9), For target local image , For target local disparity image , , for The mean and standard deviation, , for The mean and standard deviation, yes and covariance, The constant used to maintain stability, L, is the dynamic range of pixel values. =0.01, =0.

03.

8. A target recognition and localization method based on self-switching of a multi-view vision system according to claim 1, characterized in that, In step seven, the binocular localization evaluation score is obtained. The formula is as follows: Official (10) In equation (10), a is a constant term, b, c, d, e, and f are the weights of the five evaluation indicators, and k is the number of targets identified. Let be the binocular localization evaluation score for the i-th pair of binocular cameras.

9. A target recognition and localization method based on self-switching of a multi-view vision system according to claim 1, characterized in that, In step eight, it is first determined whether the switching of all pairs of binocular cameras has been completed. If not, the process switches to the next pair of high-focal-length binocular cameras and begins from step one. If the switching of all pairs of binocular cameras has been completed, the process is based on the vision system's self-switching logic, according to the number of target recognitions and binocular localization evaluation scores of each pair of binocular cameras. Select the optimal binocular camera and switch to the optimal binocular camera for subsequent target localization processing; The self-switching logic steps of the vision system are as follows: First, compare the number of targets recognized by all binocular cameras. If the number of targets recognized is different, switch to the binocular camera with the most targets recognized, and the self-switching ends. Second, compare the binocular positioning evaluation scores of all binocular cameras. Switch to the binocular camera with the highest binocular positioning evaluation score, and the self-switching ends. When the number of targets recognized and the binocular positioning evaluation score are the same, switch to the low-focal-length camera, and the self-switching ends.

Citation Information

Patent Citations

  • Fire hazard monitoring and positioning device based on binocular cameras and fire hazard monitoring and positioning method

    CN105741481A

  • Human body circumference measuring method based on multi-view vision system

    CN111598939A

  • Roadway multi-view vision measurement method and devicebased on line structured light scanning

    CN112945121A

  • A robot-assisted flame recognition and positioning method

    CN114494997B

  • A monocular and binocular vision system for object detection and pose measurement is constructed by a PTZ camera

    CN109308693A