Image processing methods for binocular camera systems built on smart terminals

By building a binocular camera system on a smart terminal and using the smart terminal processor and extended camera for image processing, the problems of high cost and low depth measurement accuracy of existing binocular cameras are solved, realizing high-precision, low-cost depth information acquisition and flexible use.

CN119919467BActive Publication Date: 2025-10-31GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510150685.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-10-31
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

Existing binocular camera products are expensive, mainly because they require embedded high-performance processors to calculate depth values, and the fixed camera structure results in poor scalability and a short baseline, leading to low depth measurement accuracy.

Method used

A binocular camera system is built based on a smart terminal. The smart terminal's processor and a camera, along with an extended camera, form a binocular camera. Depth information is generated through calibration, stereo matching, and disparity calculation. A deep learning module is combined to improve accuracy and reduce costs.

Benefits of technology

It enables high-precision, low-cost depth information acquisition on smart terminals, enhancing the flexibility of usage scenarios and user experience, reducing hardware costs, and improving the accuracy of depth measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919467B_ABST
    Figure CN119919467B_ABST
Patent Text Reader

Abstract

This invention discloses an image processing method for a binocular camera system built on a smart terminal, applicable to a smart terminal and an extended camera. The smart terminal includes a memory, a processor, a depth information processing module, and a camera assembly. The camera assembly includes at least one camera. The depth information processing module runs on the processor, and the memory stores the depth information processing module. The method includes: forming a binocular camera system using the camera and the extended camera; acquiring calibration parameters of the binocular camera system and inputting the calibration parameters into the depth information processing module; the binocular camera system capturing images of a target object to obtain a first target image and a second target image, and inputting the first and second target images into the depth information processing module; the depth information processing module generating depth information. This application utilizes the computing power of the smart terminal itself to process images and obtain depth information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to an image processing method for a binocular camera system built on a smart terminal. Background Technology

[0002] In related technologies, stereo cameras often require embedded high-performance processors to calculate depth values, which increases the cost of stereo camera products. Summary of the Invention

[0003] This invention aims to at least solve one of the technical problems existing in the prior art. Therefore, one objective of this invention is to propose an image processing method for a binocular camera system built on a smart terminal. This method constructs a binocular camera system based on a single camera and an extended camera on the smart terminal, and utilizes the computing power of the smart terminal itself to process the images from both cameras, generating more accurate and real-time depth information.

[0004] This invention provides an image processing method for a binocular camera system built on a smart terminal, applicable to the smart terminal and an extended camera. The smart terminal includes a memory, a processor, a depth information processing module, and a camera assembly. The camera assembly includes at least one camera. The depth information processing module runs on the processor, and the memory stores the depth information processing module. The method includes:

[0005] The camera and the extended camera form a binocular camera system;

[0006] Obtain the calibration parameters of the binocular camera system and input the calibration parameters into the depth information processing module;

[0007] The binocular camera system captures images of the target object to obtain a first target image and a second target image, and inputs the first target image and the second target image into the depth information processing module;

[0008] The depth information processing module generates disparity values ​​based on the first target image and the second target image, and uses the disparity values ​​and the calibration parameters to generate depth information.

[0009] In some embodiments, the smart terminal and the extended camera form a binocular camera system, including:

[0010] The extended camera is communicatively connected to the smart terminal;

[0011] Adjust the positions of the extended camera and the camera lens;

[0012] Make the pixels of the images captured by the extended camera and the camera consistent, and align the rows of the images captured by the extended camera and the camera.

[0013] In some embodiments, aligning the rows of images captured by the extended camera and the camera lens includes:

[0014] The camera is used to acquire a first aligned image, and the extended camera is used to acquire a second aligned image;

[0015] Align the rows of the first aligned image and the second aligned image.

[0016] In some embodiments, the smart terminal includes a display device for displaying the first alignment image and the second alignment image;

[0017] The step of aligning the rows of the first aligned image and the second aligned image further includes:

[0018] The lens of the extended camera and the lens of the camera are located on the same physical plane;

[0019] Align the extended camera and the camera vertically, and align corresponding pixels of the same object in the first aligned image and the second aligned image in the vertical direction.

[0020] In some embodiments, the smart terminal further includes an auxiliary alignment module; the auxiliary alignment module is used to obtain the overlap between the first aligned image and the second aligned image, thereby assisting in aligning the extended camera and the camera.

[0021] In some embodiments, the depth information processing module has a first preset baseline value and a second preset baseline value; the positions of the extended camera and the camera are adjusted so that the distance between the extended camera and the camera is one of the first preset baseline value and the second preset baseline value, thereby enabling the measurement of depth information at different distances.

[0022] In some embodiments, obtaining the calibration parameters of the binocular camera system includes:

[0023] Calibrate using a calibration plate of known dimensions;

[0024] Obtain the calibration intrinsic parameter matrix and calibration extrinsic parameter matrix of the binocular camera system;

[0025] The intrinsic and extrinsic parameters are used for epipolar correction in stereo matching.

[0026] In some embodiments, the depth information processing module includes:

[0027] An image preprocessing module is used to perform grayscale conversion, noise reduction, and histogram equalization on the first target image and the second target image to obtain a first preprocessed image and a second preprocessed image.

[0028] A stereo matching module is used to perform stereo matching on the first preprocessed image and the second preprocessed image to obtain multiple corresponding feature points;

[0029] A disparity calculation module, which is used to obtain the disparity value of each feature point;

[0030] A depth calculation module, which obtains the depth value corresponding to each feature point based on the disparity value and the calibration parameters.

[0031] In some embodiments, the depth information processing module further includes a depth map imaging module; the depth map imaging module is used to convert the depth value into a depth map.

[0032] In some embodiments, the binocular camera system further includes a mounting bracket and an extended power supply; the mounting bracket is used to support and fix the extended camera and the smart terminal in position; the extended power supply is configured to supply power to the extended camera when the power consumption of the extended camera exceeds the power supply threshold of the smart terminal.

[0033] Based on the technical solutions, the embodiments provided by the present invention have the following advantages: (1) By combining an extended camera and a camera in a smart terminal to form a binocular camera system, the binocular camera system acquires images of the target object from different perspectives, namely the first target image and the second target image, thereby obtaining the disparity value. The disparity value information is used to assist in depth estimation. This method can improve the accuracy of depth estimation for smart terminals with only a single camera. (2) The image processing method provided in this application can be applied to a binocular camera system built on a smart terminal. The binocular camera system built on a smart terminal and the processor embedded in the smart terminal perform data processing and calculation on the image of the target object, thereby obtaining the depth information of the target object. The computing power of the processor of most smart terminals is sufficient to handle the calculation and processing of binocular disparity, which can greatly reduce the cost and make it more convenient for users to use. It has great flexibility in terms of usage scenarios. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a flowchart of an image processing method for a binocular camera system based on a smart terminal, according to an embodiment of the present invention.

[0036] Figure 2 This is a simplified model diagram of depth measurement using a binocular camera;

[0037] Figure 3 It is a geometric model diagram for obtaining depth values ​​using triangulation.

[0038] Figure 4 This is a schematic diagram of a binocular camera system built on a smart terminal according to an embodiment of the present invention;

[0039] Figure 5 This is a flowchart of an image processing method for a binocular camera system based on a smart terminal, according to an embodiment of the present invention.

[0040] Figure 6 This is a flowchart of an image processing method for a binocular camera system based on a smart terminal, according to an embodiment of the present invention.

[0041] Figure 7 This is a schematic diagram of the first and second aligned images in their row alignment state;

[0042] Figure 8 This is a flowchart of an image processing method for a binocular camera system based on a smart terminal, according to an embodiment of the present invention. Detailed Implementation

[0043] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0044] The camera industry for smart devices is constantly introducing new technologies and features to meet consumers' demands for high-quality shooting experiences. Technological innovations include higher pixel counts, advanced image processing technologies, multi-camera systems, and variable aperture technology. This technological innovation continues to accelerate. Smart devices include smartphones, tablets, laptops, smartwatches, and smart TVs. The pixel count of cameras on smart devices is constantly increasing, image processing technology is advancing, and multi-camera systems are becoming an industry trend.

[0045] Specifically, with the popularization and application of AI technology, smart terminal cameras have also made significant progress in areas such as intelligent recognition and scene optimization. Smart terminals using monocular cameras capture image information through a single camera, based on two-dimensional planar imaging; their structure is simple and their cost is relatively low. Smart terminals using binocular cameras capture images through two cameras, simulating human binocular vision, and using parallax to calculate the three-dimensional spatial coordinates of target objects, thus achieving stereoscopic vision.

[0046] Monocular cameras are widely used in the machine vision industry for tasks such as object recognition, location positioning, and 2D planar measurement. Binocular cameras are more suitable for scenarios requiring 3D spatial information, such as robot navigation, object grasping, and autonomous driving.

[0047] Binocular cameras have achieved significant breakthroughs in 3D reconstruction and scene understanding, providing more accurate spatial coordinate information to help systems make more precise decisions. They also demonstrate unique advantages in applications requiring high depth information, such as face recognition and gesture recognition.

[0048] With the continuous development of machine vision technology, both monocular and binocular cameras are constantly evolving. Advances in algorithm optimization and image processing technology have gradually improved the recognition accuracy and stability of monocular cameras. Binocular cameras have made continuous breakthroughs in hardware design and algorithm optimization, achieving higher spatial resolution and more accurate depth measurement.

[0049] In summary, the technological background for extending monocular cameras into binocular cameras involves technological innovation in the smart terminal camera industry, the technological differences between monocular and binocular cameras, the advantages and applications of binocular cameras, as well as technological challenges and future development. These factors have collectively driven the advancement of smartphone camera technology, making it possible to achieve higher-quality shooting experiences and more accurate spatial information.

[0050] Currently, smart terminals include one or more cameras. When using a smart terminal to measure the depth of a target object, a single-camera smart terminal has low depth measurement accuracy when the target object does not have obvious texture and the lighting changes are not significant. On the other hand, smart terminals with multiple cameras have close distances between the different cameras, resulting in a short baseline and small parallax value, which in turn reduces the accuracy of depth measurement.

[0051] Meanwhile, stereo cameras on the market usually consist of two specially designed cameras. The camera structure is fixed and has poor scalability. Stereo cameras need to be connected to a PC and can only be called through complex code. Moreover, the parallax calculation of stereo cameras consumes a lot of computing resources and usually requires GPU / FPGA devices for acceleration. Integrated stereo cameras on the market often embed high-performance processor modules in the product, which leads to a sharp increase in product cost.

[0052] The following is for reference. Figures 1-8 An image processing method for a binocular camera system based on a smart terminal, according to an embodiment of the present invention, is described.

[0053] Example 1

[0054] like Figure 1 and Figure 2 As shown, this embodiment provides an image processing method for a binocular camera system built on a smart terminal. The method is applied to a smart terminal 1 and an extended camera 2.

[0055] The smart terminal 1 includes a depth information processing module 11 and a camera assembly 12, the camera assembly including at least one camera. It should be noted that the smart terminal here can be a smartphone, tablet, laptop, smartwatch, or smart TV, etc., and these smart terminals have embedded processors capable of data processing and computation.

[0056] The image processing method of the binocular camera system 100 built on a smart terminal includes the following steps:

[0057] S1: The camera and the extended camera 2 form a binocular camera system 100;

[0058] S2: Obtain the calibration parameters of the binocular camera system 100 and input the calibration parameters into the depth information processing module 11;

[0059] S3: The binocular camera system 100 takes pictures of the target object to obtain a first target image and a second target image, and inputs the first target image and the second target image into the depth information processing module 11;

[0060] S4: The depth information processing module 11 generates disparity values ​​based on the first target image and the second target image, and generates depth information using the disparity values ​​and calibration parameters.

[0061] In a specific application scenario, a camera on a smart terminal and an extended camera 2 form a binocular camera system 100. The calibration parameters of the binocular camera system 100 need to be redefined. The calibration parameters include at least the actual baseline distance between the camera and the extended camera 2 (i.e., the distance between the camera and the extended camera 2), the focal length of the camera, and the focal length of the extended camera 2. Both the camera and the extended camera 2 simultaneously capture images of the target object. The camera acquires a first target image, while the extended camera 2 acquires a second target image. The processor of the smart terminal runs a depth information processing module 11. The depth information processing module 11 generates a disparity value based on the first and second target images, and then generates depth information based on the disparity value and the baseline.

[0062] In the embodiments of this application (1), a binocular camera system 100 is formed by combining an extended camera 2 and a camera in a smart terminal. The binocular camera system 100 acquires images of the target object from different perspectives, namely the first target image and the second target image, thereby obtaining the disparity value. The disparity value information is used to assist in depth estimation. This method can improve the accuracy of depth estimation for smart terminals with only a single camera. (2) The image processing method provided in this application can be applied to the binocular camera system 100 built on a smart terminal. The binocular camera system 100 built on a smart terminal and the processor embedded in the smart terminal perform data processing and calculation on the image of the target object, thereby obtaining the depth information of the target object. The computing power of the processor of most smart terminals is sufficient to handle the calculation and processing of binocular disparity, which can greatly reduce the cost and make it more convenient for users to use. It has great flexibility in terms of usage scenarios.

[0063] like Figure 4 As shown, the depth information processing module 11 further includes an image preprocessing module, a stereo matching module, a disparity calculation module, and a depth calculation module. The image preprocessing module performs grayscale conversion, denoising, and histogram equalization on the first target image and the second target image to obtain a first preprocessed image and a second preprocessed image. The stereo matching module performs stereo matching on the first preprocessed image and the second preprocessed image to obtain feature point pairs. The disparity calculation module obtains the disparity value of the feature point pairs. The depth calculation module obtains the depth value corresponding to each feature point based on the disparity value and calibration parameters.

[0064] The stereo matching module utilizes stereo matching algorithms, such as semi-global matching or deep learning-based stereo matching networks, to calculate the disparity between two images captured by extended camera 2 and the camera. Stereo matching is a computational technique that analyzes image pairs captured by the two extended cameras 2 and the camera, identifying and matching corresponding feature points in the two images. The positions of these matched points differ, i.e., the disparity. Disparity information can be obtained by comparing the positional differences of corresponding points in the left and right images; disparity information is crucial for depth perception. Once an accurate disparity map is obtained through the stereo matching algorithm, the next step is to convert the disparity information into depth information, which generates a corresponding depth map. The depth map reflects the distance information of each point in the scene to extended camera 2 or the camera. The generated depth map can be used in various applications, including 3D reconstruction, augmented reality, and obstacle detection in autonomous driving. The depth map provides three-dimensional information about the scene, enabling the device to better understand its surroundings.

[0065] like Figure 2 and Figure 3 As shown, it should be further explained that disparity information and calibration parameters can be converted into depth values ​​using triangulation. The figure illustrates the imaging model of a binocular camera. , The left and right aperture centers are indicated by the squares, which represent the imaging plane. It is the focal length. and The coordinates of the imaging plane are defined according to the coordinate definition in the figure. It should be a negative number, so the distance marked on the graph is - A stereo camera typically consists of two horizontally positioned cameras: a left-eye camera and a right-eye camera. In both stereo cameras, each camera is considered a pinhole camera; their horizontal placement means that the aperture centers of both cameras lie on the x-axis. The distance between them is called the baseline of the stereo camera (denoted as b), which is an important calibration parameter for the stereo camera system. Now, consider a spatial point P, which forms an image on both the left-eye and right-eye cameras, denoted as... , Due to the existence of the camera baseline (the distance between the two cameras), the two imaging positions are different. Ideally, since the left and right cameras only have displacement along the x-axis, the image of P also only differs along the x-axis (corresponding to the u-axis of the image). Let its left coordinate be... The coordinates on the right are The geometric relationship is shown on the right side of the diagram above. According to... and Similarity relationships:

[0066] =

[0067] After processing, the depth value can be obtained from the parallax value, the focal length in the calibration parameters, and the baseline in the calibration parameters using the following formula:

[0068] = ,d = -

[0069] in, is the focal length in the calibration parameters, b is the baseline (i.e., the distance between the two cameras) in the calibration parameters, and d is the parallax value (i.e., the horizontal difference of the same object in the two images).

[0070] This invention converts the differences in images captured by a binocular camera system 100 into depth information through calibration, stereo matching, disparity calculation, and triangulation, thereby generating a depth map and providing an interface for processing various tasks, such as using API interfaces on smartphones. These tasks include, but are not limited to, 3D reconstruction, virtual reality imaging, and autonomous driving tasks.

[0071] In a specific example, a deep learning module is stored in memory, and the processor runs this module. The deep learning module includes convolutional neural networks that can automatically learn complex feature representations from raw data, reducing the need for manual feature design. Furthermore, leveraging deep learning modules can improve accuracy in many vision tasks, including image classification, object detection, semantic segmentation, and stereo matching, demonstrating good generalization capabilities and adaptability to diverse scenarios and conditions. With the development of GPUs and other dedicated hardware on smart terminals, the training and inference processes of deep learning models can be further accelerated.

[0072] Example 2

[0073] like Figure 5 As shown, the smart terminal and the extended camera 2 form a binocular camera system 100. This step S1 includes the following specific steps:

[0074] S11: Extended camera 2 communicates with the smart terminal, so that the images acquired by extended camera 2 can be transmitted to the smart terminal. Extended camera 2 and smart terminal can be connected by wired or wireless connection. No specific connection method is limited here.

[0075] S12: Adjust the positions of the extended camera 2 and the camera lens. In related technologies, smartphones with multiple cameras have a baseline of approximately 10mm between two cameras (i.e., the distance between the two cameras is approximately 10mm). The close distance between different cameras results in a short baseline and a small parallax value, which in turn reduces the accuracy of depth measurement. In this application, however, one camera of the smart terminal and the extended camera 2 form a binocular camera system 100. The extended camera 2 and the camera lens maintain a certain baseline distance, between 1cm and 20cm, which allows for a larger parallax and thus improves the accuracy of depth estimation.

[0076] S13: Make the pixels of the images captured by the extended camera 2 and the camera consistent, and align the rows of the images captured by the extended camera 2 and the camera.

[0077] Specifically, pixels are the basic units that make up a digital image. Making the pixels of the images captured by the extended camera 2 and the camera consistent means that the images captured by the extended camera 2 and the camera must have the same number of pixels in both the horizontal and vertical directions. For example, if the image captured by the camera is 1920×1080 pixels (1920 pixels horizontally and 1080 pixels vertically), then the image captured by the extended camera 2 also needs to be adjusted to the same 1920×1080 pixels.

[0078] It should also be noted that image row alignment means that pixels in corresponding rows of the images captured by extended camera 2 and the camera should represent information about the same height position in the actual scene. Each row of the two images must strictly correspond, and for a target object in the actual scene, its projection information can be found in the corresponding rows of both images. For example, for a can in the actual scene, the pixel appearing in row 50 of the image captured by extended camera 2 should appear in the image captured by the camera, and the corresponding pixel should be approximately located in row 50 of the image captured by the camera. This satisfies the epipolar geometry relationship of stereo vision, ensuring that the matching point pair is located in the same image row, reducing the two-dimensional search to one-dimensional, significantly improving the efficiency of stereo matching, and thus reducing the computational load on the processor.

[0079] Example 3

[0080] like Figure 6 As shown, further, aligning the rows of images captured by the extended camera 2 and the camera lens, this step S13 includes the following sub-steps:

[0081] S131: The camera is used to acquire the first aligned image, and the extended camera 2 is used to acquire the second aligned image;

[0082] S132: Align the rows of the first and second aligned images.

[0083] This means that by adjusting the hardware mounting angle or using software-assisted correction, the imaging planes of the camera and the extended camera 2 are made to be on the same horizontal plane, ensuring that the same scene point only has horizontal displacement (X-axis difference) in the two images, eliminating vertical misalignment (Y-axis), thereby satisfying the epipolar geometric relationship of stereo vision, ensuring that the matching point pair is located in the same image row, reducing the two-dimensional search to one dimension, greatly improving the stereo matching efficiency, and thus reducing the computational load of the processor.

[0084] Example 4

[0085] This embodiment provides a specific implementation method for adjusting the hardware installation angle based on Embodiment 3.

[0086] like Figure 7 and Figure 8 As shown, the smart terminal 1 further includes a display device 13, which is used to display a first aligned image and a second aligned image; the step of aligning the rows of the first aligned image and the second aligned image further includes the following sub-steps:

[0087] S1321: Display device 13 displays the first aligned image and the second aligned image.

[0088] S1322: Position the lens of the extended camera 2 and the lens of the camera on the same physical plane. For example, the user first fixes the position of the extended camera 2, and then adjusts the lens of the camera using the extended camera 2 as a reference, so that the lenses of the extended camera 2 and the camera are on the same physical plane. It is understood that the user can also fix the position of the camera (i.e., fix the position of the smart terminal), and then adjust the lens of the extended camera 2 using the camera lens as a reference, so that the lenses of the extended camera 2 and the camera are on the same physical plane. During this process, both the extended camera 2 and the camera can capture images of the target object. The user can adjust the position by observing the image on the display device 13 to ensure that the same object appears simultaneously on the imaging planes of both the camera and the extended camera 2.

[0089] It should be further explained that this step ensures that the extended camera 2 and the main camera are located on the same plane as much as possible, simplifying the acquisition of the rotation matrix. The rotation matrix describes the rotational relationship between the main camera and the extended camera 2. When both are on the same plane, their rotation is relatively simple, and the parameters and computational load required to calculate the rotation matrix are reduced.

[0090] S1323: Align the extended camera 2 and the camera vertically, and align corresponding pixels of the same object in the first aligned image and the second aligned image in the vertical direction.

[0091] In real three-dimensional space, the direction perpendicular to the ground is usually defined as the vertical direction. The vertical position of the extended camera 2 and the camera is their position in this vertical direction. For example, if the smart terminal is placed flat on a table, with the table as the horizontal reference plane, then the direction perpendicular to the table, either upwards or downwards, is the vertical direction as discussed here. The extended camera 2 may be directly above the camera, directly below it, or have a certain height difference in the vertical direction; these different states represent their different vertical positions. On the display device 13 of the smart terminal, the pixels corresponding to the same object in the first aligned image (captured by the camera) and the second aligned image (captured by the extended camera 2) will form a line. If these lines are not horizontal, it means that the two are not aligned in the vertical direction, that is, their vertical positions are deviated. The user can manually move the extended camera 2 or the smart terminal (including the camera) to change their vertical positions. For example, if it is found that the line connecting corresponding pixels on the display device 13 is tilted, the higher-positioned camera is moved downwards, or the lower-positioned camera is moved upwards, until the corresponding pixels of the same object in the two images are on the same horizontal plane, at which point the vertical alignment adjustment is completed.

[0092] In a specific example, two image frames can generate lines connecting the same pixels corresponding to the same object, as shown in the figure. The user can assist in alignment by adjusting the relative vertical positions of the camera and extended camera 2. By manually aligning the two cameras vertically, and placing the same pixels on the same horizontal plane, the user simplifies the acquisition of the translation matrix.

[0093] It should be further explained that this step simplifies the acquisition of the translation matrix. The translation matrix is ​​used to describe the translation relationship between the camera and the extended camera 2. When the two are vertically aligned, their translation in the vertical direction becomes simpler, making it easier to calculate the translation matrix.

[0094] In summary, the calibration parameters include the calibration intrinsic parameter matrix and the calibration extrinsic parameter matrix. The calibration extrinsic parameter matrix includes the translation and rotation matrix parameters between the camera and the extended camera 2. Therefore, after this step, the user manually aligns the calibration extrinsic parameter matrix of the simplified binocular camera system 100, optimizing user experience and reducing the computational load on the processor.

[0095] Example 5

[0096] This embodiment provides a specific implementation method for software-assisted correction based on Embodiment 3.

[0097] Furthermore, the smart terminal also includes an auxiliary alignment module; the auxiliary alignment module is used to obtain the overlap between the first aligned image and the second aligned image, thereby assisting in the alignment of the extended camera 2 and the camera.

[0098] The alignment assistance module feeds back the calculated overlap information to the user. For example, it displays the overlap value on the smart terminal's display device 13, uses intuitive graphics (such as progress bars or color indicators) to show the degree of overlap, or provides voice prompts on the smart terminal. Simultaneously, it provides suggestions for adjusting the positions of the extended camera 2 and the main camera based on the overlap, prompting the user to move the extended camera 2 up, down, left, or right by a certain distance.

[0099] During the process of the user adjusting the positions of the extended camera 2 and the camera based on the feedback information, the auxiliary alignment module will monitor the overlap changes of the first and second aligned images in real time until the overlap reaches a satisfactory level, thereby achieving precise alignment of the extended camera 2 and the camera.

[0100] Example 6

[0101] Furthermore, the depth information processing module 11 has a first preset baseline value and a second preset baseline value; by adjusting the positions of the extended camera 2 and the camera so that the distance between the extended camera 2 and the camera is one of the first preset baseline value and the second preset baseline value, it is possible to measure depth information at different distances.

[0102] The first and second preset baseline values ​​are different baseline values. In a specific example, the preset baseline values ​​can also include a third, fourth, and even nth preset baseline value. Different application scenarios may involve depth measurement of objects at different distances. A binocular vision system with a single baseline value has limitations in its measurement range, and may only be able to effectively measure objects within a certain distance range. Setting a first and second preset baseline value can expand the measurement range, enabling the measurement of both near and far objects. For example, in robot navigation in advanced task areas, the robot may need to simultaneously perceive nearby obstacles and the location of distant targets. By switching between different preset baselines, the depth measurement requirements for objects at different distances can be met.

[0103] It should be further explained that after the user selects the preset baseline value in the depth information processing module 11, the distance between the extended camera 2 and the camera is adjusted according to the preset baseline value. During this process, the distance between the two is measured using a measuring tool until the distance between the extended camera 2 and the camera is the same as the preset baseline value.

[0104] Example 7

[0105] This embodiment is basically the same as Embodiment 1, except that:

[0106] The step of obtaining the calibration parameters of the binocular camera system 100 further includes the following sub-steps:

[0107] Calibrate using a calibration plate of known dimensions;

[0108] Obtain the intrinsic and extrinsic parameter matrices of the binocular camera system 100;

[0109] Epipolar correction for stereo matching is performed based on intrinsic and extrinsic parameter matrices.

[0110] In a specific scenario, users require more precise depth information, which necessitates a series of manual operations. Here, calibration using a calibration board precisely determines the camera's intrinsic parameter matrix. This matrix includes information such as the principal point position, distortion parameters, and focus of the binocular camera system 100. The extrinsic parameter matrix includes the rotation and translation matrices of the binocular camera system 100. Users can position the two cameras according to their measurement scenario, without needing simplified operations such as camera alignment.

[0111] It should be noted that although obtaining the calibration parameters of the binocular camera system 100 based on a calibration plate of known size is a more complex process, it can yield more accurate depth information.

[0112] In a specific example, the user needs to prepare a chessboard grid, the size of the inner points, and the length of the black and white squares measured by a ruler must be known. The user also needs to take multiple pictures from different perspectives as a dataset, and then perform calculations in the smart terminal to obtain the calibration parameters of the camera and the extended camera.

[0113] Example 8

[0114] Furthermore, the depth information processing module 11 also includes a depth map imaging module, which is used to convert depth values ​​into a depth map. The depth map can be a grayscale image, where the grayscale value represents the distance from the camera. The smaller the grayscale value, the closer the object is to the camera; the larger the grayscale value, the farther the object is from the camera.

[0115] Example 9

[0116] Furthermore, the binocular camera system 100 also includes a mounting bracket and an extended power supply; the mounting bracket is used to fix the position of the camera and the smart terminal; the extended power supply is configured to supply power to the extended camera 2 when the power consumption of the extended camera 2 exceeds the power supply threshold of the smart terminal.

[0117] The extended power supply can be a battery of a certain capacity or the extended camera 2 can be directly plugged in to provide power to the extended camera 2. To ensure the relative position stability of the camera and the extended camera 2, a fixing bracket can be used to secure the smart terminal and the extended camera 2. The design of the fixing bracket needs to ensure that the extended camera 2 and the camera of the smart terminal have the correct relative position in physical space for accurate alignment. The fixing bracket can be designed as a clamp or a magnetic bracket. Ideally, the extended camera 2 should maintain a certain baseline distance (usually a few centimeters to tens of centimeters) from the camera and be perpendicular to the field of view of the camera.

[0118] The following describes a specific implementation example using a smartphone as an example of a smart terminal 1.

[0119] (I) Hardware level

[0120] The Extended Camera 2 can be used with existing cameras on smartphones and is compatible with different mobile operating systems (Android, HarmonyOS, or iOS), correctly recognizing and supporting image transmission. The Extended Camera 2 also requires a USB-C interface and supports OTG connectivity, as well as high-speed Bluetooth and Wi-Fi transmission standards.

[0121] The extended camera 2 also connects to an extended power source, which is used to charge the extended camera 2. If the extended camera 2 is connected to a smartphone via USB, the smartphone's battery may run out quickly. Therefore, it is necessary to ensure that the extended camera 2 has sufficient power. The extended power source can use a battery of a certain capacity and can charge both the smartphone and the extended camera 2.

[0122] To ensure stable relative positioning between the extended camera 2 and the smartphone, a mounting bracket is used to hold the extended camera 2 in place. The bracket design must ensure the correct relative position of the extended camera bracket and the smartphone's camera in physical space for accurate calibration. The bracket can be designed as a clamp or magnetic bracket, allowing for easy assembly with the extended camera 2 and the smartphone. For more precise binocular vision, the relative angle and position of the extended camera 2 and the smartphone's camera need to be known and remain fixed throughout use. Ideally, the extended camera 2 should maintain a baseline distance (typically a few centimeters to tens of centimeters) from the camera and be perpendicular to the camera's field of view.

[0123] (ii) Software level

[0124] This invention integrates a corresponding UI interface into a smartphone. After connecting the extended camera 2, line alignment is required. At the algorithm level, several known preset baseline values ​​'b' are defined, corresponding to several different settings for measuring depth information at different distances. The UI interface provides a simple process to guide the user through the camera alignment operation, such as moving the external camera or adjusting the angle, so that the phone screen can display the overlap or alignment status between the extended camera 2 and the camera's captured image. During calibration, the phone screen displays the alignment effect of the binocular camera system 100 in real time, helping the user ensure that the two cameras are correctly aligned and reducing errors.

[0125] The UI displays real-time images from both cameras, along with processed depth images. The depth map is presented to the user in various formats, such as color mapping, grayscale, and 3D point clouds, allowing the user to choose the appropriate display method. To aid user understanding of depth information, specific areas (such as objects closer to or farther from the camera) can be highlighted in the depth map. The depth map is updated in real-time as the user moves the binocular camera system, providing feedback on changes in depth information.

[0126] It should be further clarified that a user-moving binocular camera system refers to a system where the relative positions of the main camera and the extended camera are fixed, while the main camera and the extended camera as a whole move relative to the target object. In other words, the relative positions of the smartphone and the extended camera do not change; they move synchronously.

[0127] Since calculating depth values ​​using a binocular camera system requires significant computational resources, the UI can also provide real-time feedback on system performance, such as CPU and GPU load, and memory usage. This allows users to know if the system is currently overloaded and take appropriate measures (such as closing other applications or reducing computational load). Simultaneously, frame rate information can be displayed on the screen to help users understand whether the application is running smoothly.

[0128] When designing the UI for converting a smartphone external camera extension 2 into a binocular camera system 100, the key is to help users correctly install, calibrate, and operate the system, while providing an easy-to-understand and usable display of depth information. The UI should be simple, intuitive, and able to guarantee a smooth user experience under high-performance requirements, especially when real-time depth calculations and image processing are involved.

[0129] Other configurations and operations according to embodiments of the present invention are known to those skilled in the art and will not be described in detail here. In the description of the present invention, "first feature" and "second feature" may include one or more of the features. The up-down direction, left-right direction, and front-back direction are defined according to the up-down direction, left-right direction, and front-back direction shown in the figures.

[0130] In the description of this invention, unless otherwise expressly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features not in direct contact but through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature.

[0131] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0132] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. An image processing method for a binocular camera system built on a smart terminal, characterized in that, The invention is applied to smart terminals and extended cameras. The smart terminal includes a memory, a processor, a depth information processing module, and a camera assembly. The camera assembly includes at least one camera. The depth information processing module is used to run on the processor, and the memory is used to store the depth information processing module. The method includes: The camera and the extended camera form a binocular camera system, including: the extended camera being communicatively connected to the smart terminal; adjusting the positions of the extended camera and the camera; making the pixels of the images captured by the extended camera and the camera consistent, and aligning the rows of the images captured by the extended camera and the camera; The step of aligning the rows of images captured by the extended camera and the camera includes: the camera being used to acquire a first aligned image, and the extended camera being used to acquire a second aligned image; aligning the rows of the first aligned image and the second aligned image; The smart terminal includes a display device for displaying the first aligned image and the second aligned image; the step of aligning the first aligned image and the second aligned image further includes: positioning the lens of the extended camera and the lens of the camera on the same physical plane; aligning the extended camera and the camera vertically; and aligning corresponding pixels of the same object in the first aligned image and the second aligned image in the vertical direction. The smart terminal also includes an auxiliary alignment module; the auxiliary alignment module is used to obtain the overlap between the first aligned image and the second aligned image, thereby assisting in aligning the extended camera and the camera. Obtain the calibration parameters of the binocular camera system and input the calibration parameters into the depth information processing module; The binocular camera system captures images of the target object to obtain a first target image and a second target image, and inputs the first target image and the second target image into the depth information processing module; The depth information processing module generates disparity values ​​based on the first target image and the second target image, and uses the disparity values ​​and the calibration parameters to generate depth information.

2. The image processing method for a binocular camera system based on a smart terminal according to claim 1, characterized in that, The depth information processing module has a first preset baseline value and a second preset baseline value; The positions of the extended camera and the camera are adjusted so that the distance between the extended camera and the camera is one of the first preset baseline value and the second preset baseline value, thereby enabling the measurement of depth information at different distances.

3. The image processing method for a binocular camera system based on a smart terminal according to claim 1, characterized in that, The acquisition of calibration parameters for the binocular camera system includes: Calibrate using a calibration plate of known dimensions; Obtain the calibration intrinsic parameter matrix and calibration extrinsic parameter matrix of the binocular camera system; The intrinsic and extrinsic parameters are used for epipolar correction in stereo matching.

4. The image processing method for a binocular camera system based on a smart terminal according to claim 1, characterized in that, The depth information processing module includes: The image preprocessing module is used to perform grayscale conversion, noise reduction, and histogram equalization on the first target image and the second target image to obtain a first preprocessed image and a second preprocessed image; A stereo matching module is used to perform stereo matching on the first preprocessed image and the second preprocessed image to obtain multiple corresponding feature points; A disparity calculation module, which is used to obtain the disparity value of each feature point; A depth calculation module, which obtains the depth value corresponding to each feature point based on the disparity value and the calibration parameters.

5. The image processing method for a binocular camera system based on a smart terminal according to claim 4, characterized in that, The depth information processing module also includes a depth map imaging module; The depth map imaging module is used to convert the depth value into a depth map.

6. The image processing method for a binocular camera system based on a smart terminal according to claim 1, characterized in that, The binocular camera system also includes a mounting bracket and an extended power supply; The mounting bracket is used to support and fix the extended camera and the smart terminal in their positions. The extended power supply is configured to supply power to the extended camera when the power consumption of the extended camera exceeds the power supply threshold of the smart terminal.

Citation Information

Patent Citations

  • Intelligent terminal and 3D imaging method thereof, 3D imaging system

    CN109328459A

  • Image shooting terminal and image shooting method

    US20170064174A1