Image shooting method, apparatus, computer device, and readable storage medium
By using the coordinated shooting and image processing of the first and second cameras, the problem of poor display effect in the existing technology has been solved, achieving high-quality close-up display of the speaker, applicable to various camera types, and expanding the scope of application.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GUANGZHOU SHIYUAN ELECTRONICS CO LTD
- Filing Date
- 2026-01-14
- Publication Date
- 2026-07-30
AI Technical Summary
In existing conferencing software, the speaker close-up function implemented by the camera module of the conferencing tablet has poor optical performance, resulting in poor display effect and is not applicable to ordinary conferencing tablets and optical zoom cameras on the market.
The target object is captured by both a first camera and a second camera. By acquiring initial feature point pairs, rotation angles, and scaling ratio functions, the rotation and scaling of the rotatable lens of the second camera are controlled to improve the display effect.
It effectively improves the display effect of the screen, is suitable for a wide range of camera types, is not affected by differences in shooting depth, and does not require a fixed camera position, thus expanding the scope of application.
Smart Images

Figure CN2026072588_30072026_PF_FP_ABST
Abstract
Description
Image capturing method, apparatus, computer equipment, and readable storage medium
[0001] Related applications
[0002] This application claims priority to Chinese patent application filed on January 26, 2025, with application number 2025101219283, entitled "Image capturing method, apparatus, computer equipment and readable storage medium", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of image display technology, and in particular to an image capturing method, apparatus, computer device, and computer-readable storage medium. Background Technology
[0004] In modern meeting settings, speaker close-up functionality has become an important feature of meeting software. This feature automatically detects and tracks the speaker's face or body, magnifying and displaying it to allow remote participants to see the speaker's expressions and body language more clearly, thereby improving meeting communication efficiency and engagement.
[0005] In existing meeting scenarios, the speaker close-up function of meeting software is mainly implemented based on the camera module of the meeting tablet. To achieve the speaker close-up function, the camera module performs facial recognition on the speaker and obtains the positional relationship between the speaker and the camera module. Based on the facial recognition results and positional relationship, a close-up image of the speaker is displayed in the meeting software. However, because the optical performance of the camera modules of meeting tablets is generally poor, the display quality of the close-up image is poor. Summary of the Invention
[0006] Therefore, it is necessary to provide an image capturing method, apparatus, computer device, and computer-readable storage medium that can improve the display effect of the displayed image, in order to address the above-mentioned technical problems.
[0007] Firstly, this application provides an image capturing method, the method comprising:
[0008] When both the first camera and the second camera can capture the target object, the first current image captured by the first camera, the second current image captured by the second camera, and the current number of stepper motors of the second camera are acquired. An initial feature point is obtained from the first current image and the second current image respectively to form a target feature point pair. The lens of the second camera is a rotatable lens.
[0009] A first relationship function is obtained between the rotation angle of the second camera, the position coordinates of pixels in the image captured by the first camera, and the pixel value of pixels in the image captured by the second camera; wherein, the first position relationship of pixels in the image captured by the first camera is consistent with the second position relationship of pixels in the image captured by the second camera, the first position relationship is the position relationship between any pixel in the image captured by the first camera and the remaining pixels in the image captured by the first camera, and the second position relationship is the position relationship between any pixel in the image captured by the second camera and the remaining pixels in the image captured by the second camera;
[0010] Input the position coordinates of the corresponding initial feature point in the first current image and the pixel value of the corresponding initial feature point in the second current image into the first relational function, and output the target rotation angle;
[0011] Obtain a second relationship function between the scaling ratio, the proportion of the object captured by the first camera in the image captured by the first camera, and the number of stepper motors of the second camera. Input the proportion of the target object in the first current image and the current number of stepper motors into the second relationship function, and output the target scaling ratio.
[0012] The rotation of the second camera's rotatable lens is controlled based on the target rotation angle, and the second camera, after the lens rotation is controlled based on the target scaling ratio, captures the target object.
[0013] In one embodiment, an initial feature point is obtained from a first current image and a second current image respectively to form a target feature point pair, including:
[0014] Get the pixel values of pixels in the first and second current images;
[0015] Based on the pixel value of any pixel, obtain the pixel value difference between the pixel and the remaining pixels within a preset range around the pixel, and determine the total number of remaining pixels corresponding to pixel value differences greater than a first difference threshold.
[0016] If the total number is greater than the first preset number, the pixel is determined as the initial feature point, and a target feature point pair is formed based on the initial feature points in the first current image and the initial feature points in the second current image.
[0017] In one embodiment, a target feature point pair is formed based on initial feature points in the first current image and initial feature points in the second current image, including:
[0018] The initial feature point in the first current image is determined as the first initial feature point, and the first initial positional relationship between any one of the first initial feature points and the remaining first initial feature points in the first current image is obtained.
[0019] The initial feature point in the second current image is determined as the second initial feature point, and the second initial positional relationship between any one of the second initial feature points and the remaining second initial feature points in the second current image is obtained;
[0020] When there is a first initial positional relationship and a second initial positional relationship that are consistent, the first initial feature points and the second initial feature points corresponding to the consistent first initial positional relationship and the second initial positional relationship are determined as initial feature point pairs;
[0021] Select target feature point pairs from the initial feature point pairs.
[0022] In one embodiment, the process of obtaining the first relational function includes:
[0023] With the first camera at a first reference position, the second camera at a second reference position, and the rotatable lens at a reference angle position, at least one first reference image captured by the first camera within a preset time period and multiple second reference images captured by the second camera within the preset time period are acquired; wherein, the rotatable lens rotates continuously within the preset time period.
[0024] Select a first reference feature point from the pixels in the first reference image, and obtain the first reference positional relationship between any first reference feature point and the remaining first reference feature points in the first reference image;
[0025] Select second reference feature points from the pixels in the second reference image, and obtain the second reference positional relationship between any second reference feature point and the remaining second reference feature points in the second reference image;
[0026] When there is a first reference position relationship and a second reference position relationship that are consistent, the first reference feature points and the second reference feature points corresponding to the consistent first reference position relationship and the second reference position relationship are determined as a reference feature point pair;
[0027] A third relationship function is used to obtain the relationship between the rotation angle, the position coordinates of the pixels in the image captured by the first camera, the pixel value of the pixels in the image captured by the second camera, and the focal length of the second camera.
[0028] The first relation function is obtained based on the benchmark feature point pairs and the third relation function.
[0029] In one embodiment, obtaining a third relationship function among the rotation angle, the position coordinates of pixels in the image captured by the first camera, the pixel value of pixels in the image captured by the second camera, and the focal length of the second camera includes:
[0030] The fourth relationship function is obtained between the angle value of a pixel in the image captured by the first camera and the position coordinate of the pixel in the image captured by the first camera; the fifth relationship function is obtained between the angle value of a pixel in the image captured by the second camera, the pixel value of the pixel in the image captured by the second camera and the focal length of the second camera; and the sixth relationship function is obtained between the rotation angle, the angle value of a pixel in the image captured by the first camera and the angle value of a pixel in the image captured by the second camera.
[0031] Substituting the fourth and fifth relation functions into the sixth relation function yields the third relation function.
[0032] In one embodiment, obtaining the first relation function based on the baseline feature point pair and the third relation function includes:
[0033] For a pair of reference feature points, the current focal length of the second camera is obtained based on the position coordinates of the first reference feature point, the pixel value of the second reference feature point, and the third relational function.
[0034] Based on the current focal length of the second camera and the third relation function, obtain the first relation function.
[0035] In one embodiment, the process of obtaining the first reference position, the second reference position, and the reference angle position includes:
[0036] When both the first camera and the second camera can capture the target object, acquire a first preset image captured by the first camera and a second preset image captured by the second camera.
[0037] Select a first preset feature point from the pixels in the first preset image, and obtain a first preset positional relationship between any one first preset feature point and the remaining first preset feature points in the first preset image;
[0038] Select a second preset feature point from the pixels in the second preset image, and obtain a second preset positional relationship between any one second preset feature point and the remaining second preset feature points in the second preset image;
[0039] When there is a first preset positional relationship and a second preset positional relationship that are consistent, the first preset feature points and the second preset feature points corresponding to the consistent first preset positional relationship and the second preset positional relationship are determined as preset feature point pairs;
[0040] If the number of preset feature point pairs is greater than the second preset number, for each preset feature point pair, the pixel value difference between the first preset feature point and the second preset feature point is obtained, and the total number of preset feature point pairs whose pixel value difference is less than the second difference threshold is determined.
[0041] When the total number is greater than the third preset number, the first preset position of the first camera, the second preset position of the second camera, and the preset angle position of the rotatable lens are obtained. The first preset position is determined as the first reference position, the second preset position is determined as the second reference position, and the preset angle position is determined as the reference angle position.
[0042] Secondly, this application also provides an image capturing device, the device comprising:
[0043] The first acquisition module is used to acquire the first current image captured by the first camera, the second current image captured by the second camera, and the current number of stepper motors of the second camera when both the first camera and the second camera can capture the target object, and to acquire an initial feature point from the first current image and the second current image respectively to form a target feature point pair. The lens of the second camera is a rotatable lens.
[0044] The second acquisition module is used to acquire a first relationship function between the rotation angle of the second camera, the position coordinates of pixels in the image captured by the first camera, and the pixel value of pixels in the image captured by the second camera; wherein the first position relationship of pixels in the image captured by the first camera is consistent with the second position relationship of pixels in the image captured by the second camera, the first position relationship is the position relationship between any pixel in the image captured by the first camera and the remaining pixels in the image captured by the first camera, and the second position relationship is the position relationship between any pixel in the image captured by the second camera and the remaining pixels in the image captured by the second camera;
[0045] The input module is used to input the position coordinates of the corresponding initial feature points of the first current image and the pixel values of the corresponding initial feature points of the second current image into the first relational function, and output the target rotation angle;
[0046] The third acquisition module is used to acquire a second relationship function between the scaling ratio, the proportion of the object captured by the first camera in the image captured by the first camera, and the number of stepper motors of the second camera. The proportion of the target object in the first current image and the current number of stepper motors are input into the second relationship function, and the target scaling ratio is output.
[0047] The control module is used to control the rotation of the rotatable lens of the second camera based on the target rotation angle, and to control the second camera to capture the target object after the lens is rotated based on the target scaling ratio.
[0048] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the methods in any of the above embodiments.
[0049] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the methods in any of the above embodiments.
[0050] The aforementioned image capturing method, apparatus, computer equipment, and computer-readable storage medium, when both the first camera and the second camera can capture the target object, acquire a first current image captured by the first camera, a second current image captured by the second camera, and the current number of stepper motors of the second camera; and acquire an initial feature point from each of the first and second current images to form a target feature point pair; the lens of the second camera is a rotatable lens; and acquire a first relationship function between the rotation angle of the second camera, the position coordinates of pixels in the image captured by the first camera, and the pixel values of pixels in the image captured by the second camera; wherein the first positional relationship of pixels in the image captured by the first camera is consistent with the second positional relationship of pixels in the image captured by the second camera, and the first positional relationship is any one of the pixels in the image captured by the first camera. The first positional relationship is the positional relationship between any pixel in the image captured by the second camera and the remaining pixels in the image captured by the second camera. The position coordinates of the corresponding initial feature points in the first current image and the pixel values of the corresponding initial feature points in the second current image are input into a first relational function, which outputs the target rotation angle. A second relational function is obtained, relating the scaling ratio, the proportion of the object captured by the first camera in the image captured by the first camera, and the number of stepper motors in the second camera. The proportion of the target object in the first current image and the current number of stepper motors are input into the second relational function, which outputs the target scaling ratio. The rotatable lens of the second camera is controlled to rotate based on the target rotation angle, and the second camera, after rotating, captures the target object based on the target scaling ratio. This method of displaying the target object captured by the second camera based on the target rotation angle and target scaling ratio on the display screen effectively improves the display effect. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the disclosed drawings without creative effort.
[0052] Figure 1 is an application environment diagram of an image capturing method in one embodiment;
[0053] Figure 2 is a flowchart illustrating an image capture method in one embodiment;
[0054] Figure 3 is a pyramid diagram of the matching pairs (w0, x0) in one embodiment;
[0055] Figure 4 is a schematic diagram of the relationship between the number of stepper motors and the scaling parameters in one embodiment;
[0056] Figure 5 is a flowchart illustrating a method for obtaining target feature point pairs in one embodiment;
[0057] Figure 6 is a schematic diagram of the relationship between θ1 and θ2 in one embodiment;
[0058] Figure 7 is a schematic diagram of rotation calibration in one embodiment;
[0059] Figure 8 is a factor diagram of the first camera and the second camera during the rotation calibration process in one embodiment;
[0060] Figure 9 is a schematic diagram of rotation calibration in another embodiment;
[0061] Figure 10 is a structural block diagram of an image capturing device in one embodiment;
[0062] Figure 11 is an internal structure diagram of a computer device in one embodiment. Detailed Implementation
[0063] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0064] The problem of poor display effect of close-up images in the background technology can be caused by blurry close-up images, unstable and shaky portraits in close-up images, and severe image distortion when close-up images are magnified.
[0065] To address the aforementioned technical issues, existing technologies use mounting holes on the conference tablet and the optical zoom camera to fix the camera module of the conference tablet to an optical zoom camera using screws. This ensures that the relative spatial position between the camera module and the optical zoom camera is known. Based on this relative spatial position, this method optimizes the spatial orientation of the optical zoom camera by combining images of the speaker captured by both the camera module and the optical zoom camera. The optimized image of the speaker captured by the optical zoom camera is then displayed in the conference software, effectively improving the display quality of close-up shots. However, during the process of capturing the speaker's image, there may be situations where one of the cameras—the camera module or the optical zoom camera—fails to capture the speaker's image. Since the relative spatial position between the camera module and the optical zoom camera is fixed, this problem cannot be solved by adjusting the positions of the two cameras. Furthermore, because the shooting depths of the camera module and the optical zoom camera may differ, it becomes impossible to accurately determine the distance between the speaker and the two cameras based on the images captured by the two cameras. For example, the shooting depth of the camera module might be 2.5m, while the shooting depth of the optical zoom camera might be 3m. Additionally, this method requires fixed holes in both the conference tablet and the optical zoom camera, limiting its application to specific conference tablets and optical zoom cameras, and restricting its applicability to common conference tablets and optical zoom cameras on the market.
[0066] When the image capturing method provided in this embodiment is applied to a modern conference scenario, the first camera is a camera module, the second camera is an optical zoom camera, and the target object is the speaker. Since the relative positions between the first and second cameras are not fixed in this embodiment, adjusting their relative positions ensures that both cameras can simultaneously capture the target object. During the control process in this embodiment, the distance between the speaker and the two cameras does not need to be considered; therefore, the problem of being unable to determine the distance between the speaker and the two cameras due to different shooting depths of the two cameras will not occur. Since this embodiment does not require fixing the two cameras together by drilling holes and screws, the selection range of the first and second cameras in this embodiment is wide, and ordinary cameras on the market can be used in this embodiment.
[0067] The image capturing method provided in this application embodiment can be applied to the application environment shown in Figure 1. The first camera 104 can be a standalone camera or a camera mounted on the terminal 102 or other terminals. The terminal 102 first controls the first camera 104 to capture a first current image and then controls the second camera 106 to capture a second current image. Based on the first and second current images, it obtains rotation angle compensation values and scaling ratio compensation values. Finally, based on the rotation angle compensation values and scaling ratio compensation values, it controls the second camera 106 to capture the target object. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, smart interactive whiteboards, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc.
[0068] In an exemplary embodiment, as shown in FIG2, an image capturing method is provided. Taking the application of this method to the terminal in FIG1 as an example, the method includes the following steps 202 to 210. Wherein:
[0069] S202, when both the first camera and the second camera can capture the target object, acquire the first current image captured by the first camera, the second current image captured by the second camera, and the current number of stepper motors of the second camera, and acquire an initial feature point from the first current image and the second current image respectively to form a target feature point pair, wherein the lens of the second camera is a rotatable lens.
[0070] Here, the number of stepper motors refers to the number of pulses that the camera's stepper motor needs to receive in one complete rotation cycle; the feature points are pixels selected from the first current image and the second current image.
[0071] Optionally, in this embodiment, the lens of the first camera does not rotate. The rotatable lens of the second camera rotates around its main axis. The rotation of the rotatable lens of the second camera can be achieved by the lens itself rotating, or by the second camera rotating, thereby causing the lens to rotate. This embodiment does not specifically limit this. For example, for a second camera mounted on a gimbal, the rotation of the gimbal can be controlled to rotate the second camera, thereby causing the lens of the second camera to rotate. The gimbal is a device used to stabilize the camera and reduce shake during handheld shooting.
[0072] S204. Obtain a first relationship function between the rotation angle of the second camera, the position coordinates of the pixels in the image captured by the first camera, and the pixel value of the pixels in the image captured by the second camera; wherein, the first position relationship of the pixels in the image captured by the first camera is consistent with the second position relationship of the pixels in the image captured by the second camera, the first position relationship is the position relationship between any pixel in the image captured by the first camera and the remaining pixels in the image captured by the first camera, and the second position relationship is the position relationship between any pixel in the image captured by the second camera and the remaining pixels in the image captured by the second camera.
[0073] Pixel value refers to the color or grayscale information of a pixel, and is used to characterize the brightness of a pixel. Optionally, the first current image and the second current image can be mapped to the same image coordinate system, and the position coordinates of the feature points in this image coordinate system can be obtained.
[0074] S206. Input the position coordinates of the corresponding initial feature point of the first current image and the pixel value of the corresponding initial feature point of the second current image into the first relational function, and output the target rotation angle.
[0075] The target rotation angle is used to control the rotation of the second camera so that the second camera can capture the target object.
[0076] S208. Obtain the second relationship function between the scaling ratio, the proportion of the object captured by the first camera in the image captured by the first camera, and the number of stepper motors of the second camera. Input the proportion of the target object in the first current image and the current number of stepper motors into the second relationship function, and output the target scaling ratio.
[0077] The scaling ratio refers to the ability of a camera lens to magnify distant objects to a closer view; this ratio is usually expressed as a multiple.
[0078] Optionally, the second relational function is shown in equation (1): e w =a(x 4 -w 4 )+b(x 3 -w 3 )+c(x 2 -w 2 )+d(xw)
[0079] In equation (1), e wHere, a, b, c, and d are scaling ratios, x is the number of stepper motors for the second camera, and w is the proportion of the object captured by the first camera in the image. For example, w can be represented by the proportion of the width of the object in the image captured by the first camera to the total width of the image captured by the first camera. In a modern conference setting, the object can be the speaker.
[0080] Optionally, the second relationship function can be based on the relationship function between the scaling ratio of the second camera and the number of stepper motors, the relationship function between the scaling ratio of the first camera and the proportion of the object captured by the first camera in the image, and the scaling ratio e. w The scaling factor of the second camera is determined by the relationship function between the scaling factor of the first camera and the scaling factor of the second camera.
[0081] Optionally, the relationship between the scaling ratio of the second camera and the number of stepper motors is shown in the following equation (2): s1=ax 4 +bx 3 +cx 2 +dx+e
[0082] In equation (2), s1 is the scaling ratio of the second camera, and e is the scaling parameter.
[0083] Optionally, the relationship between the scaling ratio of the first camera and the proportion of the object captured by the first camera in the image is shown in the following equation (3): s2=aw 4 +bw 3 +cw 2 +dw+e
[0084] In equation (3), s2 is the scaling factor of the first camera.
[0085] Optionally, the scaling ratio e w The relationship between the scaling ratio of the second camera and the scaling ratio of the first camera is shown in the following equation (4): e w =s1-s2
[0086] Specifically, by substituting equations (2) and (3) into equation (4), we can obtain equation (1).
[0087] Since the scaling parameters a, b, c, and d are unknown, the second relational function needs to be scaled and calibrated before obtaining it. Specifically, at each of the multiple calibration times, a reference image taken by the first camera and a reference image taken by the second camera are obtained; the ratio w0 of the width of the target object in the reference image taken by the first camera to the total width of the reference image taken by the first camera, and the number of stepper motors x0 of the second camera during the corresponding reference image shooting process are obtained to obtain a (w0, x0) matching pair; the obtained multiple pairs of (w0, x0) are input into equation (1) to obtain the values of the scaling parameters a, b, c, and d. In the figure, the matching pairs (w0, x0) are shown as a pyramid in Figure 3. The reference image corresponding to the matching pair at the bottom of the pyramid is closer to the camera during the shooting process than the reference image corresponding to the matching pair at the top of the pyramid. The relationship curve between the number of stepper motors and the scaling parameter of the second camera is shown in Figure 4. The horizontal axis (Motor Step) represents the number of stepper motors, and the vertical axis (Scale) represents the scaling parameter. The Original Data curve is the original relationship curve after calibration, and the Fitted Data curve is the test relationship curve obtained after testing the relationship between the number of stepper motors and the scaling parameter after calibration. As can be seen from the figure, the original relationship curve and the test relationship curve are basically consistent, indicating that the scaling parameter calibration method provided in this embodiment has high accuracy.
[0088] S210: Control the rotation of the rotatable lens of the second camera based on the target rotation angle, and control the second camera after the lens rotation to capture the target object based on the target scaling ratio.
[0089] Optionally, the current reference rotation angle and reference scaling ratio of the second camera are obtained, the reference rotation angle is compensated with the target rotation angle, the rotation of the rotatable lens is controlled based on the compensated rotation angle, the reference scaling ratio is compensated with the target scaling ratio, and the second camera is controlled to capture the target object based on the compensated scaling ratio.
[0090] In the above image capturing method, when both the first camera and the second camera can capture the target object, the method acquires a first current image captured by the first camera, a second current image captured by the second camera, and the current number of stepper motors of the second camera. An initial feature point is then obtained from each of the first and second current images to form a target feature point pair. The lens of the second camera is a rotatable lens. A first relationship function is obtained between the rotation angle of the second camera, the position coordinates of pixels in the image captured by the first camera, and the pixel values of pixels in the image captured by the second camera. The first positional relationship of pixels in the image captured by the first camera is consistent with the second positional relationship of pixels in the image captured by the second camera. The first positional relationship is that any pixel in the image captured by the first camera, relative to the first phase coordinates, is a coordinate of the first phase coordinates and the second phase value. The system calculates the positional relationship between the remaining pixels in the image captured by the first camera and the remaining pixels in the image captured by the second camera. The system inputs the position coordinates of the corresponding initial feature points in the first current image and the pixel values of the corresponding initial feature points in the second current image into a first relational function, outputting the target rotation angle. It then obtains a second relational function relating the scaling ratio, the proportion of the object captured by the first camera in the image captured by the first camera, and the number of stepper motors in the second camera. The system inputs the proportion of the target object in the first current image and the current number of stepper motors into the second relational function, outputting the target scaling ratio. Based on the target rotation angle, the system controls the rotation of the rotatable lens of the second camera, and based on the target scaling ratio, controls the second camera to capture the target object after the lens rotation. This method effectively improves the display effect by displaying the target object captured by the second camera based on the target rotation angle and target scaling ratio on the display screen.
[0091] In some embodiments, as shown in FIG5, an initial feature point is obtained from a first current image and a second current image respectively to form a target feature point pair, including:
[0092] S502, Obtain the pixel values of pixels in the first current image and the second current image.
[0093] S504. Based on the pixel value of any pixel, obtain the pixel value difference between the pixel and the remaining pixels within a preset range around the pixel, and determine the total number of remaining pixels corresponding to pixel value differences greater than a first difference threshold.
[0094] S506. When the total number is greater than the first preset number, the pixel is determined as the initial feature point, and a target feature point pair is formed based on the initial feature points in the first current image and the initial feature points in the second current image.
[0095] Optionally, the pixel value difference is used to characterize the brightness difference between a pixel and each remaining pixel within a preset range. If the total number of pixel value differences greater than the first difference threshold is greater than the first preset number, it indicates that the brightness difference between the pixel and most of the surrounding pixels is relatively obvious, the pixel is more prominent in the image, and it is easy to distinguish the pixel from the image.
[0096] In this embodiment, initial feature points are selected from the pixels in the two current images based on the pixel value difference to form a target feature point pair. This makes it easier to select initial feature points and makes the rotation angle compensation value determined based on the target feature point pair more accurate.
[0097] In some embodiments, forming a target feature point pair based on initial feature points in a first current image and initial feature points in a second current image includes: determining an initial feature point in the first current image as a first initial feature point, obtaining a first initial positional relationship between any one of the first initial feature points and the remaining first initial feature points in the first current image; determining an initial feature point in the second current image as a second initial feature point, obtaining a second initial positional relationship between any one of the second initial feature points and the remaining second initial feature points in the second current image; when there is a case where the first initial positional relationship and the second initial positional relationship are consistent, determining the first initial feature point and the second initial feature point corresponding to the consistent first initial positional relationship and the second initial positional relationship as an initial feature point pair; and selecting a target feature point pair from the initial feature point pairs.
[0098] Optionally, the positional relationship between two pixels in the same image can be determined based on their coordinates in the image coordinate system. Specifically, the positional relationship can be determined based on the difference in coordinate values of the two pixels on each coordinate axis of the image coordinate system. For example, there are 9 first initial feature points in the first current image and 9 second initial feature points in the second current image. For a given first initial feature point, the difference in coordinate values between the first initial feature point and each of the remaining 8 first initial feature points is obtained based on the position coordinates of the first initial feature point and the position coordinates of the remaining 8 first initial feature points. If there exists a second initial feature point whose corresponding coordinate value difference is exactly the same as the coordinate value difference between the first initial feature points, then the first initial feature point and the second initial feature point are determined as an initial feature point pair.
[0099] Optionally, an initial feature point pair can be randomly selected as the target feature point pair, or a target feature point pair can be selected from the initial feature point pairs according to a preset selection rule. This application embodiment does not specifically limit this.
[0100] In this embodiment, based on the consistency between the first initial positional relationship and the second initial positional relationship, an initial feature point pair is determined, and a target feature point pair is selected from the initial feature point pair, so that the determined initial feature point pair is more accurate, thereby making the subsequent rotation angle compensation value determined based on the target feature point pair more accurate.
[0101] In some embodiments, the process of obtaining the first relation function includes: when the first camera is at a first reference position, the second camera is at a second reference position, and the rotatable lens is at a reference angle position, obtaining at least one first reference image captured by the first camera within a preset time period, and multiple second reference images captured by the second camera within the preset time period; wherein the rotatable lens continuously rotates within the preset time period; selecting first reference feature points from the pixels in the first reference image, and obtaining a first reference positional relationship between any one first reference feature point and the remaining first reference feature points in the first reference image; selecting second reference feature points from the pixels in the second reference image, and obtaining a second reference positional relationship between any one second reference feature point and the remaining second reference feature points in the second reference image; when there is a first reference positional relationship and a second reference positional relationship that are consistent, determining the first reference feature points and second reference feature points corresponding to the consistent first reference positional relationship and the second reference positional relationship as a reference feature point pair; obtaining a third relation function between the rotation angle, the position coordinates of the pixels in the image captured by the first camera, the pixel value of the pixels in the image captured by the second camera, and the focal length of the second camera; and obtaining the first relation function based on the reference feature point pair and the third relation function.
[0102] Optionally, the method for selecting the first and second reference feature points is the same as the method for selecting the initial feature points, and the method for determining the reference feature point pairs is the same as the method for determining the initial feature point pairs.
[0103] Alternatively, when the second camera is mounted on a gimbal, the third relational function is as shown in equation (5):
[0104] In equation (5), e r Let a0, a1, ..., a be the rotation angles. 20 Here, x is the coordinate value of the corresponding pixel of the first camera on the x-axis in the image coordinate system, y is the coordinate value of the corresponding pixel of the first camera on the y-axis in the image coordinate system, p1 is the pixel value of the corresponding pixel of the second camera, o is the pixel value of the center point of the corresponding image of the second camera, f is the focal length of the second camera, and α is the angle measured by the gimbal.
[0105] In equation (5), the rotation parameters a0, a1, ..., a 20Since both the rotation parameter and the focal length f are unknown parameters, it is necessary to perform rotation calibration on equation (5) to determine the rotation parameter and the focal length. After determining the rotation parameter and the focal length, equation (5) can be determined as the first relational function.
[0106] Optionally, the preset time period is the time period for rotation calibration. Since the lens of the first camera does not rotate, it can be assumed that the image captured by the first camera will not change during the rotation calibration process, and capturing only one first reference image is sufficient to meet the requirements of rotation calibration.
[0107] In this embodiment, a pair of reference feature points is determined based on the consistency between the first reference position relationship and the second reference position relationship. Then, a rotation calibration is performed based on the pair of reference feature points and the third relationship function to obtain the first relationship function. The first relationship function obtained in this way is more accurate, which makes the subsequently determined rotation angle compensation value more accurate.
[0108] In some embodiments, obtaining a third relationship function among the rotation angle, the position coordinates of pixels in the image captured by the first camera, the pixel value of pixels in the image captured by the second camera, and the focal length of the second camera includes: obtaining a fourth relationship function among the angle value of pixels in the image captured by the first camera and the position coordinates of pixels in the image captured by the first camera; a fifth relationship function among the angle value of pixels in the image captured by the second camera, the pixel value of pixels in the image captured by the second camera, and the focal length of the second camera; and a sixth relationship function among the rotation angle, the angle value of pixels in the image captured by the first camera, and the angle value of pixels in the image captured by the second camera; substituting the fourth and fifth relationship functions into the sixth relationship function yields the third relationship function.
[0109] The angle value of a pixel refers to the directional angle of the pixel relative to the camera's imaging plane.
[0110] Optionally, the fourth relational function is shown in equation (6): θ1=a0+a1x+a2y+a3x 2 +a4xy+a5y 2 +a6x 3 +a7x 2 y+a8xy 2 +a9y 3 +a 10 x 4 +a 11 x 3 y+a 12 x 2 y 2 +a 13 xy 3 +a 14 y4 +a 15 x 5 +a 16 x 4 y+a 17 x 3 y 2 +a 18 x 2 y 3 +a 19 xy 4 +a 20 y 5
[0111] In equation (6), θ1 is the angle value of the corresponding pixel of the first camera.
[0112] Alternatively, the fifth relational function is shown in equation (7) below:
[0113] In equation (7), θ2 is the angle value of the corresponding pixel point of the second camera.
[0114] Alternatively, when the second camera is mounted on a gimbal, the sixth relational function is as shown in equation (8): e r =θ1-θ2-α
[0115] The relationship between θ1 and θ2 in equation (8) is shown in Figure 6. In Figure 6, θ = F(p2) represents θ1, θ = H(f, p1) represents θ2, and p1 and p2 represent the corresponding pixel points of the second camera and the first camera, respectively.
[0116] Substituting equations (6) and (7) into equation (8) yields the third relational function.
[0117] In this embodiment, the fourth and fifth relation functions are substituted into the sixth relation function to obtain the third relation function, making the obtained third relation function more accurate, thereby making the first relation function obtained based on the third relation function more accurate.
[0118] In some embodiments, obtaining a first relational function based on a pair of reference feature points and a third relational function includes: for the pair of reference feature points, obtaining the current focal length of the second camera based on the position coordinates of the first reference feature point, the pixel value of the second reference feature point, and the third relational function; and obtaining the first relational function based on the current focal length and the third relational function.
[0119] Optionally, during the rotation calibration process, multiple pairs of reference feature points are obtained. The position coordinates of the first reference feature point and the pixel value of the second reference feature point in each pair are input into a third relational function to obtain a system of equations. By solving this system of equations, the rotation parameters a0, a1, ..., a0 can be obtained.20 The current focal length of the second camera, with rotation parameters a0, a1, ..., a 20 The first relation function can be obtained by substituting the current focal length of the second camera into the third relation function. A schematic diagram of the rotation calibration is shown in Figure 7. In Figure 7, a triangle represents a pair of reference feature points, the surface represents the calibration visualization surface, and the combined graphic composed of two triangles represents the second camera, which is an optical zoom camera in modern conference scenarios.
[0120] During the rotation calibration process, the convergence status of the third relation function is acquired in real time. If the convergence meets the preset convergence conditions, the calibration process stops, and the obtained rotation parameters a0, a1, ..., a... are recorded. 20 Substituting the current focal length into the third relational function yields the first relational function. Optionally, the convergence of the third relational function can be determined by calculating the Jacobian of the rotation parameters, the Jacobian of the second camera's focal length about the x-axis of the image coordinate system, and the Jacobian of the second camera's focal length about the y-axis of the image coordinate system. Convergence of the third relational function is determined if all three Jacobians are within their respective preset ranges. Here, the Jacobian is a matrix composed of first-order partial derivatives arranged in a specific order.
[0121] Specifically, let x = p1 - o, By differentiation, we can obtain:
[0122] Calculate using the chain rule. The Jacobian of the second camera's focal length with respect to the x-axis of the image coordinate system can be obtained as follows:
[0123] Similarly, the Jacobian of the focal length of the second camera with respect to the y-axis of the image coordinate system can be calculated as follows:
[0124] The Jacobian of the rotation parameter is:
[0125] Combining the three Jacobians mentioned above, we can obtain the factor diagram of the first and second cameras during the rotation calibration process, as shown in Figure 8. In Figure 8, the top triangle represents the gimbal measurement angle value and the reference feature point pair; the middle square represents the Jacobian of the rotation parameters (rotation factor), the Jacobian of the second camera's focal length about the x-axis of the image coordinate system (focal length factor), and the Jacobian of the second camera's focal length about the y-axis of the image coordinate system (focal length factor); the ellipse in the lower left corner represents the rotation parameters, with a circle within the ellipse representing a rotation parameter; and the circle in the lower right corner represents the focal length. The obtained current focal length can be verified using the prior focal length factor and the prior focal length measurement value.
[0126] In this embodiment, the current focal length of the second camera is obtained based on the position coordinates of the first reference feature point, the pixel value of the second reference feature point, and the third relation function. Based on the current focal length and the third relation function, the first relation function is obtained, making the obtained first relation function more accurate, thereby making the rotation angle compensation value determined based on the first relation function more accurate.
[0127] In some embodiments, the process of acquiring the first reference position, the second reference position, and the reference angle position includes: when both the first camera and the second camera can capture the target object, acquiring a first preset image captured by the first camera and a second preset image captured by the second camera; selecting first preset feature points from the pixels in the first preset image, and acquiring a first preset positional relationship between any one of the first preset feature points and the remaining first preset feature points in the first preset image; selecting second preset feature points from the pixels in the second preset image, and acquiring a second preset positional relationship between any one of the second preset feature points and the remaining second preset feature points in the second preset image; and when there is a case where the first preset positional relationship and the second preset positional relationship are consistent. The first and second preset feature points corresponding to the consistent first and second preset position relationships are determined as preset feature point pairs. When the number of preset feature point pairs is greater than the second preset number, for each preset feature point pair, the pixel value difference between the first and second preset feature points is obtained, and the total number of preset feature point pairs corresponding to pixel value differences less than the second difference threshold is determined. When the total number is greater than the third preset number, the first preset position of the first camera, the second preset position of the second camera, and the preset angle position of the rotatable lens are obtained, and the first preset position is determined as the first reference position, the second preset position is determined as the second reference position, and the preset angle position is determined as the reference angle position.
[0128] Since the first and second cameras can be in different positions, and the angular position of the second camera's rotatable lens can also be in different positions before rotation, these different positions will lead to different pairs of reference feature points acquired during rotation calibration, resulting in deviations in calibration accuracy. Therefore, before rotation calibration, it is necessary to determine the first reference position, second reference position, and reference angle position that maximize calibration accuracy based on the number of feature point pairs acquired under different positions and the feature point difference between the first and second feature points in each pair. The first preset position, second preset position, and preset angle position are respectively the optimal positions for the first camera, the second camera, and the rotatable lens of the second camera.
[0129] Optionally, the formula for the pixel value difference between the first preset feature point and the second preset feature point is t = pm -p l , where p m p is the pixel value of the first preset feature point. l t is the pixel value of the second preset feature point, and t is the pixel value difference.
[0130] Optionally, the selection method for the first and second preset feature points is the same as the selection method for the initial feature points, and the determination method for the preset feature point pairs is the same as the determination method for the initial feature point pairs.
[0131] In this embodiment, by determining the first preset position as the first reference position, the second preset position as the second reference position, and the preset angle position as the reference angle position, the subsequent rotation calibration results are more accurate, thereby making the first relational function determined based on the calibration results more accurate.
[0132] In one embodiment, another image-capturing method is provided for use in modern meeting scenarios. This method includes the following:
[0133] (1) For the camera module and optical zoom camera of the conference tablet, when the optical zoom camera is mounted on the gimbal, an optimized loss function of pixel-focal length-angle is constructed, where the optimized loss function is shown in Equation (5); the first feature point and the second feature point are obtained from an image taken by the camera module and multiple images taken by the optical zoom camera, respectively, and the optimized loss function is rotated and calibrated based on the first feature point and the second feature point; the rotation angle compensation value is obtained based on the optimized loss function after rotation calibration. The rotation calibration process is shown in Figure 9.
[0134] (2) Construct an error function between the width of the human face and the number of stepper motors, where the error function is shown in equation (1); scale and calibrate the error function, and obtain the scaling ratio compensation value based on the calibrated error function.
[0135] (3) Control the rotation of the optical zoom camera based on the rotation angle compensation value, thereby driving the camera lens to rotate, and control the optical zoom camera after the lens rotation based on the scaling ratio compensation value to take pictures of the speaker; display the close-up picture of the speaker on the display screen.
[0136] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0137] Based on the same inventive concept, this application also provides an image capturing apparatus for implementing the image capturing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more image capturing apparatus embodiments provided below can be found in the limitations of the image capturing method described above, and will not be repeated here.
[0138] In an exemplary embodiment, as shown in FIG10, an image capturing device 1000 is provided, including: a first acquisition module 1001, a second acquisition module 1002, an input module 1003, a third acquisition module 1004, and a control module 1005, wherein:
[0139] The first acquisition module 1001 is used to acquire a first current image captured by the first camera, a second current image captured by the second camera, and the current number of stepper motors of the second camera when both the first camera and the second camera can capture the target object, and to acquire an initial feature point from the first current image and the second current image respectively to form a target feature point pair. The lens of the second camera is a rotatable lens.
[0140] The second acquisition module 1002 is used to acquire a first relationship function between the rotation angle of the second camera, the position coordinates of the pixels in the image captured by the first camera, and the pixel value of the pixels in the image captured by the second camera; wherein the first position relationship of the pixels in the image captured by the first camera is consistent with the second position relationship of the pixels in the image captured by the second camera, the first position relationship is the position relationship between any pixel in the image captured by the first camera and the remaining pixels in the image captured by the first camera, and the second position relationship is the position relationship between any pixel in the image captured by the second camera and the remaining pixels in the image captured by the second camera.
[0141] The input module 1003 is used to input the position coordinates of the corresponding initial feature point of the first current image and the pixel value of the corresponding initial feature point of the second current image into the first relational function, and output the target rotation angle.
[0142] The third acquisition module 1004 is used to acquire a second relationship function between the scaling ratio, the proportion of the object captured by the first camera in the image captured by the first camera, and the number of stepper motors of the second camera. The proportion of the target object in the first current image and the current number of stepper motors are input into the second relationship function, and the target scaling ratio is output.
[0143] The control module 1005 is used to control the rotation of the rotatable lens of the second camera based on the target rotation angle, and to control the second camera after the lens rotation to capture the target object based on the target scaling ratio.
[0144] In some embodiments, the first acquisition module 1001 includes:
[0145] The first acquisition unit is used to acquire the pixel values of pixels in the first current image and the second current image.
[0146] The second acquisition unit is used to acquire the pixel value difference between a pixel and the remaining pixels within a preset range around the pixel based on the pixel value of any pixel, and to determine the total number of remaining pixels corresponding to pixel value differences greater than a first difference threshold.
[0147] The first determining unit is used to determine the pixel as the initial feature point when the total number is greater than the first preset number, and to form a target feature point pair based on the initial feature points in the first current image and the initial feature points in the second current image.
[0148] In some embodiments, the first determining unit is further configured to: determine an initial feature point in a first current image as a first initial feature point; obtain a first initial positional relationship between any one of the first initial feature points and the remaining first initial feature points in the first current image; determine an initial feature point in a second current image as a second initial feature point; obtain a second initial positional relationship between any one of the second initial feature points and the remaining second initial feature points in the second current image; if there is a first initial positional relationship and a second initial positional relationship that are consistent, determine the first initial feature point and the second initial feature point corresponding to the consistent first initial positional relationship and the second initial positional relationship as an initial feature point pair; and select a target feature point pair from the initial feature point pairs.
[0149] In some embodiments, the second acquisition module 1002 includes:
[0150] The third acquisition unit is configured to acquire at least one first reference image captured by the first camera within a preset time period and multiple second reference images captured by the second camera within a preset time period, when the first camera is at a first reference position, the second camera is at a second reference position, and the rotatable lens is at a reference angle position; wherein the rotatable lens rotates continuously within the preset time period.
[0151] The first selection unit is used to select a first reference feature point from the pixels in the first reference image and obtain the first reference position relationship between any one first reference feature point and the remaining first reference feature points in the first reference image.
[0152] The second selection unit is used to select second reference feature points from the pixels in the second reference image and obtain the second reference position relationship between any second reference feature point and the remaining second reference feature points in the second reference image.
[0153] The second determining unit is used to determine the first and second reference feature points corresponding to the consistent first and second reference position relationships as a pair of reference feature points when there is a first reference position relationship and a second reference position relationship that are consistent.
[0154] The fourth acquisition unit is used to acquire a third relationship function between the rotation angle, the position coordinates of the pixels in the image captured by the first camera, the pixel value of the pixels in the image captured by the second camera, and the focal length of the second camera.
[0155] The fifth acquisition unit is used to acquire the first relation function based on the benchmark feature point pair and the third relation function.
[0156] In some embodiments, the fourth acquisition unit is further configured to acquire a fourth relationship function between the angle value of a pixel in the image captured by the first camera and the position coordinates of the pixel in the image captured by the first camera, a fifth relationship function between the angle value of a pixel in the image captured by the second camera, the pixel value of the pixel in the image captured by the second camera and the focal length of the second camera, and a sixth relationship function between the rotation angle, the angle value of a pixel in the image captured by the first camera and the angle value of a pixel in the image captured by the second camera; and substitute the fourth and fifth relationship functions into the sixth relationship function to obtain the third relationship function.
[0157] In some embodiments, the fifth acquisition unit is further configured to, for a pair of reference feature points, acquire the current focal length of the second camera based on the position coordinates of the first reference feature point, the pixel value of the second reference feature point, and the third relation function; and acquire the first relation function based on the current focal length of the second camera and the third relation function.
[0158] The second acquisition module 1002 is further configured to, when both the first camera and the second camera can capture the target object, acquire a first preset image captured by the first camera and a second preset image captured by the second camera; select first preset feature points from the pixels in the first preset image, and acquire a first preset positional relationship between any one first preset feature point and the remaining first preset feature points in the first preset image; select second preset feature points from the pixels in the second preset image, and acquire a second preset positional relationship between any one second preset feature point and the remaining second preset feature points in the second preset image; and, if there exists a first preset positional relationship that is consistent with a second preset positional relationship, assign the consistent first preset positional relationship to the second preset image. The first and second preset feature points corresponding to the relationship with the second preset position are determined as preset feature point pairs. When the number of preset feature point pairs is greater than the second preset number, for each preset feature point pair, the pixel value difference between the first and second preset feature points is obtained, and the total number of preset feature point pairs corresponding to pixel value differences less than the second difference threshold is determined. When the total number is greater than the third preset number, the first preset position of the first camera, the second preset position of the second camera, and the preset angle position of the rotatable lens are obtained, and the first preset position is determined as the first reference position, the second preset position is determined as the second reference position, and the preset angle position is determined as the reference angle position.
[0159] Each module in the aforementioned image capturing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0160] In an exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram is shown in Figure 11. The computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an image capturing method.
[0161] Those skilled in the art will understand that the structure shown in Figure 11 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0162] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0163] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0164] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0165] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0166] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0167] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. An image capturing method, characterized in that, The method includes: When both the first camera and the second camera can capture the target object, the first current image captured by the first camera, the second current image captured by the second camera, and the current number of stepper motors of the second camera are acquired. An initial feature point is obtained from the first current image and the second current image respectively to form a target feature point pair. The lens of the second camera is a rotatable lens. A first relationship function is obtained between the rotation angle of the second camera, the position coordinates of pixels in the image captured by the first camera, and the pixel value of pixels in the image captured by the second camera; wherein, the first position relationship of pixels in the image captured by the first camera is consistent with the second position relationship of pixels in the image captured by the second camera, the first position relationship is the position relationship between any pixel in the image captured by the first camera and the remaining pixels in the image captured by the first camera, and the second position relationship is the position relationship between any pixel in the image captured by the second camera and the remaining pixels in the image captured by the second camera; The position coordinates of the corresponding initial feature point in the first current image and the pixel value of the corresponding initial feature point in the second current image are input into the first relational function to output the target rotation angle. Obtain a second relationship function between the scaling ratio, the proportion of the object captured by the first camera in the image captured by the first camera, and the number of stepper motors of the second camera; input the proportion of the target object in the first current image and the current number of stepper motors into the second relationship function, and output the target scaling ratio; The rotation of the rotatable lens of the second camera is controlled based on the target rotation angle, and the second camera, after the lens rotation is controlled based on the target scaling ratio, captures the target object.
2. The method according to claim 1, characterized in that, The step of obtaining an initial feature point from the first current image and the second current image respectively to form a target feature point pair includes: Obtain the pixel values of pixels in the first current image and the second current image; Based on the pixel value of any pixel, obtain the pixel value difference between the pixel and the remaining pixels within a preset range around the pixel, and determine the total number of remaining pixels corresponding to pixel value differences greater than a first difference threshold. If the total number is greater than the first preset number, the pixel is determined as the initial feature point, and a target feature point pair is formed based on the initial feature points in the first current image and the initial feature points in the second current image.
3. The method according to claim 2, characterized in that, The step of forming target feature point pairs based on initial feature points in the first current image and initial feature points in the second current image includes: The initial feature point in the first current image is determined as the first initial feature point, and the first initial positional relationship between any first initial feature point and the remaining first initial feature points in the first current image is obtained. The initial feature point in the second current image is determined as the second initial feature point, and the second initial positional relationship between any one of the second initial feature points and the remaining second initial feature points in the second current image is obtained; When there is a first initial positional relationship and a second initial positional relationship that are consistent, the first initial feature points and the second initial feature points corresponding to the consistent first initial positional relationship and the second initial positional relationship are determined as initial feature point pairs; Select the target feature point pair from the initial feature point pair.
4. The method according to claim 1, characterized in that, The process of obtaining the first relational function includes: When the first camera is at a first reference position, the second camera is at a second reference position, and the rotatable lens is at a reference angle position, at least one first reference image captured by the first camera within a preset time period and multiple second reference images captured by the second camera within the preset time period are acquired; wherein, the rotatable lens continuously rotates within the preset time period. Select a first reference feature point from the pixels in the first reference image, and obtain the first reference position relationship between any one first reference feature point and the remaining first reference feature points in the first reference image; Select a second reference feature point from the pixels in the second reference image, and obtain the second reference position relationship between any one second reference feature point and the remaining second reference feature points in the second reference image; When there is a first reference position relationship and a second reference position relationship that are consistent, the first reference feature points and the second reference feature points corresponding to the consistent first reference position relationship and the second reference position relationship are determined as a reference feature point pair; A third relationship function is used to obtain the relationship between the rotation angle, the position coordinates of the pixels in the image captured by the first camera, the pixel value of the pixels in the image captured by the second camera, and the focal length of the second camera. The first relation function is obtained based on the benchmark feature point pair and the third relation function.
5. The method according to claim 4, characterized in that, The third relationship function for obtaining the relationship between the rotation angle, the position coordinates of pixels in the image captured by the first camera, the pixel value of pixels in the image captured by the second camera, and the focal length of the second camera includes: The fourth relationship function is obtained between the angle value of a pixel in the image captured by the first camera and the position coordinate of the pixel in the image captured by the first camera; the fifth relationship function is obtained between the angle value of a pixel in the image captured by the second camera, the pixel value of the pixel in the image captured by the second camera and the focal length of the second camera; and the sixth relationship function is obtained between the rotation angle, the angle value of a pixel in the image captured by the first camera and the angle value of a pixel in the image captured by the second camera. Substituting the fourth and fifth relation functions into the sixth relation function yields the third relation function.
6. The method according to claim 4, characterized in that, The step of obtaining the first relation function based on the reference feature point pair and the third relation function includes: For the reference feature point pair, the current focal length of the second camera is obtained based on the position coordinates of the first reference feature point, the pixel value of the second reference feature point, and the third relation function; Based on the current focal length of the second camera and the third relation function, the first relation function is obtained.
7. The method according to claim 4, characterized in that, The process of obtaining the first reference position, the second reference position, and the reference angle position includes: When both the first camera and the second camera can capture the target object, acquire a first preset image captured by the first camera and a second preset image captured by the second camera; Select a first preset feature point from the pixels in the first preset image, and obtain a first preset positional relationship between any one first preset feature point and the remaining first preset feature points in the first preset image; Select a second preset feature point from the pixels in the second preset image, and obtain a second preset positional relationship between any one second preset feature point and the remaining second preset feature points in the second preset image; When there is a first preset positional relationship and a second preset positional relationship that are consistent, the first preset feature points and the second preset feature points corresponding to the consistent first preset positional relationship and the second preset positional relationship are determined as preset feature point pairs; When the number of preset feature point pairs is greater than the second preset number, for each preset feature point pair, the pixel value difference between the first preset feature point and the second preset feature point is obtained, and the total number of preset feature point pairs whose pixel value difference is less than the second difference threshold is determined. When the total number is greater than the third preset number, the first preset position of the first camera, the second preset position of the second camera, and the preset angle position of the rotatable lens are obtained. The first preset position is determined as the first reference position, the second preset position is determined as the second reference position, and the preset angle position is determined as the reference angle position.
8. An image capturing device, characterized in that, The device includes: The first acquisition module is used to acquire a first current image captured by the first camera, a second current image captured by the second camera, and the current number of stepper motors of the second camera when both the first camera and the second camera can capture the target object, and to acquire an initial feature point from the first current image and the second current image respectively to form a target feature point pair. The lens of the second camera is a rotatable lens. The second acquisition module is used to acquire a first relationship function between the rotation angle of the second camera, the position coordinates of pixels in the image captured by the first camera, and the pixel value of pixels in the image captured by the second camera; wherein the first position relationship of pixels in the image captured by the first camera is consistent with the second position relationship of pixels in the image captured by the second camera, the first position relationship is the position relationship between any pixel in the image captured by the first camera and the remaining pixels in the image captured by the first camera, and the second position relationship is the position relationship between any pixel in the image captured by the second camera and the remaining pixels in the image captured by the second camera; The input module is used to input the position coordinates of the corresponding initial feature points of the first current image and the pixel values of the corresponding initial feature points of the second current image into the first relational function, and output the target rotation angle; The third acquisition module is used to acquire a second relationship function between the scaling ratio, the proportion of the object captured by the first camera in the image captured by the first camera, and the number of stepper motors of the second camera. The proportion of the target object in the first current image and the current number of stepper motors are input into the second relationship function, and the target scaling ratio is output. The control module is used to control the rotation of the rotatable lens of the second camera based on the target rotation angle, and to control the second camera after the lens rotation to capture the target object based on the target scaling ratio.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.