Panoramic image stitching method, pan-tilt camera control method, apparatus, and electronic device

By acquiring a set of images captured by a PTZ camera with a pitch angle greater than 0, and using PTZ rotation data and calibration results, the mapping relationship between the images and the panoramic image is determined. This solves the problem of the limited vertical field of view of the PTZ camera when the pitch angle is 0, and generates a panoramic image containing the PTZ camera.

WO2026001601A1PCT designated stage Publication Date: 2026-01-02SHENZHEN OCEANWING SMART INNOVATIONS TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/099124
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-27
Filing Date
2025-06-04
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

In existing technologies, when a PTZ camera captures a panoramic image with a pitch angle of 0, the vertical field of view is limited, making it difficult to synthesize a panoramic image with a larger field of view.

Method used

By acquiring a set of images captured by a PTZ camera with a pitch angle greater than 0, and based on PTZ rotation data and calibration results, the mapping relationship between the images and the panoramic image is determined, and the images are stitched together to generate a panoramic image.

Benefits of technology

It achieves the generation of panoramic images containing PTZ cameras without increasing the vertical field of view of the lens, improves the shooting angle range, reduces the rotation restrictions on PTZ cameras, improves shooting flexibility, generates rotation data containing PTZ cameras, determines the accurate position of the image in the panoramic image, and generates a panoramic image containing the entire field of view of PTZ cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025099124_02012026_PF_FP_ABST
    Figure CN2025099124_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to a panoramic image stitching method, a pan-tilt camera control method, an apparatus, and an electronic device. The panoramic image stitching method comprises: acquiring a set of images shot by a pan-tilt camera, a calibration result of the pan-tilt camera, and pan-tilt rotation data corresponding to images in the set of images, the field-of-view range of the set of images comprising the field-of-view range of the pan-tilt camera, the set of images comprising: an image shot by the pan-tilt camera when having a tilt angle greater than 0, and the pan-tilt rotation data representing rotation data of the pan-tilt camera when shooting corresponding images; and, on the basis of the pan-tilt rotation data and the calibration result, stitching the images in the set of images to obtain a panoramic image. Thus, panoramic images having larger field-of-view ranges can be synthesized.
Need to check novelty before this filing date? Find Prior Art

Description

Panorama image splicing and control method and device of pan-tilt camera and electronic equipment

[0001] The present application claims priority to the Chinese patent application No. 202410854880.2, filed on June 27, 2024, and entitled "Panorama image splicing and control method and device of pan-tilt camera and electronic equipment", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to the field of security and protection, and in particular to a panorama image splicing and control method and device of pan-tilt camera and electronic equipment. BACKGROUND

[0003] In related technologies, it is usually necessary to control the pan-tilt camera to rotate horizontally by 360 degrees at an elevation angle of 0 (i.e., the lens is facing forward) to capture a plurality of images, and then to synthesize a panorama image.

[0004] However, the vertical field of view range of the panorama image is limited by the vertical field of view range of the lens. If a panorama image with a larger vertical field of view range is desired, it is usually necessary to increase the vertical field of view angle of the lens, such as using a fisheye lens with a field of view angle close to 180 degrees on a panoramic camera.

[0005] In practice, it is necessary to synthesize a panorama image with a larger field of view range, because the pan-tilt has a certain ability to rotate at an elevation angle. If only images with an elevation angle of 0 are used to synthesize a panorama image, it will not be possible to preview the view after the pan-tilt is rotated upward / downward.

[0006] Therefore, how to synthesize a panorama image with a larger field of view range is a technical problem worthy of attention. SUMMARY

[0007] In view of this, in order to solve one or more of the above technical problems, the present application provides a panorama image splicing and control method and device of pan-tilt camera and electronic equipment.

[0008] In a first aspect, the present application provides a panorama image splicing method, comprising:

[0009] obtaining an image set captured by a pan-tilt camera, a calibration result of the pan-tilt camera, and pan-tilt rotation data corresponding to images in the image set; wherein a field of view range of the image set contains a field of view range of the pan-tilt camera; the image set contains images captured by the pan-tilt camera at an elevation angle greater than 0; the pan-tilt rotation data represents the rotation data of the pan-tilt camera when the corresponding image is captured;

[0010] determine a mapping relationship between images in the image set and a panorama to be generated by stitching based on the pan-tilt rotation data and the calibration result;

[0011] stitch the images in the image set based on the mapping relationship to obtain the panorama.

[0012] In one possible implementation, the determining of the mapping relationship between the images in the image set and the panorama to be generated by stitching based on the pan-tilt rotation data and the calibration result comprises:

[0013] determining image rotation data of the images in the image set relative to the panorama to be generated by stitching based on the pan-tilt rotation data, wherein the image rotation data represents positions of the images in the image set relative to the panorama to be generated by stitching;

[0014] determining image intrinsic data of the images in the image set based on the calibration result;

[0015] determining the mapping relationship between the images in the image set and the panorama to be generated by stitching based on the image rotation data and the image intrinsic data.

[0016] In one possible implementation, the determining of the image rotation data of the images in the image set relative to the panorama to be generated by stitching based on the pan-tilt rotation data comprises:

[0017] determining an overlapping area between a first image and a second image in the image set based on the pan-tilt rotation data and sizes of the images in the image set; the first image and the second image are any two images in the image set having an overlapping area;

[0018] determining respective detection areas of the first image and the second image based on the overlapping area, wherein the detection areas are used to detect whether there are matching feature point pairs in the first image and the second image;

[0019] determining whether there are matching feature point pairs in the detection area of the first image and the detection area of the second image;

[0020] In the absence of the matched feature point pair, third and fourth images are determined from the first and second images, wherein the third image is an image for which image rotation data is to be determined, and the fourth image is an image for which image rotation data has been determined; based on image rotation data of the fourth image relative to the panoramic image to be generated by stitching and relative rotation data, image rotation data of the third image relative to the panoramic image to be generated by stitching is determined, wherein the relative rotation data represents a rotation relationship between the gimbal rotation data corresponding to the fourth image and the gimbal rotation data corresponding to the third image.

[0021] In one possible implementation, after the determination of whether there is a matched feature point pair in the detection region of the first image and the detection region of the second image, the method further includes:

[0022] In the presence of the matched feature point pair, based on the matched feature point pair, relative rotation data between the gimbal rotation data corresponding to the first image and the gimbal rotation data corresponding to the second image is determined;

[0023] Based on the relative rotation data between the gimbal rotation data corresponding to the first image and the gimbal rotation data corresponding to the second image, image rotation data of the first and second images relative to the panoramic image to be generated by stitching is determined respectively.

[0024] In one possible implementation, the gimbal rotation data corresponding to the third image is denoted as (h1, v1), h1 represents rotation data of the gimbal camera in the horizontal direction when the third image is captured, and v1 represents rotation data of the gimbal camera in the vertical direction when the third image is captured, the gimbal rotation data corresponding to the fourth image is denoted as (h2, v2), h2 represents rotation data of the gimbal camera in the horizontal direction when the fourth image is captured, and v2 represents rotation data of the gimbal camera in the vertical direction when the fourth image is captured; and

[0025] The determination of the relative rotation data includes:

[0026] (0, cos(v1 ÷ 180 x p), -sin(v1 ÷ 180 x p)) is determined as the first rotation axis, and h2-h1 is determined as the first rotation angle;

[0027] (1, 0, 0) is determined as the second rotation axis, and v2-v1 is determined as the second rotation angle;

[0028] determine relative rotation data between the third image corresponding cloud platform rotation data and the fourth image corresponding cloud platform rotation data based on the first rotation axis, the first rotation angle, the second rotation axis and the second rotation angle.

[0029] In one possible implementation, the determining the detection region of the first image and the detection region of the second image based on the overlapping region comprises:

[0030] determining a region in the first image containing the overlapping region and having a size greater than the size of the overlapping region as the detection region of the first image;

[0031] determining a region in the second image containing the overlapping region and having a size greater than the size of the overlapping region as the detection region of the second image.

[0032] In the second aspect, the embodiments of the present application provide a control method of a cloud platform camera, and the method comprises:

[0033] obtaining a target position selected from a panorama, wherein the panorama is obtained by using any panorama splicing method according to the first aspect;

[0034] controlling the cloud platform camera to rotate to a first mapping position of the target position in a space where the cloud platform camera is located.

[0035] In one possible implementation, the controlling the cloud platform camera to rotate to the first mapping position of the target position in the space where the cloud platform camera is located comprises:

[0036] determining a second mapping position of the target position in each fifth image in the image set;

[0037] for each fifth image, determining a distance between the second mapping position of the fifth image and an image center of the fifth image to obtain a corresponding distance of each fifth image;

[0038] taking an image center of a third image with the smallest corresponding distance in the image set as a target center, and determining target rotation data of the cloud platform camera based on a field of view angle of the target center and the target image corresponding cloud platform rotation data, wherein the target rotation data represents a rotation angle of the cloud platform camera in a case that an optical axis of the cloud platform camera moves to the first mapping position;

[0039] controlling the cloud platform camera to rotate according to the target rotation data, so that the cloud platform camera rotates to the first mapping position of the target position in the space where the cloud platform camera is located.

[0040] In a third aspect, an embodiment of the present application provides a panoramic image stitching device, the device comprising:

[0041] a first obtaining unit configured to obtain a set of images captured by a pan-tilt camera, a calibration result of the pan-tilt camera, and pan-tilt rotation data corresponding to images in the set of images; wherein a field of view range of the set of images contains a field of view range of the pan-tilt camera; the set of images contains images captured by the pan-tilt camera when the tilt angle is greater than 0; the pan-tilt rotation data represents rotation data of the pan-tilt camera when the corresponding image is captured;

[0042] a first determining unit configured to determine a mapping relationship between images in the set of images and a panoramic image to be stitched based on the pan-tilt rotation data and the calibration result;

[0043] a stitching unit configured to stitch the images in the set of images based on the mapping relationship to obtain the panoramic image.

[0044] In one possible implementation, the determining of the mapping relationship between the images in the set of images and the panoramic image to be stitched based on the pan-tilt rotation data and the calibration result comprises:

[0045] determining image rotation data of the images in the set of images relative to the panoramic image to be stitched based on the pan-tilt rotation data, wherein the image rotation data represents a position of the images in the set of images relative to the panoramic image to be stitched;

[0046] determining image intrinsic data of the images in the set of images based on the calibration result;

[0047] determining the mapping relationship between the images in the set of images and the panoramic image to be stitched based on the image rotation data and the image intrinsic data.

[0048] In one possible implementation, the determining of the image rotation data of the images in the set of images relative to the panoramic image to be stitched based on the pan-tilt rotation data comprises:

[0049] determining an overlapping area between a first image and a second image in the set of images based on the pan-tilt rotation data and a size of the image in the set of images; the first image and the second image are any two images in the set of images having an overlapping area;

[0050] determining a detection area of each of the first image and the second image based on the overlapping area, wherein the detection area is used to detect whether there is a matching feature point pair in the first image and the second image;

[0051] determining whether there is a matched feature point pair in the detection region of the first image and the detection region of the second image;

[0052] in the absence of a matched feature point pair, determining a third image and a fourth image from the first image and the second image, wherein the third image is an image whose image rotation data is to be determined, and the fourth image is an image whose image rotation data has been determined;

[0053] determining relative rotation data, wherein the relative rotation data represents a rotation relationship between the gimbal rotation data corresponding to the fourth image and the gimbal rotation data corresponding to the third image;

[0054] determining image rotation data of the third image relative to the panoramic image to be generated based on image rotation data of the fourth image relative to the panoramic image to be generated and the relative rotation data, wherein the relative rotation data represents a rotation relationship between the gimbal rotation data corresponding to the fourth image and the gimbal rotation data corresponding to the third image.

[0055] In one possible implementation, after the determination of whether there is a matched feature point pair in the detection region of the first image and the detection region of the second image, the apparatus further comprises:

[0056] a second determination unit, configured to, in the presence of a matched feature point pair, determine, based on the matched feature point pair, relative rotation data between the gimbal rotation data corresponding to the first image and the gimbal rotation data corresponding to the second image;

[0057] a third determination unit, configured to determine, based on the relative rotation data between the gimbal rotation data corresponding to the first image and the gimbal rotation data corresponding to the second image, image rotation data of the first image and the second image relative to the panoramic image to be generated, respectively.

[0058] In one possible implementation, the gimbal rotation data corresponding to the third image is denoted as (h1, v1), h1 represents rotation data of the gimbal camera in the horizontal direction when the third image is captured, and v1 represents rotation data of the gimbal camera in the vertical direction when the third image is captured, the gimbal rotation data corresponding to the fourth image is denoted as (h2, v2), h2 represents rotation data of the gimbal camera in the horizontal direction when the fourth image is captured, and v2 represents rotation data of the gimbal camera in the vertical direction when the fourth image is captured; and

[0059] The determination of the relative rotation data comprises:

[0060] determining (0, cos(v1 ÷ 180 x p), -sin(v1 ÷ 180 x p)) as the first rotation axis and h2-h1 as the first rotation angle;

[0061] determining (1, 0, 0) as the second rotation axis and v2-v1 as the second rotation angle;

[0062] determining relative rotation data between the third image corresponding pan-tilt rotation data and the fourth image corresponding pan-tilt rotation data based on the first rotation axis, the first rotation angle, the second rotation axis and the second rotation angle.

[0063] In one possible implementation, the determining the detection area of the first image and the second image respectively based on the overlapping area comprises:

[0064] determining an area in the first image containing the overlapping area and having a size larger than that of the overlapping area as the detection area of the first image;

[0065] determining an area in the second image containing the overlapping area and having a size larger than that of the overlapping area as the detection area of the second image.

[0066] In the fourth aspect, the embodiments of the present application provide a control device of a pan-tilt camera, the device comprising:

[0067] a second acquisition unit configured to acquire a target position selected from a panorama, wherein the panorama is obtained by using any panorama stitching method according to the first aspect;

[0068] a control unit configured to control the pan-tilt camera to rotate to a first mapping position of the target position in a space where the pan-tilt camera is located.

[0069] In one possible implementation, the controlling the pan-tilt camera to rotate to the first mapping position of the target position in the space where the pan-tilt camera is located comprises:

[0070] determining a second mapping position of the target position in each fifth image in the image set;

[0071] determining a distance between the second mapping position of each fifth image and the image center of the fifth image, to obtain a distance corresponding to each fifth image;

[0072] determining target rotation data of the pan-tilt camera based on a field of view angle of the target center and the pan-tilt rotation data corresponding to the target image, wherein the target rotation data represents a rotation angle of the pan-tilt camera when the optical axis of the pan-tilt camera moves to the first mapping position;

[0073] controlling the pan-tilt camera to rotate according to the target rotation data, so that the pan-tilt camera rotates to the target position in the first mapping position in the space where the pan-tilt camera is located.

[0074] In a fifth aspect, an electronic device is provided, and the electronic device comprises:

[0075] a memory configured to store a computer program;

[0076] a processor configured to execute the computer program stored in the memory, and when the computer program is executed, implement the method of any of the embodiments of the first aspect or the second aspect.

[0077] In a sixth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and when the computer program is executed by a processor, implement the method of any of the embodiments of the first aspect or the second aspect.

[0078] In a seventh aspect, a computer program product is provided, and the computer program product comprises computer readable code, and when the computer readable code is executed on a device, causes a processor in the device to implement the method of any of the embodiments of the first aspect or the second aspect.

[0079] The panoramic image splicing method provided in the embodiments of the present application can obtain an image set captured by a pan-tilt camera, a calibration result of the pan-tilt camera, and pan-tilt rotation data corresponding to images in the image set, wherein a field of view range of the image set contains a field of view range of the pan-tilt camera; the image set contains images captured by the pan-tilt camera when the tilt angle is greater than 0; the pan-tilt rotation data represents rotation data of the pan-tilt camera when the corresponding images are captured; then, a mapping relationship between the images in the image set and a panoramic image to be generated is determined based on the pan-tilt rotation data and the calibration result; and finally, the images in the image set are spliced based on the mapping relationship to obtain the panoramic image. In this way, by obtaining the images captured when the tilt angle is greater than 0, the image capturing capability of the pan-tilt camera in the vertical direction can be fully utilized; the mapping relationship between the images and the panoramic image is determined based on the pan-tilt rotation data and the calibration result, so as to determine the accurate position of the images in the panoramic image and obtain the panoramic image containing the entire field of view range of the pan-tilt camera; and the horizontal rotation is not required as the control target of the pan-tilt camera, thereby reducing the rotation limitation of the pan-tilt camera in the process of splicing the images for capturing the panoramic image. BRIEF DESCRIPTION OF DRAWINGS

[0080] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the application.

[0081] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced here. Obviously, those skilled in the art can obtain other drawings according to these drawings without any creative effort.

[0082] One or more embodiments are exemplarily illustrated by images in the drawings corresponding thereto, and these exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, unless otherwise specified. The drawings in the drawings do not constitute a proportional limitation.

[0083] FIG. 1 is a flowchart of a panoramic image splicing method according to an embodiment of the present application;

[0084] FIG. 2 is a flowchart of another panoramic image splicing method according to an embodiment of the present application;

[0085] FIG. 3 is a flowchart of a control method of a pan-tilt camera according to an embodiment of the present application;

[0086] FIG. 4 is a flowchart of another control method of a pan-tilt camera according to an embodiment of the present application;

[0087] FIG. 5A is a structural schematic diagram of a panoramic image splicing device according to an embodiment of the present application;

[0088] FIG. 5B is a structural schematic diagram of a control device of a pan-tilt camera according to an embodiment of the present application;

[0089] FIG. 6 is a structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0090] Various exemplary embodiments of the present application will now be described in detail by referring to the drawings, obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. It should be noted that: unless otherwise specifically stated, the relative arrangement, numerical expression and values of the components and steps set forth in these embodiments do not limit the scope of the present application.

[0091] Those skilled in the art can understand that the terms "first", "second" and the like in the embodiments of the present application are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor represent the logical order between them.

[0092] It should also be understood that in the present embodiment, "a plurality of" can mean two or more, and "at least one" can mean one, two or more.

[0093] It should also be understood that for any component, data or structure mentioned in the embodiments of the present application, unless specifically limited or given the opposite implication by the context, it can be understood as one or more in general.

[0094] In addition, the term "and / or" in the present application is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. In addition, the character " / " in the present application generally represents that the front and rear associated objects are in an "or" relationship.

[0095] It should also be understood that the description of various embodiments of the present application emphasizes the differences between various embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, will not be repeated.

[0096] The following description of at least one exemplary embodiment is merely illustrative in nature and does not in any way limit the application or its application or use.

[0097] The techniques, methods and devices known to those skilled in the related art can not be discussed in detail, but in appropriate cases, the above-mentioned techniques, methods and devices should be regarded as part of the specification.

[0098] It should be noted that like reference numerals and letters refer to like items throughout the accompanying drawings, and once an item is defined in one drawing, further discussions of the same item in subsequent drawings is not required.

[0099] It should be noted that the embodiments and the features in the embodiments in the present application can be combined with each other without conflicts. In order to understand the embodiments of the present application, the embodiments of the present application will be described in detail below with reference to the drawings and in combination with the embodiments. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0100] In order to solve the technical problem of how to synthesize a panorama with a larger field of view range in the prior art, the present application provides a panorama splicing method and a control method and device of a pan-tilt camera and electronic equipment, which can synthesize a panorama with a larger field of view range.

[0101] FIG. 1 is a flowchart of a panorama splicing method provided by an embodiment of the present application. The method can be applied to one or more electronic devices such as a smart phone, a notebook computer, a desktop computer, a portable computer, and a server. In addition, the execution subject of the method can be hardware or software. When the execution subject is hardware, the execution subject can be one or more of the above-mentioned electronic devices. For example, a single electronic device can execute the method, or multiple electronic devices can cooperate with each other to execute the method. When the execution subject is software, the method can be implemented as multiple software or software modules, or as a single software or software module. No specific limitation is made herein. In some scenarios, the panorama splicing method can be implemented via an application (APP) in a mobile phone and / or a server, that is, the execution subject of the panorama splicing method can not include a pan-tilt camera.

[0102] As shown in FIG. 1, the method specifically includes the following steps.

[0103] In step 101, a set of images captured by a pan-tilt camera, a calibration result of the pan-tilt camera, and pan-tilt rotation data corresponding to images in the set of images are obtained; a field of view range of the set of images contains a field of view range of the pan-tilt camera; the set of images contains images captured by the pan-tilt camera when the pitch angle is greater than 0; and the pan-tilt rotation data represents rotation data of the pan-tilt camera when the corresponding image is captured.

[0104] In the embodiment, the PTZ camera can rotate in horizontal and vertical directions, so as to adjust the monitoring angle of view and reduce the monitoring dead angle. A user or the like can remotely operate the PTZ camera through a network or the like, so as to conveniently adjust the monitoring angle in real time. Here, the lens of the PTZ camera can be a wide-angle lens, a long-focus lens, a macro lens, or a normal lens without wide-angle, long-focus, macro or the like. In practice, the normal lens arranged on the PTZ camera can reduce the production cost of the panoramic image, and compared with the wide-angle lens, the normal lens will not reduce the field of view of the panoramic image.

[0105] When the PTZ camera rotates in horizontal and vertical directions to obtain images in the image set, the rotation data of the PTZ camera when each image is captured can be recorded, such as the rotation angle in horizontal and vertical directions, so as to obtain the PTZ rotation data corresponding to the image.

[0106] The PTZ rotation data can include at least one of the yaw angle, the pitch angle and the roll angle. The yaw angle represents the rotation angle of the PTZ camera in the horizontal plane. The pitch angle represents the rotation angle of the PTZ camera in the vertical plane. The roll angle represents the rotation angle of the PTZ camera around its own axis.

[0107] The image set can include a plurality of images captured by the PTZ camera. The field of view of the image set contains the entire field of view of the PTZ camera, so that a panoramic image with the entire field of view of the PTZ camera can be synthesized subsequently. Since the PTZ camera can rotate in the vertical direction, the image set contains images captured by the PTZ camera when the pitch angle of the PTZ camera is greater than 0 (further, greater than 10 degrees, 30 degrees, 60 degrees or the like).

[0108] It should be noted that the image set can also contain images captured by the PTZ camera when the pitch angle of the PTZ camera is equal to 0. That is, the image set can contain images captured by the PTZ camera when the pitch angle of the PTZ camera is equal to 0, and images captured by the PTZ camera when the pitch angle of the PTZ camera is greater than 0.

[0109] The PTZ of the PTZ camera can generally rotate horizontally and up and down (i.e. rotate in the vertical direction). For the PTZ camera with a horizontal rotation angle of a, a vertical rotation angle of b, a horizontal field of view angle of a single image of c, and a vertical field of view angle of a single image of d, the image set can be captured in the following manner, so that the field of view of the image set contains the entire field of view of the PTZ camera:

[0110] In the first step, after the gimbal camera captures an image at the initial position, the gimbal camera captures an image every time it rotates less than or equal to c in the horizontal direction until the horizontal rotation angle of the gimbal camera in the horizontal direction reaches a.

[0111] In the second step, after the gimbal camera rotates less than or equal to d in the vertical direction, the first step is performed until the vertical rotation angle of the gimbal camera in the vertical direction reaches b.

[0112] It should be noted that the above method of obtaining an image set is only exemplary, and other methods can be used to obtain an image set containing the entire field of view of the gimbal camera within the field of view range under the inspiration of the above method of obtaining an image set. For example, after the gimbal camera captures an image at the initial position, the gimbal camera can first be controlled to capture an image every time it rotates less than or equal to d in the vertical direction until the horizontal rotation angle of the gimbal camera in the horizontal direction reaches b (hereinafter referred to as the shooting step). Then, the gimbal camera is controlled to perform the above shooting step every time it rotates less than or equal to c in the horizontal direction until the horizontal rotation angle of the gimbal camera in the horizontal direction reaches a. Alternatively, the gimbal camera can be controlled to move alternately in the horizontal direction and the vertical direction to capture images.

[0113] The above calibration results can include at least one of the following: an intrinsic matrix, an extrinsic matrix, distortion coefficients, etc.

[0114] The intrinsic matrix, extrinsic matrix, and distortion coefficients of the gimbal camera can be determined through gimbal camera calibration. In practical applications, camera calibration tools and libraries can be used to simplify the calibration process. In addition, the accuracy of the calibration results is also affected by factors such as the quality of the calibration board, the shooting environment, and the precision of the camera. Therefore, when calibrating the camera, the calibration board and shooting conditions need to be carefully selected, and multiple calibrations and verifications need to be performed to obtain reliable calibration results.

[0115] As an example, the intrinsic matrix, extrinsic matrix, and distortion coefficients of the gimbal camera can be determined as follows:

[0116] First, prepare the calibration board: the calibration board is a planar object with known feature points, such as a checkerboard. The positions of the feature points on the calibration board in the world coordinate system are known.

[0117] Then, capture calibration board images: use the camera to capture multiple calibration board images at different angles and positions.

[0118] Subsequently, extract feature points: in the captured calibration board images, use image processing algorithms to extract the pixel coordinates of the feature points.

[0119] Then, establish the equation set: establish the equation set according to the camera imaging model and the feature point position of the calibration board. These equation sets link the intrinsic, extrinsic and distortion coefficients of the camera with the pixel coordinates of the feature points.

[0120] Next, solve the equation set: use a numerical optimization algorithm to solve the equation set to obtain the estimated values of the intrinsic, extrinsic and distortion coefficients of the gimbal camera.

[0121] After that, optimize the calibration results: optimize the calibration results obtained by solving, for example, using a nonlinear optimization algorithm to improve the calibration accuracy.

[0122] Finally, verify the calibration results: use known three-dimensional objects or scenes to verify the calibration results and check the measurement accuracy and accuracy of the camera.

[0123] Step 102, based on the gimbal rotation data and the calibration results, determine the mapping relationship between the images in the image set and the panoramic image to be generated by stitching.

[0124] In this embodiment, the above-mentioned mapping relationship can represent the positional relationship and pixel transformation relationship between the images in the image set and the panoramic image to be generated by stitching. Each image in the image set can correspond to a mapping relationship. The mapping relationship can be a transformation matrix, through which the local image can be transformed into a global image to stitch the images and obtain the panoramic image.

[0125] Here, a variety of ways can be used to implement the above-mentioned step 102.

[0126] As an example, the above-mentioned step 102 can be implemented by a formula, a corresponding relationship table, and a machine learning algorithm.

[0127] Step 103, based on the mapping relationship, stitch the images in the image set to obtain the panoramic image.

[0128] In this embodiment, after obtaining the above-mentioned mapping relationship, the projection from the image to the panoramic image can be calculated by projection transformation, thereby obtaining the panoramic image

[0129] For example, after obtaining the mapping relationship from the image to the panoramic image, the step of obtaining the panoramic image by projection transformation can include:

[0130] 1. Determine the projection type: common projection types include equirectangular projection, spherical projection, cube projection, etc. Here, the appropriate projection method can be selected according to specific needs and application scenarios.

[0131] 2. Calculate projection parameters: According to the selected projection type and the characteristics of the image, calculate the corresponding projection parameters, such as the center point, radius, angle, etc.

[0132] 3. Establish pixel mapping: For each input image, establish the mapping relationship between the image pixels and the panoramic image pixels according to the projection parameters.

[0133] 4. Interpolation processing: During the pixel mapping process, non-integer coordinate positions may occur. At this time, interpolation processing is needed to obtain accurate pixel values. Common interpolation methods include bilinear interpolation, bicubic interpolation, etc.

[0134] 5. Image fusion: When multiple images are mapped to a panoramic image, there may be overlapping areas. Image fusion processing is needed to eliminate stitching marks and discontinuities. Weighted averaging, median filtering, etc. can be used for fusion.

[0135] 6. Color correction: Due to different shooting conditions of different images, the colors may not be consistent. Color correction is needed to make the colors of the panoramic image look more natural and unified.

[0136] 7. Border processing: For the borders of the panoramic image, there may be incomplete or discontinuous situations. Proper border processing such as filling, cropping or expanding is needed.

[0137] Further, exposure compensation can also be performed on the panoramic image. For example, adjust the exposure of each image in the image set in blocks, so that the brightness of each part of the stitched panoramic image can match, achieving the effect of smooth transition of the brightness of the overlapping area.

[0138] In the process of finding the seams between images, the minimum cut algorithm can be used to find the optimal seam. Then, a multi-layer Gaussian pyramid fusion can be used to obtain the panoramic image.

[0139] Specifically, the minimum cut algorithm can be used to find the optimal seam first:

[0140] Step 1, calculate the energy function of the image. For example, color difference, gradient, etc. can be used as energy measurement.

[0141] Step 2, construct the graph structure of the image, treat each pixel as a node of the graph, and connect the edges between adjacent pixels.

[0142] Step 3, use the minimum cut algorithm to find the path with the minimum energy in the graph, which is the optimal seam.

[0143] Step 4, according to the optimal seam, the images are spliced.

[0144] Then, the following method can be used to obtain the panoramic image using multi-layer Gaussian pyramid fusion.

[0145] First, build a Gaussian pyramid for the stitched image and the target image respectively.

[0146] Second, extend the Gaussian pyramid of the target image to the same size as the stitched image.

[0147] Third, fuse the stitched image and the extended target image according to the weight of the mask image.

[0148] Fourth, up-sample and reconstruct the fused image using a multi-level Gaussian pyramid to obtain the final panoramic image.

[0149] In the above process, the min-cut algorithm is used to find the optimal seam to reduce the stitching marks and ghosting phenomenon; the multi-level Gaussian pyramid fusion is used to achieve smooth transition and fusion of the image, making the panoramic image more natural. The above method can improve the quality and effect of image stitching

[0150] The panoramic image stitching method provided by the embodiments of the present application can obtain an image set captured by a gimbal camera, a calibration result of the gimbal camera, and gimbal rotation data corresponding to images in the image set; wherein a field of view range of the image set contains a field of view range of the gimbal camera; the image set contains images captured by the gimbal camera when the pitch angle is greater than 0; the gimbal rotation data represents the rotation data of the gimbal camera when the corresponding image is captured, then, based on the gimbal rotation data and the calibration result, a mapping relationship between the images in the image set and a panoramic image to be stitched is determined, and then, based on the mapping relationship, the images in the image set are stitched to obtain the panoramic image. Thus, by obtaining images captured when the pitch angle is greater than 0, the image capturing capability of the gimbal camera in the vertical direction can be fully utilized, the mapping relationship between the images and the panoramic image is determined based on the gimbal rotation data and the calibration result, to determine the accurate position of the images in the panoramic image, and the panoramic image containing the entire field of view range of the gimbal camera is obtained, without taking the horizontal rotation as the control target of the gimbal camera, thereby reducing the rotation restriction of the gimbal camera in the process of capturing the panoramic image.

[0151] FIG. 2 is a flowchart of another panoramic image stitching method provided by the embodiments of the present application. As shown in FIG. 2, the method specifically includes:

[0152] In step 201, a set of images captured by a pan-tilt camera, a calibration result of the pan-tilt camera, and pan-tilt rotation data corresponding to the images in the set of images are obtained; a field of view range of the set of images contains a field of view range of the pan-tilt camera; the set of images contains images captured by the pan-tilt camera when the pan-tilt camera has a tilt angle greater than 0; and the pan-tilt rotation data represents rotation data of the pan-tilt camera when the corresponding image is captured.

[0153] In this embodiment, step 201 is basically the same as step 101 in the corresponding embodiment of FIG. 1, and thus is not described here again.

[0154] In step 202, based on the pan-tilt rotation data, image rotation data of the images in the set of images relative to a panorama to be generated by stitching is determined, wherein the image rotation data represents a position of the images in the set of images relative to the panorama to be generated by stitching.

[0155] In this embodiment, the image rotation data can specifically represent a rotation position relationship in a camera coordinate system between the images in the set of images and the panorama to be generated by stitching.

[0156] Here, the above step 202 can be implemented in various ways.

[0157] As an example, the above step 202 can be implemented by a formula, a corresponding relationship table, or a machine learning algorithm.

[0158] In some optional implementations of this embodiment, the image rotation data of the images in the set of images relative to the panorama to be generated by stitching can be determined based on the pan-tilt rotation data in the following way:

[0159] First, based on the pan-tilt rotation data and a size of the images in the set of images, an overlapping area between a first image and a second image in the set of images is determined.

[0160] The size of the image can include a length and a width of the image. The first image and the second image are any two images in the set of images that have an overlapping area.

[0161] The first image and the second image can be any two different images in the set of images. Here, the first image and the second image can have an overlapping area, or can not have an overlapping area, i.e., the set of pixels of the overlapping area is an empty set.

[0162] Specifically, the horizontal and vertical pan rotation angles corresponding to the first image and the horizontal and vertical pan rotation angles corresponding to the second image can be determined from the pan rotation data corresponding to the images, and then the difference between the horizontal and vertical pan rotation angles of the first image and the second image can be calculated.

[0163] Then, the size of the overlapping region (in pixel units) can be calculated using the following formula: |field of view angle of the image - difference between the pan rotation angles| ÷ field of view angle x size of the image.

[0164] Then, the top-left corner coordinates of the overlapping region can be calculated: 0 or the difference between the image size and the overlapping region size can be determined according to the relative position.

[0165] In the case where the relative position indicates that the top-left corner coordinates of the overlapping region coincide with the top-left corner coordinates of the image (e.g., the first image or the second image), the top-left corner coordinates of the overlapping region can be determined as 0. In the case where the relative position indicates that the top-left corner coordinates of the overlapping region do not coincide with the top-left corner coordinates of the image (e.g., the first image or the second image), the top-left corner coordinates of the overlapping region can be determined as the difference between the image size of the image and the size of the overlapping region.

[0166] After obtaining the top-left corner coordinates of the overlapping region and the size of the overlapping region, the positions of the overlapping region in the first image and the second image can be determined.

[0167] Secondly, based on the overlapping region, the detection region of the first image and the detection region of the second image are determined.

[0168] The detection region is used to detect whether there is a matched feature point pair in the first image and the second image.

[0169] The matched feature point pair can represent the same position or the same object in space, and the images in the first image and the second image, respectively.

[0170] Here, the overlapping region can be used as the detection region, or the overlapping region can be expanded by a predetermined size, and the expanded overlapping region can be used as the detection region.

[0171] In some application scenarios of the above optional implementation manners, the following method (including sub-step one and sub-step two) can be used to determine the detection region of the first image and the detection region of the second image based on the overlapping region:

[0172] Sub-step one: the region in the first image that contains the overlapping region and has a size larger than the size of the overlapping region is determined as the detection region of the first image.

[0173] Sub-step two, determining, in the second image, a region containing the overlapping region and having a size larger than that of the overlapping region, as a detection region of the second image.

[0174] It can be understood that, in the above application scenarios, whether there is a matched feature point pair in the two images can be detected in a larger region relative to the overlapping region, so that the accuracy of panoramic image stitching can be improved by improving the feature point matching accuracy.

[0175] Step three, determining whether there is a matched feature point pair in the detection region of the first image and the detection region of the second image.

[0176] Here, whether there is a matched feature point pair in the detection region of the first image and the detection region of the second image can be determined based on ORB (Oriented FAST and Rotated BRIEF, an algorithm for fast feature point extraction and description), nearest neighbor matching, and RANSAC (random sample consensus).

[0177] ORB feature is a feature descriptor used in image recognition and computer vision tasks, which has rotation invariance and scale invariance. It consists of two parts: key point detection and descriptor calculation. Key point detection uses FAST algorithm or its improved versions such as oFAST (oriented FAST) or AGAST (Adaptive and Generic corner detection based on the Accelerated Segment Test) to detect corner points or key points in the image. These key points usually have high contrast and local features. Descriptor calculation uses BRIEF (Binary Robust Independent Elementary Features) algorithm or its improved versions such as rBRIEF (Rotation-Aware BRIEF) to calculate binary descriptors around the key points. These descriptors are obtained by comparing and encoding the pixels around the key points, which have low dimension and high computational efficiency. The main advantages of ORB feature include fast calculation speed, certain robustness to noise and illumination changes, and good performance in real-time applications and resource-constrained environments.

[0178] The nearest neighbor matching is a feature matching method used to find the most similar corresponding points between two images or two sets of feature points. It determines the matching relationship by calculating the distance between feature descriptors. In nearest neighbor matching, for each feature point in image A, the distance to all feature points in image B is calculated, and the nearest feature point is selected as the matching point. This method is simple and intuitive.

[0179] RANSAC is an algorithm for estimating model parameters from data containing outliers. The basic idea is to randomly select some data points to assume a model, and then use other data points to verify the model. If there are enough data points supporting the model, it is considered valid. By repeatedly repeating this process, the optimal model parameters are finally found. RANSAC algorithm has strong robustness and can effectively estimate model parameters in the presence of a large number of outliers or noise. In image stitching, RANSAC algorithm can filter out correct matches from feature point pairs that may contain errors, resulting in more accurate stitching results.

[0180] In the fourth step, in the case where there is no matched feature point pair (for example, the detection area of at least one of the first image and the second image lacks features or only has repetitive textures), a third image and a fourth image are determined from the first image and the second image. The third image is an image whose image rotation data is to be determined, and the fourth image is an image whose image rotation data has been determined.

[0181] In the fifth step, relative rotation data is determined. The relative rotation data represents the rotation relationship between the gimbal rotation data corresponding to the fourth image and the gimbal rotation data corresponding to the third image.

[0182] In the sixth step, based on the image rotation data of the fourth image relative to the panorama image to be generated by stitching, and the relative rotation data, the image rotation data of the third image relative to the panorama image to be generated by stitching is determined.

[0183] Here, the image whose image rotation data is not determined in the first image and the second image can be taken as the third image, and the image whose image rotation data is determined in the first image and the second image can be taken as the fourth image.

[0184] In a case that the detection region of only one of the first image and the second image lacks features or only has repetitive textures, the image whose detection region lacks features or only has repetitive textures can be taken as the third image. In a case that the detection regions of the first image and the second image both lack features or only have repetitive textures, the image of the first image and the second image, for which the corresponding image rotation data has been determined through the other image, can be taken as the fourth image; the other image of the first image and the second image can be taken as the third image. For example, if the first image is image A and the second image is image B, both image A and image B lack features, and image A has determined the image rotation data corresponding to image A through image C (i.e., the image rotation data of image A relative to the panorama to be generated), then image A can be taken as the fourth image and image B can be taken as the third image.

[0185] The image rotation data of the first image in the image set relative to the panorama to be generated can be a unit matrix, i.e., the elements on the diagonal line from the upper left corner to the lower right corner (referred to as the main diagonal line) in the matrix are all 1. All other elements are 0. Thus, the image rotation data of the other images in the image set relative to the panorama can be determined based on the first matrix.

[0186] As an example, the gimbal rotation data corresponding to the first image can be denoted as (h1, v1), where h1 represents the rotation data of the gimbal camera in the horizontal direction when the first image is captured, and v1 represents the rotation data of the gimbal camera in the vertical direction when the first image is captured. The gimbal rotation data corresponding to the second image can be denoted as (h2, v2), where h2 represents the rotation data of the gimbal camera in the horizontal direction when the second image is captured, and v2 represents the rotation data of the gimbal camera in the vertical direction when the second image is captured.

[0187] Then, (0, cos(v1 ÷ 180 x p), -sin(v1 ÷ 180 x p)) can be determined as the first rotation axis, and h2-h1 can be determined as the first rotation angle; (1, 0, 0) can be determined as the second rotation axis, and v2-v1 can be determined as the second rotation angle.

[0188] Subsequently, the relative rotation data of the first image and the second image is determined based on the first rotation axis, the first rotation angle, the second rotation axis, and the second rotation angle.

[0189] Further, the image rotation data of the third image relative to the panorama to be generated by stitching can be determined based on the image rotation data of the fourth image relative to the panorama to be generated by stitching and the relative rotation data described above. For example, in the case where the image rotation data and the relative rotation data are represented by matrices, the result of multiplying the two matrices can be determined as the image rotation data of the third image relative to the panorama to be generated by stitching.

[0190] It can be understood that, for a part image, if the detection region lacks features or only has repeated textures, resulting in the inability to calculate the image rotation data by using matched feature points, the relative rotation matrix of the image relative to another image can be calculated by using the change in the gimbal angle of the image, and then the rotation matrix of the other image (i.e., the image rotation data of the other image relative to the panorama to be generated by stitching) is multiplied by the relative rotation matrix to obtain the rotation matrix of the image, and the intrinsic matrix and the focal length of the image are set according to the calibration result of the gimbal camera, so as to obtain the mapping matrix of the image and take the mapping matrix as the image rotation data of the image relative to the panorama to be generated by stitching.

[0191] In addition, in the case where there are matched feature points, the relative rotation data between the gimbal rotation data corresponding to the first image and the gimbal rotation data corresponding to the second image can be determined based on the matched feature points first, and then the image rotation data of the first image and the image rotation data of the second image relative to the panorama to be generated by stitching are determined based on the relative rotation data between the gimbal rotation data corresponding to the first image and the gimbal rotation data corresponding to the second image.

[0192] Here, the matched feature points in the detection region of an image (e.g., the first image or the second image) can be projected into another image (e.g., the second image or the first image), so as to obtain the image rotation data of the image relative to the panorama to be generated by stitching.

[0193] Optionally, joint optimization can be performed by using bundle adjustment. Bundle adjustment is used to simultaneously optimize the pose of a camera and the coordinates of a three-dimensional point. The basic idea is to adjust the camera parameters and the positions of the three-dimensional points by minimizing the re-projection error, so that they can better fit the observation data.

[0194] It can be understood that, in the case where there are matched feature points, the image rotation data of an image relative to the panorama to be generated by stitching can be determined based on the matched feature points, so as to improve the accuracy of the stitched panorama.

[0195] In step 203, the image intrinsic data of the images in the image set is determined based on the calibration result.

[0196] In this embodiment, the distortion correction can be performed on each image in the image set based on the camera intrinsic parameters and distortion coefficients in the gimbal camera calibration result. The intrinsic parameters after distortion correction are determined as the image intrinsic parameter data of each image.

[0197] Specifically, the distortion correction can be performed on the image based on the camera intrinsic parameters and distortion coefficients in the gimbal camera calibration result in the following manner:

[0198] Firstly, the camera intrinsic parameters and distortion coefficients are read. For example, the focal length, principal point coordinates, radial distortion coefficients (such as k1, k2, k3) and tangential distortion coefficients (such as p1, p2) of the camera can be obtained from the calibration result.

[0199] Secondly, the mapping relationship between the pixel coordinates and ideal coordinates is established. For each pixel coordinate (u, v) in the image, the corresponding ideal non-distorted coordinate (u_ideal, v_ideal) is calculated.

[0200] Thirdly, the radial distortion and tangential distortion are considered. The radial distortion correction and tangential distortion correction are performed using a preset formula.

[0201] Fourthly, the corrected image is obtained by resampling. For example, the pixel value in the original image is obtained by interpolation algorithm (such as bilinear interpolation) according to the corrected coordinate (x_corrected, y_corrected), so as to obtain the image after distortion correction.

[0202] For example, assuming that the ideal coordinate (x_ideal, y_ideal) of a pixel coordinate (u, v) in an image is (100, 200), the radial distortion coefficient k1 = 0.1, the tangential distortion coefficients p1 = 0.01 and p2 = 0.02, and the image center coordinate is (cx, cy) = (500, 500), the distance r = sqrt((x_ideal-cx) 2 +(y_ideal-cy) 2 ) is calculated, the radial distortion correction is performed first, then the tangential distortion correction is performed, and finally the corrected coordinate is obtained, and the pixel value is obtained by interpolation.

[0203] Through the above steps, the image can be corrected by using the calibration parameters of the gimbal camera, so as to improve the quality and accuracy of the image.

[0204] Therefore, the intrinsic parameters after distortion correction can be determined as the image intrinsic parameter data of the corresponding image. The intrinsic parameters after distortion correction can be represented by a matrix, which can include the focal length and principal point coordinates of the gimbal camera, pixel size and other parameters.

[0205] In addition, in some cases, the image intrinsic data can be obtained directly without distortion correction.

[0206] Step 204, determining a mapping relationship between the images in the image set and the panoramic image to be generated based on the image rotation data and the image intrinsic data.

[0207] In this embodiment, the image rotation data and the image intrinsic data can be taken as the mapping relationship between the images in the image set and the panoramic image to be generated. Alternatively, the image rotation data, the image intrinsic data, and a homography matrix between the images and the panoramic image to be generated can be taken as the mapping relationship between the images in the image set and the panoramic image to be generated.

[0208] Step 205, stitching the images in the image set based on the mapping relationship to obtain the panoramic image.

[0209] In this embodiment, step 205 is basically the same as step 103 in the corresponding embodiment of FIG. 1, which will not be described here.

[0210] It should be noted that, in addition to the above, the present embodiment can also include the corresponding technical features described in the embodiment corresponding to FIG. 1, thereby achieving the technical effects of the panoramic image stitching method shown in FIG. 1. For details, please refer to the description of FIG. 1. For brevity, no further description is given here.

[0211] The panoramic image stitching method provided by the present embodiment determines the mapping relationship between the images and the panoramic image to be generated by using the image rotation data of the images relative to the panoramic image to be generated and the image intrinsic data of the images. Thus, the accuracy of panoramic image stitching can be further improved.

[0212] FIG. 3 is a flowchart of a control method of a gimbal camera according to an embodiment of the present application. The method can be applied to one or more electronic devices such as a gimbal camera, a smart phone, a notebook computer, a desktop computer, a portable computer, and a server. In addition, the execution subject of the method can be hardware or software. When the execution subject is hardware, the execution subject can be one or more of the above-mentioned electronic devices. For example, a single electronic device can execute the method, or multiple electronic devices can cooperate with each other to execute the method. When the execution subject is software, the method can be implemented as multiple software or software modules, or as a single software or software module. No specific limitation is imposed herein.

[0213] As shown in FIG. 3, the method specifically includes the following steps.

[0214] Step 301, obtaining a target position selected from a panoramic image.

[0215] In the present embodiment, the panoramic image can be obtained by using any of the panoramic image stitching methods described above, and can also be obtained by using other panoramic image stitching methods.

[0216] As an example, the other panoramic image stitching methods described above can include:

[0217] Feature point detection and matching algorithms: such as SIFT (Scale-Invariant Feature Transform) algorithm, SURF (Speeded Up Robust Features) algorithm and ORB (Oriented FAST and Rotated BRIEF) algorithm, etc. These algorithms are used to detect feature points with uniqueness and invariance in different images and match them. For example, using SIFT algorithm to detect key points in the image, and then matching by descriptor to find the corresponding relationship between different images.

[0218] Image registration algorithm: feature-based registration, which uses the feature points detected in the previous step to estimate the geometric transformation of the image. Region-based registration, which determines the registration parameters by comparing the pixel regions of the images.

[0219] Image fusion algorithm: for example, direct average fusion can simply average the pixel values of the overlapping region. For another example, multi-band fusion can decompose the image into different frequency bands and then fuse them separately. For another example, cross-fade fusion: smooth transition in the overlapping region.

[0220] Global optimization algorithm: used to optimize the cumulative error in the stitching process to make the panoramic image more accurate and natural. For example, using Bundle Adjustment to optimize the parameters of the camera and the positions of the feature points.

[0221] The target position can be a position selected by a user or the like from a screen displaying the panoramic image.

[0222] As an example, after displaying the panoramic image in the screen, the user can view the panoramic image, and can appreciate the panoramic image at different viewing angles and zoom levels through various devices (such as computers, mobile phones, tablets). For example, changing the viewing angle by sliding or rotating the device with the fingers on the mobile phone. During the user's viewing of the panoramic image, the target position can be selected from the screen displaying the panoramic image by clicking the screen or the like, so as to obtain the target position.

[0223] Step 302: control the pan-tilt camera to rotate to a first mapping position of the target position in the space where the pan-tilt camera is located.

[0224] In the embodiment, the first mapping position can be a mapping position of the target position in a space where the pan-tilt camera is located.

[0225] In some cases, the space where the pan-tilt camera is located can be a space represented by a panoramic image.

[0226] In some application scenarios of the embodiment, the pan-tilt camera can be controlled to rotate to the first mapping position of the target position in the space where the pan-tilt camera is located in the following manner:

[0227] First, determine the second mapping position of the target position in each fifth image in the image set.

[0228] The third image can be any image in the image set. For example, each image in the image set can be sequentially taken as the third image.

[0229] The second mapping position can represent the mapping position of the target position in the third image. Each third image corresponds to a second mapping position.

[0230] Second, for each fifth image, determine the distance between the second mapping position of the fifth image and the image center of the fifth image, to obtain the distance corresponding to each fifth image.

[0231] The distance corresponding to the third image can be the distance between the second mapping position of the third image and the image center of the third image.

[0232] Third, take the image center of the third image with the minimum corresponding distance in the image set as the target center, and determine the target rotation data of the pan-tilt camera based on the field of view angle of the target center and the pan-tilt rotation data corresponding to the target image.

[0233] The target rotation data represents the rotation angle of the pan-tilt camera when the optical axis of the pan-tilt camera moves to the first mapping position. Different target rotation data can also result in different field of view angles of the image center after the pan-tilt camera rotates. Therefore, since the second rotation angle is determined based on the field of view angle of the target center, the field of view angle of the image center after the pan-tilt camera rotates can be larger.

[0234] The optical axis of the pan-tilt camera can be the center line of the camera lens, which is the symmetry axis of the camera imaging system. The optical axis can be regarded as a virtual line passing through the center of the lens and perpendicular to the surface of the lens, which determines the direction of the image captured by the camera.

[0235] Specifically, after stitching the panorama successfully, the mapping relationship between each image and the panorama (which can be represented by a mapping matrix, for example) and the projection coordinate system origin coordinates can be saved. After determining the target position (u, v), the reverse projection of the target position (u, v) corresponding to each image is calculated one by one, and the projection point that does not exceed the value range (0, width, height) is selected as a reasonable mapping point. Then, the mapping point (u_map, v_map) closest to the center of the image screen is selected from the reasonable mapping points.

[0236] Then, according to the input image coordinates (u_map, v_map) and the camera's intrinsic parameters (included in the calibration result), the corresponding field of view angles relative to the center of the image screen (sub_angle_x, sub_angle_y) are calculated and output. sub_angle_x = atan((u_map-W / 2) / fx), sub_angle_y = atan((v_map-H / 2) / fy), where (W, H) represents the width and height of the image, and (fx, fy) represents the focal length of the gimbal camera in the calibration result.

[0237] Then, according to the gimbal angle ((abs_angle_x, abs_angle_y) when the image is taken and the relative angle (sub_angle_x, sub_angle_y), the angle (angle_x, angle_y) after rotating (sub_angle_x, sub_angle_y) based on (abs_angle_x, abs_angle_y) is calculated in the rectangular coordinate system, and it is taken as the positioning result. Wherein, the gimbal angle ((abs_angle_x, abs_angle_y) corresponds to the gimbal rotation data corresponding to the target image, and the relative angle (sub_angle_x, sub_angle_y) represents the first rotation data.

[0238] Subsequently, the relative angle (sub_angle_x, sub_angle_y) is converted from the camera coordinate system to the spherical coordinate system P, and the reference point O at the center of the camera in the spherical coordinate system is set.

[0239] Then, the coordinates of P and O after rotating abs_angle_y degrees around the x-axis are calculated respectively. This step can be written as a rotation matrix Ry with a dimension of 3X3 in the rectangular coordinate system, and then the product of RyP and RyO is obtained. Next, the angle relative to the reference point M = P-O is calculated.

[0240] Finally, the final gimbal angle M+(abs_angle_x, abs_angle_y) is calculated, which is the target rotation data.

[0241] In the fifth step, the pan-tilt camera is controlled to rotate according to the target rotation data, so that the pan-tilt camera rotates to a first mapping position of the target position in a space where the pan-tilt camera is located.

[0242] It can be understood that in the above application scenarios, the pan-tilt camera can be rotated to a position with the first mapping position as the center of the field of view, so as to monitor the first mapping position more targeted and comprehensively.

[0243] The control method of the pan-tilt camera provided by the embodiments of the present application can control the pan-tilt camera to rotate to a space position corresponding to the target position, so that the monitoring of the above space position can be realized.

[0244] The embodiments of the present application will be described below, but it should be noted that the embodiments of the present application can have the features described below, but the following description does not constitute a limitation on the protection scope of the embodiments of the present application.

[0245] Before introducing the present scheme, the concepts involved in the scheme are first explained as follows:

[0246] Panorama: refers to an image with a horizontal viewing angle of 360°.

[0247] Panorama stitching: stitching two or more images with overlapping parts into a seamless panorama.

[0248] Panorama positioning: the user clicks on any position on the image, calculates the corresponding pan-tilt angle, and rotates the pan-tilt around the position.

[0249] In the related art, when performing panorama stitching, a plurality of images are shot by rotating horizontally by 360° under the condition that the pitch angle is 0 (i.e., the lens is facing forward), and a panorama is obtained by combining the images.

[0250] The problem with this is that the vertical field of view angle of the panorama is limited by the vertical field of view angle of the lens. If a panorama with a larger vertical field of view angle is desired, the vertical field of view angle of the lens must be increased, for example, a fisheye lens with a field of view angle close to 180° is used on a panoramic camera.

[0251] For the panorama positioning function, it is necessary to synthesize a panorama with a larger vertical field of view angle, because the pan-tilt has a certain pitch angle rotation capability. If only images with a pitch angle of 0 are used to synthesize a panorama, the view after the pan-tilt is rotated upward / downward cannot be completely previewed, and even if the pitch angle of the pan-tilt exceeds the vertical field of view angle of the lens, the user cannot click on the corresponding pan-tilt position.

[0252] In addition, the panoramic map positioning technology in the related art is based on a simple estimation of a proportion of panoramic map coordinates in the panoramic map, and a specific calculation method is as follows:

[0253] Suppose that a user clicks a coordinate u, v, a panoramic map width and height w, h, and a horizontal field of view angle and a vertical field of view angle of the panoramic map are θx and θy respectively. Then, target gimbal angles are calculated as follows:

[0254] Such a method is relatively simple to calculate, but has low precision and cannot fully meet the demand of users for high-precision positioning.

[0255] Therefore, the present scheme can adapt to panoramic map positioning control algorithms of gimbal cameras of different specifications. The present scheme mainly includes two parts, a first part being panoramic map stitching and a second part being panoramic map positioning.

[0256] Please refer to FIG. 4, which is a flowchart of a control method of a gimbal camera according to an embodiment of the present application. In the panoramic map stitching, the input is the calibration result corresponding to the gimbal camera, a plurality of images (i.e. the image set described above) captured by the gimbal camera and the corresponding gimbal angles (i.e. the gimbal rotation data described above), and the output is the panoramic map and the parameters corresponding to each image (i.e. the mapping relationship between the image and the panoramic map to be generated).

[0257] The following is a detailed scheme of panoramic map stitching:

[0258] According to the camera intrinsic parameters and the distortion coefficients in the calibration result, each input image is subjected to distortion correction.

[0259] According to the image intrinsic parameter data of the distortion-corrected image, the rotation data in the horizontal direction and the rotation data in the vertical direction of the distortion-corrected image are calculated.

[0260] For each pair of input images (i.e. the first image and the second image described above), the position of the overlapping region is calculated.

[0261] The difference between the gimbal angles in the horizontal and vertical directions of the two images is calculated = the angle of image 1 - the angle of image 2. The angle of image 1 corresponds to the gimbal rotation data corresponding to the first image, and the angle of image 2 corresponds to the gimbal rotation data corresponding to the second image.

[0262] The size (in pixel units) of the overlapping region is calculated = |field of view angle - difference between gimbal angles| * size of image / field of view angle.

[0263] The size of the overlapping region is appropriately enlarged: the size of the image overlapping region is multiplied by a small proportionality coefficient to obtain the detection region.

[0264] Calculate the top-left coordinate of the overlap region, which is determined by the relative position as 0 or image size - overlap size.

[0265] For each pair of input images, calculate the feature points and their matching relationship in the corresponding detection region, thus reducing the false matching caused by repeated patterns. Open-source feature point detection and matching algorithms are used, mainly based on ORB features, nearest neighbor matching, and RANSAC to filter out matching feature point pairs.

[0266] Find the images containing matching feature point pairs, calculate the mapping relationship (e.g., can include intrinsic matrix and rotation matrix) between each image and the generated panoramic image to be stitched based on the matching results. And use bundle adjustment for joint optimization.

[0267] For some images, if the detection region lacks features or only has repeated textures, resulting in the inability to calculate the mapping matrix, the relative rotation matrix can be calculated using the corresponding image's gimbal angle change relative to another image. Then multiply the relative rotation matrix by the rotation matrix of "another image" to get the rotation matrix of "corresponding image". And set the image intrinsic data of "corresponding image" according to the camera calibration results, thus obtaining the mapping matrix of the corresponding image.

[0268] The following is the method for calculating the rotation matrix, the input is the shooting gimbal coordinates of two images (h1, v1) (h2, v2), and the output is the rotation matrix.

[0269] If h1 is not equal to h2 and v1 is not equal to v2, decompose the rotation into horizontal rotation and pitch rotation respectively:

[0270] For horizontal rotation, the camera coordinate system: the rotation axis is (0, cos(v1 ÷ 180 × π), -sin(v1 ÷ 180 × π)), and the rotation angle is h2 - h1.

[0271] For pitch rotation, the camera coordinate system: the rotation axis is (1, 0, 0), and the rotation angle is v2 - v1.

[0272] Convert the axis-angle representation of rotation to the rotation matrix representation, for example, through the Rodrigues rotation formula.

[0273] Horizontal correction, correct the rotation matrix to make the stitched image lie on a straight line to achieve alignment.

[0274] Then perform projection transformation, calculate the projection from the image to the panoramic image, and adjust the exposure of each image to match the brightness after stitching.

[0275] Finally, use the minimum cut algorithm to find the optimal seam, and use the multi-layer Gaussian pyramid to fuse each image, thus obtaining the panoramic image.

[0276] In the panoramic image positioning, the input of the algorithm includes the mapping relationship between the images and the panoramic image to be generated, and the target position clicked by the user, and the output is the target coordinate of the rotation of the holder, i.e. the first mapping position described above. The calculated coordinate is sent to the holder camera to rotate the holder camera to the target coordinate.

[0277] The following is a detailed scheme of panoramic image positioning:

[0278] After the successful stitching of the panoramic image, the mapping relationship between each image and the panoramic image to be generated (for example, a mapping matrix can be used to represent) and the projection coordinate system origin coordinate can be saved. After the target position (u, v) is determined, the reverse projection of the target position (u, v) corresponding to each image is calculated one by one, and the projection point that does not exceed the value range (0, width, height) is selected as a reasonable mapping point. Then, the mapping point (u_map, v_map) closest to the center of the image screen is selected from the reasonable mapping points.

[0279] Then, according to the input image coordinates (u_map, v_map) and the camera's intrinsic parameters (included in the calibration result), the corresponding field of view angle relative to the center of the image screen (sub_angle_x, sub_angle_y) is calculated and output. sub_angle_x = atan((u_map-W / 2) / fx), sub_angle_y = atan((v_map-H / 2) / fy), where (W, H) represents the width and height of the image, and (fx, fy) represents the focal length of the holder camera in the calibration result.

[0280] Then, according to the holder angle ((abs_angle_x, abs_angle_y) when the image is taken and the relative angle (sub_angle_x, sub_angle_y), the angle (angle_x, angle_y) after rotating (sub_angle_x, sub_angle_y) based on (abs_angle_x, abs_angle_y) is calculated in the rectangular coordinate system, and it is used as the positioning result. Wherein, the holder angle ((abs_angle_x, abs_angle_y) corresponds to the holder rotation data corresponding to the target image, and the relative angle (sub_angle_x, sub_angle_y) represents the first rotation data described above.

[0281] Subsequently, the relative angle (sub_angle_x, sub_angle_y) is converted from the camera coordinate system to the spherical coordinate system P, and the reference point O of the camera center in the spherical coordinate system is set.

[0282] Then, the coordinates of P and O after rotating abs_angle_y degrees around the x-axis are calculated respectively. This step can be written as a rotation matrix Ry with a dimension of 3X3 in the rectangular coordinate system, and then the product of RyP and RyO is calculated to obtain the coordinates of P and O after rotating abs_angle_y degrees around the x-axis. Next, the angle M = P-O relative to the reference point is calculated.

[0283] Finally, the final target rotation data M+(abs_angle_x,abs_angle_y) is calculated.

[0284] It should be noted that in addition to the above, the present embodiment can also include the technical features described in the above embodiments, and further achieve the technical effects of the panoramic image stitching method shown above. Please refer to the above description for details. For brevity, no further description is given here.

[0285] The panoramic image splicing method provided by the embodiments of the present application can synthesize an image with a larger vertical field of view by separately shooting and splicing at multiple pitch angles, only using a normal lens with a regular field of view, and optimizing the splicing success rate under more complex inputs. The shooting angle of each image does not need to be limited; the matching success rate is higher by calculating the overlapping area, feature points, and excluding unmatched feature areas; when there is no matching feature point pair, the rotation matrix can be calculated by the angle change of the holder. Thus, the more accurate positioning points are gradually calculated by using the mapping matrix (containing data such as perspective deformation and position relationship), the optical center deviation in the intrinsic parameter matrix, and the error caused by the coordinate system transformation, and the positioning accuracy is higher. The images shot by multiple layers are spliced, so that the splicing result can cover the entire field of view corresponding to the rotation range of the holder. When each image is shot, the holder angle at the corresponding shooting time is recorded, which is used for optimizing splicing and subsequent positioning. When the mapping matrix is calculated, the overlapping area of each pair of images is calculated according to the holder angle, and the feature points and matching of the pair of images are calculated only in the corresponding overlapping area, so as to reduce the mismatch caused by repeated patterns. For some images, if the overlapping area lacks features or only has repeated textures, the mapping matrix cannot be calculated, the relative rotation matrix is calculated by using the holder angle change of the corresponding image relative to another image, then the relative rotation matrix is multiplied by the rotation matrix of the "another image" to obtain the rotation matrix of the "corresponding image", and the intrinsic parameter matrix of the "corresponding image" is set according to the camera calibration result, so as to obtain the mapping matrix of the "corresponding image". After splicing is successful, the mapping relationship (composed of the intrinsic parameter matrix and the rotation matrix) and the projection coordinate system origin coordinates between each image and the panoramic image to be generated are saved, and after the user clicks the coordinates, the reverse projection of the point corresponding to each image is calculated, the reasonable mapping points are selected from the points not exceeding the value range, then the mapping point closest to the image picture center is selected from the reasonable mapping points, the field of view angle relative to the image picture center is calculated, and finally the angle after rotation based on the image shooting angle is converted to the angle in the rectangular coordinate system as the positioning result.

[0286] FIG. 5A is a structural schematic diagram of a panoramic image splicing device provided by an embodiment of the present application. Specifically, it comprises:

[0287] The first acquisition unit 401 is configured to acquire an image set shot by a holder camera, a calibration result of the holder camera, and holder rotation data corresponding to images in the image set; wherein the field of view range of the image set contains the field of view range of the holder camera; the image set contains images shot by the holder camera when the pitch angle is greater than 0; the holder rotation data represents the rotation data of the holder camera when the corresponding image is shot;

[0288] The first determining unit 402 is configured to determine a mapping relationship between images in the image set and a panorama to be generated by image stitching based on the gimbal rotation data and the calibration result.

[0289] The stitching unit 403 is configured to stitch the images in the image set based on the mapping relationship to obtain the panorama.

[0290] In a possible implementation, the determination of the mapping relationship between the images in the image set and the panorama to be generated by image stitching based on the gimbal rotation data and the calibration result comprises:

[0291] determining image rotation data of the images in the image set relative to the panorama to be generated by image stitching based on the gimbal rotation data, wherein the image rotation data represents positions of the images in the image set relative to the panorama to be generated by image stitching.

[0292] determining image intrinsic data of the images in the image set based on the calibration result.

[0293] determining the mapping relationship between the images in the image set and the panorama to be generated by image stitching based on the image rotation data and the image intrinsic data.

[0294] In a possible implementation, the determination of the image rotation data of the images in the image set relative to the panorama to be generated by image stitching based on the gimbal rotation data comprises:

[0295] determining an overlapping area between a first image and a second image in the image set based on the gimbal rotation data and sizes of the images in the image set; the first image and the second image are any two images in the image set having an overlapping area.

[0296] determining respective detection areas of the first image and the second image based on the overlapping area, wherein the detection areas are used to detect whether there are matching feature point pairs in the first image and the second image.

[0297] determining whether there are matching feature point pairs in the detection area of the first image and the detection area of the second image.

[0298] In a case where there are no matching feature point pairs, determining a third image and a fourth image from the first image and the second image, wherein the third image is an image for which image rotation data is to be determined, and the fourth image is an image for which image rotation data has been determined.

[0299] determining relative rotation data, wherein the relative rotation data represents a rotation relationship between the gimbal rotation data corresponding to the fourth image and the gimbal rotation data corresponding to the third image;

[0300] determining, based on the image rotation data of the fourth image relative to the panorama to be generated by stitching and the relative rotation data, the image rotation data of the third image relative to the panorama to be generated by stitching, wherein the relative rotation data represents a rotation relationship between the gimbal rotation data corresponding to the fourth image and the gimbal rotation data corresponding to the third image.

[0301] In one possible implementation, after the determination of whether there are matched feature point pairs in the detection region of the first image and the detection region of the second image, the apparatus further includes:

[0302] a second determination unit (not shown in the figure) configured to, in the case where there are matched feature point pairs, determine, based on the matched feature point pairs, relative rotation data between the gimbal rotation data corresponding to the first image and the gimbal rotation data corresponding to the second image;

[0303] a third determination unit (not shown in the figure) configured to determine, based on the relative rotation data between the gimbal rotation data corresponding to the first image and the gimbal rotation data corresponding to the second image, image rotation data of the first image and the second image relative to the panorama to be generated by stitching respectively.

[0304] In one possible implementation, the gimbal rotation data corresponding to the third image is denoted as (h1, v1), h1 represents rotation data of the gimbal camera in the horizontal direction when the third image is captured, and v1 represents rotation data of the gimbal camera in the vertical direction when the third image is captured, the gimbal rotation data corresponding to the fourth image is denoted as (h2, v2), h2 represents rotation data of the gimbal camera in the horizontal direction when the fourth image is captured, and v2 represents rotation data of the gimbal camera in the vertical direction when the fourth image is captured; and

[0305] The determination of the relative rotation data includes:

[0306] determining (0, cos(v1 ÷ 180 x π), -sin(v1 ÷ 180 x π)) as the first rotation axis and h2-h1 as the first rotation angle;

[0307] determining (1, 0, 0) as the second rotation axis and v2-v1 as the second rotation angle;

[0308] Determine relative rotation data between the third image corresponding pan-tilt rotation data and the fourth image corresponding pan-tilt rotation data based on the first rotation axis, the first rotation angle, the second rotation axis and the second rotation angle.

[0309] In a possible implementation, the determining the detection region of the first image and the detection region of the second image based on the overlapping region comprises:

[0310] Determining a region in the first image containing the overlapping region and having a size greater than the size of the overlapping region as the detection region of the first image.

[0311] Determining a region in the second image containing the overlapping region and having a size greater than the size of the overlapping region as the detection region of the second image.

[0312] The panoramic image stitching device provided by the embodiment can be the panoramic image stitching device shown in FIG. 5A, can perform all steps of the above-described panoramic image stitching method, and thus realizes the technical effects of the above-described panoramic image stitching method. For brevity, details are not described herein.

[0313] FIG. 5B is a structural schematic diagram of a pan-tilt camera device provided by an embodiment of the present application. Specifically, the pan-tilt camera device comprises:

[0314] The second acquisition unit 411 is configured to acquire a target position selected from a panoramic image, wherein the panoramic image is obtained by using any panoramic image stitching method according to the first aspect described above.

[0315] The control unit 412 is configured to control the pan-tilt camera to rotate to a first mapping position of the target position in a space where the pan-tilt camera is located.

[0316] In a possible implementation, the controlling the pan-tilt camera to rotate to the first mapping position of the target position in the space where the pan-tilt camera is located comprises:

[0317] Determining a second mapping position of the target position in each fifth image in the image set.

[0318] For each fifth image, determining a distance between the second mapping position of the fifth image and an image center of the fifth image to obtain a distance corresponding to each fifth image.

[0319] determining target rotation data of the pan-tilt camera based on a field of view angle of the target center and the pan-tilt rotation data corresponding to the target image, wherein the target rotation data represents a rotation angle of the pan-tilt camera in a case that an optical axis of the pan-tilt camera moves from a current position to the first mapping position;

[0320] controlling the pan-tilt camera to rotate according to the target rotation data, so as to rotate the pan-tilt camera to the target position in the first mapping position in the space where the pan-tilt camera is located.

[0321] The control device of the pan-tilt camera provided by the embodiment can be the control device of the pan-tilt camera as shown in FIG. 5B, can execute all steps of the control method of the pan-tilt camera, and further realizes the technical effects of the control method of the pan-tilt camera. For brevity, details are not described herein.

[0322] FIG. 6 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. The electronic device 500 shown in FIG. 6 includes at least one processor 501, a memory 502, at least one network interface 504, and other user interfaces 503. The various components in the electronic device 500 are coupled together through a bus system 505. It can be understood that the bus system 505 is used to realize the connection and communication between the components. The bus system 505 includes not only a data bus, but also a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all kinds of buses are marked as the bus system 505 in FIG. 6.

[0323] The user interface 503 can include a display, a keyboard, or a clicking device (for example, a mouse, a trackball, a touchpad, or a touch screen, etc.).

[0324] It is to be understood that the memory 502 in embodiments of the present application can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. Among them, the nonvolatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM can be used, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 502 described herein is intended to include, without being limited to, these and any other suitable types of memory.

[0325] In some embodiments, the memory 502 stores the following elements, executable units or data structures, or a subset of them, or an extended set of them: an operating system 5021 and an application program 5022.

[0326] Among them, the operating system 5021 contains various system programs, such as framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program 5022 contains various application programs, such as Media Player, Browser, etc., for implementing various application services. The program for implementing the method of the embodiments of the present application can be contained in the application program 5022.

[0327] In the present embodiment, by calling the program or instruction stored in the memory 502, specifically, the program or instruction stored in the application program 5022, the processor 501 is used to execute the method steps provided by each method embodiment, for example, including:

[0328] obtain an image set captured by a pan-tilt camera, a calibration result of the pan-tilt camera, and pan-tilt rotation data corresponding to images in the image set; wherein a field of view range of the image set contains a field of view range of the pan-tilt camera; the image set contains images captured by the pan-tilt camera when the pan-tilt camera has a tilt angle greater than 0; the pan-tilt rotation data represents rotation data of the pan-tilt camera when the corresponding image is captured;

[0329] determine a mapping relationship between the images in the image set and a panorama to be generated based on the pan-tilt rotation data and the calibration result;

[0330] stitch the images in the image set based on the mapping relationship to obtain the panorama.

[0331] alternatively,

[0332] obtain a target position selected from a panorama, wherein the panorama is obtained by using any panorama stitching method according to the first aspect;

[0333] control the pan-tilt camera to rotate to a first mapping position of the target position in a space where the pan-tilt camera is located.

[0334] The method disclosed in the embodiments of the present application can be applied to the processor 501 or implemented by the processor 501. The processor 501 can be an integrated circuit chip with a signal processing capability. In the implementation process, each step of the above method can be completed by an integrated logic circuit or an instruction in the form of software in the processor 501. The processor 501 described above can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware coding processor for execution, or a combination of hardware and software units in the coding processor for execution. The software unit can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, and other mature storage media in the art. The storage medium is located in the storage 502, and the processor 501 reads the information in the storage 502 and combines the hardware to complete the steps of the above method.

[0335] It can be understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing units can be implemented within one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSP Devices, DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general purpose processors, controllers, micro-controllers, microprocessors, other electronic units designed to perform the functions described above, or a combination thereof.

[0336] For software implementation, the techniques described herein can be implemented with a processing unit that executes program code that performs the functions described above. The program code can be stored in a storage medium and executed by the processing unit. The storage medium can be implemented within the processing unit or external to the processing unit.

[0337] The electronic device provided by the embodiment can be an electronic device as shown in FIG. 6, can execute all steps of each panoramic image splicing method described above, and further achieve the technical effects of each panoramic image splicing method described above. For brevity, the relevant description is not repeated here.

[0338] The embodiment of the present application further provides a storage medium (computer readable storage medium). The storage medium stores one or more programs. The storage medium can include a volatile memory such as a random access memory, and the memory can also include a non-volatile memory such as a read-only memory, a flash memory, a hard disk or a solid state disk, and the memory can also include a combination of the above kinds of memories.

[0339] When the one or more programs stored in the storage medium can be executed by one or more processors to implement the panoramic image splicing method executed at the electronic device side described above.

[0340] The processor described above is used to execute the panoramic image splicing program stored in the memory to implement the following steps of the panoramic image splicing method executed at the electronic device side:

[0341] obtain an image set captured by a pan-tilt camera, a calibration result of the pan-tilt camera, and pan-tilt rotation data corresponding to images in the image set; wherein a field of view range of the image set contains a field of view range of the pan-tilt camera; the image set contains images captured by the pan-tilt camera when the pan-tilt camera has a tilt angle greater than 0; the pan-tilt rotation data represents rotation data of the pan-tilt camera when corresponding images are captured;

[0342] based on the pan-tilt rotation data and the calibration result, determine a mapping relationship between images in the image set and a panoramic image to be generated by stitching;

[0343] based on the mapping relationship, stitch the images in the image set to obtain the panoramic image.

[0344] alternatively,

[0345] obtain a target position selected from a panoramic image, wherein the panoramic image is obtained by using any panoramic image stitching method according to the first aspect described above;

[0346] control the pan-tilt camera to rotate to a first mapping position of the target position in a space where the pan-tilt camera is located.

[0347] The skilled person should further appreciate that units and algorithm steps of various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, various components and steps have been described above in general terms, with the understanding that such components and steps can be implemented in either hardware or software, or a combination of both. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0348] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented by hardware, software executed by a processor, or a combination of both. The software modules can be stored in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0349] It is to be understood that the terminology used herein is for the purpose of describing particular example embodiments only and is not intended to be limiting. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises", "comprising", "includes", "including" and "has" are inclusive and therefore specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order

[0350] The above description is that of current embodiments of the application. Various modifications and changes can be made thereto without departing from the spirit and scope of the application as set forth. The scope of the application is not to be limited to the exact details shown above.

Claims

1. A panoramic image stitching method, wherein, The method includes: Acquire the image set captured by the PTZ camera, the calibration result of the PTZ camera, and the PTZ rotation data corresponding to the images in the image set; The field of view of the image set includes the field of view of the pan-tilt camera; the image set includes images captured by the pan-tilt camera when the pitch angle is greater than 0; the pan-tilt rotation data represents the rotation data of the pan-tilt camera when capturing the corresponding image. Based on the gimbal rotation data and the calibration results, the mapping relationship between the images in the image set and the panoramic image to be stitched together is determined; Based on the mapping relationship, the images in the image set are stitched together to obtain the panoramic image.

2. The method according to claim 1, wherein, The step of determining the mapping relationship between the images in the image set and the panoramic image to be stitched together, based on the gimbal rotation data and the calibration results, includes: Based on the gimbal rotation data, the image rotation data of the images in the image set relative to the panoramic image to be stitched together is determined, wherein the image rotation data represents the position of the images in the image set relative to the panoramic image to be stitched together. Based on the calibration results, determine the image intrinsic parameter data of the images in the image set; Based on the image rotation data and the image intrinsic parameter data, the mapping relationship between the images in the image set and the panoramic image to be stitched together is determined.

3. The method according to claim 2, wherein, The step of determining the image rotation data of the images in the image set relative to the panoramic image to be stitched together, based on the gimbal rotation data, includes: Based on the gimbal rotation data and the dimensions of the images in the image set, the overlapping area between the first image and the second image in the image set is determined; the first image and the second image are any two images in the image set that have an overlapping area. Based on the overlapping region, the detection regions of the first image and the second image are determined respectively; Based on the detection regions of the first image and the second image, the image rotation data of the images in the image set relative to the panoramic image to be stitched together is determined.

4. The method according to claim 3, wherein, The step of determining the image rotation data of the images in the image set relative to the panoramic image to be stitched together, based on the detection regions of the first image and the second image respectively, includes: Determine whether there are matching feature point pairs in the detection region of the first image and the detection region of the second image; In the absence of a matching feature point pair, a third image and a fourth image are determined from the first image and the second image, wherein the third image is the image whose rotation data is to be determined, and the fourth image is the image whose rotation data has been determined; Based on the third and fourth images, the image rotation data of the images in the image set relative to the panoramic image to be stitched together is determined.

5. The method according to claim 4, wherein, The step of determining the image rotation data of the images in the image set relative to the panoramic image to be stitched together, based on the third and fourth images, includes: Determine relative rotation data, wherein the relative rotation data represents the rotation relationship between the gimbal rotation data corresponding to the fourth image and the gimbal rotation data corresponding to the third image; Based on the image rotation data of the fourth image relative to the panoramic image to be stitched together, and the relative rotation data, the image rotation data of the third image relative to the panoramic image to be stitched together are determined.

6. The method according to claim 4, wherein, After determining whether a matching pair of feature points exists in the detection region of the first image and the detection region of the second image, the method further includes: In the case where a matching pair of feature points exists, the relative rotation data between the gimbal rotation data corresponding to the first image and the gimbal rotation data corresponding to the second image is determined based on the matching pair of feature points. Based on the relative rotation data between the gimbal rotation data corresponding to the first image and the gimbal rotation data corresponding to the second image, the image rotation data of the first image and the second image relative to the panoramic image to be stitched together are determined.

7. The method according to claim 5, wherein, The gimbal rotation data corresponding to the third image is denoted as (h1, v1), where h1 represents the horizontal rotation data of the gimbal camera when the third image is captured, and v1 represents the vertical rotation data of the gimbal camera when the third image is captured. The gimbal rotation data corresponding to the fourth image is denoted as (h2, v2), where h2 represents the horizontal rotation data of the gimbal camera when the fourth image is captured, and v2 represents the vertical rotation data of the gimbal camera when the fourth image is captured. as well as The determination of relative rotation data includes: Let (0, cos(v1÷180×π), -sin(v1÷180×π)) be the first axis of rotation, and let h2-h1 be the first angle of rotation; Define (1, 0, 0) as the second rotation axis and v2-v1 as the second rotation angle; Based on the first rotation axis, the first rotation angle, the second rotation axis, and the second rotation angle, the relative rotation data between the gimbal rotation data corresponding to the third image and the gimbal rotation data corresponding to the fourth image is determined.

8. The method according to claim 4, wherein, Determining whether there are matching feature point pairs in the detection region of the first image and the detection region of the second image includes: A fast feature point extraction and description algorithm is used to extract features from the detection region of the first image to obtain the first feature; The fast feature point extraction and description algorithm is used to extract features from the detection region of the second image to obtain the second feature; Based on the first feature and the second feature, determine whether there are matching feature point pairs in the detection region of the first image and the detection region of the second image.

9. The method according to claim 4, wherein, Determining whether there are matching feature point pairs in the detection region of the first image and the detection region of the second image includes: The nearest neighbor matching algorithm is used to determine whether there are matching feature point pairs in the detection regions of the first image and the second image.

10. The method according to claim 4, wherein, Determining whether there are matching feature point pairs in the detection region of the first image and the detection region of the second image includes: A random sampling consensus algorithm is used to determine whether there are matching feature point pairs in the detection regions of the first image and the second image.

11. The method according to claim 3, wherein, Determining the detection regions of the first image and the second image based on the overlapping region includes: The region in the first image that contains the overlapping region and whose size is larger than the size of the overlapping region is determined as the detection region of the first image; The region in the second image that contains the overlapping region and whose size is larger than the size of the overlapping region is determined as the detection region of the second image.

12. The method according to claim 3, wherein, Determining the detection regions of the first image and the second image based on the overlapping region includes: The size of the overlapping region is increased to obtain an enlarged overlapping region; Based on the overlapping area after size expansion, the detection areas of the first image and the second image are determined respectively.

13. The method according to any one of claims 1-12, wherein, The image set was obtained in the following manner: The following image acquisition steps are performed: After controlling the pan-tilt camera to capture images at the initial position, the pan-tilt camera is controlled to rotate in the horizontal direction by an angle less than angle c and then capture images until the horizontal rotation angle of the pan-tilt camera in the horizontal direction reaches angle a. The image acquisition step is executed after the pan-tilt camera rotates by an angle d in the vertical direction until the vertical rotation angle of the pan-tilt camera reaches angle b. Wherein, angle a is the maximum horizontal rotation angle of the pan-tilt camera, angle b is the maximum vertical rotation angle of the pan-tilt camera, angle c is the horizontal field of view angle of the pan-tilt camera corresponding to a single image, and angle d is the vertical field of view angle of the pan-tilt camera corresponding to a single image.

14. The method according to any one of claims 1-12, wherein, The calibration results include the camera intrinsic parameters of the PTZ camera and the distortion coefficient of the PTZ camera.

15. A control method for a pan-tilt-zoom (PTZ) camera, wherein, The method includes: Obtain the target location selected from the panoramic image, wherein the panoramic image is obtained by stitching together the method described in any one of claims 1-14; Control the PTZ camera to rotate to the first mapped position of the target position in the space where the PTZ camera is located.

16. The method according to claim 15, wherein, Controlling the pan-tilt camera to rotate to the first mapped position of the target position in the space where the pan-tilt camera is located includes: Determine the second mapping position of the target location in each fifth image of the image set, wherein the fifth image is any image in the image set; For each fifth image, determine the distance between the second mapping position of the fifth image and the image center of the fifth image to obtain the distance corresponding to each fifth image; Based on the distance corresponding to the fifth image, the PTZ camera is controlled to rotate to the first mapped position of the target position in the space where the PTZ camera is located.

17. The method according to claim 16, wherein, The step of controlling the PTZ camera to rotate to the first mapped position of the target position in the space where the PTZ camera is located, based on the distance corresponding to the fifth image, includes: The image center of the third image with the smallest corresponding distance in the image set is taken as the target center. Based on the field of view angle of the target center and the gimbal rotation data corresponding to the target image, the target rotation data of the gimbal camera is determined. The target rotation data represents the rotation angle of the gimbal camera when the optical axis of the gimbal camera moves from the current position to the first mapping position. The PTZ camera is rotated according to the target rotation data so that it rotates to the first mapped position of the target position in the space where the PTZ camera is located.

18. A panoramic image stitching device, wherein, The device includes: The first acquisition unit is used to acquire an image set captured by a PTZ camera, the calibration result of the PTZ camera, and PTZ rotation data corresponding to the images in the image set; wherein, the field of view of the image set includes the field of view of the PTZ camera; the image set includes: images captured by the PTZ camera when the pitch angle is greater than 0; the PTZ rotation data represents the rotation data of the PTZ camera when capturing the corresponding image; The first determining unit is used to determine the mapping relationship between the images in the image set and the panoramic image to be stitched together, based on the gimbal rotation data and the calibration results. The stitching unit is used to stitch the images in the image set based on the mapping relationship to obtain the panoramic image.

19. A control device for a pan-tilt-zoom (PTZ) camera, wherein, The device includes: The second acquisition unit is used to acquire the target location selected from the panoramic image, wherein the panoramic image is obtained by stitching together the method described in any one of claims 1-14; The control unit is used to control the pan-tilt camera to rotate to the first mapped position of the target position in the space where the pan-tilt camera is located.

20. An electronic device, wherein, include: Memory, used to store computer programs; A processor for executing a computer program stored in the memory, wherein when the computer program is executed, it implements the method of any one of claims 1-17.

Citation Information

Patent Citations

  • System and method for obtaining spherical panorama image

    CN109362234A

  • Panoramic splicing and dynamic updating method for monitoring area of two-dimensional holder

    CN117671223A

  • Camera panoramic stitching method and system based on ONVIF protocol

    CN117676340A

  • Panoramic image capture method and device

    WO2021238317A1