Image processing method and device and related product

By projecting video frames onto a panoramic sphere and performing pixel coordinate system transformation and stitching, the problem of low quality in panoramic image stitching is solved, improving the quality of panoramic images and user experience.

CN121120371APending Publication Date: 2025-12-12BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410757214.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-12
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

The panoramic images generated by existing technologies have low quality at the stitching points of individual video frames, resulting in poor quality of the stitched panoramic image.

Method used

By acquiring multiple video frames from the video to be processed as keyframes, projecting them onto a panoramic sphere, determining the reference image and non-reference images, and then transforming the non-reference image to the pixel coordinate system of the reference image for stitching, a panoramic image is generated.

Benefits of technology

It improves the stitching quality of panoramic images and enhances the user's panoramic image browsing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120371A_ABST
    Figure CN121120371A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and device and a related product, and the method comprises the steps: obtaining a to-be-processed video, selecting a plurality of video frames from the to-be-processed video as key frames, projecting the key frames to a panoramic sphere, and obtaining a projection image corresponding to the key frames; determining a first image and a second image in each projection image; the first image is a reference image in the projection images; the second image is an image except the reference image in each projection image; and converting the second image into a pixel point coordinate system of the first image, and splicing the first image and the converted second image to obtain a panoramic image. According to the embodiment of the invention, a high-quality panoramic image can be generated, and the panoramic image browsing experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of image processing, and in particular, to an image processing method, device and related product. BACKGROUND

[0002] In the prior art, the information of things can be efficiently and comprehensively recorded through a panoramic image, thereby helping users better understand the related things. Currently, a panoramic image can be generated by splicing multiple video frames in a video data recorded by a mobile phone. However, the panoramic image generated by the prior art may have quality problems at the splicing positions of the video frames, resulting in low image quality of the panoramic image obtained by splicing. SUMMARY

[0003] Embodiments of the present disclosure provide an image processing method, device and related product, which can generate a high-quality panoramic image and improve the panoramic image browsing experience of users.

[0004] In a first aspect, the embodiments of the present disclosure provide an image processing method, comprising:

[0005] obtaining a to-be-processed video, selecting multiple video frames as key frames in the to-be-processed video, and projecting the key frames onto a panoramic sphere to obtain projection images corresponding to the key frames;

[0006] determining a first image and a second image in each of the projection images; the first image is a reference image in each of the projection images; and the second image is an image other than the reference image in each of the projection images;

[0007] transforming the second image into a pixel point coordinate system of the first image, splicing the first image and the transformed second image to obtain a panoramic image.

[0008] In a second aspect, the embodiments of the present disclosure provide an image processing device, comprising:

[0009] a projection module configured to obtain a to-be-processed video, select multiple video frames as key frames in the to-be-processed video, and project the key frames onto a panoramic sphere to obtain projection images corresponding to the key frames;

[0010] a determination module configured to determine a first image and a second image in each of the projection images; the first image is a reference image in each of the projection images; and the second image is an image other than the reference image in each of the projection images;

[0011] The stitching module is configured to transform the second image into a pixel point coordinate system of the first image, stitch the first image and the transformed second image, and obtain a panoramic image.

[0012] In a third aspect, an electronic device is provided, including: a processor; and a memory configured to store computer-executable instructions that, when executed, cause the processor to implement the steps of the method of the first aspect.

[0013] In a fourth aspect, a computer-readable storage medium is provided, configured to store computer-executable instructions that, when executed by a processor, implement the steps of the method of the first aspect.

[0014] In a fifth aspect, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method of the first aspect.

[0015] In the embodiments of the present disclosure, first, a to-be-processed video is acquired, a plurality of video frames are selected as key frames in the to-be-processed video, and the key frames are projected onto a panoramic sphere to obtain projection images corresponding to the key frames; a first image and a second image are determined in each projection image; the first image is a reference image in each projection image; the second image is an image other than the reference image in each projection image; the second image is transformed into a pixel point coordinate system of the first image, and the first image and the transformed second image are stitched to obtain a panoramic image. It can be seen that, through the embodiments, by projecting each key frame onto a sphere to obtain each projection image, the panoramic image is generated based on the projection image, which can make the finally generated panoramic image have a good panoramic playing effect and improve the panoramic image browsing experience of the user. In addition, by determining the first image and the second image in each projection image, taking the first image as the reference image, transforming each second image into the pixel point coordinate system of the first image, and stitching the first image and the transformed second image, each projection image can be stitched based on the pixel point coordinate system of the first image, thereby avoiding the image quality problem caused by the stitching of multiple images, improving the image quality of the stitched panoramic image, and improving the panoramic image browsing experience of the user. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to make one or more embodiments of the present disclosure or the technical solutions in the prior art clearer, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.

[0017] Figure 1 A flowchart of an image processing method provided by an embodiment of the present disclosure is shown in the figure;

[0018] Figure 2a A schematic diagram of a splicing area provided by an embodiment of the present disclosure is shown in the figure;

[0019] Figure 2b A schematic diagram of another splicing area provided by an embodiment of the present disclosure is shown in the figure;

[0020] Figure 2c A schematic diagram of still another splicing area provided by an embodiment of the present disclosure is shown in the figure;

[0021] Figure 3 A structural schematic diagram of an image processing device provided by an embodiment of the present disclosure is shown in the figure;

[0022] Figure 4 A structural schematic diagram of an electronic device provided by an embodiment of the present disclosure is shown in the figure. DETAILED DESCRIPTION

[0023] In order to make one or more embodiments of the present disclosure or the technical solutions in the prior art clearer, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.

[0024] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0025] For example, when responding to the active request of the user, the user is sent prompt information to explicitly prompt the user that the operation requested to be executed will need to obtain and use the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the electronic device, application program, server or storage medium, etc. software or hardware that executes the operation of the technical solution of the present disclosure according to the prompt information.

[0026] As an optional but non-limiting implementation, in response to receiving the active request of the user, the manner of sending the prompt information to the user may be, for example, a pop-up window manner in which the prompt information may be presented in a text manner. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.

[0027] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0028] Figure 1 A flowchart of an image processing method provided by an embodiment of the present disclosure is shown in FIG. 1, Figure 1 The method in FIG. 1 can be executed by a server, which can be a standalone server, a server cluster composed of multiple servers, or a cloud server for cloud computing. The method can also be executed by an electronic device, which can be a mobile device such as a mobile phone with a shooting function. As shown in FIG. 2, Figure 1 The method includes the following steps:

[0029] In step S102, a to-be-processed video is obtained, a plurality of video frames are selected as key frames in the to-be-processed video, and the key frames are projected onto a panoramic sphere to obtain projection images corresponding to the key frames;

[0030] In step S104, a first image and a second image are determined in each projection image; the first image is a reference image in each projection image; and the second image is an image other than the reference image in each projection image;

[0031] In step S106, the second image is transformed into a pixel point coordinate system of the first image, and the first image and the transformed second image are spliced to obtain a panoramic image.

[0032] In the embodiments of the present disclosure, first, a to-be-processed video is acquired, a plurality of video frames are selected as key frames in the to-be-processed video, and the key frames are projected onto a panorama sphere to obtain projection images corresponding to the key frames; a first image and a second image are determined in each projection image; the first image is a reference image in each projection image; the second image is an image other than the reference image in each projection image; the second image is transformed into a pixel point coordinate system of the first image, and the first image and the transformed second image are spliced to obtain a panorama image. It can be seen that, by projecting each key frame onto a sphere to obtain each projection image, and generating a panorama image based on the projection images, the panorama image generated finally can have a good panorama playing effect, and the user's panorama image browsing experience can be improved. In addition, by determining the first image and the second image in each projection image, taking the first image as the reference image, transforming each second image into the pixel point coordinate system of the first image, and splicing the first image and the transformed second image, each projection image can be spliced based on the pixel point coordinate system of the first image, so that the image quality problem caused by the splicing of multiple images can be avoided, the image quality of the panorama image obtained by splicing can be improved, and the user's panorama image browsing experience can be improved.

[0033] The panorama image involved in the embodiments, also referred to as a panorama photo or a 360-degree panorama photo, in one example, has an aspect ratio of 2:1, and can include scene information in a horizontal 360-degree and a vertical 180-degree range, so that the panorama image can show the overall appearance of a scene. When the scene information captured by a user is incomplete or the user does not capture according to a rotation angle of horizontal 360 degrees and vertical 180 degrees, the panorama image can have missing areas in some regions, and the missing areas in the panorama image can be replaced by a black background when the panorama image is displayed.

[0034] In the step S102, the to-be-processed video is acquired. The to-be-processed video can be a video captured by shooting a target scene, for example, an exhibition hall, a museum, a house, an automobile interior, etc. The to-be-processed video can be a video captured by a device with a wide-angle lens or an ultra-wide-angle lens.

[0035] In one example, the wide-angle lens is a lens with a focal length less than 35 mm, and the focal length range is generally between 14 mm and 35 mm. The view angle range of the ultra-wide-angle lens can reach 80° to 110°, and the focal length is between 15 mm and 20 mm.

[0036] In the step S102, a plurality of video frames are also selected as key frames in the to-be-processed video. In one embodiment, selecting a plurality of video frames as key frames in the to-be-processed video includes:

[0037] selecting a plurality of initial key frames in the to-be-processed video;

[0038] obtaining first camera pose information corresponding to the initial key frame and second camera pose information corresponding to a neighboring initial key frame of the initial key frame;

[0039] selecting a final key frame from the initial key frames according to the first camera pose information and the second camera pose information.

[0040] When selecting multiple video frames as key frames from the to-be-processed video frames, one video frame can be selected as an initial key frame every first number of video frames in the to-be-processed video frames, and multiple initial key frames can be obtained. For example, one video frame can be selected as an initial key frame every 30 video frames in the to-be-processed video frames, and multiple initial key frames can be obtained. Then, after obtaining the initial key frames, first camera pose information corresponding to each initial key frame and second camera pose information corresponding to a neighboring key frame of each initial key frame can be obtained by using a pose estimation technique, and a final key frame can be selected from the initial key frames according to the first camera pose information and the second camera pose information.

[0041] In the embodiments of the present disclosure, the camera pose information corresponding to each initial key frame obtained is a 3*3 rotation matrix from a world coordinate system to a camera coordinate system. The world coordinate system is composed of three mutually perpendicular and intersecting coordinate axes X, Y and Z.

[0042] When obtaining the camera pose information of each initial key frame, there are two prerequisites, one is that the camera intrinsic parameters are fixed, and the other is that the video shooting action is approximated as a pure rotation action. The camera intrinsic parameters include the focal length and the principal point position of the camera. It can be understood that when shooting a scene video, the handheld camera rotates one circle, that is, 360 degrees horizontally or other angles, and in this process, the camera has a small translation action, so the translation action can be ignored, and the video shooting action is approximated as a pure rotation action.

[0043] It can be seen that, by the embodiments, the final key frame can be selected from the initial key frames according to the first camera pose information corresponding to the initial key frame and the second camera pose information corresponding to the neighboring initial key frame of the initial key frame, so that the selection of the key frame can be more accurate.

[0044] In some embodiments, selecting a final key frame from the initial key frames according to the first camera pose information and the second camera pose information includes:

[0045] determining a first rotation angle in a first rotation direction when the device shooting the to-be-processed video shoots the initial key frame according to the first camera pose information, and determining a second rotation angle in the first rotation direction when the device shooting the to-be-processed video shoots the neighboring key frame according to the second camera pose information;

[0046] For each initial keyframe, if the angle difference between the first rotation angle and the second rotation angle is greater than the angle difference threshold, then the initial keyframe is determined as the final keyframe.

[0047] When selecting the final keyframe from each initial keyframe based on the pose information of the first and second cameras, firstly, the first rotation angle in the first rotation direction when the camera captures the initial keyframe can be determined based on the first camera pose information. Then, the second rotation angle in the first rotation direction when the camera captures adjacent keyframes of the initial keyframe can be determined based on the second camera pose information. Next, for each initial keyframe, the angle difference between the first and second rotation angles can be determined to be greater than an angle difference threshold. If the angle difference between the first and second rotation angles is greater than the angle difference threshold, then that initial keyframe is determined as the final keyframe.

[0048] The adjacent keyframe of the initial keyframe can be the preceding keyframe. This is understandable because when shooting scene video, the user uses a camera to shoot around the scene. The camera rotates at different angles in the first rotation direction for each frame, and each frame reflects scene information. However, the similarity of scene information reflected in two or more adjacent video frames can be extremely high. Therefore, when selecting keyframes, the angle difference between the initial keyframe and its preceding adjacent keyframe can be used to determine whether the initial keyframe is a keyframe. When the angle difference between the initial keyframe and its preceding adjacent keyframe is greater than an angle difference threshold, it indicates that the scene information reflected by the initial keyframe is less similar to that reflected by its preceding adjacent keyframe, and this initial keyframe can reflect more new scene information.

[0049] For example, the first rotation direction is the vertical direction, i.e., the Y-axis in the coordinate system. Based on the first camera pose information, the rotation angle of the camera relative to the Y-axis when capturing the initial keyframe is determined to be 25 degrees. Based on the second camera pose information, the rotation angle of the camera relative to the Y-axis when capturing the adjacent keyframe of the initial keyframe is determined to be 35 degrees. Assuming the angle difference threshold is 5 degrees, then the angle difference between the first rotation angle and the second rotation angle is 10 degrees, which is greater than the angle difference threshold of 5 degrees. Therefore, the initial keyframe can be determined as the final keyframe.

[0050] As can be seen, through this embodiment, the rotation angle of the device capturing the video to be processed when capturing the initial keyframe and adjacent keyframes can be determined according to the camera pose information. The keyframes are selected according to the rotation angle between the initial keyframe and the adjacent keyframes. By selecting the initial keyframe with a significantly changed rotation angle as the final keyframe, the selection of keyframes can be made more accurate.

[0051] After obtaining the keyframes, the intrinsic parameters and rotation matrix of the shooting device can be used to project each keyframe onto the panoramic sphere, obtaining the projected image corresponding to each keyframe. Then, a first image and a second image can be determined from the multiple projected images. The first image is the reference image among the various projected images, and the second image is the image among the various projected images excluding the reference image.

[0052] In step S104 above, a first image and a second image are determined in each projected image. The first image is a reference image in each projected image; the second image is any image in each projected image other than the reference image. In one embodiment, determining the first image in each projected image includes:

[0053] The image number of each projected image is determined according to the order of the frame numbers of each keyframe.

[0054] Based on the image sequence number of each projected image, the image whose sequence number is in the middle position among all projected images is determined as the first image.

[0055] After extracting frames from the captured video, a frame number can be assigned to each video frame. When determining the first image in each projected image, the image number of each projected image can be determined first based on the order of the frame numbers of the keyframes. It's understood that keyframes are selected video frames from multiple videos, and the frame numbers of the keyframes are not consecutive. Therefore, the image number of each projected image can be determined based on the order of the frame numbers of the keyframes, resulting in consecutive image numbers for each projected image. After determining the image numbers of each projected image, the image with the middle image number can be selected as the first image. Correspondingly, the second images are all the other projected images except for the one with the middle image number.

[0056] For example, the video to be processed includes 125 video frames, each numbered from 1 to 125 in chronological order. The selected keyframes are numbered 1, 32, 63, 94, and 125. Based on the order of the keyframe numbers, the image numbers of the corresponding projection images are determined to be 1, 2, 3, 4, and 5. Then, projection image 3 is used as the first image, and projection images 1, 2, 4, and 5 are used as the second images.

[0057] As can be seen, through this embodiment, the image with the middle position of the image sequence number in each projected image can be determined as the first image, and image stitching and transformation can be performed based on the first image with the middle position, thereby improving the efficiency and accuracy of image stitching.

[0058] In step S106 above, after determining that the first image and the second image are obtained, the second image is transformed into the pixel coordinate system of the first image. In one embodiment, transforming the second image into the pixel coordinate system of the first image includes:

[0059] Determine the pixel coordinate transformation matrix between the projected image and its adjacent projected images; the keyframe corresponding to the projected image is adjacent to the keyframe corresponding to the adjacent projected image.

[0060] Based on the pixel coordinate transformation matrix, the second image is transformed into the pixel coordinate system of the first image.

[0061] When transforming the second image into the pixel coordinate system of the first image, it is first necessary to determine the pixel coordinate transformation matrix between each projected image and its adjacent projected images. Here, the keyframes corresponding to the projected images are adjacent to the keyframes corresponding to their adjacent projected images. For example, if the frame numbers of the keyframes are 1, 32, 63, 94, 125, and the image numbers of the projected images corresponding to each keyframe are 1, 2, 3, 4, 5, then keyframe 1 and keyframe 32 are two adjacent keyframes, and therefore, projected image 1 and projected image 2 are adjacent.

[0062] The pixel coordinate transformation matrix is ​​a 3x3 matrix used to transform the coordinates of an image, transferring the original pixel coordinates of one image to the pixel coordinate system of another image.

[0063] Next, based on the determined pixel coordinate transformation matrix between each projected image and its adjacent projected images, the second image is transformed into the pixel coordinate system of the first image.

[0064] As can be seen, through this embodiment, by determining the pixel coordinate transformation matrix between the projected image and its adjacent projected images, and transforming the second image into the pixel coordinate system of the first image according to the pixel coordinate transformation matrix, the second image and the first image can be spatially aligned, achieving effective stitching of the second image and the first image, and improving the efficiency and quality of image stitching.

[0065] In one embodiment, determining the pixel coordinate transformation matrix between the projected image and its adjacent projected images includes:

[0066] Determine feature point matching pairs between the projected image and adjacent projected images;

[0067] Based on the feature point matching pairs, determine the pixel coordinate transformation matrix between the projected image and the adjacent projected images.

[0068] When determining the pixel coordinate transformation matrix between a projected image and its adjacent projected images, we can first determine the feature point matching pairs between the projected image and its adjacent projected images. Then, based on the feature point matching pairs, we can determine the pixel coordinate transformation matrix between the projected image and its adjacent projected images. A feature point matching pair between a projected image and its adjacent projected images consists of a feature point in the projected image and a feature point in the adjacent projected image. The feature points are obtained by feature point extraction from the projected image. Feature points in the projected image can reflect information such as brightness, color, and texture, and are points or regions with specific characteristics composed of one or more pixels.

[0069] As can be seen, through this embodiment, by determining the feature point matching pairs between the projected image and adjacent projected images, and based on the feature point matching pairs, determining the pixel coordinate transformation matrix between the projected image and adjacent projected images, the correspondence between the projected image and adjacent projected images can be accurately determined, making the pixel coordinate transformation matrix between the projected image and adjacent projected images more accurate, thereby improving the quality of the panoramic image obtained by stitching.

[0070] In one embodiment, determining feature point matching pairs between a projected image and adjacent projected images includes:

[0071] Determine initial feature point matching pairs between the projected image and adjacent projected images; an initial feature point matching pair means that a first feature point in the projected image matches a second feature point in an adjacent projected image; the second feature point is the feature point in the adjacent projected image with the highest feature similarity to the first feature point; the first feature point is the feature point in the projected image with the highest feature similarity to the second feature point.

[0072] The initial feature point matching pairs between the projected image and its neighboring projected images are filtered and sampled to obtain the final feature point matching pairs between the projected image and its neighboring projected images.

[0073] When determining feature point matching pairs between a projected image and its neighboring projected images, initial feature point matching pairs can be determined first. Each projected image has multiple initial feature point matching pairs with its corresponding neighboring projected images. Each initial feature point matching pair includes a first feature point in the projected image and a second feature point in the adjacent projected image. An initial feature point matching pair indicates that the first feature point in the projected image matches the second feature point in the adjacent projected image. Specifically, the second feature point is the feature point in the adjacent projected image with the highest feature similarity to the first feature point; the first feature point is the feature point in the projected image with the highest feature similarity to the second feature point. In other words, the first and second feature points are the most similar feature points to each other.

[0074] Considering that there may be erroneous feature point matching pairs in the initial feature point matching pairs, and that a large number of unevenly distributed feature point matching pairs will affect the calculation of the pixel coordinate transformation matrix between the projected images, after determining the initial feature point matching pairs between the projected image and its adjacent projected images, the initial feature point matching pairs between the projected image and its adjacent projected images can be filtered and sampled to obtain the final feature point matching pairs between the projected image and its adjacent projected images. The pixel coordinate transformation matrix between the projected image and its adjacent projected images is then calculated using the final feature point matching pairs.

[0075] As can be seen, this embodiment can determine the initial feature point matching pairs between the projected image and adjacent projected images. By filtering and sampling the initial feature point matching pairs between the projected image and adjacent projected images, the final feature point matching pairs can be obtained, making the obtained feature point matching pairs more accurate and more evenly distributed.

[0076] In one embodiment, determining an initial feature point matching pair between a projected image and adjacent projected images includes:

[0077] Feature points are extracted from the projected image to obtain multiple first feature points in the projected image;

[0078] In adjacent projected images, identify the second feature point that has the highest feature similarity to the first feature point;

[0079] For each first feature point, if the first feature point is the feature point with the highest feature similarity to the second feature point in the projected image, then an initial feature point matching pair is established based on the first feature point and the corresponding second feature point.

[0080] First, feature points are extracted from the projected image to obtain multiple first feature points. Then, feature points are extracted from adjacent projected images to obtain multiple second feature points. When extracting feature points from the projected image, the SIFT (Scale-Invariant Feature Transform) algorithm can be used. Alternatively, the ORB (Oriented Fast and Rotated BRIEF) algorithm and depth feature extraction algorithms can also be used.

[0081] Next, after obtaining multiple first feature points from the projected image, a brute-force matching method can be used to match the second feature point with the highest similarity to each first feature point among multiple second feature points in adjacent projected images. For each first feature point and the second feature point with the highest similarity to the first feature point, if the first feature point is the feature point with the highest feature similarity to the second feature points in the adjacent projected images (i.e., the first feature point and the second feature point are mutually the most similar), then the first feature point and the corresponding second feature point constitute an initial feature point matching pair. If the first feature point is not the feature point with the highest feature similarity to the second feature points in the adjacent projected images (i.e., the first feature point and the second feature point are not mutually the most similar), then the first feature point and the corresponding second feature point do not constitute an initial feature point matching pair.

[0082] As can be seen, this embodiment can establish initial feature point matching pairs between the projected image and adjacent projected images through bidirectional matching, which can improve the accuracy of feature point matching, thereby making the initial feature point matching pairs between the projected image and adjacent projected images more accurate.

[0083] After determining the initial feature point matching pairs between the projected image and its neighboring projected images, considering that there may be erroneous feature point matching pairs in the initial feature point matching pairs, it is necessary to filter and sample the initial feature point matching pairs between the projected image and its neighboring projected images to obtain the final feature point matching pairs between the projected image and its neighboring projected images.

[0084] In one embodiment, initial feature point matching pairs between the projected image and adjacent projected images are filtered and sampled to obtain final feature point matching pairs between the projected image and adjacent projected images, including:

[0085] Based on the feature point coordinate difference between the first and second feature points in each initial feature point matching pair, each initial feature point matching pair is filtered.

[0086] Based on the confidence level of each initial feature point matching pair obtained through screening, the initial feature point matching pairs are sampled to obtain the final feature point matching pairs.

[0087] Image features typically include brightness, color, and grayscale. When matching feature points between a projected image and adjacent projected images, incorrect matching may occur because a certain feature in a certain region of the two projected images is the same. For example, if the video scene is a car interior, and the steering wheel and center console are the same color, when matching feature points between the projected image and adjacent projected images, a feature point in the steering wheel area of ​​the projected image and a feature point in the center console area of ​​the adjacent projected image might be considered as a feature point matching pair, resulting in an incorrect feature point matching pair. Therefore, when filtering initial feature point matching pairs, the difference in feature point coordinates between the first and second feature points in each initial feature point matching pair can be used. Coordinates reflect the positional information of feature points or pixels. If the first and second feature points are a correct feature point matching pair, then the difference in coordinates between the first and second feature points should be within a reasonable range.

[0088] The number of feature point matching pairs after screening is still large, and the uneven distribution of feature point matching pairs will affect the calculation of the pixel coordinate transformation matrix between the projected images. After screening each initial feature point matching pair, the confidence level of each initial feature point matching pair obtained by screening can be used to sample each initial feature point matching pair to obtain the final feature point matching pair.

[0089] As can be seen, through this embodiment, each initial feature point matching pair is filtered based on the feature point coordinate difference between the first and second feature points in each initial feature point matching pair. Based on the confidence level corresponding to each initial feature point matching pair obtained by the filtering, each initial feature point matching pair is sampled to obtain the final feature point matching pair. This can filter out accurate feature point matching pairs and sample more evenly distributed feature point matching pairs, thereby improving the accuracy and uniformity of feature point matching.

[0090] In one embodiment, the initial feature point matching pairs are filtered based on the feature point coordinate difference between the first and second feature points in each initial feature point matching pair, including:

[0091] The mean value of the feature point coordinate difference between each first feature point in the projected image and each corresponding second feature point in the adjacent image is calculated.

[0092] For each initial feature point matching pair, if the difference in feature point coordinates between the first and second feature points in the initial feature point matching pair meets the numerical requirement corresponding to the mean, then the initial feature point matching pair is retained; otherwise, the initial feature point matching pair is discarded.

[0093] When filtering initial feature point matching pairs based on the feature point coordinate differences between the first and second feature points, the process begins by obtaining the difference between the X-coordinate (diffX) and the difference between the Y-coordinate (diffY) of each first feature point and its corresponding second feature point in the adjacent image. Next, the obtained diffX and diffY values ​​are sorted in ascending order to obtain the X-coordinate difference sequence diffX_sort_list and the Y-coordinate difference sequence diffY_sort_list.

[0094] Then, in the X-coordinate difference sequence `diffX_sort_list` and the Y-coordinate difference sequence `diffY_sort_list`, the first 10% and the last 10% of the coordinate differences are discarded, respectively. The mean of the X-coordinate differences, `vote_diffX`, is calculated using the remaining X-coordinate differences, and the mean of the Y-coordinate differences, `vote_diffY`, is calculated using the remaining Y-coordinate differences. This yields the mean values ​​of the X-coordinate differences, `vote_diffX` and `vote_diffY`.

[0095] Next, the coordinate difference requirements can be determined based on the mean difference of the X coordinates (vote_diffX) and the mean difference of the Y coordinates (vote_diffY), and the initial feature point matching pairs can be filtered. It can be set that when the X coordinate difference is within the range of ((1-α)*vote_diffX, (1+α)*vote_diffX), and the Y coordinate difference is within the range of ((1-α)*vote_diffY, (1+α)*vote_diffY), the corresponding initial feature point matching pair is considered a correct feature matching pair. For example, the value of α can be 0.2.

[0096] For each initial feature point matching pair, if the feature point coordinate difference between the first feature point and the second feature point in the initial feature point matching pair meets the above coordinate difference requirement, then the initial feature point matching pair is a correct feature point matching pair and is retained; otherwise, the initial feature point matching pair is discarded.

[0097] As can be seen, through this embodiment, it is possible to determine whether the feature point coordinate difference between the first feature point and the second feature point in the initial feature point matching pair meets the numerical requirement corresponding to the mean value based on the mean value of the feature point coordinate difference between each first feature point in the projected image and each corresponding second feature point in the adjacent image, thereby making the screening results more accurate and thus obtaining the initial feature point matching pair that meets the requirements.

[0098] After obtaining the initial feature matching pairs after filtering, the filtered feature point matching pairs are further sampled to obtain the final feature point matching pairs.

[0099] In one embodiment, based on the confidence level corresponding to each initial feature point matching pair obtained through screening, sampling is performed on each initial feature point matching pair to obtain the latest feature point matching pair, including:

[0100] Divide the projected image into multiple grids;

[0101] Select the initial feature point matching pair with the highest confidence in each grid;

[0102] The selected initial feature point matching pairs are determined as the final feature point matching pairs.

[0103] Understandably, if multiple feature point matching pairs between the projected image and its adjacent projected images are concentrated in a certain region, then when transforming the second image to the pixel coordinate system of the first image, that region will be aligned first. Conversely, when the feature point matching pairs are more evenly distributed, the pixel coordinate transformation matrix will obtain an optimal solution that can align the regions where all feature point matching pairs are distributed.

[0104] Based on this, when sampling each initial feature point matching pair obtained from the screening based on the confidence level corresponding to each initial feature point matching pair, the projected image can first be divided into many grids, and the side length of the grid can be set to L. For example, the value of L can be 100 pixels. Then, the pair of initial feature point matching pairs with the highest confidence level is selected in each grid, and the selected initial feature point matching pairs are determined as the final feature point matching pairs.

[0105] As can be seen, through this embodiment, by selecting the initial feature point matching pair with the highest confidence in each grid corresponding to the projected image, and determining the selected initial feature point matching pair as the final feature point matching pair, the feature point matching pair with the highest confidence in each grid can be obtained, making the distribution of the obtained feature point matching pairs more uniform.

[0106] In one embodiment, it also includes:

[0107] For the first feature point in the initial feature point matching pair, determine the third feature point with the second highest similarity to the first feature point in the adjacent projected images;

[0108] The confidence level of the initial feature point matching pair is determined based on the first feature similarity between the first feature point and the corresponding second feature point, and the second feature similarity between the first feature point and the corresponding third feature point.

[0109] During feature point matching or when calculating the confidence of a feature point matching pair, for the first feature point in the initial feature point matching pair, among multiple second feature points in adjacent projected images, a third feature point with the second highest similarity to the first feature point can be matched. The ratio of the similarity between the first and third feature points to the similarity between the first and second feature points is used as the confidence of the feature point matching pair formed by the first and second feature points. The similarity between the first and second feature points can be obtained by calculating the Euclidean distance between them, and the similarity between the first and third feature points can also be obtained by calculating the Euclidean distance between them. For example, if the Euclidean distance between the first and second feature points is 1, and the Euclidean distance between the first and third feature points is 3, then the confidence of the feature point matching pair formed by the first and second feature points is 3, i.e., 3 / 1 = 3.

[0110] As can be seen, through this embodiment, by determining the third feature point, and based on the first feature similarity between the first feature point and the corresponding second feature point, and the second feature similarity between the first feature point and the corresponding third feature point, the confidence level of the initial feature point matching pair can be determined, which can more accurately assess the reliability of the initial feature point matching pair and improve the accuracy of the confidence level.

[0111] After obtaining the final feature point matching pairs, the pixel coordinate transformation matrix between the projected image and the adjacent projected image is determined based on the final feature point matching pairs. In one example, it can be that the homography matrix between the projected image and the adjacent projected image is determined based on the final feature point matching pairs, and the homography matrix is ​​used as the pixel coordinate transformation matrix. For the specific process of determining the homography matrix between the projected image and the adjacent projected image, please refer to relevant technologies, which will not be elaborated here.

[0112] After determining the pixel coordinate transformation matrix between the projected image and the adjacent projected images, the second image can be transformed into the pixel coordinate system of the first image according to the pixel coordinate transformation matrix.

[0113] In one embodiment, transforming the second image to the pixel coordinate system of the first image according to the pixel coordinate transformation matrix includes:

[0114] When the second image is adjacent to the first image, the second image is transformed into the pixel coordinate system of the first image according to the pixel coordinate transformation matrix between the second image and the first image;

[0115] When the second image is not adjacent to the first image, the transformation matrices of each pixel coordinate between the second image and the first image are concatenated. Based on the concatenated pixel coordinate transformation matrices, the second image is transformed into the pixel coordinate system of the first image.

[0116] Considering that each second image contains both adjacent and non-adjacent images to the first image, the pixel coordinate transformation matrix between two adjacent projected images can be calculated using the corresponding final feature point matching pairs. However, the pixel coordinate transformation matrix between two non-adjacent projected images needs to be obtained by concatenating the pixel coordinate transformation matrices of multiple adjacent projected images. Therefore, when transforming the second image to the pixel coordinate system of the first image based on the pixel coordinate transformation matrix, if the second image is adjacent to the first image (i.e., the second image is an adjacent projected image of the first image), then the second image is transformed to the pixel coordinate system of the first image based on the pixel coordinate transformation matrix between the second and first images. If the second image is not adjacent to the first image, the pixel coordinate transformation matrices between the second and first images are concatenated, and the second image is transformed to the pixel coordinate system of the first image based on the concatenated pixel coordinate transformation matrix.

[0117] For example, there are 5 projected images, numbered 1, 2, 3, 4, and 5. Projected image 3 is the first image, and projected images 1, 2, 4, and 5 are the second images. Projected images 2 and 4 are adjacent images of the first image, while projected images 1 and 5 are not adjacent images of the first image. When transforming the second image to the pixel coordinate system of the first image, for projected images 2 and 4, the pixel coordinate transformation matrix between them and the first image can be used to transform projected images 2 and 4 to the pixel coordinate system of projected image 3, respectively. For projected image 1, the pixel coordinate transformation matrix between projected image 1 and projected image 2, and the pixel coordinate transformation matrix between projected image 2 and projected image 3 need to be multiplied to obtain the pixel coordinate transformation matrix between projected image 1 and projected image 3. Then, based on the pixel coordinate transformation matrix between projected image 1 and projected image 3, projected image 1 is transformed to the pixel coordinate system of projected image 3. The process of transforming the projection image 5 to the pixel coordinate system of the projection image 3 can be referred to as the process of transforming the projection image 1 to the pixel coordinate system of the projection image 3, and will not be repeated here.

[0118] As can be seen, through this embodiment, the second image adjacent to the first image and the second image not adjacent to the first image can be transformed into the pixel coordinate system of the first image according to the pixel coordinate transformation matrix between the second image and the first image, so that the second image and the first image can be aligned at the pixel level, thereby making the image stitching more accurate.

[0119] In step S106 above, after transforming the second image into the pixel coordinate system of the first image, the first image and the transformed second image are stitched together to obtain a panoramic image.

[0120] In one embodiment, the first image and the transformed second image are stitched together to obtain a panoramic image, including:

[0121] Using an image region selection algorithm, stitching regions are selected in both the first image and the transformed second image.

[0122] By using an image stitching algorithm, the stitching area of ​​the first image and the stitching area of ​​the transformed second image are stitched together to obtain a panoramic image.

[0123] After transforming the second image into the pixel coordinate system of the first image, an image region selection algorithm, such as the seam-finder stitching selection algorithm, can be used to select the stitching regions that will participate in the final image stitching in both the first image and the transformed second image. Figure 2a This is a schematic diagram of an image stitching area provided in an embodiment of the present disclosure. Figure 2b This is a schematic diagram of another image stitching area provided in an embodiment of this disclosure. Figure 2c This is a schematic diagram of yet another image stitching region provided in an embodiment of this disclosure. In one example... Figure 2a , Figure 2b , Figure 2c The image shows the stitched area in three projected images selected from a car interior scene. The white area represents the stitched area, and the black area represents the non-stitched area.

[0124] After obtaining the stitching region, an image stitching algorithm is used to stitch the stitching region of the first image and the stitching region of the transformed second image together to obtain a panoramic image. For example, the multiblend fusion algorithm can be used to stitch the stitching region of the first image and the stitching region of the transformed second image together, and the colors of each image are blended during the stitching process to weaken the seams between adjacent images.

[0125] As can be seen, through this embodiment, by using the image region selection algorithm to select stitching regions in the first image and the transformed second image respectively, and by using the image stitching algorithm to stitch the stitching regions of the first image and the transformed second image together, the stitching regions of the images can be accurately determined, and the colors between adjacent images can be smoothly blended, resulting in a higher quality panoramic image.

[0126] It should be noted that the above values ​​are for illustrative purposes only and should not be used as specific limitations. In actual application scenarios, the values ​​can be set according to the actual needs of the scenario.

[0127] In one or more embodiments of this disclosure, firstly, a video to be processed is acquired; multiple video frames are selected as keyframes from the video to be processed; and the keyframes are projected onto a panoramic sphere to obtain projected images corresponding to the keyframes. A first image and a second image are determined from each projected image; the first image is a reference image among the projected images; and the second image is the image among the projected images excluding the reference image. The second image is transformed to the pixel coordinate system of the first image, and the first image and the transformed second image are stitched together to obtain a panoramic image. It can be seen that through this embodiment, by projecting each keyframe onto a sphere to obtain each projected image, and generating a panoramic image based on the projected images, the final generated panoramic image can have a good panoramic playback effect, improving the user's panoramic image browsing experience. Furthermore, by determining the first image and the second image in each projected image, using the first image as the reference image, transforming each second image to the pixel coordinate system of the first image, and stitching the first image and the transformed second image together, it is possible to stitch each projected image with the pixel coordinate system of the first image as the reference, thereby avoiding image quality problems caused by stitching multiple images, improving the image quality of the stitched panoramic image, and enhancing the user's panoramic image browsing experience.

[0128] Based on the same technical concept, one or more embodiments of this disclosure also provide an image processing apparatus corresponding to the image processing method described above. Figure 3 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of the present disclosure, as shown below. Figure 3 As shown, the device includes:

[0129] Projection module 31 is used to acquire the video to be processed, select multiple video frames as key frames in the video to be processed, and project the key frames onto the panoramic sphere to obtain the projected image corresponding to the key frames.

[0130] The determining module 32 is used to determine a first image and a second image in each of the projected images; the first image is a reference image in each of the projected images; the second image is an image in each of the projected images other than the reference image;

[0131] The stitching module 33 is used to transform the second image into the pixel coordinate system of the first image, and stitch the first image and the transformed second image together to obtain a panoramic image.

[0132] In one embodiment, the projection module 31 is specifically used for:

[0133] Select multiple initial keyframes from the video to be processed;

[0134] Obtain the pose information of the first camera corresponding to the initial keyframe and the pose information of the second camera corresponding to the adjacent initial keyframes of the initial keyframe;

[0135] Based on the pose information of the first camera and the pose information of the second camera, the final keyframe is selected from each of the initial keyframes.

[0136] In one embodiment, the projection module 31 is further specifically used for:

[0137] Based on the first camera pose information, determine the first rotation angle in the first rotation direction when the device capturing the initial keyframe of the video to be processed captures the video; based on the second camera pose information, determine the second rotation angle in the first rotation direction when the device capturing the adjacent keyframe of the video to be processed captures the video.

[0138] For each initial keyframe, if the angle difference between the first rotation angle and the second rotation angle is greater than the angle difference threshold, then the initial keyframe is determined as the final keyframe.

[0139] In one embodiment, the determining module 32 is specifically used for:

[0140] The image number of each projected image is determined according to the order of the frame numbers of each keyframe;

[0141] Based on the image sequence number of each of the projected images, the image whose image sequence number is located in the middle position among the projected images is determined to be the first image.

[0142] In one embodiment, the splicing module 33 is specifically used for:

[0143] Determine the pixel coordinate transformation matrix between the projected image and its adjacent projected images; the keyframe corresponding to the projected image is adjacent to the keyframe corresponding to the adjacent projected images;

[0144] The second image is transformed into the pixel coordinate system of the first image according to the pixel coordinate transformation matrix.

[0145] In one embodiment, the splicing module 33 is further specifically used for:

[0146] Determine feature point matching pairs between the projected image and the adjacent projected images;

[0147] Based on the feature point matching pairs, determine the pixel coordinate transformation matrix between the projected image and the adjacent projected images.

[0148] In one embodiment, the splicing module 33 is further specifically used for:

[0149] An initial feature point matching pair is determined between the projected image and the adjacent projected images; the initial feature point matching pair indicates that a first feature point in the projected image matches a second feature point in the adjacent projected image; the second feature point is the feature point in the adjacent projected image with the highest feature similarity to the first feature point; the first feature point is the feature point in the projected image with the highest feature similarity to the second feature point.

[0150] The initial feature point matching pairs between the projected image and the adjacent projected images are filtered and sampled to obtain the final feature point matching pairs between the projected image and the adjacent projected images.

[0151] In one embodiment, the splicing module 33 is further specifically used for:

[0152] Feature points are extracted from the projected image to obtain multiple first feature points in the projected image;

[0153] In the adjacent projected images, determine the second feature point that has the highest feature similarity to the first feature point;

[0154] For each of the first feature points, if the first feature point is the feature point in the projected image with the highest feature similarity to the second feature point, then the initial feature point matching pair is established based on the first feature point and the corresponding second feature point.

[0155] In one embodiment, the splicing module 33 is further specifically used for:

[0156] Based on the feature point coordinate difference between the first feature point and the second feature point in each initial feature point matching pair, each initial feature point matching pair is filtered;

[0157] Based on the confidence level of each initial feature point matching pair obtained through screening, each initial feature point matching pair is sampled to obtain the final feature point matching pair.

[0158] In one embodiment, the splicing module 33 is further specifically used for:

[0159] The mean value of the feature point coordinate difference between each first feature point in the projected image and each corresponding second feature point in the adjacent image is calculated.

[0160] For each initial feature point matching pair, if the difference in feature point coordinates between the first feature point and the second feature point in the initial feature point matching pair meets the numerical requirement corresponding to the mean, then the initial feature point matching pair is retained; otherwise, the initial feature point matching pair is discarded.

[0161] In one embodiment, the splicing module 33 is further specifically used for:

[0162] The projected image is divided into multiple grids;

[0163] In each of the grids, the initial feature point matching pair with the highest confidence level is selected;

[0164] The selected initial feature point matching pair is determined as the final feature point matching pair.

[0165] In one embodiment, the splicing module 33 is further configured to:

[0166] For the first feature point in the initial feature point matching pair, a third feature point with the second highest similarity to the first feature point is determined in the adjacent projected image;

[0167] The confidence level of the initial feature point matching pair is determined based on the first feature similarity between the first feature point and the corresponding second feature point, and the second feature similarity between the first feature point and the corresponding third feature point.

[0168] In one embodiment, the splicing module 33 is further specifically used for:

[0169] When the second image is adjacent to the first image, the second image is transformed into the pixel coordinate system of the first image according to the pixel coordinate transformation matrix between the second image and the first image;

[0170] When the second image is not adjacent to the first image, the pixel coordinate transformation matrices between the second image and the first image are concatenated, and the second image is transformed into the pixel coordinate system of the first image according to the concatenated pixel coordinate transformation matrices.

[0171] In one embodiment, the splicing module 33 is specifically used for:

[0172] The image region selection algorithm is used to select a stitching region in the first image and the transformed second image, respectively.

[0173] A panoramic image is obtained by stitching together the stitching region of the first image and the stitching region of the transformed second image using an image stitching algorithm.

[0174] In this embodiment, firstly, a video to be processed is acquired. Multiple video frames are selected as keyframes from the video and projected onto a panoramic sphere to obtain projected images corresponding to the keyframes. A first image and a second image are determined from each projected image. The first image is a reference image among the projected images; the second image is the image among the projected images excluding the reference image. The second image is transformed to the pixel coordinate system of the first image, and the first image and the transformed second image are stitched together to obtain a panoramic image. Therefore, through this embodiment, by projecting each keyframe onto a sphere to obtain projected images, and generating a panoramic image based on these projected images, the final generated panoramic image can have a good panoramic playback effect, improving the user's panoramic image browsing experience. Furthermore, by determining the first image and the second image in each projected image, using the first image as the reference image, transforming each second image to the pixel coordinate system of the first image, and stitching the first image and the transformed second image together, it is possible to stitch each projected image with the pixel coordinate system of the first image as the reference, thereby avoiding image quality problems caused by stitching multiple images, improving the image quality of the stitched panoramic image, and enhancing the user's panoramic image browsing experience.

[0175] The image processing apparatus in this embodiment can implement the various processes of the above-described image processing method embodiments and achieve the same effects and functions, which will not be repeated here.

[0176] One embodiment of this disclosure also provides an electronic device. Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, such as... Figure 4As shown, electronic devices can vary considerably due to differences in configuration or performance. They may include one or more processors 401 and memories 402, with the memory 402 storing one or more application programs or data. The memory 402 can be temporary or persistent storage. The application programs stored in the memory 402 may include one or more modules (not shown), each module including a series of computer-executable instructions within the electronic device. Furthermore, the processor 401 may be configured to communicate with the memory 402, executing the series of computer-executable instructions stored in the memory 402 on the electronic device. The electronic device may also include one or more power supplies 403, one or more wired or wireless network interfaces 404, one or more input or output interfaces 405, one or more keyboards 406, etc.

[0177] In this embodiment of the disclosure, the electronic device includes a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the following processes:

[0178] Acquire the video to be processed, select multiple video frames as key frames from the video to be processed, and project the key frames onto the panoramic sphere to obtain the projected image corresponding to the key frames.

[0179] A first image and a second image are determined in each of the projected images; the first image is a reference image in each of the projected images; the second image is an image in each of the projected images other than the reference image.

[0180] The second image is transformed into the pixel coordinate system of the first image, and the first image and the transformed second image are stitched together to obtain a panoramic image.

[0181] The electronic device in this embodiment can implement the various processes of the above-described image processing method embodiments and achieve the same effects and functions, which will not be repeated here.

[0182] Another embodiment of this disclosure also provides a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the following process:

[0183] Acquire the video to be processed, select multiple video frames as key frames from the video to be processed, and project the key frames onto the panoramic sphere to obtain the projected image corresponding to the key frames.

[0184] A first image and a second image are determined in each of the projected images; the first image is a reference image in each of the projected images; the second image is an image in each of the projected images other than the reference image.

[0185] The second image is transformed into the pixel coordinate system of the first image, and the first image and the transformed second image are stitched together to obtain a panoramic image.

[0186] The storage medium in this embodiment can implement the various processes of the above-described image processing method embodiments and achieve the same effects and functions, which will not be repeated here.

[0187] Another embodiment of this disclosure also provides a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the following process:

[0188] Acquire the video to be processed, select multiple video frames as key frames from the video to be processed, and project the key frames onto the panoramic sphere to obtain the projected image corresponding to the key frames.

[0189] A first image and a second image are determined in each of the projected images; the first image is a reference image in each of the projected images; the second image is an image in each of the projected images other than the reference image.

[0190] The second image is transformed into the pixel coordinate system of the first image, and the first image and the transformed second image are stitched together to obtain a panoramic image.

[0191] The computer program product in this disclosure embodiment can implement the various processes of the above-described image processing method embodiment and achieve the same effects and functions, which will not be repeated here.

[0192] In various embodiments of this disclosure, the computer-readable storage medium includes read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc.

[0193] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using a hardware physical module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0194] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0195] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0196] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing the embodiments of this disclosure, the functions of each unit can be implemented in one or more software and / or hardware.

[0197] Those skilled in the art will understand that one or more embodiments of this disclosure can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0198] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0199] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0200] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0201] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0202] One or more embodiments of this disclosure can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.

[0203] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0204] The above description is merely an embodiment of this disclosure and is not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.

Claims

1. An image processing method, characterized in that, include: Acquire the video to be processed, select multiple video frames as key frames from the video to be processed, and project the key frames onto the panoramic sphere to obtain the projected image corresponding to the key frames. A first image and a second image are determined from each of the projected images; the first image is a reference image among the projected images. The second image is any image other than the reference image among the various projected images; The second image is transformed into the pixel coordinate system of the first image, and the first image and the transformed second image are stitched together to obtain a panoramic image.

2. The method according to claim 1, characterized in that, The step of selecting multiple video frames as keyframes from the video to be processed includes: Select multiple initial keyframes from the video to be processed; Obtain the pose information of the first camera corresponding to the initial keyframe and the pose information of the second camera corresponding to the adjacent initial keyframes of the initial keyframe; Based on the pose information of the first camera and the pose information of the second camera, the final keyframe is selected from each of the initial keyframes.

3. The method according to claim 2, characterized in that, The step of selecting the final keyframe from each of the initial keyframes based on the first camera pose information and the second camera pose information includes: Based on the first camera pose information, determine the first rotation angle in the first rotation direction when the device capturing the initial keyframe of the video to be processed captures the video; based on the second camera pose information, determine the second rotation angle in the first rotation direction when the device capturing the adjacent keyframe of the video to be processed captures the video. For each initial keyframe, if the angle difference between the first rotation angle and the second rotation angle is greater than the angle difference threshold, then the initial keyframe is determined as the final keyframe.

4. The method according to claim 1, characterized in that, Determining the first image among the various projected images includes: The image number of each projected image is determined according to the order of the frame numbers of each keyframe; Based on the image sequence number of each of the projected images, the image whose image sequence number is located in the middle position among the projected images is determined to be the first image.

5. The method according to claim 1, characterized in that, The step of transforming the second image into the pixel coordinate system of the first image includes: Determine the pixel coordinate transformation matrix between the projected image and its adjacent projected images; the keyframe corresponding to the projected image is adjacent to the keyframe corresponding to the adjacent projected images; The second image is transformed into the pixel coordinate system of the first image according to the pixel coordinate transformation matrix.

6. The method according to claim 5, characterized in that, Determining the pixel coordinate transformation matrix between the projected image and its adjacent projected images includes: Determine feature point matching pairs between the projected image and the adjacent projected images; Based on the feature point matching pairs, determine the pixel coordinate transformation matrix between the projected image and the adjacent projected images.

7. The method according to claim 6, characterized in that, Determining the feature point matching pairs between the projected image and the adjacent projected images includes: An initial feature point matching pair is determined between the projected image and the adjacent projected images; the initial feature point matching pair indicates that a first feature point in the projected image matches a second feature point in the adjacent projected image; the second feature point is the feature point in the adjacent projected image with the highest feature similarity to the first feature point; the first feature point is the feature point in the projected image with the highest feature similarity to the second feature point. The initial feature point matching pairs between the projected image and the adjacent projected images are filtered and sampled to obtain the final feature point matching pairs between the projected image and the adjacent projected images.

8. The method according to claim 7, characterized in that, Determining the initial feature point matching pair between the projected image and the adjacent projected images includes: Feature points are extracted from the projected image to obtain multiple first feature points in the projected image; In the adjacent projected images, determine the second feature point that has the highest feature similarity to the first feature point; For each of the first feature points, if the first feature point is the feature point in the projected image with the highest feature similarity to the second feature point, then the initial feature point matching pair is established based on the first feature point and the corresponding second feature point.

9. The method according to claim 7, characterized in that, The step of filtering and sampling the initial feature point matching pairs between the projected image and the adjacent projected images to obtain the final feature point matching pairs between the projected image and the adjacent projected images includes: Based on the feature point coordinate difference between the first feature point and the second feature point in each initial feature point matching pair, each initial feature point matching pair is filtered; Based on the confidence level of each initial feature point matching pair obtained through screening, each initial feature point matching pair is sampled to obtain the final feature point matching pair.

10. The method according to claim 9, characterized in that, The step of filtering each initial feature point matching pair based on the feature point coordinate difference between the first feature point and the second feature point in each initial feature point matching pair includes: The mean value of the feature point coordinate difference between each first feature point in the projected image and each corresponding second feature point in the adjacent image is calculated. For each initial feature point matching pair, if the difference in feature point coordinates between the first feature point and the second feature point in the initial feature point matching pair meets the numerical requirement corresponding to the mean, then the initial feature point matching pair is retained; otherwise, the initial feature point matching pair is discarded.

11. The method according to claim 9, characterized in that, The step of sampling each of the initial feature point matching pairs obtained from the screening based on the confidence level corresponding to each initial feature point matching pair obtained from the screening to obtain the final feature point matching pair includes: The projected image is divided into multiple grids; In each of the grids, the initial feature point matching pair with the highest confidence level is selected; The selected initial feature point matching pair is determined as the final feature point matching pair.

12. The method according to claim 11, characterized in that, The method further includes: For the first feature point in the initial feature point matching pair, a third feature point with the second highest similarity to the first feature point is determined in the adjacent projected image; The confidence level of the initial feature point matching pair is determined based on the first feature similarity between the first feature point and the corresponding second feature point, and the second feature similarity between the first feature point and the corresponding third feature point.

13. The method according to claim 5, characterized in that, The step of transforming the second image into the pixel coordinate system of the first image according to the pixel coordinate transformation matrix includes: When the second image is adjacent to the first image, the second image is transformed into the pixel coordinate system of the first image according to the pixel coordinate transformation matrix between the second image and the first image; When the second image is not adjacent to the first image, the pixel coordinate transformation matrices between the second image and the first image are concatenated, and the second image is transformed into the pixel coordinate system of the first image according to the concatenated pixel coordinate transformation matrices.

14. The method according to claim 1, characterized in that, The step of stitching the first image and the transformed second image together to obtain a panoramic image includes: The image region selection algorithm is used to select a stitching region in the first image and the transformed second image, respectively. A panoramic image is obtained by stitching together the stitching region of the first image and the stitching region of the transformed second image using an image stitching algorithm.

15. An image processing apparatus, characterized in that, include: The projection module is used to acquire the video to be processed, select multiple video frames as key frames in the video to be processed, and project the key frames onto the panoramic sphere to obtain the projected image corresponding to the key frames. A determining module is configured to determine a first image and a second image among the various projected images; the first image is a reference image among the various projected images; The second image is any image other than the reference image among the various projected images; The stitching module is used to transform the second image into the pixel coordinate system of the first image, and stitch the first image and the transformed second image together to obtain a panoramic image.

16. An electronic device, characterized in that, include: processor; as well as, A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the steps of the method as described in any one of claims 1-14.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store computer-executable instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-14.

18. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1-14.