Image generation method, image processing method, and apparatus, medium, and program
Patent Information
- Application Number
- CN202510335993.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2026-09-22
AI Technical Summary
这种方法耗时耗力,并且效果难以保证,可能会导致生成的图像视角有限并清晰度不足
[0015]根据本发明实施例的上述图像生成方法、装置、计算机可读存储介质和程序,能够将所拍摄的全景视频分别实施具有不同畸变区域的不同的投影变换方式,从而使得拼接后的拍摄对象的二维视图能够体现出没有畸变的区域的真实图像,以展示出清晰且准确的二维视图。
Smart Images

Figure CN122802795A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, and more particularly to an image generation method, apparatus, computer-readable medium, and computer program product, as well as an image processing method, apparatus, computer-readable medium, and computer program product. Background Technology
[0002] Methods for precise imaging of specific objects (such as vehicle chassis, objects within narrow slits, etc.) are widely needed in practical applications. For example, for vehicle chassis, there is a need to obtain precise two-dimensional views of the chassis during vehicle sales, repair, or inspection to better understand the specific condition of the chassis, such as whether there is rust, scratches, or to check whether illegal items are hidden.
[0003] However, capturing two-dimensional images of these specific objects, especially those that are difficult to approach or cannot be directly photographed, presents significant challenges. For example, photographing the undercarriage of a vehicle might require parking in a special spot (such as a space with a hollow area underneath), or using tools to lift the vehicle to inspect or photograph the undercarriage. This method is time-consuming and labor-intensive, and the results are unpredictable, potentially leading to images with limited perspective and insufficient clarity.
[0004] Therefore, there is a need for a method and apparatus that can conveniently and quickly obtain a clear two-dimensional view of a specific subject and obtain accurate local images of a specific area of the obtained two-dimensional view. Summary of the Invention
[0005] To address the aforementioned technical problems, according to one aspect of the present invention, an image generation method is provided, comprising: acquiring a panoramic video of a subject captured by a moving panoramic camera; performing a first projection transformation on multiple video frames in the panoramic video to generate multiple first projected video frames, and performing a second projection transformation on the multiple video frames in the panoramic video to generate multiple second projected video frames, wherein a first distortion region in the first projected video frames is different from a second distortion region in the second projected video frames; stitching together first undistorted regions in the multiple first projected video frames to generate a first two-dimensional view of the subject, and stitching together second undistorted regions in the multiple second projected video frames to generate a second two-dimensional view of the subject.
[0006] According to another aspect of the present invention, an image processing method is provided, comprising: acquiring a first two-dimensional view and a second two-dimensional view of a photographed object, wherein the first two-dimensional view and the second two-dimensional view are generated by performing a first projection transformation and a second projection transformation respectively on multiple video frames in a panoramic video of the photographed object, wherein a first distortion region in the first projection video frame is different from a second distortion region in the second projection video frame, wherein the first two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames, and stitching together a first non-distorted region in the multiple first projection video frames, and the second two-dimensional view is generated by... Multiple second-projection video frames are generated by performing a second projection transformation on the multiple video frames, and the second undistorted regions in the multiple second-projection video frames are stitched together to generate the second-projection video frame. The position of the first region in the subject is obtained, and it is determined whether the first region is located in the first undistorted region of the first-projection video frame or in the second undistorted region of the second-projection video frame. If the first region is located in the first undistorted region of the first-projection video frame, the first region is associated with the corresponding position in the first two-dimensional view. If the first region is located in the second undistorted region of the second-projection video frame, the first region is associated with the corresponding position in the second two-dimensional view.
[0007] According to another aspect of the present invention, an image generation apparatus is provided, comprising: a shooting unit configured to acquire a panoramic video of a subject captured by a moving panoramic camera; a generation unit configured to perform a first projection transformation on a plurality of video frames in the panoramic video to generate a plurality of first projected video frames, and to perform a second projection transformation on the plurality of video frames in the panoramic video to generate a plurality of second projected video frames, wherein a first distortion region in the first projected video frames is different from a second distortion region in the second projected video frames; and a stitching unit configured to stitch together a first non-distorted region in the plurality of first projected video frames to generate a first two-dimensional view of the subject, and to stitch together a second non-distorted region in the plurality of second projected video frames to generate a second two-dimensional view of the subject.
[0008] According to another aspect of the present invention, an image generation apparatus is provided, comprising: a processor; and a memory storing computer program instructions, wherein, when the computer program instructions are executed by the processor, the processor performs the following steps: acquiring a panoramic video of a subject captured using a moving panoramic camera; performing a first projection transformation on a plurality of video frames in the panoramic video to generate a plurality of first projected video frames, and performing a second projection transformation on the plurality of video frames in the panoramic video to generate a plurality of second projected video frames, wherein a first distortion region in the first projected video frames is different from a second distortion region in the second projected video frames; stitching together a first non-distorted region in the plurality of first projected video frames to generate a first two-dimensional view of the subject, and stitching together a second non-distorted region in the plurality of second projected video frames to generate a second two-dimensional view of the subject.
[0009] According to another aspect of the present invention, a computer-readable storage medium is provided having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, perform the following steps: acquiring a panoramic video of a subject captured by a moving panoramic camera; performing a first projection transformation on a plurality of video frames in the panoramic video to generate a plurality of first projected video frames, and performing a second projection transformation on the plurality of video frames in the panoramic video to generate a plurality of second projected video frames, wherein a first distortion region in the first projected video frames is different from a second distortion region in the second projected video frames; stitching together a first non-distorted region in the plurality of first projected video frames to generate a first two-dimensional view of the subject, and stitching together a second non-distorted region in the plurality of second projected video frames to generate a second two-dimensional view of the subject.
[0010] According to another aspect of the present invention, a computer program product is provided, comprising computer program instructions, wherein the computer program instructions, when executed by a processor, implement the steps of the image generation method described in any of the preceding claims.
[0011] According to another aspect of the present invention, an image processing apparatus is provided, comprising: an acquisition unit configured to acquire a first two-dimensional view and a second two-dimensional view of a photographed object, wherein the first two-dimensional view and the second two-dimensional view are generated by performing a first projection transformation and a second projection transformation on multiple video frames in a panoramic video of the photographed object, respectively, wherein a first distortion region in the first projection video frame is different from a second distortion region in the second projection video frame, wherein the first two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames, and stitching together a first non-distorted region in the multiple first projection video frames, and the second ... the second two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames, and the second two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames, and the second two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames, and the second two-dimensional Multiple video frames undergo a second projection transformation to generate multiple second projected video frames, and the second undistorted regions in the multiple second projected video frames are stitched together to generate the second projected video frame. A judgment unit is configured to obtain the position of a first region in the captured object and determine whether the first region is located in the first undistorted region of the first projected video frame or in the second undistorted region of the second projected video frame. An association unit is configured to associate the first region with a corresponding position in the first two-dimensional view if the first region is located in the first undistorted region of the first projected video frame; and to associate the first region with a corresponding position in the second two-dimensional view if the first region is located in the second undistorted region of the second projected video frame.
[0012] According to another aspect of the present invention, an image processing apparatus is provided, comprising: a processor; and a memory storing computer program instructions, wherein, when the computer program instructions are executed by the processor, the processor performs the following steps: acquiring a first two-dimensional view and a second two-dimensional view of a photographed object, the first two-dimensional view and the second two-dimensional view being generated by performing a first projection transformation and a second projection transformation respectively on multiple video frames in a panoramic video of the photographed object, wherein a first distortion region in the first projection video frame is different from a second distortion region in the second projection video frame, wherein the first two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames, and the multiple first projection video frames are then transformed by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames. The first undistorted region in the projected video frames is stitched together to generate the second two-dimensional view. The second two-dimensional view is generated by performing a second projection transformation on the multiple video frames to generate multiple second projected video frames, and then stitching together the second undistorted regions in the multiple second projected video frames. The position of the first region in the captured object is obtained, and it is determined whether the first region is located in the first undistorted region of the first projected video frame or in the second undistorted region of the second projected video frame. If the first region is located in the first undistorted region of the first projected video frame, the first region is associated with the corresponding position in the first two-dimensional view. If the first region is located in the second undistorted region of the second projected video frame, the first region is associated with the corresponding position in the second two-dimensional view.
[0013] According to another aspect of the present invention, a computer-readable storage medium is provided, having stored thereon computer program instructions, wherein the computer program instructions, when executed by a processor, perform the following steps: acquiring a first two-dimensional view and a second two-dimensional view of a photographed object, wherein the first two-dimensional view and the second two-dimensional view are generated by performing a first projection transformation and a second projection transformation respectively on multiple video frames in a panoramic video of the photographed object, wherein a first distortion region in the first projection video frame is different from a second distortion region in the second projection video frame, wherein the first two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames, and a first non-distorted region in the multiple first projection video frames is selected. The second two-dimensional view is generated by stitching together the domains. It is generated by performing a second projection transformation on the multiple video frames to generate multiple second projected video frames, and then stitching together the second non-distorted regions within these multiple second projected video frames. The position of the first region in the captured object is obtained, and it is determined whether the first region is located in the first non-distorted region of the first projected video frame or in the second non-distorted region of the second projected video frame. If the first region is located in the first non-distorted region of the first projected video frame, it is associated with the corresponding position in the first two-dimensional view. If the first region is located in the second non-distorted region of the second projected video frame, it is associated with the corresponding position in the second two-dimensional view.
[0014] According to another aspect of the present invention, a computer program product is provided, comprising computer program instructions, wherein the computer program instructions, when executed by a processor, implement the steps of any of the preceding image processing methods.
[0015] According to the above-described image generation method, apparatus, computer-readable storage medium, and program of the present invention, different projection transformation methods can be applied to the captured panoramic video with different distortion regions, so that the stitched two-dimensional view of the captured object can reflect the true image of the distortion-free region, thereby displaying a clear and accurate two-dimensional view.
[0016] Furthermore, the image processing method, apparatus, computer-readable storage medium, and program according to embodiments of the present invention can also establish a correspondence between a specified first region and a local two-dimensional view of a non-distorted region in a corresponding projection mode, and display the first region and the corresponding non-distorted two-dimensional view accordingly, greatly improving the user experience. Attached Figure Description
[0017] The above and other objects, features, and advantages of the present invention will become clearer from the detailed description of the embodiments of the present invention in conjunction with the accompanying drawings.
[0018] Figure 1 A flowchart illustrating an image generation method according to an embodiment of the present invention is shown;
[0019] Figure 2 A flowchart illustrating an image processing method according to an embodiment of the present invention is shown;
[0020] Figure 3 A schematic diagram showing video frames of an example panoramic video according to an embodiment of the present invention;
[0021] Figure 4 A schematic diagram of a user interface for panoramic video uploading and region-related parameter configuration is shown as an example according to an embodiment of the present invention;
[0022] Figure 5 A schematic diagram illustrating an example of a region configuration according to an embodiment of the present invention;
[0023] Figure 6 A schematic diagram illustrating an example of the area configuration related to the model of the vehicle being photographed, according to an embodiment of the present invention;
[0024] Figure 7 A schematic diagram illustrating an example of the area configuration related to the model of the vehicle being filmed and the video generation time, according to an embodiment of the present invention;
[0025] Figure 8 A schematic diagram of a user interface for panoramic video uploading and region-related parameter configuration is shown as an example according to an embodiment of the present invention;
[0026] Figure 9 This diagram illustrates two consecutive adjacent keyframes extracted in an example according to an embodiment of the present invention.
[0027] Figure 10 This diagram illustrates an example of splicing two adjacent keyframes according to an embodiment of the present invention.
[0028] Figure 11 This document illustrates a specific step in stitching together projected video frames to obtain a corresponding two-dimensional view, according to an embodiment of the present invention.
[0029] Figure 12 This diagram illustrates how feature points between perspective video frames are stitched together in one example of an embodiment of the present invention.
[0030] Figure 13 This diagram illustrates, in one example of an embodiment of the present invention, the stitching of feature points between equally spaced columnar video frames.
[0031] Figure 14An example of the image generation result of an embodiment of the present invention is shown;
[0032] Figure 15 A user interface schematic diagram showing an overview of an example according to an embodiment of the present invention;
[0033] Figure 16 A schematic diagram illustrating first region information is shown as an example of an embodiment of the present invention;
[0034] Figure 17 A flowchart illustrating a partial image file generation process according to an embodiment of the present invention is shown.
[0035] Figure 18 A table example showing the association of a partial image of a first region with different projected image files according to an embodiment of the present invention;
[0036] Figure 19 An example of a user interface displaying a partial detail image according to an embodiment of the present invention;
[0037] Figure 20 A block diagram of an image generation apparatus according to an embodiment of the present invention is shown;
[0038] Figure 21 A block diagram of an image processing apparatus according to an embodiment of the present invention is shown;
[0039] Figure 22 A block diagram of an image generation apparatus according to an embodiment of the present invention is shown;
[0040] Figure 23 A block diagram of an image processing apparatus according to an embodiment of the present invention is shown. Detailed Implementation
[0041] Image generation methods, apparatus, computer-readable media, and computer program products according to embodiments of the present invention will now be described with reference to the accompanying drawings. In the drawings, the same reference numerals denote the same elements throughout. It should be understood that the embodiments described herein are merely illustrative and should not be construed as limiting the scope of the invention.
[0042] For subjects that are difficult to approach or whose two-dimensional views cannot be directly captured, obtaining a quick and accurate two-dimensional view is a challenging problem. To address this, this invention provides an image generation method and an image processing method that use different projection transformations to stitch and display distortion-free two-dimensional views of the subject's panoramic video.
[0043] Figure 1A flowchart of an image generation method 100 according to an embodiment of the present invention is shown. Referring below... Figure 1 An image generation method according to an embodiment of the present invention is described.
[0044] In step S101, a panoramic video of the subject is acquired using a moving panoramic camera.
[0045] In this embodiment of the invention, a panoramic camera can be used to move at a specific speed and traverse one or more surfaces of the subject to obtain a panoramic video of the subject. Optionally, the panoramic camera can move at a preset, uniform speed during shooting. Furthermore, the panoramic camera can optionally shoot along a straight line. Of course, considering the requirements of the shooting angle and the size of the subject, the panoramic camera can also shoot sequentially in multiple directions, and the resulting videos can be segmented and stitched together according to the shooting route to ultimately obtain a complete two-dimensional view of the subject.
[0046] Optionally, when the subject is difficult to photograph directly, a vehicle-mounted panoramic camera or a robotic arm equipped with a panoramic camera can be used to capture the subject. In one example, when using a vehicle-mounted panoramic camera for panoramic video recording, the vehicle's route and speed can be preset to traverse the subject and obtain panoramic video. The above-described methods and settings for panoramic video recording of the subject are merely examples. In practical applications, any video recording method equipped with a panoramic camera can be used according to specific needs, and no restrictions are imposed here.
[0047] In S102, a first projection transformation is performed on multiple video frames in the panoramic video to generate multiple first projection video frames, and a second projection transformation is performed on multiple video frames in the panoramic video to generate multiple second projection video frames, wherein the first distortion region that produces distortion in the first projection video frame is different from the second distortion region that produces distortion in the second projection video frame.
[0048] In this embodiment of the invention, the acquired panoramic video of the subject can be divided into multiple video frames, and at least two projection transformations can be performed on each of the acquired video frames. Specifically, at least a first projection transformation and a second projection transformation can be performed on the multiple video frames in the panoramic video, and multiple first projection video frames and multiple second projection video frames can be generated accordingly. Of course, in another embodiment of the invention, the projection transformation performed on the video frames of the panoramic video can also be more than two projection transformations, for example, it can include a third projection transformation, which is not limited here. In one example, the distortion region of the third projection transformation can be the same as or different from the distortion region of the first projection transformation / second projection transformation. In the subsequent selection of the two-dimensional view, it can be selected or combined according to the specific application scenario and image projection effect.
[0049] In this embodiment of the invention, in order to obtain accurate two-dimensional views without distortion for different parts of the subject being photographed, the first distorted region in the first projected video frame needs to be different from the second distorted region in the second projected video frame. In this embodiment, the first projection transformation and the second projection transformation can be one of various projection transformation methods, such as perspective transformation, equidistant cylindrical projection transformation, cylindrical projection transformation, Mercator transformation, etc. Depending on the projection transformation method, the corresponding distorted and undistorted regions are different. Therefore, the first projection transformation and the second projection transformation can be selected from these projection transformation methods that have different distorted and undistorted regions. For example, in one example of this embodiment, the first projection transformation is a perspective transformation, and the first undistorted region is the middle region of the first projected video frame; the second projection transformation is an equidistant cylindrical projection transformation, and the second undistorted region is the two ends of the second projected video frame; or the first projection transformation is an equidistant cylindrical projection transformation, and the first undistorted region is the two ends of the first projected video frame; the second projection transformation is a perspective transformation, and the second undistorted region is the middle region of the second projected video frame. In one example, the distortion region of the third projection transformation may be the same as or different from the distortion region of the first / second projection transformation. During the subsequent selection of the two-dimensional view, the selection or combination can be made depending on the specific application scenario and image projection effect. Furthermore, in one example of this embodiment, the middle region of each projected video frame can be the region in the middle part along the direction of motion of the panoramic camera in that projected video frame, while the two end regions of each projected video frame can be the regions in the two end parts along the direction of motion of the panoramic camera in that projected video frame.
[0050] In one example of this invention, the region configuration between the undistorted and distorted regions in each projection transformation method can be preset, for example, it can be specified by the user. In another example of this invention, the region configuration between the undistorted and distorted regions in each projection transformation method can also be set by region-related parameters. Optionally, the region-related parameters may include one or more parameters among the type, size, product model, shooting position, and shooting time of the object being photographed. In one example, when the object being photographed is a vehicle chassis, the corresponding region-related parameters may include the vehicle model, the specific size of the vehicle chassis, the shooting position of the panoramic camera used for shooting, the shooting speed during movement, etc. For example, for a certain vehicle model, the first undistorted region corresponding to the projection can be the middle 25%-75% of the video frame in the panoramic video, and the remaining area is the first distorted region. As another example, for another vehicle model, the first undistorted region corresponding to the projection can be the middle 35%-70% of the video frame in the panoramic video, and the remaining area is the first distorted region. The above-described configuration methods for different regions in various projection transformations are merely examples. In practical applications, any corresponding region configuration can be selected according to specific needs, and no restrictions are imposed here. Furthermore, in this embodiment of the invention, the definitions of the various non-distorted regions and distorted regions are only relative. That is, each non-distorted region indicates a region with a relatively small degree of distortion after projection relative to other regions, and is not limited to regions without distortion, while each distorted region indicates a region with a relatively large degree of distortion after projection relative to other regions. In one example of this embodiment, when the ranges of the non-distorted and distorted regions are specified according to user input, they refer to the specific regions defined by the user. In one example of this embodiment, the distorted and distorted regions can be continuous and non-overlapping; however, in another example, the distorted and distorted regions can be discontinuous or overlapping, and no restrictions are imposed here.
[0051] In fact, these regions can be discontinuous or overlapping.
[0052] In step S103, a first two-dimensional view of the photographed object is generated by stitching together the first non-distorted regions in the plurality of first projected video frames, and a second two-dimensional view of the photographed object is generated by stitching together the second non-distorted regions in the plurality of second projected video frames.
[0053] After obtaining each video frame from the panoramic video, optionally, all video frames or some of the video frames may be selected for subsequent two-dimensional view stitching. According to the embodiment of the present invention, optionally, all first projection video frames obtained by performing a first projection transformation on the panoramic video can be stitched; in addition, optionally, considering that there may be a large amount of overlapping regions between the first projection video frames leading to redundant calculation, a plurality of first key frames can also be selected from the plurality of first projection video frames for stitching.
[0054] Of course, similarly, the same processing method can also be adopted for the second projection transformation, that is, all the second projection video frames obtained by performing the second projection transformation on the panoramic video can be stitched, and a plurality of second key frames can also be selected from the plurality of second projection video frames for stitching.
[0055] In the specific key frame extraction process, optionally, one frame may be extracted as a key frame from every N projected video frames, for example, N may be a fixed value. In addition, optionally, key frames can also be extracted by using corresponding image processing algorithms according to the shooting parameters of the panoramic camera during shooting. In an example, the shooting parameters may be various parameters such as the shooting speed of the panoramic video, the type, size, and product model of the shooting object. After obtaining the shooting parameters, optionally, various image processing algorithms such as a similarity-based method and a clustering method can be used to extract key frames.
[0056] In an example of the embodiment of the present invention, when key frame extraction is performed by a similarity calculation method, optionally, a plurality of parts at different positions can be selected in each frame of the video, and along a direction of, for example, a straight line parallel to the shooting direction of the panoramic camera, within N frames, the similarity between a part in the first frame and another part in the n-th (1 < n ≤ N) frame is calculated for multiple times respectively, and finally the frame with the largest similarity difference is selected as the key frame in the interval of the N frames. In addition, by the same token, the key frames in the next N-frame interval can be continuously calculated until all projected video frames are traversed.
[0057] After acquiring multiple projected video frames corresponding to various projection methods in the panoramic video, or selecting multiple keyframes to be stitched from the projected video frames, optionally, adjacent first / second projected video frames or first / second keyframes can be stitched together using feature point detection and matching to obtain a two-dimensional view of the captured object. Specifically, the overlapping area between two adjacent frames along the direction of panoramic camera movement can be calculated first, and the offset between the two frames in the movement direction can be calculated from the overlapping area. Then, image fusion can be performed on the overlapping area of the two frames based on the pyramid multi-resolution method to generate a stitched image. In this example, feature detection and matching can be performed on the two images first when calculating the overlapping area. To improve the accuracy of feature matching, the selection of feature points can be specifically processed according to the characteristics of different projection transformations.
[0058] In one example, when the first projection transformation is a perspective transformation and the second projection transformation is an isometric cylindrical projection transformation: For the first projection transformation, considering that the first undistorted region is the middle region of the image, feature points located in the middle region have higher reliability. The relevant values (such as the mean, median, etc.) of feature matching in the middle region can be used to calculate the offset between adjacent frames, or the weight of the feature matching results in the middle region can be increased. Furthermore, for the second projection transformation, considering that the second undistorted region is the regions at both ends of the image, feature points located at both ends have higher reliability. The relevant values of feature matching in the regions at both ends can be used to calculate the offset between adjacent frames, or the weight of the feature matching results in the regions at both ends can be increased.
[0059] According to the above-described image generation method of the present invention, different projection transformation methods are applied to the captured panoramic video with different distortion regions, so that the stitched two-dimensional view of the captured object can reflect the true image of the distortion-free region, thereby displaying a clear and accurate two-dimensional view.
[0060] Figure 2 A flowchart of an image processing method 200 according to an embodiment of the present invention is shown. Referring below... Figure 2 An image processing method according to an embodiment of the present invention is described.
[0061] In step S201, a first two-dimensional view and a second two-dimensional view of the subject are obtained. The first two-dimensional view and the second two-dimensional view are generated by performing a first projection transformation and a second projection transformation on multiple video frames in the panoramic video captured by the subject, respectively. The first distortion region in the first projection video frame is different from the second distortion region in the second projection video frame. The first two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames and stitching together the first non-distorted regions in the multiple first projection video frames. The second two-dimensional view is generated by performing a second projection transformation on the multiple video frames to generate multiple second projection video frames and stitching together the second non-distorted regions in the multiple second projection video frames.
[0062] In this embodiment of the invention, optionally, the first and second two-dimensional views of the subject can be obtained by projecting and transforming a panoramic video of the subject. Specifically, the panoramic video of the subject can be obtained by using a panoramic camera that moves at a specific speed and traverses one or more surfaces of the subject. Optionally, the panoramic camera can move at a preset, uniform speed during shooting. Furthermore, optionally, the panoramic camera can shoot along a straight line. Of course, considering the requirements of the shooting angle and the size of the subject, the panoramic camera can also shoot sequentially in multiple directions, and the resulting videos can be segmented and stitched together according to the shooting route to finally obtain a complete two-dimensional view of the subject.
[0063] Optionally, when the subject is difficult to photograph directly, a vehicle-mounted panoramic camera or a robotic arm equipped with a panoramic camera can be used to capture the subject. In one example, when using a vehicle-mounted panoramic camera for panoramic video recording, the vehicle's route and speed can be preset to traverse the subject and obtain panoramic video. The above-described methods and settings for panoramic video recording of the subject are merely examples. In practical applications, any video recording method equipped with a panoramic camera can be used according to specific needs, and no restrictions are imposed here.
[0064] In this embodiment of the invention, the acquired panoramic video of the subject can be divided into multiple video frames, and at least two projection transformations can be performed on each of the acquired video frames. Specifically, at least a first projection transformation and a second projection transformation can be performed on the multiple video frames in the panoramic video, and multiple first projection video frames and multiple second projection video frames can be generated accordingly. Of course, in another embodiment of the invention, the projection transformation performed on the video frames of the panoramic video can also be more than two projection transformations, for example, it can include a third projection transformation, and a third two-dimensional view can be generated accordingly, for subsequent selection of partial views of the subject depending on the specific shooting scene, which is not limited here.
[0065] In this embodiment of the invention, in order to obtain accurate two-dimensional views without distortion for different parts of the subject being photographed, the first distorted region in the first projected video frame needs to be different from the second distorted region in the second projected video frame. In this embodiment, the first projection transformation and the second projection transformation can be one of various projection transformation methods, such as perspective transformation, equidistant cylindrical projection transformation, cylindrical projection transformation, Mercator transformation, etc. Depending on the projection transformation method, the corresponding distorted and undistorted regions are different. Therefore, the first projection transformation and the second projection transformation can be selected from these projection transformation methods that have different distorted and undistorted regions. For example, in one example of this embodiment, the first projection transformation is a perspective transformation, and the first undistorted region is the middle region of the first projected video frame; the second projection transformation is an equidistant cylindrical projection transformation, and the second undistorted region is the two ends of the second projected video frame; or the first projection transformation is an equidistant cylindrical projection transformation, and the first undistorted region is the two ends of the first projected video frame; the second projection transformation is a perspective transformation, and the second undistorted region is the middle region of the second projected video frame. In one example, the distortion region of the third projection transformation may be the same as or different from the distortion region of the first / second projection transformation. During the subsequent selection of the two-dimensional view, the selection or combination can be made depending on the specific application scenario and image projection effect. Furthermore, in one example of this embodiment, the middle region of each projected video frame can be the region in the middle part along the direction of motion of the panoramic camera in that projected video frame, while the two end regions of each projected video frame can be the regions in the two end parts along the direction of motion of the panoramic camera in that projected video frame.
[0066] In one example of this invention, the region configuration between the undistorted and distorted regions in each projection transformation method can be preset, for example, it can be specified by the user. In another example of this invention, the region configuration between the undistorted and distorted regions in each projection transformation method can also be set by region-related parameters. Optionally, the region-related parameters may include one or more parameters among the type, size, product model, shooting position, and shooting time of the object being photographed. In one example, when the object being photographed is a vehicle chassis, the corresponding region-related parameters may include the vehicle model, the specific size of the vehicle chassis, the shooting position of the panoramic camera used for shooting, the shooting speed during movement, etc. For example, for a certain vehicle model, the first undistorted region corresponding to the projection can be the middle 25%-75% of the video frame in the panoramic video, and the remaining area is the first distorted region. As another example, for another vehicle model, the first undistorted region corresponding to the projection can be the middle 35%-70% of the video frame in the panoramic video, and the remaining area is the first distorted region. The above-described configuration methods for different regions in various projection transformations are merely examples. In practical applications, any corresponding region configuration can be selected according to specific needs, and no restrictions are imposed here. Furthermore, in this embodiment of the invention, the definitions of the various non-distorted regions and distorted regions are only relative. That is, each non-distorted region indicates a region with a relatively small degree of distortion after projection relative to other regions, and is not limited to regions without distortion, while each distorted region indicates a region with a relatively large degree of distortion after projection relative to other regions. In one example of this embodiment, when the ranges of the non-distorted and distorted regions are specified according to user input, they refer to the specific regions defined by the user. In one example of this embodiment, the distorted and distorted regions can be continuous and non-overlapping; however, in another example, the distorted and distorted regions can be discontinuous or overlapping, and no restrictions are imposed here.
[0067] After acquiring each video frame from the panoramic video, optionally, all video frames or a portion of them can be selected for subsequent stitching of the two-dimensional view. According to an embodiment of the present invention, optionally, all first projected video frames obtained by the panoramic video through a first projection transformation can be stitched together; furthermore, optionally, considering that there may be a large amount of overlapping areas between the first projected video frames leading to redundant calculations, multiple first keyframes can also be selected from the plurality of first projected video frames for stitching.
[0068] Of course, similarly, the same processing can also be adopted for the second projection transformation, that is, all the second projection video frames obtained by performing the second projection transformation on the panoramic video can be stitched, or a plurality of second key frames can be selected from the plurality of second projection video frames for stitching.
[0069] In the specific key frame extraction process, optionally, one key frame can be extracted from every N projection video frames, for example, N may be a fixed value. In addition, optionally, key frames can also be extracted by using a corresponding image processing algorithm according to the shooting parameters when the panoramic camera shoots. In an example, the shooting parameters may be various parameters such as the shooting speed of the panoramic video, the type, size and product model of the shooting object. After obtaining the shooting parameters, optionally, various image processing algorithms such as a similarity-based method and a clustering method can be used to extract key frames.
[0070] In an example of the embodiment of the present invention, when key frame extraction is performed by a similarity calculation method, optionally, a plurality of parts at different positions can be selected in each frame of the video, and along a direction parallel to the shooting direction of the panoramic camera, for example, within N frames, the similarity between a part in the first frame and another part in the n-th frame (1<n≤N) can be calculated respectively for multiple times, and finally the frame with the largest similarity difference is selected as the key frame in the interval of these N frames. In addition, by parity of reasoning, the key frames in the next N-frame interval can be continuously calculated until all projection video frames are traversed.
[0071] After obtaining a plurality of projection video frames corresponding to each projection mode in the panoramic video, or selecting a plurality of key frames to be stitched from the projection video frames, optionally, feature point detection and matching is used to successively stitch each adjacent first / second projection video frame or first / second key frame, so as to obtain a two-dimensional view of the shooting object obtained by stitching. Specifically, the overlapping area between two adjacent frames along the moving direction of the panoramic camera can be calculated first, so that the offset of the two frames in the moving direction can be obtained through calculation of the overlapping area. Then, image fusion can be performed on the overlapping area of the two frames based on a pyramid multi-resolution method to generate a stitched image. In this example, when calculating the overlapping area, feature detection and matching can be performed on the two images first. In order to improve the accuracy of feature matching, targeted processing can be performed on the screening of feature points according to the characteristics of different projection transformations.
[0072] In one example, when the first projection transformation is a perspective transformation and the second projection transformation is an isometric cylindrical projection transformation: For the first projection transformation, considering that the first undistorted region is the middle region of the image, feature points located in the middle region have higher reliability. The relevant values (such as the mean, median, etc.) of feature matching in the middle region can be used to calculate the offset between adjacent frames, or the weight of the feature matching results in the middle region can be increased. Furthermore, for the second projection transformation, considering that the second undistorted region is the regions at both ends of the image, feature points located at both ends have higher reliability. The relevant values of feature matching in the regions at both ends can be used to calculate the offset between adjacent frames, or the weight of the feature matching results in the regions at both ends can be increased.
[0073] In step S202, the position of the first region in the shooting object is obtained, and it is determined whether the first region is located in the first non-distorted region in the first projected video frame or in the second non-distorted region in the second projected video frame.
[0074] In this embodiment of the invention, optionally, after obtaining the first two-dimensional view and the second two-dimensional view of the entire subject, a distortion-free two-dimensional view of a local area of the subject can be further obtained. Considering that the distorted areas of the first and second two-dimensional views are different, the corresponding distorted and undistorted areas of each two-dimensional view obtained after stitching are also different. Therefore, it is necessary to determine which undistorted area of the two-dimensional view the local area of the subject is located in, so as to minimize distortion in the corresponding two-dimensional view portion.
[0075] Therefore, in this embodiment of the invention, the location of the first region can be obtained first and then associated with a non-distorted region of a certain projected video frame. Optionally, the first region can be specified by the user or obtained through various information (such as coordinate information). The first region can be the corresponding area of a defined graphic in the subject being photographed, or it can be a component or part of the subject being photographed, etc., and there are no restrictions here.
[0076] After obtaining the location of the first region, it can be determined which non-distorted region of the first region lies within in the projected video frame based on its specific location information (such as coordinates, region range, region size, etc.). Optionally, when the first region is entirely located within the non-distorted region of a single projected video frame, this determination can be made directly. Furthermore, optionally, when the first region is large or spans the non-distorted regions of two different projected video frames, the non-distorted region of the main corresponding projected video frame can be selected, or the non-distorted region of the projected video frame corresponding to a key part of the first region can be selected. Alternatively, optionally, where feasible, the corresponding parts of the first non-distorted region of the first projected video frame and the corresponding parts of the second non-distorted region of the second projected video frame can be fused and stitched together subsequently.
[0077] In step S203, if the first region is located in the first non-distorted region of the first projected video frame, the first region is associated with the corresponding position in the first two-dimensional view; if the first region is located in the second non-distorted region of the second projected video frame, the first region is associated with the corresponding position in the second two-dimensional view.
[0078] In this embodiment of the invention, after determining whether the first region is located in the first undistorted region of the first projected video frame or the second undistorted region of the second projected video frame as described above, the first region can be associated with the first undistorted region of the first projected video frame and / or the second undistorted region of the second projected video frame, so as to provide relevant image information or display an undistorted two-dimensional view.
[0079] Optionally, embodiments of the present invention may further include: displaying a two-dimensional view of a corresponding position in the first two-dimensional view or a two-dimensional view of a corresponding position in the second two-dimensional view associated with the first region, based on the obtained indication regarding the first region.
[0080] According to the image processing method of the present invention, a correspondence can be established between a specified first region and a local two-dimensional view of a non-distorted region in a corresponding projection mode, and the first region and the corresponding non-distorted two-dimensional view can be displayed accordingly, which greatly improves the user experience.
[0081] The following illustrates the specific implementation steps of an image generation method and an image processing method according to an embodiment of the present invention.
[0082] In an example of an embodiment of the present invention, the object photographed by the panoramic camera can be a vehicle chassis. Since the vehicle chassis is difficult to approach or cannot be directly photographed, the image generation method illustrated in the embodiment of the present invention can be used to generate a two-dimensional view, and the image processing method illustrated in the embodiment of the present invention can be used to locate local areas or components and to associate and display the two-dimensional view.
[0083] In this example, a panoramic video of the subject is captured using a moving panoramic camera.
[0084] In an example of an embodiment of the present invention, a panoramic camera can be mounted on a vehicle that can enter the vehicle chassis and move at a specific speed to traverse the surface of the vehicle chassis to obtain a panoramic video of the vehicle chassis. Optionally, the panoramic camera can move at a preset constant speed during shooting. Furthermore, optionally, the panoramic camera can shoot approximately along the centerline of the vehicle chassis to ultimately obtain a two-dimensional view of the entire subject.
[0085] Subsequently, a first projection transformation is performed on multiple video frames in the panoramic video to generate multiple first projected video frames, and a second projection transformation is performed on multiple video frames in the panoramic video to generate multiple second projected video frames, wherein the first distortion region in the first projected video frame that produces distortion is different from the second distortion region in the second projected video frame.
[0086] In this embodiment of the invention, the panoramic video of the captured object can be split into multiple video frames, and at least two projection transformations can be performed on the acquired video frames respectively. Figure 3 A schematic diagram of video frames from a panoramic video, representing an example according to an embodiment of the present invention, is shown. Figure 3 As shown, the video frames of the panoramic video can each include a part of the vehicle chassis, and can include different parts of the vehicle chassis as the panoramic camera moves.
[0087] After acquiring the video frames of the panoramic video, at least a first projection transformation and a second projection transformation can be performed on the multiple video frames in the panoramic video, respectively, and multiple first projection video frames and multiple second projection video frames can be generated accordingly. Of course, in another embodiment of the present invention, the projection transformation performed on the video frames of the panoramic video can also be more than two projection transformations, for example, it can include a third projection transformation, which is not limited here. In one example, the distortion region of the third projection transformation can be the same as or different from the distortion region of the first projection transformation / second projection transformation. In the subsequent selection of the two-dimensional view, it can be selected or combined according to the specific application scenario and image projection effect.
[0088] In this embodiment of the invention, in order to obtain accurate two-dimensional views without distortion for different parts of the subject being photographed, the first distorted region in the first projected video frame needs to be different from the second distorted region in the second projected video frame. In this embodiment, the first projection transformation and the second projection transformation can be one of a variety of projection transformation methods. Depending on the projection transformation method, the corresponding distorted and undistorted regions are different. Therefore, the first projection transformation and the second projection transformation can be selected from these projection transformation methods that have different distorted and undistorted regions. For example, in one example of this embodiment, the first projection transformation can be a perspective transformation, and the first undistorted region is the middle region of the first projected video frame; the second projection transformation can be an equidistant cylindrical projection transformation, and the second undistorted region is the two end regions of the second projected video frame.
[0089] In one example of this invention embodiment, the region configuration between the undistorted and distorted regions in each projection transformation method can be preset, for example, it can be preset by the user. In another example of this invention embodiment, the region configuration between the undistorted and distorted regions in each projection transformation method can also be set by region-related parameters. Optionally, region-related parameters may include one or more parameters among the type, size, product model, shooting location, and shooting time of the object being photographed. Figure 4 A schematic diagram of a user interface for panoramic video uploading and region-related parameter configuration, according to an embodiment of the present invention, is shown. Figure 4 As shown, panoramic videos can be uploaded using this user interface. Furthermore, depending on user-specified or region-related parameter settings, the range of distortion and non-distortion areas that will be limited during subsequent projection transformation of the panoramic video can be adjusted. For example, this can be achieved using... Figure 4 The percentage of space occupied by the top, middle, and bottom sections of the image shown can be specified and adjusted. Of course, in another example, other methods such as pixel count can also be used to specify and adjust this.
[0090] Figure 5 A schematic diagram illustrating an example of a region configuration according to an embodiment of the present invention is shown. Figure 5 As shown, the specific configuration range of the upper feature point region, the middle feature point region, and the lower feature point region can be shown by the proportion of the regions. Figure 5 The regional configuration information shown is for illustrative purposes only. In actual applications, any regional configuration method that matches the application scenario can be adopted.
[0091] In an example of an embodiment of the present invention, when the object to be photographed is a vehicle chassis, the corresponding area-related parameters may include the vehicle model, the specific dimensions of the vehicle chassis, the shooting position of the panoramic camera used for shooting, the shooting speed during movement, and other parameters.
[0092] Optionally, the regional configuration information in the examples of this embodiment may be related to the vehicle model. In this example, various regional configuration information related to the vehicle model can be pre-set for users to select or for the system to automatically retrieve. Figure 6 A schematic diagram illustrating the area configuration related to the model of the vehicle being photographed, as an example of an embodiment of the present invention, is shown. Figure 6 The region configuration information can include multiple tables showing region configuration information corresponding to different camera model types. For example... Figure 6 As shown, the configuration range of the feature point area corresponding to the subject model M01 can be as follows: the upper feature point area is 80% to 100% of the vertical length; the middle feature point area is 25% to 80% of the vertical length; and the lower feature point area is 0% to 25% of the vertical length. Furthermore, the configuration range of the feature point area corresponding to the subject model N02 can be as follows: the upper feature point area is 70% to 100% of the vertical length; the middle feature point area is 20% to 70% of the vertical length; and the lower feature point area is 0% to 20% of the vertical length.
[0093] Furthermore, optionally, the area configuration information in the examples of this embodiment can be further related to the video generation time. In this example, various area configuration information related to parameters such as vehicle model and generation time can also be pre-set for users to select or for the system to automatically extract. Figure 7 A schematic diagram illustrating the area configuration information related to the model of the vehicle being filmed and the video generation time, as an example of an embodiment of the present invention, is shown. Figure 7 As shown, different subject IDs (0001, 0002, 0003) can correspond to different subject models (M01, N02, and M01), and can have different video generation times. For example, the configuration range of the feature point region corresponding to subject ID 0001 and subject model M01 can be as follows: the upper feature point region is 80%–100% of the vertical length; the middle feature point region is 25%–80% of the vertical length; and the lower feature point region is 0%–25% of the vertical length. The generation time of the video captured in this case can be 2024.02.15. For this example, when… Figure 7The table shown illustrates the correspondence between the subject ID, subject model, and generation time. During the process of determining the region configuration information, the user can select the appropriate region configuration information for subsequent operations. For example, the user can select the corresponding region configuration information for the subject model. Furthermore, among multiple data points for the same subject model, the user can select the region configuration information with the closest generation time to obtain the latest region configuration information.
[0094] Figure 8 A schematic diagram of a user interface for panoramic video uploading and region-related parameter configuration, according to an embodiment of the present invention, is shown. Figure 8 As shown, panoramic video can be uploaded using this user interface. Furthermore, depending on the selected subject model or user specification, region-related parameters can be configured. For example, after the user selects from the subject models M01, N01, and N02, the configuration can be based on, for example… Figure 7 The diagram shows the correspondence between the model of the photographed object and the region configuration parameters, from which the specific limitations of the corresponding region configuration information can be retrieved. For example, when in... Figure 8 When selecting model N02 as the subject in the user interface shown, you can... Figure 7 The table settings area configuration parameters are displayed in the table. Furthermore, when in... Figure 8 When selecting the subject model M01 in the user interface shown, you can... Figure 7 The table allows you to select one of several regional configuration parameters from a list of available parameters. For example, you can set the region configuration parameter with the most recent generation time as the selected region configuration parameter.
[0095] Subsequently, a first two-dimensional view of the photographed object can be generated by stitching together the first non-distorted regions in the plurality of first projected video frames, and a second two-dimensional view of the photographed object can be generated by stitching together the second non-distorted regions in the plurality of second projected video frames.
[0096] After acquiring each video frame from the panoramic video, optionally, all or a subset of the video frames can be selected for subsequent stitching of the 2D view. In this example, considering the potential for significant regional overlap between the first and second projected video frames, leading to redundant calculations, multiple first and second keyframes can be selected from the plurality of first and second projected video frames for stitching. This ensures the correctness and integrity of the stitched image while significantly improving the efficiency of the entire processing.
[0097] In a specific key frame extraction process, optionally, one key frame may be extracted from every N projected video frames. Optionally, the key frame may be extracted by using a corresponding image processing algorithm according to the shooting parameters of the panoramic camera during shooting. For example, the shooting parameters may be various parameters such as the shooting speed of the panoramic video, the type, size and product model of the shooting object. After obtaining the shooting parameters, optionally, in this example, a similarity-based method may be used to extract key frames.
[0098] When the similarity calculation method is used for key frame extraction, optionally, for example, 3 small parts may be selected in each frame of the video. Considering the moving direction of the panoramic camera, these three parts may be respectively located at the front, middle and rear parts of the panoramic video frame along the moving process of the panoramic camera. Assume that the size of each part is w×h, where w is the empirical width of the set part and h is the empirical height of the set part. In the stitching process, along the moving direction of the panoramic camera, for example, with the first frame as the starting point, within N frames, the similarity between the middle part of the first frame and the front part of the nth frame (1 < n ≤ N), as well as the similarity between the rear part of the first frame and the middle part of the nth frame are calculated. Finally, the frame with the largest similarity difference may be selected as the key frame in the N-frame interval. Subsequently, with this key frame as the starting point, key frames may be continuously extracted in the next N-frame interval until the panoramic video ends. Figure 9 shows a schematic diagram of two consecutive adjacent key frames extracted in an example according to an embodiment of the present invention. As Figure 9 shows, the extracted key frames respectively shown in the left and right figures can be located at different positions on the vehicle chassis and may include different components for later stitching of the two-dimensional view.
[0099] After a plurality of key frames to be stitched are selected from the projected video frames, optionally, the adjacent first and second key frames are sequentially stitched by means of feature point detection and matching to obtain a two-dimensional view of the stitched shooting object. Specifically, the overlapping area between two adjacent frames along the moving direction of the panoramic camera may be calculated first, so that the offset of the two frames along the moving direction can be obtained through calculation based on the overlapping area. Figure 10 shows a schematic diagram of stitching adjacent key frames in an example according to an embodiment of the present invention. As Figure 10 shows, the overlapping area between the non-distorted areas of two key frames along the moving direction of the panoramic camera may be extracted as the feature point matching area, so that the offset of the two frames along the moving direction of the panoramic camera can be calculated through the feature point matching area.
[0100] Figure 11 shows a specific example of steps for stitching each projected video frame to obtain a corresponding stitched two-dimensional view in an example according to an embodiment of the present invention. As Figure 11As shown, at the beginning of the steps, in S11, all video frames of the panoramic images to be stitched can be converted into perspective projection video frames through perspective projection transformation. Subsequently, in S12, the region configuration information of the corresponding feature points can be obtained, and in S13, the frame number n = 1 is first set. When n = 1, in S14, perspective projection video frames A1 and B1, converted from the nth and (n+1th)th panoramic video frames to be stitched, can be obtained respectively. Then, in S15, A1 and B1 can be stitched together according to the feature points in the middle feature point region of A1 and B1 in a direction parallel to the shooting motion direction, obtaining the stitched two-dimensional view C1 of the perspective video frames. After this, in S16, the stitched two-dimensional view C1 can be used as the perspective two-dimensional view A1 converted from the (n+1th)th panoramic video frame. After stitching, in S17, it can be determined whether there are any unstitched perspective video frames, and after the determination, it can be decided whether to set n to n = n+1 and return to the previous stitching steps, or directly proceed to step S3.
[0101] Similarly, in S21, all the video frames of the panoramic images to be stitched can be converted into equidistant cylindrical projection video frames through equidistant cylindrical projection transformation. Then, in S22, the corresponding feature point region configuration information can be obtained, and in S23, the frame number n = 1 is first set. When n = 1, in S24, the equidistant cylindrical projection video frames A2 and B2, converted from the nth and (n+1th)th panoramic video frames to be stitched, can be obtained respectively. Then, in S25, A2 and B2 can be stitched together according to the feature points in the upper and lower feature point regions of A2 and B2 in a direction parallel to the shooting motion direction, obtaining the stitched two-dimensional view C2 of the equidistant cylindrical video frames. After this, in S26, the stitched equidistant cylindrical two-dimensional view C2 can be used as the equidistant cylindrical two-dimensional view A2 converted from the (n+1th)th panoramic video frame. After stitching, step S27 can be used to determine if there are any unstitched equidistant columnar video frames. Following this determination, it can be decided whether to set n to n = n + 1 and return to the previous stitching step, or proceed directly to step S3. After proceeding to step S3, the perspective 2D view A1 and the equidistant columnar 2D view A2 can be saved as the final results of perspective stitching and equidistant columnar stitching, respectively, and associated with the subject recognition information. The method then concludes.
[0102] Figure 12 Specifically, a schematic diagram is shown in one example of an embodiment of the present invention, illustrating the stitching of feature points between perspective video frame A1 and perspective video frame B1. For example... Figure 12 As shown, since the non-distorted region in the perspective projection video frame is the middle region of the video frame (i.e., Figure 12 The middle feature point region in the video frame), while the distortion region is the region at both ends of the video frame (i.e., the middle feature point region). Figure 12The upper and lower feature point regions in A1 and B1 are used to stitch A1 and B1 together in a direction parallel to the shooting motion, thus obtaining a stitched two-dimensional view C1 of the perspective video frames. Furthermore, Figure 13 This diagram illustrates, in one example of an embodiment of the present invention, the stitching of feature points between equidistant columnar video frames A2 and B2. Figure 13 As shown, since the non-distorted region in the equidistant cylindrical projection video frame is the region at both ends of the video frame (i.e. Figure 13 The upper and lower feature point regions in the video frame, while the distortion region is the middle region of the video frame (i.e., the upper and lower feature point regions). Figure 13 The upper and lower feature point regions in A2 and B2 can be used to stitch A2 and B2 together in a direction parallel to the shooting motion direction to obtain a stitched two-dimensional view C2 of equidistant columnar video frames. Of course, the above stitching process based on the non-distorted region is only an example. In practical applications, the confidence level of the non-distorted region can be set to be greater than that of the distorted region, so that the stitching between video frames can be completed by comprehensively considering and calculating the non-distorted region with higher confidence and the distorted region with lower confidence.
[0103] Figure 14 An example of the image generation result of an embodiment of the present invention is shown. In step S3 above, the perspective 2D view A1 and the equidistant cylindrical 2D view A2 can be saved as the final results of perspective stitching and equidistant cylindrical stitching, respectively (e.g., ...). Figure 14 The perspective view files and isometric bar chart files listed in the table) are associated with subject identification information (such as subject ID). Additionally, one can be selected from the perspective 2D view A1 and the isometric bar chart 2D view A2 as the overview file in the processing results and saved (e.g., ...). Figure 14 The table contains overview images corresponding to different subject IDs, allowing users to observe and select data.
[0104] An example of an embodiment of the present invention further illustrates an image processing method for the various two-dimensional views generated above.
[0105] First, a first two-dimensional view and a second two-dimensional view of the subject can be obtained. In this embodiment of the invention, the first two-dimensional view can be the aforementioned perspective two-dimensional view A1, and the second two-dimensional view can be an equidistant cylindrical two-dimensional view A2.
[0106] Subsequently, the position of the first region in the photographed object can be obtained, and it can be determined whether the first region is located in the first non-distorted region in the first projected video frame or in the second non-distorted region in the second projected video frame.
[0107] In this embodiment of the invention, optionally, after obtaining the overall perspective two-dimensional view A1 and the equidistant cylindrical two-dimensional view A2 of the photographed object, a distortion-free two-dimensional view of a local area of the photographed object can be further obtained. Considering that the distorted areas of the first and second two-dimensional views are different, the corresponding distorted and undistorted areas of each two-dimensional view obtained after stitching are also different. Therefore, it is necessary to determine which undistorted area of the two-dimensional view the local area of the photographed object is located in, so as to minimize distortion in the corresponding two-dimensional view portion.
[0108] Therefore, in this embodiment of the invention, the location of a first region can be obtained first and then associated with a non-distorted region of a certain projected video frame. Optionally, the first region can be specified by the user or obtained through various information (such as coordinate information). For example, the first region can be specified by the user in the overview view. Figure 15 A user interface diagram showing an overview of an example according to an embodiment of the present invention is provided. Figure 15 As shown, the middle and both ends can be associated with the undistorted regions of different projection methods. For example, Figure 15 The middle section of the overview diagram can be associated with the perspective 2D view corresponding to the perspective projection, while the two ends can be associated with the equidistant cylindrical 2D view corresponding to the equidistant cylindrical projection. Figure 15 During the specific operation of the user interface shown, it can be based on Figure 14 The generated results are as follows: when the user selects "view details of the middle part", the perspective view file corresponding to the photographed object is obtained and displayed based on this perspective view file; when the user selects "view details of both ends", the isometric bar chart file corresponding to the photographed object is obtained and displayed based on this isometric bar chart file.
[0109] In an example of an embodiment of the present invention, the location of the first area may be specified by the user through the user interface of the overview map. Figure 16 A schematic diagram of first region information is shown as an example of an embodiment of the present invention. For example... Figure 16 As shown, the first region can be represented using local coordinates in the overview view. By dividing the overview view of the vehicle base into multiple local regions of X×Y, it is possible to... Figure 16 The first shaded region is represented as a 3-4 region with a horizontal x-coordinate of 3 and a vertical y-coordinate of 4 (i.e., local image ID 3-4), and is represented using corresponding coordinates or the proportion of the image in the horizontal and vertical directions. Subsequently, each local region can be further defined, for example... Figure 16 The coordinates are associated and represented in the local region definition in the lower right corner. For example... Figure 16As shown, since the overview image is horizontally divided into 10 equal parts, the horizontal length of each part is 10% of the total horizontal pixels. Similarly, since the overview image is vertically divided into 8 equal parts, the vertical length of each part is 12.5% of the total vertical pixels. According to... Figure 16 The horizontal and vertical lengths of the 10×8 local regions in the overall map can be used to calculate the coordinates of each local region, and can be used as follows: Figure 16 As shown in the lower right corner, it is associated with the local image ID of that local area and the subject ID.
[0110] According to the image generation method in the above example of the embodiment of the present invention, after obtaining Figure 14 Based on the generated results shown, it is necessary to further generate local image files. Figure 17 A flowchart illustrating a partial image file generation process according to an embodiment of the present invention is shown. Figure 17 As shown, in step S41, for example, Figure 5 The display area configuration information is shown, and in step S42, according to, for example... Figure 16 The region definitions and coordinates of each local image shown are used to determine the region type of the local image, such as whether it is a middle display region or an upper / lower display region. After determining the type of the local image, the file source for obtaining the local image can be determined in either step S43 or step S44. That is, based on the coordinates and region configuration of each local region, it can be determined whether to use a perspective view file or an equidistant histogram file to generate the local image file corresponding to that local region. For example, in step S44, an equidistant histogram file can be used to generate a local image file for local regions belonging to the upper and lower display regions; while in step S43, a perspective view file can be used to generate a local image file for local regions belonging to the middle display region. Finally, in step S45, the generated local image file can be associated with the local image ID and saved in relation to the captured image ID, and appended to [the appropriate database / system]. Figure 14 The generated results are shown and updated to obtain... Figure 18 The new generation result is shown.
[0111] Figure 18 A table example illustrating the association of a partial image of a first region with different projected image files, according to an embodiment of the present invention, is shown. Figure 18 As shown, based on the correspondence between the first region and the non-distorted regions of different projection methods, the local image files can be associated with the corresponding perspective view files and isometric bar chart files respectively, for subsequent image acquisition and display operations.
[0112] In an example of this embodiment of the invention, a two-dimensional view of the corresponding position in the first two-dimensional view or the corresponding position in the second two-dimensional view associated with the first region can also be displayed based on the obtained indication regarding the first region. Specifically, after obtaining the coordinates of the first region selected by the user, the association between the coordinates of the first region selected by the user and the local image can be obtained, and the undistorted image corresponding to the associated first region can be displayed. Figure 19 An example of a user interface displaying a partial detail image is shown, according to an embodiment of the present invention. Figure 19 As shown, a related local detail image in the undistorted region of the corresponding projection method can be obtained based on a local image of the first region selected in the overview view. Specifically, the coordinates specified by the user in the overview view can be obtained first, and then... Figure 16 The local region definition shown obtains the local image ID of the first region corresponding to the coordinates, and from... Figure 18 The local image file corresponding to the local image ID is obtained from the generated result shown. Figure 19 It is displayed as a partial detail image.
[0113] Below, refer to Figure 20 The image generation apparatus according to embodiments of the present invention will be described. Figure 20 A block diagram of an image generation apparatus 2000 according to an embodiment of the present invention is shown. Figure 20 As shown, the image generation apparatus 2000 includes a capturing unit 2010, a generation unit 2020, and a stitching unit 2030. Besides these units, the image generation apparatus 2000 may also include other components; however, since these components are not relevant to the content of this embodiment, their illustrations and descriptions are omitted here. Furthermore, since the specific details of the following operations performed by the image generation apparatus 2000 according to this embodiment are the same as those referred to above... Figure 1 The details described are the same, so repeated descriptions of the same details are omitted here to avoid repetition.
[0114] Figure 20 The imaging unit 2010 of the image generation device 2000 in the image generation device 2000 acquires panoramic video of the subject captured by a moving panoramic camera.
[0115] In this embodiment of the invention, a panoramic camera can be used to move at a specific speed and traverse one or more surfaces of the subject to obtain a panoramic video of the subject. Optionally, the panoramic camera can move at a preset, uniform speed during shooting. Furthermore, the panoramic camera can optionally shoot along a straight line. Of course, considering the requirements of the shooting angle and the size of the subject, the panoramic camera can also shoot sequentially in multiple directions, and the resulting videos can be segmented and stitched together according to the shooting route to ultimately obtain a complete two-dimensional view of the subject.
[0116] Optionally, when the subject is difficult to photograph directly, a vehicle-mounted panoramic camera or a robotic arm equipped with a panoramic camera can be used to capture the subject. In one example, when using a vehicle-mounted panoramic camera for panoramic video recording, the vehicle's route and speed can be preset to traverse the subject and obtain panoramic video. The above-described methods and settings for panoramic video recording of the subject are merely examples. In practical applications, any video recording method equipped with a panoramic camera can be used according to specific needs, and no restrictions are imposed here.
[0117] The generation unit 2020 performs a first projection transformation on multiple video frames in the panoramic video to generate multiple first projection video frames, and performs a second projection transformation on multiple video frames in the panoramic video to generate multiple second projection video frames, wherein the first distortion region in the first projection video frame that produces distortion is different from the second distortion region in the second projection video frame.
[0118] In this embodiment of the invention, the acquired panoramic video of the subject can be divided into multiple video frames, and at least two projection transformations can be performed on each of the acquired video frames. Specifically, at least a first projection transformation and a second projection transformation can be performed on the multiple video frames in the panoramic video, and multiple first projection video frames and multiple second projection video frames can be generated accordingly. Of course, in another embodiment of the invention, the projection transformation performed on the video frames of the panoramic video can also be more than two projection transformations, for example, it can include a third projection transformation, which is not limited here. In one example, the distortion region of the third projection transformation can be the same as or different from the distortion region of the first projection transformation / second projection transformation. In the subsequent selection of the two-dimensional view, it can be selected or combined according to the specific application scenario and image projection effect.
[0119] In this embodiment of the invention, in order to obtain accurate two-dimensional views without distortion for different parts of the subject being photographed, the first distorted region in the first projected video frame needs to be different from the second distorted region in the second projected video frame. In this embodiment, the first projection transformation and the second projection transformation can be one of various projection transformation methods, such as perspective transformation, equidistant cylindrical projection transformation, cylindrical projection transformation, Mercator transformation, etc. Depending on the projection transformation method, the corresponding distorted and undistorted regions are different. Therefore, the first projection transformation and the second projection transformation can be selected from these projection transformation methods that have different distorted and undistorted regions. For example, in one example of this embodiment, the first projection transformation is a perspective transformation, and the first undistorted region is the middle region of the first projected video frame; the second projection transformation is an equidistant cylindrical projection transformation, and the second undistorted region is the two ends of the second projected video frame; or the first projection transformation is an equidistant cylindrical projection transformation, and the first undistorted region is the two ends of the first projected video frame; the second projection transformation is a perspective transformation, and the second undistorted region is the middle region of the second projected video frame. In one example, the distortion region of the third projection transformation may be the same as or different from the distortion region of the first projection transformation / second projection transformation. In the subsequent selection of the two-dimensional view, the selection or combination can be made depending on the specific application scenario and image projection effect.
[0120] In one example of this invention, the region configuration between the undistorted and distorted regions in each projection transformation method can be preset, for example, it can be specified by the user. In another example of this invention, the region configuration between the undistorted and distorted regions in each projection transformation method can also be set by region-related parameters. Optionally, the region-related parameters may include one or more parameters among the type, size, product model, shooting position, and shooting time of the object being photographed. In one example, when the object being photographed is a vehicle chassis, the corresponding region-related parameters may include the vehicle model, the specific size of the vehicle chassis, the shooting position of the panoramic camera used for shooting, the shooting speed during movement, etc. For example, for a certain vehicle model, the first undistorted region corresponding to the projection can be the middle 25%-75% of the video frame in the panoramic video, and the remaining area is the first distorted region. As another example, for another vehicle model, the first undistorted region corresponding to the projection can be the middle 35%-70% of the video frame in the panoramic video, and the remaining area is the first distorted region. The above-described configuration methods for different regions in various projection transformations are merely examples. In practical applications, any corresponding region configuration can be selected according to specific needs, and no restrictions are imposed here. Furthermore, in this embodiment of the invention, the definitions of the various non-distorted regions and distorted regions are only relative. That is, each non-distorted region indicates a region with a relatively small degree of distortion after projection relative to other regions, and is not limited to regions without distortion, while each distorted region indicates a region with a relatively large degree of distortion after projection relative to other regions. In one example of this embodiment, when the ranges of the non-distorted and distorted regions are specified according to user input, they refer to the specific regions defined by the user. In one example of this embodiment, the distorted and distorted regions can be continuous and non-overlapping; however, in another example, the distorted and distorted regions can be discontinuous or overlapping, and no restrictions are imposed here.
[0121] The stitching unit 2030 stitches together the first non-distorted regions in the plurality of first projected video frames to generate a first two-dimensional view of the subject being photographed, and stitches together the second non-distorted regions in the plurality of second projected video frames to generate a second two-dimensional view of the subject being photographed.
[0122] After obtaining each video frame from the panoramic video, optionally, all video frames or some of the video frames can be selected for subsequent splicing of two-dimensional views. According to embodiments of the present invention, optionally, all first projection video frames obtained by performing a first projection transformation on the panoramic video can be spliced; besides, optionally, considering that there may be a large amount of overlapping areas between the first projection video frames which leads to redundant calculation, a plurality of first key frames can also be selected from the plurality of first projection video frames for splicing.
[0123] Of course, similarly, the same processing method can also be adopted for the second projection transformation, that is, all second projection video frames obtained by performing a second projection transformation on the panoramic video can be spliced, and a plurality of second key frames can also be selected from the plurality of second projection video frames for splicing.
[0124] In a specific key frame extraction process, optionally, one frame can be extracted as a key frame from every N projected video frames, for example, N can be a fixed value. In addition, optionally, key frames can also be extracted by using corresponding image processing algorithms according to the shooting parameters of the panoramic camera during shooting. In one example, the shooting parameters may be various parameters such as the shooting speed of the panoramic video, the type, size and product model of the shooting object. After obtaining the shooting parameters, optionally, various image processing algorithms such as a similarity-based method and a clustering method can be used to extract key frames.
[0125] In an example of the embodiment of the present invention, when key frame extraction is performed by a similarity calculation method, optionally, a plurality of parts at different positions can be selected in each frame of the video, and along a direction, for example, a direction of a straight line parallel to the shooting direction of the panoramic camera, within N frames, the similarity between a part in the first frame and another part in the n-th (1 < n ≤ N) frame is calculated for multiple times respectively, and finally the frame with the largest similarity difference is selected as the key frame in the interval of the N frames. In addition, the key frames in the next N-frame interval can be calculated continuously by analogy until all projected video frames are traversed.
[0126] After acquiring multiple projected video frames corresponding to various projection methods in the panoramic video, or selecting multiple keyframes to be stitched from the projected video frames, optionally, adjacent first / second projected video frames or first / second keyframes can be stitched together using feature point detection and matching to obtain a two-dimensional view of the captured object. Specifically, the overlapping area between two adjacent frames along the direction of panoramic camera movement can be calculated first, and the offset between the two frames in the movement direction can be calculated from the overlapping area. Then, image fusion can be performed on the overlapping area of the two frames based on the pyramid multi-resolution method to generate a stitched image. In this example, feature detection and matching can be performed on the two images first when calculating the overlapping area. To improve the accuracy of feature matching, the selection of feature points can be specifically processed according to the characteristics of different projection transformations.
[0127] In one example, when the first projection transformation is a perspective transformation and the second projection transformation is an isometric cylindrical projection transformation: For the first projection transformation, considering that the first undistorted region is the middle region of the image, feature points located in the middle region have higher reliability. The relevant values (such as the mean, median, etc.) of feature matching in the middle region can be used to calculate the offset between adjacent frames, or the weight of the feature matching results in the middle region can be increased. Furthermore, for the second projection transformation, considering that the second undistorted region is the regions at both ends of the image, feature points located at both ends have higher reliability. The relevant values of feature matching in the regions at both ends can be used to calculate the offset between adjacent frames, or the weight of the feature matching results in the regions at both ends can be increased.
[0128] According to the above-described image generation apparatus of the present invention, different projection transformation methods are applied to the captured panoramic video with different distortion regions, so that the two-dimensional view of the captured object after stitching can reflect the true image of the area without distortion, thereby displaying a clear and accurate two-dimensional view.
[0129] Below, refer to Figure 21 The image processing apparatus according to embodiments of the present invention will be described. Figure 21 A block diagram of an image processing apparatus 2100 according to an embodiment of the present invention is shown. Figure 21As shown, the image processing apparatus 2100 includes an acquisition unit 2110, a judgment unit 2120, and an association unit 2130. Furthermore, the image processing apparatus 2100 may also include a user interface control unit 2140. The user interface control unit 2140 can interact with the acquisition unit 2110, the judgment unit 2120, and the association unit 2130 respectively, and the user interface control unit 2140 is used to control the user interface so that the image processing apparatus transmits user interface data to the client and displays it on the client's browser. Besides these units, the image processing apparatus 2100 may also include other components; however, since these components are not relevant to the content of this embodiment, their illustrations and descriptions are omitted here. Furthermore, since the specific details of the following operations performed by the image processing apparatus 2100 according to this embodiment are the same as those referred to above... Figure 2 The details described are the same, so repeated descriptions of the same details are omitted here to avoid repetition.
[0130] Figure 21 The image processing device 2100 in the image processing device 2100 acquires a first two-dimensional view and a second two-dimensional view of the subject being photographed. The first two-dimensional view and the second two-dimensional view are generated by performing a first projection transformation and a second projection transformation on multiple video frames in a panoramic video of the subject being photographed. The first distortion region in the first projection video frame is different from the second distortion region in the second projection video frame. The first two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames and stitching together the first non-distorted regions in the multiple first projection video frames. The second two-dimensional view is generated by performing a second projection transformation on the multiple video frames to generate multiple second projection video frames and stitching together the second non-distorted regions in the multiple second projection video frames.
[0131] In this embodiment of the invention, optionally, the first and second two-dimensional views of the subject can be obtained by projecting and transforming a panoramic video of the subject. Specifically, the panoramic video of the subject can be obtained by using a panoramic camera that moves at a specific speed and traverses one or more surfaces of the subject. Optionally, the panoramic camera can move at a preset, uniform speed during shooting. Furthermore, optionally, the panoramic camera can shoot along a straight line. Of course, considering the requirements of the shooting angle and the size of the subject, the panoramic camera can also shoot sequentially in multiple directions, and the resulting videos can be segmented and stitched together according to the shooting route to finally obtain a complete two-dimensional view of the subject.
[0132] Optionally, when the subject is difficult to photograph directly, a vehicle-mounted panoramic camera or a robotic arm equipped with a panoramic camera can be used to capture the subject. In one example, when using a vehicle-mounted panoramic camera for panoramic video recording, the vehicle's route and speed can be preset to traverse the subject and obtain panoramic video. The above-described methods and settings for panoramic video recording of the subject are merely examples. In practical applications, any video recording method equipped with a panoramic camera can be used according to specific needs, and no restrictions are imposed here.
[0133] In this embodiment of the invention, the acquired panoramic video of the subject can be divided into multiple video frames, and at least two projection transformations can be performed on each of the acquired video frames. Specifically, at least a first projection transformation and a second projection transformation can be performed on the multiple video frames in the panoramic video, and multiple first projection video frames and multiple second projection video frames can be generated accordingly. Of course, in another embodiment of the invention, the projection transformation performed on the video frames of the panoramic video can also be more than two projection transformations, for example, it can include a third projection transformation, and a third two-dimensional view can be generated accordingly, for subsequent selection of partial views of the subject depending on the specific shooting scene, which is not limited here.
[0134] In this embodiment of the invention, in order to obtain accurate two-dimensional views without distortion for different parts of the subject being photographed, the first distorted region in the first projected video frame needs to be different from the second distorted region in the second projected video frame. In this embodiment, the first projection transformation and the second projection transformation can be one of various projection transformation methods, such as perspective transformation, equidistant cylindrical projection transformation, cylindrical projection transformation, Mercator transformation, etc. Depending on the projection transformation method, the corresponding distorted and undistorted regions are different. Therefore, the first projection transformation and the second projection transformation can be selected from these projection transformation methods that have different distorted and undistorted regions. For example, in one example of this embodiment, the first projection transformation is a perspective transformation, and the first undistorted region is the middle region of the first projected video frame; the second projection transformation is an equidistant cylindrical projection transformation, and the second undistorted region is the two ends of the second projected video frame; or the first projection transformation is an equidistant cylindrical projection transformation, and the first undistorted region is the two ends of the first projected video frame; the second projection transformation is a perspective transformation, and the second undistorted region is the middle region of the second projected video frame. In one example, the distortion region of the third projection transformation may be the same as or different from the distortion region of the first projection transformation / second projection transformation. In the subsequent selection of the two-dimensional view, the selection or combination can be made depending on the specific application scenario and image projection effect.
[0135] In one example of this invention, the region configuration between the undistorted and distorted regions in each projection transformation method can be preset, for example, it can be specified by the user. In another example of this invention, the region configuration between the undistorted and distorted regions in each projection transformation method can also be set by region-related parameters. Optionally, the region-related parameters may include one or more parameters among the type, size, product model, shooting position, and shooting time of the object being photographed. In one example, when the object being photographed is a vehicle chassis, the corresponding region-related parameters may include the vehicle model, the specific size of the vehicle chassis, the shooting position of the panoramic camera used for shooting, the shooting speed during movement, etc. For example, for a certain vehicle model, the first undistorted region corresponding to the projection can be the middle 25%-75% of the video frame in the panoramic video, and the remaining area is the first distorted region. As another example, for another vehicle model, the first undistorted region corresponding to the projection can be the middle 35%-70% of the video frame in the panoramic video, and the remaining area is the first distorted region. The above-described configuration methods for different regions in various projection transformations are merely examples. In practical applications, any corresponding region configuration can be selected according to specific needs, and no restrictions are imposed here. Furthermore, in this embodiment of the invention, the definitions of the various non-distorted regions and distorted regions are only relative. That is, each non-distorted region indicates a region with a relatively small degree of distortion after projection relative to other regions, and is not limited to regions without distortion, while each distorted region indicates a region with a relatively large degree of distortion after projection relative to other regions. In one example of this embodiment, when the ranges of the non-distorted and distorted regions are specified according to user input, they refer to the specific regions defined by the user. In one example of this embodiment, the distorted and distorted regions can be continuous and non-overlapping; however, in another example, the distorted and distorted regions can be discontinuous or overlapping, and no restrictions are imposed here.
[0136] After acquiring each video frame from the panoramic video, optionally, all video frames or a portion of them can be selected for subsequent stitching of the two-dimensional view. According to an embodiment of the present invention, optionally, all first projected video frames obtained by the panoramic video through a first projection transformation can be stitched together; furthermore, optionally, considering that there may be a large amount of overlapping areas between the first projected video frames leading to redundant calculations, multiple first keyframes can also be selected from the plurality of first projected video frames for stitching.
[0137] Of course, similarly, the same processing can be performed for the second projection transformation, that is, all second projection video frames obtained by subjecting the panoramic video to the second projection transformation can be stitched together, or a plurality of second key frames can be selected from the plurality of second projection video frames for stitching.
[0138] In a specific key frame extraction process, optionally, one frame can be extracted as a key frame from every N projected video frames, for example, N can be a fixed value. In addition, optionally, key frames can also be extracted by using a corresponding image processing algorithm according to shooting parameters when the panoramic camera shoots. In one example, the shooting parameters can be various parameters such as the shooting speed of the panoramic video, the type, size and product model of the shooting object. After obtaining the shooting parameters, optionally, various image processing algorithms such as a similarity-based method and a clustering method can be used to extract key frames.
[0139] In an example of the embodiment of the present invention, when key frame extraction is performed by a similarity calculation method, optionally, a plurality of parts at different positions can be selected in each frame of the video, and along a direction, for example, a direction of a straight line parallel to the shooting direction of the panoramic camera, within N frames, the similarity between a part in a first frame and another part in an n-th frame (1<n≤N) is calculated multiple times separately, and finally the frame with the largest similarity difference is selected as the key frame in the interval of the N frames. In addition, by the same token, the key frames in the next N-frame interval can be calculated continuously until all projected video frames are traversed.
[0140] After obtaining a plurality of projection video frames corresponding to each projection mode in the panoramic video, or selecting a plurality of key frames to be stitched from the projection video frames, optionally, feature point detection and matching is used to sequentially stitch each adjacent first / second projection video frame or first / second key frame, so as to obtain a two-dimensional view of the shooting object obtained by stitching. Specifically, the overlapping area between two adjacent frames along the moving direction of the panoramic camera can be calculated first, so that the offset of the two frames in the moving direction can be obtained through calculation of the overlapping area. Then, image fusion can be performed on the overlapping area of the two frames based on a pyramid multi-resolution method to generate a stitched image. In this example, when calculating the overlapping area, feature detection and matching can be performed on the two images first. In order to improve the accuracy of feature matching, targeted processing can be performed on the screening of feature points according to the characteristics of different projection transformations.
[0141] In one example, when the first projection transformation is a perspective transformation and the second projection transformation is an isometric cylindrical projection transformation: For the first projection transformation, considering that the first undistorted region is the middle region of the image, feature points located in the middle region have higher reliability. The relevant values (such as the mean, median, etc.) of feature matching in the middle region can be used to calculate the offset between adjacent frames, or the weight of the feature matching results in the middle region can be increased. Furthermore, for the second projection transformation, considering that the second undistorted region is the regions at both ends of the image, feature points located at both ends have higher reliability. The relevant values of feature matching in the regions at both ends can be used to calculate the offset between adjacent frames, or the weight of the feature matching results in the regions at both ends can be increased.
[0142] The judgment unit 2120 obtains the position of the first region in the shooting object and determines whether the first region is located in the first non-distorted region in the first projected video frame or in the second non-distorted region in the second projected video frame.
[0143] In this embodiment of the invention, optionally, after obtaining the first two-dimensional view and the second two-dimensional view of the entire subject, a distortion-free two-dimensional view of a local area of the subject can be further obtained. Considering that the distorted areas of the first and second two-dimensional views are different, the corresponding distorted and undistorted areas of each two-dimensional view obtained after stitching are also different. Therefore, it is necessary to determine which undistorted area of the two-dimensional view the local area of the subject is located in, so as to minimize distortion in the corresponding two-dimensional view portion.
[0144] Therefore, in this embodiment of the invention, the location of the first region can be obtained first and then associated with a non-distorted region of a certain projected video frame. Optionally, the first region can be specified by the user or obtained through various information (such as coordinate information). The first region can be the corresponding area of a defined graphic in the subject being photographed, or it can be a component or part of the subject being photographed, etc., and there are no restrictions here.
[0145] After obtaining the location of the first region, it can be determined which non-distorted region of the first region lies within in the projected video frame based on its specific location information (such as coordinates, region range, region size, etc.). Optionally, when the first region is entirely located within the non-distorted region of a single projected video frame, this determination can be made directly. Furthermore, optionally, when the first region is large or spans the non-distorted regions of two different projected video frames, the non-distorted region of the main corresponding projected video frame can be selected, or the non-distorted region of the projected video frame corresponding to a key part of the first region can be selected. Alternatively, optionally, where feasible, the corresponding parts of the first non-distorted region of the first projected video frame and the corresponding parts of the second non-distorted region of the second projected video frame can be fused and stitched together subsequently.
[0146] When the first region is located in the first non-distorted region of the first projected video frame, the association unit 2130 associates the first region with the corresponding position in the first two-dimensional view; when the first region is located in the second non-distorted region of the second projected video frame, the association unit 2130 associates the first region with the corresponding position in the second two-dimensional view.
[0147] In this embodiment of the invention, after determining whether the first region is located in the first undistorted region of the first projected video frame or the second undistorted region of the second projected video frame as described above, the first region can be associated with the first undistorted region of the first projected video frame and / or the second undistorted region of the second projected video frame, so as to provide relevant image information or display an undistorted two-dimensional view.
[0148] Optionally, embodiments of the present invention may further include: displaying a two-dimensional view of a corresponding position in the first two-dimensional view or a two-dimensional view of a corresponding position in the second two-dimensional view associated with the first region, based on the obtained indication regarding the first region.
[0149] The image processing apparatus according to embodiments of the present invention can establish a correspondence between a specified first region and a local two-dimensional view of a non-distorted region in a corresponding projection mode, and display the first region and the corresponding non-distorted two-dimensional view accordingly, thereby greatly improving the user experience.
[0150] Below, refer to Figure 22 The image generation apparatus according to embodiments of the present invention will be described. Figure 22 A block diagram of an image generation apparatus 2200 according to an embodiment of the present invention is shown. Figure 22 As shown, the device 2200 can be a computer or a server.
[0151] like Figure 22 As shown, the image generation apparatus 2200 includes one or more processors 2210 and a memory 2220. In addition, the image generation apparatus 2200 may also include input devices, output devices (not shown), etc., and these components can be interconnected via a bus system and / or other forms of connection mechanisms. It should be noted that... Figure 22 The components and structure of the image generation apparatus 2200 shown are merely exemplary and not limiting. The image generation apparatus 2200 may also have other components and structures as needed.
[0152] The processor 2210 may be a central processing unit (CPU) or other processing unit with data processing and / or instruction execution capabilities, and may utilize computer program instructions stored in memory 2220 to perform desired functions, including: acquiring panoramic video of a subject captured by a moving panoramic camera; performing a first projection transformation on multiple video frames in the panoramic video to generate multiple first projected video frames, and performing a second projection transformation on the multiple video frames in the panoramic video to generate multiple second projected video frames, wherein the first distortion region in the first projected video frames is different from the second distortion region in the second projected video frames; stitching together the multiple first undistorted regions in the multiple first projected video frames to generate a first two-dimensional view of the subject, and stitching together the multiple second undistorted regions in the multiple second projected video frames to generate a second two-dimensional view of the subject.
[0153] The memory 2220 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 2210 may execute the program instructions to implement the functions of the image generation apparatus of the embodiments of the present invention described above, and / or other desired functions, and / or to execute the image generation method according to the embodiments of the present invention. Various application programs and various data may also be stored in the computer-readable storage medium.
[0154] The following describes a computer-readable storage medium according to an embodiment of the present invention, having stored thereon computer program instructions, wherein the computer program instructions, when executed by a processor, perform the following steps: acquiring a panoramic video of a subject captured by a moving panoramic camera; performing a first projection transformation on a plurality of video frames in the panoramic video to generate a plurality of first projected video frames, and performing a second projection transformation on the plurality of video frames in the panoramic video to generate a plurality of second projected video frames, wherein a first distortion region in the first projected video frames is different from a second distortion region in the second projected video frames; stitching together a first non-distorted region in the plurality of first projected video frames to generate a first two-dimensional view of the subject, and stitching together a second non-distorted region in the plurality of second projected video frames to generate a second two-dimensional view of the subject.
[0155] The following describes a computer program product according to an embodiment of the present invention, including computer program instructions, wherein the computer program instructions, when executed by a processor, implement the steps of the aforementioned image generation method.
[0156] Below, refer to Figure 23 The image processing apparatus according to embodiments of the present invention will be described. Figure 23 A block diagram of an image processing apparatus 2300 according to an embodiment of the present invention is shown. Figure 23 As shown, the image processing device 2300 can be a computer or a server.
[0157] like Figure 23 As shown, device 2300 includes one or more processors 2310 and memory 2320. Furthermore, device 2300 may also include a transceiver 2330 for communication with clients. Of course, in addition to these, device 2300 may also include input devices, output devices (not shown), etc., and these components can be interconnected via a bus system and / or other forms of connection mechanisms. It should be noted that... Figure 23 The components and structure of the device 2300 shown are merely exemplary and not limiting; the device 2300 may also have other components and structures as needed.
[0158] The processor 2310 may be a central processing unit (CPU) or other processing unit with data processing and / or instruction execution capabilities, and may utilize computer program instructions stored in memory 2320 to perform desired functions, including: acquiring a first two-dimensional view and a second two-dimensional view of a photographed object, wherein the first two-dimensional view and the second two-dimensional view are generated by performing a first projection transformation and a second projection transformation respectively on multiple video frames in a panoramic video of the photographed object, wherein the first distortion region in the first projection video frame is different from the second distortion region in the second projection video frame, wherein the first two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames, and the multiple first projection video frames are transformed by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames, and the first distortion region in the first projection video frame is different from the second distortion region in the second projection video frame. The first undistorted region in the projected video frames is stitched together to generate the second two-dimensional view. The second two-dimensional view is generated by performing a second projection transformation on the multiple video frames to generate multiple second projected video frames, and then stitching together the second undistorted regions in the multiple second projected video frames. The position of the first region in the captured object is obtained, and it is determined whether the first region is located in the first undistorted region of the first projected video frame or in the second undistorted region of the second projected video frame. If the first region is located in the first undistorted region of the first projected video frame, the first region is associated with the corresponding position in the first two-dimensional view. If the first region is located in the second undistorted region of the second projected video frame, the first region is associated with the corresponding position in the second two-dimensional view.
[0159] The memory 2320 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 2310 may execute the program instructions to implement the functions of the image processing apparatus of the embodiments of the present invention described above, and / or other desired functions, and / or to execute the image processing method according to the embodiments of the present invention. Various application programs and various data may also be stored in the computer-readable storage medium.
[0160] The following describes a computer-readable storage medium according to an embodiment of the present invention, having stored thereon computer program instructions, wherein the computer program instructions, when executed by a processor, perform the following steps: acquiring a first two-dimensional view and a second two-dimensional view of a photographed object, wherein the first two-dimensional view and the second two-dimensional view are generated by performing a first projection transformation and a second projection transformation respectively on multiple video frames in a panoramic video of the photographed object, wherein a first distortion region in the first projection video frame is different from a second distortion region in the second projection video frame, wherein the first two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames, and a first non-distorted region in the multiple first projection video frames is obtained. The second two-dimensional view is generated by stitching together multiple video frames. This is achieved by performing a second projection transformation on the multiple video frames to generate multiple second projected video frames, and then stitching together the second non-distorted regions within these multiple second projected video frames. The position of a first region in the captured object is obtained, and it is determined whether the first region is located in the first non-distorted region of the first projected video frame or in the second non-distorted region of the second projected video frame. If the first region is located in the first non-distorted region of the first projected video frame, it is associated with the corresponding position in the first two-dimensional view. If the first region is located in the second non-distorted region of the second projected video frame, it is associated with the corresponding position in the second two-dimensional view.
[0161] The following describes a computer program product according to an embodiment of the present invention, including computer program instructions, wherein the computer program instructions, when executed by a processor, implement the steps of the aforementioned image processing method.
[0162] Of course, the specific embodiments described above are merely examples and not limitations. Those skilled in the art can combine and integrate some steps and devices from the various embodiments described separately above to achieve the effects of the present invention. Such combined and integrated embodiments are also included in the present invention, but will not be described one by one here.
[0163] Note that the advantages, benefits, and effects mentioned in this invention are merely examples and not limitations, and should not be considered as essential features of every embodiment of the invention. Furthermore, the specific details described above are for illustrative and illustrative purposes only, and are not intended to limit the invention. These details do not limit the invention from being implemented solely by employing these specific details.
[0164] The block diagrams of devices, apparatuses, devices, and systems involved in this invention are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0165] The flowcharts and method descriptions in this invention are merely illustrative examples and are not intended to require or imply that the steps of the various embodiments must be performed in the given order. As those skilled in the art will recognize, the steps in the above embodiments can be performed in any order. Words such as "then," "next," etc., are not intended to limit the order of steps; these words are only used to guide the reader through the description of these methods. Furthermore, any reference to a singular element, such as the use of the articles "a," "one," or "the," is not to be construed as limiting that element to the singular.
[0166] Furthermore, the steps and apparatus in the various embodiments herein are not limited to any one embodiment. In fact, new embodiments can be conceived by combining relevant steps and apparatus in the various embodiments herein with the concepts of the present invention, and these new embodiments are also included within the scope of the present invention.
[0167] Each operation described above can be performed by any suitable means capable of performing the corresponding function. Such means may include various hardware and / or software components and / or modules, including but not limited to circuits, application-specific integrated circuits (ASICs), or processors.
[0168] The various exemplified logic blocks, modules, and circuits described herein can be implemented or performed using a general-purpose processor, digital signal processor (DSP), ASIC, field-programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but alternatively, it may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.
[0169] The steps of the methods or algorithms described in this invention can be directly embedded in hardware, in a software module executed by a processor, or a combination of both. The software module can reside in any form of tangible storage medium. Some examples of usable storage media include random access memory (RAM), read-only memory (ROM), flash memory, EPROM, EEPROM, registers, hard disks, removable disks, CD-ROMs, etc. The storage medium can be coupled to the processor so that the processor can read information from and write information to the storage medium. Alternatively, the storage medium can be integral with the processor. The software module can be a single instruction or many instructions, and can be distributed across several different code segments, different programs, and across multiple storage media.
[0170] The method of this invention includes one or more actions for implementing the method. The methods and / or actions may be interchanged without departing from the scope of the claims. In other words, unless a specific order of actions is specified, the order and / or use of specific actions may be modified without departing from the scope of the claims.
[0171] The described functionality can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality can be stored as one or more instructions on a tangible computer-readable medium. The storage medium can be any available tangible medium that can be accessed by a computer. By way of example, and not limitation, such a computer-readable medium can include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other tangible medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, a disc includes a compact disc (CD), a laser disc, an optical disc, a digital universal disc (DVD), a floppy disk, and a Blu-ray disc.
[0172] Therefore, a computer program product can perform the operations described herein. For example, such a computer program product can be a computer-readable tangible medium having instructions tangibly stored (and / or encoded) thereon, which can be executed by one or more processors to perform the operations described herein. The computer program product may include packaging materials.
[0173] Software or instructions can also be transmitted via a transmission medium. For example, software can be transmitted from a website, server, or other remote source using transmission media such as coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, or microwave.
[0174] Furthermore, modules and / or other suitable means for carrying out the methods and techniques described herein can be downloaded and / or obtained by user terminals and / or base stations as appropriate. For example, such a device can be coupled to a server to facilitate the transmission of means for carrying out the methods described herein. Alternatively, the various methods described herein can be provided via storage components (e.g., RAM, ROM, physical storage media such as CDs or floppy disks) so that user terminals and / or base stations can obtain the various methods when coupled to the device or when providing storage components to the device. Furthermore, any other suitable techniques for providing the methods and techniques described herein to the device can be utilized.
[0175] Other examples and implementations are within the scope and spirit of this invention and the appended claims. For example, due to the nature of software, the functions described above can be implemented using software executed by a processor, hardware, firmware, hardwired, or any combination thereof. Features implementing the functions can also be physically located in various places, including being distributed so that parts of the functions are implemented at different physical locations. Moreover, as used herein, including as used in the claims, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the described examples are preferred or better than other examples.
[0176] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims is not limited to the specific aspects of the processes, machines, manufacturing processes, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufacturing processes, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufacturing processes, events, means, methods, or actions within their scope.
[0177] The above description of aspects of the invention is provided to enable any person skilled in the art to make or use the invention. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the invention. Therefore, the invention is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features of the invention herein.
[0178] The above description has been given for illustrative and descriptive purposes. Furthermore, this description is not intended to limit the embodiments of the invention to the forms described herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. An image generation method, comprising: Acquire panoramic video of the subject using a moving panoramic camera; Multiple video frames in the panoramic video are subjected to a first projection transformation to generate multiple first projection video frames, and multiple video frames in the panoramic video are subjected to a second projection transformation to generate multiple second projection video frames, wherein the first distortion region in the first projection video frame that produces distortion is different from the second distortion region in the second projection video frame. A first two-dimensional view of the photographed object is generated by stitching together the first non-distorted regions in the plurality of first projected video frames, and a second two-dimensional view of the photographed object is generated by stitching together the second non-distorted regions in the plurality of second projected video frames.
2. The method according to claim 1, wherein, Acquiring panoramic video of a subject using a moving panoramic camera includes: Obtain the panoramic video generated by moving the panoramic camera at a constant speed through all areas of the subject being photographed.
3. The method according to claim 1, wherein, The first projection transformation is a perspective transformation, and the first undistorted region is the middle region of the first projected video frame; the second projection transformation is an equidistant cylindrical projection transformation, and the second undistorted region is the regions at both ends of the second projected video frame; or The first projection transformation is an equidistant cylindrical projection transformation, and the first undistorted region is the two ends of the first projected video frame; the second projection transformation is a perspective transformation, and the second undistorted region is the middle region of the second projected video frame.
4. The method according to claim 1, wherein, Generating a first two-dimensional view of the photographed object by stitching together the first non-distorted regions from the plurality of first projected video frames includes: selecting a plurality of first keyframes from the plurality of first projected video frames, and generating a first two-dimensional view of the photographed object based on the plurality of first keyframes; or The process of stitching together the second non-distorted regions in the plurality of second projected video frames to generate a second two-dimensional view of the photographed object includes: selecting a plurality of second keyframes from the plurality of second projected video frames, and generating a second two-dimensional view of the photographed object based on the plurality of second keyframes.
5. The method according to claim 4, wherein, The method further includes: Based on the shooting parameters of the panoramic camera, an image processing algorithm is used to select multiple first keyframes from the multiple first projected video frames, and / or select multiple second keyframes from the multiple second projected video frames.
6. The method according to claim 1, wherein, Generating a first two-dimensional view of the photographed object by stitching together the first non-distorted regions in the plurality of first projected video frames includes: performing feature point detection and matching between the first non-distorted regions of adjacent frames of the plurality of first projected video frames; stitching together the adjacent frames based on the feature point detection and matching results to generate a first two-dimensional view of the photographed object; and / or The process of stitching together the second non-distorted regions in the plurality of second projected video frames to generate a second two-dimensional view of the photographed object includes: performing feature point detection and matching between the second non-distorted regions of adjacent frames of the plurality of second projected video frames, and stitching together the adjacent frames based on the feature point detection and matching results to generate a second two-dimensional view of the photographed object.
7. The method according to claim 1, wherein, The stitching method is pyramid multi-resolution image fusion.
8. The method according to claim 1, wherein, The region configuration of the first distorted region and the first undistorted region in the first projected video frame, and / or the region configuration of the second distorted region and the second undistorted region in the second projected video frame, are preset or set according to region-related parameters.
9. The method according to claim 8, wherein, The relevant parameters for the region include: The type, size, product model, shooting location, shooting time, and panoramic video shooting speed of the object being photographed are among one or more parameters.
10. An image processing method, comprising: A first two-dimensional view and a second two-dimensional view of the subject are obtained. The first two-dimensional view and the second two-dimensional view are generated by performing a first projection transformation and a second projection transformation on multiple video frames in a panoramic video of the subject. The first distortion region in the first projection video frame is different from the second distortion region in the second projection video frame. The first two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames and stitching together the first non-distorted regions in the multiple first projection video frames. The second two-dimensional view is generated by performing a second projection transformation on the multiple video frames to generate multiple second projection video frames and stitching together the second non-distorted regions in the multiple second projection video frames. Obtain the position of the first region in the shooting object, and determine whether the first region is located in the first non-distorted region in the first projected video frame or in the second non-distorted region in the second projected video frame; If the first region is located in the first non-distorted region of the first projected video frame, the first region is associated with the corresponding position in the first two-dimensional view; if the first region is located in the second non-distorted region of the second projected video frame, the first region is associated with the corresponding position in the second two-dimensional view.
11. The method according to claim 10, wherein, The first projection transformation is a perspective transformation, and the first undistorted region is the middle region of the first projected video frame; the second projection transformation is an equidistant cylindrical projection transformation, and the second undistorted region is the regions at both ends of the second projected video frame; or The first projection transformation is an equidistant cylindrical projection transformation, and the first undistorted region is the two ends of the first projected video frame; the second projection transformation is a perspective transformation, and the second undistorted region is the middle region of the second projected video frame.
12. The method according to claim 10, wherein, The region configuration of the first distorted region and the first undistorted region in the first projected video frame, and / or the region configuration of the second distorted region and the second undistorted region in the second projected video frame, are preset or set according to region-related parameters.
13. The method according to claim 12, wherein, The relevant parameters for the region include: The type, size, product model, shooting location, shooting time, and panoramic video shooting speed of the object being photographed are among one or more parameters.
14. The method of claim 10, wherein, The method further includes: Based on the obtained instructions regarding the first region, display a two-dimensional view of the corresponding location in the first two-dimensional view or a two-dimensional view of the corresponding location in the second two-dimensional view associated with the first region.
15. An image generation apparatus, comprising: The shooting unit is configured to acquire panoramic video of the subject captured by a moving panoramic camera. The generation unit is configured to perform a first projection transformation on multiple video frames in the panoramic video to generate multiple first projection video frames, and to perform a second projection transformation on multiple video frames in the panoramic video to generate multiple second projection video frames, wherein the first distortion region in the first projection video frame that produces distortion is different from the second distortion region in the second projection video frame. The stitching unit is configured to stitch together a first non-distorted region from the plurality of first projected video frames to generate a first two-dimensional view of the subject being photographed, and to stitch together a second non-distorted region from the plurality of second projected video frames to generate a second two-dimensional view of the subject being photographed.
16. An image generation apparatus, comprising: processor; and a memory, in which computer program instructions are stored. When the computer program instructions are executed by the processor, the processor performs the following steps: Acquire panoramic video of the subject using a moving panoramic camera; Multiple video frames in the panoramic video are subjected to a first projection transformation to generate multiple first projection video frames, and multiple video frames in the panoramic video are subjected to a second projection transformation to generate multiple second projection video frames, wherein the first distortion region in the first projection video frame that produces distortion is different from the second distortion region in the second projection video frame. A first two-dimensional view of the photographed object is generated by stitching together the first non-distorted regions in the plurality of first projected video frames, and a second two-dimensional view of the photographed object is generated by stitching together the second non-distorted regions in the plurality of second projected video frames.
17. A computer-readable storage medium having stored thereon computer program instructions, wherein, When the computer program instructions are executed by the processor, the following steps are performed: Acquire panoramic video of the subject using a moving panoramic camera; Multiple video frames in the panoramic video are subjected to a first projection transformation to generate multiple first projection video frames, and multiple video frames in the panoramic video are subjected to a second projection transformation to generate multiple second projection video frames, wherein the first distortion region in the first projection video frame that produces distortion is different from the second distortion region in the second projection video frame. A first two-dimensional view of the photographed object is generated by stitching together the first non-distorted regions in the plurality of first projected video frames, and a second two-dimensional view of the photographed object is generated by stitching together the second non-distorted regions in the plurality of second projected video frames.
18. A computer program product comprising computer program instructions, wherein, When the computer program instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-9.
19. An image processing apparatus, comprising: The acquisition unit is configured to acquire a first two-dimensional view and a second two-dimensional view of the subject being photographed. The first two-dimensional view and the second two-dimensional view are generated by performing a first projection transformation and a second projection transformation on multiple video frames in a panoramic video of the subject being photographed. The first distortion region in the first projection video frame is different from the second distortion region in the second projection video frame. The first two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames and stitching together the first non-distorted regions in the multiple first projection video frames. The second two-dimensional view is generated by performing a second projection transformation on the multiple video frames to generate multiple second projection video frames and stitching together the second non-distorted regions in the multiple second projection video frames. The judgment unit is configured to obtain the position of the first region in the shooting object and determine whether the first region is located in the first non-distorted region in the first projected video frame or in the second non-distorted region in the second projected video frame. The association unit is configured to associate the first region with a corresponding position in the first two-dimensional view if the first region is located in a first non-distorted region in the first projected video frame; and to associate the first region with a corresponding position in the second two-dimensional view if the first region is located in a second non-distorted region in the second projected video frame.
20. An image processing apparatus, comprising: processor; and a memory, in which computer program instructions are stored. When the computer program instructions are executed by the processor, the processor performs the following steps: A first two-dimensional view and a second two-dimensional view of the subject are obtained. The first two-dimensional view and the second two-dimensional view are generated by performing a first projection transformation and a second projection transformation on multiple video frames in a panoramic video of the subject. The first distortion region in the first projection video frame is different from the second distortion region in the second projection video frame. The first two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames and stitching together the first non-distorted regions in the multiple first projection video frames. The second two-dimensional view is generated by performing a second projection transformation on the multiple video frames to generate multiple second projection video frames and stitching together the second non-distorted regions in the multiple second projection video frames. Obtain the position of the first region in the shooting object, and determine whether the first region is located in the first non-distorted region in the first projected video frame or in the second non-distorted region in the second projected video frame; If the first region is located in the first non-distorted region of the first projected video frame, the first region is associated with the corresponding position in the first two-dimensional view; if the first region is located in the second non-distorted region of the second projected video frame, the first region is associated with the corresponding position in the second two-dimensional view.
21. A computer-readable storage medium having stored thereon computer program instructions, wherein, When the computer program instructions are executed by the processor, the following steps are performed: A first two-dimensional view and a second two-dimensional view of the subject are obtained. The first two-dimensional view and the second two-dimensional view are generated by performing a first projection transformation and a second projection transformation on multiple video frames in a panoramic video of the subject. The first distortion region in the first projection video frame is different from the second distortion region in the second projection video frame. The first two-dimensional view is generated by performing a first projection transformation on the multiple video frames to generate multiple first projection video frames and stitching together the first non-distorted regions in the multiple first projection video frames. The second two-dimensional view is generated by performing a second projection transformation on the multiple video frames to generate multiple second projection video frames and stitching together the second non-distorted regions in the multiple second projection video frames. Obtain the position of the first region in the shooting object, and determine whether the first region is located in the first non-distorted region in the first projected video frame or in the second non-distorted region in the second projected video frame; If the first region is located in the first non-distorted region of the first projected video frame, the first region is associated with the corresponding position in the first two-dimensional view; if the first region is located in the second non-distorted region of the second projected video frame, the first region is associated with the corresponding position in the second two-dimensional view.
22. A computer program product comprising computer program instructions, wherein, When the computer program instructions are executed by a processor, they implement the steps of the method according to any one of claims 10-14.