A method, apparatus and device for video image stitching

By identifying overlapping areas and the boundaries and expansion areas of moving objects in video image stitching, and employing an improved stitching algorithm, the problem of poor stitching caused by moving objects is solved, thus improving the stitching quality of wide-angle video images.

CN115293969BActive Publication Date: 2026-01-09SHANGHAI SMAWAVE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210879185.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-25
Publication Date
2026-01-09
Estimated Expiration
2042-07-25

AI Technical Summary

Technical Problem

In wide-angle video image stitching, when moving objects pass through the stitching line of the overlapping area, the stitching effect is poor, with obvious seams and blurring.

Method used

By acquiring image features from several video clips, the boundaries and expansion regions of overlapping areas and moving objects are determined. An improved stitching algorithm is used to avoid moving objects. SIFT, FLANN, and SORT models are used for feature matching and tracking. The intensity threshold of pixels is set to determine the optimal stitching line.

Benefits of technology

It effectively avoids the impact of moving objects on the splicing effect, improves splicing quality and performance, and reduces the generation of ghosting and splicing seams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115293969B_ABST
    Figure CN115293969B_ABST
Patent Text Reader

Abstract

The application provides a technical scheme for video image splicing. The method comprises the following steps: acquiring a plurality of video clips, wherein each video clip comprises a plurality of frames of continuous video images; determining image features of each frame of video image based on each video clip, and determining an overlapping area of each frame of video image of each video clip based on the image features; determining a boundary and an inflation area of a moving object in the overlapping area; and splicing synchronous frame video images of corresponding video clips based on the overlapping area, the boundary and the inflation area of the moving object in the overlapping area, so as to obtain a large-view-angle video image after splicing. Through the method, different processing can be performed on the boundary and the inflation area of the moving object in the predicted overlapping area of the video image when the synchronous frame video images of the corresponding video clips are spliced, the influence of the moving object in the overlapping area on the splicing effect can be avoided, and better splicing quality and effect can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video image processing, and in particular to a technology for video image splicing. BACKGROUND

[0002] In an application scenario requiring large-angle video images, such as large-area monitoring or remote control, a plurality of cameras are usually arranged at corresponding positions to form a camera array, each camera captures a scene of a fixed field of view, or a scene of a corresponding field of view is captured through synchronous slow rotation of the camera, so that the scenes captured by all the cameras can cover the entire scene area to be monitored, wherein the video images captured by the related cameras should include overlapping areas. Then, the video images in each video captured by each camera are spliced to synthesize a large-angle video image, so that the entire area can be monitored in the large-angle video image.

[0003] Due to the difference in camera angles, the video images captured by different cameras for the same area have a certain parallax, so that the spliced video image will have obvious seams in the overlapping area, as well as blurring, ghosting and other effects that affect the quality of the video image. Especially when the video image captured includes a moving object and the moving object passes through the stitching line of the overlapping area, this phenomenon will be more obvious. SUMMARY

[0004] The purpose of the present application is to provide a technical solution for video image splicing to at least partially solve the technical problem of poor splicing effect of video images containing moving objects in the overlapping area in the prior art.

[0005] According to one aspect of the present application, a method for video image splicing is provided, wherein the method comprises:

[0006] obtaining a plurality of video segments, wherein each video segment includes a plurality of frames of continuous video images;

[0007] determining the image features of each frame of video images of each video segment based on each video segment, and determining the overlapping area of each frame of video images of each video segment based on the image features of the synchronous frame video images of the corresponding video segment;

[0008] determining the boundary and the inflation area of the moving object in the overlapping area of each frame of video images of each video segment;

[0009] splicing the synchronous frame video images of the corresponding video segment based on the overlapping area of each frame of video images of each video segment, the boundary and the inflation area of the moving object in the overlapping area, to obtain a spliced large-angle video image.

[0010] Optionally, the determining the image features of the synchronous frame video images of each video segment based on each video segment comprises:

[0011] Optionally, the determining the image features of the synchronous frame video images of each video segment based on each video segment comprises:

[0012] Optionally, the determining the image features of the synchronous frame video images of each video segment based on each video segment comprises:

[0013] Optionally, the determining the image features of the synchronous frame video images of each video segment based on each video segment comprises:

[0014] Optionally, the image feature matching model comprises a FLANN model.

[0015] Optionally, the determining the boundary and the inflation area of the moving object in the overlapping area of each frame video image of each video segment comprises:

[0016] Optionally, the determining the boundary and the inflation area of the moving object in the overlapping area of each frame video image of each video segment comprises:

[0017] Optionally, the multi-target tracking model comprises a SORT model.

[0018] Optionally, the splicing the synchronous frame video images of each video segment based on the overlapping area of each frame video image of each video segment, the boundary and the inflation area of the moving object in the overlapping area of each frame video image of each video segment comprises:

[0019] Optionally, the splicing the synchronous frame video images of each video segment based on the overlapping area of each frame video image of each video segment, the boundary and the inflation area of the moving object in the overlapping area of each frame video image of each video segment comprises:

[0020] Optionally, the splicing the synchronous frame video images of each video segment based on the overlapping area of each frame video image of each video segment, the boundary and the inflation area of the moving object in the overlapping area of each frame video image of each video segment comprises:

[0021] Optionally, the splicing the synchronous frame video images of each video segment based on the overlapping area of each frame video image of each video segment, the boundary and the inflation area of the moving object in the overlapping area of each frame video image of each video segment comprises:

[0022] set the intensity value of the pixel point on the boundary of the motion object in the overlapping region of each frame of video image of each video segment to a first preset intensity threshold value, and set the intensity value of the pixel point on the boundary of the expanded region of the motion object in the overlapping region of each frame of video image of each video segment to a second preset intensity threshold value;

[0023] set the intensity value of each pixel point in the expanded region based on the first preset intensity threshold value, the second preset intensity threshold value and a preset smoothing function;

[0024] based on the intensity values of the pixel points in the overlapping region of the synchronous frame video images of the corresponding video segment, determine a pixel line formed by adjacent pixel points across the overlapping region and having the minimum intensity value change as the optimal stitching line between the synchronous frame video images of the corresponding video segment.

[0025] Optionally, wherein the setting the intensity value of each pixel point in the expanded region based on the first preset intensity threshold value, the second preset intensity threshold value and a preset smoothing function comprises:

[0026] based on the first preset intensity threshold value, the second preset intensity threshold value and a preset smoothing function, set the intensity value of the pixel point on the boundary of the region corresponding to each prediction state in the expanded region, wherein the boundary of the region corresponding to each prediction state in the expanded region is marked with a preset number of prediction states of the motion object;

[0027] based on the intensity values of the pixel points on the boundaries of the regions corresponding to adjacent prediction states and the preset smoothing function, set the intensity value of each pixel point in the region corresponding to the corresponding prediction state.

[0028] According to another aspect of the present application, there is also provided a device for video image stitching, wherein the device comprises:

[0029] a first module for acquiring a plurality of video segments, wherein each video segment comprises a plurality of frames of continuous video images;

[0030] a second module for determining the image features of each frame of video image of each video segment based on each video segment, and determining the overlapping region of each frame of video image of each video segment based on the image features of the synchronous frame video images of the corresponding video segment;

[0031] a third module for determining the boundary and the expanded region of the motion object in the overlapping region of each frame of video image of each video segment;

[0032] a fourth module for stitching the synchronous frame video images of the corresponding video segment based on the overlapping region, the boundary and the expanded region of the motion object in the overlapping region of each frame of video image of each video segment, to obtain a large-view-angle video image after stitching.

[0033] Compared with existing technologies, this application provides a technical solution for video image stitching, the method of which includes: firstly, acquiring several video segments, wherein each video segment includes several consecutive video images; then, based on each video segment, determining the image features of each frame of the video image in each video segment, and based on the image features of the synchronous frame video images of the corresponding video segments, determining the overlapping region of each frame of the video images in each video segment; then, determining the boundary and expansion region of moving objects in the overlapping region of each frame of the video images in each video segment; finally, based on the overlapping region of each frame of the video images in each video segment, the boundary and expansion region of moving objects in the overlapping region, stitching the synchronous frame video images of the corresponding video segments to obtain a stitched wide-angle video image. Optionally, the intensity value of the pixel on the boundary of the moving object in the overlapping region of each frame of each video image of each video segment is set to a first preset intensity threshold, and the intensity value of the pixel on the boundary of the expansion region of the moving object in the overlapping region of each frame of each video image of each video segment is set to a second preset intensity threshold; based on the first preset intensity threshold, the second preset intensity threshold, and a preset smoothing function, the intensity value of each pixel in the expansion region is set; based on the intensity value of each pixel in the overlapping region of the synchronous frame video images of the corresponding video segment, the pixel line that crosses the overlapping region and is composed of adjacent pixels with the smallest change in intensity value is determined as the optimal stitching line between the synchronous frame video images of the corresponding video segment; then, based on the optimal stitching line, the synchronous frame video images of the corresponding video segment are stitched together to obtain a stitched wide-angle video image.

[0034] The technical effects of the video image stitching solution provided in this application are as follows:

[0035] In the process of stitching synchronous frame video images of corresponding video segments, different processing is applied to the boundaries of moving objects and the dilated regions in the overlapping areas of the predicted video images. This avoids the impact of moving objects in the overlapping areas on the stitching effect, resulting in better stitching quality and effect. Optionally, the intensity value of each pixel in the dilated region is smoothly set by incorporating the predicted information of moving objects when determining the optimal stitching line, thereby improving the performance of video image stitching. Attached Figure Description

[0036] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0037] Figure 1 A flowchart of a method for video image stitching according to one aspect of this application is shown;

[0038] Figure 2Fig. 1 shows a schematic diagram of an apparatus for video image stitching according to another aspect of the present application;

[0039] The same or similar reference signs in the drawings represent the same or similar components. DETAILED DESCRIPTION

[0040] The application will be further described below in conjunction with the drawings.

[0041] In one typical configuration of the present application, the execution subject of the method, each trusted party of the system and / or each module of the apparatus can include one or more processors (CPUs), input / output interfaces, network interfaces and memories.

[0042] The memory can include non-persistent memory in computer readable media, random access memory (RAM), and / or non-volatile memory such as read only memory (ROM) or flash memory. The memory is an example of computer readable media.

[0043] Computer readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic disk storage or other magnetic storage device, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer readable media does not include non-transitory computer readable media such as modulated data signals and carriers.

[0044] In order to further illustrate the technical means adopted by the present application and the effects achieved, the technical solutions of the present application will be described clearly and completely in conjunction with the drawings and preferred embodiments.

[0045] Figure 1 Fig. 1 shows a schematic diagram of an apparatus for video image stitching according to another aspect of the present application;

[0046] S101 acquiring a plurality of video segments, wherein each video segment includes a plurality of frames of continuous video images;

[0047] S102, based on each video segment, determining image features of each frame of video image of each video segment, and based on the image features of the synchronous frame of video image of the corresponding video segment, determining the overlapping area of each frame of video image of each video segment;

[0048] S103, determining the boundary and inflation area of the moving object in the overlapping area of each frame of video image of each video segment;

[0049] S104, based on the overlapping area of each frame of video image of each video segment, the boundary and inflation area of the moving object in the overlapping area, splicing the synchronous frame of video image of the corresponding video segment to obtain the large-view-angle video image after splicing.

[0050] In the present application, the method is executed by a device 100, which is a computer device and / or a cloud. The computer device includes but is not limited to a personal computer, a notebook computer, an industrial computer, a network host, a single network server, and a plurality of network server sets. The cloud is composed of a large number of computers or network servers based on cloud computing, which is a kind of distributed computing composed of a virtual supercomputer formed by a loose-coupled computer set.

[0051] Herein, the computer device and / or the cloud are only examples, and other existing or future devices and / or resource sharing platforms suitable for the present application should also be included in the protection scope of the present application, which are hereby included by reference.

[0052] In this embodiment, the plurality of video segments are captured by a camera array composed of a plurality of cameras, wherein the video segment captured by each camera covers a fixed field of view or a preset field of view range of a scene, so that the video segments captured by all the cameras can cover the entire scene area to be monitored, wherein the synchronous frame of video image of the video segment captured by the related camera should include an overlapping area.

[0053] In step S101, the device 100 synchronously acquires a plurality of video segments captured by a camera array, wherein each video segment includes a plurality of frames of continuous video images. The synchronous frame of video image of the corresponding video segment should include an overlapping area.

[0054] The device 100 can acquire each video segment through real-time video acquisition, network transmission, offline cache copy, etc. without limitation, and other existing or future video segment acquisition formats suitable for the present application should also be included in the protection scope of the present application.

[0055] In this embodiment, in the step S102, the device 100 determines the image features of each frame of video image of each video segment based on each video segment, and determines the overlapping area of each frame of video image of each video segment based on the image features of the synchronous frame of video image of the corresponding video segment.

[0056] In this embodiment, in the step S102, the device 100 determines the image features of each frame of video image of each video segment based on each video segment, and determines the overlapping area of each frame of video image of each video segment based on the image features of the synchronous frame of video image of the corresponding video segment.

[0057] In this embodiment, in the step S102, the device 100 determines the image features of each frame of video image of each video segment based on each video segment, and determines the overlapping area of each frame of video image of each video segment based on the image features of the synchronous frame of video image of the corresponding video segment.

[0058] In this embodiment, in the step S102, the device 100 determines the image features of each frame of video image of each video segment based on each video segment, and determines the overlapping area of each frame of video image of each video segment based on the image features of the synchronous frame of video image of the corresponding video segment.

[0059] Optionally, the step of determining the image features of the synchronous frame of video image of each video segment based on each video segment comprises:

[0060] In this embodiment, in the step S102, the device 100 determines the image features of each frame of video image of each video segment based on each video segment, and determines the overlapping area of each frame of video image of each video segment based on the image features of the synchronous frame of video image of the corresponding video segment.

[0061] In this embodiment, in the step S102, the device 100 determines the image features of each frame of video image of each video segment based on each video segment, and determines the overlapping area of each frame of video image of each video segment based on the image features of the synchronous frame of video image of the corresponding video segment.

[0062] The device 100 inputs each acquired video clip into a SIFT model to acquire SIFT features of each frame of video image of each video clip. For example, the SIFT model can be built by using OpenCV library and Python, and each acquired video clip is input into the SIFT model to acquire SIFT features of each frame of video image of each video clip.

[0063] Optionally, the determining of the overlapping area of each frame of video image of each video clip based on the image features of the synchronized frames of video image of the corresponding video clip comprises:

[0064] The SIFT features of the synchronized frames of video image of the corresponding video clip are input into an image feature matching model to determine the overlapping area of each frame of video image of each video clip.

[0065] Each SIFT feature (or SIFT feature descriptor) of a video image is actually a feature vector based on a key point and a direction, and the SIFT features of the synchronized frames of video image of the corresponding video clip are input into the image feature matching model to traverse each SIFT feature vector of the video image, perform matching comparison, and determine whether the key points corresponding to the SIFT features of the synchronized frames of video image of the corresponding video clip are matching points, so as to determine whether the synchronized frames of video image of the corresponding video clip have overlapping areas.

[0066] The image feature matching model is used to perform matching comparison on local image features of corresponding video images to determine whether the corresponding video images have overlapping areas.

[0067] The image feature matching model can be used to determine the overlapping area of each frame of video image of each video clip, and perform affine transformation on the synchronized frames of video image of the related video clip to obtain synchronized frames of video image of the same space of the corresponding video clip in a unified perspective, which are used for subsequent video image stitching.

[0068] Optionally, the image feature matching model comprises a FLANN model.

[0069] The FLANN (Fast Library for Approximate Nearest Neighbors) is a feature matching algorithm model, which is a collection of algorithms for nearest neighbor search on large datasets and high-dimensional features, and these algorithms have been optimized.

[0070] The image feature matching model in this embodiment can also use other image feature matching models such as a BF (Brute Force) model, and any existing or future image feature matching model that is applicable to the present application should be included in the protection scope of the present application.

[0071] In this embodiment, in step S103, the device 100 determines the boundary and the inflation area of the moving object in the overlapping area of each frame of video image of each video segment.

[0072] If the video image in the overlapping area includes a moving object, the seam line in the conventional splicing will pass through the moving object area, and the misalignment of the moving object and other situations may occur. In order to avoid such situations, the boundary of the moving object in the overlapping area of the video image can be determined first, and the area corresponding to the subsequent motion state of the moving object, i.e. the inflation area, can be predicted, so as to avoid the moving object and its inflation area during video image splicing.

[0073] Optionally, the step S103 includes:

[0074] The overlapping area of each frame of video image of each video segment is cropped, and a multi-target tracking model is inputted to determine the boundary and the inflation area of the moving object in the overlapping area of each frame of video image of each video segment.

[0075] Since the splicing line of the synchronous frame video images of the same space of the uniform perspective of the corresponding video segment is across the overlapping area during splicing, the moving object in the non-overlapping area of the video image will not affect the splicing quality and effect, and the moving object in the non-overlapping area does not need to be involved in the determination of the splicing line. In order to reduce the amount of calculation and improve the operation performance, only the influence of the moving object in the overlapping area on the quality and effect of image splicing needs to be considered. Therefore, after the image feature matching model determines the overlapping area of each frame of video image of each video segment, the overlapping area of each frame of video image of each video segment can be cropped, and the overlapping area can be inputted into a multi-target tracking model to determine the boundary and the inflation area of the moving object in the overlapping area of each frame of video image of each video segment.

[0076] The multi-target tracking model usually includes a Kalman filter part, and the prediction information provided by the Kalman filter can be used to determine the inflation area of the moving object.

[0077] Optionally, the multi-target tracking model includes a SORT model.

[0078] The SORT (Simple Online and Realtime Tracking) model is a MOT (Multiple Object Tracking) algorithm.

[0079] The SORT model is usually used to realize real-time tracking of multiple targets in a video image, determine the boundary of a moving object, and avoid the moving object in subsequent video image stitching, but ghosting and stitching seams may still occur, affecting the stitching quality and effect.

[0080] To avoid ghosting and stitching seams, in an embodiment of the present application, the SORT model is also used to determine the inflation area of a moving object in a current frame of video image based on the state information of the moving object in the current frame of video image predicted from a previous frame of video image by Kalman filtering. Thus, the area to be avoided in subsequent video image stitching includes not only the moving object but also the inflation area of the moving object.

[0081] For example, the state modeling of a moving object in a current frame of video image can be represented as:

[0082] X = [u, v, s, r, u', v', s'] T

[0083] wherein u and v respectively represent the horizontal and vertical pixel positions of the center of a moving object in a current frame of video image; s represents the proportion (or area) of the bounding box of the moving object in the current frame of video image; u' and v' respectively represent the horizontal and vertical pixel positions of the center of the moving object in a predicted next frame of video image; s' represents the proportion (or area) of the bounding box of the moving object in the predicted next frame of video image; and r represents the aspect ratio of the bounding box of the moving object, which is an external parameter and can be considered as a constant in combination with the actual application scenario and the pre-setting of the SORT model.

[0084] In the multi-target tracking model, each moving object (target) in a current frame of video image is not only determined by target detection to determine the boundary of the moving object, but also records the predicted information of the moving object in a next frame of video image obtained after Kalman filtering as the inflation area of the moving object in the current frame of video image, for subsequent video image stitching.

[0085] Continuing in this embodiment, in the step S104, the device 100 stitches the synchronous frames of video image of each video segment based on the overlapping area of each frame of video image of each video segment, the boundary and inflation area of the moving object in the overlapping area, to obtain a large-angle video image after stitching.

[0086] The continuous frame large-view video images are outputted based on time domain, and the large-view video after splicing can be obtained.

[0087] Optionally, the step S104 comprises:

[0088] The best seam line between the synchronous frame video images of the corresponding video segments is determined based on the overlapping region, the boundary of the moving object in the overlapping region and the inflation region of each frame video image of each video segment.

[0089] The synchronous frame video images of the corresponding video segments are spliced based on the best seam line to obtain the large-view video image after splicing.

[0090] In the splicing of the synchronous frame video images of the corresponding video segments, the improved best seam line algorithm is adopted, and in the determination of the best seam line, the corresponding splicing strategy is adopted for the boundary of the moving object in the overlapping region and the inflation region and the other part of the overlapping region, so that the finally determined best seam line can avoid the moving object and no ghosting and obvious seam line is generated.

[0091] In an optional embodiment, the device 100 first acquires each video segment of the camera array, and acquires the SIFT image features of each frame video image of each video segment through the SIFT model; then the synchronous frame video images and the overlapping region of each frame video image in the same image space are obtained through the SIFT image feature matching and affine transformation through the FLANN model; then the overlapping region is cropped out and input into the SORT model, and the inflation region of the moving object can be determined through the Kalman filtering part of the SORT model, and the boundary of the moving object in the overlapping region is detected through the Kalman filtering and the Hungarian algorithm of the SORT model; the intensity of the pixel points on the boundary of the moving object and the inflation region is set differently through the improved best seam line algorithm, a pixel line composed of adjacent pixel points with the minimum intensity change across the overlapping region is determined as the best seam line, and the splicing of the synchronous frame video images of the corresponding video segments is completed.

[0092] Optionally, the determination of the best seam line between the synchronous frame video images of the corresponding video segments based on the overlapping region, the boundary of the moving object in the overlapping region and the inflation region of each frame video image of each video segment comprises:

[0093] The intensity value of the pixel point on the boundary of the moving object in the overlapping region of each frame video image of each video segment is set as a first preset intensity threshold, and the intensity value of the pixel point on the boundary of the inflation region of the moving object in the overlapping region of each frame video image of each video segment is set as a second preset intensity threshold.

[0094] set intensity values of the pixel points in the expansion region based on the first preset intensity threshold, the second preset intensity threshold and a preset smoothing function;

[0095] based on the intensity values of the pixel points in the overlapping region of the synchronous frame video images of the corresponding video segments, a pixel line composed of adjacent pixel points across the overlapping region and having the least intensity value change is determined as the optimal seam line between the synchronous frame video images of the corresponding video segments.

[0096] The existing video image stitching is usually performed by determining an optimal seam line composed of continuous adjacent pixel points having the least intensity value change in the overlapping region, and then performing video image stitching. The color difference of the pixel points in the region near the optimal seam line can be made to realize natural transition through image fusion. The optimal seam line thus determined cannot avoid the motion object, which affects the final stitching effect.

[0097] In this embodiment, in order to make the optimal seam line across the overlapping region avoid the motion object, the intensity values of the pixel points on the boundary of the motion object can be set to a large first intensity threshold after the boundary of the motion object in the video image is determined by the multi-target tracking model, so that the region surrounded by the boundary of the motion object can be bypassed when the pixel points constituting the optimal seam line are determined. The expansion region is a transition buffer region between the other part of the overlapping region and the boundary of the motion object, and the periphery of the expansion region is allowed to be crossed by the optimal seam line, while the part close to the boundary of the motion object should not be crossed by the optimal seam line. Therefore, the intensity values of the pixel points in the expansion region should be set to gradually decrease in the gradient direction from the boundary of the motion object to the boundary of the expansion region, so that the expansion region plays a buffering role.

[0098] In this embodiment, in order to make the optimal seam line across the overlapping region avoid the motion object, the intensity values of the pixel points on the boundary of the motion object can be set to a large first intensity threshold after the boundary of the motion object in the video image is determined by the multi-target tracking model, so that the region surrounded by the boundary of the motion object can be bypassed when the pixel points constituting the optimal seam line are determined. The expansion region is a transition buffer region between the other part of the overlapping region and the boundary of the motion object, and the periphery of the expansion region is allowed to be crossed by the optimal seam line, while the part close to the boundary of the motion object should not be crossed by the optimal seam line. Therefore, the intensity values of the pixel points in the expansion region should be set to gradually decrease in the gradient direction from the boundary of the motion object to the boundary of the expansion region, so that the expansion region plays a buffering role.

[0099] Optionally, the setting of the intensity values of the pixel points in the expansion region based on the first preset intensity threshold, the second preset intensity threshold and the preset smoothing function comprises:

[0100] set the intensity values of the pixel points on the region boundary corresponding to each prediction state in the expansion region based on the first preset intensity threshold, the second preset intensity threshold and a preset smoothing function, wherein the region boundary corresponding to a preset number of prediction states in the expansion region is marked with the motion object.

[0101] Based on the intensity values ​​of pixels on the boundary of the region corresponding to adjacent prediction states and the preset smoothing function, the intensity values ​​of each pixel in the region corresponding to the corresponding prediction state are set.

[0102] The number of predicted states within the dilation region can be preset according to the specific application scenario. After the multi-object tracking model determines the boundaries of moving objects and their dilation regions within the overlapping areas of the video images, the region boundaries corresponding to each predicted state within the dilation region can be clarified by combining the preset number of predicted states. For example, the preset dilation region may include N predicted states, S1 to S2. N E1 to E represent the boundary pixels corresponding to the predicted state. N S0 represents the intensity value of the boundary pixel corresponding to the predicted state; while S0 represents the boundary pixel of the moving object, E max The intensity value of the boundary pixels representing the moving object, i.e., the first preset intensity threshold; S N+1 E represents the pixel representing the boundary (outer boundary) of the expansion region. min The intensity value of the boundary pixels representing the expansion region, i.e., the second preset intensity threshold. To ensure that S0~S N+1 Intensity E0~E N+1 For a smooth transition, the following smoothing function can be preset:

[0103] E n =2 (N-n) ﹒ (E max -E min ) / (2 N -1)+(2 N E min -E max ) / (2 N -1)

[0104] Among them, E n It is the intensity value of the pixel on the boundary corresponding to the nth predicted state within the expansion region.

[0105] Since the boundaries between adjacent predicted states within the dilation region may include several pixels, to minimize the pixel intensity variation along the determined optimal stitching line, the intensity value of each pixel between the boundaries of adjacent predicted states can be set based on the preset smoothing function according to the gradient direction. Here, N is the number of pixels between the boundaries of adjacent predicted states based on the gradient direction, and E... max E is the intensity value of the pixel on the inner boundary corresponding to the adjacent predicted state. min E is the intensity value of the pixel on the outer boundary corresponding to the adjacent predicted state. nis the intensity value of the nth pixel point from inside to outside between the boundaries corresponding to the adjacent prediction states based on the gradient direction.

[0106] Figure 2 Fig. 1 shows a schematic diagram of an apparatus for video image stitching according to another aspect of the present application, wherein the apparatus comprises:

[0107] The first module 210 is configured to acquire a plurality of video segments, wherein each video segment comprises a plurality of continuous video images;

[0108] The second module 220 is configured to determine image features of each video image of each video segment based on each video segment, and determine overlapping areas of each video image of each video segment based on the image features of the synchronous video images of the corresponding video segment;

[0109] The third module 230 is configured to determine boundaries and inflation areas of moving objects in the overlapping areas of each video image of each video segment;

[0110] The fourth module 240 is configured to stitch the synchronous video images of the corresponding video segment based on the overlapping areas of each video image of each video segment, the boundaries and inflation areas of moving objects in the overlapping areas, to obtain a large-view-angle video image after stitching.

[0111] In this embodiment, the apparatus is integrated in the device 100. The device 100 has the same software and hardware environment as the device 100 in the foregoing method embodiments.

[0112] In this embodiment, the first module 210 of the apparatus can acquire a plurality of video segments captured by the camera array through real-time video acquisition, network transmission, offline cache copying, etc. Each video segment comprises a plurality of continuous video images. The synchronous video images of the corresponding video segment should comprise overlapping areas.

[0113] In this embodiment, the second module 220 of the apparatus processes each video segment acquired to determine image features of each video image of each video segment, and determines overlapping areas of each video image of each video segment according to the image features of each video image of each video segment.

[0114] The purpose of determining the image features of the video image is to determine the overlapping areas in the video image, mainly by comparing and matching the local image features of the video image to determine that there are the same local areas, i.e., the overlapping areas, in the relevant video images.

[0115] Since the different video clips captured by the relevant cameras in the camera array cover part of the same scene, the synchronous frame video images of the corresponding video clips include part of the same area, i.e., the overlapping area. The local image features in the overlapping area should be the same.

[0116] In this embodiment, the third module 230 of the device determines the boundary and the inflation area of the moving object in the overlapping area of each frame video image of each video clip.

[0117] If the video image in the overlapping area includes a moving object, the stitching line in the conventional splicing will pass through the moving object area, and the misalignment of the moving object will occur. In order to avoid such a situation, the boundary of the moving object in the overlapping area of the video image can be determined first, and the area corresponding to the subsequent motion state of the moving object, i.e., the inflation area, can be predicted, so as to avoid the moving object and its inflation area in the video image splicing.

[0118] In this embodiment, the fourth module 240 of the device splices the synchronous frame video images of the corresponding video clips based on the overlapping area of each frame video image of each video clip, the boundary of the moving object in the overlapping area, and the inflation area, to obtain the large-view-angle video image after splicing.

[0119] According to another aspect of the present application, a computer readable medium is also provided, which stores computer readable instructions executable by a processor to implement the foregoing method.

[0120] According to another aspect of the present application, a device for video image splicing is also provided, wherein the device comprises:

[0121] one or more processors; and

[0122] a memory storing computer readable instructions that, when executed, cause the processor to perform operations as in the foregoing method.

[0123] For example, the computer readable instructions, when executed, cause the one or more processors to: based on each video clip, determine the image features of each frame video image of each video clip, and based on the image features of the synchronous frame video images of the corresponding video clips, determine the overlapping area of each frame video image of each video clip; determine the boundary and the inflation area of the moving object in the overlapping area of each frame video image of each video clip; and splice the synchronous frame video images of the corresponding video clips based on the overlapping area of each frame video image of each video clip, the boundary of the moving object in the overlapping area, and the inflation area, to obtain the large-view-angle video image after splicing.

[0124] It will be obvious to a person skilled in the art that the application is not limited to the details of the above-described exemplary embodiments, but that the application can be implemented in other concrete forms without deviating from the spirit or the basic characteristics of the application. The embodiments are therefore to be considered in all respects as illustrative and not restrictive, the scope of the application being defined by the appended claims rather than by the above Description, which is therefore intended merely as a specification. All changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein. Any reference signs in the claims should not be construed as limiting the claim concerned. Furthermore, it is to be noted that the term "comprising" does not exclude other elements or steps, that the term "a" or "an" does not exclude a plurality, and that a single processor or other unit can fulfil the functions of several units recited in the claims. The terms first, second and the like do not denote any ordering, but rather are used as names for naming different units.

Claims

1. A method for video image stitching, characterized in that, The method comprises: acquiring a plurality of video clips, wherein each video clip comprises a plurality of frames of continuous video images; based on each video clip, determining the image features of each frame of video image of each video clip, and based on the image features of the synchronous frame of video image of the corresponding video clip, determining the overlapping area of each frame of video image of each video clip; determining the boundary and the inflation area of the moving object in the overlapping area of each frame of video image of each video clip, wherein the inflation area is the area corresponding to the subsequent motion state of the predicted moving object in the overlapping area, and the inflation area is marked with the area boundary corresponding to the preset number of prediction states of the moving object; setting the intensity value of the pixel point on the boundary of the moving object in the overlapping area of each frame of video image of each video clip to a first preset intensity threshold, and setting the intensity value of the pixel point on the boundary box of the inflation area of the moving object in the overlapping area of each frame of video image of each video clip to a second preset intensity threshold, setting the intensity value of each pixel point in the inflation area based on the first preset intensity threshold, the second preset intensity threshold and the preset smoothing function, based on the intensity value of each pixel point in the overlapping area of the synchronous frame of video image of the corresponding video clip, determining the pixel line composed of adjacent pixel points with the minimum change of intensity value across the overlapping area as the best stitching line between the synchronous frames of video image of the corresponding video clip, and based on the best stitching line, splicing the synchronous frames of video image of the corresponding video clip to obtain the large-view-angle video image after splicing.

2. The method of claim 1, wherein, The method comprises: inputting each frame of video image of each video clip into a SIFT model to obtain the SIFT features of each frame of video image of each video clip.

3. The method of claim 2, wherein, The method comprises: inputting the SIFT features of the synchronous frames of video image of the corresponding video clip into an image feature matching model to determine the overlapping area of each frame of video image of each video clip.

4. The method of claim 3, wherein, The image feature matching model comprises a FLANN model.

5. The method of claim 1, wherein, The method comprises: cropping the overlapping area of each frame of video image of each video clip and inputting it into a multi-target tracking model to determine the boundary and the inflation area of the moving object in the overlapping area of each frame of video image of each video clip.

6. The method of claim 5, wherein, The multi-target tracking model comprises a SORT model.

7. The method of claim 1, wherein, The method comprises: based on the first preset intensity threshold, the second preset intensity threshold and the preset smoothing function, setting the intensity value of the pixel point on the boundary of each prediction state corresponding to the area boundary in the inflation area, wherein the inflation area is marked with the area boundary corresponding to the preset number of prediction states of the moving object; The intensity values of the pixels in the expansion region are set based on the intensity values of the pixels on the region boundary corresponding to the adjacent prediction state and the preset smoothing function.

8. An apparatus for video image stitching, the apparatus comprising: The apparatus comprises: A first module configured to acquire a plurality of video segments, wherein each video segment comprises a plurality of consecutive video images; A second module configured to determine, based on each video segment, image features of each video image of the video segment, and determine, based on the image features of the synchronous video images of the corresponding video segment, an overlap region of each video image of each video segment; A third module configured to determine a boundary of a moving object and an expansion region in the overlap region of each video image of each video segment, wherein the expansion region is a region corresponding to a subsequent motion state of the predicted moving object in the overlap region, and the expansion region is marked with a preset number of region boundaries corresponding to the prediction states of the moving object; A fourth module configured to set the intensity values of the pixels on the boundary of the moving object in the overlap region of each video image of each video segment to a first preset intensity threshold, and set the intensity values of the pixels on the boundary frame of the expansion region of the moving object in the overlap region of each video image of each video segment to a second preset intensity threshold, set the intensity values of the pixels in the expansion region based on the first preset intensity threshold, the second preset intensity threshold, and a preset smoothing function, determine a best stitching line between the synchronous video images of the corresponding video segment based on the intensity values of the pixels in the overlap region of the synchronous video images of the corresponding video segment, the best stitching line being composed of adjacent pixels that span the overlap region and have the smallest change in intensity value, and stitch the synchronous video images of the corresponding video segment based on the best stitching line to obtain a large-view-angle video image after stitching.

9. A computer-readable medium, comprising: computer-readable instructions stored thereon, the computer-readable instructions being executed by a processor to implement the method of any one of claims 1-7.

10. An apparatus for video image stitching, the apparatus comprising: The device comprises: one or more processors; and a memory storing computer-readable instructions that, when executed, cause the processor to perform the operations of the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Image stitching method and equipment

    CN107038686A

  • Video image mosaic method, device, terminal device and storage medium

    CN109146832A