Video motion vector determination method and device, equipment and storage medium
By splicing and optical flow calculation of the frame images of multiple videos, the motion vectors of multiple videos are determined, which solves the problem that the motion vector of multiple videos cannot be estimated in the prior art, and realizes the convenience and accuracy of global stability analysis of multiple videos.
Patent Information
- Application Number
- CN202410084695.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-19
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art cannot estimate the motion vector of the complete content of multiple videos at the same time, resulting in the inability to conveniently and quickly perform global stability comparison analysis of videos captured by multiple photography devices.
By splicing the frame video images of multiple videos, the forward and backward optical flows of the feature points of each adjacent two frames of video images in the spliced video are calculated, and the average displacement amount of the matching feature point pairs in each adjacent two frames of video images is determined, and the motion vectors of each of the multiple videos are determined based on the average displacement amount.
It realizes one-time determination of the complete content motion vectors of multiple videos, and conducts global stability comparison and analysis conveniently and quickly to meet the needs of equipment performance analysis and application research and development.
Smart Images

Figure CN120358319A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of the Internet of Things, and particularly to a method, apparatus, device, and storage medium for determining video motion vectors. Background Art
[0002] Video motion vector estimation is one of the important technologies in the field of video processing, and can be used to determine the motion vector of a video on a two-dimensional plane from adjacent frames of a video image sequence. This technology has a wide range of applications in many fields, such as video tracking, driverless driving, and video compression.
[0003] In related technologies, motion vectors are usually estimated only for the main contents such as people and vehicles in a single video, and it is impossible to simultaneously estimate the motion vectors of the complete contents of multiple videos, and thus it is impossible to conveniently and quickly perform global stability comparison and analysis on the videos captured by multiple photographic devices, that is, it cannot meet the requirements of scenarios such as device performance analysis and application research and development. Summary of the Invention
[0004] To overcome the problems existing in related technologies, embodiments of the present disclosure provide a method, apparatus, device, and storage medium for determining video motion vectors to solve the defects in related technologies.
[0005] According to a first aspect of an embodiment of the present disclosure, a method for determining a video motion vector is provided, and the method includes:
[0006] In response to obtaining a plurality of videos for which motion vectors are to be determined, splicing each frame of video image of the plurality of videos to obtain a spliced video, and each frame of video image of the spliced video includes the video images of the corresponding frames of the plurality of videos;
[0007] Calculating forward optical flow and backward optical flow for feature points of every two adjacent frames of video images in the spliced video to obtain a pair of matching feature points for every two adjacent frames of video images;
[0008] Determining an average displacement amount of the pair of matching feature points of each video image in every two adjacent frames of video images;
[0009] Determining the motion vectors of the plurality of videos respectively based on the average displacement amount.
[0010] In some embodiments, the method further includes:
[0011] In response to detecting a target subject in each video image of each frame of video image, dividing each video image into a main body region and a background region, and the main body region is the region where the target subject is located;
[0012] Determining the average displacement amount of the matching feature point pairs of each video frame in each adjacent two video frames includes:
[0013] Determining the average displacement amount of the matching feature point pairs in a specified area in each video frame, where the specified area is the main area and / or the background area;
[0014] Based on the average displacement amount, determining the motion vectors of each of the multiple videos includes:
[0015] Based on the average displacement amount of the matching feature point pairs in the specified area, determining the motion vectors of each of the multiple videos.
[0016] In some embodiments, the feature points of each adjacent two video frames include the feature points in the valid area of each adjacent two video frames.
[0017] In some embodiments, the method further includes pre-determining the valid area based on the following method:
[0018] Performing binarization processing on a first video image selected from the spliced video to obtain a binary image;
[0019] Finding the contours of each foreground in the binary image;
[0020] Among the found contours of each foreground, removing the contours with an enclosed area less than or equal to the pixel quantity threshold to obtain remaining contours;
[0021] Based on the remaining contours, determining the first valid area of each video frame in the first video image;
[0022] Determining the area corresponding to the first valid area in each adjacent two video frames as the valid area of each adjacent two video frames.
[0023] In some embodiments, based on the remaining contours, determining the first valid area of each video frame in the first video image includes:
[0024] Determining the first area in each video frame that contains the remaining contours;
[0025] Reducing the first area by a preset ratio to obtain the first valid area of each video frame.
[0026] In some embodiments, before finding the contours of each foreground in the binary image, it further includes:
[0027] Eliminating the noise and / or holes in the binary image through morphological transformation.
[0028] In some embodiments, calculating the forward optical flow and backward optical flow for the feature points of every two adjacent video images in the spliced video to obtain the matching feature point pairs of every two adjacent video images includes:
[0029] Determining a second feature point in the latter video image of every two adjacent video images from a first feature point in the former video image of every two adjacent video images;
[0030] Determining a third feature point in the former video image from the second feature point;
[0031] Determining the pixel error between the first feature point and the third feature point;
[0032] Deleting the feature point pairs in every two adjacent video images where the pixel error is greater than or equal to a threshold to obtain the matching feature point pairs of every two adjacent video images.
[0033] In some embodiments, determining the average displacement amount of the matching feature point pairs of each video frame within every two adjacent video images includes:
[0034] Determining the sum of the displacement amounts of the matching feature point pairs of every two adjacent video images in the spliced video;
[0035] Determining the average displacement amount based on the sum of the displacement amounts and the number of the matching feature point pairs of every two adjacent video images.
[0036] In some embodiments, the method further includes:
[0037] Inputting the motion vectors of the multiple videos into a preset application program for differential display.
[0038] According to a second aspect of the embodiments of the present disclosure, there is provided a video motion vector determination device, the device includes:
[0039] A video acquisition module, configured to splice each video frame of the multiple videos in response to acquiring the multiple videos for which motion vectors are to be determined, to obtain a spliced video, and each video image of the spliced video includes the video frames corresponding to the multiple videos;
[0040] A point pair determination module, configured to calculate the forward optical flow and backward optical flow for the feature points of every two adjacent video images in the spliced video to obtain the matching feature point pairs of every two adjacent video images;
[0041] A displacement determination module, configured to determine the average displacement amount of the matching feature point pairs of each video frame within every two adjacent video images;
[0042] A vector determination module, configured to determine motion vectors of each of the multiple videos based on the average displacement amount.
[0043] In some embodiments, the apparatus further includes:
[0044] A screen division module, configured to divide each of the video screens into a main body area and a background area in response to detecting a target subject in each video screen within each frame of video image, where the main body area is the area where the target subject is located;
[0045] The displacement determination module is further configured to determine an average displacement amount of matching feature point pairs in a specified area in each of the video screens, where the specified area is the main body area and / or the background area;
[0046] The vector determination module is further configured to determine motion vectors of each of the multiple videos based on the average displacement amount of the matching feature point pairs in the specified area.
[0047] In some embodiments, the feature points of every two adjacent frames of video images include the feature points within the valid area of every two adjacent frames of video images.
[0048] In some embodiments, the apparatus further includes a region determination module;
[0049] The region determination module includes:
[0050] A binary processing unit, configured to perform binary processing on a first video image selected from the spliced video to obtain a binary image;
[0051] A contour search unit, configured to search for contours of each foreground in the binary image;
[0052] A contour removal unit, configured to remove contours with a surrounding area less than or equal to a pixel quantity threshold from the searched contours of each foreground to obtain remaining contours;
[0053] A first determination unit, configured to determine a first valid area of each video screen in the first video image based on the remaining contours;
[0054] A region determination unit, configured to determine the area corresponding to the first valid area in every two adjacent frames of video images as the valid area of every two adjacent frames of video images.
[0055] In some embodiments, the first determination unit is further configured to:
[0056] Determine a first area in each of the video screens that contains the remaining contours;
[0057] Reduce the first region by a preset ratio to obtain the first effective region of each video frame.
[0058] In some embodiments, the region determination module further includes:
[0059] A morphological transformation unit for eliminating noise and / or holes in the binary image through morphological transformation.
[0060] In some embodiments, the point pair determination module includes:
[0061] A forward calculation unit for determining the second feature points of the latter video frame in each adjacent pair of video frames from the first feature points of the former video frame in each adjacent pair of video frames;
[0062] A backward calculation unit for determining the third feature points of the former video frame from the second feature points;
[0063] An error determination unit for determining the pixel error between the first feature points and the third feature points;
[0064] A point pair determination unit for deleting the feature point pairs with pixel errors greater than or equal to a threshold in each adjacent pair of video frames to obtain the matching feature point pairs of each adjacent pair of video frames.
[0065] In some embodiments, the displacement determination module includes:
[0066] A sum determination unit for determining the sum of the displacement amounts of the matching feature point pairs of each adjacent pair of video frames in the spliced video;
[0067] A displacement determination unit for determining the average displacement amount based on the sum of the displacement amounts and the number of the matching feature point pairs of each adjacent pair of video frames.
[0068] In some embodiments, the apparatus further includes:
[0069] A visualization module for inputting the motion vectors of the multiple videos into a preset application program for differential display.
[0070] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, the device includes:
[0071] A processor and a memory for storing a computer program;
[0072] Wherein, the processor is configured to implement:
[0073] In response to obtaining a plurality of videos for which motion vectors are to be determined, stitching each frame of video images of the plurality of videos to obtain a stitched video, where each frame of video image of the stitched video includes the video images of the corresponding frames of the plurality of videos;
[0074] Calculating forward optical flow and backward optical flow for feature points of every two adjacent frames of video images in the stitched video to obtain matching feature point pairs of every two adjacent frames of video images;
[0075] Determining the average displacement amount of the matching feature point pairs of each video image within every two adjacent frames of video images;
[0076] Determining the respective motion vectors of the plurality of videos based on the average displacement amount.
[0077] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the following is implemented:
[0078] In response to obtaining a plurality of videos for which motion vectors are to be determined, stitching each frame of video images of the plurality of videos to obtain a stitched video, where each frame of video image of the stitched video includes the video images of the corresponding frames of the plurality of videos;
[0079] Calculating forward optical flow and backward optical flow for feature points of every two adjacent frames of video images in the stitched video to obtain matching feature point pairs of every two adjacent frames of video images;
[0080] Determining the average displacement amount of the matching feature point pairs of each video image within every two adjacent frames of video images;
[0081] Determining the respective motion vectors of the plurality of videos based on the average displacement amount.
[0082] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0083] The present disclosure obtains a plurality of videos for which motion vectors are to be determined, stitches the video frames of each of the plurality of videos to obtain a stitched video, where each video image frame of the stitched video includes the video frames corresponding to the plurality of videos, calculates the forward optical flow and backward optical flow for the feature points of every two adjacent video images in the stitched video to obtain the matching feature point pairs of every two adjacent video images, then determines the average displacement amount of the matching feature point pairs of each video frame within every two adjacent video images, and further determines the motion vectors of each of the plurality of videos based on the average displacement amount. This can achieve the determination of the motion vectors of the complete content of a plurality of videos at one time, which is beneficial for subsequent convenient and rapid comparative analysis of the global stability of the videos captured by multiple photographic devices, thereby meeting the requirements of scenarios such as device performance analysis and application research and development.
[0084] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. Brief Description of the Drawings
[0085] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0086] Figure 1 is a flowchart of a method for determining a video motion vector according to an exemplary embodiment of the present disclosure;
[0087] Figure 2A is a flowchart of a method for determining a video motion vector according to another exemplary embodiment of the present disclosure;
[0088] Figure 2B is a schematic diagram of a determination method for a human body region according to an exemplary embodiment of the present disclosure;
[0089] Figure 3A is a flowchart of how to determine the valid region of every two adjacent video images according to an exemplary embodiment of the present disclosure;
[0090] Figure 3B is a schematic diagram of a determination method for the valid region of every two adjacent video images according to an exemplary embodiment of the present disclosure;
[0091] Figure 4 is a flowchart of how to determine the first valid region of each video frame in the first video image based on the remaining contour according to an exemplary embodiment of the present disclosure;
[0092] Figure 5A is a flowchart of how to determine the matching feature point pairs of every two adjacent video images according to an exemplary embodiment of the present disclosure;
[0093] Figure 5B It is a schematic process diagram showing the matching feature point pairs of every two adjacent video images according to an exemplary embodiment of the present disclosure;
[0094] Figure 6A It is a flowchart of a video motion vector determination method according to another exemplary embodiment of the present disclosure;
[0095] Figure 6B It is a schematic diagram of the display interface of an interactive visualization tool according to an exemplary embodiment of the present disclosure;
[0096] Figure 6C It is a schematic diagram of the enlarged effect of the ROI option according to an exemplary embodiment of the present disclosure;
[0097] Figure 7 It is a block diagram of a video motion vector determination device according to an exemplary embodiment of the present disclosure;
[0098] Figure 8 It is a block diagram of another video motion vector determination device according to an exemplary embodiment of the present disclosure;
[0099] Figure 9 It is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0100] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0101] Figure 1 It is a flowchart of a video motion vector determination method according to an exemplary embodiment; the method of this embodiment can be executed by a video motion vector determination device, and the video motion vector determination device can be configured in an electronic device with video data processing functions, such as a server, a workstation, a personal computer, a mobile terminal (such as a mobile phone, a tablet computer, etc.), a wearable device (such as glasses, a watch, etc.), etc. As Figure 1 shown, the method includes the following steps S101 - S104:
[0102] In step S101, in response to obtaining a plurality of videos for which motion vectors are to be determined, each frame of the video pictures of the plurality of videos is spliced to obtain a spliced video.
[0103] In this embodiment, when the electronic device acquires multiple videos for which motion vectors are to be determined, it can splice each frame of video images of the multiple videos to obtain a spliced video.
[0104] For example, the above-mentioned multiple videos can be videos taken by multiple photographic devices that currently need to perform video stability comparative analysis, etc.
[0105] Each frame of video image of the above-mentioned spliced video can include the video images of the corresponding frames of the multiple videos.
[0106] As an example, the Nth (N≥1) frame video images of the multiple videos can be spliced into the Nth frame video image of the above-mentioned spliced video. That is to say, the Nth frame video image of the above-mentioned spliced video includes the Nth frame video images of each of the multiple videos. Further, in order to realize subsequent separate processing of different video images in each frame of the spliced video, a preset interval area, such as an interval black edge, etc., can be set between the multiple video images in each frame of the spliced video. This embodiment does not limit this.
[0107] In step S102, forward optical flow and backward optical flow are calculated for the feature points of every two adjacent frames of video images in the spliced video to obtain the matching feature point pairs of every two adjacent frames of video images.
[0108] In this embodiment, when each frame of video images of the multiple videos is spliced to obtain a spliced video, forward optical flow and backward optical flow can be calculated for the feature points of every two adjacent frames of video images in the spliced video to obtain the matching feature point pairs of every two adjacent frames of video images. It can be understood that the matching feature point pairs of every two adjacent frames of video images can include the matching feature point pairs of every two adjacent frames of video images of each of the multiple videos.
[0109] For example, assuming that the matching feature point pairs of every two adjacent frames of video images in the spliced video include point pair 1, point pair 2, point pair 3, and point pair 4, then point pair 1 can belong to the matching feature point pairs of every two adjacent frames of video images of the first video among the multiple videos, point pair 2 can belong to the matching feature point pairs of every two adjacent frames of video images of the second video among the multiple videos, point pair 3 can belong to the matching feature point pairs of every two adjacent frames of video images of the third video among the multiple videos, and point pair 4 can belong to the matching feature point pairs of every two adjacent frames of video images of the fourth video among the multiple videos.
[0110] In some embodiments, in order to improve the accuracy of determining the video motion vector, the effective region of each video frame in the video image may be determined in advance (e.g., the field of view FOV of the video), and then the feature points of every two adjacent video images may be set as the feature points within the effective region. That is to say, the feature points considered in this embodiment may be the feature points within the effective region of each video frame. Among them, the determination method of the above effective region may be selected from related technologies according to the scene requirements, and this embodiment does not limit this. In some other embodiments, reference may also be made to the following Figure 3A illustrated embodiment, which will not be elaborated here for the time being.
[0111] As an example, the method of calculating the forward optical flow and the backward optical flow for the feature points of every two adjacent video images in the spliced video to obtain the matching feature point pairs of every two adjacent video images may refer to the following Figure 5A illustrated embodiment, which will not be elaborated here for the time being.
[0112] In step S103, determine the average displacement of the matching feature point pairs of each video frame within every two adjacent video images.
[0113] In this embodiment, after obtaining the matching feature point pairs of every two adjacent video images, the average displacement of the matching feature point pairs of each video frame within every two adjacent video images may be determined.
[0114] In some embodiments, the sum of the displacements of the matching feature point pairs of every two adjacent video images in the spliced video may be determined, and then the average displacement may be determined based on the sum of the displacements and the number of the matching feature point pairs of every two adjacent video images.
[0115] For example, assume that the matching feature point pairs of the i-th group of adjacent two video images are prevpts i (x, y) and nextpts i (x, y), where prevpts i (x, y) is the feature point of the previous video image in the i-th group of adjacent two video images, and nextpts i (x, y) is the feature point of the subsequent video image in the i-th group of adjacent two video images. Then, the displacement (dx, dy) of the matching feature point pairs of each video frame within the i-th group of adjacent two video images may be determined based on the following formula (1-1) i :
[0116] (dx, dy) i = prevpts i (x, y) - nextpts i (x, y); (1-1)
[0117] On this basis, the average displacement amount mean in the x-axis direction of the matching feature point pairs of each video frame in every two adjacent video images can be determined respectively based on the following formulas (1-2) and (1-3) x and the average displacement amount mean in the y-axis direction y :
[0118]
[0119]
[0120] where N is the number of matching feature point pairs of every two adjacent video images
[0121] In step S104, the motion vectors of the respective multiple videos are determined based on the average displacement amount
[0122] In this embodiment, after determining the average displacement amount of the matching feature point pairs of each video frame in every two adjacent video images, the motion vectors of the respective multiple videos can be determined based on the average displacement amount
[0123] In some embodiments, to ensure the accuracy of determining the motion vectors of the respective multiple videos, when determining the average displacement amount mean in the above x-axis direction x and the average displacement amount mean in the y-axis direction y if it is detected that the number of matching feature point pairs of each video frame in two adjacent video images (hereinafter simply referred to as "the number of feature point pairs") is greater than or equal to the number threshold, the differences between dx, dy and the corresponding means can be sorted from small to large to retain the feature point pairs in the top set proportion (named "FeatureRatio") closest to the corresponding means, so as to recalculate the average displacement amount based on the retained feature point pairs, and then the mean x and mean y of each video frame in every two adjacent video images of the spliced video can be output, and thus the motion vectors of the respective multiple videos can be obtained. As an example, FeatureRatio in this embodiment can be set to 0.8
[0124] It can be understood that if it is detected that the number of feature point pairs is less than the above number threshold, the calculation results of the average displacement amounts of all feature point pairs can be retained, and then the mean x and mean y of each video frame in every two adjacent video images of the spliced video can be output, and thus the motion vectors of the respective multiple videos can be obtained
[0125] As can be seen from the above description, the method of this embodiment splices each frame of video pictures of the multiple videos in response to obtaining the multiple videos for which the motion vectors are to be determined, to obtain a spliced video. Each frame of video image of the spliced video includes the video pictures of the corresponding frames of the multiple videos, and calculates the forward optical flow and the backward optical flow for the feature points of every two adjacent frames of video images in the spliced video, to obtain the matching feature point pairs of every two adjacent frames of video images. Then, the average displacement amount of the matching feature point pairs of each video picture within every two adjacent frames of video images is determined, and further, the motion vectors of the multiple videos are determined based on the average displacement amount, which can achieve the determination of the motion vectors of the complete contents of the multiple videos at one time, facilitating the convenient and fast comparative analysis of the global stability of the videos captured by multiple photographic devices, and thus can meet the requirements of scenarios such as device performance analysis and application research and development.
[0126] Figure 2A FIG. 4 is a flowchart of a method for determining video motion vectors according to another exemplary embodiment of the present disclosure; the method of this embodiment can be executed by a video motion vector determination device, and the video motion vector determination device can be configured in an electronic device with video data processing functions, such as a server, a workstation, a personal computer, a mobile terminal (such as a mobile phone, a tablet computer, etc.), a wearable device (such as glasses, a watch, etc.). As Figure 2A shown, the method includes the following steps S201-S207:
[0127] In step S201, in response to obtaining multiple videos for which the motion vectors are to be determined, each frame of video pictures of the multiple videos is spliced to obtain a spliced video.
[0128] Each frame of video image of the above spliced video may include the video pictures of the corresponding frames of the multiple videos.
[0129] In step S202, the forward optical flow and the backward optical flow are calculated for the feature points of every two adjacent frames of video images in the spliced video, to obtain the matching feature point pairs of every two adjacent frames of video images.
[0130] In step S203, it is determined whether a target subject is detected in each video picture within each frame of video image: if yes, step S204 is executed; if not, step S206 is executed.
[0131] In step S204, each video picture is divided into a subject area and a background area. Among them, the subject area is the area where the target subject is located.
[0132] In step S205, the average displacement amount of the matching feature point pairs within a specified area in each video picture is determined, and the specified area is the subject area and / or the background area.
[0133] In step S206, determine the average displacement amount of the matching feature point pairs of each video frame within every two adjacent video images.
[0134] Among them, for the relevant explanations and descriptions of the above steps S201 - S202 and S206, reference can be made to the steps S101 - S103 in the above Figure 1 illustrated embodiments, which will not be elaborated here.
[0135] In step S207, determine the motion vectors of the respective multiple videos based on the average displacement amount.
[0136] Based on the above Figure 1 illustrated embodiments, in order to improve the flexibility and diversity of determining the video motion vectors in this embodiment, it is also possible to determine the motion vectors of local regions of multiple videos, so as to improve the speed of subsequent comparative analysis of the stability of local regions of multiple videos. As an example, in this embodiment, it can be determined whether a target object is detected in each video frame within each video image. Among them, the target object may include set objects such as a human body, an animal, a tree, a building, a statue, a billboard, etc., and this embodiment does not limit this.
[0137] Taking the target object as a human body as an example, it is possible to detect whether there is a target object in each video frame within each video image based on an image recognition algorithm in the related art. For example, a preset face recognition algorithm can be used to detect the face region in each video frame to obtain the feature points of the face region; and then, based on the feature points of the face region, expansion in multiple directions such as up, down, left, and right can be performed to obtain the human body region, which is used as the above-mentioned main body region. For example, Figure 2B is a schematic diagram of a method for determining a human body region shown in an exemplary embodiment of the present disclosure; as Figure 2B shown, after obtaining the face region (a rectangular region with width w and height h in the figure), it is possible to expand 0.05*w in width on both the left and right sides, and expand 0.4*h and 0.8*h in height on the upper and lower sides respectively. In addition, regions lower than the bottom edge of the current expanded region can also be divided into the human body region, and thus the final human body region (i.e., Figure 2B the region enclosed by the lower part of the dashed line in the figure) can be obtained.
[0138] On this basis, each of the video frames can be divided into a main region and a background region (i.e., the region other than the main region in the video frame), and the average displacement of the matching feature point pairs within a specified region in each of the video frames is determined. Herein, the specified region may be the above-mentioned main region and / or background region. It can be understood that when the specified region is the main region and the background region, this embodiment can achieve the determination of the video motion vectors of the main region and the background region respectively, so that subsequent comparison and analysis can be performed based on the video motion vectors of the main region and the background region of multiple videos.
[0139] Further, after determining the average displacement of the matching feature point pairs within the above-mentioned specified region, the motion vectors of the respective specified regions of the multiple videos can be determined based on this average displacement. This can facilitate subsequent different device manufacturers to formulate different video image stabilization strategies (such as formulating a stabilization strategy for the local stability or global stability of the video image, etc.).
[0140] For example, if the manufacturer pays more attention to the stability of the background region of the video image, this embodiment can be used to determine the average displacement of the matching feature point pairs within the background region of each video frame, and then determine the motion vectors of each video; if the manufacturer pays more attention to the stability of the global region of the video image, then Figure 1 the illustrated embodiment can be used to determine the average displacement of the matching feature point pairs within the "foreground + background" region of each video frame, and then determine the motion vectors of each video.
[0141] Figure 3A is a flowchart showing how to determine the effective region of each adjacent two video frames according to an exemplary embodiment of the present disclosure; this embodiment takes how to determine the effective region of each adjacent two video frames as an example for exemplary illustration on the basis of the above embodiment.
[0142] As Figure 3A shown, the video motion vector determination method of this embodiment may further include determining the effective region of each adjacent two video frames based on the following steps S301 - S305:
[0143] In step S301, the first video image selected from the spliced video is binarized to obtain a binary image.
[0144] In this embodiment, in order to determine the effective region of each adjacent two video frames, the effective region of the first video image selected from the spliced video (such as the first frame video image of the spliced video, etc.) can be determined first, and then according to the effective region of the first video image, the effective region of each video frame in the spliced video is determined, that is, the effective region of each adjacent two video frames in the spliced video can be obtained.
[0145] Therefore, in this embodiment, the first video image selected from the spliced video can be binarized to obtain a binary image. For example, the RGB values of the first video image can be converted into a grayscale image of a preset size (such as 8 bits, etc.), and then, based on a specified fixed threshold (Binary Thresholding), the pixels in the image can be divided into two categories. As an example, the fixed threshold can be selected as 3. Furthermore, the pixel grayscales in the grayscale image that are less than or equal to the above fixed threshold can be set to 0, while the pixel grayscales that are greater than the above fixed threshold can be set to 255, thereby realizing the binarization of the first video image and obtaining a binary image.
[0146] In some embodiments, in order to make the edges of the binarized image smoother and clearer, before finding the contours of each foreground in the binary image, the noise and / or holes in the binary image can be eliminated through morphological transformation. For example, the area of the binary image can be reduced by eliminating boundary points based on an erosion operation, and / or the boundary points of the binary image can be expanded based on a dilation operation.
[0147] In step S302, find the contours of each foreground in the binary image.
[0148] In this embodiment, after obtaining the binary image, the contours of each foreground can be found in the binary image.
[0149] In step S303, among the found contours of each foreground, remove the contours whose enclosed area is less than or equal to the pixel quantity threshold to obtain the remaining contours.
[0150] In this embodiment, after finding the contours of each foreground in the binary image, among the found contours of each foreground, remove the contours whose enclosed area is less than or equal to the pixel quantity threshold to obtain the remaining contours.
[0151] Among them, the above pixel quantity threshold can be freely set according to the requirements of the application scenario, such as being set to 5% of the total number of image pixels, etc. This embodiment does not limit this.
[0152] In step S304, determine the first effective area of each video frame in the first video image based on the remaining contours.
[0153] In this embodiment, after obtaining the remaining contours, the first effective area of each video frame in the first video image can be determined based on the remaining contours. For example, the area in each video frame that contains the remaining contours can be determined as the first effective area.
[0154] For example, Figure 3BIt is a schematic diagram of a method for determining the valid region of every two adjacent video images shown according to an exemplary embodiment of the present disclosure; as Figure 3B shown, the above first video image contains 4 video frames from Video 1 to Video 4, and the interval black border between two adjacent video frames is the preset interval region. On this basis, the first video image can be first binarized to obtain Figure 3B the binary image shown in the lower left corner of Figure 3B and then through morphological transformation operations such as erosion and dilation to obtain Figure 3B the image shown in the lower right corner of
[0155] and then through the search and screening of the foreground contours, the valid regions in each video frame in Figure 4 the upper right corner of
[0156] are obtained (that is, the regions within the white rectangular frames).
[0157] In this embodiment, after determining the first valid region of each video frame in the first video image based on the remaining contours, the regions in every two adjacent video images corresponding to the first valid region can be determined as the valid regions of every two adjacent video images.
[0158] For example, after determining the first valid region of the m-th video frame in the first video image, the region in the m-th video frame of other video images that is the same as the first valid region can be determined as the valid region of the m-th video frame in these other video images. That is to say, the valid regions of the video frames of the same video are the same.
[0159] As can be seen from the above description, in this embodiment, by performing binarization processing on the first video image selected in the spliced video to obtain a binary image, and searching for the contours of each foreground in the binary image, and then removing the contours with an enclosed area less than or equal to the pixel quantity threshold among the searched foreground contours to obtain the remaining contours, and determining the first valid region of each video frame in the first video image based on the remaining contours, and further determining the regions in every two adjacent video images corresponding to the first valid region as the valid regions of every two adjacent video images, it is possible to reasonably and accurately determine the valid regions of every two adjacent video images in the spliced video.
[0160] Figure 4It is a flowchart showing how to determine the first valid region of each video frame in the first video image based on the remaining contour according to an exemplary embodiment of the present disclosure; this embodiment takes how to determine the first valid region of each video frame in the first video image based on the remaining contour as an example for exemplary illustration on the basis of the above embodiment.
[0161] As Figure 4 shown, the determination of the first valid region of each video frame in the first video image based on the remaining contour in step S304 described above may include the following steps S401 - S402:
[0162] In step S401, determine the first region in each video frame that contains the remaining contour.
[0163] In this embodiment, when removing the contours with an enclosed area less than or equal to the pixel quantity threshold from the contours of each foreground found and obtaining the remaining contour, the first region in each video frame that contains the remaining contour can be determined. That is to say, for each frame of video image, the region containing the remaining contour in each of its video frames can be determined to obtain the above - mentioned first region.
[0164] In step S402, reduce the first region by a preset ratio to obtain the first valid region of each video frame.
[0165] In this embodiment, when the first region in each video frame that contains the remaining contour is determined, the first region can be reduced by a preset ratio (hereinafter referred to as "Field - of - View Ratio FOV Ratio") to obtain the first valid region of each video frame.
[0166] It is worth noting that this embodiment takes into account the feature points that are prone to detection errors at the image edge (for example, misidentifying the points on the preset interval region between two video frames as feature points). Therefore, by reducing the first region of each video frame in this embodiment, the edge regions of each video frame can be appropriately cut off, and further, the detection error of the feature points at the edge of each video frame can be reduced.
[0167] Among them, the value of the above - mentioned FOV Ratio can be freely set according to application requirements, such as set to 0.9, etc., and this embodiment does not limit this.
[0168] From the above description, it can be seen that in this embodiment, by determining the first region in each video frame that contains the remaining contour and reducing the first region by a preset ratio to obtain the first valid region of each video frame, the detection error of the feature points at the edge of the frame can be reduced by reducing the contour, and further, the accuracy of determining the subsequent video motion vector can be improved.
[0169] Figure 5A It is a flowchart showing how to determine the matching feature point pairs of each adjacent two-frame video images according to an exemplary embodiment of the present disclosure; in this embodiment, on the basis of the above embodiment, taking how to determine the matching feature point pairs of each adjacent two-frame video images as an example for exemplary illustration.
[0170] As Figure 5A shown, the determination of the matching feature point pairs of each adjacent two-frame video images in the above step S102 may include the following steps S501-S504:
[0171] In step S501, the second feature points of the latter frame video image in each adjacent two-frame video images are determined from the first feature points of the former frame video image in each adjacent two-frame video images;
[0172] In step S502, the third feature points of the former frame video image are determined from the second feature points;
[0173] In step S503, the pixel error between the first feature points and the third feature points is determined;
[0174] In step S504, the feature point pairs in each adjacent two-frame video images with the pixel error greater than or equal to the threshold are deleted to obtain the matching feature point pairs of each adjacent two-frame video images.
[0175] It should be noted that the optical flow method can be used to find the correspondence between the previous frame and the next frame (i.e., the previous frame and the next frame in two adjacent frames of images) by using the change of pixels in the time domain in the image sequence and the correlation between adjacent frames, so as to calculate the motion information of the object between adjacent frames. Among them, the instantaneous change rate of gray scale at a specific coordinate point on the two-dimensional image plane can be defined as the optical flow vector. Forward optical flow refers to the pixel motion information from the current frame t to the next frame t+1, and backward optical flow is the pixel motion information from the current frame t to the previous frame t-1. Generally speaking, forward optical flow can be regarded as predicting the possible position of the object in the next frame image, while backward optical flow can be regarded as inferring the past state according to the motion that has occurred.
[0176] In this embodiment, by calculating the forward optical flow and the backward optical flow, more accurate results can be obtained to break through the limitations of using single-direction optical flow calculation. That is to say, if only forward optical flow calculation is used to determine the matching feature point pairs, due to the inability to obtain past information, the adaptability to dynamic environments is poor; similarly, if only backward optical flow calculation is used to determine the matching feature point pairs, it may be interfered by future information, affecting the accuracy of the calculation results.
[0177] For example,Figure 5B It is a schematic diagram of the process of determining the matching feature point pairs of every two adjacent video images shown according to an exemplary embodiment of the present disclosure; as Figure 5B shown, in this embodiment, forward optical flow calculation can be performed first to obtain the second feature points NextPts of the subsequent video image from the first feature points PrevPts of the previous video image; then backward optical flow calculation can be performed to obtain the third feature points PrevPtsRvs of the previous video image from the second feature points NextPts of the subsequent video image. Then, the pixel error between the first feature points PrevPts and the third feature points PrevPtsRvs can be determined. Furthermore, the feature point pairs with pixel errors greater than or equal to the threshold in every two adjacent video images can be deleted to obtain the matching feature point pairs of every two adjacent video images, that is, one PrevPts and a corresponding NextPts form a matching feature point pair. Among them, the above threshold can be set according to application requirements, such as set to 1, etc., and this embodiment does not limit this.
[0178] It can be seen from the above description that in this embodiment, by determining the second feature points of the subsequent video image in every two adjacent video images from the first feature points of the previous video image in every two adjacent video images, and determining the third feature points of the previous video image from the second feature points, then determining the pixel error between the first feature points and the third feature points, and further deleting the feature point pairs with pixel errors greater than or equal to the threshold in every two adjacent video images to obtain the matching feature point pairs of every two adjacent video images, it is possible to accurately determine the matching feature point pairs of every two adjacent video images, which is beneficial to subsequent determination of the average displacement amount of the matching feature points in each video frame within every two adjacent video images, and determining the motion vectors of each of the multiple videos based on the average displacement amount, that is, it is possible to determine the motion vectors of the complete content of multiple videos at one time, which is beneficial to subsequent convenient and fast comparative analysis of the global stability of the videos captured by multiple photographic devices, and meets the requirements of scenarios such as equipment performance analysis and application research and development.
[0179] Figure 6A It is a flowchart of a method for determining video motion vectors shown according to another exemplary embodiment of the present disclosure; the method of this embodiment can be executed by a video motion vector determination device, and this video motion vector determination device can be configured in an electronic device with video data processing functions, such as a server, a workstation, a personal computer, a mobile terminal (such as a mobile phone, a tablet computer, etc.), a wearable device (such as glasses, a watch, etc.), etc. As Figure 6A shown, the method includes the following steps S601 - S605:
[0180] In step S601, in response to obtaining multiple videos for which motion vectors are to be determined, each frame of video of the multiple videos is spliced to obtain a spliced video, and each frame of video image of the spliced video includes the video of the corresponding frame of the multiple videos.
[0181] In step S602, for the feature points of every two adjacent frames of video images in the spliced video, forward optical flow and backward optical flow are calculated to obtain the matching feature point pairs of every two adjacent frames of video images.
[0182] In step S603, the average displacement amount of the matching feature point pairs of each video in every two adjacent frames of video images is determined.
[0183] In step S604, based on the average displacement amount, the motion vectors of the multiple videos are determined respectively.
[0184] Among them, for the relevant explanations and descriptions of steps S601 - S604, reference can be made to steps S101 - S104 in the above Figure 1 illustrated embodiments, which will not be elaborated here.
[0185] In step S605, based on the average displacement amount, the motion vectors of the multiple videos are determined respectively.
[0186] In step S606, the motion vectors of the multiple videos are input into a preset application program for differential display.
[0187] In this embodiment, on the display page of the above preset application program, by setting at least one of the color, line type, and indication mark of the motion vector curves of the multiple videos, differential display of the motion vectors of the multiple videos can be achieved.
[0188] In some embodiments, the above preset application program can be an interactive visualization tool pre-written for implementing functions such as data adaptive scaling of video motion vectors and rapid positioning of video time nodes. Exemplarily, this visualization tool can be used to display the motion vectors of the multiple videos determined in the above embodiments in the form of a chart, and can also achieve precise positioning from the motion vector data to the video time axis, thereby providing users with a friendly interactive experience and precise data analysis functions.
[0189] For example, Figure 6B is a schematic diagram of the display interface of an interactive visualization tool shown according to an exemplary embodiment of the present disclosure; as Figure 6BAs shown, the display interface of the visualization tool includes the following parts: a. list bar, b. bottom bar, c. chart box, d. top bar, e. video player, f. video controller. Exemplarily, first, the user can click on the "Select File" option shown in the lower left corner to select a document (such as a document in json format, etc.) storing the motion vector data of multiple videos, and then all visualizable video options can be displayed in the list bar on the left (i.e., "a. list bar").
[0190] Secondly, the user can trigger the visualization tool to display the motion vectors of multiple videos in the x and y axis directions respectively in the chart box on the right (i.e., "c. chart box") by clicking on the video options in the list bar. As an example, the chart box is divided into two parts, where the content in the upper part is the motion vector of each video frame in the x axis direction, and the content in the lower part is the motion vector of each video frame in the y axis direction. Among them, the vertical axis unit of the upper and lower charts is the number of pixels, the horizontal axis is the video playback time, and the four curves 0, 1, 2, and 3 respectively correspond to the motion vectors of different video frame effective regions (FOV) in the spliced video. On this basis, the user can also perform operations such as up and down sliding and free scaling using the scroll wheel as needed, and can also pan the content in the chart by dragging with the mouse.
[0191] Furthermore, the top bar of the visualization tool (i.e., "d. top bar") contains two function options: "Y-axis synchronization" and "ROI (Region of Interest)". Among them, the function of the "Y-axis synchronization" option is to synchronize the scaling of the y-axis scales of the upper and lower charts to facilitate subsequent comparison of the amplitudes of the motion vectors in the x-axis or y-axis directions. The function of the "ROI" option is to magnify the area selected by the user. Exemplarily, the magnification effect of the "ROI" option is as Figure 6C shown. As Figure 6C shown, when the user selects the data in the rectangular shaded area in the left half of the figure based on the "ROI" option, the visualization tool can magnify the selected part to the effect in the right half of the right figure for the user to view the content of this part more clearly.
[0192] Again, "x" and "y" in the bottom bar of the visualization tool (i.e., "b. bottom bar") can be used to indicate the data at the current position of the mouse; when there is a selected area by the user, "△X" and "△Y" can be used to indicate the time difference and pixel difference of the current selected area respectively.
[0193] Next, when the user clicks Figure 6BWhen the "Open Video" option in the lower left corner is selected, the video player (i.e., "e. Video Player") can be opened. This player can interact with the chart in the chart frame to achieve precise positioning of motion vector data. Exemplarily, the user can double-click on the data in the chart, causing the visualization tool to send the corresponding data coordinates to the video player program, so that the played video jumps to the corresponding progress.
[0194] In addition, there are various jump functions at the bottom of the player (i.e., "f. Video Controller"). The user can click on options such as "Draw Curve", "Play", "Stop", "Previous Frame", "Next Frame", and "1.0x (Playback Speed)" to control the playback of the video player.
[0195] As can be seen from the above description, the device of this embodiment splices each frame of video images of the multiple videos in response to obtaining the multiple videos for which the motion vectors are to be determined, obtains a spliced video, calculates the forward optical flow and backward optical flow for the feature points of each pair of adjacent frames of video images in the spliced video to obtain the matching feature point pairs of each pair of adjacent frames of video images, then determines the average displacement amount of the matching feature point pairs of each video image within each pair of adjacent frames of video images, and determines the motion vectors of each of the multiple videos based on the average displacement amount. Furthermore, the motion vectors of each of the multiple videos are input into a preset application program for differential display, which can achieve the determination of the motion vectors of the complete content of multiple videos at one time, facilitating the convenient and rapid comparative analysis of the global stability of videos captured by multiple photographic devices, and enabling vivid and flexible visual display, which can improve the display effect of video motion vector data, thereby meeting the requirements of scenarios such as device performance analysis and application research and development.
[0196] Figure 7 is a block diagram of a video motion vector determination device shown according to an exemplary embodiment of the present disclosure; the device of this embodiment can be configured in an electronic device with video data processing capabilities, such as a server, a workstation, a personal computer, a mobile terminal (such as a mobile phone, a tablet computer, etc.), a wearable device (such as glasses, a watch, etc.). As Figure 7 shown, the device may include: a video acquisition module 110, a point pair determination module 120, a displacement determination module 130, and a vector determination module 140, where:
[0197] The video acquisition module 110 is configured to splice each frame of video images of the multiple videos in response to obtaining the multiple videos for which the motion vectors are to be determined, and obtain a spliced video, where each frame of video image of the spliced video includes the video images of the corresponding frames of the multiple videos;
[0198] A point pair determination module 120 is configured to calculate the forward optical flow and the backward optical flow for the feature points of every two adjacent video images in the stitched video, so as to obtain the matching feature point pairs of every two adjacent video images;
[0199] A displacement determination module 130 is configured to determine the average displacement amount of the matching feature point pairs of each video frame in every two adjacent video images;
[0200] A vector determination module 140 is configured to determine the motion vectors of each of the multiple videos based on the average displacement amount.
[0201] As can be seen from the above description, the device in this embodiment stitches each video frame of the multiple videos in response to obtaining the multiple videos for which motion vectors are to be determined, to obtain a stitched video. Each video image of the stitched video includes the video frames corresponding to the multiple videos. Then, the forward optical flow and the backward optical flow are calculated for the feature points of every two adjacent video images in the stitched video to obtain the matching feature point pairs of every two adjacent video images. Next, the average displacement amount of the matching feature point pairs of each video frame in every two adjacent video images is determined. Furthermore, the motion vectors of each of the multiple videos are determined based on the average displacement amount, so as to realize the determination of the motion vectors of the complete content of multiple videos at one time, which is beneficial to the subsequent convenient and fast comparative analysis of the global stability of the videos captured by multiple photographic devices, and thus can meet the requirements of scenarios such as device performance analysis and application research and development.
[0202] Figure 8 is a block diagram of a video motion vector determination device shown according to an exemplary embodiment of the present disclosure; the device in this embodiment can be configured in an electronic device with video data processing functions, such as a server, a workstation, a personal computer, a mobile terminal (such as a mobile phone, a tablet computer, etc.), a wearable device (such as glasses, a watch, etc.). Among them, the video acquisition module 210, the point pair determination module 220, the displacement determination module 230, and the vector determination module 240 have the same functions as the video acquisition module 110, the point pair determination module 120, the displacement determination module 130, and the vector determination module 140 in the foregoing Figure 7 shown embodiment, and will not be described in detail here.
[0203] As Figure 8 shown, the device may further include:
[0204] A frame division module 250 is configured to divide each video frame into a main body area and a background area in response to detecting a target main body in each video frame of each video image, where the main body area is the area where the target main body is located;
[0205] Furthermore, the displacement determination module 230 can also be used to determine the average displacement amount of the matching feature point pairs in the specified area within each of the video frames, where the specified area is the main area and / or the background area;
[0206] The vector determination module 240 can also be used to determine the motion vectors of the respective multiple videos based on the average displacement amount of the matching feature point pairs within the specified area.
[0207] In some embodiments, the feature points of every two adjacent video frames described above can include the feature points within the valid area of every two adjacent video frames.
[0208] On this basis, the above device can further include a region determination module 260;
[0209] The region determination module 260 can include:
[0210] The binary processing unit 261 is used to perform binary processing on the first video frame selected from the spliced video to obtain a binary image;
[0211] The contour search unit 262 is used to search for the contours of each foreground in the binary image;
[0212] The contour removal unit 263 is used to remove the contours with a surrounding area less than or equal to the pixel quantity threshold from the found contours of each foreground to obtain the remaining contours;
[0213] The first determination unit 264 is used to determine the first valid area of each video frame in the first video frame based on the remaining contours;
[0214] The region determination unit 265 is used to determine the area corresponding to the first valid area in every two adjacent video frames as the valid area of every two adjacent video frames.
[0215] In some embodiments, the above first determination unit 264 can also be used to:
[0216] Determine the first area containing the remaining contours in each video frame;
[0217] Reduce the first area by a preset ratio to obtain the first valid area of each video frame.
[0218] In some embodiments, the above region determination module 260 can further include:
[0219] The morphological transformation unit 266 is used to eliminate the noise and / or holes in the binary image through morphological transformation.
[0220] In some embodiments, the above-mentioned point pair determination module 220 may include:
[0221] A forward calculation unit 221, configured to determine second feature points in the latter video image of each adjacent two video images from first feature points in the former video image of each adjacent two video images;
[0222] A backward calculation unit 222, configured to determine third feature points in the former video image from the second feature points;
[0223] An error determination unit 223, configured to determine the pixel error between the first feature points and the third feature points;
[0224] A point pair determination unit 224, configured to delete feature point pairs with a pixel error greater than or equal to a threshold in each adjacent two video images, to obtain matching feature point pairs in each adjacent two video images.
[0225] In some embodiments, the above-mentioned displacement determination module 230 may include:
[0226] A sum value determination unit 231, configured to determine the sum of displacement amounts of matching feature point pairs in each adjacent two video images in the spliced video;
[0227] A displacement determination unit 232, configured to determine the average displacement amount based on the sum of the displacement amounts and the number of matching feature point pairs in each adjacent two video images.
[0228] In some embodiments, the above-mentioned apparatus may further include:
[0229] A visualization module 270, configured to input the motion vectors of the respective multiple videos into a preset application program for differential display.
[0230] Regarding the apparatus in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0231] Figure 9 is a block diagram of an electronic device shown according to an exemplary embodiment. For example, device 900 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0232] Referring to Figure 9 , device 900 may include one or more of the following components: a processing component 902, a memory 904, a power component 906, a multimedia component 908, an audio component 910, an input / output (I / O) interface 912, a sensor component 914, and a communication component 916.
[0233] The processing component 902 generally controls the overall operation of the device 900, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 902 may include one or more processors 920 to execute instructions to complete all or part of the steps of the above-described video motion vector determination method. In addition, the processing component 902 may include one or more modules to facilitate the interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate the interaction between the multimedia component 908 and the processing component 902.
[0234] The memory 904 is configured to store various types of data to support the operation of the device 900. Examples of such data include instructions for any application or method operating on the device 900, contact data, phone book data, messages, pictures, videos, etc. The memory 904 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0235] The power component 906 provides power to various components of the device 900. The power component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 900.
[0236] The multimedia component 908 includes a screen that provides an output interface between the device 900 and the user. In some embodiments, the screen may include a liquid crystal display panel and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the device 900 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0237] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC) that is configured to receive external audio signals when the device 900 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 904 or transmitted via the communication component 916. In some embodiments, the audio component 910 further includes a speaker for outputting audio signals.
[0238] The I / O interface 912 provides an interface between the processing component 902 and peripheral interface modules, and the peripheral interface modules may be a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0239] The sensor component 914 includes one or more sensors for providing an assessment of various aspects of the state of the device 900. For example, the sensor component 914 can detect the on / off state of the device 900, the relative positioning of components, such as the display panel and keypad of the device 900. The sensor component 914 can also detect a change in the position of the device 900 or a component of the device 900, the presence or absence of user contact with the device 900, the orientation or acceleration / deceleration of the device 900, and the temperature change of the device 900. The sensor component 914 can also include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 914 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 914 can also include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0240] The communication component 916 is configured to facilitate communication between the device 900 and other devices in a wired or wireless manner. The device 900 can access a wireless network based on communication standards, such as WiFi, 2G or 3G, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0241] In an exemplary embodiment, the device 900 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described video motion vector determination method.
[0242] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as a memory 904 including instructions, is also provided. The above instructions may be executed by a processor 920 of the device 900 to complete the above-described video motion vector determination method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0243] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the disclosure herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0244] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes may be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for determining a video motion vector, characterized in that, The method includes: In response to obtaining multiple videos for which motion vectors are to be determined, splicing each frame of video images of the multiple videos to obtain a spliced video, where each frame of video image of the spliced video includes the video images of the corresponding frames of the multiple videos; Calculating forward optical flow and backward optical flow for the feature points of every two adjacent frames of video images in the spliced video to obtain matching feature point pairs of every two adjacent frames of video images; Determining the average displacement amount of the matching feature point pairs of each video image within every two adjacent frames of video images; Determining the motion vectors of each of the multiple videos based on the average displacement amount.
2. The method according to claim 1, wherein The method further includes: In response to detecting a target object in each video image within each frame of video images, dividing each video image into a main body area and a background area, where the main body area is the area where the target object is located; The determining the average displacement amount of the matching feature point pairs of each video image within every two adjacent frames of video images includes: Determining the average displacement amount of the matching feature point pairs within a specified area in each video image, where the specified area is the main body area and / or the background area; The determining the motion vectors of each of the multiple videos based on the average displacement amount includes: Determining the motion vectors of each of the multiple videos based on the average displacement amount of the matching feature point pairs within the specified area.
3. The method according to claim 1, characterized in that, The feature points of every two adjacent frames of video images include the feature points within the valid area of every two adjacent frames of video images.
4. The method according to claim 3, wherein The method further includes pre-determining the valid area based on the following method: Performing binarization processing on a first video image selected from the spliced video to obtain a binary image; Searching for the contours of each foreground in the binary image; Among the searched contours of each foreground, removing the contours with an enclosed area less than or equal to a pixel quantity threshold to obtain remaining contours; Determining the first valid area of each video image in the first video image based on the remaining contours; Determining the area corresponding to the first valid area in every two adjacent frames of video images as the valid area of every two adjacent frames of video images.
5. The method according to claim 4, wherein The determining the first valid area of each video image in the first video image based on the remaining contours includes: Determining the first area in each video image that contains the remaining contours; Reducing the first area by a preset ratio to obtain the first valid area of each video image.
6. The method according to claim 4, characterized in that, Before searching for the contours of each foreground in the binary image, it further includes: Eliminating the noise and / or holes in the binary image through morphological transformation.
7. The method according to claim 1, characterized in that The calculating forward optical flow and backward optical flow for the feature points of every two adjacent frames of video images in the spliced video to obtain the matching feature point pairs of every two adjacent frames of video images includes: Determining a second feature point in the subsequent frame of video image of every two adjacent frames of video images from a first feature point in the previous frame of video image of every two adjacent frames of video images; Determining a third feature point in the previous frame of video image from the second feature point; Determining the pixel error between the first feature point and the third feature point; Delete the feature point pairs with pixel errors greater than or equal to the threshold in every two adjacent video images to obtain the matching feature point pairs of every two adjacent video images.
8. The method according to claim 1, characterized in that, The determination of the average displacement of the matching feature point pairs of each video frame within every two adjacent video images includes: Determine the sum of the displacement amounts of the matching feature point pairs of every two adjacent video images in the spliced video; Based on the sum of the displacement amounts and the number of the matching feature point pairs of every two adjacent video images, determine the average displacement amount.
9. The method according to claim 1, characterized in that, The method further includes: Input the motion vectors of the multiple videos into a preset application program for differential display.
10. A video motion vector determination device, characterized in that, The device includes: A video acquisition module, configured to, in response to acquiring multiple videos for which motion vectors are to be determined, splice each video frame of the multiple videos to obtain a spliced video, where each video image of the spliced video includes the video frames corresponding to the multiple videos; A point pair determination module, configured to calculate the forward optical flow and backward optical flow for the feature points of every two adjacent video images in the spliced video to obtain the matching feature point pairs of every two adjacent video images; A displacement determination module, configured to determine the average displacement of the matching feature point pairs of each video frame within every two adjacent video images; A vector determination module, configured to determine the motion vectors of the multiple videos respectively based on the average displacement amount.
11. An electronic device, characterized in that, The device includes: A processor and a memory for storing a computer program; Wherein, the processor is configured to, when executing the computer program, implement: In response to acquiring multiple videos for which motion vectors are to be determined, splice each video frame of the multiple videos to obtain a spliced video, where each video image of the spliced video includes the video frames corresponding to the multiple videos; Calculate the forward optical flow and backward optical flow for the feature points of every two adjacent video images in the spliced video to obtain the matching feature point pairs of every two adjacent video images; Determine the average displacement of the matching feature point pairs of each video frame within every two adjacent video images; Determine the motion vectors of the multiple videos respectively based on the average displacement amount.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements: In response to acquiring multiple videos for which motion vectors are to be determined, splice each video frame of the multiple videos to obtain a spliced video, where each video image of the spliced video includes the video frames corresponding to the multiple videos; Calculate the forward optical flow and backward optical flow for the feature points of every two adjacent video images in the spliced video to obtain the matching feature point pairs of every two adjacent video images; Determine the average displacement of the matching feature point pairs of each video frame within every two adjacent video images; Determine the motion vectors of the multiple videos respectively based on the average displacement amount.