Video jitter sequence extraction method and device based on fixed scene and medium

By analyzing the optical flow data of adjacent frame images in the video, generating jitter sequence analysis diagrams and detecting peaks, extracting jitter frame sequences in the video, solving the problem of difficulty in positioning and extracting unstable segments in the prior art, and achieving efficient video dejitter processing.

CN120017967APending Publication Date: 2025-05-16SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510054217.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to quickly locate and extract unstable segments in videos, resulting in large amounts of dejitter operation calculations, affecting efficiency and potentially damaging video quality.

Method used

By obtaining the image optical flow data of two adjacent frames of images in the target video data, determining the motion intensity information, generating a jitter sequence analysis chart, performing peak detection, and extracting the target jitter frame sequence with jitter.

Benefits of technology

It realizes rapid positioning and extraction of jitter clips in video, reduces the calculation amount of the dejitter algorithm, improves the computing efficiency, and avoids unnecessary processing of non-jitter video clips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017967A_ABST
    Figure CN120017967A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video analysis, and relates to a video jitter sequence extraction method and device based on a fixed scene, and a medium. The method comprises the following steps: acquiring target video data for a target area; extracting each target image in the target video data; for each target image, acquiring image optical flow data of the target image, and determining motion intensity information of the target image according to the image optical flow data; generating a jitter sequence analysis chart according to each piece of motion intensity information; performing peak value detection on the jitter sequence analysis chart to obtain a target peak value meeting a preset peak value requirement; and determining a target jitter frame sequence with jitter in the target video data according to each target peak value. According to the method and the device, the target jitter frame sequence with jitter can be extracted, whole-section jitter elimination does not need to be carried out on the target video data when a jitter elimination algorithm is carried out subsequently, and jitter fragments in the target video data can be quickly positioned and extracted while the calculation amount of the jitter elimination algorithm is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video analysis technology, and in particular, to a method, device and medium for extracting video jitter sequences based on a fixed scene. Background Art

[0002] Video shooting scenes are divided into fixed scenes and moving scenes. Fixed scene videos, that is, the scene of the video shooting does not move significantly, and the monitoring area is fixed. Most videos in the field of video surveillance are fixed scene videos. In recent years, video surveillance has been stimulated and driven by factors such as security projects and the rapid growth of video surveillance demand in various industries, and has achieved rapid development, and the entire market size has expanded rapidly. If the monitoring system wants to play its due role, it must ensure the quality of the transmitted video, and the monitoring system needs to be operated and maintained.

[0003] Video jitter is a common fault in monitoring systems. It is usually caused by the camera not being fixed firmly enough or external force or human power, which causes the video screen to shake periodically or irregularly. With the continuous development and expansion of the monitoring market, the number of front-end monitoring cameras is increasing, and the workload of manual operation and maintenance is increasing, and the cost is increasing. Therefore, a method to automatically detect whether the video is jittery is particularly important.

[0004] After detecting that the video is shaking, the video needs to be de-shaked. However, in a video, it is rare that the video is shaking all the time. If a video de-shaking operation is performed directly on the entire video, the amount of calculation will be increased for no reason, affecting the calculation efficiency. Even de-shaking the video clips without shaking will affect the original video quality. Therefore, a method to quickly locate and extract unstable clips in the video is particularly important. Summary of the invention

[0005] The present application provides a video jitter sequence extraction method, device and medium based on a fixed scene to solve one or more technical problems existing in the prior art and at least provide a beneficial choice or create conditions.

[0006] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.

[0007] According to one aspect of an embodiment of the present application, a method for extracting a video jitter sequence based on a fixed scene is provided, the method comprising:

[0008] Acquire target video data for a target area, wherein the target video data consists of multiple frames of images;

[0009] Extracting each target image from the target video data, each of the target images is composed of two adjacent frames of images;

[0010] For each target image, acquiring image optical flow data of the target image, and determining motion intensity information of the target image according to the image optical flow data;

[0011] Generate a jitter sequence analysis graph of the target video data according to each of the motion intensity information;

[0012] Performing peak detection on the jitter sequence analysis graph to obtain multiple target peak values ​​that meet preset peak requirements;

[0013] A target jitter frame sequence with jitter in the target video data is determined according to each of the target peak values.

[0014] In one embodiment of the present application, based on the above-mentioned solution, each pixel point of each frame of the image corresponds to each position point in the target area, and the step of acquiring the image optical flow data of the target image includes:

[0015] Acquire a plurality of optical flow vectors between two adjacent frames of the target image, wherein each pixel point corresponds to one optical flow vector;

[0016] According to a preset random parameter estimation algorithm, the abnormal optical flow vectors in an abnormal state are removed from the optical flow vectors to obtain the target optical flow vectors in a normal state;

[0017] Generate the image optical flow data according to each of the target optical flow vectors;

[0018] The abnormal optical flow vector is used to indicate that there is a moving object at a position point corresponding to the abnormal optical flow vector.

[0019] In one embodiment of the present application, based on the above solution, determining the motion intensity information of the target image according to the image optical flow data includes:

[0020] For each target optical flow vector in the image optical flow data, performing displacement vector decomposition on the target optical flow vector to obtain a first displacement vector in a first preset direction and a second displacement vector in a second preset direction, and determining a pixel optical flow vector of a pixel point corresponding to the target optical flow vector according to the first displacement vector and the second displacement vector;

[0021] The motion intensity information of the target image is determined according to each of the pixel optical flow vectors.

[0022] In one embodiment of the present application, based on the above solution, determining the motion intensity information of the target image according to each of the pixel optical flow vectors includes:

[0023] The motion intensity information is obtained by averaging the sums of the optical flow vectors of each pixel.

[0024] In one embodiment of the present application, based on the above scheme, the target video data includes frame sequence information, and the frame sequence information is used to characterize the order of the images of each frame in the target video data, and the jitter sequence analysis diagram of the target video data is generated according to each of the motion intensity information, including:

[0025] The jitter sequence analysis graph is generated according to each of the motion intensity information and the frame sequence information.

[0026] In one embodiment of the present application, based on the aforementioned scheme, the preset peak requirement includes a preset height threshold and a preset protrusion threshold, the preset height threshold is the sum of the first preset threshold and the second preset threshold, and the preset protrusion threshold is one tenth of the third preset threshold; the first preset threshold is the average value of each of the motion intensity information, the second preset threshold is one half of the standard deviation of each of the motion intensity information, and the third preset threshold is the maximum value of each of the motion intensity information; the preset peak requirement is that the motion intensity information is greater than the preset height threshold and greater than the preset protrusion threshold.

[0027] In one embodiment of the present application, based on the above scheme, the target peak value includes an index position and a width, and determining a target jitter frame sequence having jitter in the target video data according to each of the target peak values ​​includes:

[0028] Searching for a jitter frame sequence with jitter in the jitter sequence analysis diagram according to each of the target peaks;

[0029] A target jittery frame sequence having jitter in the target video data is determined according to each of the jittery frame sequences.

[0030] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. The computer program includes executable instructions. When the executable instructions are executed by a processor, the method described in the above embodiment is implemented.

[0031] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; and a memory for storing executable instructions of the processors, wherein when the executable instructions are executed by the one or more processors, the one or more processors implement the methods described in the above embodiments.

[0032] Beneficial effects of the present application: The present application obtains image optical flow data of two adjacent frames of images, that is, the target image, and then determines the motion intensity information of the target image through the image optical flow data, that is, the motion intensity information of two adjacent frames of images. In this way, the motion intensity information between each two adjacent frames of images in the target video data is obtained, and then the corresponding jitter sequence analysis diagram is generated. Then, the peak value detection is performed on the jitter sequence analysis diagram to obtain multiple target peak values ​​that meet the preset peak value requirements, and then the target jitter frame sequence with jitter in the target video data can be determined according to the position of the target peak value, so as to extract the target jitter frame sequence with jitter. When performing the de-jitter algorithm later, there is no need to de-jitter the entire segment of the target video data, and only the extracted target jitter frame sequence needs to be de-jittered, which reduces the calculation amount of the de-jitter algorithm while quickly locating and extracting the jittered segments in the target video data.

[0033] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0035] Figure 1 It is a flowchart of a video jitter sequence extraction method based on a fixed scene according to an embodiment of the present application;

[0036] Figure 2 A video sequence waveform diagram formed without removing abnormal optical flow vectors using the RANSAC algorithm according to an embodiment of the present application;

[0037] Figure 3 The waveform diagram of the video sequence formed by removing abnormal optical flow vectors through the RANSAC algorithm according to the embodiment of the present application;

[0038] Figure 4 is a waveform diagram of a jitter sequence with jitter according to an embodiment of the present application;

[0039] Figure 5 is a waveform diagram of a jitter sequence without jitter according to an embodiment of the present application;

[0040] Figure 6 It is a structural diagram of an electronic device according to the present application. DETAILED DESCRIPTION

[0041] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.

[0042] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the present application.

[0043] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or micro-control node devices.

[0044] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0045] It should be noted that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0046] The following is a detailed introduction to the background technology of the embodiments of the present application:

[0047] The relevant methods are as follows:

[0048] Grayscale projection method: The image grayscale projection algorithm mainly uses grayscale to perform cumulative projection in the horizontal and vertical directions, and estimates the motion parameters of the inter-frame sequence based on the distribution law of the projection. The algorithm calculates the difference between the row and column and the row and column mean, and draws the grayscale projection curves of the current frame and the reference frame based on the difference, and then performs cross-correlation operations on the row and column projections of the two. The absolute value of the corresponding extreme value vector in the curve is taken as the displacement of the image frame. When the displacement reaches a certain set value, the current frame is identified as a jitter frame.

[0049] The grayscale projection method only targets the rows and columns of the image, can only calculate translational motion, cannot calculate rotation and scaling motion, and has a very limited range; when there are moving targets in the video sequence, the change of the projection curve will be seriously disturbed, resulting in motion estimation errors. The above shortcomings make the algorithm have strong limitations and difficult to apply to video surveillance.

[0050] Block matching method: the most commonly used algorithm in video stabilization systems. This method divides the current frame into blocks, each pixel in the block has the same motion vector, and then searches for the best match within a specific range of the reference frame for each block, thereby estimating the global motion vector of the video sequence.

[0051] The block matching method usually needs to divide the blocks and estimate the global motion vector based on the motion vector in each block. Therefore, it is not effective in detecting video jitter problems in certain specific scenes. For example, in a picture, the picture is divided into four grids, three of which are stationary and the object in one grid is moving. In addition, the block matching method usually requires Kalman filtering to process the calculated motion vector, which has high computational overhead and cannot adapt to the scenario of large displacement and strong jitter of the lens in a short time.

[0052] Feature point matching method: The core of the feature point matching algorithm is the extraction and matching of feature points. It relies heavily on feature point detection and has poor effects when feature points cannot be effectively located. For example, it cannot effectively detect monitoring scenes with relatively clean textures.

[0053] The implementation details of the technical solution of the embodiment of the present application are described in detail below:

[0054] According to one aspect of the present application, a video jitter sequence extraction method based on a fixed scene is provided. Figure 1 Flowchart of a video jitter sequence extraction method based on a fixed scene according to an embodiment of the present application. The video jitter sequence extraction method based on a fixed scene at least includes steps S1 to S6, which are described in detail as follows:

[0055] In step S1, target video data for a target area is acquired, where the target video data consists of multiple frames of images.

[0056] Specifically, the target video data may be video surveillance data of a certain time period, for example, if the fixed scene (i.e., the target area) is the school gate, then the target video data may be video surveillance data of the school gate from 7 a.m. to 9 a.m. Then the target video data is composed of multiple frames of images, that is, the video surveillance data composed of each frame of images captured from 7 a.m. to 9 a.m.

[0057] In step S2, each target image in the target video data is extracted, and each target image is composed of two adjacent frame images.

[0058] Specifically, the target image in the embodiment of the present application does not refer to a single image, but to two adjacent frames of images. For the convenience of subsequent description, the target image refers to two adjacent frames of images. Then, to find each target image in the target video, it means to extract each two adjacent frames of images in the target video data as the analysis object (i.e., calculate the optical flow vector of the two adjacent frames of images).

[0059] In step S3, for each target image, image optical flow data of the target image is acquired, and motion intensity information of the target image is determined according to the image optical flow data.

[0060] In one embodiment of the present application, each pixel point of each frame of the image corresponds to each position point in the target area, and acquiring the image optical flow data of the target image includes:

[0061] Acquire a plurality of optical flow vectors between two adjacent frames of the target image, wherein each pixel point corresponds to one optical flow vector;

[0062] According to a preset random parameter estimation algorithm, the abnormal optical flow vectors in an abnormal state are removed from the optical flow vectors to obtain the target optical flow vectors in a normal state;

[0063] Generate the image optical flow data according to each of the target optical flow vectors;

[0064] The abnormal optical flow vector is used to indicate that there is a moving object at a position point corresponding to the abnormal optical flow vector.

[0065] Specifically, each frame of the image is composed of multiple pixels, and each pixel corresponds to a position point in the target area. The Farneback optical flow method is used to estimate the optical flow of each consecutive frame of the video, and the optical flow vector corresponding to each pixel is obtained. Multiple optical flow vectors between two adjacent frames of the target image are obtained, that is, the optical flow vector between each pixel in the two adjacent frames of the image is obtained from the previous frame to the current frame.

[0066] The preset random parameter estimation algorithm is specifically the RANSAC algorithm, which is the random sampling consensus algorithm. The RANSAC algorithm infers abnormal data in the data through function fitting. If there is object movement in the image, then the optical flow vector generated by the moving object is abnormal for the entire image optical flow. The RANSAC algorithm can be used to identify abnormal pixels where there are moving objects, that is, the abnormal optical flow vectors corresponding to the abnormal pixels need to be removed to avoid being affected by the moving object. Figure 2 and Figure 3 As shown, Figure 2 is a video sequence waveform diagram formed without removing abnormal optical flow vectors through the RANSAC algorithm (i.e., the jitter sequence waveform diagram described in this application), Figure 3 The video sequence waveform diagram formed by removing abnormal optical flow vectors through the RANSAC algorithm (i.e., the jitter sequence waveform diagram described in this application) is obtained. Figure 2 and Figure 3 The comparison shows that Figure 2 The presence of moving objects in the image may cause high intensity motion, which may be misjudged as camera shaking. Figure 3 After removing the abnormal optical flow vectors corresponding to the moving object, a relatively stable waveform is obtained. Therefore, the waveform obtained by removing the abnormal optical flow vectors through the RANSAC algorithm (ie, the jitter sequence analysis diagram mentioned in the embodiment of the present application) is more accurate.

[0067] In one embodiment of the present application, determining the motion intensity information of the target image according to the image optical flow data includes:

[0068] For each target optical flow vector in the image optical flow data, performing displacement vector decomposition on the target optical flow vector to obtain a first displacement vector in a first preset direction and a second displacement vector in a second preset direction, and determining a pixel optical flow vector of a pixel point corresponding to the target optical flow vector according to the first displacement vector and the second displacement vector;

[0069] The motion intensity information of the target image is determined according to each of the pixel optical flow vectors.

[0070] Specifically, the optical flow vector can be decomposed into displacement vectors x (i.e., the first displacement vector described in the present application) and y (i.e., the second displacement vector described in the present application) in the horizontal direction (i.e., the first preset direction described in the present application) and the vertical direction (i.e., the second preset direction described in the present application), and the displacement vectors are Magnite is the modulus of the optical flow vector, that is, the size of each pixel optical flow vector, and then the motion intensity information is obtained by averaging the sum of the pixel optical flow vectors, that is, Figures 2 to 5 The vertical axis in any graph is the intensity of exercise. Figures 2 to 5The horizontal axis in is the frame sequence, that is, the number of frames, in order from the first frame to the Nth frame, and the number N is obtained according to the duration of the target video data and the frame rate of the camera.

[0071] In step S4, a jitter sequence analysis diagram of the target video data is generated according to each piece of motion intensity information.

[0072] Specifically, the motion intensity information between each two adjacent frames and their own frame sequence information (such as the order of the two adjacent frames in all the frames) are used to generate a curve graph (i.e., a jitter sequence analysis graph), such as Figure 4 and Figure 5 As shown in the figure, it can be seen that when jitter occurs in the video, it is manifested as a peak in the jitter sequence analysis diagram, and when there is no jitter in the entire video, it is manifested as a smooth curve or periodic fluctuation in the jitter sequence analysis diagram ( Figures 2 to 5 The horizontal axis is the video frame sequence (frame sequence information), and the vertical axis is the motion intensity (i.e., the motion intensity information described in this application). Figure 4 It can be seen that when jitter occurs in the video sequence (ie, the target video data), it is reflected as a significant peak in the jitter sequence analysis graph. Figure 5 The entire target video data has no jitter, showing periodic fluctuations, and the motion intensity is relatively low.

[0073] In step S5, peak detection is performed on the jitter sequence analysis graph to obtain a plurality of target peak values ​​that meet preset peak value requirements.

[0074] In one embodiment of the present application, the preset peak requirement includes a preset height threshold and a preset protrusion threshold, the preset height threshold is the sum of the first preset threshold and the second preset threshold, and the preset protrusion threshold is one tenth of the third preset threshold; the first preset threshold is the average value of each of the motion intensity information, the second preset threshold is one half of the standard deviation of each of the motion intensity information, and the third preset threshold is the maximum value of each of the motion intensity information.

[0075] Specifically, the peak value of the jitter sequence analysis chart is detected, and the peak value in the motion signal sequence (i.e., the corresponding target video data) is identified by using the peak value detection function corresponding to the preset peak value requirement. The preset height threshold is set as the average value of each motion intensity information + 0.5 × the standard deviation of each motion intensity information. The detected target peak value needs to be greater than the preset height threshold, which ensures that the detected target peak value is relatively prominent, because its height (i.e., motion intensity information) is significantly higher than the average level, reducing the sensitivity to noise.

[0076] Preset prominence threshold=0.1×the maximum value of each of the motion intensity information, that is, 10% of the maximum signal value. The preset prominence threshold is used to ensure that a sufficiently prominent target peak is detected, that is, the target peak is sufficiently conspicuous in its surrounding background (multiple peaks adjacent to it).

[0077] In step S6, a target jitter frame sequence having jitter in the target video data is determined according to each of the target peak values.

[0078] In one embodiment of the present application, based on the above scheme, the target peak value includes an index position and a width, and determining a target jitter frame sequence having jitter in the target video data according to each of the target peak values ​​includes:

[0079] Searching for a jitter frame sequence with jitter in the jitter sequence analysis diagram according to each of the target peaks;

[0080] A target jittery frame sequence having jitter in the target video data is determined according to each of the jittery frame sequences.

[0081] Specifically, according to each of the target peaks, a jitter frame sequence with jitter is searched in the jitter sequence analysis diagram, such as Figure 2 As shown in the figure, the target peak in the red box corresponds to the 75th to 140th frames of the target video data (i.e., the width of the index), the index position is the 110th frame, and the corresponding motion intensity is about 1.92. At this time, the jitter frame sequence (i.e., Figure 2 The red motion signal box in the middle is determined as the target jitter frame sequence. The target jitter frame sequence is extracted and input into the de-jittering algorithm for calculation, so that the target video data can be de-jittered without performing a large amount of de-jittering calculation on the entire target video data like the existing method. This reduces the calculation cost and improves the calculation efficiency, and can quickly locate the jittered video clips.

[0082] As another aspect, the present application further provides a computer-readable storage medium on which a program product capable of implementing the method provided above in this specification is stored. In some possible implementations, various aspects of the present application may also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary implementations of the present application described in the above "Embodiment Method" section of this specification.

[0083] According to the embodiment of the present application, the program product for implementing the above method can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present application is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.

[0084] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0085] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0086] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.

[0087] Program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0088] Refer to the following Figure 6 The electronic device 400 according to this embodiment of the present application is described. Figure 6 The electronic device 400 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0089] like Figure 6 As shown, the electronic device 400 is in the form of a general computing device. The components of the electronic device 400 may include but are not limited to: at least one processing unit 410, at least one storage unit 420, and a bus 430 connecting different system components (including the storage unit 420 and the processing unit 410).

[0090] The storage unit stores program codes, which can be executed by the processing unit 410, so that the processing unit 410 executes the steps described in the above “Example Method” section of this specification according to various exemplary implementations of the present application.

[0091] The storage unit 420 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 421 and / or a cache memory unit 422 , and may further include a read-only memory unit (ROM) 423 .

[0092] The storage unit 420 may also include a program / utility 424 having a set (at least one) of program modules 425, such program modules 425 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0093] Bus 430 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller node, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0094] The electronic device 400 may also communicate with one or more external devices 1200 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 400, and / or communicate with any device that enables the electronic device 400 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 450. In addition, the electronic device 400 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 460. As shown, the network adapter 460 communicates with other modules of the electronic device 400 via a bus 430. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0095] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation methods of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation methods of the present application.

[0096] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.

[0097] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be performed without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A video jitter sequence extraction method based on a fixed scene, characterized in that: The method comprises: Acquire target video data for a target area, wherein the target video data consists of multiple frames of images; Extracting each target image from the target video data, each of the target images is composed of two adjacent frames of images; For each target image, acquiring image optical flow data of the target image, and determining motion intensity information of the target image according to the image optical flow data; Generate a jitter sequence analysis graph of the target video data according to each of the motion intensity information; Performing peak detection on the jitter sequence analysis graph to obtain multiple target peak values ​​that meet preset peak requirements; A target jitter frame sequence with jitter in the target video data is determined according to each of the target peak values.

2. The video jitter sequence extraction method based on a fixed scene according to claim 1 is characterized in that: Each pixel point of each frame of the image corresponds to each position point in the target area, and the acquiring of the image optical flow data of the target image includes: Acquire a plurality of optical flow vectors between two adjacent frames of the target image, wherein each pixel point corresponds to one optical flow vector; According to a preset random parameter estimation algorithm, the abnormal optical flow vectors in an abnormal state are removed from the optical flow vectors to obtain the target optical flow vectors in a normal state; Generate the image optical flow data according to each of the target optical flow vectors; The abnormal optical flow vector is used to indicate that there is a moving object at a position point corresponding to the abnormal optical flow vector.

3. The video jitter sequence extraction method based on fixed scene according to claim 2 is characterized in that: The determining the motion intensity information of the target image according to the image optical flow data comprises: For each target optical flow vector in the image optical flow data, performing displacement vector decomposition on the target optical flow vector to obtain a first displacement vector in a first preset direction and a second displacement vector in a second preset direction, and determining a pixel optical flow vector of a pixel point corresponding to the target optical flow vector according to the first displacement vector and the second displacement vector; The motion intensity information of the target image is determined according to each of the pixel optical flow vectors.

4. The video jitter sequence extraction method based on fixed scenes according to claim 3 is characterized in that: The determining the motion intensity information of the target image according to each of the pixel optical flow vectors includes: The motion intensity information is obtained by averaging the sums of the optical flow vectors of each pixel.

5. The video jitter sequence extraction method based on fixed scene according to claim 4 is characterized in that: The target video data includes frame sequence information, and the frame sequence information is used to characterize the order of the images of each frame in the target video data. The jitter sequence analysis diagram of the target video data is generated according to each of the motion intensity information, including: The jitter sequence analysis graph is generated according to each of the motion intensity information and the frame sequence information.

6. The video jitter sequence extraction method based on fixed scenes according to claim 5 is characterized in that: The preset peak requirement includes a preset height threshold and a preset protrusion threshold, the preset height threshold is the sum of the first preset threshold and the second preset threshold, and the preset protrusion threshold is one tenth of the third preset threshold; The first preset threshold is the average value of each of the motion intensity information, the second preset threshold is half of the standard deviation of each of the motion intensity information, and the third preset threshold is the maximum value of each of the motion intensity information; the preset peak requirement is that the motion intensity information is greater than the preset height threshold and greater than the preset protrusion threshold.

7. The video jitter sequence extraction method based on fixed scenes according to claim 6 is characterized in that: The target peak value includes an index position and a width, and determining a target jitter frame sequence having jitter in the target video data according to each of the target peak values ​​includes: Searching for a jitter frame sequence with jitter in the jitter sequence analysis diagram according to each of the target peaks; A target jittery frame sequence having jitter in the target video data is determined according to each of the jittery frame sequences.

8. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the video jitter sequence extraction method based on a fixed scene according to any one of claims 1 to 7 when executing the computer program.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for extracting a video jitter sequence based on a fixed scene according to any one of claims 1 to 7 is implemented.