Video anti-shake processing method, medium and equipment

By downsampling and grid division of video frames, the problem of missing pictures caused by violent shaking in port container monitoring is solved, and the output of stable video frames is achieved, ensuring the safety of port operations.

CN120358419AActive Publication Date: 2025-07-22BROAD VISION (XIAMEN) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510825446.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-22
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The existing video anti-shake algorithm cannot be applied to video images that shake violently in port container monitoring scenarios, resulting in missing surveillance images and affecting operational safety.

Method used

By downsampling and extracting feature points of video frames, calculating the homographic transformation matrix between adjacent frames, generating a global transformation matrix for frame alignment, and combining adaptive detail grid division and uniform grid division to generate a comprehensive grid block set, and using the weighted fusion recombination method for pixel-level fusion, outputting stable video frames.

Benefits of technology

The video frame image processing in violent shaking scenes is realized to avoid missing surveillance images and ensure port operation safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358419A_ABST
    Figure CN120358419A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video processing, and particularly discloses a video anti-shake processing method, a medium and equipment, and the method comprises the steps: carrying out the down-sampling of a first video frame, extracting feature points, calculating a homography transformation matrix between adjacent frames, and generating a global transformation matrix to achieve the frame alignment; generating a comprehensive grid block set covering the key area and the overall situation in combination with self-adaptive detailed grid division and uniform grid division, and marking the type of each grid block; and carrying out pixel-level fusion based on the distance weight and the type weight coefficient through a weighted fusion recombination method, and outputting a stable video frame. According to the method, through combination of global motion compensation and local grid optimization, video frame images in a violent shaking scene can be processed, picture content can be restored, stable video frame images can be obtained, monitoring picture missing is avoided, and port operation safety is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to a video anti-shake processing method, medium, and device. Background Art

[0002] In the field of port container transportation and logistics operations, by means of cameras installed on the container trolley frame, the entire process of container loading, unloading, stacking, and transfer operations can be monitored and recorded in real time, which is of great significance for ensuring operation safety and improving operation efficiency. However, compared with ordinary monitoring scenarios, the port container monitoring scenario has significant particularities: First, the scene presents fixed and periodic repetitive characteristics. Most of the monitoring cameras in the port container yard are fixedly installed, monitoring the same fixed area for a long time, and the picture content is highly repetitive and has strong predictability. This characteristic provides the possibility for the system to memorize and learn historical scenes and realize scene prediction.

[0003] Second, there are special shaking patterns. During port operations, when heavy mechanical equipment such as cranes and gantry cranes are running, sudden and large-amplitude shaking will be caused, and the shaking duration is usually less than 1 second. Different from common continuous slight jitters, such shaking is instantaneous and violent, which will cause the monitoring picture to become severely blurred instantly, greatly affecting the monitoring effect.

[0004] Third, the real-time monitoring requirement is strict. Given the strict requirements for safety and efficiency in port container operations, the monitoring system must achieve real-time processing of video streams. The traditional offline processing method has a high-latency solution and cannot meet the actual needs of port operations.

[0005] As Figure 1 shown, the blue part shows the shaking patterns targeted by the usual anti-shake algorithms, while the red part shows the sudden large-amplitude shaking in the port container operation scenario concerned in this application. There are significant differences in shaking characteristics and scene performances between the two, which results in the fact that traditional video anti-shake algorithms cannot be applied to the processing of video pictures in the port container monitoring scenario, easily leading to the missing of monitoring pictures and leaving potential safety hazards. Summary of the Invention

[0006] In view of the above problems, the present invention provides a video anti-shake processing method, medium, and device to solve the technical problems that the existing video anti-shake algorithms cannot be applied to the anti-shake processing of video pictures with severe shaking, cannot meet the application in the port container monitoring scenario, and are prone to missing monitoring pictures and affecting safe operations.

[0007] To achieve the above object, in a first aspect, the present application provides a video anti-shake processing method, and the method includes the following steps: S1: Receive the first video frame, downsample the first video frame, extract feature points from the downsampled frame, calculate the homography transformation matrix between adjacent frames based on the feature points, cumulatively generate the global transformation matrix relative to the reference frame, and perform an alignment operation on the first video frame according to the global transformation matrix to obtain the second video frame; S2: Adopt an adaptive detailed grid division method to generate the first set of grid blocks based on the corner detection algorithm, and the first set of grid blocks covers the key areas on the second video frame; S3: Generate the second set of grid blocks with global coverage using the uniform grid division method; S4: Take the union of the first set of grid blocks and the second set of grid blocks to generate the comprehensive set of grid blocks, add type tags to all the grid blocks in the comprehensive set of grid blocks, and output the list of grid block information corresponding to all the grid blocks in the comprehensive set of grid blocks; S5: According to the list of grid block information, apply the weighted fusion and recombination method to perform pixel-level fusion based on the distance weight function and the type tag weight coefficient, and output the stable video frame.

[0008] Furthermore, step S1 includes: Downsample the first video frame to , and the downsampling ratio , ; Obtain the consecutive frames and after downsampling, extract feature points and establish the corresponding relationship between pixel points ; Use the RANSAC algorithm to estimate the homography transformation matrix such that the pixels of the consecutive frames after downsampling satisfy: , where represents the geometric transformation from frame t to frame t + 1; Select the reference frame , and for the frame t after downsampling, calculate its cumulative transformation matrix relative to the reference frame, and the calculation formula is as follows: ; Convert the cumulative transformation matrix at the downsampling resolution to the transformation matrix at the original resolution. For the first video frame at the original resolution, apply the transformation matrix for alignment to obtain the second video frame , and the calculation formula is as follows: .

[0009] Further, denote the image size of the second video frame as , the grid block size as , and the allowable overlap size between grid blocks as A. Then calculate the minimum distance of feature points . Step S2 includes: Convert the second video frame into a grayscale image; Use the Shi-Tomasi corner detection algorithm to detect in the grayscale image feature points, and represent the set of feature points as , where each point , and the Euclidean distance between any two points and is at least ; For each feature point , generate a grid block centered at with a size of . The coverage area of the grid block is: ; Only retain the grid blocks that are completely within the boundaries of the grayscale image. The center point of the grid block satisfies the following conditions: and . Output the coordinate list of all valid grid block center points to obtain the first grid block set.

[0010] Further, step S3 includes: Place the second video frame in a two-dimensional spatial coordinate system and calculate the center of the grid block on the X-axis: until , and calculate the center of the grid block on the Y-axis: until , ensuring that: and so that the grid blocks at the edges can completely cover the edges of the second video frame, obtaining the second grid block set: ; For each pixel in the image, calculate the number of grid blocks covering the pixel, create a list for each pixel to record the indices of all grid blocks covering the pixel, and obtain the pixel-to-grid block coverage mapping map; Output the second grid block set and the pixel-to-grid block coverage mapping map.

[0011] Further, the size A of the overlapping area between adjacent grid blocks is dynamically adjusted according to the shaking intensity of the video frame. The adjustment formula is as follows: ; where is the base overlap size; is the shaking intensity factor, is the adjustment coefficient.

[0012] Further, step S4 includes: Receiving the list of center coordinates of the first set of grid blocks output by step S2 and the list of center coordinates of the second set of grid blocks output by step S3 , taking the union of and to obtain a comprehensive set of grid blocks, and adding type tags to all grid blocks in the comprehensive set of grid blocks, where the type tag of the grid blocks originally belonging to the first set of grid blocks is the adaptive grid type tag, and the type tag of the grid blocks originally belonging to the second set of grid blocks and not belonging to the first set of grid blocks is the uniform grid type tag; Outputting the merged list of grid block information, where the list of grid block information includes the center coordinates of the grid blocks, at least two edge coordinates of the grid blocks, and the type tags corresponding to the grid blocks.

[0013] Further, step S5 includes: Each pixel position , calculating according to the following formula to output a stable video frame: ; where, is the pixel value at position on the stable video frame, is the pixel value of the th grid block at position , is the distance-based weight function value, is the weight adjustment coefficient of the grid block type, is the total number of grid blocks covering position ; where, , d is the Euclidean distance from pixel to the center of the grid block, and K is the half-width of the overlapping area A of the grid block.

[0014] Further, the weight adjustment coefficient of the grid block with the uniform grid type tag .

[0015] In a second aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the video anti-shake processing method as described in the first aspect of the present application.

[0016] In a third aspect, the present application provides an electronic device storing a computer program, including a processor and a storage medium. The storage medium stores a computer program, and when the computer program is executed by the processor, it implements the video anti-shake processing method as described in the first aspect of the present application.

[0017] Different from the prior art, the above technical solution provides a video anti-shake processing method, medium and device. The method includes: extracting feature points by downsampling a first video frame, calculating a homography transformation matrix between adjacent frames, and generating a global transformation matrix to achieve frame alignment; combining adaptive detail grid division and uniform grid division to generate a set of comprehensive grid blocks covering key areas and the whole, and marking the types of each grid block; through a weighted fusion and recombination method, performing pixel-level fusion based on distance weights and type weight coefficients, and outputting a stable video frame. By combining global motion compensation and local grid optimization, the method can process video frame images in a violently shaking scene, restore the picture content, obtain a stable video frame image, avoid the missing of surveillance pictures, and ensure the safety of port operations.

[0018] The above relevant descriptions of the invention content are only an overview of the technical solution of the present invention. In order to enable those of ordinary skill in the art to more clearly understand the technical solution of the present invention, and then to implement it according to the content recorded in the description and the drawings, and in order to make the above objects, other objects, features and advantages of the present invention more easily understood, the following is described in conjunction with the specific embodiments and drawings of the present invention. Description of the Drawings

[0019] The drawings are only used to show the principles, implementation methods, applications, features and effects of the specific embodiments of the present invention and other related contents, and should not be regarded as a limitation of the present invention.

[0020] In the drawings of the specification: Figure 1 It is a comparison schematic diagram of the shaking modes corresponding to the conventional anti-shake algorithm and the anti-shake algorithm involved in the present application; Figure 2 It is a first flowchart of the video anti-shake processing method involved in the specific embodiment; Figure 3 It is a second flowchart of the video anti-shake processing method involved in the specific embodiment; Figure 4 It is a distribution schematic diagram of the feature points detected by the corner detection algorithm involved in the specific embodiment; Figure 5 It is a schematic diagram of the grid block distribution in the first set of grid blocks involved in the specific embodiment; Figure 6 It is a flowchart of the specific implementation steps of step S2 involved in the specific embodiment; Figure 7 Schematic diagram of the distribution of the central coordinate points of the grid blocks involved in the specific implementation manner; Figure 8 Schematic diagram of the distribution of the grid blocks in the second grid block set involved in the specific implementation manner; Figure 9 Flowchart of the specific implementation steps of step S3 involved in the specific implementation manner; Figure 10 Flowchart of the specific implementation steps of step S4 involved in the specific implementation manner; Figure 11 Schematic diagram of the distribution after the central points of the uniform grid blocks and the adaptive grid blocks involved in the specific implementation manner are fused; Figure 12 Schematic diagram of the distribution of the grid blocks in the comprehensive grid block set involved in the specific implementation manner; Figure 13 Flowchart of the specific implementation steps of step S5 involved in the specific implementation manner; Figure 14 Module schematic diagram of the electronic device described in the specific implementation manner; The descriptions of the reference numerals involved in the above-mentioned various drawings are as follows: 10. Electronic device; 101. Processor; 102. Storage medium. Specific implementation manner

[0021] To describe in detail the possible application scenarios, technical principles, implementable specific solutions, achievable purposes and effects of the present invention, etc., the following will be described in detail in conjunction with the listed specific embodiments and with reference to the accompanying drawings. The embodiments described herein are only used to more clearly illustrate the technical solutions of the present invention, and therefore are only examples and cannot be used to limit the protection scope of the present invention.

[0022] Referring to "embodiment" in this article means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present invention. The term "embodiment" that appears in various positions in the specification does not necessarily refer to the same embodiment, nor does it particularly limit its independence or relevance to other embodiments. In principle, in the present invention, as long as there is no technical contradiction or conflict, the various technical features mentioned in each embodiment can be combined in any way to form the corresponding implementable technical solutions.

[0023] Unless otherwise defined, the meanings of the technical terms used herein are the same as those generally understood by those skilled in the technical field to which the present invention belongs; the use of the relevant terms herein is only for describing specific embodiments and is not intended to limit the present invention.

[0024] In the description of the present invention, the term "and / or" is an expression used to describe the logical relationship between objects, indicating that there can be three relationships. For example, A and / or B means: there is A, there is B, and there is both A and B at the same time. In addition, the character " / " in this text generally represents an "or" logical relationship between the associated objects before and after.

[0025] In the present invention, terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantitative, primary-secondary, or sequential relationships between these entities or operations.

[0026] Without further limitation, in the present invention, the terms "comprising", "including", "having", or other similar expressions used in a statement are intended to cover non-exclusive inclusion. These expressions do not exclude the possibility that there may be additional elements in the process, method, or product that includes the said elements. Thus, a process, method, or product that includes a series of elements may include not only those defined elements, but also other elements not explicitly listed, or elements inherent to such a process, method, or product.

[0027] In the present invention, expressions such as "greater than", "less than", "exceeding", etc. are understood not to include the number itself; expressions such as "above", "below", "within", etc. are understood to include the number itself. In addition, in the description of the embodiments of the present invention, the meaning of "a plurality of" is two or more (including two). Similar expressions related to "many", such as "multiple groups", "multiple times", etc., are understood in the same way, unless otherwise specifically defined.

[0028] In the description of the embodiments of the present invention, the spatially related expressions used, such as "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "perpendicular", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the specific embodiment or the drawing. It is only for the convenience of describing the specific embodiments of the present invention or for the reader's understanding, and does not indicate or imply that the device or component referred to must have a specific position, a specific orientation, or be constructed or operated in a specific orientation. Therefore, it should not be construed as a limitation to the embodiments of the present invention.

[0029] Unless otherwise clearly specified or defined, in the description of the embodiments of the present invention, terms such as "installed", "connected", "linked", "fixed", "set", etc. shall be understood in a broad sense. For example, the "connection" may be a fixed connection, a detachable connection, or an integral setting; it may be a mechanical connection, an electrical connection, or a communication connection; it may be a direct connection, or an indirect connection through an intermediate medium; it may be the communication inside two components or the interaction relationship between two components. For those skilled in the art to which the present invention pertains, the specific meanings of the above terms in the embodiments of the present invention can be understood according to specific circumstances.

[0030] As Figure 2 shown, in the first aspect, the present application provides a video anti-shake processing method, including the following steps: S1: Receive a first video frame, downsample the first video frame, extract feature points from the downsampled frame, calculate the homography transformation matrix between adjacent frames based on the feature points, cumulatively generate a global transformation matrix relative to a reference frame, and perform an alignment operation on the first video frame according to the global transformation matrix to obtain a second video frame; S2: Adopt an adaptive detailed grid division method to generate a first grid block set based on a corner detection algorithm, and the first grid block set covers key areas on the second video frame; S3: Adopt a uniform grid division method to generate a second grid block set covering the whole; S4: Take the union of the first grid block set and the second grid block set to generate a comprehensive grid block set, add type tags to all grid blocks in the comprehensive grid block set, and output a grid block information list corresponding to all grid blocks in the comprehensive grid block set; S5: According to the grid block information list, apply a weighted fusion and recombination method to perform pixel-level fusion based on a distance weight function and a type tag weight coefficient, and output a stable video frame.

[0031] The above solution combines global motion compensation and local grid optimization, can process video frame images in a violently shaking scene, restore the picture content, obtain a stable video frame image, avoid the lack of monitoring pictures, and ensure the safety of port operations.

[0032] As Figures 3 - 13 shown, in some embodiments, step S1 includes: Downsample the first video frame to , with a downsampling ratio , ; Obtain consecutive frames and , extract feature points and establish the correspondence between pixel points ; Use the RANSAC algorithm to estimate the homography transformation matrix , such that the pixels of the downsampled consecutive frames satisfy: , where represents the geometric transformation from frame t to frame t + 1; Select a reference frame , for the downsampled frame t, calculate its cumulative transformation matrix relative to the reference frame , and the calculation formula is as follows: ; Convert the cumulative transformation matrix at the downsampled resolution to the transformation matrix at the original resolution , for the first video frame at the original resolution , apply the transformation matrix for alignment to obtain the second video frame , and the calculation formula is as follows: .

[0033] The above steps can be implemented by a fast anti-shake preprocessing module. The fast anti-shake preprocessing module is designed specifically for the instantaneous severe shaking problem in port container monitoring videos (usually with a duration less than 1 second). This module uses image downsampling and transformation matrix estimation techniques to achieve efficient preliminary global alignment, laying a foundation for subsequent fine grid processing. Its key features include: significantly reducing the computational complexity using the downsampling strategy, accurately estimating the geometric transformation relationship between consecutive frames, accumulating the transformation matrix to represent the total displacement relative to the stable state, and applying the scaled transformation matrix to align the original resolution frames. Different from traditional video anti-shake algorithms, this method does not require complex matrix path smoothing processing because the shaking that occurs in port container monitoring is of a transient nature, and the system will naturally return to a stable state within an instant.

[0034] Taking the size of the original port monitoring video frame (i.e., the first video frame) as (such as 1920×1080) as an example, the size of the downsampled video frame can be set to , that is, while keeping the aspect ratio of the original video frame unchanged, the length and width of the original video frame are reduced as a whole.

[0035] The RANSAC algorithm estimates the model parameters by randomly sampling a subset of data, and uses inliers (data points that conform to the model) and outliers (data points that do not conform to the model) to optimize the model, so as to accurately estimate the homography transformation matrix in the presence of noise and outliers. In this embodiment, is preferably matrix.

[0036] In this embodiment, the reference frame is the stable frame before shaking. For the frames restored after shaking, the cumulative matrix will naturally approach the identity transformation matrix, indicating that the system returns to the stable state. The cumulative transformation matrix at the downsampled resolution is converted to the transformation matrix at the original resolution which can be achieved through the following steps: For the translation component in the perspective transformation matrix, perform the following transformation and keep other matrix elements unchanged: ; .

[0037] In this embodiment, after aligning the applied transformation matrix bilinear or bicubic interpolation can also be used to ensure the quality of the transformed image.

[0038] After processing through the above steps, an aligned sequence of frames at the original resolution can be generated while maintaining the details of the original image. These aligned frames (i.e., the second video frames) can be used as the input to the subsequent mesh block division and processing module.

[0039] Based on the fast anti-shake preprocessing module to process the instantaneous severe shaking scenes in the port container monitoring video, it has the following advantages: (1) Computational efficiency: The computational complexity can be reduced by about times through the downsampling strategy; for 1080P input, the compression factor is times; significantly reducing the computational burden of feature extraction, matching, and matrix estimation.

[0040] (2) Fast response ability: It can adapt to the characteristics of short-term severe shaking (usually the timing time < 1 second) in the port environment; without the need for complex path smoothing algorithms, simplifying the calculation process; the matrix accumulation naturally captures the total displacement relative to the stable state.

[0041] (3) Spatiotemporal coherence guarantee: Global alignment ensures that the subsequent divided mesh blocks are coherent in the time dimension, enabling the mesh blocks to contain meaningful spatiotemporal information of the front and back frames.

[0042] (4) Resolution adaptability: The system can process various port monitoring resolutions from 720P to 4K; the downsampling ratio can be flexibly adjusted according to the actual computing resources and real-time requirements.

[0043] In some embodiments, as Figures 4 - 6 shown, denote the image size of the second video frame as , and the mesh block size as If the allowable overlap size between grid blocks is A, then calculate the minimum distance of feature points Step S2 includes: Convert the second video frame into a grayscale image; Use the Shi-Tomasi corner detection algorithm to detect in the grayscale image feature points, and the feature point set is represented as where each point For any two points and the Euclidean distance between them is at least This can ensure a reasonable distribution of feature points in the port environment; For each feature point generate a grid block centered on with a size of The coverage area of the grid block is: ; Only keep the grid blocks that are completely within the boundary of the grayscale image. The center point of the grid block satisfies the following conditions: and Output the coordinate list of the center points of all valid grid blocks to obtain the first grid block set.

[0044] Adaptive detail grid division is particularly suitable for the key areas (such as container identification, operating equipment, etc.) of the anti-shake algorithm for port container monitoring videos. This method is based on the feature point detection algorithm to identify the points with significant features in the video. For example, the maximum number of feature points can be set to 200, the quality threshold can be set to 0.01 (the lower this value, the more feature points are detected), and the minimum distance between feature points is The distribution of the obtained feature points is as shown in Figure 4 Then, generate grid blocks of a fixed size centered on these feature points to achieve precise coverage of the key areas of the port scene. The distribution of the generated grid blocks is as shown in Figure 5 shown.

[0045] When performing boundary inspection on the port container monitoring video image, only keep the grid blocks that are completely within the boundary of the port monitoring image. Specifically, for each detected feature point calculate the boundary coordinates of the grid block, that is, the upper left corner coordinate and the lower right corner coordinate Compare these two boundary coordinates of the grid block with the boundary coordinates of the video image to determine whether the current grid block is completely within the boundary of the port container monitoring video image. If it is within the boundary, then Add it to the list of the central coordinates of the grid blocks, and then return the list of the coordinates of the central points of all valid grid blocks . The entire implementation process of step S2 refers to Figure 6 as shown below

[0046] In some embodiments, step S3 includes: Place the second video frame in the two-dimensional space coordinate system, and calculate the center of the grid blocks on the X-axis: until , and calculate the center of the grid blocks on the Y-axis: until , ensuring that: and , so that the grid blocks at the edges can completely cover the edges of the second video frame, obtaining the second set of grid blocks: ; For each pixel in the image , calculate the number of grid blocks covering the pixel, create a list for each pixel, record the indices of all grid blocks covering the pixel, and obtain the coverage mapping map from pixels to grid blocks; Output the second set of grid blocks and the coverage mapping map from pixels to grid blocks

[0047] In this embodiment, uniform grid division is a systematic image processing technique for ensuring complete coverage of port surveillance videos. Different from adaptive segmentation, uniform grid division emphasizes global balance, ensuring that each pixel in the surveillance image is covered by an appropriate number of grid blocks, and is particularly suitable for processing large areas and backgrounds in the port container operation environment

[0048] In this embodiment, the grid block size parameter can be set to 64 or 128 pixels. The distribution of the midpoints of the grid blocks obtained by the above method is as Figure 7 shown, and the corresponding grid distribution schematic diagram is as Figure 8 shown. Through boundary processing, it can be ensured that the grid blocks at the grid edges can completely cover the boundaries of the port container surveillance video image. If the edge of the last grid block does not reach the image boundary, additional grid block centers are added to ensure complete coverage. The entire implementation process of step S2 refers to Figure 9 as shown below

[0049] In some embodiments, the size A of the overlapping area between adjacent grid blocks is dynamically adjusted according to the shaking intensity of the video frame, and the adjustment formula is as follows: ; where is the basic overlapping size; is the shaking intensity factor, is the adjustment coefficient

[0050] When performing anti - shake processing on the port container monitoring video image, the parameter selection has an important impact on the system performance. The following are the design principles of the key parameters: (1) Adaptive grid block size: The basic size is pixels (preferably or or ), and the specific size selection of S needs to be adapted to the corresponding anti - shake algorithm.

[0051] (2) Intelligent overlap mechanism: The size A of the overlapping area of adjacent grid blocks (preferably the basic value A = 16 or 32), and it can be dynamically adjusted according to the shaking intensity. The value range of the shaking intensity factor M is preferably [0, 1], and the adjustment coefficient is preferably 0.5.

[0052] Taking different common shaking situations in port container monitoring as an example, when setting pixels, , the specific adjustment calculation method of the size A of the grid block overlapping area is as follows: Slight shaking ( ): pixels; Medium shaking ( ): pixels; Severe shaking ( ): pixels; The shaking intensity is determined by analyzing the magnitude of the optical flow field between adjacent frames. For the severe shaking generated by gantry cranes and cranes in port operations, the system will automatically increase the overlapping area to improve the jitter compensation effect.

[0053] Taking the port container monitoring video image with a resolution of 1080P (1920×1080 pixels) as an example, when S = 64 (block size), A = 16 (overlapping area size), and the effective step size = S - A = 48 pixels, then when the image is divided into grid blocks, the following specific results can be obtained: The number of blocks per row in the horizontal direction = 1920 / 48 = 40 blocks; The number of blocks per column in the vertical direction = 1080 / 48 = 23 blocks; The total number of blocks = 40×23 = 920 blocks.

[0054] When S = 128 (i.e., larger block) and A = 32 (i.e., larger overlapping area), at this time the effective step size = 96 pixels, then when the image is divided into grid blocks, the following specific results can be obtained: Total number of blocks = (1920 / 96) x (1080 / 96) = 20 × 12 = 240 blocks.

[0055] In some embodiments, step S4 includes: Receiving the list of central coordinates of the first set of grid blocks output by step S2 and the list of central coordinates of the second set of grid blocks output by step S3 , and taking the union of them to obtain a comprehensive set of grid blocks, and adding type tags to all the grid blocks in the comprehensive set of grid blocks, where the type tag of the grid blocks originally belonging to the first set of grid blocks is the adaptive grid type tag, and the type tag of the grid blocks originally belonging to the second set of grid blocks and not belonging to the first set of grid blocks is the uniform grid type tag; Outputting the merged list of grid block information, where the list of grid block information includes the central coordinates of the grid blocks, at least two edge coordinates of the grid blocks, and the type tags corresponding to the grid blocks.

[0056] In this embodiment, the adaptive grid block and uniform grid block merging module directly combines the results of the two grid division strategies without complex redundancy removal. The adaptive grid blocks correspond to the feature-rich areas (such as container identification, operating equipment, etc.) in the port container monitoring video image, while the uniform grid blocks ensure the complete coverage of the monitoring screen. By simply and directly merging these two types of grid blocks, the system can provide more refined processing for key areas while ensuring global coverage, thereby improving the overall anti-shake effect.

[0057] In the actual port monitoring scene image, the number of adaptive grid blocks (based on feature points) is usually much less than that of uniform grid blocks, generally only accounting for 20% - 30% of the total number of blocks. This moderate redundancy will not significantly increase the computational burden. For the key areas (such as container identification, operating equipment, etc.) in the port monitoring video image, the repeated coverage of multiple grid blocks actually helps to improve the robustness and accuracy of the anti-shake processing. Removing the redundancy detection step greatly simplifies the algorithm process, reduces the system complexity, and improves the processing efficiency and maintainability.

[0058] In the weighted fusion and recombination stage, the present application designs a weight mechanism dedicated to processing multi-grid block coverage, which can naturally handle the redundant coverage situation.

[0059] Assume that the set of central coordinates of the grid blocks generated by adaptive grid division is , and the set of central coordinates of the grid blocks generated by uniform grid division is , then the merging operation is simplified to: The set of central coordinates of the comprehensive grid blocks ; Total number of merged grid blocks , and add a type tag to each grid block, where the type tag of the adaptive grid block is 1 and the type tag of the uniform grid block is 0.

[0060] Taking the port container surveillance video image in 1080P (resolution 1920×1080 pixels) as an example, assuming the grid block size S = 128 pixels and the overlapping area size A = 32 pixels, 150 grid blocks concentrated around the port containers and equipment (i.e., covering the key areas) and 240 uniform grid blocks covering the entire surveillance screen can be obtained through adaptive grid division. Merging these two types of grid blocks can get 390 grid blocks.

[0061] The complete implementation process of step S4 is as Figure 10 shown. As Figure 11 and Figure 12 shown, through this simple and effective direct merging strategy, the system makes full use of the accurate feature capture ability of the adaptive grid blocks and the comprehensive coverage characteristics of the uniform grid blocks, providing an optimal grid division basis for subsequent anti-shake processing. In actual port surveillance applications, the efficiency and robustness of this method can effectively handle various complex shaking scenarios. From Figure 12 , it is not difficult to see that the average coverage of the key area of the image is 4 - 6 grid blocks / pixel, indicating that this area is covered by both adaptive grid blocks and uniform grid blocks, while the average coverage of the non-key area (i.e., the background area) is 2 - 3 grid blocks / pixel, indicating that this area is mainly covered by uniform grid blocks.

[0062] In some embodiments, step S5 includes: Each pixel position , calculates according to the following formula and outputs a stable video frame: ; where, is the pixel value at position on the stable video frame, is the pixel value of the th grid block at position , is the distance-based weight function value, is the weight adjustment coefficient of the grid block type, is the total number of grid blocks covering position ; where, , d is the Euclidean distance from pixel to the center of the grid block, and K is the half-width of the grid block overlapping area A.

[0063] Weighted fusion recombination is an advanced image reconstruction method used to merge overlapping grid blocks that have undergone independent anti-shake processing in port container surveillance video images into a seamless and high-quality overall image. This algorithm specifically considers the characteristics of two different types of grid blocks, adaptive and uniform grids, and ensures the best fusion effect through intelligent weight assignment, eliminating the blurring of the video caused by severe shaking during port operations.

[0064] The weight adjustment coefficient of the grid block marked with the uniform grid type , and the weight adjustment coefficient of the grid block marked with the adaptive grid type . Adaptive grid blocks are given higher weights because they usually cover areas in port container surveillance video images that contain more details and textures. By assigning higher weights, these details and textures can be better preserved.

[0065] Specifically, the weight can be set to the maximum at the center of the grid block , and the weight can be set to the minimum at the edge of the grid block . A smooth transition occurs in the intermediate area to ensure seamless fusion of each grid block in the port container surveillance video image.

[0066] In this embodiment, for each pixel position in the port container surveillance video image , a list of all grid block indices covering this pixel can be obtained. Then, for each grid block that covers this pixel obtained, the distance from the pixel to the center of the grid block is calculated: . At the same time, the distance weight is calculated, and the weight adjustment coefficient of the grid block type is determined . Higher weights are given to adaptive grid blocks, and the calculation process is as follows: Accumulate the weighted pixel values: ; Accumulate the weights: ; Calculate the final pixel value: ; For each frame image in the port container surveillance video sequence, by updating the position information of the adaptive grid blocks (i.e., tracking the positions of key feature points), the mapping relationship between the pixels and the grid blocks is recalculated to obtain the final stable frame image.

[0067] For example, for the pixel position in the port container surveillance video image, assume that this pixel is covered by two adaptive grid blocks and two uniform grid blocks, specifically as follows: Adaptive grid block 1 (this grid block covers the container identification): The center coordinates are , distance , , , pixel value ; Adaptive grid block 2 (this gateway goose block covers the feature area): The center coordinates are , distance , , , pixel value ; Uniform grid block 1 (this grid block covers the background area): The center coordinates are , distance , , , pixel value ; Uniform grid block 2 (this grid block covers the background area): The center coordinates are , distance , , , pixel value ; According to the foregoing formula, the final pixel value corresponding to this pixel on the stable frame image can be calculated as: .

[0068] The specific implementation process of step S5 is as Figure 13 shown. This step combines the distance weight and the grid block type weight, which can more precisely control the fusion process, adaptively balance the detail retention and smooth transition in the port surveillance video image. Specifically, it provides the maximum weight at the center of the grid block to ensure the retention of key features in port surveillance, and smoothly decays to zero at the edge of the grid block to avoid visible splicing traces. Intelligently allocate importance weights according to the grid block characteristics, strengthen the detail-rich areas in the port surveillance video while maintaining overall consistency, and can handle the grid block changes in the moving scenes of port container surveillance, adapt to different inter-frame feature point changes, and maintain stable output.

[0069] The anti-shake method for port container surveillance video proposed in this application effectively solves the problem of severe shaking caused by the operation of heavy equipment such as cranes and gantry cranes in the port operation environment, resulting in blurred surveillance images and missing images that are difficult to trace. This method realizes a high-quality video stabilization effect by intelligently identifying the shaking intensity, dynamically adjusting the grid parameters, and combining the distance weight function and the grid block type weight.

[0070] In a second aspect, the present invention further provides a computer-readable storage medium having stored thereon a computer program, which when executed by a processor implements the video anti-shake processing method as described in the first aspect of the present invention.

[0071] Among them, the computer-readable storage medium may be a volatile memory or a non-volatile memory, or may include both a volatile and a non-volatile memory.

[0072] The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD ROM); the magnetic surface memory may be a disk memory or a tape memory.

[0073] The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), direct rambus random access memory (DRRAM). The computer-readable storage medium described in the embodiments of the present invention is intended to include these and any other suitable types of memory.

[0074] As Figure 14 shown, in a third aspect, the present invention provides an electronic device 10, including a processor 101 and a storage medium 102, where a computer program is stored on the storage medium, and when the computer program is executed by the processor, the video anti-shake processing method described in the first aspect of the present invention is implemented.

[0075] In some embodiments, the processor may be implemented by software, hardware, firmware, or a combination thereof, and at least one of a circuit, a single or multiple application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), central processing units (CPUs), controllers, microcontrollers, and microprocessors may be used, so that the processor can execute some steps, all steps, or any combination of the steps in the video anti-shake processing method described in the various embodiments of the present application.

[0076] Finally, it should be noted that although the above embodiments have been described in the text and drawings of the specification of the present application, the patent protection scope of the present application cannot be limited thereby. Any technical solutions obtained by equivalent structure or equivalent process substitution or modification based on the substantial concept of the present application and using the content recorded in the text and drawings of the specification of the present application, as well as any technical solutions directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, are all included in the patent protection scope of the present application.

Claims

1. A video anti-shake processing method, characterized in that, Including the following steps: S1: Receive a first video frame, downsample the first video frame, extract feature points from the downsampled frame, calculate a homography transformation matrix between adjacent frames based on the feature points, cumulatively generate a global transformation matrix relative to a reference frame, and perform an alignment operation on the first video frame according to the global transformation matrix to obtain a second video frame; S2: Adopt an adaptive detail grid division method to generate a first set of grid blocks based on a corner detection algorithm, where the first set of grid blocks covers key regions on the second video frame; S3: Adopt a uniform grid division method to generate a second set of grid blocks that globally covers; S4: Take the union of the first set of grid blocks and the second set of grid blocks to generate a comprehensive set of grid blocks, add type tags to all grid blocks in the comprehensive set of grid blocks, and output a list of grid block information corresponding to all grid blocks in the comprehensive set of grid blocks; S5: According to the list of grid block information, apply a weighted fusion and recombination method to perform pixel-level fusion based on a distance weight function and a type tag weight coefficient, and output a stable video frame.

2. The video anti-shake processing method according to claim 1, wherein Step S1 includes: Downsample the first video frame to , with a downsampling ratio of , ; Obtain the downsampled consecutive frames and , extract feature points and establish the corresponding relationships between pixel points ; Estimate the homography transformation matrix using the RANSAC algorithm such that the pixels of consecutive downsampled frames satisfy: where, represents the geometric transformation from frame t to frame t+1; Select a reference frame , for the downsampled frame t, calculate its cumulative transformation matrix relative to the reference frame , and the calculation formula is as follows: ; Convert the cumulative transformation matrix at the downsampled resolution to the transformation matrix at the original resolution . For the first video frame at the original resolution , apply the transformation matrix for alignment to obtain the second video frame . The calculation formula is as follows: .

3. The video anti-shake processing method according to claim 2, wherein Denote the image size of the second video frame as , the grid block size as , and the allowable overlap size between grid blocks as A. Then calculate the minimum distance of feature points . Step S2 includes: Convert the second video frame into a grayscale image; Detect feature points in the grayscale image using the Shi-Tomasi corner detection algorithm The set of feature points is denoted as , where each point , and the Euclidean distance between any two points and is at least ; For each feature point , generate a -centered -sized grid block, and the covered area of the grid block is: ; Only retain the grid blocks that are completely within the boundaries of the grayscale image, and the center points of the grid blocks meet the following conditions: and , output a list of the coordinates of the center points of all valid grid blocks , and obtain the first set of grid blocks.

4. The video anti-shake processing method according to claim 3, characterized in that Step S3 includes: Place the second video frame in a two-dimensional spatial coordinate system and calculate the centers of the grid blocks on the X-axis: until , and calculate the centers of the grid blocks on the Y-axis: until , ensuring that: and , so that the grid blocks at the edges can completely cover the edges of the second video frame, obtaining the second set of grid blocks: ; For each pixel in the image , calculate the number of grid blocks covering the pixel, create a list for each pixel to record the indices of all grid blocks covering the pixel, and obtain the coverage mapping map from pixels to grid blocks; Output the second set of grid blocks and the coverage mapping map from pixels to grid blocks.

5. The video anti-shake processing method according to claim 3 or 4, characterized in that, The size A of the overlapping region between adjacent grid blocks is dynamically adjusted according to the shaking intensity of the video frame, and the adjustment formula is as follows: ; Among them, is the basic overlapping size; is the shaking intensity factor, is the adjustment coefficient.

6. The video anti-shake processing method according to claim 4, characterized in that Step S4 includes: Receive the list of center coordinates of the first set of grid blocks output by step S2 and the list of center coordinates of the second set of grid blocks output by step S3 , and for and take the union to obtain a comprehensive set of grid blocks, and add type markers to all grid blocks in the comprehensive set of grid blocks. Among them, the type marker of the grid blocks originally belonging to the first set of grid blocks is the adaptive grid type marker, and the type marker of the grid blocks originally belonging to the second set of grid blocks and not belonging to the first set of grid blocks is the uniform grid type marker; Output the merged list of grid block information, where the list of grid block information includes the center coordinates of the grid block, at least two edge coordinates of the grid block, and the type tag corresponding to the grid block.

7. The video anti-shake processing method according to claim 5, wherein, Step S5 includes: Each pixel position , is calculated according to the following formula to output a stable video frame: ; Among them, is the pixel value at position on the stable video frame, is the pixel value of the th grid block at position ; is the distance-based weight function value, is the weight adjustment coefficient of the grid block type, is the total number of grid blocks covering position . Among them, , d is the Euclidean distance from the pixel to the center of the grid block, and K is the half-width of the overlapping area A of the grid block.

8. The video anti-shake processing method according to claim 7, wherein Weight adjustment coefficient of grid blocks marked with uniform grid type , weight adjustment coefficient of grid blocks marked with adaptive grid type .

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the video anti-shake processing method according to any one of claims 1 to 8.

10. An electronic device on which a computer program is stored, characterized in that, Including a processor and a storage medium, where a computer program is stored on the storage medium, and when the computer program is executed by the processor, it implements the video anti-shake processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Video image stabilization method based on feature tracking and grid path motion

    CN110753181A

  • Video image stabilization method and device, equipment and storage medium

    CN114095659A

  • Video anti-shake system for eliminating interference of dynamic feature points

    CN119211730A

  • Carbon steel pipe welding quality evaluation system

    CN119747954A

  • Method and system for real-time geo referencing stabilization

    US20240098367A1

Cited By

  • Finished product camera video real-time splicing method and system based on heterogeneous calculation

    CN121788342A

  • Real-time stitching method and system for finished camera video based on heterogeneous computing

    CN121788342B