Law enforcement recorder image anti-shake method and system based on visual inertia fusion

Through the visual inertial fusion method, the law enforcement recorder achieves accurate compensation for complex motion, solves the problems of image artifacts and errors in traditional methods, and improves the stability of the image and the reliability of evidence.

CN120602779AInactive Publication Date: 2025-09-05SHENZHEN YULONG MOBILE INTERNET
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511102077.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-09-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The image anti-shake method of traditional law enforcement recorders cannot accurately describe complex three-dimensional motion, resulting in artifacts and cumulative errors in the image, affecting the objectivity and reliability of the evidence.

Method used

Using a method based on visual inertial fusion, the image sequence and inertial measurement data are collected, multi-scale feature extraction and motion inertial estimation are performed, and the visual jitter features and inertial compensation constraints are determined to achieve the confidence fusion compensation of the image.

Benefits of technology

Improves the anti-shake robustness of law enforcement recorders in complex scenarios, ensures image stability and reliability of evidence, and maintains smooth transitions and rapid recovery in the event of sudden movement or sensor failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602779A_ABST
    Figure CN120602779A_ABST
Patent Text Reader

Abstract

The invention provides a law enforcement recorder image anti-shake method and system based on visual inertia fusion, and the method comprises the steps: carrying out the motion inertia estimation of each shake feature point in an image frame, obtaining an inertia matching relation between adjacent image frames in a law enforcement recorder, and determining the visual shake features of the image frames in the law enforcement recorder according to the inertia matching relation; performing inertial constraint compensation on the image frame in the law enforcement recorder through inertial compensation constraint between the visual characteristic and the inertial characteristic and the visual jitter characteristic to obtain a plurality of motion compensation components of the image frame in the law enforcement recorder, and further determining the constraint confidence of each motion compensation component; and performing jitter restoration on the next frame of image in the law enforcement recorder based on the constraint confidence of each motion compensation component to obtain a restored image of the next frame of image in the law enforcement recorder. Based on the scheme, confidence fusion compensation of the law enforcement recorder image based on visual characteristics and inertial characteristics can be realized, so that the anti-shake robustness in a complex scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image stabilization, and more specifically, to a method and system for stabilizing the image of a law enforcement recorder based on visual-inertial fusion. Background Art

[0002] The anti-shake technology of law enforcement recorders originates from the rigid demand for image clarity in mobile law enforcement scenarios. Early law enforcement recorders caused severe image shaking due to movement while wearing the camera, sudden situations, etc., and key evidence was easily lost. With the maturity of artificial intelligence and micro-gimbal mechanical anti-shake technology, the industry has gradually evolved from electronic anti-shake to optical anti-shake and micro-gimbal mechanical anti-shake. The latter uses a gyroscope to compensate for movement in real time to achieve lossless and stable shooting of the original video, meeting the strict requirements of the judicial field for the authenticity of evidence.

[0003] Traditional anti-shake methods usually simplify the movement of law enforcement recorders into a global two-dimensional translation model. This simplified processing has significant defects in practical applications and cannot accurately describe the common complex three-dimensional movements of law enforcement recorders, especially the rotational jitter around the optical axis and the perspective changes caused by forward and backward movement. The simplified model will cause obvious artifacts in the compensated image. When the rotational motion is incorrectly compensated as translation, image shear distortion will occur; when the pitch / yaw motion is forcibly converted into a plane displacement, the content in the edge area will be distorted. When the law enforcement recorder has multiple degrees of freedom compound motion, the pure translation model will produce cumulative errors, which will eventually lead to a jello effect or deformation of key actions in the picture, directly affecting the objectivity and evidentiary effectiveness of the law enforcement record. Therefore, how to achieve confidence fusion compensation of law enforcement recorder images based on visual characteristics and inertial characteristics, so as to improve the anti-shake robustness in complex scenes has become a difficult problem faced by the industry. Summary of the Invention

[0004] The present application provides a method and system for stabilizing law enforcement recorder images based on visual-inertial fusion, which can realize confidence fusion compensation of law enforcement recorder images based on visual characteristics and inertial features, thereby improving the stabilization robustness in complex scenarios.

[0005] In a first aspect, the present application provides a method for image stabilization of law enforcement recorders based on visual-inertial fusion, comprising: Collecting an image sequence of a law enforcement recorder and inertial measurement data in an inertial unit, performing multi-scale feature extraction on the image sequence, and obtaining a plurality of jitter feature points of an image frame in the law enforcement recorder; Estimating the motion inertia of each jitter feature point in the image frame using the inertial measurement data to obtain an inertia matching relationship between adjacent image frames in the law enforcement recorder, and determining the visual jitter characteristics of the image frame in the law enforcement recorder based on the inertia matching relationship and the motion inertia parameters of each jitter feature point; Determining an inertia compensation constraint between visual characteristics and inertia features of an image frame from a law enforcement recorder during motion estimation, performing inertia constraint compensation on the image frame from the law enforcement recorder using the inertia compensation constraint and the visual jitter feature, obtaining multiple motion compensation components of the image frame from the law enforcement recorder, and then determining a constraint confidence level for each motion compensation component; Based on the constrained confidence of each motion compensation component, the jitter of the next frame image in the law enforcement recorder is repaired to obtain a repaired image of the next frame image in the law enforcement recorder.

[0006] In some embodiments, performing multi-scale feature extraction on the image sequence to obtain multiple jitter feature points of the image frame in the law enforcement recorder specifically includes: Performing multi-scale sampling on each image frame in the image sequence to obtain a Gaussian pyramid for each image frame; Determine multiple difference feature points between adjacent image frames in the law enforcement recorder through all Gaussian pyramids, and then determine the jitter degree of each difference feature point; According to all jitter degrees, multiple jitter feature points of the image frame in the law enforcement recorder are screened out from all difference feature points.

[0007] In some embodiments, the inertial measurement data is used to estimate the motion inertia of each jitter feature point in the image frame to obtain the inertial matching relationship between adjacent image frames in the law enforcement recorder, specifically including: Extracting angular velocity features and acceleration features of each jitter feature point in the image frame between adjacent image frames from the inertial measurement data; The inertia of each jitter feature point is estimated through all angular velocity features and acceleration features, and the inertia matching relationship between adjacent image frames in the law enforcement recorder is obtained.

[0008] In some embodiments, determining the visual jitter characteristics of the image frame in the law enforcement recorder based on the inertia matching relationship and the motion inertia parameters of each jitter feature point specifically includes: For each jitter feature point, determining the motion inertia parameter of the jitter feature point; Determining the visual jitter value of the jitter feature point by using the motion inertia parameter and the inertia matching relationship, and then obtaining the visual jitter value of each jitter feature point; The visual jitter characteristics of the image frames in the body camera are determined based on all the visual jitter values.

[0009] In some embodiments, determining the inertia compensation constraint between the visual characteristics and the inertial characteristics of the image frame of the body camera during the motion estimation process specifically includes: Initialize a motion estimation model based on the state vector; The displacement observation values ​​of the visual feature points of the law enforcement recorder's image frame during the motion estimation process are used as the visual feature input in the motion estimation model; The inertial measurement data of the image frame of the law enforcement recorder during the motion estimation process is used as the inertial feature input in the motion estimation model; The motion estimation model is used to perform nonlinear constraints on the relationship between the visual characteristics and the inertial characteristics, thereby obtaining inertia compensation constraints between the visual characteristics and the inertial characteristics.

[0010] In some embodiments, performing inertia constraint compensation on the image frames in the law enforcement recorder using the inertia compensation constraint and the visual jitter feature to obtain multiple motion compensation components of the image frames in the law enforcement recorder specifically includes: Extracting independent motion components of adjacent image frames in the body camera in each compensation dimension from the inertia compensation constraint; Determining a consistency score for each independent motion component using the compensation parameter of each independent motion component and the visual jitter feature; Multiple motion compensation components of the image frame in the body camera are screened out from all independent motion components through respective consistency scores.

[0011] In some embodiments, determining the constraint confidence of each motion compensation component specifically includes: For each motion compensation component, determining a compensation residual of the motion compensation component in each compensation dimension; All compensation residuals are mapped to the constrained confidence of the motion compensation components, thereby obtaining the constrained confidence of each motion compensation component.

[0012] In a second aspect, the present application provides an image stabilization system for a law enforcement recorder based on visual-inertial fusion, including a shake repair unit, wherein the shake repair unit includes: An acquisition module is used to acquire an image sequence of the law enforcement recorder and inertial measurement data in the inertial unit, perform multi-scale feature extraction on the image sequence, and obtain multiple jitter feature points of the image frame in the law enforcement recorder; a processing module configured to estimate the motion inertia of each jitter feature point in the image frame using the inertial measurement data, obtain an inertia matching relationship between adjacent image frames in the law enforcement recorder, and determine the visual jitter characteristics of the image frames in the law enforcement recorder based on the inertia matching relationship and the motion inertia parameters of each jitter feature point; The processing module is further configured to determine an inertia compensation constraint between visual characteristics and inertia features of an image frame of the law enforcement recorder during a motion estimation process, perform inertia constraint compensation on the image frame in the law enforcement recorder using the inertia compensation constraint and the visual jitter feature, obtain multiple motion compensation components of the image frame in the law enforcement recorder, and further determine a constraint confidence level of each motion compensation component; The execution module is used to perform jitter repair on the next frame image in the law enforcement recorder based on the constraint confidence of each motion compensation component to obtain a repaired image of the next frame image in the law enforcement recorder.

[0013] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the above-mentioned law enforcement recorder image stabilization method based on visual-inertial fusion.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium, which stores instructions or codes. When the instructions or codes are run on a computer, the computer implements the above-mentioned law enforcement recorder image stabilization method based on visual-inertial fusion when executing the instructions or codes.

[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: The present application provides an image stabilization method and system for law enforcement recorders based on visual-inertial fusion, which collects image sequences of law enforcement recorders and inertial measurement data from an inertial unit, performs multi-scale feature extraction on the image sequences, and obtains multiple jitter feature points of the image frames in the law enforcement recorder; performs motion inertia estimation on each jitter feature point in the image frame using the inertial measurement data to obtain an inertial matching relationship between adjacent image frames in the law enforcement recorder, and determines the visual jitter features of the image frames in the law enforcement recorder based on the inertial matching relationship and the motion inertia parameters of each jitter feature point; determines the inertia compensation constraints between the visual characteristics and inertial characteristics of the image frames of the law enforcement recorder during the motion estimation process, performs inertial constraint compensation on the image frames in the law enforcement recorder using the inertia compensation constraints and the visual jitter features, obtains multiple motion compensation components of the image frames in the law enforcement recorder, and further determines the constraint confidence of each motion compensation component; performs jitter repair on the next frame image in the law enforcement recorder based on the constraint confidence of each motion compensation component to obtain a repaired image of the next frame image in the law enforcement recorder.

[0016] It can be seen that in this application, the jitter repair of the next frame image in the law enforcement recorder is performed based on the constrained confidence of each motion compensation component to obtain a repaired image of the next frame image in the law enforcement recorder; first, the visual jitter characteristics are determined to accurately quantify the degree of motion abnormality of each jitter feature point, thereby providing a key basis for multi-sensor data fusion, and determining the visual jitter characteristics can accurately identify visual tracking abnormal points caused by environmental interference, effectively overcoming the defect that traditional pure visual methods are prone to mismatching in complex scenes. At the same time, by statistically analyzing the jitter distribution characteristics of all jitter feature points in the entire frame image, the global stability of the current picture can be comprehensively evaluated, providing data support for the dynamic adjustment of subsequent compensation strategies, ensuring that even in the case of failure to track some jitter feature points, the law enforcement recorder can still maintain stable anti-shake performance based on reliable visual jitter characteristics. Then, by determining the constraint confidence, the reliability quantitative index of each motion compensation component can be obtained, thereby realizing adaptive weighted fusion and fault-tolerant processing. The constraint confidence is determined by comprehensively considering multi-dimensional indicators such as the size of the compensation residual, time continuity and spatial consistency. It can dynamically reflect the credibility of each motion component in the current environment, and can automatically reduce the influence of unstable components while enhancing the contribution of high-reliability components, so that the law enforcement recorder can maintain smooth transition and rapid recovery when facing sudden violent movements or temporary failure of the sensor. At the same time, it can ensure that no secondary distortion caused by the accumulation of sensor errors will be introduced during the compensation process, and finally achieve a robust and stable anti-shake effect in complex scenes; in summary, based on the above scheme, confidence fusion compensation of law enforcement recorder images based on visual characteristics and inertial characteristics can be realized, thereby improving the anti-shake robustness in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0018] Figure 1 This is an exemplary flow chart of a method for stabilizing an image of a body camera based on visual-inertial fusion according to some embodiments of the present application; Figure 2 is a schematic diagram of a process for determining constraint confidence according to some embodiments of the present application; Figure 3 is a schematic structural diagram of a jitter repair unit according to some embodiments of the present application; Figure 4It is a structural diagram of a computer device for implementing a law enforcement recorder image stabilization method based on visual-inertial fusion as shown in some embodiments of the present application. DETAILED DESCRIPTION

[0019] In order to better understand the technical solution of the present application, the technical solution of the present application will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0020] refer to Figure 1 This figure is an exemplary flow chart of a method for stabilizing an image of a law enforcement recorder based on visual-inertial fusion according to some embodiments of the present application. The method for stabilizing an image of a law enforcement recorder based on visual-inertial fusion mainly includes the following steps: In step 101, an image sequence of a law enforcement recorder and inertial measurement data in an inertial unit are collected, and multi-scale feature extraction is performed on the image sequence to obtain multiple jitter feature points of the image frame in the law enforcement recorder.

[0021] It should be noted that in this application, the image sequence refers to the multiple frames of video continuously taken by the law enforcement recorder, which can be used to record dynamic scenes; the inertial measurement data refers to the raw data of angular velocity and linear acceleration output by the inertial unit, which can reflect the motion state of the device.

[0022] In the specific implementation, first, the built-in camera in the law enforcement recorder continuously captures the scene at a fixed frame rate (for example, 30 frames per second) to form a temporally coherent image sequence, wherein each frame of the image is automatically exposed and white balanced, and then stored in the device cache with a specific resolution (for example, 1920×1080) and format (for example, H.264 encoding). During the acquisition, it is necessary to ensure the adaptability to the ambient light and the dynamic range to avoid overexposure or underexposure, and at the same time, the hardware anti-shake mechanism is used to preliminarily suppress the influence of mechanical shake on the imaging; then, the law enforcement recorder's habitual The inertial unit samples device motion information at a high frequency (for example, 200 Hz). The inertial unit contains a three-axis gyroscope and a three-axis accelerometer. The three-axis gyroscope measures the angular velocity around the three coordinate axes and is used to calculate the rotation change; the three-axis accelerometer detects linear acceleration and estimates the displacement change after gravity compensation. During acquisition, sensor calibration (to eliminate zero bias and scale errors) and real-time filtering (such as low-pass filtering to suppress high-frequency noise) are required to ensure that the data is strictly synchronized with the image frame timestamp, so that the collection of angular velocity and linear acceleration is used as the inertial measurement data in the inertial unit.

[0023] In some embodiments, performing multi-scale feature extraction on the image sequence to obtain multiple jitter feature points of the image frame in the body camera can be achieved by using the following steps: Performing multi-scale sampling on each image frame in the image sequence to obtain a Gaussian pyramid for each image frame; Determine multiple difference feature points between adjacent image frames in the law enforcement recorder through all Gaussian pyramids, and then determine the jitter degree of each difference feature point; According to all jitter degrees, multiple jitter feature points of the image frame in the law enforcement recorder are screened out from all difference feature points.

[0024] It should be noted that in this application, jitter feature points refer to key points with significant motion characteristics after screening, and jitter feature points can be used for subsequent motion estimation constraints; Gaussian pyramid is a set of images used to extract feature information at different scales; difference feature points refer to pixel points that show significant grayscale changes between the layers of the Gaussian pyramid, and difference feature points can reflect key position information in the image; jitter degree is an indicator that quantifies the degree of change in the feature point position between adjacent frames.

[0025] In the specific implementation, first, for each image frame in the image sequence, the image frame is Gaussian blurred, and a Gaussian kernel with a specific standard deviation is used to perform convolution operation on the image to eliminate high-frequency noise and smooth image details. The blurred image is downsampled by alternate row and column sampling to reduce the image size to a quarter of the previous layer. The above blurring and downsampling process is repeated to generate a 3-5 layer pyramid structure. Each layer of the pyramid image in the pyramid structure represents scene information of different scales. The bottom layer retains complete details, and the high layer highlights the overall structure. The scale difference between the layers of the pyramid is usually set to a multiple of 2 to ensure the comparability of features at different scales. The collection of all pyramid structures is used as the Gaussian pyramid of the image frame. The Gaussian pyramid of each image frame can be obtained by the above method; then, in each layer of the constructed Gaussian pyramid, a grayscale gradient-based method is used. The corner detection method calculates the gradient changes in the horizontal and vertical directions within the pixel neighborhood, and locates the pixel points whose gradient changes exceed the specified change threshold as feature points with obvious corner characteristics as difference feature points, so as to obtain multiple difference feature points between adjacent image frames in the law enforcement recorder; for each feature point, the pyramid images of adjacent image frames are compared layer by layer, and the displacement vector of the feature point in the pyramid image of the adjacent image frame is calculated by the optical flow method, thereby comprehensively considering the modulus and directional consistency of all displacement vectors, and thus calculating the mean of all displacement vectors as the jitter degree of the difference feature point. The jitter degree of each difference feature point can be obtained in the above method; finally, the jitter interval is obtained from the control console of the law enforcement recorder, and the difference feature points whose jitter degrees are within the jitter interval are taken as jitter feature points, so as to obtain multiple jitter feature points of the image frames in the law enforcement recorder.

[0026] In step 102, the motion inertia of each jitter feature point in the image frame is estimated using the inertial measurement data to obtain an inertial matching relationship between adjacent image frames in the law enforcement recorder. The visual jitter characteristics of the image frame in the law enforcement recorder are determined based on the inertial matching relationship and the motion inertia parameters of each jitter feature point.

[0027] In some embodiments, the following steps may be used to estimate the motion inertia of each jitter feature point in the image frame using the inertial measurement data to obtain the inertial matching relationship between adjacent image frames in the law enforcement recorder: Extracting angular velocity features and acceleration features of each jitter feature point in the image frame between adjacent image frames from the inertial measurement data; The inertia of each jitter feature point is estimated through all angular velocity features and acceleration features, and the inertia matching relationship between adjacent image frames in the law enforcement recorder is obtained.

[0028] It should be noted that, in this application, the inertial matching relationship refers to the motion correspondence between feature points between adjacent frames; the angular velocity feature refers to the angular velocity component of the device rotating around the three coordinate axes measured by the gyroscope, and the angular velocity feature can reflect the rotational motion characteristics of the feature point; the acceleration characteristic value refers to the linear acceleration component of the accelerometer in the direction of the three coordinate axes, and the acceleration characteristic value can reflect the translational motion characteristics of the feature point.

[0029] In the specific implementation, first, for each jitter feature point, the inertial data is aligned according to the timestamp of the image frame, and the angular velocity and acceleration that completely match the adjacent frame time windows are obtained from the inertial measurement data, so that the set of all accelerations is used as the angular velocity feature of the jitter feature point between adjacent image frames, and the set of all accelerations is used as the acceleration feature of the jitter feature point between adjacent image frames. The angular velocity feature and acceleration feature of each jitter feature point in the image frame between adjacent image frames can be obtained in the above manner; then, for each jitter feature point, each angular velocity in the angular velocity feature of the jitter feature point is converted into a relative rotation quantity through quaternion integration, and at the same time, the acceleration in the acceleration feature of the jitter feature point is converted into a relative displacement quantity through quadratic integration, and all relative rotation quantities and all relative displacement quantities are used to construct a rotational motion trajectory as the local matching relationship of the jitter feature point. The local matching relationship of each jitter feature point can be obtained in the above manner, so that the set of all local matching relationships can be used as the inertial matching relationship between adjacent image frames in the law enforcement recorder.

[0030] In some embodiments, determining the visual jitter characteristics of the image frame in the law enforcement recorder based on the inertia matching relationship and the motion inertia parameters of each jitter feature point can be achieved by using the following steps: For each jitter feature point, determining the motion inertia parameter of the jitter feature point; Determining the visual jitter value of the jitter feature point by using the motion inertia parameter and the inertia matching relationship, and then obtaining the visual jitter value of each jitter feature point; The visual jitter characteristics of the image frames in the body camera are determined based on all the visual jitter values.

[0031] It should be noted that in this application, the visual jitter feature is a feature used to describe the jitter characteristics of the entire frame image; the motion inertia parameter is a physical quantity parameter used to describe the motion state of the feature point; and the visual jitter value is a numerical indicator that quantifies the degree of unexpected motion of the feature point in the image plane.

[0032] In the specific implementation, first, for each jitter feature point, the standard inertial motion trajectory of the jitter feature point is obtained from the console of the law enforcement recorder as the motion inertia parameter of the jitter feature point; then, the deviation between the inertial motion trajectory in the motion inertia parameter and the rotational motion trajectory of the jitter feature point in the inertial matching relationship is calculated through the visual feature matching algorithm as the visual jitter value of the jitter feature point. The visual jitter value of each jitter feature point can be obtained through the above method; finally, the set of all visual jitter values ​​is used as the visual jitter feature of the image frame in the law enforcement recorder.

[0033] In step 103, the inertia compensation constraints between the visual characteristics and the inertia characteristics of the image frames of the law enforcement recorder during the motion estimation process are determined, and the image frames in the law enforcement recorder are compensated for the inertia constraints using the inertia compensation constraints and the visual jitter characteristics to obtain multiple motion compensation components of the image frames in the law enforcement recorder, and then the constraint confidence of each motion compensation component is determined.

[0034] In some embodiments, determining the inertia compensation constraint between the visual characteristics and the inertial features of the image frame of the body camera during the motion estimation process can be achieved by using the following steps: Initialize a motion estimation model based on the state vector; The displacement observation values ​​of the visual feature points of the law enforcement recorder's image frame during the motion estimation process are used as the visual feature input in the motion estimation model; The inertial measurement data of the image frame of the law enforcement recorder during the motion estimation process is used as the inertial feature input in the motion estimation model; The motion estimation model is used to perform nonlinear constraints on the relationship between the visual characteristics and the inertial characteristics, thereby obtaining inertia compensation constraints between the visual characteristics and the inertial characteristics.

[0035] It should be noted that, in this application, inertia compensation constraint refers to the constraint between visual characteristics and inertial characteristics, and the inertia compensation constraint includes independent motion components of adjacent image frames in each compensation dimension. Independent motion components refer to the basic motion elements after decomposing the complex motion of the law enforcement recorder, including basic motion forms such as translation and rotation. The motion estimation model is a mathematical model that integrates visual and inertial data. The motion estimation model can be used to accurately estimate the motion state of the law enforcement recorder. The motion estimation model is centered on the state vector, which includes the position, posture and speed of the law enforcement recorder. The displacement observation value of the visual feature point provides the motion information of the image plane, and the inertial measurement data provides the physical motion parameters of the device. The motion estimation model describes the motion prediction of the inertial data by establishing a state equation, and through observation The equation expresses the measurement relationship of visual features, and uses a nonlinear optimization method to process the data differences between the two sensors. On the basis of considering their respective error characteristics, the optimal state estimation is found through iterative calculation. The motion estimation model can utilize the long-term stability of visual data and the high-frequency response characteristics of inertial data. Through spatiotemporal alignment and error distribution, a constraint relationship with physical significance is established, and finally an accurate motion trajectory and compensation parameters are output as the inertial compensation constraint between visual characteristics and inertial characteristics. In summary, the inertial compensation constraint is a visual-inertial data fusion condition obtained through optimization calculation. Its essence is to establish a mathematical association rule between the two sensor data and decompose the compound motion of the device into independent components of each dimension, which includes translation and rotation, thereby providing an accurate motion parameter compensation basis for image stabilization.

[0036] In some embodiments, performing inertial constraint compensation on the image frames in the law enforcement recorder using the inertial compensation constraint and the visual jitter feature to obtain multiple motion compensation components of the image frames in the law enforcement recorder can be achieved by the following steps: Extracting independent motion components of adjacent image frames in the body camera in each compensation dimension from the inertia compensation constraint; Determining a consistency score for each independent motion component using the compensation parameter of each independent motion component and the visual jitter feature; Multiple motion compensation components of the image frame in the body camera are screened out from all independent motion components through respective consistency scores.

[0037] In specific implementation, first, a six-degree-of-freedom motion decomposition model is established, and the composite motion parameters obtained from the inertial compensation constraint are decomposed into three translation components (the default is horizontal, vertical, and depth) and three rotation components (the default is pitch, yaw, and roll). This can obtain the independent motion components of adjacent image frames in the law enforcement recorder in each compensation dimension. Then, for each independent motion component, a motion prediction model of the independent motion component is established on the image plane. The motion trajectory in the inertial compensation constraint is compared with the actual visual tracking result. The mean and variance of the position deviation are calculated. A motion smoothness index is introduced to evaluate the continuity performance of the independent motion component in the time series. That is, the weighted sum of the visual jitter feature (the default weight is 0.6) and the compensation parameter of the independent motion component (the default weight is 0.4) is calculated. The result of the weighted sum is normalized to obtain a score value in the range of 0-1 as the consistency score of the independent motion component. In this way, the consistency score of each independent motion component can be obtained. Finally, the independent motion component with a consistency score lower than the preset compensation threshold is used as the motion compensation component, and multiple motion compensation components of the image frame in the law enforcement recorder can be obtained.

[0038] In some embodiments, the constraint confidence of each motion compensation component is determined with reference to Figure 2 As described above, this figure is a schematic diagram of the process of determining the constraint confidence in some embodiments of the present application. In this embodiment, the constraint confidence can be determined by the following steps: In step 1031, for each motion compensation component, a compensation residual of the motion compensation component in each compensation dimension is determined; In step 1032, all compensation residuals are mapped to the constrained confidence of the motion compensation components, thereby obtaining the constrained confidence of each motion compensation component.

[0039] It should be noted that in this application, the constraint confidence is a quantitative indicator used to measure the reliability of the motion compensation component; the compensation residual refers to the difference between the theoretical prediction value and the actual observation value of the motion compensation component, and the compensation residual can reflect the compensation accuracy.

[0040] In the specific implementation, first, for each motion compensation component, a compensation effect evaluation model is established on the image plane, and the position of the feature point after compensation of the component and the corresponding position in the reference frame are used as the comparison object of the compensation effect evaluation model. The compensation effect evaluation model is used to calculate the pixel-level coordinate deviation. For the translation component, the residual algorithm is used to calculate the displacement residuals in the horizontal and vertical directions as the compensation residuals; for the rotation component, the residual algorithm is used to calculate the angle deviation residuals as the compensation residuals; for the scaling component, the residual algorithm is used to calculate the scale difference residuals of the scaling component as the compensation residuals. The compensation residuals of the motion compensation component in each compensation dimension can be obtained in the above manner; then, a nonlinear mapping function is used to convert the compensation residuals into confidence values ​​in the range of 0-1 as the constraint confidence of the motion compensation component. The constraint confidence of each motion compensation component can be obtained in the above manner.

[0041] It should be noted that for the translation and rotation components, an exponential decay function is used to process the residual mean and variance; for the scaling component, a logarithmic function is used for normalization. In actual use, a time sliding window mechanism can be introduced to comprehensively consider the residual performance of the current frame and the historical frame to enhance the temporal stability of the confidence assessment. At the same time, dynamic adjustment parameters are set. When the ambient lighting conditions change or the motion mode changes, the steepness of the mapping curve is automatically adjusted. The final output confidence value will reflect the immediate reliability and long-term stability of each component, providing a basis for subsequent weighted fusion. Components with a confidence level lower than 0.3 will be marked as low reliability, triggering a special exception handling mechanism.

[0042] In step 104, the next frame image in the law enforcement recorder is jitter-repaired based on the constraint confidence of each motion compensation component to obtain a repaired image of the next frame image in the law enforcement recorder.

[0043] In some embodiments, the next frame image in the law enforcement recorder is repaired for jitter based on the constrained confidence of each motion compensation component, and the repaired image of the next frame image in the law enforcement recorder can be achieved in the following manner, namely: first, all jitter feature points are weightedly fused using the constrained confidence of each motion compensation component, and the components with constrained confidence higher than 0.7 are used as high-confidence components, and the components with constrained confidence lower than 0.3 are used as low-confidence components. The high-confidence components are given larger weights. For the translation component, the confidence-weighted displacement vector is used to calculate the final compensation position; for the rotation component, quaternion spherical linear interpolation is used to achieve a smooth transition; the scaling component is fused using geometric mean, and when establishing the image transformation matrix, components with constrained confidence higher than 0.7 are applied first, and the low-confidence components are amplitude-limited. An adaptive filling strategy is used in the image boundary area, and missing pixels are intelligently generated according to the content of adjacent frames and motion trends to avoid the black edge phenomenon. The final repaired image can not only maintain the integrity of the scene, but also effectively suppress the unexpected jitter of the law enforcement recorder.

[0044] In addition, in another aspect of the present application, in some embodiments, the present application provides a law enforcement recorder image stabilization system based on visual inertial fusion, the law enforcement recorder image stabilization system based on visual inertial fusion includes a shake repair unit, reference Figure 3 , which is a schematic diagram of the structure of a jitter repair unit according to some embodiments of the present application. The jitter repair unit includes: an acquisition module 201, a processing module 202, and an execution module 203, which are described as follows: Acquisition module 201, in this application, acquisition module 201 is mainly used to collect image sequences of law enforcement recorders and inertial measurement data in the inertial unit, perform multi-scale feature extraction on the image sequences, and obtain multiple jitter feature points of the image frames in the law enforcement recorder; Processing module 202, in the present application, is used to estimate the motion inertia of each jitter feature point in the image frame using the inertial measurement data, obtain an inertia matching relationship between adjacent image frames in the law enforcement recorder, and determine the visual jitter characteristics of the image frame in the law enforcement recorder based on the inertia matching relationship and the motion inertia parameters of each jitter feature point; It should be noted that the processing module 202 is further configured to determine an inertia compensation constraint between visual characteristics and inertia features of the image frames of the law enforcement recorder during motion estimation, perform inertia constraint compensation on the image frames in the law enforcement recorder using the inertia compensation constraint and the visual jitter features, obtain multiple motion compensation components of the image frames in the law enforcement recorder, and further determine the constraint confidence of each motion compensation component; Execution module 203. In this application, execution module 203 is mainly used to perform jitter repair on the next frame image in the law enforcement recorder based on the constraint confidence of each motion compensation component to obtain a repaired image of the next frame image in the law enforcement recorder.

[0045] The above describes in detail the examples of the image stabilization method and system for law enforcement recorders based on visual-inertial fusion provided in the embodiments of the present application. It can be understood that in order to realize the above functions, the corresponding device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0046] In some embodiments, the present application also provides a computer device, which includes a memory and a processor, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the above-mentioned law enforcement recorder image stabilization method based on visual inertial fusion.

[0047] In some embodiments, reference Figure 4 , the dotted line in the figure indicates that the unit or module is optional. The figure is a structural diagram of a computer device for implementing a law enforcement recorder image stabilization method based on visual inertial fusion according to an embodiment of the present application. The law enforcement recorder image stabilization method based on visual inertial fusion described in the above embodiment can be Figure 4 The computer device shown in the figure is implemented, and the computer device includes at least one processor 301, a memory 302 and at least one communication unit 305. The computer device can be a terminal device, a server or a chip.

[0048] The processor 301 may be a general-purpose processor or a dedicated processor. For example, the processor 301 may be a central processing unit (CPU), which may be used to control the computer device, execute software programs, and process data from the software programs. The computer device may also include a communication unit 305 for inputting (receiving) and outputting (transmitting) signals.

[0049] For example, the computer device may be a chip, the communication unit 305 may be an input and / or output circuit of the chip, or the communication unit 305 may be a communication interface of the chip, and the chip may be a component of a terminal device, a network device, or other device.

[0050] For another example, the computer device may be a terminal device or a server, and the communication unit 305 may be a transceiver of the terminal device or the server, or the communication unit 305 may be a transceiver circuit of the terminal device or the server.

[0051] The computer device may include one or more memories 302, on which a program 304 is stored. The program 304 can be executed by the processor 301 to generate instructions 303, so that the processor 301 executes the method described in the above method embodiment according to the instructions 303. Optionally, data (such as a target audit model) can also be stored in the memory 302. Optionally, the processor 301 can also read data stored in the memory 302. The data can be stored at the same storage address as the program 304, or at a different storage address from the program 304.

[0052] The processor 301 and the memory 302 may be provided separately or integrated together, for example, integrated on a system on chip (SOC) of a terminal device.

[0053] It should be understood that each step of the above method embodiment can be completed by a hardware-based logic circuit or software-based instructions in the processor 301. The processor 301 can be a CPU, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, such as discrete gates, transistor logic devices, or discrete hardware components.

[0054] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0055] For example, in some embodiments, the present application also provides a computer-readable storage medium, which stores instructions or codes. When the instructions or codes are run on a computer, the computer implements the above-mentioned law enforcement recorder image stabilization method based on visual-inertial fusion when executing.

[0056] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0057] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A method for image stabilization of law enforcement recorders based on visual-inertial fusion, characterized in that: The steps include: Collecting an image sequence of a law enforcement recorder and inertial measurement data in an inertial unit, performing multi-scale feature extraction on the image sequence, and obtaining a plurality of jitter feature points of an image frame in the law enforcement recorder; Estimating the motion inertia of each jitter feature point in the image frame using the inertial measurement data to obtain an inertia matching relationship between adjacent image frames in the law enforcement recorder, and determining the visual jitter characteristics of the image frame in the law enforcement recorder based on the inertia matching relationship and the motion inertia parameters of each jitter feature point; Determining an inertia compensation constraint between visual characteristics and inertia features of an image frame from a law enforcement recorder during motion estimation, performing inertia constraint compensation on the image frame from the law enforcement recorder using the inertia compensation constraint and the visual jitter feature, obtaining multiple motion compensation components of the image frame from the law enforcement recorder, and then determining a constraint confidence level for each motion compensation component; Based on the constrained confidence of each motion compensation component, the jitter of the next frame image in the law enforcement recorder is repaired to obtain a repaired image of the next frame image in the law enforcement recorder.

2. The method according to claim 1, wherein The multi-scale feature extraction is performed on the image sequence to obtain multiple jitter feature points of the image frame in the law enforcement recorder, specifically including: Performing multi-scale sampling on each image frame in the image sequence to obtain a Gaussian pyramid for each image frame; Determine multiple difference feature points between adjacent image frames in the law enforcement recorder through all Gaussian pyramids, and then determine the jitter degree of each difference feature point; According to all jitter degrees, multiple jitter feature points of the image frame in the law enforcement recorder are screened out from all difference feature points.

3. The method according to claim 1, wherein The inertial measurement data is used to estimate the motion inertia of each jitter feature point in the image frame to obtain the inertial matching relationship between adjacent image frames in the law enforcement recorder, which specifically includes: Extracting angular velocity features and acceleration features of each jitter feature point in the image frame between adjacent image frames from the inertial measurement data; The inertia of each jitter feature point is estimated through all angular velocity features and acceleration features, and the inertia matching relationship between adjacent image frames in the law enforcement recorder is obtained.

4. The method according to claim 1, wherein Determining the visual jitter characteristics of the image frame in the law enforcement recorder according to the inertia matching relationship and the motion inertia parameters of each jitter feature point specifically includes: For each jitter feature point, determining the motion inertia parameter of the jitter feature point; Determining the visual jitter value of the jitter feature point by using the motion inertia parameter and the inertia matching relationship, and then obtaining the visual jitter value of each jitter feature point; The visual jitter characteristics of the image frames in the body camera are determined based on all the visual jitter values.

5. The method according to claim 1, wherein The inertia compensation constraints between the visual characteristics and inertial features of the image frames of the body camera during motion estimation are specifically determined as follows: Initialize a motion estimation model based on the state vector; The displacement observation values ​​of the visual feature points of the law enforcement recorder's image frame during the motion estimation process are used as the visual feature input in the motion estimation model; The inertial measurement data of the image frame of the law enforcement recorder during the motion estimation process is used as the inertial feature input in the motion estimation model; The motion estimation model is used to perform nonlinear constraints on the relationship between the visual characteristics and the inertial characteristics, thereby obtaining inertia compensation constraints between the visual characteristics and the inertial characteristics.

6. The method according to claim 1, wherein The image frames in the law enforcement recorder are subjected to inertia constraint compensation using the inertia compensation constraint and the visual jitter feature, and the multiple motion compensation components of the image frames in the law enforcement recorder are obtained, specifically including: Extracting independent motion components of adjacent image frames in the body camera in each compensation dimension from the inertial compensation constraint; Determining a consistency score for each independent motion component based on the compensation parameter of each independent motion component and the visual jitter feature; Multiple motion compensation components of the image frame in the body camera are screened out from all independent motion components through respective consistency scores.

7. The method according to claim 1, wherein Determining the constraint confidence of each motion compensation component specifically includes: For each motion compensation component, determining a compensation residual of the motion compensation component in each compensation dimension; All compensation residuals are mapped to the constrained confidence of the motion compensation components, thereby obtaining the constrained confidence of each motion compensation component.

8. A law enforcement recorder image stabilization system based on visual inertial fusion, the law enforcement recorder image stabilization system based on visual inertial fusion includes a shake repair unit, characterized in that: The jitter repair unit includes: An acquisition module is used to acquire an image sequence of the law enforcement recorder and inertial measurement data in the inertial unit, perform multi-scale feature extraction on the image sequence, and obtain multiple jitter feature points of the image frame in the law enforcement recorder; a processing module configured to estimate the motion inertia of each jitter feature point in the image frame using the inertial measurement data, obtain an inertia matching relationship between adjacent image frames in the law enforcement recorder, and determine the visual jitter characteristics of the image frames in the law enforcement recorder based on the inertia matching relationship and the motion inertia parameters of each jitter feature point; The processing module is further configured to determine an inertia compensation constraint between visual characteristics and inertia features of an image frame of the law enforcement recorder during a motion estimation process, perform inertia constraint compensation on the image frame in the law enforcement recorder using the inertia compensation constraint and the visual jitter feature, obtain multiple motion compensation components of the image frame in the law enforcement recorder, and further determine a constraint confidence level of each motion compensation component; The execution module is used to perform jitter repair on the next frame image in the law enforcement recorder based on the constraint confidence of each motion compensation component to obtain a repaired image of the next frame image in the law enforcement recorder.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the law enforcement recorder image stabilization method based on visual inertial fusion as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions or codes, and when the instructions or codes are executed on a computer, the computer implements the image stabilization method for law enforcement recorders based on visual-inertial fusion as described in any one of claims 1 to 7.