Image dithering removing method and device, terminal equipment and storage medium
By using the bounding rectangular frame of the reference image in the roadside camera to calculate and correct the position deviation caused by image jitter, the image jitter problem caused by camera vibration is solved, and the effect of object detection and tracking is improved.
Patent Information
- Application Number
- CN202311631280.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-06
AI Technical Summary
The vibration caused by roadside cameras due to strong winds and bridge vibrations, resulting in image shaking, affecting the effect of target detection and tracking.
By taking a reference image in a stationary state in advance, the boundary rectangular frame of the reference stationary target is stored as the reference boundary rectangular frame. Then, the image to be processed is detected, and the coordinate error between each initial boundary rectangle box and the reference boundary rectangle box is calculated, the minimum error is found, and the error is used to perform pixel translation processing on the image to correct the image position deviation caused by jitter.
It effectively corrects the image jitter caused by camera vibration, improves the effect of object detection and tracking, and ensures the clarity and stability of the image.
Smart Images

Figure CN120111362A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, apparatus, terminal device and storage medium for image de-jittering. Background Art
[0002] The cameras of the roadside system are generally installed on road facilities such as gantries, viaducts, roadside poles or roadside crossbars, and are mainly used to capture images of road targets such as vehicles and pedestrians in order to detect and track these road targets. However, due to reasons such as strong winds and bridge vibrations, it is easy to cause camera vibration, which will cause the continuous frame video imaging to jitter, resulting in a large deviation between the previous and next frame images, seriously affecting the effect of target detection and tracking. Summary of the invention
[0003] In view of this, the embodiments of the present application provide a method, apparatus, terminal device and storage medium for image de-shaking, which can correct image information after camera shake to achieve the effect of image de-shaking.
[0004] A first aspect of an embodiment of the present application provides a method for image de-shaking, comprising:
[0005] Acquire the image to be processed taken by the target camera;
[0006] Performing target detection processing on the image to be processed to obtain an initial bounding rectangular frame of each stationary target in the image to be processed;
[0007] Calculating the first coordinate error between each initial bounding rectangular frame and a pre-stored reference bounding rectangular frame respectively; wherein the reference bounding rectangular frame is a bounding rectangular frame of a reference stationary target in a reference image pre-photographed by the target camera in a stationary state;
[0008] The image to be processed is de-jittered according to the minimum error among the first coordinate errors.
[0009] In an embodiment of the present application, the target camera is preliminarily ordered to take a reference image in a stationary state, a reference stationary target (such as a lane line, a manhole cover or a roadblock, etc.) is selected from the reference image, and the bounding rectangle of the reference stationary target is stored as the reference bounding rectangle. Afterwards, when the target camera is used for image acquisition (at this time, the target camera may vibrate), the image to be processed acquired by the target camera is first subjected to target detection processing to obtain the initial bounding rectangle of each stationary target therein, and then the coordinate error between each initial bounding rectangle and the reference bounding rectangle is calculated respectively, the minimum error among these coordinate errors is found, and finally the image to be processed is de-jittered using the minimum error. In the above process, since the initial bounding rectangle corresponding to the minimum error can be regarded as the bounding rectangle corresponding to the reference stationary target obtained after target detection of the image to be processed, the minimum error can be regarded as the image position deviation caused by image jitter, and finally the image position deviation can be corrected by using the minimum error to perform translation and other processing on the pixels of the image to be processed, thereby correcting the image information after the camera shakes, and achieving the effect of image de-jittering.
[0010] In an implementation of the embodiment of the present application, de-jittering is performed on the image to be processed according to the minimum error among the first coordinate errors, including:
[0011] Determine the initial bounding rectangular frame whose first coordinate error with the reference bounding rectangular frame is the minimum error as the target bounding rectangular frame corresponding to the reference stationary object;
[0012] Determine the pixel translation amount according to a first coordinate error between the target bounding rectangular box and the reference bounding rectangular box;
[0013] According to the pixel translation amount, each pixel of the image to be processed is translated to obtain the image to be processed after de-jittering.
[0014] In an implementation of the embodiment of the present application, the image to be processed includes multiple frames of continuous images; determining the pixel translation amount according to the first coordinate error between the target bounding rectangular box and the reference bounding rectangular box includes:
[0015] Obtain the horizontal coordinate error and the vertical coordinate error between the target bounding rectangle and the reference bounding rectangle;
[0016] The horizontal coordinate error and the vertical coordinate error are used as the state quantity of the Kalman filter, and the state quantity is optimally estimated using the Kalman filter algorithm to obtain the optimal estimated value of the horizontal coordinate error and the optimal estimated value of the vertical coordinate error of each frame of multiple continuous images;
[0017] For each frame of the multiple continuous images, determine the pixel translation amount of the frame according to the optimal estimated value of the horizontal coordinate error and the optimal estimated value of the vertical coordinate error of the frame;
[0018] Accordingly, according to the pixel translation amount, each pixel of the image to be processed is translated, including:
[0019] For each frame of the multiple continuous image frames, each pixel of the frame is subjected to translation processing according to the pixel translation amount of the frame.
[0020] In an implementation of the embodiment of the present application, before respectively calculating the first coordinate error between each initial bounding rectangular box and the pre-stored reference bounding rectangular box, the method further includes:
[0021] In each initial bounding rectangular frame, the initial bounding rectangular frame whose corresponding target category is different from the target category corresponding to the reference bounding rectangular frame is deleted.
[0022] In an implementation of the embodiment of the present application, de-jittering is performed on the image to be processed according to the minimum error among the first coordinate errors, including:
[0023] Determine whether the minimum error among the first coordinate errors is less than a preset error threshold;
[0024] If the minimum error among the first coordinate errors is smaller than the error threshold, de-jitter processing is performed on the image to be processed according to the minimum error among the first coordinate errors.
[0025] In an implementation of the embodiment of the present application, after determining whether the minimum error among the first coordinate errors is less than a preset error threshold, the method further includes:
[0026] If the minimum error among the first coordinate errors is greater than or equal to the error threshold, a pre-stored candidate bounding rectangular box is obtained; wherein the candidate bounding rectangular box is a bounding rectangular box of a candidate stationary target different from the reference stationary target in the reference image;
[0027] Calculate the second coordinate error between each initial bounding rectangular box and the candidate bounding rectangular box respectively;
[0028] The image to be processed is de-jittered according to the minimum error among the second coordinate errors.
[0029] In one implementation of the embodiment of the present application, the error threshold is determined in the following manner:
[0030] Detecting the vibration amplitude of the target camera within a preset time period;
[0031] An error threshold is determined according to the vibration amplitude; wherein the error threshold is proportional to the vibration amplitude.
[0032] In one implementation of the embodiment of the present application, the reference stationary target is any lane line in the reference image; target detection processing is performed on the image to be processed to obtain an initial bounding rectangular frame of each stationary target in the image to be processed, including:
[0033] The lane line detection process is performed on the image to be processed to obtain the initial boundary rectangular frame of each lane line in the image to be processed.
[0034] A second aspect of an embodiment of the present application provides an image de-shaking device, including:
[0035] An image acquisition module is used to acquire the image to be processed taken by the target camera;
[0036] The target detection module is used to perform target detection processing on the image to be processed to obtain the initial bounding rectangle frame of each stationary target in the image to be processed;
[0037] A coordinate error calculation module, used to respectively calculate a first coordinate error between each initial bounding rectangular frame and a pre-stored reference bounding rectangular frame; wherein the reference bounding rectangular frame is a bounding rectangular frame of a reference stationary target in a reference image pre-photographed by a target camera in a stationary state;
[0038] The image de-jitter module is used to perform de-jitter processing on the image to be processed according to the minimum error among the first coordinate errors.
[0039] A third aspect of an embodiment of the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the image de-shaking method provided in the first aspect of the embodiment of the present application is implemented.
[0040] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the image de-shaking method provided in the first aspect of the embodiments of the present application is implemented.
[0041] A fifth aspect of the embodiments of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device executes the image de-shaking method provided in the first aspect of the embodiments of the present application.
[0042] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a flow chart of a method for image de-shaking provided by an embodiment of the present application;
[0044] Figure 2 is a structural schematic diagram of an image de-jittering device provided in an embodiment of the present application;
[0045] Figure 3 It is a schematic diagram of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures, technologies, etc. are proposed, so as to thoroughly understand the embodiments of the present application. However, it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details. In other cases, the detailed description of well-known systems, devices, circuits and methods is omitted to prevent unnecessary details from hindering the description of the present application. In addition, in the description of the present application specification and the attached claims, the terms "first", "second", "third" etc. are only used to distinguish the description, and cannot be interpreted as indicating or suggesting relative importance.
[0047] At present, roadside cameras installed on road facilities such as gantries, viaducts, roadside poles or roadside crossbars often vibrate due to strong winds and bridge vibrations, which can cause large deviations between the previous and next frame images, seriously affecting the effect of target detection and tracking. In view of this, the embodiments of the present application provide a method, device, terminal device and storage medium for image de-shaking, which can correct the image information after the camera shakes to achieve the effect of image de-shaking. For more specific technical implementation details of the embodiments of the present application, please refer to the method embodiments described below.
[0048] It should be understood that the execution subjects of the various method embodiments of the present application are various types of terminal devices or servers, such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA), large-screen TVs, etc. The embodiments of the present application do not impose any restrictions on the specific types of the terminal devices and servers.
[0049] See also Figure 1 , shows a method for image de-shaking provided by an embodiment of the present application, comprising:
[0050] 101. Acquire an image to be processed captured by a target camera;
[0051] First, the image to be processed taken by the target camera is obtained. The target camera can be any type of camera installed at any position, for example, it can be a roadside camera installed on road facilities such as a gantry, an elevated bridge, a roadside pole or a roadside crossbar that is prone to vibration. Due to the vibration of the target camera and other reasons, the video imaging of the target camera will be shaken, so it is necessary to de-shake the image taken by the target camera. The image to be processed is one or more frames of images taken by the target camera that need to be de-shaked.
[0052] 102. Performing target detection processing on the image to be processed to obtain an initial bounding rectangle frame of each stationary target in the image to be processed;
[0053] After acquiring the image to be processed, the image to be processed is subjected to target detection processing for stationary targets, thereby obtaining the initial bounding rectangle of each stationary target in the image to be processed. The initial bounding rectangle is a rectangle used to frame the target position. Stationary targets refer to various types of targets in the image to be processed that are in a stationary state. For example, for images collected by roadside cameras, stationary targets can be targets such as lane lines, roadside lines, sewer covers, zebra crossings, and roadblocks on the road, while moving targets can be targets such as vehicles and pedestrians on the road.
[0054] On the other hand, in the embodiment of the present application, the target camera is allowed to shoot in a stationary state to obtain one or more reference images in advance, and the stationary target is detected in the reference image, so as to obtain each stationary target in the reference image. Afterwards, one of the stationary targets is selected as the reference stationary target, and the bounding rectangle of the reference stationary target is stored as the reference bounding rectangle.
[0055] In one implementation of the embodiment of the present application, the reference stationary target is any lane line in the reference image; target detection processing is performed on the image to be processed to obtain an initial bounding rectangular frame of each stationary target in the image to be processed, including:
[0056] The lane line detection process is performed on the image to be processed to obtain the initial boundary rectangular frame of each lane line in the image to be processed.
[0057] In roadside scenes, lane lines are the most common static targets, and lane line detection technology is also very mature. Therefore, the embodiment of the present application can select any lane line from the reference image as the reference static target, and the bounding rectangle of the selected lane line is used as the reference bounding rectangle. Correspondingly, lane line detection processing is performed on the image to be processed, thereby obtaining the initial bounding rectangle of each lane line in the image to be processed.
[0058] 103. Calculate first coordinate errors between each initial bounding rectangular box and a pre-stored reference bounding rectangular box respectively;
[0059] After obtaining the initial bounding rectangular frames of each stationary target in the image to be processed, the first coordinate error between each initial bounding rectangular frame and the pre-stored reference bounding rectangular frame is calculated. Specifically, the distance between the center points of two bounding rectangular frames can be detected as the first coordinate error. For example, assuming that N initial bounding rectangular frames are detected, the coordinates (x 1 ,y 1 )、(x 2 ,y 2 ),…(x N ,y N ), if the center point coordinates of the reference bounding rectangle are (x g ,y g ), then the distance between the center point coordinates of each initial bounding rectangle and the center point coordinates of the reference bounding rectangle is calculated respectively, that is, (x 1 -x g ) 2 +(y 1 -y g ) 2 、(x 2 -x g ) 2 +(y 2 -y g ) 2 、…(x N -x g ) 2 +(y N -y g ) 2 , these distances can be respectively used as the first coordinate errors between each initial bounding rectangle and the reference bounding rectangle.
[0060] In an implementation of the embodiment of the present application, before respectively calculating the first coordinate error between each initial bounding rectangular box and the pre-stored reference bounding rectangular box, the method further includes:
[0061] In each initial bounding rectangular frame, the initial bounding rectangular frame whose corresponding target category is different from the target category corresponding to the reference bounding rectangular frame is deleted.
[0062] The purpose of calculating the first coordinate error corresponding to each initial bounding rectangular frame is to find the minimum error in the first coordinate error, and the initial bounding rectangular frame corresponding to the minimum error is usually a bounding rectangular frame surrounding the reference stationary target, so the target category corresponding to the initial bounding rectangular frame must be the same as the target category corresponding to the reference bounding rectangular frame. According to this feature, before calculating the first coordinate error between each initial bounding rectangular frame and the pre-stored reference bounding rectangular frame, the initial bounding rectangular frames whose corresponding target categories are different from the target categories corresponding to the reference bounding rectangular frame can be deleted in each initial bounding rectangular frame. That is, only the initial bounding rectangular frame with the same category as the reference stationary target is calculated to exclude interference from other stationary targets of different categories in the image. For example, if the reference stationary target is a lane line, that is, the target category corresponding to the reference bounding rectangular frame is a lane line, then the initial bounding rectangular frames of other stationary targets other than the lane line in the image can be deleted, which can reduce the amount of calculation to a certain extent and improve the algorithm running speed.
[0063] 104. De-jitter the image to be processed according to the minimum error among the first coordinate errors.
[0064] After calculating the first coordinate error corresponding to each initial bounding rectangle, find the minimum error of each first coordinate error, and regard the initial bounding rectangle corresponding to the minimum error as the bounding rectangle surrounding the reference stationary target. In this way, the coordinate error between the initial bounding rectangle corresponding to the minimum error and the reference bounding rectangle can be regarded as the deviation caused by image jitter, so the image to be processed can be de-jittered by using the deviation.
[0065] In an implementation of the embodiment of the present application, de-jittering is performed on the image to be processed according to the minimum error among the first coordinate errors, including:
[0066] (1) determining an initial bounding rectangular frame having a minimum first coordinate error with the reference bounding rectangular frame as a target bounding rectangular frame corresponding to the reference stationary object;
[0067] (2) determining a pixel translation amount according to a first coordinate error between the target bounding rectangular box and the reference bounding rectangular box;
[0068] (3) According to the pixel translation amount, each pixel of the image to be processed is translated to obtain the image to be processed after de-jittering.
[0069] First, find the initial bounding rectangular box whose first coordinate error with the reference bounding rectangular box is the minimum error, and determine it as the target bounding rectangular box corresponding to the reference stationary target; then, determine the corresponding pixel translation amount according to the first coordinate error between the target bounding rectangular box and the reference bounding rectangular box; finally, translate each pixel of the image to be processed according to the pixel translation amount, so as to obtain the de-jittered image to be processed.
[0070] In an implementation of the embodiment of the present application, the image to be processed includes multiple frames of continuous images; determining the pixel translation amount according to the first coordinate error between the target bounding rectangular box and the reference bounding rectangular box includes:
[0071] (1) Obtaining the horizontal coordinate error and the vertical coordinate error between the target bounding rectangular box and the reference bounding rectangular box;
[0072] (2) taking the horizontal coordinate error and the vertical coordinate error as the state quantity of the Kalman filter, and using the Kalman filter algorithm to optimally estimate the state quantity, and obtaining the optimal estimated value of the horizontal coordinate error and the optimal estimated value of the vertical coordinate error of each frame of the multiple continuous images;
[0073] (3) For each frame of the multiple continuous images, the pixel translation amount of the frame is determined according to the optimal estimated value of the horizontal coordinate error and the optimal estimated value of the vertical coordinate error of the frame.
[0074] Assuming that the image to be processed includes multiple continuous images, when determining the pixel translation amount according to the first coordinate error between the target boundary rectangular frame and the reference boundary rectangular frame, the horizontal coordinate error and the vertical coordinate error between the target boundary rectangular frame and the reference boundary rectangular frame are first obtained, and then the horizontal coordinate error and the vertical coordinate error are used as the state quantity of the Kalman filter, and the state quantity is optimally estimated using the Kalman filter algorithm, so as to obtain the optimal estimation value of the horizontal coordinate error and the optimal estimation value of the vertical coordinate error of each frame image in the multiple continuous images. Among them, the principle of using the Kalman filter algorithm to optimally estimate the state quantity can refer to the prior art and will not be repeated here. Afterwards, for each frame image in the multiple continuous images, the pixel translation amount of the frame image is determined according to the optimal estimation value of the horizontal coordinate error and the optimal estimation value of the vertical coordinate error of the frame image. Since there is a certain error in the detection results of each initial boundary rectangular frame, the Kalman filter is used to optimally estimate the deviation value of the detected target position (i.e., the target boundary rectangular frame) and the priori position (i.e., the reference boundary rectangular frame), so as to obtain a more stable deviation value, thereby improving the image de-shaking effect.
[0075] For example, suppose the coordinates of the center point of the reference bounding rectangle are (x g ,y g ), the coordinates of the center point of the target bounding rectangle are (xm ,y m ), then the horizontal coordinate error Δx=x m -x g , the vertical coordinate error Δy = y m -y g , taking Δx and Δy as the state variables of the Kalman filter, constructing the Kalman filter algorithm to estimate the error value of each frame of the image, the optimal estimated value d of the horizontal coordinate error of each frame of the multiple continuous images can be obtained. x and the optimal estimate of the ordinate error d y Obviously, each frame of the image can obtain its own optimal estimate of the horizontal coordinate error and the optimal estimate of the vertical coordinate error, and then calculate its own pixel translation amount. As an example, the optimal estimate of the horizontal coordinate error can be used as the pixel translation amount in the horizontal direction, and the optimal estimate of the vertical coordinate error can be used as the pixel translation amount in the vertical direction.
[0076] Accordingly, according to the pixel translation amount, each pixel of the image to be processed is translated, including:
[0077] For each frame of the multiple continuous image frames, each pixel of the frame is subjected to translation processing according to the pixel translation amount of the frame.
[0078] For example, the optimal estimated value d of the horizontal coordinate error of a frame image is calculated by the Kalman filter algorithm. x and the optimal estimate of the ordinate error d y Afterwards, d x As the pixel translation amount in the horizontal axis direction, d y As the pixel translation amount in the ordinate direction, each pixel of the frame image is subjected to the following translation processing: pixel[x, y]=pixel[x+dx, y+dy], thereby completing the de-shaking processing of the frame image. In addition, assuming that the width of the frame image is W and the height is H, if the offset pixel exceeds the image range, that is, when x+dx<0, x+dx>W, y+dy<0 or y+dy>H occurs, it can be filled with a specified pixel value, for example, it can be filled with zero: pixel[x, y]=0.
[0079] In an implementation of the embodiment of the present application, de-jittering is performed on the image to be processed according to the minimum error among the first coordinate errors, including:
[0080] (1) determining whether the minimum error among the first coordinate errors is less than a preset error threshold;
[0081] (2) If the minimum error among the first coordinate errors is less than the error threshold, de-jitter processing is performed on the image to be processed according to the minimum error among the first coordinate errors.
[0082] In some cases, the target bounding rectangle corresponding to the reference stationary target in the image to be processed may be blocked. For example, if the reference stationary target is a lane line, it may be blocked by a passing vehicle. Considering these situations, it is possible to first determine whether the minimum error among the first coordinate errors is less than a preset error threshold. If the minimum error is less than the error threshold, it means that the target bounding rectangle corresponding to the reference stationary target in the image to be processed is not blocked. At this time, the minimum error can be used as a pixel translation amount to de-jitter the image to be processed according to the method described above.
[0083] In an implementation of the embodiment of the present application, after determining whether the minimum error among the first coordinate errors is less than a preset error threshold, the method further includes:
[0084] (1) if the minimum error among the first coordinate errors is greater than or equal to the error threshold, obtaining a pre-stored candidate bounding rectangular box; wherein the candidate bounding rectangular box is a bounding rectangular box of a candidate stationary target different from the reference stationary target in the reference image;
[0085] (2) respectively calculating the second coordinate error between each initial bounding rectangular box and the candidate bounding rectangular box;
[0086] (3) De-jittering the image to be processed according to the minimum error among the second coordinate errors.
[0087] If the minimum error among the first coordinate errors is greater than or equal to the error threshold, it means that the target bounding rectangle box corresponding to the reference stationary target in the image to be processed is blocked, and at this time, the minimum error among the first coordinate errors cannot be used as the pixel translation amount to de-jitter the image to be processed. In view of this situation, the bounding rectangle box of the candidate stationary target different from the reference stationary target in the reference image can be obtained as the alternative bounding rectangle box, and the reference bounding rectangle box is replaced with the alternative bounding rectangle box, and the de-jitter processing of the image to be processed is completed in the same way. That is, the second coordinate error between each initial bounding rectangle box and the alternative bounding rectangle box is calculated respectively, and then the image to be processed is de-jittered according to the minimum error among the second coordinate errors. Obviously, multiple different alternative stationary targets can be set, and when the current alternative stationary target is blocked, the next alternative stationary target is continuously obtained for judgment until an alternative stationary target that is not blocked is obtained.
[0088] In one implementation of the embodiment of the present application, the error threshold is determined in the following manner:
[0089] (1) Detecting the vibration amplitude of the target camera within a preset time period;
[0090] (2) Determine an error threshold based on the vibration amplitude; wherein the error threshold is proportional to the vibration amplitude.
[0091] When setting the above error threshold, the vibration amplitude of the target camera within a preset time period can be detected. If the vibration amplitude is higher, the deviation between the position of the detected bounding rectangle frame of the reference stationary target and the actual position of the bounding rectangle frame of the reference stationary target will also be greater, so the error threshold is set larger, that is, the error threshold is proportional to the vibration amplitude. By processing in this way, the above error threshold can be set more reasonably in combination with the actual situation of camera vibration.
[0092] In an embodiment of the present application, the target camera is preliminarily ordered to take a reference image in a stationary state, a reference stationary target (such as a lane line, a manhole cover or a roadblock, etc.) is selected from the reference image, and the bounding rectangle of the reference stationary target is stored as the reference bounding rectangle. Afterwards, when the target camera is used for image acquisition (at this time, the target camera may vibrate), the image to be processed acquired by the target camera is first subjected to target detection processing to obtain the initial bounding rectangle of each stationary target therein, and then the coordinate error between each initial bounding rectangle and the reference bounding rectangle is calculated respectively, the minimum error among these coordinate errors is found, and finally the image to be processed is de-jittered using the minimum error. In the above process, since the initial bounding rectangle corresponding to the minimum error can be regarded as the bounding rectangle corresponding to the reference stationary target obtained after target detection of the image to be processed, the minimum error can be regarded as the image position deviation caused by image jitter, and finally the image position deviation can be corrected by using the minimum error to perform translation and other processing on the pixels of the image to be processed, thereby correcting the image information after the camera shakes, and achieving the effect of image de-jittering.
[0093] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0094] A method for image de-shaking is mainly described above, and a device for image de-shaking will be described below.
[0095] See also Figure 2 In one embodiment of the present application, an image de-shaking device includes:
[0096] An image acquisition module 201 is used to acquire an image to be processed taken by a target camera;
[0097] The target detection module 202 is used to perform target detection processing on the image to be processed to obtain the initial bounding rectangle of each stationary target in the image to be processed;
[0098] The coordinate error calculation module 203 is used to respectively calculate the first coordinate error between each initial bounding rectangular frame and a pre-stored reference bounding rectangular frame; wherein the reference bounding rectangular frame is a bounding rectangular frame of a reference stationary target in a reference image pre-photographed by the target camera in a stationary state;
[0099] The image de-jitter module 204 is used to perform de-jitter processing on the image to be processed according to the minimum error among the first coordinate errors.
[0100] In one implementation of the embodiment of the present application, the image de-shaking module includes:
[0101] A target bounding rectangular frame determining unit, configured to determine an initial bounding rectangular frame having a minimum first coordinate error with a reference bounding rectangular frame as a target bounding rectangular frame corresponding to the reference stationary target;
[0102] A pixel translation amount determination unit, used to determine the pixel translation amount according to a first coordinate error between the target bounding rectangular frame and the reference bounding rectangular frame;
[0103] The pixel translation unit is used to perform translation processing on each pixel of the image to be processed according to the pixel translation amount to obtain the image to be processed after de-jittering.
[0104] In an implementation of the embodiment of the present application, the image to be processed includes multiple frames of continuous images; the pixel translation amount determination unit includes:
[0105] The error acquisition subunit is used to acquire the horizontal coordinate error and the vertical coordinate error between the target bounding rectangular box and the reference bounding rectangular box;
[0106] The optimal estimation subunit is used to use the horizontal coordinate error and the vertical coordinate error as the state quantity of the Kalman filter, and use the Kalman filter algorithm to perform optimal estimation on the state quantity, so as to obtain the optimal estimation value of the horizontal coordinate error and the optimal estimation value of the vertical coordinate error of each frame of the multiple continuous images;
[0107] A pixel translation amount determination subunit is used to determine the pixel translation amount of each frame of the multiple continuous image frames according to the optimal estimated value of the horizontal coordinate error and the optimal estimated value of the vertical coordinate error of the frame of the image;
[0108] Correspondingly, the pixel shifting unit is used to: for each frame of the multiple continuous images, perform a shifting process on each pixel of the frame according to the pixel shifting amount of the frame.
[0109] In one implementation of the embodiment of the present application, the device for image de-shaking includes:
[0110] The bounding rectangle box deletion module is used to delete the initial bounding rectangle boxes whose corresponding target categories are different from the target categories corresponding to the reference bounding rectangle boxes.
[0111] In one implementation of the embodiment of the present application, the image de-shaking module includes:
[0112] A threshold judgment unit, used to judge whether the minimum error among the first coordinate errors is less than a preset error threshold;
[0113] The first image de-jittering unit is used for performing de-jittering processing on the image to be processed according to the minimum error among the first coordinate errors if the minimum error among the first coordinate errors is less than the error threshold.
[0114] In one implementation of the embodiment of the present application, the image de-shaking module further includes:
[0115] A candidate bounding rectangular frame acquisition unit is used to acquire a pre-stored candidate bounding rectangular frame if the minimum error among the first coordinate errors is greater than or equal to the error threshold; wherein the candidate bounding rectangular frame is a bounding rectangular frame of a candidate stationary target different from the reference stationary target in the reference image;
[0116] A second coordinate error calculation unit, used to respectively calculate a second coordinate error between each initial bounding rectangular box and the candidate bounding rectangular box;
[0117] The second image de-jittering unit is used to perform de-jittering processing on the image to be processed according to the minimum error among the second coordinate errors.
[0118] In one implementation of the embodiment of the present application, the image de-shaking module further includes:
[0119] A vibration amplitude detection unit, used to detect the vibration amplitude of the target camera within a preset time period;
[0120] The error threshold determination unit is used to determine the error threshold according to the vibration amplitude; wherein the error threshold is proportional to the vibration amplitude.
[0121] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image de-shaking method represented by any of the above embodiments is implemented.
[0122] An embodiment of the present application also provides a computer program product. When the computer program product is run on a terminal device, the terminal device executes the image de-shaking method represented by any of the above embodiments.
[0123] Figure 3 Schematic diagram of a terminal device provided by an embodiment of the present application. Figure 3 As shown, the terminal device 3 of this embodiment includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30. When the processor 30 executes the computer program 32, the steps in the above-mentioned embodiments of the method for de-shaking the image are implemented, for example Figure 1 Alternatively, when the processor 30 executes the computer program 32, the functions of each module / unit in the above-mentioned device embodiments are realized, for example Figure 2 Functions of modules 201 to 204 are shown.
[0124] The computer program 32 may be divided into one or more modules / units, which are stored in the memory 31 and executed by the processor 30 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 32 in the terminal device 3.
[0125] The processor 30 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0126] The memory 31 may be an internal storage unit of the terminal device 3, such as a hard disk or memory of the terminal device 3. The memory 31 may also be an external storage device of the terminal device 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 3. Further, the memory 31 may also include both an internal storage unit and an external storage device of the terminal device 3. The memory 31 is used to store the computer program and other programs and data required by the terminal device. The memory 31 may also be used to temporarily store data that has been output or is to be output.
[0127] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0128] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0129] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0130] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0131] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0132] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0133] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0134] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0135] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for image de-jittering, It is characterized in that include: Acquire the image to be processed taken by the target camera; Performing target detection processing on the image to be processed to obtain an initial bounding rectangular frame of each stationary target in the image to be processed; Calculating respectively the first coordinate error between each of the initial bounding rectangular frames and a pre-stored reference bounding rectangular frame; wherein the reference bounding rectangular frame is a bounding rectangular frame of a reference stationary target in a reference image pre-photographed by the target camera in a stationary state; De-jittering is performed on the image to be processed according to the minimum error among the first coordinate errors.
2. The method according to claim 1, It is characterized in that The de-jittering process is performed on the image to be processed according to the minimum error among the first coordinate errors, comprising: Determine the initial bounding rectangular frame whose first coordinate error with the reference bounding rectangular frame is the minimum error as the target bounding rectangular frame corresponding to the reference stationary object; Determining a pixel translation amount according to a first coordinate error between the target bounding rectangular box and the reference bounding rectangular box; According to the pixel translation amount, each pixel of the image to be processed is subjected to translation processing to obtain the image to be processed after de-jittering.
3. The method according to claim 2, It is characterized in that The image to be processed includes a plurality of frames of continuous images; and determining the pixel translation amount according to a first coordinate error between the target bounding rectangular frame and the reference bounding rectangular frame includes: Obtaining a horizontal coordinate error and a vertical coordinate error between the target bounding rectangular box and the reference bounding rectangular box; The horizontal coordinate error and the vertical coordinate error are used as state quantities of a Kalman filter, and the state quantities are optimally estimated using a Kalman filter algorithm to obtain an optimal estimation value of the horizontal coordinate error and an optimal estimation value of the vertical coordinate error of each frame of the multiple frames of continuous images; For each frame of the plurality of continuous images, determining a pixel translation amount of the frame of the image according to an optimal estimated value of abscissa error and an optimal estimated value of ordinate error of the frame of the image; Correspondingly, performing translation processing on each pixel of the image to be processed according to the pixel translation amount includes: For each frame of the multiple frames of continuous images, a translation process is performed on each pixel of the frame according to the pixel translation amount of the frame.
4. The method according to claim 1, It is characterized in that Before respectively calculating the first coordinate error between each of the initial bounding rectangular boxes and the pre-stored reference bounding rectangular box, the method further includes: In each of the initial bounding rectangular frames, the initial bounding rectangular frames corresponding to a target category different from the target category corresponding to the reference bounding rectangular frame are deleted.
5. The method according to claim 1, It is characterized in that The de-jittering process is performed on the image to be processed according to the minimum error among the first coordinate errors, comprising: Determining whether the minimum error among each of the first coordinate errors is less than a preset error threshold; If the minimum error among the first coordinate errors is smaller than the error threshold, de-jitter processing is performed on the image to be processed according to the minimum error among the first coordinate errors.
6. The method according to claim 5, It is characterized in that After determining whether the minimum error among the first coordinate errors is less than a preset error threshold, the method further includes: If the minimum error among the first coordinate errors is greater than or equal to the error threshold, a pre-stored candidate bounding rectangular box is obtained; wherein the candidate bounding rectangular box is a bounding rectangular box of a candidate stationary target different from the reference stationary target in the reference image; respectively calculating a second coordinate error between each of the initial bounding rectangular boxes and the candidate bounding rectangular boxes; De-jittering is performed on the image to be processed according to the minimum error among the second coordinate errors.
7. The method according to claim 5, It is characterized in that The error threshold is determined by: Detecting the vibration amplitude of the target camera within a preset time period; The error threshold is determined according to the vibration amplitude; wherein the error threshold is proportional to the vibration amplitude.
8. The method according to any one of claims 1 to 7, It is characterized in that The reference stationary target is any lane line in the reference image; the target detection processing is performed on the image to be processed to obtain the initial bounding rectangular frame of each stationary target in the image to be processed, including: Lane line detection processing is performed on the image to be processed to obtain an initial boundary rectangular frame of each lane line in the image to be processed.
9. A device for removing image jitter, It is characterized in that include: An image acquisition module is used to acquire the image to be processed taken by the target camera; The target detection module is used to perform target detection processing on the image to be processed to obtain an initial bounding rectangular frame of each stationary target in the image to be processed; A coordinate error calculation module, used to respectively calculate a first coordinate error between each of the initial bounding rectangular frames and a pre-stored reference bounding rectangular frame; wherein the reference bounding rectangular frame is a bounding rectangular frame of a reference stationary target in a reference image pre-photographed by the target camera in a stationary state; An image de-jittering module is used to perform de-jittering processing on the image to be processed according to the minimum error among the first coordinate errors.
10. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the image de-shaking method according to any one of claims 1 to 8 is implemented.
11. A computer-readable storage medium storing a computer program. It is characterized in that When the computer program is executed by a processor, the image de-shaking method according to any one of claims 1 to 8 is implemented.