Vehicle-mounted camera video processing method and device, electronic equipment and vehicle
By performing motion vector clustering and global motion estimation on video frames from vehicle-mounted cameras, the problem of video distortion caused by camera shake is solved, achieving video stability and reliability, and ensuring driver safety.
Patent Information
- Application Number
- CN202511356027.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-11-11
AI Technical Summary
Vehicle-mounted cameras can cause video to shake and become distorted due to vehicle movement, affecting the driver's ability to anticipate driving conditions and endangering personal safety.
By performing motion vector clustering on video frames captured by vehicle-mounted cameras, foreground and background regions are separated, feature points in the background region are detected, and global motion estimation and motion compensation are performed to eliminate jitter.
It improves the stability and reliability of in-vehicle camera video, enhancing the accuracy of driver predictions and personal safety in complex environments.
Smart Images

Figure CN120935459A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automotive video processing technology, specifically to a method, apparatus, electronic device, and vehicle for processing video from an in-vehicle camera. Background Technology
[0002] With the continuous development of science and technology and society, people's demand for driving safety is getting higher and higher.
[0003] Nowadays, in-vehicle cameras are widely used in automotive driving systems. When in-vehicle cameras capture videos, the images are often jittery due to the camera's own shooting or other external factors. Specifically, when a car travels over uneven roads, the car body will experience significant undulations, causing the in-vehicle camera fixed to the car body to move simultaneously, ultimately capturing blurry or distorted videos. If driving judgments are made based on such videos, it will seriously affect the driver's predictions and personal safety. Summary of the Invention
[0004] One objective of this invention is to provide a video processing method for vehicle-mounted cameras, in order to solve the problem in the prior art where the vehicle-mounted camera captures blurry or distorted videos as it follows the movement of the vehicle body. If driving judgments are made based on these videos, it will seriously affect the driver's prediction and personal safety. Another objective is to provide a video processing device for vehicle-mounted cameras. A third objective is to provide a system. A fourth objective is to provide an electronic device. A fifth objective is to provide a vehicle.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] A method for processing video from a vehicle-mounted camera, the method comprising:
[0007] Obtain the original video frames corresponding to the video captured by the vehicle-mounted camera; the original video frames are video frames with jitter caused by the movement of the vehicle-mounted camera.
[0008] Clustering is performed on the first motion vector of the original video frame to determine the foreground and background regions in the original video frame; the background region reflects the motion of the vehicle-mounted camera.
[0009] Feature points are detected from the background region of the original video frame, and the second motion vectors of the feature points in the background region of adjacent original video frames are clustered to determine the target feature points from the feature points;
[0010] Global motion estimation is performed on the original video frame based on the target feature points, and motion compensation is performed on the original video frame based on the global motion estimation results.
[0011] A vehicle-mounted camera video processing device, the device comprising:
[0012] The original video frame acquisition module is used to acquire the original video frame corresponding to the video captured by the vehicle-mounted camera; the original video frame is a video frame with jitter caused by the movement of the vehicle-mounted camera.
[0013] The first clustering module is used to cluster the first motion vector of the original video frame to determine the foreground region and background region in the original video frame; the background region reflects the motion of the vehicle-mounted camera.
[0014] The second clustering module is used to detect feature points from the background region of the original video frame and cluster the second motion vectors of the feature points in the background region of adjacent original video frames to determine the target feature points from the feature points.
[0015] The motion compensation module is used to perform global motion estimation on the original video frame based on the target feature points, and to perform motion compensation on the original video frame based on the global motion estimation result.
[0016] An electronic device includes: a processor; and a memory for storing processor-executable instructions.
[0017] The processor is configured to execute the instructions to implement the above-described vehicle camera video processing method.
[0018] A computer-readable storage medium, when the instructions in the storage medium are executed by the processor of a mobile terminal, enables the mobile terminal to perform the above-described vehicle camera video processing method.
[0019] A vehicle comprising the aforementioned electronic equipment.
[0020] The beneficial effects of this invention are:
[0021] In this embodiment of the invention, the original video frames corresponding to the video captured by the vehicle-mounted camera are obtained. These original video frames are jittery due to the movement of the vehicle-mounted camera. The first motion vector of the original video frames is clustered to determine the foreground and background regions. The background region reflects the movement of the vehicle-mounted camera. Feature points are detected from the background region of the original video frames. Based on the second motion vector of the feature points in the background regions of adjacent original video frames, target feature points are clustered to determine the target feature points. Global motion estimation is performed on the original video frames based on the target feature points, and motion compensation is applied to the original video frames based on the global motion estimation results. This embodiment of the invention performs multiple clusterings on the original video frames, accurately extracting target feature points (optimal clustering) that reflect the movement of the vehicle-mounted camera to perform motion compensation. This removes jittery video frames caused by the movement of the vehicle-mounted camera, making the motion-compensated original video frames stable. This improves the viewing experience of the vehicle driver for the video captured by the vehicle-mounted camera and enhances the reliability of computer analysis based on the video, ensuring the accuracy of the driver's predictions and personal safety when driving in complex environments. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating the steps of a vehicle-mounted camera video processing method provided in an embodiment of the present invention;
[0023] Figure 2 This is a flowchart of a video stabilization method for a vehicle-mounted camera provided in an embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram of a vehicle-mounted camera video stabilization system provided in an embodiment of the present invention;
[0025] Figure 4 This is a schematic diagram of the structure of the vehicle-mounted camera video stabilization system provided in an embodiment of the present invention;
[0026] Figure 5 This is a schematic diagram of the structure of a vehicle-mounted camera video processing device provided in an embodiment of the present invention;
[0027] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0028] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.
[0029] It should be noted that the embodiments of the present invention may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).
[0030] Reference Figure 1 The diagram illustrates a flowchart of a video processing method for an in-vehicle camera provided in an embodiment of the present invention, specifically including the following steps:
[0031] Step 101: Obtain the original video frame corresponding to the video captured by the vehicle-mounted camera; the original video frame is a video frame with jitter caused by the movement of the vehicle-mounted camera.
[0032] In practical implementation, one or more vehicle-mounted cameras can be installed on the vehicle. These cameras can capture video of the vehicle's surroundings while the vehicle is in motion and display the video to the driver on the vehicle's screen. Simultaneously, the video can be used to enable intelligent driving. It's understandable that because the vehicle-mounted cameras are installed on the vehicle, they move with the vehicle. For example, if there are speed bumps on the road, the vehicle will shake up and down when it goes over them. The video captured by the vehicle-mounted cameras will contain unstable (shaky) frames due to the camera's movement, resulting in blurry or distorted video frames.
[0033] In this embodiment of the invention, multiple video frames from the video captured by the vehicle-mounted camera can be obtained as the original video frames. Then, the motion path of the vehicle-mounted camera can be estimated using the original video frames. Subsequently, motion compensation can be performed on the original video frames based on the motion path, thereby eliminating the problem of video instability caused by the movement of the vehicle-mounted camera in the original video frames.
[0034] Step 102: Cluster the first motion vector of the original video frame to determine the foreground region and background region in the original video frame; the background region reflects the motion of the vehicle camera.
[0035] In practice, a motion vector is a mathematical vector that describes the distance and direction in which an image or a pixel moves from the previous video frame to the current video frame.
[0036] The original video frame can be divided into two regions: the foreground region and the background region. The foreground region is where moving objects (such as pedestrians, bicycles, or other vehicles) are located in the original video frame, while the background region is the relatively static environmental background (such as buildings, utility poles, traffic lights, or roads) relative to the moving objects. The background region appears dynamic when the video is displayed on the screen due to the movement of the vehicle-mounted camera. Therefore, the background region reflects the movement of the vehicle-mounted camera. Estimating the motion path of the vehicle-mounted camera based on the background region is used for motion compensation of the original video frame, accurately eliminating the instability in the video captured by the vehicle-mounted camera due to its movement. In some embodiments, the clustering method can be K-means or other methods, and this embodiment of the invention does not impose any limitations on this. Pixels in the original video frame are categorized into foreground and background regions. To prevent local motion from affecting global motion, the motion vectors between two original video frames are clustered to remove local motion.
[0037] In this embodiment of the invention, after obtaining the original video frame corresponding to the video captured by the vehicle-mounted camera, a first motion vector can be generated based on the original video frame. Then, the first motion vector can be clustered to determine the foreground region and background region in the original video frame based on the clustering result. Then, the foreground region in the original video frame can be eliminated, that is, only the background region is retained to estimate the motion path of the vehicle-mounted camera.
[0038] Step 103: Detect feature points from the background region of the original video frame, and cluster the second motion vectors of the feature points in the background region of adjacent original video frames to determine the target feature points from the feature points.
[0039] In this embodiment of the invention, feature points can be detected from the background region of the original video frame. Feature points are pixels or pixel blocks in the original video frame that distinguish it from the surrounding area and are frame structures that can be reliably detected, such as corners, edges, and spots. Feature points carry key information about the original video frame, such as texture, gradient, and orientation; therefore, they are pixels with crucial information in the original video frame and can be used as stable anchor points for matching and recognition. In some embodiments, the SURF (Speeded Up Robust Features) algorithm can be used to detect feature points in the original video frame.
[0040] After obtaining the feature points of each original video frame, a second motion vector can be generated based on the feature points of the background region of adjacent original video frames. Then, the second motion vector can be clustered, and the target feature points can be determined from the feature points based on the clustering results.
[0041] Step 104: Perform global motion estimation on the original video frame based on the target feature points, and then perform motion compensation on the original video frame based on the global motion estimation results.
[0042] In this embodiment of the invention, the target feature point is the feature point that most accurately reflects the motion of the vehicle camera after multiple clusterings in the original video frame. That is, unnecessary pixels that would affect the global motion estimation have been removed. Therefore, the original video frame is used for global motion estimation based on the target feature point, and motion compensation is performed on the original video frame based on the global motion estimation result. The original video frame is stable after motion compensation.
[0043] In the aforementioned vehicle-mounted camera video processing method, the original video frames corresponding to the video captured by the vehicle-mounted camera are obtained. These original video frames are jittery due to the movement of the vehicle-mounted camera. The first motion vector of the original video frames is clustered to determine the foreground and background regions, where the background region reflects the movement of the vehicle-mounted camera. Feature points are detected from the background region of the original video frames. Clustering is then performed based on the second motion vector of the feature points in the background regions of adjacent original video frames to determine target feature points. Global motion estimation is then performed on the original video frames based on the target feature points, and motion compensation is applied to the original video frames based on the global motion estimation results. This embodiment of the invention performs multiple clustering operations on the original video frames, accurately extracting target feature points that reflect the movement of the vehicle-mounted camera to perform motion compensation. This removes jittery video frames caused by the movement of the vehicle-mounted camera, making the motion-compensated original video frames stable. This improves the viewing experience for drivers of the video captured by the vehicle-mounted camera and enhances the reliability of computer analysis based on the video, ensuring the accuracy of driver predictions and personal safety when driving in complex environments.
[0044] In one embodiment of the present invention, step 102, clustering the first motion vector of the original video frame to determine the foreground and background regions in the original video frame, may include:
[0045] Each of the original video frames is divided into superpixel blocks; each superpixel block has a corresponding class label and a superpixel centroid corresponding to the class label.
[0046] Determine the first distance between the centroids of superpixels with the same class label in adjacent original video frames;
[0047] When the first distance meets the preset distance, a first motion vector between adjacent superpixel blocks is determined;
[0048] The foreground and background regions are determined from the original video frame based on the first motion vector.
[0049] The first distance can be Euclidean distance, which is a distance generated by measuring the shortest straight-line distance between two points in space. Of course, the first distance can also be cosine distance, etc., and this embodiment of the invention does not limit it.
[0050] In the specific implementation, the pixels of the original video frame are divided into multiple superpixel blocks by utilizing the similarity attributes of the image (original video frame). The pixels in the superpixel block have similar features such as color, texture, and brightness. The original video frame is divided into regions with greater semantic consistency and visual connectivity. Based on the superpixel blocks, the local features of the image can be captured better.
[0051] Each superpixel block has a corresponding class label and a corresponding superpixel centroid. The class label is a unique identifier used to identify and distinguish each superpixel block in the original video frame. For example, the class labels of superpixel blocks in the first original video frame can be 1, 2, 3, 4, ..., the class labels of superpixel blocks in the second original video frame can be 1, 2, 3, 4, ..., and so on. The superpixel centroid is a point representing the average position and average color characteristics of a superpixel block. In some embodiments, the average position and average color characteristics of a superpixel block can be determined based on the pixel coordinates and pixel color values of each superpixel block's pixels.
[0052] In this embodiment of the invention, after dividing each original video frame into multiple superpixel blocks, a first distance can be determined between the centroids of superpixels with the same class label in adjacent original video frames. For example, the first distance can be determined between the centroids of superpixel blocks with class label 1 in the first original video frame and superpixel blocks with class label 1 in the second original video frame. If the first distance meets a preset distance, for example, if the first distance is less than a preset distance, it indicates that the two superpixel blocks are a valid match. Then, a first motion vector between adjacent superpixel blocks can be determined, and the foreground and background regions can be determined from the original video frames based on the first motion vector. Of course, if the first distance does not meet the preset distance, for example, if the first distance is greater than or equal to the preset distance, it indicates that the two superpixel blocks are an invalid match, and the superpixel blocks can be eliminated and not used for determining the foreground and background regions in subsequent original video frames.
[0053] In the above embodiments, the original video frame is divided into multiple superpixel blocks, and then the superpixel blocks and their class labels are matched with the superpixel centroids and motion vectors are calculated. Furthermore, superpixel blocks with incorrect matching and abnormal displacement can be effectively filtered out by a preset distance threshold, ensuring the reliability of the first motion vector. In turn, the clustering of the first motion vector can accurately distinguish the foreground and background regions in the original video frame, so as to accurately estimate the global motion through the background region, thus ensuring the stability and accuracy of the video captured by the vehicle camera.
[0054] In one embodiment of the present invention, determining the foreground region and background region from the original video frame based on the first motion vector may include:
[0055] Two initial motion vectors are randomly selected from the first motion vectors;
[0056] Determine a second distance between each of the first motion vectors and the first initial motion vector;
[0057] Based on the second distance, the first motion vector is divided into a first motion vector cluster corresponding to the first initial motion vector;
[0058] Determine the first degree of difference of the first motion vector in each of the first motion vector clusters;
[0059] When the first difference exceeds the preset first difference, a first mean is determined based on the first motion vector in the first motion vector cluster, the first mean is used as the first initial motion vector, and the step of determining the second distance between each first motion vector and the first initial motion vector is returned.
[0060] When the first degree of difference does not exceed the preset first degree of difference, the magnitude of the first motion vector cluster is determined;
[0061] The region of the original video frame corresponding to the first motion vector cluster with the smallest magnitude of the first motion vector cluster is determined as the background region, and the region of the original video frame corresponding to the first motion vector cluster with the largest magnitude of the first motion vector cluster is determined as the foreground region.
[0062] In this embodiment of the invention, when dividing the foreground region and background region from the original video frame, two first motion vectors can be randomly initialized from multiple first motion vectors as first initial motion vectors (cluster centers). Then, by iteratively calculating the Euclidean distance between each first motion vector and the cluster center, the first motion vector is divided into the first motion vector cluster with the closest distance. Subsequently, the first degree of difference (difference degree) of the first motion vector in each first motion vector cluster is determined. The degree of difference refers to the distance between a motion vector and the cluster center of its first motion vector cluster, such as Euclidean distance. The degree of difference can be used to measure the degree of deviation of a single point from the cluster center.
[0063] If the degree of difference of the first motion vectors in the first motion vector cluster does not exceed a preset first degree of difference, it indicates that the first motion vector is likely to be assigned to an accurate first motion vector cluster. Conversely, if the degree of difference of the first motion vectors in the first motion vector cluster exceeds the preset first degree of difference, it indicates that the first motion vector is likely not assigned to an accurate first motion vector cluster. Then, the first mean value calculated from all the first motion vectors within the first motion vector cluster is used as the first initial motion vector (new cluster center). This process continues until the degree of difference between the first vectors within the first motion vector cluster and the cluster center of the first motion vector cluster is lower than the preset first degree of difference, at which point the convergence condition is considered met. Alternatively, the convergence condition can be considered met when a preset number of iterations is reached. After the convergence condition is met, the foreground and background regions are determined based on the magnitude (amplitude) of the converged first motion vector clusters. The region corresponding to the first motion vector cluster with the smaller magnitude in the original video frame is the background region, and the region corresponding to the first motion vector cluster with the larger magnitude in the original video frame is the foreground region.
[0064] In the above embodiments, the foreground and background regions in the original video frame can be accurately distinguished by clustering the first motion vector. Subsequently, global motion estimation can be performed using the background region, thereby improving the accuracy of motion compensation for the original video frame.
[0065] In one embodiment of the present invention, dividing each of the original video frames into superpixel blocks may include:
[0066] Cluster centers are evenly placed in the original video frames; each cluster center has a corresponding label.
[0067] Determine a third distance between each cluster center and pixels within a specified range of the cluster center; the third distance includes at least color, texture, brightness, and spatial distance.
[0068] Assign the label corresponding to the cluster center that is closest to the pixel in the third distance to the pixel to the pixel, and merge the pixels with the same label into a pixel cluster;
[0069] A new cluster center is determined based on the cluster center corresponding to the pixel cluster;
[0070] Determine the center shift distance between the new cluster center and the existing cluster center;
[0071] When the center movement distance is greater than or equal to the preset center movement distance, the new cluster center is taken as the cluster center, and a third distance is returned between each cluster center and the pixels within a specified range of the cluster center.
[0072] When the center movement distance is less than the preset center movement distance, the pixel cluster is determined as a superpixel block.
[0073] In this embodiment of the invention, multiple cluster centers can be evenly placed in the original video frame, for example, 100 or 200 cluster centers. A corresponding label is assigned to each cluster center. A third distance is determined between each cluster center and pixels in the original video frame within a specified range of the cluster center. This third distance includes at least color, texture, brightness, and spatial distances. The third distance can be a comprehensive distance calculated using a weighted average of these distances. The label corresponding to the cluster center with the closest third distance to the pixel is assigned to that pixel. By assigning pixel labels within a specified range around the cluster centers, computational complexity is reduced, processing speed is ensured, and thus the processing speed of the vehicle-mounted camera video can be improved.
[0074] Subsequently, pixels with the same label can be merged into a pixel cluster containing multiple pixels. Then, based on the cluster center corresponding to this pixel cluster, a new cluster center can be determined using methods such as mean calculation. The center shift distance (which can be Euclidean distance) between the new cluster center and the existing cluster center can be determined. If the center shift distance is greater than or equal to a preset center shift distance, it indicates that the new cluster center has not reached the convergence condition and exceeds the allowable error range. In this case, the new cluster center can be used as the new cluster center, and the above steps can be iterated. If the center shift distance is less than the preset center shift distance, it indicates that the new cluster center has reached the convergence condition and does not exceed the allowable error range. In this case, the pixel cluster can be determined as a superpixel block. A superpixel block is a region in the original video frame that demonstrates semantic consistency and visual connectivity. Alternatively, convergence can be considered achieved after reaching a preset number of iterations.
[0075] In the above embodiments, pixel clusters can be merged based on the color, texture, brightness, and spatial distance between pixels within a specified range of the cluster centers. Then, the center movement distance between the new cluster centers of the pixel clusters can be determined to automatically determine whether the convergence condition has been met. If so, the pixel clusters can be identified as superpixel blocks, ensuring that the superpixel blocks have good semantic consistency and visual connectivity, providing reliable data for subsequent global motion estimation and video processing.
[0076] In one embodiment of the present invention, clustering the second motion vectors of feature points in the background region of adjacent original video frames to determine target feature points from the feature points may include:
[0077] Obtain the second motion vector of feature points in the background region of adjacent original video frames;
[0078] Randomly select multiple second initial motion vectors from the second motion vector;
[0079] Determine a fourth distance between each of the second motion vectors and the second initial motion vector;
[0080] Based on the fourth distance, the second motion vector is divided into the second motion vector cluster corresponding to the second initial motion vector;
[0081] Determine the second degree of difference of the second motion vectors in each of the second motion vector clusters;
[0082] When the second difference exceeds the preset second difference, a second mean is determined based on the second motion vector in the second motion vector cluster, the second mean is used as the second initial motion vector, and the step of determining the fourth distance between each second motion vector and the second initial motion vector is returned.
[0083] When the second degree of difference does not exceed the preset second degree of difference, the magnitude of the second motion vector cluster is determined;
[0084] Feature points corresponding to the second motion vector cluster whose magnitude is less than a preset magnitude threshold are identified as target feature points.
[0085] In this embodiment of the invention, after dividing the foreground region and background region from the original video frame, a second motion vector based on the feature points of the background region of adjacent original video frames can be obtained. Then, multiple second motion vectors can be randomly initialized from multiple second motion vectors as second initial motion vectors (cluster centers). Then, by iteratively calculating the distance (e.g., Euclidean distance) between each second motion vector and the cluster center, the second motion vector is divided into the second motion vector clusters with the closest distance. Subsequently, the second degree of difference (difference degree) of the second motion vectors in each second motion vector cluster is determined, where the degree of difference refers to the distance between a motion vector and the cluster center of its second motion vector cluster, such as Euclidean distance. The degree of difference can be used to measure the degree of deviation of a single point from the cluster center.
[0086] If the second difference degree of a second motion vector in a second motion vector cluster does not exceed a preset second difference degree, it indicates that the second motion vector has been likely assigned to an accurate second motion vector cluster. Conversely, if the second difference degree of a second motion vector in a second motion vector cluster exceeds the preset second difference degree, it indicates that the second motion vector has been likely not assigned to an accurate second motion vector cluster. In this case, the second mean calculated from all second motion vectors within the second motion vector cluster can be used as the second initial motion vector (new cluster center). This process continues until the difference degree between the second vector within the second motion vector cluster and the cluster center of the second motion vector cluster is lower than the preset second difference degree, at which point the convergence condition is considered met. Alternatively, the convergence condition can be met when a preset number of iterations is reached. After the convergence condition is met, the magnitude (amplitude) of the converged second motion vector cluster is used to determine whether the corresponding feature points can be used as target feature points. Feature points corresponding to second motion vector clusters with a magnitude less than a preset magnitude threshold can be identified as target feature points, while feature points corresponding to second motion vector clusters with a magnitude greater than or equal to the preset magnitude threshold can be eliminated and not participate in subsequent global motion estimation.
[0087] In the above embodiments, the target feature points in the background region can be accurately distinguished by clustering the second motion vector. Subsequently, global motion estimation can be performed using the target feature points in the background region, thereby improving the accuracy of motion compensation for the original video frame.
[0088] In one embodiment of the present invention, step 104, performing global motion estimation on the original video frame based on the target feature points, and then performing motion compensation on the original video frame based on the global motion estimation result, may include:
[0089] The motion path of the original vehicle camera is obtained by performing global motion estimation on the target feature points matched between adjacent original video frames.
[0090] The motion path of the original vehicle-mounted camera is smoothed to obtain the motion path of the target vehicle-mounted camera.
[0091] Calculate the compensation transformation matrix based on the difference between the target vehicle camera's motion path and the original vehicle camera's motion path;
[0092] The original video frame is inversely mapped and resampled according to the compensation transformation matrix to perform motion compensation on the original video frame.
[0093] In this embodiment of the invention, feature point detection is performed on adjacent original video frames to obtain target feature points. Global motion estimation is then performed using the matched target feature points between adjacent original video frames to obtain a motion transformation model. This global motion estimation can be a homography transformation. The original vehicle camera motion path is then obtained by integrating the motion transformation model. This original vehicle camera motion path is then smoothed to construct a smooth (stable) target vehicle camera motion path. This accurately separates the intentional motion of the vehicle camera that needs to be retained from the unintentional jitter that needs to be eliminated, thus providing an ideal vehicle camera motion path for generating a stable video. A compensation transformation matrix is then calculated based on the difference between the target vehicle camera motion path and the original vehicle camera motion path. The original video frames are then inversely mapped and resampled based on the compensation transformation matrix to accurately compensate for motion, resulting in a stable video (original video frame).
[0094] In the above embodiment, global motion estimation is performed on the target feature points matched between adjacent original video frames to obtain the original vehicle camera motion path, and a balancing process is performed to obtain an ideal target vehicle camera motion path. Then, motion compensation can be performed on the original video frames based on the difference between the two paths to stabilize the video captured by the vehicle camera.
[0095] In one embodiment of the present invention, the step of performing global motion estimation on the target feature points matched between adjacent original video frames to obtain the motion path of the original vehicle camera may include:
[0096] The original video frame is divided into multiple grids;
[0097] Based on the target feature points and the motion vectors corresponding to the target feature points, the mesh is processed using a multihomography estimation strategy to obtain an initialized mesh;
[0098] The original motion path of the target feature point is reconstructed based on the initial mesh to obtain the reconstructed motion path, and the motion optimization mesh is determined based on the reconstructed motion path and the residual between the reconstructed motion path and the original motion path.
[0099] The motion paths of the motion optimization mesh are integrated to obtain the original motion path of the vehicle camera.
[0100] In one embodiment of the present invention, the step of processing the mesh using a multihomography estimation strategy based on the target feature points and the motion vectors corresponding to the target features to obtain an initialized mesh may include:
[0101] Optical flow between adjacent original video frames is extracted using a pre-set optical flow estimation model;
[0102] A pre-set keypoint detector model is used to detect feature points and target feature points from the background region; the target feature point information includes at least the position, orientation, and scale information of the feature points.
[0103] The motion vector of the target feature point is determined based on the optical flow and target feature point information;
[0104] Clustering is performed on the target feature points of each original video frame to divide the target feature points into corresponding planar clusters; wherein, when the number of target feature points in any planar cluster is lower than a preset threshold, the target feature points of the planar cluster are merged into another planar cluster;
[0105] The planar perceptual homography matrix is determined based on the motion vectors of the target feature points around the vertices of the grid and belonging to the same plane;
[0106] The initial mesh is obtained by transforming the mesh according to the plane-sensory homography matrix.
[0107] In practical implementations, traditional motion estimation algorithms rely heavily on manually extracting features to estimate motion paths. Low-texture content, occlusion areas, and dynamic objects common in complex scenes can all affect these features, thus reducing image stabilization performance. Therefore, this invention does not simply calculate motion vectors between pixels (feature points) to measure motion estimation; instead, it performs motion estimation based on a grid. Specifically, it divides the original video frame into multiple grids and then estimates the motion vectors of the corresponding feature points in all grids, thereby improving the accuracy of motion estimation.
[0108] Specifically, firstly, the motion vectors of target feature points in each original video are determined. These motion vectors are more robust to noise, lighting changes, and occlusion. Based on the target feature points and their motion vectors, a multi-homography estimation strategy is used to initialize the mesh-based motion to obtain an initialization network. Then, a motion optimization network is determined based on the initialization mesh. Finally, the motion paths of the motion optimization network are integrated to obtain the original motion path of the vehicle camera for global motion estimation. This allows for further improvement in motion estimation accuracy based on the original vehicle camera motion path. The original vehicle camera motion path can be generated by temporally associating the mesh-based motion paths. The training objective is set by utilizing the spatial and temporal consistency between target feature points and mesh vertices. The supervised network learns the feature representation of each subtask. In this way, the network can be trained in an unsupervised manner, eliminating the need to collect complex paired image stabilization datasets, thus broadening its applications.
[0109] In this embodiment of the invention, a motion initialization module is introduced. This module determines an initial grid based on target feature points and their motion vectors, aiming to reduce the impact of dynamic objects on the estimated camera motion. Since motion caused by camera shake is spatially correlated, and motion caused by dynamic objects is always different from this motion, using grid-based motion to enhance correlation is more accurate than pixel-based motion and can suppress the influence of dynamic objects.
[0110] In the motion initialization phase, an optical flow estimation model is used to extract the optical flow between the original video frames. Optical flow is the apparent motion pattern of objects, surfaces, or edges in the video on the imaging plane, and motion vectors can be determined based on optical flow. Considering that pixel-based motion includes noise and occlusion from dynamic objects, general image stabilization methods use a set threshold to eliminate the influence of dynamic objects. This requires the user to select an appropriate threshold for each jittery video, a process that is cumbersome and has limitations. Therefore, this embodiment of the invention utilizes the spatial correlation of motion caused by jitter, using sparse target feature points as an intermediate representation to estimate a grid-based motion path from the motion vector based on the target feature points, thereby reducing the influence of dynamic objects.
[0111] First, the keypoint detector model is used to detect the feature point information of the target feature point. The feature point information can include at least the position, orientation and scale of the feature point. The keypoint detector model extracts a fixed-length feature vector from the original video frame for matching. The region where the target feature point is located is used as a reliable region in the optical flow. The motion vector of the target feature point can be obtained based on the coordinate position of the target feature point.
[0112] Then, the mesh-based motion path can be initialized based on the motion path of the target feature points. Specifically, each original video frame is divided into a uniform mesh. Combining the extracted target feature point motion, clustering algorithms such as K-means clustering are used to divide each original video frame into two planar clusters. Each original video frame belongs to one planar cluster. A preset threshold is set, which can be 20% of the number of target feature points (other values can be chosen based on actual conditions; this embodiment does not impose such restrictions). If the number of target feature points in either planar cluster is less than the preset threshold, that planar cluster is merged into the other planar cluster. Then, for each mesh, a planar perceptual homography matrix is determined based on the motion vectors of the target feature points around the mesh vertices that belong to the same plane. The mesh is then transformed based on the planar perceptual homography matrix to obtain the initialized mesh.
[0113] Subsequently, to accurately describe the characteristics of target feature points, distance features and motion features are two indispensable factors. Distance features describe the distance between the target feature point and the grid vertices, while motion features represent the target object's motion direction and trend. Motion features can estimate the motion path between consecutive frames over time, improving the accuracy of motion estimation. Referring to general methods for feature learning, these methods lack feature selection mechanisms, which may lead to inaccurate attention to key features and the mixing of redundant features with key features, reducing model performance. Therefore, this embodiment proposes using a motion-aware SE (Squeeze-and-Excitation) fusion module (motion-aware fusion module) to add a feature selection mechanism for extracting and enhancing motion features. This module can adaptively select and weight feature channels during motion feature processing, improving the ability to focus on and represent motion features, reducing redundancy and noise, and facilitating the acquisition of temporally continuous motion, thereby improving model performance.
[0114] By utilizing target feature points and their motion vectors, a mesh-based motion initialization process can be completed. After initialization, the mesh can reduce the impact of dynamic objects on the motion of the original video frames. This invention proposes a motion optimization module to improve the accuracy of motion estimation. Specifically, the motion optimization module first uses the motion paths of the vertices of the initialized mesh to reconstruct the motion paths of the detected target feature points. Simultaneously, it calculates the residual between the reconstructed motion path and the original motion path of the target feature points to obtain the motion-optimized mesh. The smaller the residual, the more accurate the motion path estimation.
[0115] Since target feature points may still overlap with dynamic objects, this would minimize the residual of each target feature point in the estimated motion path, potentially introducing additional noise. Therefore, from a motion correlation perspective, it is necessary to distinguish whether target feature points are located on a static background or on a dynamic object. Thus, an assumption is proposed: motion on a static background should be consistent, while motion on a dynamic object should not follow a global consistency principle. Motion correlation is used to distinguish the object to which a target feature point belongs, thereby obtaining accurate motion estimation. Specifically, the sparse target feature points and the relatively dense grid vertices are treated as two point clouds with motion and 2D (two-dimensional) positional distribution. Several 1D (one-dimensional) convolutions are used to associate the residual motion of the target feature points with each grid vertex, thereby fine-tuning the motion path of the grid vertices. Although the previously mentioned motion-aware fusion module is introduced, since the importance of distance features may be similar at different locations or in different feature dimensions, this module adaptively selects and weights only motion features after feature integration. Specifically, motion features typically exhibit stronger temporal correlation, and their core role is to capture the dynamic changes of the input sequence in the time dimension. By introducing a one-dimensional SE attention mechanism, the model can adaptively learn and focus on feature channels related to key motions in the motion perception fusion module, thereby enhancing the representation ability of motion features and ensuring better temporal continuity of the subsequent motion estimation results. Based on this, the motion features and distance features of the target feature points are fused and summed to generate aggregated attention features. These features are further input to the motion decoder to predict a dense residual motion field, which in turn calculates the motion vector for each grid vertex. Finally, by integrating the motion paths on all grids, the original motion path of the vehicle-mounted camera for global motion estimation is obtained.
[0116] In one embodiment of the present invention, step 101, obtaining the original video frame corresponding to the video captured by the vehicle-mounted camera, may include:
[0117] Acquire video from the vehicle's onboard camera;
[0118] The video captured by the vehicle-mounted camera is processed by frame extraction to obtain the original video frames corresponding to the video captured by the vehicle-mounted camera.
[0119] Among them, frame extraction processing refers to extracting video frames from continuous video captured by vehicle-mounted cameras at certain intervals.
[0120] In this embodiment of the invention, video captured by a vehicle-mounted camera can be acquired, and then multiple original video frames can be obtained from the acquired video to perform global motion estimation. In this way, it is not necessary to process all video frames of all videos captured by all vehicle-mounted cameras, which reduces the amount of data processing, improves the efficiency of global motion estimation and motion compensation of video, and improves the efficiency of generating stable video.
[0121] In one embodiment of the present invention, the vehicle-mounted camera is an infrared vehicle-mounted camera, and the original video frames captured by the infrared vehicle-mounted camera are infrared video frames. After obtaining the original video frames corresponding to the video captured by the vehicle-mounted camera, the method further includes:
[0122] Convert the original video frame into an original video frame in a specified color space;
[0123] A five-dimensional feature vector for each pixel is determined from the original video frame converted to a specified color space. The five-dimensional feature vector includes color space component values and position coordinates. The five-dimensional feature vector is used to determine the third distance, the new cluster center, and the center movement distance.
[0124] In practical implementation, mainstream perception sensors currently include LiDAR, visible light cameras, and millimeter-wave radar. These are often combined in intelligent driving systems. While they provide drivers with a wealth of information during driving, their reliability is not guaranteed depending on the severity of the environment. To address these issues, the vehicle-mounted camera in this embodiment of the invention can be an infrared vehicle-mounted camera. Infrared vehicle-mounted cameras represent a current focus of development in intelligent driving, and their advantages of passive imaging, all-weather operation (unaffected by rain, snow, or fog), and all-time operation (unaffected by glare, backlight, or darkness) make them the optimal solution for adverse weather and complex operating conditions.
[0125] For example, adverse weather and complex operating conditions can include: Nighttime environments: Nighttime is the peak time for traffic accidents, accounting for more than half of all traffic accidents nationwide each year. Infrared vehicle cameras, with their passive imaging capabilities, do not rely on visible light imaging, allowing them to perform well even in complete darkness. Glare environments: Typical examples of glare environments are the "black hole effect" and the "white hole effect." The "black hole effect" refers to the phenomenon where, when a driver enters a tunnel and the light suddenly changes from bright to dark, their pupils dilate rapidly, and the eyes cannot adapt quickly enough, resulting in a perception of complete darkness, as if the tunnel entrance were a black hole. Conversely, when exiting the tunnel, the light suddenly changes from dark to bright, causing the pupils to constrict rapidly, and the human eye sees a white light, known as the "white hole effect." Traffic accidents can occur within these brief seconds of glare. Glare environments are not only unbearable for the human eye, but also for visible light cameras. However, infrared vehicle cameras, because they only receive long-wavelength information and not visible light information, have their perception function completely unaffected in scenarios with alternating strong and weak light. Similar severe weather and complex working conditions can include rain, fog, snow, alternating light and dark environments, and complete darkness. These are all scenarios that visible light cameras and the naked eye cannot accurately identify.
[0126] The original video frames captured by the infrared vehicle-mounted camera in this embodiment of the invention are infrared video frames (infrared images). The infrared video frames can be preprocessed. Since the infrared video frames themselves have low contrast, the SLIC (Simple Linear Iterative Clustering) algorithm is used for preprocessing based on the characteristics of the infrared video frames themselves to retain useful information such as spatial texture, color, and brightness. Before the SLIC algorithm performs superpixel segmentation on the infrared video frames, the infrared video frames need to be converted from the RGB color space to the Lab color space (specified color space). Furthermore, the RGB image needs to be converted into a five-dimensional feature vector that includes the Lab color space and XY coordinates. The five-dimensional feature vector can include color features (color space component values) and position coordinates (spatial features). The three color space component values of the color features in the Lab color space are brightness (L) and color opposite dimensions (a, b). The spatial features refer to the coordinate positions (x, y) of the pixels in the image. By using five-dimensional feature vectors, the SLIC algorithm can simultaneously measure the color similarity and spatial proximity between pixels in a unified five-dimensional space, thereby generating superpixel blocks that both fit the edges of infrared video frames and maintain spatial compactness.
[0127] After color space conversion, a preset number K is set for the number of superpixel blocks in the infrared video frame. For example, K can be 900. Based on the preset number, a suitable segmentation result of the superpixel blocks can be obtained. In some embodiments, for the infrared video frame segmented into superpixels, the compactness coefficient between adjacent superpixels can be set to 100. The compactness coefficient controls the shape and boundary fit of the generated superpixels, determining how the SLIC algorithm balances color similarity and spatial proximity.
[0128] In this embodiment of the invention, in order to obtain more accurate superpixel blocks that better match the characteristics of the infrared video frame itself, an adaptive region propagation merging algorithm is used to merge the small pixel blocks initially processed by SLIC, and the number of superpixel blocks is set. However, the superpixel to which the original pixel belongs is dynamically determined. At the same time, considering only the Lab color space, the average value of the superpixels can be calculated, and the best superpixel with the closest average value is selected as the center of region propagation to determine whether to merge superpixel blocks. In this way, some superpixel blocks after superpixel segmentation can be connected, reducing the number of superpixel blocks while reducing the amount of computation, and avoiding interference from too many pixel blocks in the calculation results.
[0129] In this embodiment of the invention, the third distance, the new cluster center, and the center movement distance can be determined based on the corresponding five-dimensional feature vector of each pixel. This allows for simultaneous and balanced consideration of the apparent similarity (color) and spatial proximity between pixels. For example, pixels with similar colors but far apart will not be assigned to the same superpixel block, and pixels that are close together but have large color differences will not be forcibly merged, thereby achieving a better superpixel block segmentation effect.
[0130] Furthermore, when the original video frame is an infrared video frame, the ORB (Oriented FAST and Rotated BRIEF) algorithm can be used to detect feature points in the original video frame. Specifically, the ORB algorithm is a fast feature point extraction and description algorithm. For infrared video frames, and specifically for driving applications, the ORB algorithm's running time is significantly better than the SIFT (Scale-Invariant Feature Transform) and SURF algorithms, offering more real-time feature detection and ensuring processing efficiency. The ORB algorithm detects feature points with scale and rotation invariance, and is also invariant to noise and perspective transformations. Its excellent performance makes ORB widely applicable in feature description scenarios. ORB feature detection is mainly divided into the following two steps: (1) Orientation FAST feature point detection, to quickly find sufficiently significant pixels (feature points) in infrared video frames and assign them an orientation; (2) BRIEF feature description, to create a unique and compact descriptor for each detected feature point. The descriptor describes the appearance of the image region around the feature point (such as texture, gradient, etc.) for subsequent matching, such as calculating Euclidean distance, etc.
[0131] To enable those skilled in the art to better understand the embodiments of the present invention, a specific example is used for illustration below.
[0132] Reference Figure 2 This is a flowchart of a video stabilization method for a vehicle-mounted camera provided in an embodiment of the present invention. Specifically, it may include the following steps: connecting to the target vehicle-mounted system and extracting the original video frames; performing superpixel processing on the original video frames to capture local features; matching adjacent original video frames to calculate superpixel motion vectors and eliminating local motion blocks through clustering; detecting and matching feature points of adjacent original video frames, calculating motion vectors to establish a motion space, clustering again to find the optimal solution, and eliminating mismatched points; global motion estimation; and smoothing and compensating the original video frames based on the results of the global motion estimation to obtain a stable video.
[0133] Reference Figure 3This is a schematic diagram of a vehicle-mounted camera video stabilization system provided in this embodiment of the invention. Specifically, it may include: a data acquisition module, which acquires video of the vehicle while it is in motion via a vehicle-mounted camera and stores it in a vehicle-mounted recorder; a data extraction module, used to extract video from the vehicle-mounted recorder, specifically through frame extraction; a pixel processing module, which divides the extracted video frames into regions, obtains the foreground and background regions of the video frames, and processes pixels that meet the corresponding requirements to accurately locate the target region; a feature detection and extraction module, used to establish the motion space of target feature points; a motion estimation module, which finds the optimal solution through two clustering processes to obtain an accurate global motion offset (global motion estimation); a motion smoothing module, which calculates the motion path of the vehicle-mounted camera through motion estimation and smooths the motion path to ensure the continuity of motion; and a motion compensation module, which uses a compensation matrix to perform geometric transformation on the smoothed motion path to obtain a stable video. The entire vehicle-mounted camera video stabilization system of this embodiment of the invention is implemented in multiple modules. This modular approach facilitates problem localization, analysis, and improvement when issues arise. In some embodiments, in the motion smoothing module, the motion path of the original vehicle camera is smoothed. Based on the difference between the smoothed original vehicle camera motion path and the original vehicle camera motion path, the difference is minimized so that the smoothed original vehicle camera motion path and the original vehicle camera motion path are as consistent as possible in terms of motion characteristics. This can avoid distortion and over-cropping after image stabilization, thus obtaining the optimal smoothing path result.
[0134] Reference Figure 4 This is a schematic diagram illustrating the process of achieving more accurate motion estimation based on superpixel segmentation in a video stabilization method for vehicle-mounted cameras provided in this embodiment of the invention. Specifically, it may include: preprocessing video frames (images) using superpixel segmentation to obtain reliable image information based on the image's inherent attributes; calculating the motion vector of the superpixel centroid points corresponding to the class labels of each group of superpixel blocks in adjacent frames; dividing the obtained motion vectors using clustering to distinguish between foreground and background regions, and removing local motion blocks corresponding to motion vectors with larger cluster center values; performing feature point detection and matching on the two frames after removing local motion blocks, calculating motion vectors and establishing a motion space; and clustering the motion vectors corresponding to the feature point pairs again, setting a reasonable threshold, and calculating the optimal clustering to obtain a more accurate global motion estimation result.
[0135] By applying the embodiments of the present invention, a preprocessing method based on superpixels is used to divide the video frames of the video captured by the vehicle camera into regions by combining clustering, and then global motion estimation is performed based on the divided superpixel blocks to achieve more accurate global motion estimation in complex scenes.
[0136] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0137] Reference Figure 5 The diagram illustrates a structural block diagram of a vehicle-mounted camera video processing device provided in an embodiment of the present invention. The device includes:
[0138] The original video frame acquisition module 501 is used to acquire the original video frame corresponding to the video captured by the vehicle-mounted camera; the original video frame is a video frame with jitter caused by the movement of the vehicle-mounted camera.
[0139] The first clustering module 502 is used to cluster the first motion vector of the original video frame to determine the foreground region and background region in the original video frame; the background region reflects the motion of the vehicle-mounted camera.
[0140] The second clustering module 503 is used to detect feature points from the background region of the original video frame and cluster the second motion vectors of the feature points in the background region of adjacent original video frames to determine target feature points from the feature points.
[0141] The motion compensation module 504 is used to perform global motion estimation on the original video frame based on the target feature points, and to perform motion compensation on the original video frame based on the global motion estimation result.
[0142] In one embodiment of the present invention, the first clustering module 502 is used for:
[0143] Each of the original video frames is divided into superpixel blocks; each superpixel block has a corresponding class label and a superpixel centroid corresponding to the class label.
[0144] Determine the first distance between the centroids of superpixels with the same class label in adjacent original video frames;
[0145] When the first distance meets the preset distance, a first motion vector between adjacent superpixel blocks is determined;
[0146] The foreground and background regions are determined from the original video frame based on the first motion vector.
[0147] In one embodiment of the present invention, the first clustering module 502 is used for:
[0148] Two initial motion vectors are randomly selected from the first motion vectors;
[0149] Determine a second distance between each of the first motion vectors and the first initial motion vector;
[0150] Based on the second distance, the first motion vector is divided into a first motion vector cluster corresponding to the first initial motion vector;
[0151] Determine the first degree of difference of the first motion vector in each of the first motion vector clusters;
[0152] When the first difference exceeds the preset first difference, a first mean is determined based on the first motion vector in the first motion vector cluster, the first mean is used as the first initial motion vector, and the step of determining the second distance between each first motion vector and the first initial motion vector is returned.
[0153] When the first degree of difference does not exceed the preset first degree of difference, the magnitude of the first motion vector cluster is determined;
[0154] The region of the original video frame corresponding to the first motion vector cluster with the smallest magnitude of the first motion vector cluster is determined as the background region, and the region of the original video frame corresponding to the first motion vector cluster with the largest magnitude of the first motion vector cluster is determined as the foreground region.
[0155] In one embodiment of the present invention, the first clustering module 502 is used for:
[0156] Cluster centers are evenly placed in the original video frames; each cluster center has a corresponding label.
[0157] Determine a third distance between each cluster center and pixels within a specified range of the cluster center; the third distance includes at least color, texture, brightness, and spatial distance.
[0158] Assign the label corresponding to the cluster center that is closest to the pixel in the third distance to the pixel to the pixel, and merge the pixels with the same label into a pixel cluster;
[0159] A new cluster center is determined based on the cluster center corresponding to the pixel cluster;
[0160] Determine the center shift distance between the new cluster center and the existing cluster center;
[0161] When the center movement distance is greater than or equal to the preset center movement distance, the new cluster center is taken as the cluster center, and a third distance is returned between each cluster center and the pixels within a specified range of the cluster center.
[0162] When the center movement distance is less than the preset center movement distance, the pixel cluster is determined as a superpixel block.
[0163] In one embodiment of the present invention, the second clustering module 503 is used for:
[0164] Obtain the second motion vector of feature points in the background region of adjacent original video frames;
[0165] Randomly select multiple second initial motion vectors from the second motion vector;
[0166] Determine a fourth distance between each of the second motion vectors and the second initial motion vector;
[0167] Based on the fourth distance, the second motion vector is divided into the second motion vector cluster corresponding to the second initial motion vector;
[0168] Determine the second degree of difference of the second motion vectors in each of the second motion vector clusters;
[0169] When the second difference exceeds the preset second difference, a second mean is determined based on the second motion vector in the second motion vector cluster, the second mean is used as the second initial motion vector, and the step of determining the fourth distance between each second motion vector and the second initial motion vector is returned.
[0170] When the second degree of difference does not exceed the preset second degree of difference, the magnitude of the second motion vector cluster is determined;
[0171] Feature points corresponding to the second motion vector cluster whose magnitude is less than a preset magnitude threshold are identified as target feature points.
[0172] In one embodiment of the present invention, the motion compensation module 504 is used for:
[0173] The motion path of the original vehicle camera is obtained by performing global motion estimation on the target feature points matched between adjacent original video frames.
[0174] The motion path of the original vehicle-mounted camera is smoothed to obtain the motion path of the target vehicle-mounted camera.
[0175] Calculate the compensation transformation matrix based on the difference between the target vehicle camera's motion path and the original vehicle camera's motion path;
[0176] The original video frame is inversely mapped and resampled according to the compensation transformation matrix to perform motion compensation on the original video frame.
[0177] In one embodiment of the present invention, the original video frame acquisition module 501 is used for:
[0178] Acquire video from the vehicle's onboard camera;
[0179] The video captured by the vehicle-mounted camera is processed by frame extraction to obtain the original video frames corresponding to the video captured by the vehicle-mounted camera.
[0180] In one embodiment of the invention, the distance includes at least Euclidean distance.
[0181] In one embodiment of the present invention, the motion compensation module 504 is used for:
[0182] The original video frame is divided into multiple grids;
[0183] Based on the target feature points and the motion vectors corresponding to the target feature points, the mesh is processed using a multihomography estimation strategy to obtain an initialized mesh;
[0184] The original motion path of the target feature point is reconstructed based on the initial mesh to obtain the reconstructed motion path, and the motion optimization mesh is determined based on the reconstructed motion path and the residual between the reconstructed motion path and the original motion path.
[0185] The motion paths of the motion optimization mesh are integrated to obtain the original motion path of the vehicle camera.
[0186] In one embodiment of the present invention, the motion compensation module 504 is used for:
[0187] Optical flow between adjacent original video frames is extracted using a pre-set optical flow estimation model;
[0188] A pre-set keypoint detector model is used to detect feature points and target feature points from the background region; the target feature point information includes at least the position, orientation, and scale information of the feature points.
[0189] The motion vector of the target feature point is determined based on the optical flow and target feature point information;
[0190] Clustering is performed on the target feature points of each original video frame to divide the target feature points into corresponding planar clusters; wherein, when the number of target feature points in any planar cluster is lower than a preset threshold, the target feature points of the planar cluster are merged into another planar cluster;
[0191] The planar perceptual homography matrix is determined based on the motion vectors of the target feature points around the vertices of the grid and belonging to the same plane;
[0192] The initial mesh is obtained by transforming the mesh according to the planar homography matrix.
[0193] In this embodiment of the invention, the original video frames corresponding to the video captured by the vehicle-mounted camera are obtained. These original video frames are jittery due to the movement of the vehicle-mounted camera. The first motion vector of the original video frames is clustered to determine the foreground and background regions. The background region reflects the movement of the vehicle-mounted camera. Feature points are detected from the background region of the original video frames. Based on the second motion vector of the feature points in the background regions of adjacent original video frames, target feature points are clustered to determine the target feature points. Global motion estimation is performed on the original video frames based on the target feature points, and motion compensation is applied to the original video frames based on the global motion estimation results. This embodiment of the invention performs multiple clustering operations on the original video frames to accurately extract target feature points that reflect the movement of the vehicle-mounted camera, thereby performing motion compensation on the original video frames. This removes jittery video frames caused by the movement of the vehicle-mounted camera, making the motion-compensated original video frames stable. This improves the viewing experience of the vehicle driver for the video captured by the vehicle-mounted camera and enhances the reliability of computer analysis based on the video, ensuring the accuracy of the driver's predictions and personal safety when driving in complex environments.
[0194] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0195] This invention also provides an electronic device, such as... Figure 6 As shown, it includes a processor 601, a device interface 602, a memory 603, and a bus 604;
[0196] Memory 603 is used to store computer programs;
[0197] The processor 601 performs the above steps when executing the program stored in the memory 603.
[0198] The bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0199] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0200] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0201] The present invention also provides a storage medium that, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to perform the vehicle camera video processing method of the foregoing embodiments.
[0202] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0203] The algorithms and displays provided herein are not inherently related to any particular computer, virtual device, or other equipment. The structure required to construct such a device is readily apparent from the above description. Furthermore, this invention is not directed to any particular programming language. It should be understood that the contents of the invention described herein can be implemented using various programming languages, and the above description of specific languages is for the purpose of disclosing the best mode of implementation of the invention.
[0204] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0205] Similarly, it should be understood that, in order to simplify the invention and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, this method of disclosure should not be construed as reflecting an intention that the claimed invention requires more features than expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into this detailed description, wherein each claim itself is a separate embodiment of the invention.
[0206] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0207] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the sorting device according to the present invention. The present invention can also be implemented as a device or apparatus program for performing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0208] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0209] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0210] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0211] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0212] It should be noted that the various data-related processes in the embodiments of this application are carried out in compliance with the relevant data protection laws and policies of the country where the location is located, and with the authorization granted by the owner of the corresponding device.
Claims
1. A video processing method for an in-vehicle camera, characterized in that, The method includes: Obtain the original video frames corresponding to the video captured by the vehicle-mounted camera; the original video frames are video frames with jitter caused by the movement of the vehicle-mounted camera. Clustering is performed on the first motion vector of the original video frame to determine the foreground and background regions in the original video frame; the background region reflects the motion of the vehicle-mounted camera. Feature points are detected from the background region of the original video frame, and the second motion vectors of the feature points in the background region of adjacent original video frames are clustered to determine the target feature points from the feature points; Global motion estimation is performed on the original video frame based on the target feature points, and motion compensation is performed on the original video frame based on the global motion estimation results.
2. The method according to claim 1, characterized in that, The step of clustering the first motion vectors of the original video frame to determine the foreground and background regions in the original video frame includes: Each of the original video frames is divided into superpixel blocks; each superpixel block has a corresponding class label and a superpixel centroid corresponding to the class label. Determine the first distance between the centroids of superpixels with the same class label in adjacent original video frames; When the first distance meets the preset distance, a first motion vector between adjacent superpixel blocks is determined; The foreground and background regions are determined from the original video frame based on the first motion vector.
3. The method according to claim 2, characterized in that, Determining the foreground and background regions from the original video frame based on the first motion vector includes: Two initial motion vectors are randomly selected from the first motion vectors; Determine a second distance between each of the first motion vectors and the first initial motion vector; Based on the second distance, the first motion vector is divided into a first motion vector cluster corresponding to the first initial motion vector; Determine the first degree of difference of the first motion vector in each of the first motion vector clusters; When the first difference exceeds the preset first difference, a first mean is determined based on the first motion vector in the first motion vector cluster, the first mean is used as the first initial motion vector, and the step of determining the second distance between each first motion vector and the first initial motion vector is returned. When the first degree of difference does not exceed the preset first degree of difference, the magnitude of the first motion vector cluster is determined; The region of the original video frame corresponding to the first motion vector cluster with the smallest magnitude of the first motion vector cluster is determined as the background region, and the region of the original video frame corresponding to the first motion vector cluster with the largest magnitude of the first motion vector cluster is determined as the foreground region.
4. The method according to claim 2, characterized in that, The step of dividing each of the original video frames into superpixel blocks includes: Cluster centers are evenly placed in the original video frames; each cluster center has a corresponding label. Determine a third distance between each cluster center and pixels within a specified range of the cluster center; the third distance includes at least color, texture, brightness, and spatial distance. Assign the label corresponding to the cluster center that is closest to the pixel in the third distance to the pixel to the pixel, and merge the pixels with the same label into a pixel cluster; A new cluster center is determined based on the cluster center corresponding to the pixel cluster; Determine the center shift distance between the new cluster center and the existing cluster center; When the center movement distance is greater than or equal to the preset center movement distance, the new cluster center is taken as the cluster center, and a third distance is returned between each cluster center and the pixels within a specified range of the cluster center. When the center movement distance is less than the preset center movement distance, the pixel cluster is determined as a superpixel block.
5. The method according to claim 1, characterized in that, The step of clustering the second motion vectors of feature points in the background region of adjacent original video frames to determine target feature points from the feature points includes: Obtain the second motion vector of feature points in the background region of adjacent original video frames; Randomly select multiple second initial motion vectors from the second motion vector; Determine a fourth distance between each of the second motion vectors and the second initial motion vector; Based on the fourth distance, the second motion vector is divided into the second motion vector cluster corresponding to the second initial motion vector; Determine the second degree of difference of the second motion vectors in each of the second motion vector clusters; When the second difference exceeds the preset second difference, a second mean is determined based on the second motion vector in the second motion vector cluster, the second mean is used as the second initial motion vector, and the step of determining the fourth distance between each second motion vector and the second initial motion vector is returned. When the second degree of difference does not exceed the preset second degree of difference, the magnitude of the second motion vector cluster is determined; Feature points corresponding to the second motion vector cluster whose magnitude is less than a preset magnitude threshold are identified as target feature points.
6. The method according to claim 1, characterized in that, The step of performing global motion estimation on the original video frame based on the target feature points, and then performing motion compensation on the original video frame based on the global motion estimation result, includes: Global motion estimation of the original vehicle camera motion path is performed on the target feature points matched between adjacent original video frames; The motion path of the original vehicle-mounted camera is smoothed to obtain the motion path of the target vehicle-mounted camera. Calculate the compensation transformation matrix based on the difference between the target vehicle camera's motion path and the original vehicle camera's motion path; The original video frame is inversely mapped and resampled according to the compensation transformation matrix to perform motion compensation on the original video frame.
7. The method according to claim 1, characterized in that, The acquisition of the original video frames corresponding to the video captured by the vehicle-mounted camera includes: Acquire video from the vehicle's onboard camera; The video captured by the vehicle-mounted camera is processed by frame extraction to obtain the original video frames corresponding to the video captured by the vehicle-mounted camera.
8. The method according to any one of claims 1-7, characterized in that, The distance must include at least the Euclidean distance.
9. The method according to claim 4, characterized in that, The vehicle-mounted camera is an infrared vehicle-mounted camera, and the original video frames captured by the infrared vehicle-mounted camera are infrared video frames. After obtaining the original video frames corresponding to the video captured by the vehicle-mounted camera, the method further includes: Convert the original video frame into an original video frame in a specified color space; A five-dimensional feature vector for each pixel is determined from the original video frame converted to a specified color space. The five-dimensional feature vector includes color space component values and position coordinates. The five-dimensional feature vector is used to determine the third distance, the new cluster center, and the center movement distance.
10. The method according to claim 6, characterized in that, The step of performing global motion estimation of the target feature points matched between adjacent original video frames to determine the motion path of the original vehicle camera includes: The original video frame is divided into multiple grids; Based on the target feature points and the motion vectors corresponding to the target feature points, the mesh is processed using a multihomography estimation strategy to obtain an initialized mesh; The original motion path of the target feature point is reconstructed based on the initial mesh to obtain the reconstructed motion path, and the motion optimization mesh is determined based on the reconstructed motion path and the residual between the reconstructed motion path and the original motion path. The motion paths of the motion optimization mesh are integrated to obtain the original motion path of the vehicle camera.
11. The method according to claim 10, characterized in that, The step of processing the mesh using a multi-homography estimation strategy based on the target feature points and the motion vectors corresponding to the target features to obtain an initialized mesh includes: Optical flow between adjacent original video frames is extracted using a pre-set optical flow estimation model; A pre-set keypoint detector model is used to detect feature points and target feature points from the background region; the target feature point information includes at least the position, orientation, and scale information of the feature points. The motion vector of the target feature point is determined based on the optical flow and target feature point information; Clustering is performed on the target feature points of each original video frame to divide the target feature points into corresponding planar clusters; wherein, when the number of target feature points in any planar cluster is lower than a preset threshold, the target feature points of the planar cluster are merged into another planar cluster; The planar perceptual homography matrix is determined based on the motion vectors of the target feature points around the vertices of the grid and belonging to the same plane; The initial mesh is obtained by transforming the mesh according to the planar homography matrix.
12. A vehicle-mounted camera video processing device, characterized in that, The device includes: The original video frame acquisition module is used to acquire the original video frame corresponding to the video captured by the vehicle-mounted camera; the original video frame is a video frame with jitter caused by the movement of the vehicle-mounted camera. The first clustering module is used to cluster the first motion vector of the original video frame to determine the foreground region and background region in the original video frame; the background region reflects the motion of the vehicle-mounted camera. The second clustering module is used to detect feature points from the background region of the original video frame and cluster the second motion vectors of the feature points in the background region of adjacent original video frames to determine the target feature points from the feature points. The motion compensation module is used to perform global motion estimation on the original video frame based on the target feature points, and to perform motion compensation on the original video frame based on the global motion estimation result.
13. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the instructions to implement the vehicle camera video processing method as described in any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the mobile terminal, the mobile terminal is able to perform the vehicle camera video processing method as described in any one of claims 1 to 11.
15. A vehicle, characterized in that, The vehicle includes the electronic equipment as described in claim 13.