An intelligent ROI detection method and system based on ultra-high-definition 8K video

By combining background subtraction, optical flow calculation, feature matching and similarity evaluation techniques, accurate tracking of moving targets in ultra-high-definition 8K videos and extraction of ROIs is achieved, which solves the problems of large calculation volume, slow processing speed and insufficient accuracy in the ultra-high-definition 8K video environment, and achieves efficient and accurate ROI detection effect.

CN119206193BActive Publication Date: 2025-05-20GUOCHUANG RUISHI ARTIFICIAL INTELLIGENCE TECHNOLOGY (SICHUAN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411701739.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-05-20
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

In ultra-high-definition 8K video environment, traditional video ROI detection methods are large in computing, slow in processing speed, and insufficient accuracy, making it difficult to effectively deal with complex prospects and backgrounds.

Method used

Combined with background subtraction, optical flow calculation, feature matching and similarity evaluation technology, accurate tracking of moving targets and ROI extraction in ultra-high-definition 8K videos are achieved. The specific steps include obtaining the pending frame, extracting the foreground area using background subtraction, calculating the optical flow to obtain motion vectors, combining feature matching to perform ROI matching, and using Kalman filtering tracking to update the ROI state.

Benefits of technology

It realizes accurate tracking of moving targets in ultra-high-definition 8K videos and efficient extraction of ROI, improves processing speed and accuracy, and can effectively deal with complex video environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206193B_ABST
    Figure CN119206193B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent ROI detection method and system based on ultra-high-definition 8K video, which relates to the field of image processing technology, including obtaining a current frame to be processed from a picture frame sequence of an ultra-high-definition 8K video; extracting a foreground area from the current frame to be processed and its preceding and following frames by using a background subtraction method; calculating the optical flow between the foreground area in the current frame to be processed and the foreground areas in the preceding and following frames, and obtaining a motion vector of each foreground area; performing ROI matching on the foreground area in the current frame to be processed and the tracking result of a preceding tracking picture frame based on the motion vector and in combination with feature matching and similarity evaluation, and obtaining an ROI matching result; when the ROI matching result determines that the current frame to be processed is a tracking picture frame, tracking is performed using the ROI to be tracked and its motion vector, determining the ROI of the current frame to be processed and updating the state, thereby realizing accurate tracking of moving targets in ultra-high-definition 8K video and extraction of ROI.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an intelligent ROI detection method and system based on ultra-high-definition 8K video. Background Art

[0002] With the rapid development of ultra-high-definition video technology, the 8K video resolution has become the top level of current video technology, providing users with an unprecedented visual experience. However, the massive data of ultra-high-definition 8K video also brings unprecedented challenges to video processing and analysis. In application scenarios such as video surveillance, motion analysis, and target tracking, how to quickly and accurately extract the region of interest (ROI) from ultra-high-definition 8K video has become an urgent problem to be solved.

[0003] Most traditional video ROI detection methods are based on low-resolution videos. In the ultra-high-definition 8K video environment, these methods face problems such as large computational complexity, slow processing speed, and insufficient accuracy. In addition, the foreground and background in ultra-high-definition 8K video are often more complex and variable, and traditional methods such as background subtraction and motion detection are difficult to effectively handle.

[0004] Therefore, it is necessary to provide an intelligent ROI detection method and system based on ultra-high-definition 8K video to solve the above technical problems. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides an intelligent ROI detection method and system based on ultra-high-definition 8K video, which realizes accurate tracking of moving targets and extraction of ROI in ultra-high-definition 8K video by combining technologies such as background subtraction method, optical flow calculation, feature matching, and similarity evaluation.

[0006] The present invention provides an intelligent ROI detection method based on ultra-high-definition 8K video, and the detection method includes the following steps:

[0007] Obtain the current frame to be processed from the picture frame sequence of the ultra-high-definition 8K video;

[0008] Use the background subtraction method to extract the foreground region from the current frame to be processed and its front and back frames;

[0009] Calculate the optical flow between the foreground regions in the current frame to be processed and the foreground regions in the front and back frames to obtain the motion vector of each foreground region;

[0010] Based on the motion vector, and combining feature matching and similarity evaluation, perform ROI matching on the foreground region in the current frame to be processed and the tracking result of the previous tracking picture frame, and obtain the ROI matching result, where the previous tracking picture frame represents the set of frames that have been processed before the current frame to be processed and whose ROI has been tracked;

[0011] When the current frame to be processed is determined as a tracking picture frame based on the ROI matching result, use the ROI to be tracked and its motion vector for tracking, determine the ROI of the current frame to be processed, and update the status.

[0012] Preferably, obtaining the current frame to be processed from the picture frame sequence of the ultra-high definition 8K video includes:

[0013] Decode the ultra-high definition 8K video stream to obtain a continuous sequence of picture frames;

[0014] In the picture frame sequence, select and extract the current frame to be processed according to a preset frame interval;

[0015] Preprocess the current frame to be processed, where the preprocessing includes color correction, noise removal, and resolution adjustment.

[0016] Preferably, using the background subtraction method to extract the foreground region from the current frame to be processed and its previous and next frames includes:

[0017] Construct a background model, where the background model is obtained based on the historical frame data in the picture frame sequence of the ultra-high definition 8K video;

[0018] Compare the current frame to be processed and its previous and next frames with the background model respectively. By calculating the difference measure between each pixel value and the pixel value at the corresponding position in the background model, and based on a preset threshold, identify and extract the region whose pixel value exceeds the threshold with the background model as the foreground region.

[0019] Preferably, calculating the optical flow between the foreground regions in the current frame to be processed and the foreground regions in the previous and next frames to obtain the motion vector of each foreground region includes:

[0020] Perform boundary segmentation on the foreground regions in the current frame to be processed and its previous and next frames;

[0021] Use the optical flow algorithm to calculate the optical flow field between each foreground region in the current frame to be processed after boundary segmentation and the corresponding foreground regions in the previous and next frames;

[0022] Extract the motion vector of each foreground region from the optical flow field, where the motion vector represents the displacement direction and speed of the foreground region from the previous frame to the current frame or from the current frame to the next frame.

[0023] Preferably, the optical flow algorithm is the Horn-Schunck algorithm.

[0024] Preferably, based on the motion vector, and in combination with feature matching and similarity evaluation, perform ROI matching on the foreground region in the current frame to be processed and the tracking result of the previous tracked picture frame, and obtain an ROI matching result, including:

[0025] Extract the feature descriptors of each foreground region in the current frame to be processed, and the feature descriptors of the ROI in the tracking result of the previous tracked picture frame;

[0026] Calculate the feature matching score between the feature descriptor of the foreground region in the current frame to be processed and the feature descriptor of the ROI in the tracking result of the previous tracked picture frame;

[0027] Determine the matching relationship according to the feature matching score, a preset similarity threshold, and the motion vector of the foreground region.

[0028] Preferably, the determination of the matching relationship includes:

[0029] If the feature matching score between the foreground region in the current frame to be processed and any ROI in the previous tracked picture frame is higher than the similarity threshold and the motion vectors are consistent, it is considered a successful match, and update the position and status of the ROI;

[0030] If the feature matching scores between the foreground regions in the current frame to be processed and the ROIs in the previous tracked picture frame are all lower than the similarity threshold, or the motion vectors are inconsistent, add this foreground region to the ROIs to be tracked;

[0031] If the feature matching scores between the ROIs in the previous tracked picture frame and the foreground regions in the current frame to be processed are all lower than the similarity threshold, or the motion vectors are inconsistent, delete this ROI from the ROIs to be tracked.

[0032] Preferably, when it is determined that the current frame to be processed is a tracked picture frame based on the ROI matching result, use the ROIs to be tracked and their motion vectors for tracking, determine the ROI of the current frame to be processed and update the status, including:

[0033] Use the existing ROIs to be tracked and their corresponding motion vectors, and adopt the Kalman filter tracking algorithm to track the foreground regions in the current frame to be processed;

[0034] According to the tracking result, determine at least one ROI as the ROI of the current frame to be processed, and update the ROI status, where the ROI status includes position, size, and motion direction.

[0035] The present invention also provides an intelligent ROI detection system based on an ultra-high-definition 8K video for executing the intelligent ROI detection method based on an ultra-high-definition 8K video, and the detection system includes:

[0036] A frame acquisition module, configured to acquire a currently to-be-processed frame from a sequence of picture frames of an ultra-high definition 8K video;

[0037] A foreground extraction module, configured to extract a foreground region from the currently to-be-processed frame and its front and back frames by using a background subtraction method;

[0038] An optical flow calculation module, configured to calculate an optical flow between the foreground region in the currently to-be-processed frame and the foreground regions in the front and back frames, and obtain a motion vector for each foreground region;

[0039] An ROI matching module, configured to perform ROI matching on the foreground region in the currently to-be-processed frame and the tracking result of a previous tracked picture frame based on the motion vector, in combination with feature matching and similarity evaluation, to obtain an ROI matching result, where the previous tracked picture frame represents a set of frames that have been processed before the currently to-be-processed frame and whose ROIs have been tracked;

[0040] A tracking and updating module, configured to, when the ROI matching result determines that the currently to-be-processed frame is a tracked picture frame, perform tracking by using the ROI to be tracked and its motion vector, determine the ROI of the currently to-be-processed frame, and update the state.

[0041] Compared with related technologies, an intelligent ROI detection method and system based on an ultra-high definition 8K video provided by the present invention have the following beneficial effects:

[0042] The present invention first obtains a continuous sequence of picture frames by decoding an ultra-high definition 8K video stream, selects and extracts a currently to-be-processed frame according to a preset frame interval, and then extracts a foreground region from the currently to-be-processed frame and its front and back frames by using a background subtraction method. Then, an optical flow algorithm is used to calculate the motion vector of the foreground region. On this basis, in combination with feature matching and similarity evaluation technologies, ROI matching is performed on the foreground region in the currently to-be-processed frame and the tracking result of a previous tracked picture frame. Finally, according to the ROI matching result, the Kalman filter tracking algorithm is used to track the foreground region in the currently to-be-processed frame, determine the ROI, and update the state. By the method of the present invention, accurate tracking of moving targets and extraction of ROIs in an ultra-high definition 8K video can be achieved. Description of the Drawings

[0043] Figure 1 It is a flowchart of an intelligent ROI detection method based on an ultra-high definition 8K video provided by the present invention;

[0044] Figure 2 It is a schematic diagram of the module structure of an intelligent ROI detection system based on an ultra-high definition 8K video provided by the present invention. Detailed Embodiments

[0045] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only the parts related to the present invention rather than all the structures are shown in the drawings. Furthermore, the embodiments in the present invention and the features in the embodiments can be combined with each other without conflict.

[0046] It should also be noted that for the sake of description, only the parts related to the present invention rather than all the content are shown in the drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0047] Embodiment 1

[0048] The present invention provides an intelligent ROI detection method based on ultra-high-definition 8K video. Referring to Figure 1 as shown, the detection method includes the following steps:

[0049] S1: Obtain the current frame to be processed from the picture frame sequence of the ultra-high-definition 8K video.

[0050] Specifically, step S1 includes the following steps:

[0051] S11: Decode the ultra-high-definition 8K video stream to obtain a continuous picture frame sequence.

[0052] In digital video processing, videos are usually stored in a compressed format to reduce the requirements for storage space and transmission bandwidth. However, before performing video analysis, these compressed video streams need to be decoded into the original picture frame sequence. For 8K videos, due to their extremely high resolution (7680x4320 pixels), the decoding process not only requires high efficiency but also needs to ensure that the image quality is not lost. In this step, the decoder reads the encoded information in the video file, such as advanced coding formats like H.265 / HEVC, and then restores it to the original image data through a decoding algorithm.

[0053] The beneficial effect of this process is that it can accurately and losslessly restore high-quality image frames, providing a solid foundation for subsequent image processing and analysis.

[0054] S12: In the picture frame sequence, select and extract the current frame to be processed according to a preset frame interval.

[0055] In this embodiment, due to the large amount of data brought by the high resolution of 8K video, directly processing each frame may lead to a great waste of computing resources. Therefore, a certain frame interval can be set to select representative frames for processing, which can not only reduce the computational burden but also maintain the coherence and representativeness of the video content. For example, one frame can be extracted every N frames as the frame to be processed, and N here can be flexibly set according to the content change rate and processing requirements of the video. This process helps to improve the processing efficiency and ensure that key information is not lost, thereby optimizing the performance of the entire system.

[0056] S13: Preprocess the current frame to be processed, where the preprocessing includes color correction, noise removal, and resolution adjustment.

[0057] In this embodiment, color correction is to correct possible color deviations to make the image colors more natural and real; noise removal is to filter various noises that may be introduced during the image acquisition process to improve the clarity of the image; and resolution adjustment is to enlarge or reduce the image when necessary to meet different processing requirements.

[0058] These preprocessing steps are crucial for improving the accuracy of ROI (Region of Interest) detection because they can improve the image quality and make subsequent operations such as background subtraction and optical flow calculation more effective. Through such preprocessing, the effect of final ROI detection can be significantly improved, ensuring that the target can be accurately identified and tracked even in complex video environments.

[0059] S2: Use the background subtraction method to extract the foreground region from the current frame to be processed and its adjacent frames.

[0060] Specifically, step S2 includes the following steps:

[0061] S21: Build a background model, where the background model is obtained based on the historical frame data in the picture frame sequence of the ultra-high-definition 8K video.

[0062] In this embodiment, the construction of the background model is to distinguish the static background and dynamic foreground objects in the video. In practical applications, the background model is usually established through statistical analysis of a series of consecutive frames at the beginning of the video. These initial frames should contain static scenes as much as possible to accurately capture the characteristics of the background.

[0063] The process of building a background model includes calculating the Gaussian mixture model statistical model of the average value of each pixel. In the application background of 8K videos, due to the extremely high image resolution, special attention needs to be paid to the calculation efficiency and memory occupancy when building the background model to avoid excessive resource consumption. The beneficial effect of this step is that it can provide a stable and reliable reference standard for subsequent foreground extraction, thereby improving the accuracy and robustness of ROI detection.

[0064] More specifically, the background model building process includes:

[0065] Select a series of consecutive frames at the beginning stage of the video. These frames should contain a static background as much as possible to ensure the accuracy of the background model.

[0066] For each pixel, calculate its pixel values in multiple historical frames. The Gaussian mixture model (GMM) can be used to model the distribution of each pixel, where GMM is a probability model used to represent the mixed distribution of data points. In the background model, the pixel value of each pixel can be represented by one or more Gaussian distributions.

[0067] Through maximum likelihood estimation (MLE), estimate the parameters of each Gaussian distribution, including the mean, variance, and weight. These parameters reflect the statistical characteristics of the pixel in different frames.

[0068] As the video plays, the background model needs to be continuously updated to adapt to environmental changes. This can be achieved through online learning methods. For example, gradually adjust the parameters of the Gaussian distribution to reflect the latest background information

[0069] S22: Compare the current frame to be processed and its adjacent frames with the background model respectively. By calculating the difference metric between each pixel value and the pixel value at the corresponding position in the background model, and based on a preset threshold, identify and extract the regions where the pixel values exceed the threshold compared to the background model as the foreground regions.

[0070] In this embodiment, this step is achieved by comparing the differences between the current frame and the background model pixel by pixel. If the difference of a certain pixel exceeds the preset threshold, then this pixel is considered to belong to the foreground region. This method can effectively separate moving objects.

[0071] For 8K videos, due to their high resolution which may result in a large number of foreground pixels, the influence of noise needs to be considered when setting the threshold to avoid false detection or missed detection caused by improper threshold setting.

[0072] More specifically, first, calculate the difference metric between each pixel in the current frame to be processed and its adjacent frames and the pixel at the corresponding position in the background model. In this application, the difference metric uses the Euclidean distance.

[0073] For each pixel, calculate the Euclidean distance between the current frame and the background model.

[0074] Next, compare the calculated difference metric with a preset threshold to determine whether the pixel belongs to the foreground region. If it is greater than the threshold, the pixel is considered to belong to the foreground region; otherwise, it is considered to belong to the background region.

[0075] Among them, the setting of the threshold can be adjusted according to the actual application scenario and video content. Usually, a suitable threshold can be determined through experiments.

[0076] After threshold processing, a binary image can be obtained, where the pixel value of the foreground region is 1 and the pixel value of the background region is 0. This binary image is the mask of the foreground region.

[0077] To further refine the foreground region, connected component analysis can be performed on the binary image to extract the connected foreground regions. This step can remove some small noise points and improve the accuracy of the foreground region.

[0078] Finally, overlay the extracted foreground region on the original image to generate an image with foreground markings. This step can be used for subsequent ROI detection and tracking

[0079] S3: Calculate the optical flow between the foreground regions in the current frame to be processed and the foreground regions in the previous and next frames, and obtain the motion vectors of each foreground region.

[0080] Specifically, step S3 includes the following steps:

[0081] S31: Perform boundary segmentation on the foreground regions in the current frame to be processed and its previous and next frames.

[0082] After the foreground regions are extracted, in order to calculate the optical flow more accurately, it is necessary to perform boundary segmentation on these foreground regions. The purpose of boundary segmentation is to clearly identify the edge contours of the foreground regions so that the subsequent optical flow calculation can more accurately capture the motion information of the objects. The specific implementation process is as follows:

[0083] In this embodiment, boundary segmentation can be implemented by the Canny edge detection algorithm, which can detect the intensity changes in the image to find the boundaries of the objects. For each foreground region, first apply the edge detection algorithm to generate a binary edge image. In this binary image, the boundaries of the foreground regions are marked as white pixels, and the other parts are black pixels. In this way, the contours of the foreground regions are clearly identified.

[0084] The beneficial effect of boundary segmentation is that it can significantly reduce the complexity of optical flow calculation because the optical flow calculation only needs to be performed on boundary pixels instead of the entire foreground area. In addition, boundary segmentation helps to improve the accuracy of optical flow calculation because boundary pixels usually contain more motion information.

[0085] S32: Use the optical flow algorithm to calculate the optical flow field between each foreground area in the current frame to be processed after boundary segmentation and its corresponding foreground area in the previous and next frames.

[0086] In this embodiment, optical flow refers to the displacement of pixel points in an image between adjacent frames, which can be used to describe the motion of an object. In this application, the Horn-Schunck algorithm is adopted, which is a classic global optical flow algorithm.

[0087] The specific implementation process is as follows:

[0088] First, initialize the optical flow field and , which represent the displacements of each pixel point in the horizontal and vertical directions respectively. Initially, the displacements of all pixel points can be set to 0.

[0089] For each pixel point , calculate the brightness change between the current frame and the previous frame and the next frame .

[0090]

[0091]

[0092] Calculate the gradients and of each pixel point in the horizontal and vertical directions.

[0093]

[0094]

[0095] Use the iterative formula of the Horn-Schunck algorithm to gradually update the optical flow field and .

[0096]

[0097]

[0098] Among them, and respectively represent at the At the second iteration, the pixel The optical flow components in the horizontal and vertical directions; and respectively represent at the second iteration, the pixel The optical flow components in the horizontal and vertical directions; respectively represent the horizontal gradient and vertical gradient of the image at the pixel location; is a smoothing parameter used to control the smoothness of the optical flow field, and the iteration process needs to be performed multiple times until the optical flow field converges.

[0099] S33: Extract the motion vector of each foreground region from the optical flow field, where the motion vector characterizes the displacement direction and speed of the foreground region from the previous frame to the current frame or from the current frame to the next frame.

[0100] In this embodiment, this step is to extract the motion vector of each foreground region from the optical flow field, where the motion vector characterizes the displacement direction and speed of the foreground region from the previous frame to the current frame or from the current frame to the next frame. After the optical flow field calculation is completed, it is necessary to extract the motion vector of each foreground region from the optical flow field for subsequent ROI matching and tracking.

[0101] In this embodiment, the specific implementation process is as follows:

[0102] First, according to the boundary segmentation result obtained in step S31, locate the boundary pixels of each foreground region.

[0103] For the boundary pixels of each foreground region, extract its corresponding horizontal displacement and vertical displacement from the optical flow field, and these displacement values together constitute the motion vector of the foreground region.

[0104] To reduce the influence of noise, the average motion vector of each foreground region can be calculated. The specific method is to average the displacement vectors of all boundary pixels within the foreground region.

[0105] The finally obtained average motion vector characterizes the overall motion direction and speed of the foreground region.

[0106] The motion vector can provide accurate motion information of the foreground region, which is very crucial for subsequent ROI matching and tracking. Through these motion vectors, the system can more accurately judge the motion state of the foreground region, thereby improving the accuracy of ROI detection and tracking.

[0107] S4: Based on the motion vectors, and in combination with feature matching and similarity evaluation, perform ROI matching on the foreground regions in the current frame to be processed and the tracking results of the previous tracking image frames, to obtain the ROI matching result, where the previous tracking image frames refer to the set of frames that have been processed before the current frame to be processed and whose ROIs have been tracked.

[0108] Specifically, step S4 includes the following steps:

[0109] S41: Extract the feature descriptors of each foreground region in the current frame to be processed, and the feature descriptors of the ROIs in the tracking results of the previous tracking image frames.

[0110] In image processing, a feature descriptor is a vector used to describe the features of a specific region in an image, and these features can be color, texture, shape, etc. In this step, the feature descriptors include but are not limited to SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), ORB (Oriented FAST and Rotated BRIEF), etc.

[0111] The specific implementation process is as follows:

[0112] First, use a feature point detection algorithm to detect feature points in the foreground regions of the current frame to be processed. These feature points are usually points with unique properties in the image, such as corner points, edge points, etc.

[0113] For each detected feature point, extract the feature descriptor of the local region around it. For example, use the SIFT algorithm to extract the gradient direction histogram of a 16x16 pixel region around the feature point to form a 128-dimensional feature vector.

[0114] Similarly, perform feature point detection and feature descriptor extraction on the already tracked ROIs in the previous tracking image frames to generate the corresponding set of feature descriptors.

[0115] S42: Calculate the feature matching scores between the feature descriptors of the foreground regions in the current frame to be processed and the feature descriptors of the ROIs in the tracking results of the previous tracking image frames.

[0116] In this embodiment, the feature matching score is used to measure the similarity between two feature descriptors, and the matching methods include but are not limited to nearest neighbor matching, ratio test, etc.

[0117] The specific implementation process is as follows:

[0118] Use a feature matching algorithm to match the feature descriptors in the current frame to be processed with the feature descriptors in the previous tracking image frames. During the matching process, calculate the distance (Euclidean distance) between each pair of feature descriptors and record the matching results.

[0119] To improve the reliability of matching, the ratio test method is used. For each feature point, find the two closest matching points in the previous frame. If the ratio of the distance between the closest matching point and the second-closest matching point is less than a certain threshold (such as 0.7), then the match is considered reliable.

[0120] According to the matching results, calculate the matching score between the foreground region in the current frame to be processed and the ROI in the previous tracked image frame. The matching score can be, but is not limited to, the number of matching feature points, the total distance of the matching feature points.

[0121] S43: Determine the matching relationship based on the feature matching score, a preset similarity threshold, and the motion vector of the foreground region.

[0122] Among them, the determination of the matching relationship includes the following situations:

[0123] If the feature matching score between the foreground region in the current frame to be processed and any ROI in the previous tracked image frame is higher than the similarity threshold and the motion vectors are consistent, then the match is considered successful, and update the position and status of the ROI.

[0124] If the feature matching scores between the foreground region in the current frame to be processed and each ROI in the previous tracked image frame are all lower than the similarity threshold, or the motion vectors are inconsistent, then add this foreground region to the ROI to be tracked.

[0125] If the feature matching scores between the ROI in the previous tracked image frame and each foreground region in the current frame to be processed are all lower than the similarity threshold, or the motion vectors are inconsistent, then delete this ROI from the ROI to be tracked.

[0126] This step can accurately determine the matching relationship of the foreground region by comprehensively considering the feature matching score and the motion vector, thereby improving the accuracy and robustness of ROI tracking.

[0127] S5: When the matching result of the ROI determines that the current frame to be processed is a tracked image frame, use the ROI to be tracked and its motion vector for tracking, determine the ROI of the current frame to be processed and update the status.

[0128] Specifically, step S5 includes the following steps:

[0129] S51: Use the existing ROI to be tracked and its corresponding motion vector, and adopt the Kalman filter tracking algorithm to track the foreground region in the current frame to be processed.

[0130] In this embodiment, Kalman filtering is a recursive prediction-correction algorithm that is widely used in the field of target tracking and performs particularly well in dealing with noise and uncertainty. In this embodiment, the Kalman filtering algorithm is used to predict and correct the position and state of the ROI to achieve continuous tracking of the foreground region.

[0131] The basic principle of Kalman filtering:

[0132] The Kalman filtering algorithm works through two main steps: prediction and update. The prediction step estimates the current state based on the previous state, and the update step corrects the predicted state using the observed data at the current time. Specifically, the Kalman filter maintains a state vector and a covariance matrix, which represent the state of the target and the uncertainty of the state estimate, respectively.

[0133] Among them, the prediction step includes:

[0134] State prediction: Predict the current state based on the previous state and the state transition matrix.

[0135] Covariance prediction: Predict the current covariance matrix based on the previous covariance matrix and the process noise covariance matrix.

[0136] The update step includes:

[0137] Calculate the Kalman gain: Calculate the Kalman gain based on the predicted covariance matrix and the observation noise covariance matrix.

[0138] State update: Update the current state based on the observed data at the current time and the Kalman gain.

[0139] Covariance update: Update the current covariance matrix based on the Kalman gain.

[0140] Through the Kalman filtering algorithm, the position and state of the ROI can be effectively predicted and corrected, thus achieving continuous tracking of the foreground region. The Kalman filter can handle noise and uncertainty, improving the robustness and accuracy of tracking. In addition, the recursive nature of the Kalman filter makes it suitable for real-time processing, enabling it to quickly respond to the movement changes of the target and ensuring the real-time and stability of tracking.

[0141] S52: According to the tracking result, determine at least one ROI as the ROI of the current frame to be processed, and update the ROI state, where the ROI state includes position, size, and movement direction.

[0142] In this embodiment, according to the tracking result of the Kalman filter, the ROI that best matches the current foreground region is selected. This can be achieved by calculating the similarity between the foreground region and the candidate ROI (including but not limited to the overlap rate, feature matching score, etc.) to select the optimal ROI.

[0143] After determining the optimal ROI, its status information is updated, including position, size, and movement direction. The position can be directly obtained from the result of the Kalman filter, while the size and movement direction can be calculated based on the bounding box of the foreground region and the optical flow information.

[0144] Through this step, the tracking result of the Kalman filter can be converted into specific ROI information, ensuring that each foreground region has a clear identification and status description. This not only facilitates subsequent ROI detection and tracking but also provides basic data support for other advanced analysis tasks (such as behavior recognition, anomaly detection, etc.). In addition, the updated ROI status can be used for the tracking initialization of the next frame, forming a closed-loop tracking system and improving the overall tracking continuity and accuracy.

[0145] Embodiment 2

[0146] The present invention also provides an intelligent ROI detection system based on an ultra-high-definition 8K video, for implementing an intelligent ROI detection method based on an ultra-high-definition 8K video. Refer to Figure 2 As shown, the detection system includes:

[0147] A frame acquisition module 100, configured to acquire the current frame to be processed from the picture frame sequence of the ultra-high-definition 8K video.

[0148] A foreground extraction module 200, configured to extract the foreground region from the current frame to be processed and its adjacent frames using the background subtraction method.

[0149] An optical flow calculation module 300, configured to calculate the optical flow between the foreground region in the current frame to be processed and the foreground regions in the adjacent frames, and obtain the motion vector of each foreground region.

[0150] An ROI matching module 400, configured to perform ROI matching on the foreground region in the current frame to be processed and the tracking result of the previous tracking picture frame based on the motion vector, combined with feature matching and similarity evaluation, to obtain the ROI matching result, where the previous tracking picture frame represents the set of frames that have been processed before the current frame to be processed and whose ROIs have been tracked.

[0151] A tracking and updating module 500, configured to perform tracking using the ROI to be tracked and its motion vector when the ROI matching result determines that the current frame to be processed is a tracking picture frame, and determine and update the ROI of the current frame to be processed.

[0152] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0153] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium, which includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.

[0154] It should also be noted that the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity, or device. Without further limitations, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity, or device including the element.

Claims

1. An intelligent ROI detection method based on ultra-high-definition 8K video, characterized in that: The detection method comprises the following steps: Get the current frame to be processed from the picture frame sequence of the ultra-high-definition 8K video; Extracting a foreground area from the current frame to be processed and its previous and next frames using a background subtraction method; Calculate the optical flow between the foreground area in the current frame to be processed and the foreground areas in the previous and next frames to obtain the motion vector of each foreground area; Based on the motion vector, and in combination with feature matching and similarity evaluation, ROI matching is performed on the foreground area in the current frame to be processed and the tracking result of the previous tracking picture frame to obtain a ROI matching result, wherein the previous tracking picture frame represents a set of frames that have been processed before the current frame to be processed and whose ROIs have been tracked; The roi matching includes: If the feature matching score between the foreground area in the current frame to be processed and any ROI in the previous tracking picture frame is higher than the similarity threshold, and the motion vectors are consistent, the match is considered successful, and the position and state of the ROI are updated; If the feature matching scores of the foreground area in the current frame to be processed and each ROI in the previous tracking picture frame are lower than the similarity threshold, or the motion vectors are inconsistent, then the foreground area is added to the ROI to be tracked; If the feature matching scores of the roi of the previous tracking picture frame and the foreground regions in the current frame to be processed are both lower than the similarity threshold, or the motion vectors are inconsistent, then the roi is deleted from the roi to be tracked; When the roi matching result determines that the current frame to be processed is a tracking picture frame, the roi to be tracked and its motion vector are used for tracking to determine the roi of the current frame to be processed and update the state.

2. According to claim 1, an intelligent ROI detection method based on ultra-high-definition 8K video is characterized in that: The step of obtaining a current frame to be processed from a picture frame sequence of an ultra-high-definition 8K video includes: Decode ultra-high-definition 8K video streams to obtain continuous image frame sequences; In the picture frame sequence, selecting and extracting the current frame to be processed according to a preset frame interval; The current frame to be processed is preprocessed, wherein the preprocessing includes color correction, noise removal and resolution adjustment.

3. According to claim 2, an intelligent ROI detection method based on ultra-high-definition 8K video is characterized in that: The extracting of the foreground area from the current frame to be processed and its previous and next frames by using the background subtraction method comprises: Constructing a background model, wherein the background model is obtained based on historical frame data in a picture frame sequence of the ultra-high-definition 8K video; The current frame to be processed and its previous and next frames are respectively compared with the background model, and the difference measurement between each pixel value and the pixel value at the corresponding position in the background model is calculated. Based on a preset threshold, the area with a pixel value exceeding the threshold value relative to the background model is identified and extracted as the foreground area.

4. According to claim 3, an intelligent ROI detection method based on ultra-high-definition 8K video is characterized in that: The step of calculating the optical flow between the foreground area in the current frame to be processed and the foreground areas in the previous and next frames to obtain the motion vector of each foreground area includes: Perform boundary segmentation on the foreground area in the current frame to be processed and its previous and next frames; The optical flow algorithm is used to calculate the optical flow field between each foreground area in the current frame to be processed after boundary segmentation and its corresponding foreground area in the previous and next frames; A motion vector of each foreground region is extracted from the optical flow field, wherein the motion vector represents a displacement direction and speed of the foreground region from a previous frame to a current frame or from a current frame to a next frame.

5. According to claim 4, an intelligent ROI detection method based on ultra-high-definition 8K video is characterized in that: The optical flow algorithm is the Horn-Schunck algorithm.

6. The intelligent ROI detection method based on ultra-high-definition 8K video according to claim 5 is characterized in that: The step of performing ROI matching on the foreground area in the current frame to be processed and the tracking result of the previous tracking picture frame based on the motion vector and combining feature matching and similarity evaluation to obtain the ROI matching result includes: Extracting feature descriptors of each foreground area in the current frame to be processed and feature descriptors of the roi in the tracking result of the previous tracking picture frame; Calculate the feature matching score between the feature descriptor of the foreground area in the current frame to be processed and the feature descriptor of the roi in the tracking result of the previous tracking picture frame; The matching relationship is determined according to the feature matching score, a preset similarity threshold, and a motion vector of the foreground area.

7. The intelligent ROI detection method based on ultra-high-definition 8K video according to claim 6 is characterized in that: When the ROI matching result determines that the current frame to be processed is a tracking picture frame, tracking is performed using the ROI to be tracked and its motion vector, determining the ROI of the current frame to be processed and updating the state, including: Using the existing ROI to be tracked and its corresponding motion vector, a Kalman filter tracking algorithm is used to track the foreground area in the current frame to be processed; According to the tracking result, at least one ROI is determined as the ROI of the current frame to be processed, and the ROI state is updated, wherein the ROI state includes position, size and motion direction.

8. An intelligent ROI detection system based on ultra-high-definition 8K video, used to execute an intelligent ROI detection method based on ultra-high-definition 8K video according to any one of claims 1 to 7, characterized in that: The detection system comprises: The frame acquisition module is used to obtain the current frame to be processed from the picture frame sequence of the ultra-high-definition 8K video; A foreground extraction module, used for extracting a foreground area from the current frame to be processed and its previous and next frames using a background subtraction method; An optical flow calculation module is used to calculate the optical flow between the foreground area in the current frame to be processed and the foreground areas in the previous and next frames to obtain the motion vector of each foreground area; A roi matching module is used to perform roi matching on the foreground area in the current frame to be processed and the tracking result of the previous tracking picture frame based on the motion vector and in combination with feature matching and similarity evaluation to obtain a roi matching result, wherein the previous tracking picture frame represents a set of frames that have been processed before the current frame to be processed and whose roi has been tracked; The tracking and updating module is used to track the ROI to be tracked and its motion vector when the ROI matching result determines that the current frame to be processed is a tracking picture frame, determine the ROI of the current frame to be processed and update the state.

Citation Information

Patent Citations

  • Stereoscopic vision comfort level evaluation method based on motion features inside area of interest

    CN103096122A

  • Head identification detection method and system in outdoor monitoring scene, terminal and medium

    CN113392726A