Video change detection method and system based on double models

By combining the dual-model method of temporal model and spatial model, low-dimensional subspace projection and probability grid calculation are used to dynamically update the background estimation, solving the accuracy problem of video change detection in a jitter environment, real-time and low-complexity video change detection is achieved.

CN120298943APending Publication Date: 2025-07-11INTELLIGENT INTER CONNECTION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510280269.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing video change detection methods are difficult to achieve accurate and real-time change detection in complex environments such as sensor jitter, pixel noise and light changes, especially when the sudden jitter or the jitter amplitude increases significantly.

Method used

A video change detection method based on a dual model is adopted, combining the temporal model and spatial model, and by constructing a temporal model and spatial model, using low-dimensional subspace projection and probability grid calculation, dynamically update the background estimation and standard deviation, and directly predict the jitter impact of a single-frame gradient, normalized residual joint detection is performed.

Benefits of technology

Real-time detection of video changes under low complexity is achieved, the accuracy of detection is improved, natural dynamic noise and sudden jitter noise can be effectively suppressed, background gradients are tracked, and the accuracy of change detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298943A_ABST
    Figure CN120298943A_ABST
Patent Text Reader

Abstract

The invention discloses a video change detection method and system based on double models, and relates to the field of intelligent traffic, and the method comprises the steps: combining a time model and a space model, achieving change detection through the judgment of a minimum normalized residual threshold, achieving the complementation of the time and space models, and achieving the inhibition of natural dynamic noise through the time model. The spatial model suppresses sudden jitter noise. Meanwhile, the time domain correlation of historical frames is captured through low-dimensional subspace projection, background estimation and standard deviation are dynamically updated, subspace base vectors can be updated in real time with low complexity, and background gradient is tracked. Pixel intensity change caused by jitter is predicted based on single-frame gradient distribution, a spatial mean value and variance are calculated through a probability grid, and adaptive adjustment is performed in combination with an inter-frame scaling factor. Historical jitter data is not needed, the jitter influence is directly predicted based on the single-frame gradient, and the detection accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation, and particularly to a video change detection method and system based on a dual model. Background Art

[0002] In the fields of image processing and computer vision, change detection is an important research direction and is widely applied in fields such as security monitoring, traffic management, environmental monitoring, military reconnaissance, etc. The core task of change detection is to identify significant changes occurring in an image sequence, such as the appearance of new objects, the movement or disappearance of objects, etc. However, change detection in practical applications faces many challenges. Especially in complex environments where there are jitters in the sensor platform, pixel noise, illumination changes, etc., it has become a difficult problem to accurately and real-time detect changes.

[0003] Traditional background modeling methods usually assume that the image sequence is perfectly aligned, that is, there is no jitter or displacement between frames. However, in practical applications, a sensor platform, such as a camera, may be affected by factors such as mechanical vibration, wind force, temperature change, etc., resulting in jitters between image frames. Such jitters will introduce changes in pixel intensity. Especially in areas where there are strong gradients (such as edges, textures, etc.) in the scene, jitters will cause significant fluctuations in pixel intensity, thereby interfering with the accuracy of change detection.

[0004] To address the jitter problem, various methods have been proposed in the prior art. By aligning the current frame with the previous frame or the background model, the influence of jitters is reduced. However, frame alignment usually requires high-precision pixel-level registration. Especially in the case of high frame rates, the computational complexity is relatively high and it is difficult to meet the requirements of real-time processing. In addition, frame alignment methods have limited effects when dealing with sudden jitters or non-stationary jitters. Especially in the case where the jitter amplitude is large or the jitter distribution is unstable. Some other methods detect changes by establishing a statistical model of pixel intensity. For example, the method based on the Gaussian mixture model can handle the multi-modal distribution of pixel intensity and is applicable to change detection in complex backgrounds. However, these methods usually assume that the jitter is stationary and the jitter amplitude is small, and it is difficult to cope with sudden jitters or situations where the jitter amplitude increases significantly. There are also some methods that use subspace projection techniques to capture the main change patterns in the image sequence, especially the pixel intensity changes caused by jitters. By projecting the image frames into a low-dimensional subspace. However, these methods usually rely on a relatively long training sequence to estimate the subspace and have limited effects when dealing with sudden jitters or non-stationary jitters. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a video change detection method and system based on a dual model, which can solve the problems of relatively high limitations and low accuracy in the existing video change detection based on a dual model.

[0006] To achieve the above object, the present invention provides a video change detection method based on a dual model, and the method includes:

[0007] Constructing a time model based on historical video frame data and dynamically updating the time model;

[0008] Predicting pixel intensity changes according to the gradient distribution of the current single-frame data, and obtaining the spatial mean and variance through a preset probability grid;

[0009] Constructing a spatial model according to the pixel intensity changes, spatial mean and variance;

[0010] Performing normalized residual joint detection of video changes according to the time model and the spatial model.

[0011] Further, the step of constructing a time model based on historical video frame data and dynamically updating the time model includes:

[0012] According to the formula B TEMPORAL (t) = W(t - 1)W T (t - 1)X(t), W(t) = FAPI_Update(W(t - 1), X(t), β1),

[0013] Constructing a time model and dynamically updating the time model, where the initial variance of each pixel is R + 7 is the number of frames for initializing the model, R is the subspace dimension, [X(k, h; u) is the intensity value of the pixel (k, h) in the u-th frame, B TEMPORAL is the mean background estimate of the first R + 7 frames, each current frame is t, the pixel intensity value of the current frame is X(t), projected onto the subspace basis vector matrix W(t - 1) of the previous frame to obtain the background estimate, the projection residual is X(t) - B TEMPORAL (t), the time-domain normalized residual is Z TEMPORAL , the subspace basis vector matrix is W(t), β1 is the decay factor, and γ is the variance forgetting factor.

[0015] Further, the step of predicting pixel intensity changes according to the gradient distribution of the current single-frame data, obtaining the spatial mean and variance through a preset probability grid, and constructing a spatial model according to the pixel intensity changes, spatial mean and variance includes:

[0016] According to the formula dr(V3 - V1) + dc(V2 - V1) + (dr)(dc)(V1 + V4 - V2 - V3),

[0018] Construct a spatial model, where the probability of displacement within each pixel neighborhood is P(i, j), is the standard Gaussian distribution function, σ is the standard deviation of the jitter, the size of the probability grid is BOX, Y = 2.806σ, each pixel is (k, h), and the conditional expected value of each pixel under jitter displacement in the t-th frame is The conditional variance is V1, V2, V3, V4 are the four pixel values in the neighborhood of the pixel (k, h), dr and dc are the displacements in the row and column directions, the scaling factor for each frame is S(t), b SPATIAL (k, h;t) is the spatial mean estimation of the pixel (k, h), m(t) is the normalized residual mean of the strong gradient pixels, Ω is the set of strong gradient pixels, and NΩ is the number of pixels in the set of strong gradient pixels Ω.

[0019] Further, the step of performing normalized residual joint detection of video changes according to the time model and the spatial model includes:

[0020] According to the formula

[0021] Z MIN (k, h;t) = min{|Z TEMPORAL (k, h;t)|, |Z SPATIAL (k, h;t)|} > T1 for detection, where the temporal normalized residual is Z TEMPORAL , and the spatial normalized residual is Z SPATIAL , if Z MIN (k, h;t) > T1, then it is determined that the pixel (k, h) is a changed pixel, and T1 is the detection threshold.

[0022] Further, the subspace dimension R is configured to be 2.5 - 3.5, the decay factors β1 and β2 are configured to be 0.9 - 0.99, and the variance forgetting factors γ1 and γ2 are configured to be 0.9 - 1.

[0023] Further, the present invention provides a video change detection system based on a dual model, and the system includes:

[0024] A construction module for constructing a time model based on historical video frame data and dynamically updating the time model;

[0025] An acquisition module, configured to predict pixel intensity changes based on the gradient distribution of current single-frame data, and obtain the spatial mean and variance through a preset probability grid;

[0026] The construction module is further configured to construct a spatial model according to the pixel intensity change, spatial mean and variance;

[0027] A detection module, configured to perform a normalized residual joint detection on video changes according to the time model and the spatial model.

[0028] Further, the construction module is specifically configured to, according to the formula B TEMPORAL (t) = W(t - 1)W T (t - 1)X(t), W(t) = FAPI_Update(W(t - 1), X(t), β1), construct a time model and dynamically update the time model, where the initial variance of each pixel is R + 7 is the number of frames for initializing the model, R is the subspace dimension, [X(k, h; u) is the intensity value of the pixel (k, h) in the u-th frame, B TEMPORAL is the mean background estimate of the first R + 7 frames, each current frame is t, the pixel intensity value of the current frame is X(t), projected onto the subspace basis vector matrix W(t - 1) of the previous frame to obtain the background estimate, and the projection residual is X(t) - B TEMPORAL (t), the time-domain normalized residual is Z TEMPORAL , the subspace basis vector matrix is W(t), β1 is the decay factor, and γ is the variance forgetting factor.

[0029] Further, the construction module is specifically further configured to, according to the formula (dr)(dc)(V1 + V4 - V2 - V3),

[0031]

[0032]

[0033] construct a spatial model, where the probability of displacement within the neighborhood of each pixel is P(i, j), is the standard Gaussian distribution function, σ is the standard deviation of the jitter, the size of the probability grid is BOX, Y = 2.806σ, each pixel is (k, h), and the conditional expected value of each pixel in the t-th frame under the jitter displacement is the conditional variance is V1, V2, V3, and V4 are the pixel values of four pixels in the neighborhood of pixel (k, h), dr and dc are the displacements of rows and columns, the scaling factor for each frame is S(t), and b SPATIAL (k,h;t) Spatial mean estimation of pixel (k, h), m(t) is the normalized residual mean of strong gradient pixels, Ω is the set of strong gradient pixels, and NΩ is the number of pixels in the set of strong gradient pixels Ω.

[0034] Furthermore, the detection module is specifically configured to use the formula Z MIN (k,h;t) = min{|Z TEMPORAL (k,h;t)|, |Z SPATIAL (k,h;t)|} > T1 for detection, where the temporal normalized residual is Z TEMPORAL , and the spatial normalized residual is Z SPATIAL . If Z MIN (k,h;t) > T1, then it is determined that pixel (k, h) is a changing pixel, and T1 is the detection threshold.

[0035] Furthermore, the subspace dimension R is configured to be 2.5 - 3.5, the decay factors β1 and β2 are configured to be 0.9 - 0.99, and the variance forgetting factors γ1 and γ2 are configured to be 0.9 - 1.

[0036] A video change detection method and system based on a dual - model provided by the present invention combines a temporal model and a spatial model, realizes change detection through minimum normalized residual threshold determination, achieves complementarity between the spatio - temporal models, where the temporal model suppresses natural dynamic noise and the spatial model suppresses sudden jitter noise. At the same time, it uses low - dimensional subspace projection to capture the temporal correlation of historical frames, dynamically updates background estimation and standard deviation, can update subspace basis vectors in real - time with low complexity, and tracks background gradual changes. It predicts pixel intensity changes caused by jitter based on single - frame gradient distribution, calculates spatial mean and variance through a probability grid, and adaptively adjusts in combination with the inter - frame scaling factor. Without historical jitter data, it directly predicts the impact of jitter based on single - frame gradient, improving detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a flowchart of a video change detection method based on a dual - model provided by the present invention;

[0038] Figure 2 is a schematic diagram of a video change detection system based on a dual - model provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The following further describes the device structure and implementation manner of the present invention in detail through the drawings and embodiments.

[0040] The present invention provides a video change detection method based on a dual model, as Figure 1 shown, which specifically includes the following steps:

[0041] 101. Construct a time model based on historical video frame data and dynamically update the time model.

[0042] Specifically, according to the formula

[0043] B TEMPORAL (t) = W(t - 1)W T (t - 1)X(t), W(t) = FAPI_Update(W(t - 1, Xt, β1),

[0044] construct a time model and dynamically update the time model, where the initial variance of each pixel is R + 7 is the number of frames for initializing the model, R is the subspace dimension, [X(k, h; u) is the intensity value of the pixel (k, h) in the u-th frame, B TEMPORAL is the mean background estimate of the first R + 7 frames, each current frame is t, the pixel intensity value of the current frame is X(t), projected onto the subspace basis vector matrix W(t - 1) of the previous frame to obtain the background estimate, and the projection residual is X(t) - B TEMPORAL (t), the time-domain normalized residual is Z TEMPORAL , the subspace basis vector matrix is W(t), β1 is the decay factor, and γ is the variance forgetting factor.

[0046] For example, initialization: Initialize the subspace basis vector using the first R + 7 frame data. First, calculate the mean of the first 8 frames as the first basis vector V1, and then generate subsequent basis vectors V2, V3,..., V R through the Gram - Schmidt orthogonalization process. Initialize the time - domain variance estimate ε TEMPORAL , and calculate the initial variance of each pixel (k, h) based on the projection residuals of the first R + 7 frames: where: R + 7: the number of frames for initializing the model. V1, V2, V3,..., V R : the subspace basis vectors, representing the main change patterns of the background. ε TEMPORAL : the time - domain variance estimate, measuring the degree of fluctuation of the pixel intensity. B TEMPORAL (k, h; R + 7): the mean background estimate of the first R + 7 frames. U: the frame index variable, representing the frame number used to initialize the model. (k, h): the pixel coordinates, representing the specific position in the image.

[0047] Then perform adaptive subspace projection: for each frame t, project the current frame X(t) onto the subspace basis vector matrix W(t - 1) of the previous frame to obtain the background estimate B TEMPORAL (t). B TEMPORAL (t) = W(t - 1)W T (t - 1)X(t), calculate the projection residual X(t) - B TEMPORAL (t), and normalize it to obtain the time-domain normalized residual Z TEMPORAL : Use the FAPI algorithm to recursively update the subspace basis vector matrix W(t), and the update formula is: W(t) = FAPI_Update(W(t - 1, Xt, β1), where β1 is the decay factor that controls the weight of historical data.

[0048] Finally, perform time-domain variance update: dynamically update the time-domain variance ε of each pixel based on the projection residual TEMPORAL : where γ is the variance forgetting factor that controls the rate of variance update.

[0049] 102. Predict the pixel intensity change based on the gradient distribution of the current single-frame data, and obtain the spatial mean and variance through the preset probability grid.

[0050] Specifically, assume that the jitter follows a Gaussian distribution, and calculate the probability P(i, j) of the displacement within the neighborhood of each pixel: where is the standard Gaussian distribution function, and σ is the standard deviation of the jitter. Calculate the size BOX of the probability grid according to the jitter standard deviation σ to ensure that it covers 99% of the jitter displacements: Y = 2.806σ, and then take

[0051] 103. Construct a spatial model based on the pixel intensity change, spatial mean, and variance.

[0052] Specifically, according to the formula

[0053] dc(V2 - V1)+(dr)(dc)(V1 + V4 - V2 - V3),

[0054]

[0055]

[0056] Construct a spatial model, where the probability of the displacement within the neighborhood of each pixel is P(i, j), is the standard Gaussian distribution function, σ is the standard deviation of the jitter, the size of the probability grid is BOX, and Y = 2.806σ, Each pixel is (k, h), and the conditional expectation value of each pixel in the t-th frame under the dither displacement is The conditional variance is V1, V2, V3, V4 are the four pixel values in the neighborhood of the pixel (k, h), dr and dc are the displacements of the row and column, the scaling factor for each frame is S(t), b SPATIAL (k, h; t) Spatial mean estimation of the pixel (k, h), m(t) is the normalized residual mean of the strong gradient pixels, Ω is the set of strong gradient pixels, and NΩ is the number of pixels in the set of strong gradient pixels Ω.

[0057] For example, first perform spatial moment estimation: for each pixel (k, h) in the t-th frame, calculate its conditional expectation value under the dither displacement and the conditional variance :

[0058] where V1, V2, V3, V4 are the four pixel values in the neighborhood of the pixel (k, h), and dr and dc are the displacements of the row and column. Calculate the spatial mean and variance using the conditional expectation value and variance:

[0059] Then the dynamic scaling factor: calculate the scaling factor S(t) for each frame to compensate for the difference between the actual dither amplitude and the model assumption: where m(t) is the normalized residual mean of the strong gradient pixels, Ω is the set of strong gradient pixels, N Ω is the number of pixels in the set of strong gradient pixels Ω.

[0060] 104. Perform joint detection of the normalized residuals for video changes according to the time model and the space model.

[0061] Specifically, according to the formula

[0062] Z MIN (k, h; t) = min{|Z TEMPORAL (k, h; t)|, |Z SPATIAL (k, h; t)|} > T1 for detection, where the time-domain normalized residual is Z TEMPORAL , and the space-domain normalized residual is Z SPATIAL . If Z MIN (k, h; t) > T1, then it is determined that the pixel (k, h) is a changed pixel, and T1 is the detection threshold.

[0063] For example, first perform normalized residual calculation: calculate the time-domain normalized residual Z TEMPORAL : Calculate the space-domain normalized residual Z SPATIAL : Then perform joint decision: take the smaller of the absolute values of the time-domain and space-normalized residuals: Z MIN (k, h; t) = min{|Z TEMPORAL (k, h; t)|, |Z SPATIAL (k, h; t)|} > T1. If Z MIN (k, h; t) > T1, then determine that the pixel (k, h) is a changing pixel.

[0064] Furthermore, the above parameters can also be set and optimized. Specifically, the subspace dimension R is usually set to 3 to balance the computational complexity and the model expression ability. The decay factor β1 controls the background update rate and is usually set to 0.90 - 0.99; β2 controls the target suppression rate and is usually set to 0.99. The variance forgetting factor γ1 controls the variance update rate and is usually set to 0.99; γ2 is set to 1.0 to prevent the variance update of abnormal pixels. The jitter standard deviation σ is set according to the actual jitter amplitude and is usually conservatively estimated as the standard deviation of the expected maximum jitter. The thresholds T1, T2, and T3 are adjusted according to the noise environment and the target signal strength. T2 and T3 are usually set to 3 - 8, and T1 is set to 10 to suppress the false detection of low-variance pixels.

[0065] A video change detection method based on a dual model provided by an embodiment of the present invention combines a time model and a space model, and realizes change detection through the determination of the minimum normalized residual threshold, achieving the complementarity of the spatio-temporal model. The time model suppresses natural dynamic noise, and the space model suppresses sudden jitter noise. At the same time, it uses low-dimensional subspace projection to capture the time-domain correlation of historical frames, dynamically updates the background estimate and the standard deviation, can update the subspace basis vectors in real time with low complexity, and tracks the background gradual change. Predicts the pixel intensity change caused by jitter based on the single-frame gradient distribution, calculates the spatial mean and variance through the probability grid, and adaptively adjusts in combination with the inter-frame scaling factor. Without historical jitter data, directly predicts the jitter impact based on the single-frame gradient, improving the detection accuracy.

[0066] As Figure 1 a specific implementation manner of the method shown, an embodiment of the present invention provides a video change detection system based on a dual model. As Figure 2 shown, the system includes: a construction module 21, configured to construct a time model according to historical video frame data and dynamically update the time model;

[0067] an acquisition module 22, configured to predict the pixel intensity change according to the gradient distribution of the current single-frame data, and obtain the spatial mean and variance through a preset probability grid;

[0068] The construction module 21 is further configured to construct a space model according to the pixel intensity change, the spatial mean, and the variance;

[0069] A detection module 23, configured to perform a normalized residual joint detection on video changes according to the time model and the space model.

[0070] Further, the construction module 21 is specifically configured to, according to the formula B TEMPORAL (t) = W(t - 1)W T (t - 1)X(t), W(t) = FAPI_Update(W(t - 1), X(t), β1), construct a time model and dynamically update the time model, wherein the initial variance of each pixel is R + 7 is the number of frames for initializing the model, R is the subspace dimension, [X(k, h; u) is the intensity value of the pixel (k, h) in the u-th frame, B TEMPORAL is the mean background estimate of the first R + 7 frames, each current frame is t, the pixel intensity value of the current frame is X(t), projected onto the subspace basis vector matrix W(t - 1) of the previous frame to obtain the background estimate, and the projection residual is X(t) - B TEMPORAL (t), the time-domain normalized residual is Z TEMPORAL , the subspace basis vector matrix is W(t), β1 is the attenuation factor, and γ is the variance forgetting factor.

[0071] Further, the construction module 21 is specifically further configured to, according to the formula

[0072]

[0073]

[0074]

[0075] construct a space model, wherein the probability of displacement within the neighborhood of each pixel is P(i, j), is the standard Gaussian distribution function, σ is the standard deviation of the jitter, the size of the probability grid is BOX, Y = 2.806σ, each pixel is (k, h), and the conditional expected value of each pixel in the jitter displacement in the t-th frame is the conditional variance is V1, V2, V3, V4 are the four pixel values in the neighborhood of the pixel (k, h), dr and dc are the displacements of the row and the column, and the scaling factor of each frame is S(t), b SPATIAL(k, h;t) Spatial mean estimation of pixel (k, h), m(t) is the normalized residual mean of strong gradient pixels, Ω is the set of strong gradient pixels, and NΩ is the number of pixels in the set of strong gradient pixels Ω.

[0076] Furthermore, the detection module 23 is specifically configured to detect according to the formula where the time-domain normalized residual is Z TEMPORAL , and the spatial normalized residual is Z SPATIAL . If Z MIN (k, h;t)>T1, then it is determined that pixel (k, h) is a changing pixel, and T1 is the detection threshold.

[0077] Furthermore, the subspace dimension R is configured to be 2.5 - 3.5, the attenuation factors β1 and β2 are configured to be 0.9 - 0.99, and the variance forgetting factors γ1 and γ2 are configured to be 0.9 - 1.

[0078] A video change detection system based on a dual model provided by the present invention combines a time model and a space model, and realizes change detection through the determination of the minimum normalized residual threshold, achieving the complementarity of the spatio-temporal model. The time model suppresses natural dynamic noise, and the space model suppresses sudden jitter noise. At the same time, it uses low-dimensional subspace projection to capture the time-domain correlation of historical frames, dynamically updates the background estimation and standard deviation, can update the subspace basis vector in real time with low complexity, and tracks the background gradual change. It predicts the pixel intensity change caused by jitter based on the single-frame gradient distribution, calculates the spatial mean and variance through the probability grid, and adaptively adjusts in combination with the inter-frame scaling factor. Without historical jitter data, it directly predicts the jitter impact based on the single-frame gradient, improving the detection accuracy.

[0079] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the protection scope of the present disclosure. The appended method claims present the elements of various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy described.

[0080] In the above detailed description, various features are combined in a single embodiment to simplify the present disclosure. This disclosure method should not be construed as reflecting an intention that the embodiments of the claimed subject matter require more features than those clearly stated in each claim. On the contrary, as reflected in the appended claims, the present invention lies in a state with fewer features than all the features of the disclosed single embodiment. Therefore, the appended claims are hereby clearly incorporated into the detailed description, where each claim stands alone as a separate preferred embodiment of the present invention.

[0081] The above-described embodiments have been presented for the purpose of enabling any person skilled in the art to make or use the present invention. For those skilled in the art, various modifications to these embodiments will be apparent, and the general principles defined herein may be applied to other embodiments without departing from the spirit and scope of the present disclosure. Thus, the present disclosure is not limited to the embodiments given herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0082] The foregoing description includes examples of one or more embodiments. Of course, it is not possible to describe all possible combinations of components or methods for the purpose of describing the above embodiments, but those of ordinary skill in the art should recognize that the various embodiments can be further combined and arranged. Thus, the embodiments described herein are intended to cover all such changes, modifications, and variations that fall within the scope of the appended claims. In addition, with respect to the term "comprising" used in the specification or claims, this term is inclusive in a manner similar to the term "including" as interpreted when employed as a transitional word in a claim. Further, any use of the term "or" in the specification or claims is intended to mean "non-exclusive or".

[0083] Those skilled in the art will also appreciate that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention may be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the various illustrative components, units, and steps have been described generally in terms of their functionality. Whether such functionality is implemented by hardware or software depends upon the particular application and design constraints of the overall system. For each particular application, those skilled in the art may implement the described functionality in a variety of ways, but such implementation should not be construed as departing from the scope of the embodiments of the present invention.

[0084] In the embodiments of the present invention, the various illustrative logical blocks or units described can be implemented or operated with the described functions by a general-purpose processor, a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of the above designs. The general-purpose processor can be a microprocessor, and optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a digital signal processor core, or any other similar configuration.

[0085] The steps of the methods or algorithms described in the embodiments of the present invention can be directly embedded in hardware, software modules executed by the processor, or a combination of both. The software modules can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be disposed in an ASIC, and the ASIC can be disposed in a user terminal. Optionally, the processor and the storage medium can also be disposed in different components of the user terminal.

[0086] In one or more exemplary designs, the functions described in embodiments of the present invention can be implemented in hardware, software, firmware, or any combination of the three. If implemented in software, these functions can be stored on a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. A computer-readable medium includes a computer storage medium and a communication medium that facilitates transfer of a computer program from one place to another. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer. For example, such computer-readable media can include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store program code in the form of instructions or data structures and other forms readable by a general-purpose or special-purpose computer, or a general-purpose or special-purpose processor. In addition, any connection can be properly defined as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source via a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless means such as infrared, wireless, and microwave, it is also included in the defined computer-readable medium. The disks and discs include compact disks, laser disks, optical disks, DVDs, floppy disks, and Blu-ray disks. Disks usually reproduce data magnetically, while discs usually reproduce data optically by laser. The above combinations can also be included in a computer-readable medium.

[0087] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A video change detection method based on a dual model, characterized in that, The method includes: Constructing a time model based on historical video frame data and dynamically updating the time model; Predicting pixel intensity changes according to the gradient distribution of the current single-frame data, and obtaining the spatial mean and variance through a preset probability grid; Constructing a spatial model according to the pixel intensity changes, spatial mean and variance; Performing a normalized residual joint detection on video changes according to the time model and the spatial model.

2. The video change detection method based on a dual model according to claim 1, characterized in that, The step of constructing a time model based on historical video frame data and dynamically updating the time model includes: According to the formula BTEMPORALt = W(t - 1)WT(t - 1)X(t), W(t) = FAPI_Update(W(t - 1), X(t), β1), Construct a temporal model and dynamically update the temporal model, where the initial variance of each pixel is R + 7 is the number of frames of the initialized model, R is the subspace dimension, [X(k, h; u) is the intensity value of the pixel (k, h) in the u-th frame, B TEMPORAL is the mean background estimate of the first R + 7 frames, each current frame is t, the pixel intensity value of the current frame is X(t), projected onto the subspace basis vector matrix W(t - 1) of the previous frame to obtain the background estimate, and the projection residual is X(t) - B TEMPORAL (t), the temporal domain normalized residual is Z TEMPORAL , the subspace basis vector matrix is W(t), β1 is the decay factor, and γ is the variance forgetting factor.

3. A video change detection method based on a dual model according to claim 1 or 2, characterized in that The step of predicting pixel intensity changes according to the gradient distribution of the current single-frame data, obtaining the spatial mean and variance through a preset probability grid, and constructing a spatial model according to the pixel intensity changes, spatial mean and variance includes: According to the formula Construct a spatial model, where the probability of displacement within each pixel neighborhood is P(i,j), is the standard Gaussian distribution function, σ is the standard deviation of the jitter, the size of the probability grid is BOX, Y = 2.806σ, each pixel is (k,h), and the conditional expectation value of each pixel in the jitter displacement in the t-th frame is The conditional variance is V1, V2, V3, V4 are the four pixel values in the neighborhood of the pixel (k,h), dr and dc are the displacements of the row and column, the scaling factor for each frame is S(t), b SPATIAL (k,h;t) Spatial mean estimation of the pixel (k,h), m(t) is the normalized residual mean of the strong gradient pixels, Ω is the set of strong gradient pixels, and NΩ is the number of pixels in the set of strong gradient pixels Ω.

4. A dual-model-based video change detection method according to claim 3, characterized in that The step of performing a normalized residual joint detection on video changes according to the time model and the spatial model includes: According to the formula Z MIN (k,h;t)=min{|Z TEMPORAL (k,h;t)|,|Z SPATIAL (k,h;t)|}>T1 is used for detection, where the time-domain normalized residual is Z TEMPORAL , and the space-domain normalized residual is Z SPATIAL . If Z MIN (k,h;t)>T1, then the pixel (k,h) is determined to be a changing pixel, and T1 is the detection threshold.

5. A dual-model-based video change detection method according to claim 4, characterized in that The method further includes: The subspace dimension R is configured to be 2.5 - 3.5, the decay factors β1 and β2 are configured to be 0.9 - 0.99, and the variance forgetting factors γ1 and γ2 are configured to be 0.9 - 1.

6. A video change detection system based on a dual model, characterized in that, The system includes: A construction module, configured to construct a time model based on historical video frame data and dynamically update the time model; An acquisition module, configured to predict pixel intensity changes according to the gradient distribution of the current single-frame data, and obtain the spatial mean and variance through a preset probability grid; The construction module is further configured to construct a spatial model according to the pixel intensity changes, spatial mean and variance; A detection module, configured to perform a normalized residual joint detection on video changes according to the time model and the spatial model.

7. The video change detection system based on a dual model according to claim 6, characterized in that The building block is specifically used to build a time model and dynamically update the time model according to the formula B TEMPORAL (t) = W(t - 1)W T (t - 1)X(t), W(t) = FAPI_Update(W(t - 1, Xt, β1), where the initial variance of each pixel is R + 7 is the number of frames for initializing the model, R is the subspace dimension, [X(k, h; u) is the intensity value of the pixel (k, h) in the u-th frame, B TEMPORAL is the mean background estimate of the first R + 7 frames, the current frame is t, the pixel intensity value of the current frame is X(t), projected onto the subspace basis vector matrix W(t - 1) of the previous frame to obtain the background estimate, and the projection residual is X(t) - B TEMPORAL (t), the time-domain normalized residual is Z TEMPORAL , the subspace basis vector matrix is W(t), β1 is the decay factor, and γ is the variance forgetting factor.

8. The video change detection system based on a dual model according to claim 6 or 7, characterized in that The building block is specifically further configured to, according to the formula xk,h;t = V1 + drV3 - V1 + dc(V2 - V1) + (dr)(dc)(V1 + V4 - V2 - V3), construct a spatial model, where the probability of displacement within each pixel neighborhood is P(i,j), is the standard Gaussian distribution function, σ is the standard deviation of the jitter, the size of the probability grid is BOX, Y = 2.806σ, each pixel is (k,h), and the conditional expected value of each pixel in the jitter displacement at the t-th frame is The conditional variance is V1, V2, V3, V4 are the pixel values of four pixels in the neighborhood of the pixel (k,h), dr and dc are the displacements of the row and column, the scaling factor for each frame is S(t), b SPATIAL (k,h;t) is the spatial mean estimation of the pixel (k,h), m(t) is the normalized residual mean of the strong gradient pixels, Ω is the set of strong gradient pixels, and NΩ is the number of pixels in the set of strong gradient pixels Ω.

9. The video change detection system based on a dual model according to claim 8, characterized in that The detection module is specifically configured to perform detection according to the formula Z MIN (k, h; t) = min{|Z TEMPORAL (k, h; t)|, |Z SPATIAL (k, h; t)|} > T1, where the time-domain normalized residual is Z TEMPORAL , and the space-domain normalized residual is Z SPATIAL . If Z MIN (k, h; t) > T1, then the pixel (k, h) is determined to be a changed pixel, and T1 is the detection threshold.

10. A video change detection system based on a dual model according to claim 9, characterized in that, The subspace dimension R is configured to be 2.5 - 3.5, the decay factors β1 and β2 are configured to be 0.9 - 0.99, and the variance forgetting factors γ1 and γ2 are configured to be 0.9 - 1.

Citation Information

Patent Citations

  • Regional average value kernel density estimation-based moving target detecting method in dynamic scene

    CN101957997A

  • Image sequence change detection method based on belief propagation algorithm with time-space joint information

    CN105551014A