Method and Application for Extracting Target Object Information from Video Based on Super-Resolution

By applying a deep expansion network based on super-resolution convolutional robust principal component analysis in X-ray angiography video sequence, combining the super-resolution module and the circulating neural network layer, the problems of low vascular information extraction efficiency and serious noise interference in the prior art are solved, and efficient and accurate vascular information extraction effect is achieved.

CN114170076BActive Publication Date: 2025-06-24SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111272433.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-06-24
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

The prior art has problems such as low time efficiency and spatial efficiency when extracting vascular information from X-ray angiography video sequences, large impact on tissue and organ movement in the background layer, and serious interference of complex mixed noise, resulting in poor extraction effect.

Method used

The deep expansion network based on super resolution is adopted to analyze the deep expansion network, combining the super resolution module and the recurrent neural network layer, through the combination of the deep expansion network and the super resolution module, the efficient segmentation of the video sequence and the extraction of target object information is achieved.

Benefits of technology

It improves the real-time and accuracy of vascular information extraction, effectively reduces the influence of background vascular structures and complex mixed noise, and significantly improves the effect of small blood vessel extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114170076B_ABST
    Figure CN114170076B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and application for extracting target object information from a video based on super-resolution. The method includes the following steps: collecting a video sequence containing a target object; dividing the video sequence into sub-blocks and inputting them into a trained deep unfolding network model for solution, and splicing the outputs to obtain a prediction result of the target object; the deep unfolding network model is a convolutional robust principal component analysis deep unfolding network, which is constructed by combining a deep unfolding algorithm based on robust principal component analysis with a super-resolution module. Compared with the prior art, the present invention has the advantages of high real-time performance, interference removal, accurate detection, etc. When the above method is applied to an X-ray angiography video, it can effectively reduce the influence of background-like vascular structures and complex mixed noises, and significantly improve the extraction effect of small blood vessels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information extraction, and in particular to a video processing method, and more particularly to a method for extracting target object information from a video based on super-resolution and its application. Background Art

[0002] In the field of information, it is often necessary to extract target object information from a video sequence. For example, an X-ray angiography video sequence is a type of video sequence, and accurate blood vessel information is the object that technicians need to obtain. Due to the X-ray projection imaging mechanism, the images in this type of video sequence contain many structures other than the blood vessels through which the contrast agent flows, such as human tissues and organs like bones, lungs, and diaphragms. In addition, various types of mixed noise will inevitably be generated during the imaging process. These background structures and mixed noise interfere with the recognition of blood vessel information, thus affecting the further analysis of blood vessel information and accurate clinical diagnosis and treatment. Therefore, it is necessary to separate the background layer and the blood vessel layer in the video sequence to obtain a blood vessel layer video sequence in which it is easier to obtain blood vessel information.

[0003] Currently, robust principal component analysis is the algorithm with the best effect for extracting the blood vessel layer from an X-ray angiography video sequence [Jin, M., Li, R., Jiang, J. and Qin, B., 2017. Extracting contrast-filled vessels in X-ray angiography by graduated RPCA with motion coherency constraint. Pattern Recognition, 63, pp. 653 - 666.]. From the perspective of motion analysis, this algorithm decomposes the video sequence into the sum of a low-rank matrix and a sparse matrix, where the low-rank matrix represents the background layer in the video sequence with greater similarity and smaller motion changes, and the sparse matrix represents the target object layer in the video sequence with sparse distribution and larger motion changes.

[0004] Traditional robust principal component analysis algorithms have limitations in extracting blood vessels from X-ray angiography video sequences. Such algorithms require a large number of iterative calculations, resulting in low time efficiency and space efficiency, which limits their clinical applications. Secondly, the human tissues and organs existing in the background layer of X-ray images are not completely stationary, and the slight motions of these structures will have a greater impact on the algorithm results. At the same time, there are a large number of complex mixed noises in X-ray images, and these mixed noises will damage blood vessel information, especially the information of small blood vessel branches. Therefore, the interference of tissues and organs in the background layer and complex mixed noises makes it impossible for traditional robust principal component analysis algorithms to accurately separate the blood vessel layer and the background layer.

[0005] In addition, some image segmentation techniques are used in vascular segmentation work to obtain the vascular region part in the image. Common methods include image enhancement techniques, deformable models, vascular tracking, etc. These methods are usually based on the vascular morphological structure or image gray value. When using these methods, the vascular-like structures in the image background and the complex Gaussian-Poisson mixed noise in the image will cause great interference to the segmentation result, making it difficult to distinguish the foreground and background in some regions. At the same time, the segmentation results of these methods focus on extracting the vascular structure contour features, while ignoring the vascular gray information in the original image.

[0006] Therefore, the existing methods cannot quickly and accurately extract vascular information from the X-ray angiography video sequence, which in turn makes it difficult to carry out diagnostic items such as quantification and functional analysis based on the shape and gray value restoration of the angiographic blood vessels. Generally speaking, the defects of the existing vascular extraction algorithms are summarized as follows:

[0007] 1. The time efficiency and space efficiency of extracting blood vessels are relatively low;

[0008] 2. The extracted angiographic blood vessel image contains tissue and organ structures and noises in the background layer;

[0009] 3. In the extracted angiographic blood vessel image, the information of small blood vessel branches cannot be retained. Summary of the Invention

[0010] The purpose of the present invention is to provide a method and application for extracting target object information from a video based on super-resolution with high real-time performance and accurate detection, so as to overcome the defects of the above-mentioned existing technologies.

[0011] The purpose of the present invention can be achieved through the following technical solutions:

[0012] A method for extracting target object information from a video based on super-resolution, which is applied to a transmission imaging video. The method includes the following steps:

[0013] Collect a video sequence containing the target object;

[0014] Segment the video sequence into sub-blocks and input them into a trained deep unfolding network model for solution, and splice the outputs to obtain the prediction result of the target object;

[0015] The deep unfolding network model is a convolutional robust principal component analysis deep unfolding network, which is constructed based on the deep unfolding algorithm of robust principal component analysis combined with a super-resolution module.

[0016] Furthermore, the specific construction process of the deep unfolding network model is as follows:

[0017] Based on video characteristics, a robust principal component analysis model is constructed and transformed into a Lagrangian form model;

[0018] The Lagrangian form model is iteratively solved to obtain the calculation formulas for each motion layer of the video;

[0019] The calculation formulas for each motion layer of the video are deeply expanded to obtain multiple iterative layers;

[0020] Multiple iterative layers and a super-resolution module are combined into the deep expansion network model.

[0021] Furthermore, the constructed robust principal component analysis model is:

[0022] min||L|| * +λ||S||1s.t.D=L+S

[0023] Among them, the matrix D represents the data matrix of the original video sequence, each column vector of which is the original video image after vectorization of one frame. The matrix L represents the low-rank matrix, which is the data matrix of the background layer to be solved. The matrix S represents the sparse matrix, which is the data matrix of the foreground layer to be solved. ‖L‖ * represents the nuclear norm of the matrix L, ‖S‖1 represents the l1 norm of the matrix S, and λ is a regularization parameter used to adjust the proportion of the foreground layer components obtained by decomposition.

[0024] Furthermore, the Lagrangian form model is:

[0025]

[0026] Among them, H1 and H2 are the metric matrices of L and S respectively. Here, H1 = H2 = I, ‖S‖ 1,2 represents the l 1,2 norm of the matrix S, and λ1 and λ2 are the regularization parameters of L and S respectively.

[0027] Furthermore, the iterative solution is implemented using a linear inverse problem solving algorithm.

[0028] Specifically, the linear inverse problem solving algorithm includes the soft thresholding iteration algorithm, the fast soft thresholding iteration algorithm, the alternating direction multiplier method, etc.

[0029] Furthermore, each motion layer of the video includes an approximately stationary background layer and a moving target object layer.

[0030] Specifically, during the iterative solution of the Lagrangian form model using the soft thresholding iteration algorithm, the low-rank matrix L and the sparse matrix S are iteratively updated until convergence. In the (k + 1)-th iteration, L k+1 and S k+1It can be updated according to the following calculation formula:

[0031]

[0032]

[0033] Among them, is the singular value threshold operator, is the soft threshold operator, and L f is the Lipschitz constant.

[0034] Furthermore, the specific implementation of the depth unfolding is as follows: Replace the coefficient terms in the calculation formula of each motion layer with convolutional layers, and replace the multiplication operation with a convolutional operation.

[0035] Specifically, the coefficient matrix terms composed of H1 and H2 can be replaced by convolutional layers, and the multiplication operation can be replaced by a convolutional operation. The calculation of the k-th layer in the unfolded network is as follows:

[0036]

[0037]

[0038] where * represents the convolutional operator, is the convolutional layer, is the regularization parameter. Both the convolutional layer parameters and the regularization parameter are obtained during training.

[0039] Furthermore, the super-resolution module includes a sampling layer and a sub-block sparse feature selection network layer. The specific implementation of combining multiple iterative layers with the super-resolution module is as follows:

[0040] Embed the sampling layer at the beginning position of the iterative layer, and embed the sub-block sparse feature selection network layer at the end position of the iterative layer.

[0041] The sampling layer is a commonly used network layer in neural networks that plays a feature selection function, and its role is to eliminate redundant information and retain effective information. Specifically, the sampling layer includes, but is not limited to, average pooling layer, max pooling layer, overlapping pooling layer, spatial pyramid pooling layer, upsampling layer, etc.

[0042] Furthermore, the sub-block sparse feature selection network layer includes a residual network layer and a recurrent neural network layer.

[0043] Among them, the recurrent neural network layer is a type of network that takes sequences as input and has a memory function. Specifically, the recurrent neural network layer includes, but is not limited to, traditional recurrent neural network, bidirectional recurrent neural network, gated recurrent neural network, long short-term memory network, convolutional long short-term memory network, etc.

[0044] Furthermore, the traditional target object extraction algorithm includes an extraction method based on background completion.

[0045] Furthermore, the extraction method based on background completion includes:

[0046] Segment the original image to obtain a background layer image with the region where the target object is located removed;

[0047] Estimate the background gray value of the target object region through background completion to obtain an estimated background layer image;

[0048] Subtract the estimated background layer image from the original image to obtain a target object gray information image.

[0049] Furthermore, the deep unfolding network model is trained with label samples, and the label samples are weakly supervised label samples, which are obtained by using a traditional target object extraction algorithm or manual annotation.

[0050] The present invention also provides an application of the method for extracting target object information from a video based on super-resolution as described above in an X-ray angiography video.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] First, the traditional robust principal component analysis algorithm uses iterative solution, with a large number of iterations, which consumes a lot of time and limits its practical application. The present invention constructs a convolutional robust principal component analysis deep unfolding network, and for the first time proposes to combine robust principal component analysis with deep unfolding. Each layer of the network represents one iteration of the iterative algorithm. Usually, the deep unfolding network can also achieve good results when the number of network layers is much smaller than the number of iterations of the traditional algorithm, greatly improving the time efficiency of the deep unfolding network compared with the original iterative algorithm. Therefore, the use of the robust principal component analysis deep unfolding network enables the present invention to have high real-time performance and enables the present invention to be applied to clinical applications such as X-ray sequence blood vessel extraction.

[0053] Second, the convolutional robust principal component analysis deep unfolding network constructed in the present invention first proposes to combine the robust principal component analysis deep unfolding network with a super-resolution module. In transmission imaging images such as X-ray sequence images, the background part contains overlapping complex anatomical structures, such as human tissues and organs like bones, lungs, vertebrae, diaphragms, etc. Due to factors such as respiratory movement and human movement, some background structures also have a certain degree of movement. At the same time, some structures in the background have morphological features and gray levels similar to blood vessels. These factors have a great impact on the extraction effects of traditional robust principal component analysis methods and deep unfolding network methods based solely on robust principal component analysis. It is difficult to distinguish the blood vessel components from the background components in some areas, especially the small blood vessel components. The present invention embeds a super-resolution module in the network layer, and the super-resolution module includes a sampling layer and a sub-block sparse feature selection network layer. Among them, the sampling layer is located before the robust principal component analysis, and can screen the input features, retain the effective features, remove the useless features, and eliminate the interference of some background-like blood vessel structures. The sub-block sparse feature selection network layer can reduce the influence of complex mixed noise, extract and enhance the detailed features of the image, and improve the detection rate of small blood vessels.

[0054] Third, the sub-block sparse feature selection network layer of the present invention includes a residual network layer and a recurrent neural network layer, and first proposes to use the recurrent neural network to realize the selection of blood vessel features. The recurrent neural network has memory, can transmit the feature information between the front and back frames of the video sequence in the network, and screen the features to improve the detection rate of the current frame.

[0055] Fourth, the present invention obtains the prediction result of the target object by segmenting the video sequence into sub-blocks for solution and then splicing. Transmission imaging images such as X-ray sequence images have complex mixed Gaussian-Poisson noise, which is caused by quantum noise and thermal noise generated by electronic devices. Among them, Poisson noise is signal-related, and the strength of the noise has a large correlation with the strength of the local signal. General global noise reduction algorithms are not suitable for reducing Poisson noise. The present invention can effectively reduce the mixed noise pattern in the sub-block area based on the sub-block method for X-ray sequence images, and effectively remove the mixed Gaussian-Poisson noise. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a schematic flow chart of the present invention;

[0057] Figure 2 is a schematic diagram of the convolutional robust principal component analysis deep unfolding network constructed by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0058] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and provides a detailed implementation manner and a specific operation process. However, the protection scope of the present invention is not limited to the following embodiments.

[0059] The present invention provides a method for extracting target object information from a video based on super-resolution, which is applied to a transmission imaging video. During the video imaging process, the target object attenuates the imaging light. Refer to Figure 1 As shown, the method includes the following steps: collecting a video sequence containing the target object; segmenting the video sequence into sub-blocks and inputting them into a trained deep unfolding network model for solution, and splicing the outputs to obtain the prediction result of the target object; the deep unfolding network model is a convolutional robust principal component analysis deep unfolding network, which is constructed by combining a deep unfolding algorithm based on Robust Principal Component Analysis (RPCA) with a super-resolution module, and the deep unfolding network model is trained with weakly supervised label samples obtained by using a traditional target object extraction algorithm.

[0060] Specifically, the specific construction process of the deep unfolding network model is as follows: based on the video characteristics, a robust principal component analysis model is constructed and transformed into a Lagrangian form model; the Lagrangian form model is solved, and the solution methods include but are not limited to Iterative Shrinkage Thresholding Algorithm (ISTA), Fast Iterative Shrinkage Thresholding Algorithm (FISTA), and Alternating Direction Method of Multipliers (ADMM) to obtain the calculation formulas for each motion layer of the video; the calculation formulas for each motion layer of the video are deeply unfolded to obtain a plurality of iterative layers. The specific deep unfolding is as follows: replacing the coefficient terms in the calculation formulas for each motion layer with convolutional layers and replacing the multiplication operations with convolutional operations; combining the plurality of iterative layers with the super-resolution module to form the deep unfolding network model.

[0061] Specifically, the super-resolution module includes a sampling layer and a sub-block sparse feature selection network layer. The combination of multiple iterative layers and the super-resolution module is specifically as follows: the sampling layer is embedded at the starting position of the iterative layer, and the sub-block sparse feature selection network layer is embedded at the ending position of the iterative layer. The sub-block sparse feature selection network layer includes a residual network layer and a recurrent neural network layer. Among them, the sampling layer includes, but is not limited to, an average pooling layer, a max pooling layer, an overlapping pooling layer, an empty pyramid pooling layer, an upsampling layer, etc.; the recurrent neural network layer includes, but is not limited to, a traditional recurrent neural network, a bidirectional recurrent neural network, a gated recurrent neural network, a long short-term memory network, a convolutional long short-term memory network, etc.

[0062] As Figure 2 shown, a convolutional robust principal component analysis depth unfolding network constructed based on the above method includes a pooling layer (Pooling layer), a robust principal component analysis depth unfolding layer (RPCAunrolling layer), and a super-resolution layer (Super-resolution module, SR module), where the super-resolution layer includes a convolutional layer (Convolutionallayer), a residual module (Residual module), a CLSTM module (convolutional long short-term memory network), and PixelShuffle (pixel shuffling).

[0063] The traditional algorithm for extracting target objects is a background completion-based method, which specifically includes: segmenting the original image to obtain a background layer image with the area where the target object is located removed; estimating the background gray value of the target object area through background completion to obtain an estimated background layer image; and obtaining the target object gray information image by subtracting the estimated background layer image from the original image.

[0064] The above method can effectively improve the real-time performance and accuracy of target object information extraction by using the constructed convolutional robust principal component analysis depth unfolding network.

[0065] When the above method is implemented in the form of software functional units and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0066] Embodiment 1

[0067] This embodiment realizes the accurate extraction of blood vessels from an X-ray angiography image sequence based on the above method, including:

[0068] 101) Construct a robust principal component analysis model for the X-ray angiography image sequence, that is:

[0069] min||L|| * +λ||S||1 s.t. D = L + S

[0070] Among them, the matrix D represents the data matrix of the original video sequence, and each column vector of it is a frame of the original video image after vectorization. The matrix L represents the low-rank matrix, which is the data matrix of the background layer to be solved. The matrix S represents the sparse matrix, which is the data matrix of the foreground layer to be solved. ‖L‖ * represents the nuclear norm of the matrix L, ‖S‖1 represents the l1 norm of the matrix S, and λ is a regularization parameter used to adjust the proportion of the foreground layer components obtained by decomposition.

[0071] 102) Write the robust principal component analysis model (the mathematical model of the RPCA algorithm) in Lagrangian form:

[0072]

[0073] Among them, H1 and H2 are the metric matrices of L and S respectively. Here, H1 = H2 = I. ‖S‖ 1,2 represents the l 1,2 norm of the matrix S, and λ1 and λ2 are the regularization parameters of L and S respectively.

[0074] 103) Solve the Lagrangian form of the model using the soft-thresholding iterative algorithm. During the iteration process, the low-rank matrix L and the sparse matrix S will be updated until convergence. In the (k + 1)-th iteration, Lk+1 and S k+1 can be updated according to the following formula:

[0075]

[0076]

[0077] where is the singular value threshold operator, is the soft threshold operator, L f is the Lipschitz constant, and are the conjugate matrices of H1 and H2.

[0078] 104) Deeply unfold the solution of robust principal component analysis. Replace the coefficient matrix terms composed of H1 and H2 with convolutional layers, and replace the multiplication operation with a convolutional operation. Therefore, the calculation of the k-th layer in the unfolded network is as follows:

[0079]

[0080]

[0081] where * represents the convolution operator, is the convolutional layer, is the regularization parameter.

[0082] 105) Connect the unfolded iterative network layers to construct a deep unfolding network. In this embodiment, the number of constructed iterative layers is 4. The size of the convolutional kernels of the first two layers is 5, and the size of the convolutional kernels of the last two layers is 3. The regularization parameter of the low-rank component is 0.4, and the regularization parameter of the sparse component is 1.8.

[0083] 106) Embed an average pooling layer at the beginning of the iterative layer for downsampling the input.

[0084] 107) Embed a residual module and a recurrent neural network module after the robust principal component analysis module. In this embodiment, the recurrent neural network module uses a convolutional long short-term memory network.

[0085] 108) Use SVS-net to segment the original vascular image sequence to obtain a sequence of vascular region segmentation maps and a sequence of background region segmentation maps.

[0086] 109) Solve the sequence of background region segmentation maps using the t-TNN tensor completion model to obtain the background layer data tensor image.

[0087] 110) Divide the original image sequence data by the elements at the same positions in the background layer data tensor to obtain the vascular data tensor image.

[0088] 111) Train a robust principal component analysis deep unfolding network using the vascular data tensor image to obtain a trained model.

[0089] 112) Segment the image sequence of the blood vessels to be extracted into sub - blocks and input them into the trained model, and splice the outputs to obtain a complete output image. In this embodiment, the size of the segmented sub - blocks is 64×64×20 (64 in length and width, and 20 frames), and 50% of the adjacent sub - blocks overlap.

[0090] The overall implementation process of the accurate extraction of blood vessels from the above - mentioned X - ray angiography image sequence is as Figure 1 shown. The structure of the deep network iteration layer in this embodiment is as Figure 2 shown.

[0091] In this embodiment, 43 clinical angiography image sequences are taken as examples to further illustrate the above - mentioned method. Each angiography image sequence contains 30 - 140 frames of images, the resolution of each frame of image is 512×512, the actual size represented by each pixel is 0.3mm×0.3mm, and the bit - depth of the pixel is 8.

[0092] The process of extracting blood vessels from the above - mentioned X - ray angiography image sequence has the advantages of effectively reducing the influence of background - like vascular structures and complex mixed noises, and significantly improving the extraction effect of small blood vessels.

[0093] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. A method for extracting target object information from a video based on super-resolution, which is applied to a transmission imaging video, and is characterized in that, The method includes the following steps: Collect a video sequence containing a target object; Segment the video sequence into sub-blocks and input them into a trained deep unfolding network model for solution, and splice the outputs to obtain the prediction result of the target object; The deep unfolding network model is a convolutional robust principal component analysis deep unfolding network, which is constructed based on the deep unfolding algorithm of robust principal component analysis combined with a super-resolution module; The specific construction process of the deep unfolding network model is as follows: Based on the video characteristics, construct a robust principal component analysis model and transform it into a Lagrangian form model; Among them, matrix L represents a low-rank matrix, which is the data matrix of the background layer to be solved, matrix S represents a sparse matrix, which is the data matrix of the foreground layer to be solved, H1 and H2 are the metric matrices of L and S respectively, and ‖S‖ 1,2 represents the l 1,2 norm of matrix S, and λ1 and λ2 are the regularization parameters of L and S respectively; the Lagrangian form model is iteratively solved to obtain the calculation formula for each motion layer of the video; Deeply unfold the calculation formulas of each motion layer of the video to obtain multiple iterative layers; Combine multiple iterative layers with a super-resolution module to form the deep unfolding network model; The specific deep unfolding is as follows: Replace the coefficient terms in the calculation formulas of each motion layer with convolutional layers, that is, replace the coefficient matrix terms composed of H1 and H2 with convolutional layers, and replace the multiplication operation with a convolutional operation. The calculation of the k-th layer in the unfolding network is as follows: Among them, is the singular value threshold operator, is the soft threshold operator, and * represents the convolution operator, is the convolutional layer, and the matrix D represents the data matrix of the original video sequence; The super-resolution module includes a sampling layer and a sub-block sparse feature selection network layer. The combination of multiple iterative layers with the super-resolution module is specifically: Embed the sampling layer at the start position of the iterative layer and embed the sub-block sparse feature selection network layer at the end position of the iterative layer; The sub-block sparse feature selection network layer includes a residual network layer and a recurrent neural network layer.

2. The method for extracting target object information from a video based on super-resolution according to claim 1, wherein, The iterative solution is implemented by a linear inverse problem solution algorithm.

3. The method for extracting target object information from a video based on super-resolution according to claim 1, characterized in that Each motion layer of the video includes an approximately stationary background layer and a moving target object layer.

4. The method for extracting target object information from a video based on super-resolution according to claim 1, wherein The traditional target object extraction algorithm includes an extraction method based on background completion.

5. The method for extracting target object information from a video based on super-resolution according to claim 1, characterized in that The deep unfolding network model is trained with label samples, and the label samples are weakly supervised label samples, which are obtained by using the traditional target object extraction algorithm or manual annotation.

6. Application of a method for extracting target object information from a video based on super-resolution as claimed in any one of claims 1-5 in an X-ray angiography video.