A method for image and video recognition, analysis and evaluation based on spatio-temporal comparison

Through a method based on space-time comparison, the backbone network combined with ResNet-101 and Transformer and the hierarchical perceptual attention module are used to extract and compare the spatiotemporal characteristics of image videos, solving the problem of difficult to capture subtle differences in videos in the prior art and improving the accuracy of image video recognition and analysis.

CN116682044BActive Publication Date: 2025-05-30FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310703244.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-14
Publication Date
2025-05-30
Estimated Expiration
2043-06-14

AI Technical Summary

Technical Problem

The prior art is difficult to capture subtle differences in videos under abnormal situations such as large number of similar image videos or low-light blur, resulting in insufficient accuracy of image video recognition analysis.

Method used

Using image video recognition analysis and evaluation methods based on space-time comparison, the backbone network combined with ResNet-101 and Transformer is constructed to extract the spatiotemporal features of paired image videos, and a hierarchical perceptual attention module and a differential encoder based on space-time comparison are designed to compare and differential encoding of space-time features.

Benefits of technology

Effectively extract the global spatio-temporal features and the differential features between video pairs in image videos, and improve the accuracy of image video recognition analysis and evaluation, especially in similar scenes and low-light conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116682044B_ABST
    Figure CN116682044B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for image and video recognition, analysis and evaluation based on spatio-temporal comparison. The method includes the following steps: Step S1, obtaining the image and video of the relevant application scenario, and annotating the target boxes, categories or quality scores included in the images to construct a data set; Step S2, constructing a backbone network combining ResNet-101 pre-trained based on a large data set and Transformer, and training according to the data set constructed in Step S1 to extract the spatio-temporal features of paired image and video; Step S3, designing a hierarchical perception attention module and a difference encoder based on spatio-temporal comparison to compare the spatio-temporal feature differences of the paired videos in Step S2; Step S4, performing recognition and analysis on the spatio-temporal features of a single image and video and the difference features between paired image and video, and outputting the final analysis and evaluation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of image and video processing, and computer vision, in particular to a method for identifying, analyzing and evaluating video images based on spatio-temporal contrast. Background Art

[0002] In recent years, with the rapid development of artificial intelligence and multimedia technologies, automatically identifying and analyzing video images by machines has great application value and academic value, and has attracted wide attention from researchers. With the iterative development of this type of technology and the continuous penetration of specific application scenarios into reality, the demand in video image-related fields such as medical image diagnosis and human motion analysis has further expanded. How to capture the subtle differences in videos to identify accurate and detailed information in a large number of similar images or abnormal scenarios such as low-light and blurred images has become one of the research focuses and challenges in image analysis.

[0003] Although great progress has been made in image and video recognition and analysis technologies, the existing functions are only limited to the analysis of a single video, often only focusing on the global information of the current video, and unable to learn more local features that are beneficial to fine recognition and analysis. Using the detail comparison between two similar videos can provide more levels of fine features and detail importance, and assist the machine in learning to distinguish the true differences between different videos. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for identifying, analyzing and evaluating video images based on spatio-temporal contrast, which can effectively detect and extract the global spatio-temporal features in video images, use spatio-temporal contrast learning between videos to distinguish fine differences, and improve the accuracy of image recognition, analysis and evaluation.

[0005] To achieve the above purpose, the technical solution of the present invention is: A method for identifying, analyzing and evaluating video images based on spatio-temporal contrast, comprising the following steps:

[0006] Step S1: Obtain video images of relevant application scenarios, and annotate the target boxes, categories or quality scores contained in the images to construct a data set;

[0007] Step S2: Construct a backbone network combining ResNet-101 pre-trained on a large data set and Transformer, and train according to the data set constructed in Step S1 to extract the spatio-temporal features of paired video images;

[0008] Step S3: Design a hierarchical perception attention module and a spatio-temporal contrast-based difference encoder to compare the spatio-temporal feature differences of the paired video images in Step S2;

[0009] Step S4: Identify and analyze the spatio-temporal features of a single-image video and the differential features between paired-image videos, and output the final analysis and evaluation results.

[0010] In an embodiment of the present invention, step S1 specifically includes the following steps:

[0011] Step S11: Obtain videos or consecutive pictures of relevant image application scenarios from the network to form a preliminary data set;

[0012] Step S12: Clean the preliminary data set, retain the existing data annotations, and supplement the annotation information of the target box, category, or quality score according to the requirements of the application scenario to complete the construction of the data set;

[0013] Step S13: Divide the constructed data set into a training set and a test set according to a predetermined ratio.

[0014] In an embodiment of the present invention, step S2 specifically includes the following steps:

[0015] Step S21: For a given input query video Randomly select a sample video ε from the training set to form a video pair where T, H, and W are the number of temporal frames, height, and width of the video, respectively;

[0016] Step S22: Pre-train the ResNet-101 network structure based on the large-scale image data set ImageNet in the scene to obtain a 2D convolutional network model with excellent spatial information extraction ability, and then combine it with Transformer to form a backbone network for feature extraction;

[0017] Step S23: Use the backbone network constructed in step S22 to extract features from the video pair in step S21. For a video, first use ResNet-101 to extract the 2D spatial information contained in each frame of the video to obtain the feature f T×C×H×w , C is the current channel dimension, and then perform 2D sine-cosine position encoding on f T×C×H×w and send it to the encoder of Transformer, and use Transformer to encode in the temporal T dimension to learn the temporal information in the video to obtain f′ T×C×H×w ; The above process is expressed as follows:

[0018]

[0019] where V is the processed video, and N, P, and E are ResNet-101, 2D sine-cosine position encoding, and the encoder of Transformer parameterized by Θ, φ, and respectively.

[0020] In an embodiment of the present invention, step S3 specifically includes the following steps:

[0021] Step S31: Use 2D global average pooling operation to convert f' obtained in step S2 T×C×H×w into f T×C , and finally obtain the global features containing the spatio-temporal information of the video;

[0022] Step S32: Design a hierarchical perception attention module to further enhance the global features in multiple dimensions. Specifically, for f T×C , first perform global attention in the T dimension to focus on the image information at all time points and extract the temporal relationship between the information. The specific operation is to perform a fully connected layer and a Sigmoid activation function on f T×C in the T dimension to obtain the global attention weight and then multiply it pixel by pixel with the original f T×C to obtain Next, the attention in the C dimension focuses on the importance of different channel information to extract the key image information at the channel level. The specific operation is to perform global max pooling and global average pooling on in the C dimension respectively and then add them to , send them into the MLP module and the Sigmoid activation function to obtain the channel attention weight and then multiply it pixel by pixel with the original to obtain Finally, perform attention in the T dimension again to extract the image information at the detailed level, so as to deeply understand the internal rules and information distribution expressed by the video. The specific operation is to perform global max pooling and global average pooling on in the T dimension respectively and then splice them with , pass through a fully connected layer and a Sigmoid activation function to obtain the temporal detail attention weight and then multiply it pixel by pixel with the original to obtain f' T×C ; The above process is expressed as follows:

[0023]

[0024]

[0025]

[0026]

[0027] Among them, and respectively represent the weight coefficients of the first to third levels of hierarchical perception attention, FC 1 and FC 3 are the fully connected layers of the first level and the third level respectively. MaxPool and AvgPool represent global max pooling and global average pooling respectively. Concat is the concatenation operation along the dimension. Sigmoid is the Sigmoid activation function. Then, through an MLP module, the potential embedding of the feature f′ T×C is obtained, and finally the deep feature of the input video pair is obtained d is the feature dimension;

[0028] Step S33: Design a spatio-temporal contrast-based difference encoder, which is composed of 3 network modules with multi-head cross-attention and MLP blocks. Input the into the difference encoder to learn the corresponding spatio-temporal relationship between the query video and the sample video ε, perform spatio-temporal contrast and focus on the consistent feature information and mine the difference information, and generate a new feature The specific formula is as follows:

[0029]

[0030] Among them, and f ε are used as the query and key-value pairs respectively. W Q , W K and W V are weight matrices, is the normalization factor, softmax is the normalization exponential function, and (·) T is the transpose operation.

[0031] In an embodiment of the present invention, the step S4 specifically includes the following steps:

[0032] Step S41: Perform recognition and analysis on the global feature of the input query video , and perform recognition and analysis on the difference feature obtained by spatio-temporal contrast between paired image videos. It is formalized as follows:

[0033]

[0034]

[0035] Among them, is the recognition and analysis network required for relevant application scenarios, and are the result features predicted after recognizing and analyzing the global feature and the difference feature of the video respectively;

[0036] Step S42: Using the true annotation information of the sample video ε as the reference information, further analyze the differential result features obtained in Step S41 to obtain the result features of the query video based on spatio-temporal comparison, which is formalized as follows: As follows:

[0037]

[0038] where F ε is the true annotation information of the sample video ε, is the information fusion operation, specifically element-wise addition;

[0039] Step S43: Fuse the results obtained from the recognition and analysis based on the global video features and differential features to obtain the final result of the query video as follows:

[0040]

[0041] where is the final image video recognition, analysis and evaluation result.

[0042] In an embodiment of the present invention, the recognition and analysis network includes a classifier and a regressor.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] 1. Aiming at the problem that traditional image video recognition and analysis methods only analyze the spatio-temporal features of a single video and are difficult to capture fine and important detail features, the present invention proposes an image video recognition, analysis and evaluation method based on spatio-temporal comparison, which is a semantic representation for fine spatio-temporal feature comparison and identification of differential information between video pairs in similar scenarios, and a method for learning to distinguish which information is more important in image recognition. The present invention can effectively extract key differential information from image videos, effectively use the reliable information of existing sample videos as a reference to compare the features of query videos, reduce the difficulty of identifying subtle differences and improve the analysis accuracy.

[0045] 2. The present invention innovatively proposes a hierarchical perception attention module. Aiming at the global spatio-temporal features concerned by traditional methods, the hierarchical perception attention module learns the importance of feature information at each time node and learns the information representation of different importance levels under each channel to further help aggregate temporal features into more fine-grained representative features. The present invention conducts progressive attention learning at three levels, enabling the model to effectively learn how to distinguish important information in image videos in various application scenarios.

[0046] 3. The method of the present invention combines traditional single video recognition and analysis with innovative spatio-temporal contrast analysis between video pairs, which can effectively analyze various difficult scenarios. Currently, many image videos have a large number of similar images with only very subtle differences, and there are also many low-light, dark, and blurred image videos. The present invention first analyzes the global features of a single video to learn the important scene information described by the video, captures the basic information distribution, and then uses the fine spatio-temporal contrast between video pairs to capture and distinguish subtle differences, and combines the basic information to analyze and learn the changes expressed by these differences, effectively identifying image videos in difficult scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a flowchart of the method implementation of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The technical solutions of the present invention will be specifically described below in conjunction with the accompanying drawings.

[0049] It should be noted that the following detailed description is illustrative and is intended to provide further description of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0050] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0051] As Figure 1 shown, the present invention provides a method for recognizing, analyzing, and evaluating image videos based on spatio-temporal contrast, including the following steps:

[0052] Step S1: Obtain the image videos of relevant application scenarios, and label the target boxes, categories, or quality scores contained in the images to construct a data set. Specifically, it includes the following steps:

[0053] Step S11: Obtain videos or coherent pictures of relevant image application scenarios from the network to initially form a data set;

[0054] Step S12: Clean the data set, retain the existing data labels, and supplement the annotation information such as target boxes, categories, or quality scores according to the requirements of the application scenario to complete the construction of the data set;

[0055] Step S13: Divide the constructed data set into a training set and a test set according to a certain ratio.

[0056] Step S2: Construct a backbone network combining ResNet-101 pre-trained on a large dataset and Transformer, and train it according to the dataset constructed in Step S1 to extract the spatio-temporal features of paired image videos. The specific steps are as follows:

[0057] Step S21: For a given input query video Randomly select a sample video ε from the training set to form a video pair where T, H, and W are the number of temporal frames, height, and width of the video, respectively;

[0058] Step S22: Pre-train the ResNet-101 network structure on the large-scale ImageNet image dataset with rich scenes to obtain a 2D convolutional network model with excellent spatial information extraction ability, and then combine it with Transformer to form a backbone network for feature extraction;

[0059] Step S23: Use the backbone network constructed in Step S22 to extract features from the video pair in Step S21. Taking one video as an example, first use the excellent spatial recognition ability of the ResNet-101 convolutional network to extract the 2D spatial information contained in each frame of the video to obtain the feature f T×C×H×w , C is the current channel dimension, and then perform 2D sine-cosine position encoding on f T×C×H×w and send it into the Transformer encoder to encode in the temporal T dimension using the powerful global information capture ability of Transformer to learn the temporal information in the video and obtain f′ T×C×H×w , and the above process is summarized as follows:

[0060]

[0061] where V is the processed video, and N, P, and E are the ResNet-101 convolutional network, 2D sine-cosine position encoding, and Transformer encoder parameterized by Θ, φ, and respectively.

[0062] Step S3: Design a hierarchical perception attention module and a spatio-temporal contrast-based difference encoder to compare the spatio-temporal feature differences of the paired videos in Step S2. The specific steps are as follows:

[0063] Step S31: Use 2D global average pooling operation to convert f′ obtained in Step S2 T×C×H×w into f T×C , and finally obtain the global feature containing the spatio-temporal information of the image video;

[0064] Step S32: Design a hierarchical perception attention module to further enhance the global features comprehensively in multiple dimensions. Specifically, for f T×C , first perform global attention in the T dimension to focus on the image information at all time points and extract the temporal relationship between the information. The specific operation is to perform a fully connected layer and a Sigmoid activation function on f T×C in the T dimension to obtain the global attention weight and then multiply it pixel by pixel with the original f T×C to obtain Next, the attention in the C dimension focuses on the importance of different channel information to extract the key image information at the channel level. The specific operation is to perform global max pooling and global average pooling on in the C dimension respectively and then add them to , send them into the MLP module and the Sigmoid activation function to obtain the channel attention weight and then multiply it pixel by pixel with the original to obtain Finally, we perform attention in the T dimension again to extract the image information at the detail level, so as to deeply understand the internal rules and information distribution expressed by the image video. The specific operation is to perform global max pooling and global average pooling on in the T dimension respectively and then concatenate them with , pass through a fully connected layer and the Sigmoid activation function to obtain the temporal detail attention weight and then multiply it pixel by pixel with the original to obtain f′ T×C . The above process is described as follows:

[0065]

[0066]

[0067]

[0068]

[0069] Among them, and respectively represent the first to third weight coefficients of the hierarchical perception attention. FC 1 and FC 3 are the fully connected layers of the first layer and the third layer respectively. MaxPool and AvgPool represent global max pooling and global average pooling respectively. Concat is the concatenation operation along the dimension. Sigmoid is the Sigmoid activation function. Then, through an MLP module, the potential embedding of the feature f′ T×C is further obtained, and finally the input video pair is obtained Deep features d is the feature dimension;

[0070] Step S33: Design a difference encoder based on spatio-temporal contrast. This encoder consists of 3 network modules with multi-head cross-attention and MLP blocks. Input into this difference encoder to learn the corresponding spatio-temporal relationship between the query video and the sample video ε, perform spatio-temporal contrast, focus on the consistent feature information and mine the difference information, and generate new features The specific formula is as follows:

[0071]

[0072] Among them, and f ε are used as the query and key-value pair respectively, and W Q , W K and W V are weight matrices, is the normalization factor, softmax is the normalization exponential function, and (·) T is the transpose operation.

[0073] Step S4: Identify and analyze the spatio-temporal features of the single-image video and the difference features between the paired-image videos, and output the final analysis and evaluation results. Specifically, it includes the following steps:

[0074] Step S41: Different from the traditional image video recognition and analysis method, the present invention not only performs recognition and analysis on the global features of the input query video , but also performs spatio-temporal contrast between the paired-image videos to obtain difference features and also performs recognition and analysis on this difference feature. The formalization is as follows:

[0075]

[0076]

[0077] Among them, is the recognition and analysis network required for the relevant application scenario (such as classifier, regressor, etc.), and are the predicted result features after recognizing and analyzing the global features and difference features of the video respectively;

[0078] Step S42: Use the true annotation information of the sample video ε as the reference information to further analyze the difference result features obtained in Step S41, and obtain the result features of the query video based on spatio-temporal contrast Formalize as follows:

[0079]

[0080] Among them, F ε is the true annotation information of the sample video ε, is an information fusion operation, specifically element-wise addition;

[0081] Step S43: Fuse the results obtained from the recognition and analysis based on the global features and differential features of the video to obtain the final result of the query video The final result is expressed as follows:

[0082]

[0083] Among them, That is, the final video recognition, analysis and evaluation result.

[0084] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention, when the functions and effects generated do not exceed the scope of the technical solutions of the present invention, shall fall within the protection scope of the present invention.

Claims

1. A method for image and video recognition analysis and evaluation based on spatio-temporal contrast, characterized in that, it includes the following steps: S1. Obtain the image and video of the relevant application scenario, and label the target box, category or quality score contained in the image to construct a data set; S2. Construct a backbone network combining ResNet-101 pre-trained on a large data set and Transformer, and train and extract the spatio-temporal features of paired image and video according to the data set constructed in step S1; S3. Design a hierarchical perception attention module and a difference encoder based on spatio-temporal contrast to compare the spatio-temporal feature differences of the paired image and video in S2; S4. Identify and analyze the spatio-temporal features of a single image and video and the difference features between paired image and video, and output the final analysis and evaluation results; S2 includes the following steps: S21. For a given input query video Randomly select sample videos ε from the training set to form video pairs T, H, and W are the number of temporal frames, height, and width of the video, respectively; S22. Pre-train the ResNet-101 network structure based on the image data set ImageNet under a large-scale scenario to obtain a 2D convolutional network model with excellent spatial information extraction ability, and then combine it with Transformer to form a backbone network for feature extraction; S23. Use the backbone network constructed by S22 to extract features from the video pair of S21. For a video, first use ResNet-101 to extract the 2D spatial information contained in each frame of the video to obtain the feature f T×C×H×W , where C is the current channel dimension. Then, after performing 2D sine-cosine positional encoding on f T×C×H×W , send it into the encoder of the Transformer, and use the Transformer to perform encoding in the temporal dimension T to learn the temporal information in the video and obtain f' T×C×H×W : V is the processed video, and N, P, and E are the ResNet-101 parameterized by Θ, the 2D sine-cosine positional encoding, and the encoder of the Transformer, respectively; and the encoder of the Transformer, respectively; S3 includes the following steps: S31. Use 2D global average pooling operation to convert f' obtained in step S2 T×C×H×W into f T×c , and finally obtain the global features containing the spatio-temporal information of the video with images S32. Design a hierarchical perception attention module to further enhance the global features comprehensively in multiple dimensions for f T×c , first perform global attention in the T dimension to focus on the image information at all time points and extract the temporal relationship between the information. Specifically, for f T×C , obtain the global attention weights through a fully connected layer and a Sigmoid activation function in the T dimension , and then multiply it pixel by pixel with the original f T×C to obtain . Next, the attention in the C dimension focuses on the importance of different channel information to extract the key image information at the channel level. Specifically, for , perform global max pooling and global average pooling in the C dimension respectively, and then add them to , send them into the MLP module and the Sigmoid activation function to obtain the channel attention weights , and then multiply it pixel by pixel with the original to obtain . Finally, perform attention in the T dimension again to extract the image information at the detailed level, so as to deeply understand the internal laws and information distribution expressed by the image video. Specifically, for , perform global max pooling and global average pooling in the T dimension respectively, and then concatenate them with , pass through a fully connected layer and a Sigmoid activation function to obtain the temporal detail attention weights , and then multiply it pixel by pixel with the original to obtain f' T×C .

2. A method for image and video recognition analysis and evaluation based on spatio-temporal contrast according to claim 1, characterized in that, the specific steps of S1 are as follows: S11. Obtain the video or coherent pictures of the relevant image application scenario from the network to form a preliminary data set; S12. Clean the preliminary data set, retain the existing data labels, and supplement the annotation information of the target box, category or quality score according to the requirements of the application scenario to complete the construction of the data set; S13. Divide the constructed data set into a training set and a test set according to a predetermined ratio.

3. A method for image and video recognition analysis and evaluation based on spatio-temporal contrast according to claim 1, characterized in that, the process of S32 is expressed as follows: Among them, and respectively represent the first to third weight coefficients of the hierarchical perception attention. FC 1 and FC 3 are the fully connected layers of the first layer and the third layer respectively. MaxPool and AvgPool represent global max pooling and global average pooling respectively. Concat is the concatenation operation along the dimension. Sigmoid is the Sigmoid activation function. Then, a feature f' T×C of the potential embedding is further obtained through an MLP module, and finally the deep feature of the input video pair d is the feature dimension; And after completing S32, S33 is also executed, and the specific content of S33 is as follows: S33. Design a differential encoder based on spatio-temporal contrast. The differential encoder consists of 3 network modules with multi-head cross-attention and MLP blocks, and input it into the differential encoder to learn the corresponding spatio-temporal relationship between the query video and the sample video ε, conduct spatio-temporal contrast, focus on the consistent feature information and mine the differential information, and generate new features The specific formula is as follows: Among them, and f ε are used as query and key-value pair respectively, and W Q , W K and W V are weight matrices, is a normalization factor, softmax is a normalization exponential function, and (·) T is a transpose operation.

4. A method for image and video recognition analysis and evaluation based on spatio-temporal contrast according to claim 1, characterized in that, the specific steps of S4 are as follows: S41. Identify and analyze the global features of the input query video and identify and analyze the differential features obtained by performing spatio-temporal comparison between paired image videos as follows: ​ Among them, is the recognition and analysis network required for the relevant application scenarios, and are the result features predicted after recognizing and analyzing the global features and difference features of the video, respectively; S42. Using the true annotation information of the sample video ε as the reference information, further analyze the difference result features obtained in step S41 to obtain the query video and the result features based on spatio-temporal comparison are formalized as follows: Among them, F ε is the true annotation information of the sample video ε, is an information fusion operation, specifically element-wise addition; S43. Integrate the results obtained from the recognition and analysis based on the global features and differential features of the video to obtain the final result of the query video, which is expressed as follows: The final result is as follows: Among them, that is, the final image video recognition analysis and evaluation results.

5. A method for image and video recognition analysis and evaluation based on spatio-temporal contrast according to claim 4, characterized in that, the recognition and analysis network includes a classifier or a regressor.

Citation Information

Patent Citations

  • Video human body behavior recognition method based on time sequence enhancement module

    CN112464835A

  • Self-supervised action recognition method and device based on hierarchical multi-view

    CN115147676A