Video auditing method, electronic equipment, medium and product
By integrating video image features, audio features, and barrage emotional features, the problem of insufficient video review accuracy in existing technologies is solved, and more efficient recognition of abnormal video content is achieved.
Patent Information
- Application Number
- CN202510820012.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-19
AI Technical Summary
In existing technologies, video review relies solely on image and audio feature analysis, resulting in insufficient review accuracy and robustness, and inability to effectively identify illegal videos.
By extracting the image features, audio features and emotional features of the barrage of comments from the video to be reviewed, a fusion analysis is performed, and cross-modal association and temporal association are used to integrate visual, auditory and textual information for video review.
Improves the accuracy and reliability of video review and can more effectively identify abnormal content in videos.
Smart Images

Figure CN120676206A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a video review method, electronic equipment, medium, and product. Background Art
[0002] With the development of Internet technology, user-generated content (UGC) such as live broadcasts and short videos has enriched Internet content. However, illegal videos may also appear, which require video review. Among the related technologies, the image and audio features of the video are mainly analyzed based on network models to determine whether the video is illegal. However, this method only relies on images and audio, and the analysis modality is relatively limited, which reduces the accuracy of video review. Summary of the Invention
[0003] The present disclosure provides a video review method, electronic equipment, medium and product.
[0004] In a first aspect, an embodiment of the present disclosure provides a video review method, the method comprising:
[0005] Determine the emotional characteristics of the barrage content of the video to be reviewed;
[0006] Obtaining fusion features based on the image features, audio features, and emotional features of the video to be reviewed;
[0007] According to the fusion feature, a first video review result is obtained to determine whether there is any abnormal content in the video to be reviewed. In a second aspect, an embodiment of the present disclosure provides an electronic device comprising a memory and a processor; the memory stores a computer program that can be executed by the processor, and when the computer program is executed by the processor, the above-mentioned video review method is implemented.
[0008] In a third aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned video review method when executed by a processor.
[0009] In a fourth aspect, an embodiment of the present disclosure provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned video review method.
[0010] In the disclosed embodiment, the image features and audio features of the video to be reviewed are extracted, and the barrage emotional features of the barrage content corresponding to the video to be reviewed are determined, and then the image features, audio features and barrage emotional features are fused to obtain fused features. Based on the fused features, the first video review result is obtained. In this way, multiple modal features such as video frames, audio and barrage text information can be fused to review the video, and the barrage emotional features can be temporally correlated and fused with the image features and audio features. Cross-modal correlation analysis can be performed, which can integrate visual, auditory, text and real-time barrage emotional analysis to improve the accuracy and reliability of video review. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In the accompanying drawings of the embodiments of the present disclosure:
[0012] Figure 1 A flowchart of a video review method provided in an embodiment of the present disclosure;
[0013] Figure 2 A flow chart of spatiotemporal joint anomaly location provided by an embodiment of the present disclosure;
[0014] Figure 3 An architectural diagram of a video review system provided in an embodiment of the present disclosure;
[0015] Figure 4 A block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0016] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0017] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, but the illustrated embodiments may be embodied in different forms, and the present disclosure should not be construed as limited to the embodiments set forth below. Rather, these embodiments are provided so that the present disclosure will be thorough and complete and will fully understand the scope of the present disclosure to those skilled in the art.
[0018] The accompanying drawings of the embodiments of the present disclosure are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the detailed embodiments, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent to those skilled in the art by describing the detailed embodiments with reference to the accompanying drawings.
[0019] The present disclosure may be described with reference to plan views and / or cross-sectional views by way of ideal schematic views of the present disclosure. Therefore, the exemplary illustrations may be modified according to manufacturing techniques and / or tolerances.
[0020] In the absence of conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.
[0021] The terms used in this disclosure are only used to describe specific embodiments and are not intended to limit the disclosure. As used in this disclosure, the term "and / or" includes any and all combinations of one or more related enumerated items. As used in this disclosure, the singular forms "a" and "the" are also intended to include plural forms, unless the context clearly indicates otherwise. As used in this disclosure, the terms "comprising" and "made of" specify the presence of the features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof.
[0022] Unless otherwise defined, all terms (including technical and scientific terms) used in this disclosure have the same meanings as those commonly understood by those skilled in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined in this disclosure.
[0023] The present disclosure is not limited to the embodiments shown in the drawings, but includes modifications of the configurations formed based on the manufacturing process. Therefore, the regions illustrated in the drawings have schematic properties, and the shapes of the regions shown in the drawings illustrate the specific shapes of the regions of the elements, but are not intended to be limiting.
[0024] To facilitate understanding, several concepts involved in this disclosure are briefly described below.
[0025] Optical flow map: In computer vision, optical flow maps can reflect pixel-level motion information between video frames. By analyzing the changes in pixels in consecutive video frames, motion information can be inferred. Usually, an optical flow map can be obtained by calculating the vector optical flow between every two adjacent video frames. Optical flow is a vector field that represents the motion of pixels in the video and is used to capture motion information.
[0026] Heatmap: Heatmaps can display the distribution and density of data through different color intensities, and can show the hot spots or significant areas that users pay attention to in the video.
[0027] Frame rate: Indicates the number of image frames displayed per second. The higher the frame rate, the smoother the video picture.
[0028] Long Short-Term Memory (LSTM): A recurrent neural network with memory units that can handle dependencies in time series data.
[0029] Spatio-Temporal Multi-scale Convolutional Neural Network (ST-MCNN): can be used to capture motion saliency features in videos.
[0030] Salient area: The area in a video that is most noticeable or important, usually related to factors such as motion and contrast.
[0031] Conditional Diffusion Model: A generative model that can restore the data distribution through gradual denoising and generate new video frames under given conditions (such as initial images or semantic information).
[0032] With the development of Internet technology, illegal videos may appear and need to be reviewed. Among the related technologies, the image and audio features of the video are mainly analyzed based on the network model to determine whether the video is illegal. This method only relies on images and audio, and the analysis modality is relatively limited, which reduces the accuracy and robustness of video review.
[0033] In the embodiment of the present disclosure, for a video to be reviewed, image features and audio features of the video to be reviewed are extracted, and the emotional features of the barrage corresponding to the barrage content are determined, and then the image features, audio features and barrage emotional features are fused to obtain fused features. Based on the fused features, a first video review result is obtained. In this way, multiple modal features such as video frames, audio and barrage text information can be fused to review the video, and the barrage text information is not simply input as static text. Instead, the barrage content is subjected to emotional analysis, and real-time barrage emotional features are integrated. Moreover, the barrage emotional features are temporally correlated and fused with image features and audio features. Cross-modal correlation analysis can integrate visual, auditory, textual and real-time barrage emotional analysis to solve the cross-modal semantic gap problem. The unstructured text characteristics of the barrage content are fused with video images and audio to construct a multimodal emotional map, thereby assisting in correcting misjudgments of visual or auditory modalities to improve accuracy. In addition, the behavioral characteristics of the barrage content itself have strong temporal correlation. The temporal consistency of the barrage content and the video and audio can be utilized to improve the accuracy of abnormal content positioning and improve the accuracy of video review.
[0034] It should be noted that the present disclosure is applicable to any scenario that requires efficient monitoring of illegal video content, such as online short video platforms, live interactive platforms, security monitoring systems, etc., and is not limited to this. In these scenarios, users generate video content, which will be accompanied by barrage comments, music / audio and other information.
[0035] See Figure 1 As shown, an embodiment of the present disclosure provides a flowchart of a video review method, the method comprising:
[0036] S101: Determine the emotional characteristics of the barrage content based on the barrage content of the video to be reviewed.
[0037] In the embodiment of the present disclosure, taking the live broadcast scenario as an example, the live video stream within a certain time period can be obtained in real time as the video to be reviewed, and the barrage content of the video to be reviewed within the time period can be obtained synchronously, and the video to be reviewed can be detected for abnormalities to determine whether there is any abnormal and illegal content.
[0038] For example, the bidirectional encoder representation from transformers (BERT) model can be pre-trained, and the barrage content can be semantically analyzed based on the BERT model to obtain the emotional features of the barrage.
[0039] S102: Obtain fusion features based on the image features, audio features, and barrage emotional features of the video to be reviewed.
[0040] For example, obtain the video to be reviewed (such as the video to be reviewed is H.264 encoded and has a resolution ≥ 1080p) and the barrage text stream, where the barrage text stream can be in JSON format and contain a timestamp and barrage content. Then, the video to be reviewed can be decoded, and the video frames and video frame image features can be extracted at a first frame rate. In addition, a preset optical flow algorithm can be used to calculate the X / Y direction optical flow between adjacent video frames to obtain the first optical flow image features, and extract the audio features of the video to be reviewed, such as the audio features using a logarithmic Mel spectrum graph, etc., which is not restricted.
[0041] In the disclosed embodiment, the emotional features of the barrage can be used as independent modal input and dynamically aligned with the video to be reviewed through the timestamp, that is, the emotional features of the barrage, image features, and audio features are time-aligned and then fused to improve the fusion accuracy and enhance the ability to analyze the correlation of abnormal events.
[0042] S103: Obtain a first video review result of whether there is any abnormal content in the video to be reviewed based on the fusion features.
[0043] In the disclosed embodiment, sentiment analysis is performed on the barrage content of the video to be reviewed to determine the emotional characteristics of the barrage, and fusion characteristics are obtained based on the image characteristics, audio characteristics and emotional characteristics of the barrage of the video to be reviewed. The abnormal content of the video is reviewed based on the fusion characteristics to obtain the first video review result. In this way, multiple modal information such as video, audio, and emotional characteristics of barrage can be integrated to fully utilize the complementary characteristics of each modality, which can improve the accuracy of identifying abnormal content in the video.
[0044] In some possible embodiments, barrages are usually real-time comments sent by users when watching live broadcasts or videos. The content text of such barrages usually has the following characteristics: time correlation: each barrage has a strict timestamp, which is usually closely related to a frame or event in the video; semantic fragmentation: a single barrage is usually short (such as 1 to 10 words), and it is common to omit components such as the subject and predicate, and the semantic independence is poor; high repetitiveness: screen-swiping phenomena are common, such as repeatedly sending Internet buzzwords or emoticons such as "hahaha" and "666"; strong event dependence and weak language context: barrages lack semantic context, but are usually closely related to the video events and collective atmosphere at the time; diverse language styles: including colloquialisms, pinyin abbreviations, Internet terms, dialects, and even self-created symbols, with varied styles; real-time interactivity: users post instantly while the video is playing, usually with strong emotional expression and quick reactions.
[0045] The present disclosure provides a possible embodiment for determining the emotional characteristics of the barrage content based on the characteristics of the barrage content. The determination of the emotional characteristics of the barrage content based on the barrage content of the video to be reviewed in step S101 includes:
[0046] 1) Segment the bullet comment content of the video to be reviewed according to the time sequence and time window, and obtain bullet comment content segments corresponding to multiple time windows.
[0047] For example, if the time window is 5 seconds, the barrage content of the video to be reviewed is segmented in chronological order, and the barrage content within every 5 seconds is merged into a corpus segment, thereby obtaining multiple barrage content segments.
[0048] Furthermore, in the present disclosure, the barrage content can also be pre-processed before being input into the first model. For example, the pre-processing of the barrage content can include: constructing a barrage word table and an expression mapping table, and then performing word conversion on the barrage content. For example, "666", "awsl", "hahahaha" and the like in the barrage content are converted into unified labels such as "[Cheers]", "[Surprised]", etc. In this way, the generalization ability of the model can be improved, noise can be reduced, and accuracy can be improved.
[0049] 2) Based on the first model, semantic analysis is performed on each barrage content segment to determine the barrage emotional features corresponding to each barrage content segment, wherein the barrage emotional features of each barrage content segment correspond to each video frame in the time window.
[0050] Among them, the first model is obtained by training based on the barrage content sample of the video sample, or based on the barrage content sample of the video sample and related information. The first model is used to identify the emotional type of the barrage.
[0051] In the disclosed embodiment, the emotional features of the bullet screen corresponding to the bullet screen content segments are obtained, which can be associated with each video frame in the time window and can be used for time alignment during subsequent fusion.
[0052] In a possible embodiment, the first model is trained in the following manner: taking the barrage content sample of the video sample as input, or taking the barrage content sample and related information of the video sample as input, obtaining the predicted barrage emotion type of the barrage content sample through the first model, and training the first model based on the predicted barrage emotion type and the corresponding barrage emotion type label; wherein the related information includes at least one of the following: image features of the video frame in the video sample, audio features of the video frame in the video sample, user information corresponding to the barrage content sample, and corresponding identified event type in the video sample.
[0053] In the disclosed embodiment, the accuracy and reliability of model training can be further improved by training the first model through the content of the bullet comment and combining a variety of different related information. In one embodiment, the image features and audio features of the video frame can be combined, or the subtitle content, other comment content, etc. can be combined to construct a multimodal input to train the first model. For example, the video frame image is converted into a semantic vector using a model such as Contrastive Language–Image Pre-training (CLIP) or Bootstrapped Language-Image Pre-training (BLIP) and input into the model together with the bullet comment content to enhance the semantic correspondence capability. In another embodiment, pseudo labels can be used for corpus enhancement. For example, the bullet comment content within a certain time period is clustered, and semantic embedding analysis and feature clustering are performed to enhance the feature representation of the bullet comment. Signals such as event popularity and user likes in the video sample can be introduced to generate emotion or topic pseudo labels, and then sentiment classifiers or multi-task learning are performed based on the pseudo labels to improve the accuracy of the barrage sentiment analysis. In another embodiment, user information such as user identity, user activity, and historical user behavior can be incorporated into the training of the first model to enhance the semantic expression of the characteristics of the barrage content. This can also support higher-level analytical models such as user behavior prediction and influence analysis. In another embodiment, the event type identified in the video can be bound to the barrage content. For example, if the current event type is "product flash sale," the first model can be trained to identify whether the barrage content contains emotional characteristics such as joy, disappointment, complaints, and questions, which can improve accuracy.
[0054] In the disclosed embodiment, there is no restriction on the training of the first model and the use of associated information during training. It can be customized according to needs through various methods such as time context modeling, multimodal information fusion, and language style standardization. Combined with the model structure and optimization method, it can improve the ability to understand the content of the barrage, thereby providing better basic support for the barrage sentiment analysis and improving the accuracy of the barrage sentiment analysis.
[0055] In some possible embodiments, obtaining fusion features based on the image features, audio features, and comment emotion features of the video to be reviewed in step S102 may include:
[0056] 1) Concatenate the video frame image features and the first optical flow features to determine the query features.
[0057] For example, an attention network can be used to fuse image features, audio features, and barrage emotion features to obtain fused features.
[0058] The image features include video frame image features and first optical flow features of a first optical flow map, and the first optical flow map is obtained by calculating the optical flow between adjacent video frames in the video to be reviewed.
[0059] For example, the video frame image feature is (T, 1024), and the first optical flow feature is (T, 1024), where T is the feature matrix representation and 1024 is the dimension size. The video frame image feature and the first optical flow feature are concatenated and input into the encoding model (for example, ResNet50, with an output dimension of 1024) to determine the query feature (Query).
[0060] 2) Determine key features and value features based on audio features and barrage emotional features.
[0061] For example, based on the attention network, the audio features and barrage emotional features are linearly transformed to determine the key features and value features. In order to adapt the dimensions of each feature, the audio features can be represented as (T, 512) after the convolution layer, and the barrage emotional features can be represented as (T, 512) after the linear layer. The key features (Key) and value features (Value) of the attention network are represented as audio features + barrage emotional features.
[0062] 3) Based on the query features and key features, determine the cross-attention weights among the video frame image features, the first optical flow features, the audio features, and the barrage emotion features.
[0063] For example, the query feature and the key feature are multiplied and activated to obtain the cross attention weight.
[0064] 4) Obtain fusion features based on the cross attention weights and value features.
[0065] In addition, in the embodiments of the present disclosure, other fusion strategies can also be used when fusing multimodal features, such as a series method of fusing image features with barrage emotional features and then fusing them with audio features. For example, cross-modal relationship modeling based on graph neural networks can be used to replace attention networks for feature fusion. This is not limited in the embodiments of the present disclosure.
[0066] In addition, in the embodiment of the present disclosure, the time point at which abnormal content exists in the video to be reviewed can also be located based on the updated cross-attention weight. For example, for any video frame in the video to be reviewed, when its image features, audio features and barrage emotional features are fused, the cross-attention weight of each feature can be obtained. At the same time, if an abnormality is detected in the video frame image content by performing anomaly detection on the video frame image content, and the cross-attention weight of the image feature corresponding to the video frame is greater than the set threshold, the abnormal content priority mark of the video frame can be directly triggered to locate the abnormal content of the video. This is because the video frame screen usually contains richer information. If the image features account for a large proportion of the fused features during fusion, and it is determined that the video frame screen contains abnormal content, it can be considered that the first video review result determined based on the fused features is also likely to contain abnormal content. In this way, the time point at which abnormal content appears in the video to be reviewed can be quickly located to a certain extent, thereby achieving rapid anomaly location.
[0067] In some possible embodiments, in the above step S103, based on the fusion features, a first video review result of whether there is abnormal content in the video to be reviewed is obtained. Specifically, the fusion features can be input into a pre-trained network model to obtain the first video review result of whether there is abnormal content. The network model is used to classify and identify abnormal or normal content.
[0068] In some possible embodiments, obtaining the first video review result of whether there is abnormal content in the video to be reviewed based on the fusion features in step S103 may include:
[0069] S1041: Determine the abnormal time period in which the abnormality occurs in the video to be reviewed based on the fusion features.
[0070] S1042: Perform content anomaly detection based on the first video segment corresponding to the abnormal time period in the video to be reviewed, and obtain a first video review result of whether the video to be reviewed contains abnormal content.
[0071] The following describes in detail each step of the above embodiment.
[0072] In some possible embodiments, see Figure 2FIG. 1 is a flowchart of spatiotemporal joint anomaly location in an embodiment of the present disclosure. The flowchart includes: determining the abnormal time period in which the abnormality occurs in the video to be reviewed based on the fusion features in step S1041 above; and
[0073] S1: Perform traffic anomaly detection based on the fusion features corresponding to each video frame in the video to be reviewed, and determine the first time period in which the traffic anomaly occurs in the video to be reviewed.
[0074] This step specifically includes:
[0075] 1) According to the duration of a preset time step, a sequence of video frame segments arranged in time order is determined, wherein each time step corresponds to a video frame segment.
[0076] For example, if the preset time step is 5 seconds, the video to be reviewed can be segmented every 5 seconds to obtain a sequence of video frame segments. In addition, the sequence of video frame segments can be standardized to eliminate dimensional effects.
[0077] 2) Determine the target hidden state features corresponding to each video frame segment based on the fusion features corresponding to the video frames in each video frame segment, where the target hidden state features represent the dependency relationship with the video frame segments at the time step before the current time step and the time step after the current time step.
[0078] For example, Figure 2 As shown, the LSTM method can be used. The LSTM parameter settings are, for example, a hidden layer dimension of 256, a time window of 5 seconds, a sliding step of 1 second, and a forward LSTM processes a sequence of video frame fragments from left to right, which can capture the historical dependency of the video frame fragments with the time step before the current time step. The backward LSTM processes a sequence of video frame fragments from right to left, which can capture the future dependency of the video frame fragments with the time step after the current time step, thereby splicing the hidden state obtained by the forward processing and the hidden state obtained by the backward processing to obtain the target hidden state feature corresponding to the video frame fragment, such as a dimension of 2×256=512.
[0079] 3) According to the target hidden state features corresponding to each video frame segment, the similarity between the target hidden state features corresponding to the video frame segments at adjacent time steps is determined.
[0080] In the disclosed embodiment, a bidirectional LSTM method is used to output target hidden state features of the video frame segment at each time step. The output layer usually does not use a display classification layer, but instead calculates similarity through the hidden state, which is more accurate.
[0081] For example, if the adjacent time steps are the i-th step and the i+1-th step, the cosine similarity method can be used for similarity. Then, the similarity between the target hidden state features corresponding to the i-th step and the i+1-th step can be expressed as:
[0082]
[0083] Among them, V i is the target hidden state feature of the i-th step, V i+1 is the target hidden state feature of the i+1th step.
[0084] 4) According to the similarity between the target hidden state features corresponding to the video frame segments of adjacent time steps, the first time period corresponding to the traffic anomaly of the video to be reviewed is determined.
[0085] In one possible implementation, if the similarities between a plurality of consecutive adjacent time steps are all less than a preset threshold, a middle time step among the plurality of consecutive adjacent time steps may be marked as an abnormal first time period.
[0086] For example, the preset threshold is 0.85, and the number of consecutive times is three. If the similarities between time steps i and i+1, i+1 and i+2, and i+2 and i+3 are all less than 0.85, then the i+2 can be determined as a mutation point, that is, the first abnormal time period.
[0087] S2: Determine the video frames with significant areas based on the heat map of each video frame, and determine the second time period with significant areas corresponding to the video to be reviewed based on the video frames with significant areas.
[0088] For this step, the present disclosure provides a possible implementation method for determining the heat map. Based on multiple convolution branches in the convolutional network, convolution features are extracted for each video frame to obtain multiple feature maps corresponding to each video frame, and the convolution kernel sizes of the multiple convolution branches are different. For each video frame in each video frame, the multiple feature maps corresponding to each video frame are fused to obtain a heat map of each video frame.
[0089] For example, the ST-MCNN method can be used, which uses three parallel convolution branches and uses convolution kernels of different sizes, such as 3×3, 5×5, and 7×7 convolution kernel sizes. Among them, the small-scale (3×3) convolution branch is mainly used to capture local details and edge information, the medium-scale (5×5) convolution branch is mainly used to balance local and global features, and the large-scale (7×7) convolution branch is mainly used to extract broader contextual information, which is suitable for detecting large-scale anomalies.
[0090] Then, based on each convolution branch, convolution features are extracted from the video frame to generate a feature map. The feature maps output by each convolution branch may have different resolutions and need to be adjusted to the same resolution as the original video frame through upsampling or interpolation. The feature maps of each convolution branch are then spliced (i.e., fused) to obtain multi-scale fused features. The multi-scale fused features are converted into a saliency heat map with the same resolution as the original video frame through additional convolution layers or fully connected layers. Each pixel in the heat map corresponds to a saliency score, and the range is usually normalized to [0,1]. The saliency score of the pixel can be expressed as:
[0091] S(x,y)=α·Local(x,y)+β·Mid(x,y)+γ·Global(x,y)
[0092] Among them, α+β+γ=1, the parameter value can be optimized and adjusted, S(x,y) represents the significance score of the pixel point (x,y), Local(x,y) represents the convolution feature obtained by the small-scale convolution branch, Mid(x,y) represents the convolution feature obtained by the medium-scale convolution branch, and Global(x,y) represents the convolution feature obtained by the large-scale convolution branch.
[0093] S3: Determine the abnormal time period in which the abnormality occurs in the video to be reviewed based on the overlapping time and the length of the overlapping time between the first time period and the second time period.
[0094] For example, if there is overlapping time between the first time period and the second time period, and the duration of the overlapping time is greater than or equal to a set value, such as the set value is 0.5 seconds, the overlapping time is determined to be an abnormal time period.
[0095] For example, Figure 2 As shown, according to the fusion features of each video frame, LSTM can be used for time series analysis to determine the first time period where the traffic anomaly occurs, that is, to determine the timestamp of the mutation point, and ST-MCNN can be used for spatial analysis to determine the second time period of the video frame with a significant area. Among them, the execution order of LSTM traffic anomaly detection and ST-MCNN significant area detection is not restricted, and can also be executed in parallel to improve efficiency. Then, the positioning results of the two are combined to jointly determine the final abnormal time period. Through the joint time-space analysis and positioning, the accuracy of abnormal time period positioning can be improved.
[0096] In the disclosed embodiment, a method for joint positioning of LSTM and ST-MCNN is used, where LSTM can capture long-term traffic mutation events (such as a surge in the amount of barrage, a sudden increase in audio energy, etc.), and ST-MCNN can locate short-term visual abnormal events (such as abnormal actions, abnormal images, etc.), thereby combining the global temporal modeling of bidirectional LSTM with the local saliency detection of ST-MCNN to achieve spatiotemporal dual-channel complementarity, thereby breaking through the limitations of a single modality, and solving the problem of false detection caused by single-modality noise through spatial alignment of time window sliding and saliency heat map.
[0097] When performing joint positioning, abnormal positioning decisions can also be made based on preset events. For example, visually significant events are detected separately (for example, they can be detected based on the ST-MCNN method), and barrage or audio abnormal mutation events are detected (for example, they can be detected based on the LSTM method). If the time overlap rate of visually significant events and barrage / audio abnormal mutation events is greater than or equal to 80%, a joint alarm is triggered, that is, the overlapping time is positioned as an abnormal time period. In one possible embodiment, the confidence level of visually significant events and barrage / audio abnormal mutation events is greater than the confidence threshold in at least n consecutive time steps, for example, n is 3 and the confidence threshold is 0.9.
[0098] The traffic mutation module (LSTM predicts historical traffic baseline) and the motion saliency module (improved ST-MCNN, adding channel attention) independently output confidence scores, and reduce the false trigger rate through weighted fusion (the weight is adaptively adjusted by the anomaly type).
[0099] In the embodiment of the present disclosure, the abnormality localization decision logic in the joint localization method is a dynamically adjustable process. For example, for the spatial weight of the ST-MCNN method, the visual modality weight can be dynamically adjusted according to the confidence score of the salient area (such as the entropy value of the saliency map). For the temporal weight of the bidirectional LSTM method, the temporal modality weight can be adjusted based on the mutation probability gradient output by the LSTM (such as the mutation probability difference between adjacent video frames). For example, the weight detected by the ST-MCNN method is The weight detected by the LSTM method is Specifically:
[0100]
[0101] Among them, S t is the significance score of the detected significant region, ΔP t is the probability change rate of detecting traffic anomalies, and α, β, and γ are hyperparameters.
[0102] Furthermore, in the embodiments of the present disclosure, when abnormal events are detected based on different modalities, weighted fusion can be performed based on their corresponding weights and the probability values of the detected abnormal events to obtain the abnormal confidence level, and then a final judgment can be made based on the abnormal confidence level whether to mark it as an abnormal time period. For example, if a visually significant event is detected but there is no sudden change in the timing, that is, a cross-modal conflict is detected, since visually significant events can usually indicate that abnormal images have appeared in the video frame, and there may be no sudden abnormalities in the audio or barrage, in this case it can still be considered that the video frame contains abnormal content. At this time, the weight adaptive adjustment can be triggered to increase the weight value of the visually significant event, which can reduce missed detections and improve accuracy.
[0103] In addition, in the embodiments of the present disclosure, in addition to using LSTM and ST-MCNN, a timing model based on a 3D convolutional network or a Transformer network can also be used to capture motion changes, while combining methods such as optical flow mutation detection. This is not limited in the embodiments of the present disclosure.
[0104] In some possible embodiments, possible implementations are further provided for the above step S1042, specifically:
[0105] S1051: For a first video segment corresponding to an abnormal time period in the video to be reviewed, extract multiple key frames according to a first frame rate, and perform motion compensation on the multiple key frames to generate compensated frames.
[0106] S1052: Generate a second video segment according to the multiple key frames and the compensation frames, where a second frame rate of the second video segment is greater than the first frame rate.
[0107] S1053: Perform content anomaly detection on the second video clip to obtain the first video review result of whether the second video clip contains abnormal content.
[0108] In the disclosed embodiment, based on the fusion features, the abnormal time period in the video to be reviewed is first located, and then only the first video segment of the suspected abnormal time period is interpolated. Through motion compensation, a new second video segment with a higher frame rate is generated. This can reduce the amount of calculation and be more efficient, and also retain important key frames. Then, content anomaly detection is performed on the second video segment with a higher frame rate. The frame rate is higher and the motion is smoother and more fluent, similar to a slow motion effect, which can enhance detail analysis, facilitate content anomaly detection, and improve detection accuracy.
[0109] In some possible embodiments, for the first video segment corresponding to the abnormal time period in the video to be reviewed in step S1051, extracting multiple key frames according to the first frame rate, and performing motion compensation on the multiple key frames to generate compensated frames include:
[0110] 1) Extend the abnormal time period in the video to be reviewed to generate an extended abnormal time period, wherein the extended abnormal time period does not exceed the time range of the video to be reviewed.
[0111] For example, the abnormal time period is [T_start, T_end], which is extended by 2 seconds forward and backward to generate an extended abnormal time period [T_start-2s, T_end+2s]. If the extended time period exceeds the time range of the video to be reviewed, it will be truncated to the time range of the video to be reviewed so as not to exceed the time range.
[0112] In this way, when processing the abnormal time period, expanding the before and after time windows can capture the actions before and after the abnormality occurs, obtain more content details, and improve the accuracy of subsequent abnormal content detection.
[0113] 2) Extracting a plurality of key frames from a first video segment corresponding to the extended abnormal time period of the video to be reviewed at a first frame rate.
[0114] 3) Determine pixel motion vector information between adjacent key frames in the plurality of key frames, and generate compensation frames between the adjacent key frames based on the pixel motion vector information between the adjacent key frames.
[0115] For example, the first frame rate is 25fps. In the first video clip of the extended abnormal time period, multiple key frames are extracted according to the first frame rate. Then, the optical flow method can be used to calculate the pixel motion vector information between adjacent key frames, and the deformable convolutional network can be used to generate compensation frames between adjacent key frames, which can ensure motion consistency and continuity, which is equivalent to generating more video frames. The second video clip obtained based on the key frames and the compensation frames has an increased frame rate due to the increase in the number of video frames it contains, and its duration is the total duration of the extended abnormal time period. For example, the duration of the abnormal time period is 3 seconds, and the extended abnormal time period is 7 seconds. The second frame rate of the second video clip is 50fps, and the total number of video frames is 7×50=350 frames.
[0116] In the disclosed embodiment, by extracting key frames and generating compensation frames for the first video clip, a second video clip with a higher frame rate can be obtained, and a second video clip with a slow-motion effect can be obtained. The frame rate is higher, and more details in the same video frame sequence will be restored and captured. Motion compensation is performed on the key frames, and more key frames are retained, which effectively saves storage and computing resources and optimizes resource usage while ensuring the integrity of the key frames.
[0117] In some possible embodiments, generating a second video segment according to the plurality of key frames and compensation frames in step S1052 includes:
[0118] 1) Generate a third video segment by combining multiple key frames and compensation frames in chronological order.
[0119] Furthermore, in the embodiment of the present disclosure, motion blur detection can also be performed on the third video clip. If the blur exceeds a threshold (for example, the gradient variance is less than 0.1), a compensation frame can be regenerated to obtain a third video clip based on the regenerated compensation frame and key frame. In this way, the quality and effect of motion compensation can be improved.
[0120] 2) Based on the diffusion model, noise is added to each video frame in the third video segment to obtain a noise image of each video frame in the third video segment.
[0121] In the disclosed embodiment, after obtaining the third video segment based on the key frames and the compensation frames, the video frames in the third video segment may be further enhanced to further improve the quality of the video frames, thereby improving the accuracy of subsequent abnormal content detection.
[0122] For example, a diffusion model can be used to perform noise addition and denoising processing on the third video clip to achieve reconstruction of the video frames in the third video clip. Furthermore, in order to improve accuracy and reliability, constraints can be added during the noise addition and denoising processing. For example, the barrage emotional features and / or the second optical flow features corresponding to the third video clip are used as constraints of the diffusion model, wherein the second optical flow features represent the optical flow information between adjacent video frames in the third video clip. Of course, other constraints can also be used, such as preset emotion types as constraints, etc., which are not limited in the embodiments of the present disclosure. In this way, it is possible to guide the generation of slow-motion video frames that conform to physical laws and are more semantically consistent, and make the generated video frames more consistent with the actual situation, which can further improve the robustness.
[0123] 3) De-noising is performed on the noise image of each video frame in the third video segment to obtain a new video frame generated corresponding to each video frame in the third video segment.
[0124] The noise intensity of the added noise can be dynamically adjusted. For example, the initial noise intensity is 0.5 and can be linearly attenuated. This is not limited in the embodiments of the present disclosure.
[0125] During the denoising and denoising process, for example, the loss function of the diffusion model is L1 reconstruction loss (weight = 0.7) + VGG19 perception loss (weight = 0.3), which is mainly used to adjust and enhance the visual authenticity and naturalness of the newly generated video.
[0126] In another possible embodiment, in the embodiment of the present disclosure, a preset review task label corresponding to the video to be reviewed can also be obtained. For example, the review task is to determine whether the video to be reviewed is a "Class A" video, that is, the review task label is "Class A". "Class A" can be encoded to obtain a text embedding vector, and then spliced with the second optical flow feature as a constraint condition of the diffusion model.
[0127] 4) Obtain a second video segment based on the generated new video frame.
[0128] Optionally, the new video frames obtained through the diffusion model can be further tested for spatiotemporal consistency, mainly to detect whether the actions or pictures in the generated second video clip are coherent, such as whether there are abrupt or unconventional action switches. If detected, inconsistent video frames can be sampled for regeneration, which can improve the smoothness and quality of the generated video.
[0129] In the disclosed embodiment, a slow-motion video frame sequence can be generated for a third video clip with a high frame rate based on a conditional diffusion model, and based on the diffusion model, the emotional features of the barrage and the second optical flow features are used as constraints to reconstruct the video frames, thereby enhancing the visualization effect of hidden details, reflecting more information on data blind spots or blurred areas, improving generalization capabilities, and allowing subtle movements that are difficult to detect in the original video to be clearly captured, thereby improving the review accuracy and reducing the false alarm rate.
[0130] In a possible embodiment, the second video clip obtained based on the diffusion model in the embodiment of the present disclosure can be further fed back into the process of obtaining fusion features. For example, based on the image features in the second video clip, and fused with the corresponding audio features and barrage emotional features to obtain fusion features, the accuracy of positioning and review of abnormal time periods can be improved.
[0131] In addition, in the embodiments of the present disclosure, content anomaly detection is performed after processing based on motion compensation and diffusion model, which can improve detection accuracy. There may be a loss in detection time. In order to comprehensively consider the detection cost and performance, dynamic feedback can also be performed on the entire detection process. For example, the abnormal time period in the video to be reviewed is determined, and then motion compensation is performed to increase the frame rate. Enhanced video frames are generated based on the diffusion model, and then content anomaly detection is performed to obtain the first video review result. Then, the parameters of the modified diffusion model can be fed back to improve the performance of the diffusion model. This is not limited in the embodiments of the present disclosure.
[0132] In some possible embodiments, performing content anomaly detection on the second video clip in step S1053 to obtain a first video review result of whether the second video clip contains abnormal content includes:
[0133] 1) For each video frame in the second video clip, randomly mask a non-salient area in the video frame to generate a first noise video frame, and perform noise processing on each video frame in the second video clip to generate a second noise video frame.
[0134] For example, the heat map of the video frame can be calculated to determine the non-significant areas and significant areas in the video frame, and then the non-significant areas in the video frame can be randomly blocked, where the size of the blocking block can be set according to needs. In one optional method, the non-significant area is blocked by less than 30% of the non-significant area, which can remove the background area in the video frame and highlight the abnormal picture content. In addition, the abnormal content of the video frame can be enhanced by adding noise to the video frame, such as adding salt and pepper noise with a probability of 0.05. There is no restriction on this.
[0135] In this way, in the embodiment of the present disclosure, adversarial samples of video frames can be generated through random occlusion and noise addition processing to improve the accuracy of anomaly detection.
[0136] 2) According to each video frame in the second video segment and the corresponding first and second noise video frames, obtain a probability value of the presence of abnormal content in each video frame in the second video segment and a saliency intensity value of the presence of a salient area.
[0137] For example, the feature vectors of the video frame, the first noisy video frame, and the second noisy video frame can be extracted respectively, and after fusion processing, the probability value of the presence of abnormal content can be output based on the anomaly classification model, and its significance intensity value can be calculated through ST-MCNN.
[0138] 3) Based on the probability value of the existence of abnormal content in each video frame in the second video clip and the significant intensity value of the existence of significant areas, determine the second video review result of whether there is abnormal content in each video frame in the second video clip, and the confidence score of the second video review result.
[0139] For this step, a possible embodiment is specifically provided. For each video frame in the second video clip, in response to the probability value of the existence of abnormal content in the video frame in the second video clip being greater than a first threshold and the significant intensity value of the presence of a significant area being greater than a second threshold, the second video review result of the video frame in the second video clip is determined to be the existence of abnormal content; based on the first weight value of the probability value of the abnormal content and the second weight value of the significant intensity value, the probability value and the significant intensity value of the video frame in the second video clip are weightedly summed to obtain a confidence score for the second video review result.
[0140] For example, the probability value of obtaining abnormal content in a video frame is P, and the significance intensity value is S, where P∈[0,1], S∈[0,1], the first threshold is 0.9, the second threshold is 0.7, the first weight value is 0.6, and the second weight value is 0.4. Then, when P is greater than 0.9 and S is greater than 0.7, the second video review result is determined to be the presence of abnormal content; otherwise, the second video review result is determined to be the absence of abnormal content, and the confidence score of the second review result can be expressed as 0.6P+0.4S.
[0141] 4) For each video frame in the second video clip, in response to the second video review result being that abnormal content exists and the confidence score is greater than or equal to the score threshold, determine that the first video review result of the video frame in the second video clip is that abnormal content exists.
[0142] In addition, in the embodiment of the present disclosure, other methods may be used to generate noise video frames, for example, based on PGD, FGSM methods, etc., and other threshold determination logics may be used, which are not limited in the embodiment of the present disclosure.
[0143] In this way, by introducing adversarial sample training and dual-threshold decision-making, the accuracy of the judgment of the second video review result can be improved based on the classification probability value and significance intensity value, and the confidence score of the second video review result can be determined based on the probability value and significance intensity value, which can further determine the reliability of the second video review result, thereby improving the accuracy and robustness of the final first video review result.
[0144] Further, in some possible embodiments, in response to the first video review result of more than a preset number of video frames being continuously abnormal, the video to be reviewed is processed accordingly according to a preset processing strategy; or, in response to the first video review result of more than a preset number of video frames being continuously abnormal, the user account of the video to be reviewed is processed accordingly.
[0145] For example, in a live streaming scenario, if abnormal content is detected in the video to be reviewed multiple times in a row, the live streaming account can be frozen or the live streaming can be shut down. Exceptions can be handled in a timely manner to improve security, and manual review can be performed to improve processing accuracy and avoid mishandling.
[0146] In a possible example, a specific application scenario is used below to illustrate the video review method in the embodiment of the present disclosure. Figure 3 The figure shows the architecture diagram of the video review in the embodiment of the present disclosure.
[0147] like Figure 3As shown, the multimodal data input layer includes multimodal data such as the video stream, audio stream and barrage content of the video to be reviewed, and can obtain image features through the video frame extraction module, audio features through the audio processing module, and barrage emotional features through the barrage emotional analysis module.
[0148] Multimodal fusion and preprocessing layer: Based on the attention network, image features, audio features and barrage emotional features are fused to obtain fusion features.
[0149] Spatiotemporal anomaly localization layer: The fused features are input into the spatiotemporal anomaly localization layer to perform flow anomaly detection and significance detection respectively, which can be executed in parallel. Then, based on the first time period determined by the flow anomaly detection and the second time period determined by the significance detection, the abnormal time period where the anomaly occurs in the video to be reviewed is determined, which can improve the accuracy of anomaly localization.
[0150] Dynamic sampling and enhancement layer: The dynamic sampling module is used to extract multiple key frames from the first video segment corresponding to the abnormal time period in the video to be reviewed, and perform motion compensation on the multiple key frames to generate compensation frames. It can automatically expand the time window of the abnormal time period to obtain an extended abnormal time period, and increase the frame rate by generating compensation frames. At the same time, key frames are selected based on importance.
[0151] The diffusion model enhancement module is used to generate a second video clip based on multiple key frames and compensation frames. For example, in one possible embodiment, based on the diffusion model, new video frames are reconstructed and generated with the emotional features of the barrage and the second optical flow features as constraints to obtain the second video clip. Higher quality slow motion frames can be generated to enhance detailed content analysis.
[0152] Intelligent review and output layer: The adversarial sample generation module is mainly used to generate a first noisy video frame and a second noisy video frame through random occlusion and noise processing to construct an adversarial sample.
[0153] The decision module is mainly used to detect and make decisions on abnormal content based on the video frame and the corresponding first noise video frame and second noise video frame. For example, in the embodiment of the present disclosure, a dual threshold decision can be adopted based on the probability value of the existence of abnormal content and the significant intensity value of the significant area to determine the final video review result.
[0154] In the embodiment of the present disclosure, multiple modal information is used, and through fusion processing, joint anomaly positioning, motion compensation and diffusion model enhancement, adversarial sample generation and multi-strategy decision-making, such as Figure 3 The collaborative processing of various modules can improve the efficiency and accuracy of video review.
[0155] Figure 4 A block diagram of an electronic device provided in an embodiment of the present disclosure. Figure 4An embodiment of the present disclosure provides an electronic device, which includes: at least one processor 401; at least one memory 402, and one or more I / O interfaces 403, connected between the processor 401 and the memory 402; wherein the memory 402 stores one or more computer programs that can be executed by the at least one processor 401, and the one or more computer programs are executed by the at least one processor 401 so that the at least one processor 401 can execute the above-mentioned video review method.
[0156] Each module in the above-mentioned electronic device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0157] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned video review method when executed by a processor. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.
[0158] An embodiment of the present disclosure also provides a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned video review method.
[0159] Among them, the processor is a device with data processing capabilities, including but not limited to the central processing unit (CPU); the memory is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically such as SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read-write interface) is connected between the processor and the memory, which can realize information exchange between the memory and the processor, including but not limited to the data bus (Bus), etc.
[0160] Those skilled in the art will appreciate that all or some of the steps, systems, and functional modules / units disclosed above may be implemented as software, firmware, hardware, or a suitable combination thereof. In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division between physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components.
[0161] Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit (CPU), a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH) or other disk storage; compact disc (CD-ROM), digital versatile disc (DVD) or other optical disc storage; magnetic cassettes, tapes, disk storage or other magnetic storage; any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0162] The present disclosure has disclosed example embodiments, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.
Claims
1. A video review method, comprising: Determine the emotional characteristics of the barrage content of the video to be reviewed; Obtaining fusion features based on the image features, audio features, and emotional features of the video to be reviewed; A first video review result is obtained based on the fusion feature to determine whether abnormal content exists in the video to be reviewed.
2. The method according to claim 1, wherein Determining the emotional characteristics of the barrage content of the video to be reviewed includes: Segment the bullet comment content of the video to be reviewed according to the time sequence and time window to obtain bullet comment content segments corresponding to multiple time windows; Based on the first model, semantic analysis is performed on each of the barrage content fragments to determine the barrage emotional features corresponding to each of the barrage content fragments, wherein the barrage emotional features of each barrage content fragment correspond to each video frame in the time window. The first model is a barrage content sample based on a video sample, or is obtained by training based on the barrage content sample and associated information based on the video sample. The first model is used to identify the barrage emotional type.
3. The method according to claim 2, wherein: The first model is trained in the following way: Taking a video sample's barrage content sample as input, or taking a video sample's barrage content sample and associated information as input, obtaining a predicted barrage emotion type of the barrage content sample through the first model, and training the first model based on the predicted barrage emotion type and the corresponding barrage emotion type label; The associated information includes at least one of the following: image features of the video frame in the video sample, audio features of the video frame in the video sample, user information corresponding to the barrage content sample, and the corresponding identified event type in the video sample.
4. The method according to claim 1, wherein The image features include video frame image features and first optical flow features of a first optical flow map, where the first optical flow map is obtained by calculating the optical flow between adjacent video frames in the video to be reviewed. The fusion features are obtained based on the image features, audio features, and emotional features of the barrage of comments of the video to be reviewed, including: Splicing the video frame image feature and the first optical flow feature to determine a query feature; Determine key features and value features based on the audio features and the emotional features of the barrage; Determining, based on the query feature and the key feature, a cross-attention weight among the video frame image feature, the first optical flow feature, the audio feature, and the barrage emotion feature; A fusion feature is obtained according to the cross attention weight and the value feature.
5. The method according to any one of claims 1 to 4, wherein: Obtaining a first video review result of whether abnormal content is present in the video to be reviewed based on the fusion feature includes: Determining, based on the fusion features, an abnormal time period in which an abnormality occurs in the video to be reviewed; A content anomaly detection is performed based on a first video segment corresponding to the abnormal time period in the video to be reviewed to obtain a first video review result of whether the video to be reviewed contains abnormal content.
6. The method according to claim 5, wherein: Determining, based on the fusion features, an abnormal time period in which an abnormality occurs in the video to be reviewed, includes: Performing traffic anomaly detection based on the fusion features corresponding to each video frame in the video to be reviewed, and determining the first time period in which the traffic anomaly occurs in the video to be reviewed; Determining a video frame having a significant area based on the heat map of each video frame, and determining a second time period having a significant area corresponding to the video to be reviewed based on the video frame having the significant area; An abnormal time period in which an abnormality occurs in the video to be reviewed is determined according to the overlapping time and the duration of the overlapping time between the first time period and the second time period.
7. The method according to claim 6, wherein: The performing of traffic anomaly detection based on the fusion features corresponding to each video frame in the video to be reviewed, and determining a first time period corresponding to the occurrence of traffic anomaly in the video to be reviewed, includes: Determining a sequence of video frame segments arranged in time order according to a duration of a preset time step, wherein each time step corresponds to a video frame segment; Determining, based on fusion features corresponding to video frames in each video frame segment, a target hidden state feature corresponding to each video frame segment, wherein the target hidden state feature represents a dependency relationship with video frame segments at a time step before a current time step and at a time step after the current time step; Determining the similarity between target hidden state features corresponding to video frame segments at adjacent time steps based on the target hidden state features corresponding to each of the video frame segments; According to the similarity between the target hidden state features corresponding to the video frame segments of adjacent time steps, the first time period in which the traffic anomaly occurs corresponding to the video to be reviewed is determined.
8. The method according to claim 6, wherein: The method further comprises: Based on multiple convolution branches in a convolutional network, convolution features are extracted for each video frame to obtain multiple feature maps corresponding to each video frame, where the convolution kernels of the multiple convolution branches have different sizes; For each of the video frames, a plurality of feature maps corresponding to each of the video frames are fused to obtain a heat map of each of the video frames.
9. The method according to claim 5, wherein: The performing of content anomaly detection based on the first video segment corresponding to the abnormal time period in the video to be reviewed to obtain a first video review result of whether the video to be reviewed contains abnormal content includes: Extracting a plurality of key frames according to a first frame rate from a first video segment corresponding to the abnormal time period in the video to be reviewed, and performing motion compensation on the plurality of key frames to generate compensated frames; generating a second video segment according to the plurality of key frames and the compensation frame, wherein a second frame rate of the second video segment is greater than the first frame rate; The second video clip is subjected to content anomaly detection to obtain a first video review result of whether the second video clip contains abnormal content.
10. The method according to claim 9, wherein: The step of extracting a plurality of key frames according to a first frame rate from a first video segment corresponding to the abnormal time period in the video to be reviewed, and performing motion compensation on the plurality of key frames to generate compensated frames includes: Extending the abnormal time period in the video to be reviewed to generate an extended abnormal time period, wherein the extended abnormal time period does not exceed the time range of the video to be reviewed; extracting, according to the first frame rate, a plurality of key frames from a first video segment corresponding to the extended abnormal time period of the video to be reviewed; Pixel motion vector information between adjacent key frames among the plurality of key frames is determined, and compensation frames between the adjacent key frames are generated according to the pixel motion vector information between the adjacent key frames.
11. The method according to claim 9, wherein Generating a second video segment according to the multiple key frames and the compensation frame includes: Generate a third video segment by using the multiple key frames and the compensation frames in chronological order; adding noise to each video frame in the third video segment based on a diffusion model to obtain a noise image of each video frame in the third video segment; performing denoising on the noise image of each video frame in the third video segment to obtain a new video frame generated corresponding to each video frame in the third video segment; A second video segment is obtained according to the generated new video frame.
12. The method according to claim 9, wherein The performing content anomaly detection on the second video clip to obtain a first video review result of whether the second video clip contains abnormal content includes: For each video frame in the second video clip, randomly masking a non-salient area in the video frame to generate a first noise video frame, and performing noise processing on each video frame in the second video clip to generate a second noise video frame; Obtaining, based on each video frame in the second video segment and the corresponding first and second noise video frames, a probability value of the presence of abnormal content and a saliency intensity value of a salient region in each video frame in the second video segment; Determining a second video review result of whether each video frame in the second video clip contains abnormal content and a confidence score of the second video review result based on a probability value of the presence of abnormal content and a saliency intensity value of a salient area in each video frame in the second video clip; For each video frame in the second video clip, in response to the second video review result being that abnormal content exists and the confidence score being greater than or equal to a score threshold, it is determined that the first video review result of the video frame in the second video clip is that abnormal content exists.
13. The method according to claim 12, wherein: The second video review result of determining whether each video frame in the second video clip has abnormal content and a saliency intensity value of a saliency region in each video frame in the second video clip, and a confidence score of the second video review result, comprising: For each video frame in the second video segment, in response to a probability value of the presence of abnormal content in the video frame in the second video segment being greater than a first threshold and a saliency intensity value of the presence of a salient region being greater than a second threshold, determining a second video review result of the video frame in the second video segment as the presence of abnormal content; According to the first weight value of the probability value of the abnormal content and the second weight value of the significant intensity value, the probability value and the significant intensity value of the video frame in the second video clip are weighted and summed to obtain the confidence score of the second video review result.
14. The method according to any one of claims 1 to 13, wherein: The method further comprises: In response to the first video review result of more than a preset number of video frames being abnormal, the video to be reviewed is processed accordingly according to a preset processing strategy, or the user account of the video to be reviewed is processed accordingly.
15. An electronic device comprising a memory and a processor; the memory stores a computer program that can be executed by the processor, and when the computer program is executed by the processor, it implements the video review method described in any one of claims 1 to 14.
16. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the video review method according to any one of claims 1 to 14.
17. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for video review according to any one of claims 1 to 14 is implemented.