Video detection method, video detection device, electronic device, and storage medium
By using artificial intelligence technology to extract, segment, and fuse features from videos, the problem of high false detection risk in existing video detection methods is solved, achieving higher detection accuracy and improved video quality.
Patent Information
- Application Number
- CN202310834081.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-07
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2043-07-07
AI Technical Summary
Existing video detection methods rely on manual inspection, which carries a significant risk of false positives and affects the accuracy of video detection.
Artificial intelligence technology is used for video detection. By acquiring target video frame images, segmenting them into multiple video region images, extracting image content features, and performing feature fusion and video detection, the time points of abnormal images are determined.
It improves the accuracy of video detection, enabling rapid detection and improvement of video quality, ensuring that insurance product recommendation videos meet requirements, and enhancing transaction security.
Smart Images

Figure CN117095325B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology, and in particular to a video detection method, a video detection device, an electronic device, and a storage medium. Background Technology
[0002] With the development of computer technology and artificial intelligence, traditional offline businesses are gradually migrating online, and this trend has become irreversible. For example, online shopping, live streaming, and online transactions benefit institutions such as banks and online merchants.
[0003] In the financial sector, online sales of insurance products are often conducted through audio and video recordings. To enhance transaction security, regulatory agencies have stringent requirements for the content of marketing videos for insurance products. They frequently conduct video inspections of product recommendation videos to identify any non-compliant elements and promptly instruct relevant insurance institutions or agents to revise the videos accordingly.
[0004] Current video detection methods often rely on manual inspection to discover anomalies in videos. This method carries a significant risk of false positives, affecting the accuracy of video detection. Therefore, improving the accuracy of video detection has become an urgent technical problem to be solved. Summary of the Invention
[0005] The main objective of this application is to provide a video detection method, a video detection device, an electronic device, and a storage medium, with the aim of improving the accuracy of video detection.
[0006] To achieve the above objectives, a first aspect of this application provides a video detection method, the method comprising:
[0007] Acquire the target video;
[0008] The target video is subjected to video frame extraction to obtain the target video frame image of the target video;
[0009] The target video frame image is segmented to obtain multiple video region images and image position data of each video region image;
[0010] Feature extraction is performed on the video region image and the image location data to obtain the image content features of each target video frame image;
[0011] The video content features of the target video are obtained by fusing multiple image content features.
[0012] Based on the video content features, video detection is performed on the target video to obtain video detection data, which is used to characterize the time points when abnormal images appear in the target video.
[0013] In some embodiments, the step of extracting features from the video region image and the image location data to obtain the image content features of each target video frame image includes:
[0014] The video region image and the image location data are input into a preset feature extraction model, which includes an encoder and a decoder;
[0015] Based on the encoder, the video region image and the image location data are encoded to obtain the image content latent vector;
[0016] Based on the decoder, the latent vector of the image content is extracted to obtain the image content features.
[0017] In some embodiments, the feature fusion of multiple image content features to obtain the video content features of the target video includes:
[0018] Obtain the image time of each target video frame;
[0019] The video content features are obtained by concatenating multiple image content features based on a preset temporal model and the image time.
[0020] In some embodiments, the step of performing video detection on the target video based on the video content features to obtain video detection data includes:
[0021] Based on a preset video quality inspection model, video content features are detected to obtain quality inspection index data of the target video, wherein the quality inspection index data includes at least one actual quality inspection index.
[0022] By comparing the actual quality inspection indicators with the preset reference quality inspection indicators, abnormal images of the target video are obtained;
[0023] Extract the image time of the abnormal image to obtain the time point;
[0024] The video detection data is obtained based on the abnormal image and the time point.
[0025] In some embodiments, after performing video detection on the target video based on the video content features to obtain video detection data, the method further includes:
[0026] Based on the video detection data, abnormal images are extracted from the target video;
[0027] Based on the image content of the abnormal image, an intermediate image is selected from a set of preset reference images;
[0028] The target video is optimized based on the intermediate image to update the target video, and the updated target video is used to publish to the target platform.
[0029] In some embodiments, after performing video detection on the target video based on the video content features to obtain video detection data, the method further includes:
[0030] Based on the video detection data, abnormal images are extracted from the target video;
[0031] The total number of abnormal images is obtained by counting the number of abnormal images.
[0032] If the total number of abnormal images exceeds a preset image number threshold, the target video will be returned to the target object.
[0033] In some embodiments, acquiring the target video includes:
[0034] Retrieve the original video uploaded by the target object;
[0035] The original video is resized according to preset size parameters to obtain the target video.
[0036] To achieve the above objectives, a second aspect of this application provides a video detection apparatus, the apparatus comprising:
[0037] The video acquisition module is used to acquire the target video;
[0038] The image extraction module is used to extract video frames from the target video to obtain target video frame images of the target video;
[0039] The image segmentation module is used to segment the target video frame image to obtain multiple video region images and image position data of each video region image;
[0040] An image feature extraction module is used to extract features from the video region image and the image location data to obtain the image content features of each target video frame image;
[0041] The feature fusion module is used to fuse multiple image content features to obtain the video content features of the target video;
[0042] The video detection module is used to perform video detection on the target video based on the video content features to obtain video detection data, which is used to characterize the time points when abnormal images appear in the target video.
[0043] To achieve the above objectives, a third aspect of the present application provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described in the first aspect.
[0044] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0045] The video detection method, device, electronic equipment, and storage medium proposed in this application acquire a target video; extract video frames from the target video to obtain target video frame images, which facilitates the acquisition of target video frame images and improves video detection efficiency. Furthermore, the target video frame images are segmented to obtain multiple video region images and image position data for each video region image; features are extracted from the video region images and image position data to obtain image content features for each target video frame image, which facilitates the extraction of content feature information from the video region images. Furthermore, feature fusion is performed on multiple image content features to obtain video content features of the target video. These video content features represent all video content information of the target video, enabling content understanding of the target video and allowing video detection based on video content features to determine whether the target video contains anomalies. Finally, video detection is performed on the target video based on video content features to obtain video detection data. This data is used to characterize the time points when abnormal images appear in the target video. This method enables rapid detection of the target video and allows for the implementation of corresponding solutions based on the video detection data to improve the video quality. This significantly improves the accuracy of video detection and facilitates the detection of non-compliance in insurance product recommendation videos published by insurance institutions or agents. This enhances the video quality and compliance of insurance product recommendation videos, thereby effectively improving transaction security. Attached Figure Description
[0046] Figure 1 This is a flowchart of the video detection method provided in the embodiments of this application;
[0047] Figure 2 yes Figure 1 The flowchart of step S101 in the text;
[0048] Figure 3 yes Figure 1 The flowchart of step S104 in the process;
[0049] Figure 4 yes Figure 1 The flowchart of step S105 in the process;
[0050] Figure 5 yes Figure 1 The flowchart of step S106 in the process;
[0051] Figure 6 This is another flowchart of the video detection method provided in the embodiments of this application;
[0052] Figure 7 This is another flowchart of the video detection method provided in the embodiments of this application;
[0053] Figure 8 This is a schematic diagram of the structure of the video detection device provided in the embodiments of this application;
[0054] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0056] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0058] First, let's analyze some of the terms used in this application:
[0059] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0060] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, intent recognition, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.
[0061] Long Short-Term Memory (LSTM) networks are a type of recurrent neural network specifically designed to address the long-term dependency problem inherent in general Recurrent Neural Networks (RNNs). All RNNs have a chain-like structure of repeating neural network modules. In a standard RNN, this repeating structure has only a very simple configuration, such as a single tanh layer. LSTM is a type of neural network containing LSTM blocks or other similar structures. In literature and other resources, LSTM blocks may be described as intelligent network units because they can memorize values of indefinite duration. Each block contains a gate that determines whether the input is important enough to be memorized and whether it can be output.
[0062] Recurrent Neural Networks (RNNs): RNNs are a type of recursive neural network that takes sequence data as input, recurses along the direction of sequence evolution, and connects all nodes (recurrent units) in a chain-like manner. The core of an RNN is a directed graph. Elements connected in a chain-like manner in the unfolded directed graph are called recurrent units (RNN cells). Typically, the chain-like connections formed by recurrent units are analogous to hidden layers in feedforward neural networks. However, in different discussions, an "RNN layer" may refer to the recurrent units at a single time step or all recurrent units; therefore, for general introduction, the concept of "hidden layer" is avoided here. Given learning data X = {X1, X2, ..., Xτ} input in sequence, the unfolded length of the RNN is τ. The sequence to be processed is usually a time series, where the direction of sequence evolution is called a "time step".
[0063] Gated Recurrent Unit (GRU): A type of Recurrent Neural Network (RNN), and like LSTM (Long-Short Term Memory), it was proposed to solve problems related to long-term memory and gradients in backpropagation.
[0064] The Transformer model, similar to the Attention model, also employs an encoder-decoder architecture. However, its structure is more complex than Attention, typically involving multiple stacked encoders and decoder layers. The encoder consists of a self-attention layer and a feedforward neural network. Self-attention helps the current node focus beyond just the current word, allowing it to grasp the semantic context. The decoder includes not only the self-attention layer and the feedforward neural network but also an attention layer positioned between them. This attention layer helps the current node focus on the most important information at hand.
[0065] Transformer Layer: The neural network consists of an embedding layer (which can be called an input embedding layer) and at least one transformer layer. There can be N transformer layers (N is a positive integer). The embedding layer includes an input embedding layer and a positional encoding layer. In the input embedding layer, word embeddings are performed on each word in the current input to obtain word embedding vectors. In the positional encoding layer, the position of each word in the current input is obtained, and a position vector is generated for each word. Each transformer layer includes sequentially adjacent attention layers, add and normalize layers, feedforward layers, and add and normalize layers. In the input embedding layer, the current input is embedded to obtain multiple feature vectors. In the attention layer, P input vectors are obtained from the layer above the transformer layer. Using any first input vector among these P vectors as the center, intermediate vectors are obtained based on the correlation between each input vector within a predefined attention window and that first input vector. This process determines P intermediate vectors corresponding to the P input vectors. In the pooling layer, the P intermediate vectors are merged into Q output vectors, where at least one of the output vectors obtained from the last transformer layer is used as the feature representation of the current input. In the embedding layer, the current input (which can be text input, such as a paragraph or a sentence; the text can be Chinese / English or other languages) is embedded to obtain multiple feature vectors. After obtaining the current input, the embedding layer can embed each word in the current input to obtain the feature vectors of each word.
[0066] Attention mechanism: The attention mechanism enables neural networks to focus on a subset of their input (or features), selecting specific inputs. It can be applied to any type of input regardless of its shape. In situations with limited computational power, the attention mechanism is a primary resource allocation scheme for addressing information overload problems, allocating computational resources to more important tasks.
[0067] Encoder: Transforms an input sequence into a fixed-length vector.
[0068] Decoding: This involves transforming a previously generated fixed vector into an output sequence; the input sequence can be text, speech, image, or video; the output sequence can be text or image.
[0069] Softmax function: The Softmax function is a normalization exponential function that can "compress" a K-dimensional vector z containing arbitrary real numbers into another K-dimensional real vector σ(z), such that each element is in the range (0,1) and the sum of all elements is 1. This function is often used in multi-class classification problems.
[0070] With the development of computer technology and artificial intelligence, traditional offline businesses are gradually migrating online, and this trend has become irreversible. For example, online shopping, live streaming, and online transactions benefit institutions such as banks and online merchants.
[0071] In the financial sector, online sales of insurance products are often conducted through audio and video recordings. To enhance transaction security, regulatory agencies have stringent requirements for the content of marketing videos for insurance products. They frequently conduct video inspections of product recommendation videos to identify any non-compliant elements and promptly instruct relevant insurance institutions or agents to revise the videos accordingly.
[0072] Current video detection methods often rely on manual inspection to discover anomalies in videos. This method carries a significant risk of false positives, affecting the accuracy of video detection. Therefore, improving the accuracy of video detection has become an urgent technical problem to be solved.
[0073] Based on this, embodiments of this application provide a video detection method, a video detection device, an electronic device, and a storage medium, aiming to improve the accuracy of video detection.
[0074] The video detection method, video detection device, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the video detection method in this application is described.
[0075] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0076] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0077] The video detection method provided in this application relates to the field of artificial intelligence technology. The video detection method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the video detection method, but is not limited to the above forms.
[0078] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0079] Figure 1 This is an optional flowchart of the video detection method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S106.
[0080] Step S101: Obtain the target video;
[0081] Step S102: Extract video frames from the target video to obtain the target video frame image;
[0082] Step S103: Segment the target video frame image to obtain multiple video region images and image position data of each video region image;
[0083] Step S104: Extract features from the video region image and image location data to obtain the image content features of each target video frame image;
[0084] Step S105: Perform feature fusion on multiple image content features to obtain the video content features of the target video;
[0085] Step S106: Perform video detection on the target video based on video content features to obtain video detection data. The video detection data is used to characterize the time points when abnormal images appear in the target video.
[0086] Steps S101 to S106 of this embodiment involve acquiring a target video; extracting video frames from the target video to obtain target video frame images, which facilitates the acquisition of target video frame images and improves video detection efficiency. Further, the target video frame images are segmented to obtain multiple video region images and image position data for each video region image; and feature extraction is performed on the video region images and image position data to obtain image content features for each target video frame image, which facilitates the extraction of content feature information from the video region images. Furthermore, feature fusion is performed on multiple image content features to obtain video content features of the target video. These video content features represent all video content information of the target video, enabling content understanding of the target video and allowing video detection based on video content features to determine whether the target video contains anomalies. Finally, video detection is performed on the target video based on video content features to obtain video detection data. The video detection data is used to characterize the time points when abnormal images appear in the target video. This method can achieve rapid detection of the target video and take corresponding measures based on the video detection data to improve the video quality of the target video, thus significantly improving the accuracy of video detection.
[0087] Please see Figure 2 In some embodiments, step S101 may include, but is not limited to, steps S201 to S202:
[0088] Step S201: Obtain the original video uploaded by the target object;
[0089] Step S202: The original video is resized according to preset size parameters to obtain the target video.
[0090] In step S201 of some embodiments, the original video uploaded by the target object can be crawled based on a preset computer program or through a web crawler. The target object can be staff from various business fields, internet users, etc., without limitation. The original video may include product introductions, knowledge popularization, etc., filmed by the target object. Taking the insurance field as an example, the target object can be an insurance agent, and the original video can be a product introduction of an insurance product filmed by the insurance agent. The original video contains descriptions of the insurance product's type, cost, applicable population, etc. In this scenario, it is necessary to detect whether the insurance agent in the original video is properly dressed, whether they have left the filming frame, and whether the audio and video of the original video are synchronized.
[0091] In step S202 of some embodiments, the preset size parameters can be set according to actual conditions and are not limited. For example, the size parameter is 1920*1080, which means adjusting the resolution of the original video to 1920*1080 to obtain the target video.
[0092] Through the above steps S201 to S202, original videos of different sizes can be processed into target videos of the same size, which can effectively improve the video standardization of the target videos and help improve video quality.
[0093] In step S102 of some embodiments, a preset time interval can be used. After each preset time interval, a preset code program is used to extract video frames from the target video, extracting each video frame as the target video frame image. The preset time interval can be set according to actual conditions and is not limited. For example, a preset time interval of 5 seconds means that each frame of the target video is extracted every 5 seconds, and each extracted frame is used as a target video frame image. The time corresponding to each target video frame image is the time point of that target video frame image. This method can conveniently extract target video frame images, which helps improve video detection efficiency.
[0094] In step S103 of some embodiments, when segmenting the target video frame image, the target video frame image can be cut into nine small regions of equal area according to the image size of the target video frame image, and the image of each region can be used as a video region image.
[0095] Furthermore, based on the location of each video region image within the target video frame image, positional encoding is performed on each video region image to determine the location of each feature element within that image, thus obtaining the image position data for each video region image. Specifically, this positional encoding can be either absolute or relative, without restriction. When performing absolute encoding on the video region image, an absolute positional code for each feature element is generated using sine and cosine functions. This absolute positional code is then used to mark the position of each feature element, serving as its position label. When performing relative encoding on the video region image, the distance between every two feature elements is calculated. This distance can be Euclidean distance, Manhattan distance, etc. Relationship numbers are assigned to each pair of feature elements based on the magnitude of these distance values. These relationship numbers can be used to characterize the semantic order of the feature elements. This method facilitates the determination of the semantic order of each feature element within the video region image. This semantic order can serve as a basis for adjusting the position of image content features generated in subsequent feature extraction processes, resulting in better semantic coherence in the generated image content features.
[0096] Please see Figure 3 In some embodiments, step S104 may include, but is not limited to, steps S301 to S303:
[0097] Step S301: Input the video region image and image location data into a preset feature extraction model, which includes an encoder and a decoder;
[0098] Step S302: Based on the encoder, the video region image and image position data are encoded to obtain the image content latent vector;
[0099] Step S303: Extract content from the latent vector of image content based on the decoder to obtain image content features.
[0100] In step S301 of some embodiments, the preset feature extraction model can be built based on a transformer encoder, which includes an encoder and a decoder. Video region images and image location data can be directly input into the feature extraction model via scripts or similar programs. The feature extraction model then understands the video content information contained in the video region images and outputs corresponding content feature information.
[0101] In step S302 of some embodiments, the video region image is embedded based on the encoder to obtain initial image embedding features, and the image position data is embedded to obtain position coding features corresponding to each video region image. The initial image embedding features and position coding features are concatenated to obtain target image embedding features for each video region image.
[0102] Furthermore, the target image embedding features are layer-normalized based on the encoder so that the mean and variance of the target image embedding features meet the preset normalization conditions, thereby obtaining the first intermediate image features. The preset normalization conditions can be such that the mean of the target image embedding features is 1 and the variance is 0.
[0103] Furthermore, attention calculations are performed on the first intermediate image features based on the encoder's attention layer. The key matrix, value matrix, and query matrix for each first intermediate image feature are calculated. The key matrix, value matrix, and query matrix are then weighted using a softmax function to obtain the key image features for each target video frame. This attention calculation process strengthens the mapping of important feature information in the first intermediate image features while reducing the mapping of secondary feature information.
[0104] Furthermore, activation processing is performed on the key image features of each target video frame to standardize these key image features. Then, the activated key image features are mapped to a low-dimensional vector space to obtain the image content latent vector. Since low-dimensional features have better semantic representation, the image content latent vector can possess richer semantic content.
[0105] The encoder described above can easily extract content feature information from video region images, and perform importance analysis on the extracted content feature information to identify important video content information among these content feature information, thereby outputting an image content latent vector containing video content feature information.
[0106] In step S303 of some embodiments, the image content latent vector in vector form is converted into sequence form based on the decoder, thereby obtaining image content features that can characterize all content features of each target video frame image.
[0107] By fully understanding and analyzing the video content information contained in each video region image through the above steps S301 to S303, and obtaining the content feature information of each video region image, it is possible to obtain the complete image content information corresponding to each target video frame image, thereby improving the completeness and accuracy of video image information acquisition.
[0108] Please see Figure 4In some embodiments, step S105 may include, but is not limited to, steps S401 to S402:
[0109] Step S401: Obtain the image time of each target video frame image;
[0110] Step S402: Based on a preset temporal model and image time, multiple image content features are stitched together to obtain video content features.
[0111] In step S401 of some embodiments, when obtaining the image time of each target video frame image, the time point of each target video frame image recorded in the video frame extraction stage can be directly called and the time point is used as the image time of each target video frame image.
[0112] In step S402 of some embodiments, the preset temporal model can be constructed based on a recurrent neural network (RNN), a long short-term memory network (LSTM), or a gated recurrent unit (GRU). Taking the LSTM temporal model as an example, each image content feature is sequentially input into the temporal model according to the order of image time. The image information of each image content feature is extracted and spliced based on the input gate, forget gate, and output gate of the temporal model to obtain the video content features.
[0113] Through the above steps S401 to S402, multiple image content features can be conveniently spliced into a complete content feature according to the chronological order of the images. This complete content feature is then used as the video content feature, which can represent all the video content information of the target video. This method enables the understanding of the target video content, allowing video detection based on the video content feature to determine whether the target video has any anomalies, thus effectively improving the accuracy of video detection.
[0114] Please see Figure 5 In some embodiments, step S106 may include, but is not limited to, steps S501 to S504:
[0115] Step S501: Based on the preset video quality inspection model, video content features are detected to obtain quality inspection index data of the target video, wherein the quality inspection index data includes at least one actual quality inspection index.
[0116] Step S502: Compare the actual quality inspection indicators with the preset reference quality inspection indicators to obtain the abnormal images of the target video;
[0117] Step S503: Extract the image time of the abnormal image to obtain the time point;
[0118] Step S504: Based on the abnormal image and time point, obtain video detection data.
[0119] In step S501 of some embodiments, the preset video quality inspection model includes multiple quality inspection modules. Each quality inspection module is used to detect different video indicators of the target video. That is, the quality inspection modules can be set based on actual business needs. For example, the video quality inspection model includes three quality inspection modules: a first module, a second module, and a third module. The first module is used to detect whether the target video is synchronized with the audio based on video content features. The second module is used to detect whether there are scenes in the target video where people's clothing is not neat based on video content features. The third module is used to detect whether there are scenes in the target video where people's clothing is detached from the video frame based on video content features. Since the above quality inspection modules all detect whether there are anomalies in the target video, the quality inspection modules can be constructed based on softmax classifiers. The softmax classifier of each quality inspection module performs binary classification video detection on the video content features to obtain the actual quality inspection indicators corresponding to each quality inspection module. All the actual quality inspection indicators are integrated to obtain the quality inspection indicator data of the target video.
[0120] In step S502 of some embodiments, the preset reference quality inspection index can be the detection result displayed when each index is normal. For example, when the quality inspection content is to detect whether there are scenes of people with disheveled clothing in the target video, the reference quality inspection index is no scenes of people with disheveled clothing. Therefore, the actual quality inspection index and the preset reference quality inspection index can be compared. If the actual quality inspection index and the reference quality inspection index are inconsistent, it is considered that there is an index anomaly in the target video, and the corresponding abnormal image is extracted from the target video based on the abnormal actual quality inspection index.
[0121] In step S503 of some embodiments, when obtaining the time point of the abnormal image, the time point of each target video frame image recorded in the video frame extraction stage can be directly called, and the image time corresponding to the abnormal image can be used as the time point of the abnormal image.
[0122] In step S504 of some embodiments, each abnormal image and its time point are combined to obtain multiple abnormal images with time tags, and the abnormality of each abnormal image is described in text. Each abnormal image and its text description are used as a detection package. Multiple detection packages are integrated to obtain the final video detection data. The video detection data includes detection packages, and each detection package includes an abnormal image and the corresponding time point and abnormality description.
[0123] Through the above steps S501 to S504, it is possible to determine whether there are abnormal images in the target video based on index comparison, which can achieve rapid detection of the target video and take corresponding solutions based on the video detection data to improve the video quality of the target video and improve the accuracy of video detection.
[0124] Please see Figure 6 After step S106 in some embodiments, the video detection method may include, but is not limited to, steps S601 to S603:
[0125] Step S601: Extract abnormal images from the target video based on video detection data;
[0126] Step S602: Based on the image content of the abnormal image, select an intermediate image from a set of multiple reference images;
[0127] Step S603: Optimize the target video based on the intermediate image to update the target video. The updated target video is then published to the target platform.
[0128] In step S601 of some embodiments, after obtaining the video detection data, the abnormal images are extracted from the target video based on the time points of the abnormal images contained in the video detection data.
[0129] In step S602 of some embodiments, the image content of the abnormal image can be obtained based on the image content features of the aforementioned target video frame image. That is, the image content features of the target video frame image at the same time point as the abnormal image are called, and text is generated from these image content features to obtain image content text. Further, keywords are extracted from the image content text using methods such as named entity algorithms to obtain content keywords. Intermediate images that match the semantic content of the content keywords are retrieved and filtered from the reference images using these content keywords.
[0130] In step S603 of some embodiments, when optimizing the target video based on the intermediate image, the intermediate image can be used to replace the abnormal image in the target video, thereby updating the content of the target video, eliminating the abnormal image in the target video, so that the updated target video meets the publishing requirements, and the updated target video is published to the target platform. The target platform can be different social platforms or websites, forums, financial trading platforms, banking business processing platforms, etc., without limitation.
[0131] Through the above steps S601 to S603, the target video can be fine-tuned using intermediate images, thereby updating the content of the target video, eliminating abnormal images in the target video, and making the target video meet the publishing requirements.
[0132] Please see Figure 7 After step S106 in some embodiments, the video detection method may include, but is not limited to, steps S701 to S703:
[0133] Step S701: Extract abnormal images from the target video based on video detection data;
[0134] Step S702: Count the number of abnormal images to obtain the total number of abnormal images;
[0135] Step S703: If the total number of abnormal images is greater than the preset image number threshold, the target video is returned to the target object.
[0136] In step S701 of some embodiments, after obtaining the video detection data, the abnormal images are extracted from the target video based on the time points of the abnormal images contained in the video detection data.
[0137] In step S702 of some embodiments, the total number of abnormal images is obtained by performing a count on all extracted abnormal images based on the sum function or other statistical functions.
[0138] In step S703 of some embodiments, if the total number of abnormal images is greater than the image number threshold, it indicates that the target video needs a lot of modification. Modifying the target video requires a lot of manpower and time. Therefore, the target video is directly returned to the target object so that the target object can re-record a new video based on the text description of the abnormal images and then upload it. The specific value of the image number threshold can be set according to actual business needs.
[0139] Through the above steps S701 to S703, when there are many anomalies in the target video, the target video can be returned to the target object for re-recording. This can effectively save the time cost of modifying the target video, and also enable the target object to determine the cause of the recording error based on the received abnormal image and text descriptions. This helps to reduce the possibility of anomalies in subsequent video recordings by the target object and improves the standardization of the target object's video recording operation.
[0140] The video detection method of this application embodiment acquires a target video; extracts video frames from the target video to obtain target video frame images, which facilitates the acquisition of target video frame images and improves video detection efficiency. Further, the target video frame images are segmented to obtain multiple video region images and image position data for each video region image; and feature extraction is performed on the video region images and image position data to obtain image content features for each target video frame image, which facilitates the extraction of content feature information from the video region images. Further, feature fusion is performed on multiple image content features to obtain video content features of the target video. These video content features represent all video content information of the target video, enabling content understanding of the target video and allowing video detection based on video content features to determine whether the target video contains anomalies. Finally, video detection is performed on the target video based on the video content features to obtain video detection data. The video detection data is used to characterize the time points when abnormal images appear in the target video. This method enables rapid detection of the target video, and corresponding solutions can be taken based on the video detection data to improve the video quality of the target video, thus significantly improving the accuracy of video detection. Meanwhile, the video detection method in this application combines the excellent image understanding ability of the transformer model with the continuous understanding ability of the temporal model for temporal data. It can enhance the understanding of the video content of the target video as a whole, effectively improve the detection accuracy of anomalies in the video, and thus more easily detect whether there are any non-compliant parts in the recommended videos of insurance products released by relevant insurance institutions or insurance agents. This can improve the video quality and compliance of the recommended videos of insurance products, thereby effectively improving transaction security.
[0141] Please see Figure 8 This application also provides a video detection device that can implement the above-described video detection method. The device includes:
[0142] Video acquisition module 801 is used to acquire the target video;
[0143] Image extraction module 802 is used to extract video frames from the target video to obtain the target video frame image of the target video;
[0144] The image segmentation module 803 is used to segment the target video frame image to obtain multiple video region images and image position data of each video region image;
[0145] The image feature extraction module 804 is used to extract features from video region images and image location data to obtain the image content features of each target video frame image;
[0146] The feature fusion module 805 is used to fuse multiple image content features to obtain the video content features of the target video.
[0147] The video detection module 806 is used to perform video detection on the target video based on video content features to obtain video detection data. The video detection data is used to characterize the time points when abnormal images appear in the target video.
[0148] In some embodiments, the image extraction module 802 includes:
[0149] The input unit is used to input video region images and image location data into a preset feature extraction model, which includes an encoder and a decoder;
[0150] The encoding unit is used to encode video region images and image location data based on the encoder to obtain image content latent vectors;
[0151] The content extraction unit is used to extract image content features based on the latent vector of the decoder.
[0152] In some embodiments, the feature fusion module 805 includes:
[0153] The timing acquisition unit is used to acquire the image timing of each target video frame.
[0154] The stitching unit is used to stitch together multiple image content features based on a preset temporal model and image time to obtain video content features.
[0155] In some embodiments, the video detection module 806 includes:
[0156] The detection unit is used to perform video detection on video content features based on a preset video quality inspection model to obtain quality inspection index data of the target video, wherein the quality inspection index data includes at least one actual quality inspection index.
[0157] The comparison unit is used to compare the actual quality inspection indicators with the preset reference quality inspection indicators to obtain abnormal images of the target video.
[0158] The time extraction unit is used to extract the image time of abnormal images to obtain time points;
[0159] The detection data generation unit is used to obtain video detection data based on abnormal images and time points.
[0160] In some embodiments, the video detection device further includes a video optimization module, specifically comprising:
[0161] The first image extraction unit is used to extract abnormal images from the target video based on video detection data;
[0162] The image filtering unit is used to filter out intermediate images from a set of preset reference images based on the image content of the abnormal images.
[0163] The content optimization unit is used to optimize the target video based on the intermediate image to update the target video. The updated target video is then published to the target platform.
[0164] In some embodiments, the video detection device further includes a video playback module, specifically comprising:
[0165] The second image extraction unit is used to extract abnormal images from the target video based on video detection data;
[0166] The image count unit is used to count the number of abnormal images and obtain the total number of abnormal images.
[0167] The rollback unit is used to roll back the target video to the target object if the total number of abnormal images exceeds a preset image number threshold.
[0168] In some embodiments, the video acquisition module 801 includes:
[0169] The video acquisition unit is used to acquire the original video uploaded by the target object;
[0170] The size transformation unit is used to transform the original video according to preset size parameters to obtain the target video.
[0171] The specific implementation of this video detection device is basically the same as the specific implementation of the video detection method described above, and will not be repeated here.
[0172] This application also provides an electronic device, which includes: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communication between the processor and the memory. When the program is executed by the processor, it implements the aforementioned video detection method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0173] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0174] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0175] The memory 902 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the video detection method of the embodiments of this application.
[0176] The input / output interface 903 is used to implement information input and output;
[0177] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0178] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);
[0179] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.
[0180] This application also provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the video detection method described above.
[0181] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0182] The video detection method, video detection device, electronic device, and computer-readable storage medium provided in this application acquire a target video; extract video frames from the target video to obtain target video frame images, which can conveniently obtain target video frame images and improve video detection efficiency. Further, the target video frame images are segmented to obtain multiple video region images and image position data for each video region image; and feature extraction is performed on the video region images and image position data to obtain image content features for each target video frame image, which can conveniently extract content feature information from the video region images. Further, feature fusion is performed on multiple image content features to obtain video content features of the target video. These video content features can represent all video content information of the target video. This method enables content understanding of the target video, allowing video detection based on video content features to determine whether the target video contains anomalies. Finally, video detection is performed on the target video based on video content features to obtain video detection data. This data is used to characterize the time points when abnormal images appear in the target video. This method enables rapid detection of the target video, and corresponding solutions can be taken based on the video detection data to improve the video quality and significantly enhance the accuracy of video detection. Furthermore, the video detection method in this embodiment combines the excellent image understanding capabilities of the transformer model with the continuous understanding of time-series data by the temporal model. This strengthens the overall understanding of the target video's content and effectively improves the accuracy of detecting anomalies in the video. Consequently, it more easily detects whether the recommended videos of insurance products published by relevant insurance institutions or agents have any non-compliance, improving the video quality and compliance of recommended insurance product videos and thus effectively enhancing transaction security.
[0183] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0184] It will be understood by those skilled in the art that Figure 1-7 The technical solutions shown do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0185] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0186] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0187] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0188] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0189] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0190] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0191] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0192] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0193] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A video detection method, characterized by, The method comprises: acquiring a target video; performing video frame extraction on the target video to obtain a target video frame image of the target video; performing segmentation processing on the target video frame image to obtain a plurality of video region images and image position data of each video region image; performing feature extraction on the video region image and the image position data to obtain image content features of each target video frame image; performing feature fusion on a plurality of image content features to obtain video content features of the target video; performing video detection on the target video based on the video content features to obtain video detection data, wherein the video detection data is used to represent a time point at which an abnormal image appears in the target video; the performing feature fusion on a plurality of image content features to obtain video content features of the target video comprises: acquiring an image time of each target video frame image; inputting a plurality of image content features into a preset time sequence model in sequence according to the image time; extracting and splicing image information of each image content feature based on an input gate, a forgetting gate and an output gate of the time sequence model to obtain the video content features; the performing video detection on the target video based on the video content features to obtain video detection data comprises: performing video detection on the video content features based on a preset video quality detection model to obtain quality detection index data of the target video, wherein the quality detection index data comprises at least one actual quality detection index; comparing the actual quality detection index with a preset reference quality detection index to obtain an abnormal image of the target video; extracting an image time of the abnormal image as the time point; combining the abnormal image and the time point to obtain a plurality of abnormal images with time labels, and describing an abnormal situation of each abnormal image in words; taking each abnormal image and the word description as a detection package, integrating a plurality of detection packages to obtain video detection data.
2. The video detection method of claim 1, wherein, the performing feature extraction on the video region image and the image position data to obtain image content features of each target video frame image comprises: inputting the video region image and the image position data into a preset feature extraction model, wherein the feature extraction model comprises an encoder and a decoder; performing encoding processing on the video region image and the image position data based on the encoder to obtain an image content hidden vector; performing content extraction on the image content hidden vector based on the decoder to obtain the image content features.
3. The video detection method of claim 1, wherein, after the performing video detection on the target video based on the video content features to obtain video detection data, the method further comprises: extracting an abnormal image from the target video based on the video detection data; screening an intermediate image from a plurality of reference images according to image content of the abnormal image; performing content optimization on the target video based on the intermediate image to update the target video, and the updated target video is used to be published to a target platform.
4. The video detection method of claim 1, wherein, After the video detection based on the video content feature is performed on the target video to obtain video detection data, the method further comprises: extracting an abnormal image from the target video based on the video detection data; counting the number of the abnormal image to obtain an abnormal image total number; if the abnormal image total number is greater than a preset image number threshold, returning the target video to the target object.
5. The video detection method of any one of claims 1 to 4, characterized in that, The target video is obtained by: obtaining an original video uploaded by a target object; performing size transformation on the original video according to a preset size parameter to obtain the target video.
6. A video detection apparatus characterized by comprising: The device comprises: a video obtaining module for obtaining a target video; an image extracting module for performing video frame extraction on the target video to obtain target video frame images of the target video; an image segmenting module for performing segmentation processing on the target video frame images to obtain a plurality of video region images and image position data of each video region image; an image feature extracting module for performing feature extraction on the video region images and the image position data to obtain image content features of each target video frame image; a feature fusing module for fusing a plurality of the image content features to obtain a video content feature of the target video; a video detecting module for performing video detection on the target video based on the video content feature to obtain video detection data, wherein the video detection data is used to represent a time point at which an abnormal image appears in the target video; the fusing of a plurality of the image content features to obtain the video content feature of the target video comprises: obtaining an image time of each target video frame image; inputting a plurality of the image content features into a preset time sequence model in sequence according to the image time; extracting and splicing image information of each image content feature based on an input gate, a forgetting gate and an output gate of the time sequence model to obtain the video content feature; the video detection based on the video content feature on the target video to obtain video detection data comprises: performing video detection on the video content feature based on a preset video quality inspection model to obtain quality inspection index data of the target video, wherein the quality inspection index data comprises at least one actual quality inspection index; comparing the actual quality inspection index with a preset reference quality inspection index to obtain an abnormal image of the target video; extracting an image time of the abnormal image as the time point; combining the abnormal image and the time point to obtain a plurality of abnormal images with time labels, and describing an abnormal situation of each abnormal image in words; taking each abnormal image and the word description as a detection package, integrating a plurality of detection packages to obtain video detection data.
7. An electronic device, comprising: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the video detection method of any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 7. The computer program is executed by the processor to implement the video detection method of any one of claims 1 to 5.
Citation Information
Patent Citations
Image data processing method and device, computer equipment and storage medium
CN111914811A