Key-frame-based Visual Semantic Detection Method and System
By calculating the histogram feature values of video frame images, performing video sampling and keyframe recognition, combining dimensional transformation and graphical analysis, the problem of video semantic information detection in the prior art is solved, and efficient compliance judgment of video data streams is achieved.
Patent Information
- Application Number
- CN202111609817.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-12-27
AI Technical Summary
It is difficult for prior art to accurately identify and parse semantic information contained in videos, especially at keyframes.
By calculating the eigenvalue of the histogram of the frame image in the gradient direction, obtaining the eigenvalue jump points, performing video sampling, vectorization and convolution operations to obtain the keyframe, dimensional conversion and plane division at this keyframe, the frame images of the same plane are graphically analyzed, and whether the frame images and data flow are compliant.
It realizes effective sampling and keyframe identification of video data streams, can accurately judge the compliance of frame images and data streams, and improves the detection accuracy of semantic information in videos.
Smart Images

Figure CN115527138B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network multimedia, and particularly to a visual semantic detection method and system based on key frames. Background Art
[0002] There are many videos in the existing network that do not contain or contain a small amount of semantic information, and it is difficult to accurately identify and analyze them according to the existing semantic detection methods. However, each of these videos will have several key frames, and the key frame is the video frame that best represents its content. If the key frames are utilized well for video detection, it has become a key research direction for those skilled in the art.
[0003] Therefore, there is an urgent need for a targeted visual semantic detection method and system based on key frames. Summary of the Invention
[0004] The purpose of the present invention is to provide a visual semantic detection method and system based on key frames. By calculating the eigenvalue of the histogram of the frame image in the gradient direction, obtaining the eigenvalue jump point, performing video sampling, vectorization, and convolution operation on the video data stream to obtain the key frame, performing dimensional conversion and plane division at the key frame, and performing graphic analysis on the frame images in the same plane to determine whether the frame image and the data stream are compliant.
[0005] In the first aspect, this application provides a visual semantic detection method based on key frames, and the method includes:
[0006] Obtain the video data stream, calculate the eigenvalue of the histogram of each frame image in the gradient direction. When the difference between the eigenvalues of the frames is greater than a preset threshold, perform video sampling on the video data stream. The video sampling uses a basic filtering unit to extract the first image feature, vectorize the first image feature, determine several key points according to the size of the vectorized eigenvalue, perform clustering operation on the several key points, and map them to the corresponding visual dictionary for quantization;
[0007] Input the quantized result into the N-layer convolution unit, obtain the first intermediate result according to the output result of the N-layer convolution unit, perform smoothing processing on the first intermediate result to obtain a high-dimensional image carrying boundary and regional local features, and define the frame of the high-dimensional image as the key frame;
[0008] Vectorize the image of the key frame and then perform dimensional conversion, input the high-dimensional sample set, call the classifier to classify the high-dimensional sample set. The pairwise eigenvalues of the set within a preset difference range are grouped into one group, calculate the average eigenvalue of the group, and divide the different groups with the difference between the average eigenvalues of the groups greater than the threshold into different planes;
[0009] Input frame images on the same plane into a graphic analysis model, identify the object information contained in the frame images, obtain the key object features, and determine whether the frame images include non-compliant graphic content;
[0010] If the graphic content included in the frame image is non-compliant, delete the segment of video data stream.
[0011] Combined with the first aspect, in the first possible implementation manner of the first aspect, the eigenvalue of the histogram in the gradient direction includes detecting the change in the gray centroid position of the image.
[0012] Combined with the first aspect, in the second possible implementation manner of the first aspect, the step of calling the classifier to classify the high-dimensional sample set includes the inner product operation between the input vector and the high-dimensional sample set.
[0013] Combined with the first aspect, in the third possible implementation manner of the first aspect, the kernels of both the semantic analysis model and the graphic analysis model use neural network models.
[0014] In a second aspect, the present application provides a key-frame-based visual semantic detection system, which includes a processor and a memory:
[0015] The memory is used to store program codes and transmit the program codes to the processor;
[0016] The processor is configured to execute the method according to any one of the four possible aspects described in the first aspect according to the instructions in the program code.
[0017] In a third aspect, the present application provides a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the method according to any one of the four possible aspects described in the first aspect.
[0018] The present invention provides a key-frame-based visual semantic detection method and system. By calculating the eigenvalue of the histogram of the frame image in the gradient direction, obtaining the eigenvalue jump points, performing video sampling, vectorization and convolution operations to obtain key frames, performing dimensionality conversion and plane division at the key frames, and performing graphic analysis on the frame images on the same plane to determine whether the frame images and data streams are compliant. Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a flowchart of the method of the present invention. Detailed implementation mode
[0021] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making the protection scope of the present invention more clearly defined.
[0022] Figure 1 It is a flowchart of the visual semantic detection method based on key frames provided by this application, including:
[0023] Obtain the video data stream, calculate the eigenvalue of the histogram of each frame image in the gradient direction. When the difference between the eigenvalues between the frames is greater than a preset threshold, perform video sampling on the video data stream. The video sampling uses a basic filtering unit to extract the first image feature, vectorize the first image feature, determine several key points according to the size of the vectorized eigenvalue, perform clustering operations on the several key points, and map them to the corresponding visual dictionary for quantization;
[0024] Input the quantized result into the N-layer convolution unit. According to the output result of the N-layer convolution unit, obtain the first intermediate result, perform smoothing processing on the first intermediate result to obtain a high-dimensional image carrying boundary and regional local features, and define the frame of the high-dimensional image as a key frame;
[0025] After vectorizing the image of the key frame, perform dimensionality conversion, input the high-dimensional sample set, call the classifier to classify the high-dimensional sample set. The pairwise eigenvalues of the set within a preset difference range are grouped into one group, calculate the average eigenvalue of the group, and divide the different groups with the difference between the average eigenvalues of the groups greater than the threshold into different planes;
[0026] Input the frame images on the same plane into the graphic analysis model, identify the object information contained in the frame image, obtain the key object features, and determine whether the frame image includes non-compliant graphic content;
[0027] If the graphic content included in the frame image is non-compliant, delete the segment of the video data stream.
[0028] In some preferred embodiments, the eigenvalue of the histogram in the gradient direction includes the change in the gray centroid position of the detected image.
[0029] In some preferred embodiments, the calling of the classifier to classify the high-dimensional sample set includes the inner product operation between the input vector and the high-dimensional sample set.
[0030] In some preferred embodiments, the kernels of both the semantic analysis model and the graphic analysis model use neural network models.
[0031] This application provides a visual semantic detection system based on key frames. The system includes a processor and a memory:
[0032] The memory is used to store program code and transmit the program code to the processor;
[0033] The processor is used to execute the method described in any one of all the embodiments of the first aspect according to the instructions in the program code.
[0034] This application provides a computer-readable storage medium. The computer-readable storage medium is used to store program code, and the program code is used to execute the method described in any one of all the embodiments of the first aspect.
[0035] In specific implementation, the present invention also provides a computer storage medium. The computer storage medium can store a program, and when the program is executed, it may include some or all of the steps in various embodiments of the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (abbreviation: ROM), or a random access memory (abbreviation: RAM), etc.
[0036] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0037] For the same and similar parts between the embodiments of this specification, reference can be made to each other. In particular, for the embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description in the method embodiments.
[0038] The above-described embodiments of the present invention do not constitute a limitation on the protection scope of the present invention.
Claims
1. A key-frame based visual semantic detection method, characterized in that The method includes: Obtaining a video data stream, calculating the eigenvalue of the histogram of each frame image in the gradient direction. When the difference between the eigenvalues between frames is greater than a preset threshold, video sampling is performed on the video data stream. The video sampling uses a basic filtering unit to extract the first image feature, vectorize the first image feature, determine a number of key points according to the magnitude of the vectorized eigenvalue, perform a clustering operation on the number of key points, and map them to the corresponding visual dictionary for quantization; Inputting the quantized result into an N-layer convolutional unit, obtaining a first intermediate result according to the output result of the N-layer convolutional unit, performing smoothing processing on the first intermediate result to obtain a high-dimensional image carrying boundary and regional local features, and defining the frame of the high-dimensional image as a key frame; After vectorizing the image of the key frame, performing dimensionality conversion, inputting a high-dimensional sample set, and calling a classifier to classify the high-dimensional sample set. Samples with pairwise eigenvalues within a preset difference range are grouped into one group, and the average eigenvalue of the group is calculated. Different groups with an average eigenvalue difference greater than the threshold between groups are divided into different planes; Inputting the frame images of the same plane into a graphic analysis model, identifying the object information contained in the frame image, obtaining key object features, and determining whether the frame image includes non-compliant graphic content; If the graphic content included in the frame image is non-compliant, the above video data stream is deleted.
2. The method according to claim 1, wherein: The eigenvalue of the histogram in the gradient direction includes detecting the change in the gray centroid position of the image.
3. The method according to any one of claims 1-2, characterized in that: The step of calling a classifier to classify the high-dimensional sample set includes an inner product operation between the input vector and the high-dimensional sample set.
4. A key-frame based visual semantic detection system, characterized in that, The system includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the method described in any one of claims 1-3 according to the instructions in the program code.
5. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program code, and the program code is used to execute the method described in any one of claims 1-3.
Citation Information
Patent Citations
Video key frame extraction method, system and device based on multi-view features
CN110472484A
Visual inspection method and system, electronic equipment and medium
CN112991281A