Video stream detection method and device, equipment and medium
By performing keyframe screening and multi-dimensional data analysis on the video stream, the problem of recognition of AI face-changing video streams is solved, the accuracy and efficiency of detection are improved, and the security of the video stream is ensured.
Patent Information
- Application Number
- CN202510806547.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-08-08
AI Technical Summary
The existing technology cannot effectively identify video streams forged through AI face swap technology, resulting in threats to personal privacy data and property security.
By performing keyframe screening of the video stream, distortion detection and continuity detection, and combining multi-dimensional data analysis, we determine whether the video stream is an AI face-changing video.
It improves the accuracy and efficiency of video stream detection, reduces the number of video keyframes to be processed, and enhances the recognition ability of AI face-changing videos.
Smart Images

Figure CN120455780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of data processing technology and financial technology, and in particular to a video stream detection method, device, equipment and medium. Background Art
[0002] When a facial recognition device performs facial recognition, users must follow the device's instructions and perform corresponding actions to verify their identity. With the recent development of artificial intelligence (AI) technology, videos using artificial intelligence (AI) face-swapping have emerged to forge user identities and steal their personal data or assets. Therefore, it is necessary to monitor the video streams captured by facial recognition devices.
[0003] Currently, identity information can only be identified for video streams obtained by facial recognition devices that forge users by using user photos. Identity information cannot be identified for video streams obtained by facial recognition devices that forge users by using AI technology to swap faces. Summary of the Invention
[0004] The present invention provides a video stream detection method, device, equipment and medium to improve the accuracy of video stream detection.
[0005] In a first aspect, an embodiment of the present invention provides a video stream detection method, the method comprising:
[0006] Determine the video stream to be detected based on the obtained start verification time and end verification time;
[0007] Perform key frame screening on the video stream to be detected to determine a key frame sequence, where the key frame sequence includes at least one video key frame;
[0008] Performing distortion detection on the key frame sequence to determine the amount of distortion corresponding to the key frame sequence;
[0009] Perform continuity detection on the key frame sequence to determine the pixel position change data corresponding to the key frame sequence;
[0010] The video stream detection result is determined based on the distortion amount and pixel position change data.
[0011] In a second aspect, an embodiment of the present invention further provides a video stream detection device, the device comprising:
[0012] The video stream acquisition module is used to determine the video stream to be detected based on the acquired start verification time and end verification time;
[0013] A key frame screening module is used to screen the key frames of the video stream to be detected and determine a key frame sequence, where the key frame sequence includes at least one video key frame;
[0014] A distortion detection module is used to perform distortion detection on a key frame sequence and determine the amount of distortion corresponding to the key frame sequence;
[0015] A continuity detection module is used to perform continuity detection on the key frame sequence and determine the pixel position change data corresponding to the key frame sequence;
[0016] The video detection module is used to determine the video stream detection result based on the distortion amount and pixel position change data.
[0017] In a third aspect, an embodiment of the present invention further provides a video stream detection device, the video stream detection device comprising:
[0018] at least one processor; and
[0019] a memory communicatively connected to at least one processor; wherein,
[0020] The memory stores a computer program that can be executed by at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the video stream detection method according to any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions, which are used to enable a processor to implement the video stream detection method according to any embodiment of the present invention when executed.
[0022] The technical solution of the embodiment of the present invention improves the efficiency of video stream detection by screening key frames of the video stream to be detected, determining the key frame sequence, reducing the number of video key frames to be processed, and performing distortion detection and continuity detection on the key frame sequence, and detecting the video stream through multi-dimensional data, thereby improving the accuracy and comprehensiveness of video stream detection.
[0023] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0025] Figure 1 This is a flowchart of a video stream detection method provided according to the first embodiment of the present invention;
[0026] Figure 2 This is a flowchart of a video stream detection method provided according to the second embodiment of the present invention;
[0027] Figure 3 is a structural diagram of a video stream detection device provided according to an embodiment of the present invention;
[0028] Figure 4 It is a structural diagram of a video stream detection device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0031] In the technical solutions of the embodiments of the present invention, the acquisition, storage and application of the operating information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0032] Example 1
[0033] Figure 1 This is a flow chart of a video stream detection method provided in Embodiment 1 of the present invention. This embodiment of the present invention is applicable to situations where video stream detection is required. This method can be performed by a video stream detection device, which can be implemented in the form of hardware and / or software.
[0034] See also Figure 1 The video stream detection method shown includes:
[0035] S101. Determine a video stream to be detected according to the obtained start verification time and end verification time.
[0036] The start verification time may be the time when the video stream of the liveness detection action is started to be collected. The end verification time may be the time when the video stream of the liveness detection action is finished to be collected. The video stream to be detected may be the collected video stream to be subjected to liveness detection.
[0037] Specifically, liveness detection is an identity authentication technology that uses computer vision and deep learning algorithms to analyze the user's real-time actions, such as nodding, shaking the head, or turning the head left and right, and combines the dynamic changes in facial features to verify whether the user is a real living person and not a fake (such as a photo, video, or 3D model, etc.). There may be a liveness detection device, which is equipped with a camera. When the liveness detection device is ready to send an action prompt message, the current time of the liveness detection device is obtained, which is determined as the start verification time, and the camera starts to collect video streams. The preset duration is obtained, and the start verification time is added to the preset duration to determine the end verification time. The camera starts collecting video streams from the start verification time, and when the end verification time is reached, the video stream is no longer collected. The video stream corresponding to the start verification time and the end verification time is determined as the video stream to be detected, which is used to determine whether it is an AI face-changing video stream.
[0038] S102: Screen the key frames of the video stream to be detected and determine a key frame sequence, where the key frame sequence includes at least one video key frame.
[0039] The key frame sequence may be a set of video frames in a video stream, and the video key frame may be a video frame extracted from the video stream.
[0040] Specifically, the video stream to be detected is a time-series sequence of continuous video frames. Video frames containing user actions are screened from the video stream to be detected and identified as video key frames. There is at least one video key frame. Each video key frame is sorted in the order of acquisition time to generate a key frame sequence.
[0041] S103: Perform distortion detection on the key frame sequence to determine the distortion amount corresponding to the key frame sequence.
[0042] The distortion detection may be a face similarity detection between key frames in the key frame sequence, and the distortion number may be the number of distortion detection results in the key frame sequence that are distortion results.
[0043] Specifically, the process of AI face-changing generally includes steps such as recognition, replacement and fusion rendering. If the AI face-changing technology is mature, the algorithm is optimized and the computing power is sufficient, and the original video and the face-changing material are of high quality and have good feature matching, then the face-changed video may be able to achieve good fusion on most video frames. Only a very small number of video frames may have temporary facial inconsistencies due to factors such as complex scene changes or sudden changes in light. For example, in some scenes with fast motion or drastic changes in light, there will be some video frames with facial differences; in videos where facial posture or expression changes too quickly, the facial features of some frames may be inaccurately extracted, resulting in facial differences; or if the facial image is not well fused with the original video frame, the faces of some frames will be dissimilar. In this case, the probability that the video stream to be detected is a face-changing video is relatively high. Therefore, distortion detection is performed on each video key frame in the key frame sequence to detect whether there is facial dissimilarity between the video key frames in the key frame sequence, and the number of distorted video key frames is counted to determine the amount of distortion.
[0044] S104: Perform continuity detection on the key frame sequence to determine pixel position change data corresponding to the key frame sequence.
[0045] The continuity test may be a test of whether the content, motion, time, or visual perception of each key frame in the key frame sequence is continuous. The pixel position change data may be information describing the degree of natural pixel connection between each key frame in the key frame sequence.
[0046] Specifically, video continuity refers to the degree to which details such as facial movements, expressions, lighting, or textures in adjacent video keyframes flow naturally together after the face-swap video, as well as the coordination and consistency between the face-swapped area and the surrounding environment (such as head movement or background changes). Videos with poor continuity will exhibit problems such as facial jumps, broken expressions, or sudden changes in lighting and shadows, and can be identified as "fake," meaning face-swapped videos. For example, when the head turns from left to right, the rotation angle and pitch angle of the face after the face swap should be consistent with the original video character. Otherwise, there will be a sense of disharmony of "the head moves but the face does not move" or "the facial rotation trajectory suddenly changes". In this case, the video continuity is poor and it is identified as a face-swapped video. The amplitude of the corners of the mouth when smiling, the frequency and duration of blinking must conform to the timing rules of real expressions (such as blinking usually lasts 100-400 milliseconds). The gradual process of micro-expressions is not captured in the video, and there is a fault of "instantaneous switching of expressions". In this case, the video continuity is poor and it is identified as a face-swapped video. If the face of the video is in the sun, the face should have obvious light and dark transitions. If it is detected that the brightness of the face in the video is always uniform, or the light and shadow suddenly change, the video continuity is poor and it is identified as a face-swapped video. When the video continuity is poor, the degree of connection between the pixel position changes in each key frame of the video is poor, and the pixel position will suddenly change. The key frame sequence is tested for continuity to determine the pixel position change data corresponding to the key frame sequence.
[0047] S105: Determine the video stream detection result according to the distortion amount and pixel position change data.
[0048] Among them, the video stream detection result can be detection data of whether the video stream is a face-swapped video with a fake user.
[0049] Specifically, based on the similarities between the key frames in the corresponding key frame sequence of the video stream, the number of dissimilar video key frame sets is counted to determine the distortion level of the key frame sequence. Based on the distortion level, the distortion result of the video to be tested is determined. Pixel position change data between the key frames in the key frame sequence is detected in chronological order. The continuity result of the key frame sequence is determined. Based on the distortion and continuity results, the video stream detection result of the video to be tested is determined to determine whether the video stream to be tested is an AI face-swapped video stream.
[0050] The technical solution of the embodiment of the present invention performs key frame screening on the video stream to be detected, determines the key frame sequence, reduces the video key frames to be processed, and improves the efficiency of video stream detection; performs distortion detection and continuity detection on the key frame sequence, and can determine whether the video stream to be detected contains a large number of video key frames of dissimilar faces through distortion detection, and can determine whether the video key frames in the video stream to be detected can be naturally connected through continuity detection. Through multi-dimensional detection, the accuracy and comprehensiveness of video stream detection are improved.
[0051] Example 2
[0052] Figure 2 This is a flow chart of a video stream detection method provided in the second embodiment of the present invention. Based on the above embodiments, the embodiment of the present invention optimizes and improves the video stream detection operation.
[0053] Furthermore, "determining the video stream detection result based on the distortion amount and pixel position change data" is refined into "obtaining at least one distortion range and the distortion weight corresponding to each distortion range; comparing the distortion amount with each distortion range to determine the target range corresponding to the distortion amount and the target weight corresponding to the target range; calculating the video stream detection value based on the distortion amount, the target weight and the pixel position change data; obtaining the video detection threshold; determining the video stream detection result based on the comparison result between the video stream detection value and the video detection threshold" to improve the video stream detection operation.
[0054] It should be noted that for parts not described in detail in the embodiments of the present invention, reference may be made to the descriptions of other embodiments.
[0055] See also Figure 2 The video stream detection method shown includes:
[0056] S201. Determine a video stream to be detected according to the obtained start verification time and end verification time.
[0057] S202: Perform key frame screening on the video stream to be detected to determine a key frame sequence, where the key frame sequence includes at least one video key frame.
[0058] S203: Perform distortion detection on the key frame sequence to determine the distortion amount corresponding to the key frame sequence.
[0059] S204: Perform continuity detection on the key frame sequence to determine pixel position change data corresponding to the key frame sequence.
[0060] S205: Obtain at least one distortion range and a distortion weight corresponding to each distortion range.
[0061] The distortion range may be a data range of the distortion amount corresponding to the key frame sequence, and the distortion weight may be used to describe the degree of influence of the distortion range on the video detection result.
[0062] Specifically, based on industry experience, at least one distortion range and corresponding distortion weights are preset. The larger the distortion range to which the distortion quantity belongs, the greater the distortion quantity corresponding to the key frame sequence, and the greater the probability that the video stream to be detected corresponding to the key frame sequence is a face-swapped video stream; the smaller the distortion range to which the distortion quantity belongs, the smaller the distortion quantity corresponding to the key frame sequence, and the lower the probability that the video stream to be detected corresponding to the key frame sequence is a face-swapped video stream. The larger the distortion range, the greater the corresponding distortion weight, and the smaller the distortion range, the smaller the corresponding distortion weight.
[0063] S206: Compare the distortion amount with each distortion range to determine a target range corresponding to the distortion amount and a target weight corresponding to the target range.
[0064] Specifically, when performing distortion detection between key frames in a key frame sequence, the distortion amount corresponding to the key frame sequence is statistically calculated. The distortion amount is compared with each distortion range to determine the distortion range to which the distortion amount belongs. The distortion range corresponding to the distortion amount is determined as the target range, and the distortion weight corresponding to the target range is determined as the target weight.
[0065] S207: Calculate the video stream detection value according to the distortion amount, target weight and pixel position change data.
[0066] Specifically, the video stream detection value is calculated by multiplying the distortion amount, target weight, and pixel position change data. The target weight is used to adjust the influence of the distortion amount on the video stream detection value. When the face-swap video is continuous and the movement amplitude specified by the detection device is small, the pixel position change data value is very small. In this case, the distortion amount of the key frame sequence must be counted and adjusted using the target weight to determine the video stream detection value corresponding to the key frame sequence.
[0067] S208: Obtain a video detection threshold.
[0068] The video detection threshold may be a threshold corresponding to a preset video stream detection value.
[0069] Specifically, the video detection threshold is obtained in a manner including, but not limited to, mouse selection, keyboard input, or voice input, etc., which is not limited in the embodiment of the present invention.
[0070] S209: Determine a video stream detection result according to a comparison result between the video stream detection value and the video detection threshold.
[0071] Specifically, when the video stream detection value is greater than or equal to the video detection threshold, it indicates that the number of distortions in the video stream to be detected is large, the distortion weight is large, or the continuity between the video key frames is poor, and the video to be detected is detected as a face-changing video. When the video stream detection value is less than the video detection threshold, it indicates that the number of distortions in the video stream to be detected is small, the distortion weight is small, or the continuity between the video key frames is poor, and the video to be detected is detected as the user's real face video.
[0072] The embodiment of the present invention obtains a video stream detection value by calculating the distortion amount, target weight and pixel position change data, calculates the video stream detection value based on multi-dimensional data, improves the accuracy of the video stream detection value calculation, determines the video stream detection result based on the comparison result between the video stream detection value and the video detection threshold, and determines whether the video stream detection value is within the normal range through the video detection threshold, thereby determining whether the video to be detected is a face-changing video.
[0073] Optionally, key frame screening is performed on the video stream to be detected to determine a key frame sequence, including: obtaining a verification action, a moment step and a number of moment acquisitions; in response to identifying a verification action in the video stream to be detected, determining the moment data to be acquired, the moment data to be acquired including: the start moment of the action, the middle moment of the action and the end moment of the action; determining at least one acquisition moment based on the moment step, the number of moment acquisitions and the moment data to be acquired; key frame screening is performed on the video stream to be detected according to each acquisition moment to obtain at least one video key frame, and a key frame sequence is generated.
[0074] The verification action can be an action to be identified in the video to be detected. The time step can be the interval between adjacent video key frames. The number of time acquisitions can be the number of video key frames acquired in the left and right neighborhoods at a specified time.
[0075] Specifically, the verification action, moment step and moment acquisition number are obtained. The video stream to be detected is identified, and the recognition method can be a face tracking algorithm, etc. When the verification action is identified, the start time and end time of the verification action are determined. The middle moment between the start time and the end time of the action is selected and determined as the middle moment of the action. According to the moment step, the moment acquisition number and the start time of the action, at least one acquisition moment is determined; according to the moment step, the moment acquisition number and the middle moment of the action, at least one acquisition moment is determined; according to the moment step, the moment acquisition number and the end time of the action, at least one acquisition moment is determined; according to the moment step, the moment acquisition number and the end time of the action, key frame screening is performed on the video stream to be detected according to each acquisition moment, at least one video key frame is obtained, and a key frame sequence is generated. For example: the moment acquisition number is 4, the moment step is 1 second, and the action start time is 3:20, then the acquisition moments are determined to be 3.18, 3.19, 3.20, 3.21 and 3.22. In the verification video stream, capture the video key frames corresponding to times 3.18, 3.19, 3.20, 3.21, and 3.22. If the middle of the action is 3:50, the acquisition times are determined to be 3.48, 3.49, 3.50, 3.51, and 3.52. In the verification video stream, capture the video key frames corresponding to times 3.48, 3.49, 3.50, 3.51, and 3.52. If the middle of the action is 4:10, the acquisition times are determined to be 4.08, 4.09, 4.10, 4.11, and 4.12. In the verification video stream, capture the video key frames corresponding to times 4.08, 4.09, 4.10, 4.11, and 4.12. Generate a key frame sequence based on each video key frame.
[0076] By responding to the recognition of the verification action in the video stream to be detected, the data at the time to be collected is determined, and at least one collection time is determined according to the time step, the number of time collections and the data at the time to be collected; the key frames of the video stream to be detected are screened according to each collection time to obtain at least one video key frame, and a key frame sequence is generated to screen the video key frames, thereby reducing the number of video key frames to be processed and improving the efficiency of video stream detection.
[0077] Optionally, distortion detection is performed on the key frame sequence to determine the amount of distortion corresponding to the key frame sequence, including: combining each video key frame in the key frame sequence in pairs to obtain at least one image group; performing similarity calculation on the two video key frames in each image group to obtain an image similarity value corresponding to each image group; obtaining a similarity threshold; for each image group, comparing the image similarity value corresponding to the image group with the similarity threshold to obtain a distortion result of the image group; and according to the distortion result of each image group, obtaining the amount of distortion corresponding to the key frame sequence by statistics.
[0078] The image group may be a set of two video key frames whose similarity is to be compared. The image similarity value may be a similarity value between the two video key frames in the image group. The similarity threshold may be a preset threshold of the image similarity value.
[0079] Specifically, each video key frame in the key frame sequence is combined in pairs to obtain at least one image group; the two video key frames in each image group are similarly calculated to obtain the image similarity value corresponding to each image group. The similarity algorithm includes but is not limited to: structural similarity algorithm, cosine similarity algorithm or deep learning algorithm, etc., and the embodiment of the present invention does not limit this. A similarity threshold is obtained; for each image group, the image similarity value corresponding to the image group is compared with the similarity threshold. If the image similarity value corresponding to the image group is greater than or equal to the similarity threshold, the distortion result of the image group is undistorted; if the image similarity value corresponding to the image group is less than the similarity threshold, the distortion result of the image group is distorted. According to the distortion result of each image group, the number of distortion results corresponding to the key frame sequence is obtained by counting, and the distortion number of the key frame sequence is determined.
[0080] By calculating the similarity between the two video key frames in each image group, the image similarity value corresponding to each image group is obtained; for each image group, the image similarity value corresponding to the image group is compared with the similarity threshold to obtain the distortion result of the image group. The distortion number corresponding to the key frame sequence is statistically obtained to determine whether the faces of each key frame in the key frame sequence are consistent.
[0081] Optionally, a continuity check is performed on the key frame sequence to determine the pixel position change data corresponding to the key frame sequence, including: obtaining at least one trained continuity recognition network; randomly selecting a video key frame in the key frame sequence to determine its corresponding image resolution; screening each continuity recognition network according to the image resolution to determine the target recognition network; inputting the key frame sequence into the target recognition network for continuity check to determine the image pixel change data corresponding to each image key frame in the key frame sequence; and calculating the pixel position change data based on the image pixel change data corresponding to each image key frame.
[0082] The continuity recognition network may be a network that identifies the degree of connectivity between video frames in a video stream to be detected. The image resolution may be a resolution corresponding to a video keyframe. The image pixel change data may be information describing the degree of natural pixel connectivity between a subsequent video keyframe and a previous video keyframe in a keyframe sequence.
[0083] Specifically, at least one trained continuity recognition network is obtained, where different continuity recognition networks have different corresponding image resolution identifiers. A video keyframe is randomly selected in the keyframe sequence to determine its corresponding image resolution. Based on the matching of the image resolution with the image resolution identifiers corresponding to each continuity recognition network, the target recognition network is screened within each continuity recognition network to determine the target recognition network. Each video keyframe in the keyframe sequence is input into the target recognition network in chronological order for continuity detection, and image pixel change data corresponding to each image keyframe in the keyframe sequence is determined. The image pixel change data corresponding to each image keyframe is then added together to calculate pixel position change data.
[0084] By randomly selecting a video key frame in the key frame sequence, its corresponding image resolution is determined; according to the image resolution, each continuity recognition network is screened to determine the target recognition network; the key frame sequence is input into the target recognition network for continuity detection to determine the image pixel change data corresponding to each image key frame in the key frame sequence; the pixel position change data is calculated based on the image pixel change data corresponding to each image key frame. Key frame sequences with different image resolutions can be identified through different continuity recognition networks, thereby improving the accuracy of continuity recognition of key frame sequences.
[0085] Optionally, each continuity recognition network is obtained through the following steps: obtaining multiple training sample groups and inputting them into the corresponding initial recognition network for training respectively; the image resolution corresponding to each training sample group is different, and the image resolution of each initial recognition network corresponds to the image resolution of the input training sample group.
[0086] The training sample group may be a set of video key frames with the same image resolution.
[0087] Specifically, the initial recognition network comprises a feature extraction module, a temporal information detection module, and a feature classification module. The feature extraction module can be an EfficientNet, which extracts facial features and generates feature vectors corresponding to each video keyframe. The temporal information detection module can be a two-layer long short-term memory (LSTM) network, which captures temporal information between frames. Each LSTM network has 512 hidden layer neurons, and a dropout layer is introduced between the two LSTM networks to prevent overfitting. The hidden state of the LSTM at the last moment is output. The features extracted by the LSTM network are input into the first fully connected layer of the feature classification module, mapped to an intermediate dimension, and activated using a Reluctant Unit (ReLU) function. The output of the first fully connected layer is then connected to a dropout layer to prevent overfitting. The second fully connected layer outputs the classification results. For face-swap detection tasks, a softmax function is used as the activation function. Cross-entropy loss is used to guide model training. Multiple training sample groups are obtained, each corresponding to a different image resolution. Within each training sample group, the image resolution of each video keyframe is the same. An initial recognition network only recognizes video key frames of one image resolution. Different initial recognition networks recognize different image resolutions. Each training sample group is input into the initial recognition network of the corresponding image resolution for training, thereby obtaining at least one continuous recognition network.
[0088] By training each initial recognition network with training sample groups of different image resolutions, we obtain continuity recognition networks corresponding to different image resolutions. Each continuity recognition network has the highest accuracy in identifying the key frame sequence of its corresponding image resolution, which refines the continuity recognition scheme and improves the accuracy of continuity recognition.
[0089] Optionally, the image pixel change data includes: background pixel change data and face pixel change data.
[0090] Specifically, AI face-changing technology usually replaces or modifies the facial area in the video, while the background area remains relatively stable. If a video is processed with AI face-changing, some unnatural pixel changes may appear where the face and the background merge, such as inconsistent lighting, color differences, or imperfect edge fusion. By analyzing the statistical characteristics, texture information, and inter-frame change patterns of background pixels, these abnormal clues can be found and the background pixel change data can be determined. During the AI face-changing process, the pixel values of the face will change. Whether it is replacing the source face with the target face or editing the facial features, the pixel distribution or texture structure of the face area will change. At the same time, the movement and expression changes of real faces in videos have certain regularities, and AI face-changing may destroy these regularities. The facial pixel change data of each key frame of the video is detected for face-changing detection.
[0091] The image pixel change data is determined by the background pixel change data and the face pixel change data, and whether the video stream to be detected is a face-changing video stream is detected through multi-dimensional data, thereby improving the accuracy of face-changing detection.
[0092] Example 3
[0093] Figure 3 This is a schematic diagram of the structure of a video stream detection device provided in Embodiment 3 of the present invention. This embodiment of the present invention is applicable to situations where video stream detection is performed. The device can execute a video stream detection method and can be implemented in the form of hardware and / or software.
[0094] See also Figure 3 The video stream detection device shown includes: a video stream acquisition module 301, a key frame screening module 302, a distortion detection module 303, a continuity detection module 304 and a video detection module 305, wherein:
[0095] The video stream acquisition module 301 is used to determine the video stream to be detected based on the acquired start verification time and end verification time;
[0096] A key frame screening module 302 is configured to screen key frames of the video stream to be detected and determine a key frame sequence, wherein the key frame sequence includes at least one video key frame;
[0097] The distortion detection module 303 is used to perform distortion detection on the key frame sequence and determine the amount of distortion corresponding to the key frame sequence;
[0098] The continuity detection module 304 is used to perform continuity detection on the key frame sequence and determine pixel position change data corresponding to the key frame sequence;
[0099] The video detection module 305 is used to determine the video stream detection result according to the distortion amount and pixel position change data.
[0100] The technical solution of the embodiment of the present invention improves the efficiency of video stream detection by screening key frames of the video stream to be detected, determining the key frame sequence, reducing the number of video key frames to be processed, and performing distortion detection and continuity detection on the key frame sequence, and detecting the video stream through multi-dimensional data, thereby improving the accuracy and comprehensiveness of video stream detection.
[0101] Optionally, the video detection module 305 is specifically configured to:
[0102] Obtaining at least one distortion range and a distortion weight corresponding to each distortion range;
[0103] Compare the distortion amount with each distortion range to determine the target range corresponding to the distortion amount and the target weight corresponding to the target range;
[0104] The video stream detection value is calculated based on the distortion amount, target weight and pixel position change data;
[0105] Get the video detection threshold;
[0106] The video stream detection result is determined based on a comparison result between the video stream detection value and the video detection threshold.
[0107] Optionally, the key frame screening module 302 is specifically configured to:
[0108] The library data screening subunit is used to screen each scheme parameter item in the historical probability library corresponding to the current cluster according to the current cluster, operation information, scheme parameter items corresponding to the emergency plan, and parameter values corresponding to the scheme parameter items, to obtain the historical probability corresponding to the parameter values corresponding to the scheme parameter items;
[0109] The product acquisition subunit is used to multiply the historical probabilities corresponding to the parameter values corresponding to the parameter items of each plan of the emergency plan to calculate the initial probability corresponding to the emergency plan.
[0110] Optionally, the distortion detection module 303 is specifically configured to:
[0111] Combining each video key frame in the key frame sequence in pairs to obtain at least one image group;
[0112] Calculate the similarity between the two video key frames in each image group to obtain the image similarity value corresponding to each image group;
[0113] Get similarity threshold;
[0114] For each image group, the image similarity value corresponding to the image group is compared with the similarity threshold to obtain the distortion result of the image group;
[0115] According to the distortion results of each image group, the distortion quantity corresponding to the key frame sequence is obtained by statistics.
[0116] Optionally, the continuity detection module 304 includes:
[0117] A network acquisition unit, configured to acquire at least one trained continuity recognition network;
[0118] A resolution determination unit, configured to select any video key frame from the key frame sequence and determine the corresponding image resolution;
[0119] A network screening unit, used to screen various continuity recognition networks according to image resolution and determine a target recognition network;
[0120] A network detection unit is used to input the key frame sequence into the target recognition network for continuity detection, and determine the image pixel change data corresponding to each key frame in the key frame sequence;
[0121] The data calculation unit is used to calculate the pixel position change data according to the image pixel change data corresponding to each image key frame.
[0122] Optionally, each continuity identification network is obtained by the following steps:
[0123] A plurality of training sample groups are obtained and respectively input into corresponding initial recognition networks for training; each training sample group corresponds to a different image resolution, and the image resolution of each initial recognition network corresponds to the image resolution of the input training sample group.
[0124] Optionally, the image pixel change data includes: background pixel change data and face pixel change data.
[0125] The video stream detection device provided in the embodiment of the present invention can execute the video stream detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the video stream detection method.
[0126] Example 4
[0127] Figure 4 FIG. 4 is a schematic structural diagram of a video stream detection device 400 that can be used to implement an embodiment of the present invention.
[0128] like Figure 4As shown, the video stream detection device 400 includes at least one processor 401 and a memory connected to the at least one processor 401, such as a read-only memory (ROM) 402, a random access memory (RAM) 403, etc., wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 401 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 402 or the computer program loaded from the storage unit 408 into the random access memory (RAM) 403. Various programs and data required for the operation of the video stream detection device 400 can also be stored in the RAM 403. The processor 401, ROM 402 and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0129] Multiple components in the video stream detection device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the video stream detection device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0130] The processor 401 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 401 executes the various methods and processes described above, such as the video stream detection method.
[0131] In some embodiments, the video stream detection method can be implemented as a computer program that is tangibly contained in a computer-readable storage medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the video stream detection device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by processor 401, one or more steps of the video stream detection method described above can be performed. Alternatively, in other embodiments, processor 401 can be configured to perform the video stream detection method by any other appropriate means (e.g., by means of firmware).
[0132] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0133] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0134] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0135] To provide interaction with a user, the systems and techniques described herein can be implemented on a video stream detection device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the video stream detection device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0136] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0137] A computing system may include clients and servers. The clients and servers are generally remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS (Virtual Private Server) services.
[0138] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0139] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A video stream detection method, characterized in that: The method comprises: Determine the video stream to be detected based on the obtained start verification time and end verification time; Performing key frame screening on the video stream to be detected to determine a key frame sequence, wherein the key frame sequence includes at least one video key frame; Performing distortion detection on the key frame sequence to determine the amount of distortion corresponding to the key frame sequence; Performing continuity detection on the key frame sequence to determine pixel position change data corresponding to the key frame sequence; A video stream detection result is determined according to the distortion amount and the pixel position change data.
2. The method according to claim 1, characterized in that Determining a video stream detection result according to the distortion amount and the pixel position change data includes: Obtaining at least one distortion range and a distortion weight corresponding to each distortion range; Comparing the distortion amount with each of the distortion ranges, and determining a target range corresponding to the distortion amount and a target weight corresponding to the target range; Calculating a video stream detection value according to the distortion amount, the target weight, and the pixel position change data; Get the video detection threshold; A video stream detection result is determined according to a comparison result between the video stream detection value and the video detection threshold.
3. The method according to claim 1, characterized in that The step of screening the key frames of the video stream to be detected and determining a key frame sequence includes: Get the verification action, time step and time collection quantity; In response to identifying the verification action in the video stream to be detected, determining the time data to be collected, the time data to be collected including: the action start time, the action middle time and the action end time; Determining at least one collection moment according to the moment step, the number of time collections, and the time data to be collected; Key frame screening is performed on the video stream to be detected according to each of the acquisition moments to obtain at least one video key frame, and a key frame sequence is generated.
4. The method according to claim 1, wherein The performing distortion detection on the key frame sequence to determine the distortion amount corresponding to the key frame sequence includes: Combining the video key frames in the key frame sequence in pairs to obtain at least one image group; Calculating the similarity between two video key frames in each of the image groups to obtain a similarity value of the corresponding images in each of the image groups; Get similarity threshold; For each of the image groups, comparing the image similarity value corresponding to the image group with the similarity threshold to obtain a distortion result of the image group; According to the distortion results of each of the image groups, the distortion quantity corresponding to the key frame sequence is obtained by counting.
5. The method according to claim 1, wherein The performing continuity detection on the key frame sequence to determine pixel position change data corresponding to the key frame sequence includes: Obtain at least one trained continuity recognition network; Selecting any one video key frame in the key frame sequence and determining its corresponding image resolution; Screening the continuity recognition networks according to the image resolution to determine a target recognition network; Inputting the key frame sequence into the target recognition network for continuity detection, and determining image pixel change data corresponding to each image key frame in the key frame sequence; The pixel position change data is calculated based on the image pixel change data corresponding to each of the image key frames.
6. The method according to claim 5, characterized in that Each of the continuity identification networks is obtained by the following steps: A plurality of training sample groups are obtained and respectively input into corresponding initial recognition networks for training; the image resolution corresponding to each training sample group is different, and the image resolution of each initial recognition network corresponds to the image resolution of the input training sample group.
7. The method according to claim 5, characterized in that The image pixel change data includes background pixel change data and face pixel change data.
8. A video stream detection device, characterized in that: The device comprises: The video stream acquisition module is used to determine the video stream to be detected based on the acquired start verification time and end verification time; A key frame screening module, configured to screen the video stream to be detected for key frames and determine a key frame sequence, wherein the key frame sequence includes at least one video key frame; a distortion detection module, configured to perform distortion detection on the key frame sequence and determine the amount of distortion corresponding to the key frame sequence; a continuity detection module, configured to perform continuity detection on the key frame sequence and determine pixel position change data corresponding to the key frame sequence; The video detection module is used to determine a video stream detection result according to the distortion amount and the pixel position change data.
9. A video stream detection device, characterized in that: The video stream detection device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the video stream detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the video stream detection method according to any one of claims 1 to 7 when executed.