Recharge video consistency detection method, device and equipment and storage medium

By comparing the re-fed video frames with the template images in real time, identifying and eliminating template frames, and comparing the re-fed video with the original video frame by frame, the problem of template frame interference in the re-fed video is solved, and the accuracy of autonomous driving simulation testing is improved.

CN121767902APending Publication Date: 2026-03-31广州视晟科技有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In simulation testing of autonomous driving systems, template frames inserted into the re-feedback video interfered with the consistency of the video data and affected the accuracy of the simulation test results.

Method used

By acquiring the re-implemented video frames in real time and comparing them with template images, template frames are identified and excluded. The re-implemented video is compared with the original video frame by frame, and the consistency of the frames is judged using indicators such as structural similarity index and peak signal-to-noise ratio, so as to ensure the accuracy of the effective video frames.

Benefits of technology

It improves the accuracy of simulation test results, eliminates template frame interference before and after the effective video in the re-feedback video, and realizes high-fidelity, transparent data capture and real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767902A_ABST
    Figure CN121767902A_ABST
Patent Text Reader

Abstract

The invention discloses a recharge video consistency detection method, device and equipment and a storage medium, for each video frame of a recharge video acquired in real time, the video frame is compared with a template image according to the sequence of the video frame, and when continuous K video frames are matched with the template image, the video frame is compared with the template image according to the sequence of the video frame. Taking the next frame of the continuous K video frames as an initial frame of consistency detection, performing consistency comparison on the video frames and the original video frame by frame according to the sequence of the video frames from the initial frame, and when detecting that the picture consistency result of the original frames with the corresponding sequence numbers in the original video is inconsistent, performing the consistency comparison on the video frames and the original frames with the corresponding sequence numbers in the original video. And taking the video frame as an abnormal frame, comparing the abnormal frame with the template image, and ending consistency detection when the abnormal frame is matched with the template image or the last frame of the original video is compared. According to the invention, the interference of the template frames before and after the effective video in the recharge video can be eliminated, and the accuracy of the simulation test result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to autonomous driving simulation testing technology, and more particularly to a method, apparatus, device, and storage medium for backfeeding video consistency detection. Background Technology

[0002] In the research and development and simulation testing of autonomous driving systems, a closed-loop verification model of "real-vehicle road data acquisition + laboratory data re-injection" is commonly adopted. By re-injecting the raw video data collected during actual driving into the domain controller in the laboratory, complex scenarios can be reproduced, and the accuracy and robustness of the perception algorithm can be verified.

[0003] To ensure the validity of the test, the re-acquired video must be highly consistent with the original acquired video in terms of content, timing, and format. However, actual re-acquired systems typically insert template frames before and after the valid video data for system synchronization, status indication, or process control. The presence of these non-original contents can interfere with automated verification and affect the accuracy of simulation test results. Summary of the Invention

[0004] This invention provides a method, apparatus, device, and storage medium for detecting the consistency of reloaded video, so as to eliminate the interference of template frames before and after the effective video in the reloaded video and improve the accuracy of simulation test results.

[0005] In a first aspect, the present invention provides a method for detecting consistency in re-implemented video, comprising:

[0006] For each video frame of the real-time acquired feedback video, the video frames are compared with the template image in the order of the video frames;

[0007] When K consecutive video frames match the template image, the next frame of the K consecutive video frames is taken as the starting frame for consistency detection, where K is a positive integer greater than 1.

[0008] Starting from the initial frame, the video frames are compared frame by frame with the original video in the order of the video frames to determine the consistency of the video frames with the original frames of the corresponding sequence number in the original video.

[0009] When a video frame is detected that does not match the image consistency result of the original frame with the corresponding sequence number in the original video, the video frame is regarded as an abnormal frame.

[0010] The abnormal frame is compared with the template image;

[0011] The consistency detection ends when the abnormal frame matches the template image, or when the last frame of the original video has been compared.

[0012] Optionally, before comparing the video frames with the template image in the order of the video frames, the method further includes:

[0013] The relay device is connected in series on a high-speed serial link between the recharge system and the vehicle domain controller. The relay device copies the video stream output by the recharge system into two video streams. One video stream is output to the vehicle domain controller, and the other video stream is used as the recharge video stream.

[0014] The re-fed video stream is decoded to obtain decoded video frames.

[0015] Optionally, comparing the video frames with the template image in the order of the video frames includes:

[0016] Preload template images;

[0017] Extract the Y component from the template image to obtain the first image;

[0018] The first image is scaled to a preset resolution to obtain the second image;

[0019] The Y component is extracted from the video frames in the order they appear to be to obtain the third image.

[0020] The third image is scaled to the preset resolution to obtain the fourth image;

[0021] The structural similarity index between the template image and the video frame is calculated based on the second image and the fourth image as the first similarity.

[0022] When the first similarity is greater than a first preset value, the video frame is determined to match the template image.

[0023] Optionally, starting from the starting frame, the video frames are compared frame by frame with the original video in the order of the video frames to determine the consistency result of the video frames with the original frames of the corresponding sequence numbers in the original video, including:

[0024] Starting from the initial frame, calculate the peak signal-to-noise ratio of the video frame and the corresponding original frame in the original video according to the order of the video frames;

[0025] Starting from the initial frame, the structural similarity index between the video frame and the corresponding original frame in the original video is calculated as the second similarity, according to the order of the video frames.

[0026] When the peak signal-to-noise ratio is less than a second preset value and the second similarity is less than a third preset value, the consistency result between the video frame and the original frame with the corresponding sequence number is determined to be inconsistent.

[0027] When the peak signal-to-noise ratio is greater than or equal to the second preset value, or the second similarity is greater than or equal to the third preset value, the video frame is determined to be consistent with the original frame of the corresponding sequence number.

[0028] Optionally, the video consistency detection method also includes:

[0029] When the abnormal frame does not match the template image, the abnormal frame is treated as an abnormal scene in the valid video, and the comparison between the next video frame and the next original frame continues until the template image is matched, or the last frame of the original video is compared.

[0030] Optionally, the video consistency detection method also includes:

[0031] If an abnormal frame matching the template image is detected before the original video has been fully read, the frame before the abnormal frame is taken as the end frame of the consistency detection, and it is determined that the re-fed video has lost frames or that the re-fed video was injected into the template frame in advance.

[0032] If the last frame of the original video is read and the current video frame does not match the template image, then continue reading the next video frame;

[0033] If the next video frame matches the template image, it is determined that the number of frames in the re-fed video is the same as that in the original video;

[0034] If the next video frame does not match the template image, it is determined that the re-fed video has redundant video frames.

[0035] Optionally, after the consistency check is completed, the following may also be included:

[0036] Calculate the mean of the peak signal-to-noise ratio and the mean of the second similarity of the effective video frames in the re-injected video, wherein the effective video frames are consecutive video frames from the start frame to the end frame;

[0037] When the mean of the peak signal-to-noise ratio is less than the first threshold, or the mean of the second similarity is less than the second threshold, it is determined that the re-fed video has generalized image distortion.

[0038] Secondly, the present invention also provides a video consistency detection device for re-feedback, comprising:

[0039] The first comparison module is used to compare each video frame of the real-time acquired re-entered video with the template image in the order of the video frames.

[0040] The starting frame determination module is used to determine the next frame of the K consecutive video frames as the starting frame for consistency detection when the K consecutive video frames match the template image, where K is a positive integer greater than 1.

[0041] The second comparison module is used to compare the video frames with the original video frame by frame, starting from the starting frame and following the order of the video frames, to determine the consistency result of the video frames with the original frames with corresponding numbers in the original video.

[0042] The abnormal frame determination module is used to identify the video frame as an abnormal frame when a video frame is detected that does not match the image consistency result of the original frame with the corresponding sequence number in the original video.

[0043] The third comparison module is used to compare the abnormal frame with the template image;

[0044] The detection termination module is used to terminate the consistency detection when the abnormal frame matches the template image, or when the last frame of the original video has been compared.

[0045] Thirdly, the present invention also provides an electronic device, comprising:

[0046] One or more processors;

[0047] Storage device for storing one or more programs;

[0048] When the one or more programs are executed by the one or more processors, the one or more processors implement the video consistency detection method as described in the first aspect of the present invention.

[0049] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video backfeeding consistency detection method as described in the first aspect of the present invention.

[0050] The re-implementation video consistency detection method provided by this invention compares each video frame of the real-time acquired re-implementation video with a template image in sequence. When K consecutive video frames match the template image, the next frame after the K consecutive video frames is used as the starting frame for consistency detection. Starting from the starting frame, the video frames are compared frame by frame with the original video in sequence to determine the consistency result between the video frame and the corresponding original frame in the original video. When a video frame is detected that does not match the consistency result of the corresponding original frame in the original video, it is designated as an abnormal frame. The abnormal frame is then compared with the template image. The consistency detection ends when the abnormal frame matches the template image or when the last frame of the original video has been compared. This invention, by acquiring video frames of the re-implemented video in real time and comparing them with template images to determine the valid video frames of the re-implemented video, and performing a frame-by-frame consistency comparison between the valid video frames and the video frames of the original video, can eliminate the interference of template frames before and after the valid video in the re-implemented video, thus improving the accuracy of simulation test results.

[0051] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 A flowchart of a video consistency detection method provided by the present invention;

[0054] Figure 2 This is a schematic diagram of the structure of a video consistency detection system for re-feedback provided by the present invention;

[0055] Figure 3 This is a schematic diagram of the structure of a video consistency detection device for reflow provided by the present invention;

[0056] Figure 4 This is a schematic diagram of the structure of an electronic device provided by the present invention.

[0057] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation

[0058] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0059] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0060] Figure 1 This is a flowchart of a video re-implementation consistency detection method provided by the present invention. This embodiment can be used to solve the problem that after inserting template frames into the re-implemented video, inconsistencies arise between the template frames and the original video, making consistency comparison impossible. This method can be executed by the video re-implementation consistency detection device provided by the present invention. This device can be implemented by software and / or hardware, and is typically configured in an electronic device, such as... Figure 1 As shown, the video consistency detection method includes the following steps:

[0061] S101. For each video frame of the real-time acquired feedback video, compare the video frames with the template image in the order of the video frames.

[0062] In this embodiment of the invention, the re-uploaded video can be acquired in real time. For each video frame of the real-time acquired re-uploaded video, the video frames are compared with the template image in the order of the video frames.

[0063] As mentioned earlier, re-implantation systems typically insert template frames (e.g., black frames or static identifiers) before and after valid video data for system synchronization, status indication, or process control. Therefore, this invention compares video frames with template images in sequence to identify template frames in the re-implanted video, thereby determining the valid video frames in the re-implanted video. The template image is the same as the template frame.

[0064] Figure 2 This is a schematic diagram of the structure of a video consistency detection system provided by the present invention, as shown below. Figure 2 As shown, the reload video consistency detection system includes a reload system 101, a relay device 102, an automotive domain controller 103, and a host computer 104. The host computer 104 is used to execute the reload video consistency detection method described in this invention. The reload system 101 processes the original video captured by the camera, inserting template frames before and after the original video for system synchronization, status indication, or process control. The relay device 102 is connected in series on the high-speed serial link between the reload system 101 and the automotive domain controller 103, such as a Gigabit Multimedia Serial Link (GMSL). The relay device 102 copies the video stream output by the reload system 101 into two video streams. One video stream is output to the automotive domain controller 103, and the other video stream is output to the host computer 104 as the reload video stream. For example, the relay device 102 is composed of an FPGA and a GMSL SerDes chip. After receiving the video bitstream from the feedback system 101, it generates two electrically isolated and protocol-compatible GMSL output signals after deserialization.

[0065] Currently, vehicle-mounted cameras and domain controllers mostly use proprietary high-speed serial protocols such as high-speed serial links to transmit video data. Their point-to-point physical structure makes direct signal splitting or monitoring difficult. Traditional methods either rely on invasive modifications (such as altering the re-feedback device interface) or employ offline analysis, performing data comparison only after re-feedback is complete. This results in high detection latency, delayed anomaly feedback, and an inability to intervene in real time. This invention connects a relay device in series in the high-speed serial link between the re-feedback system and the vehicle domain controller. This copies the video stream output from the re-feedback system into two video streams: one output to the vehicle domain controller, and the other used as the re-feedback video stream. Without invasive modifications, this achieves high-fidelity, transparent data capture, ensuring the detection process does not interfere with the original system. Furthermore, the re-feedback process to the vehicle domain controller and the consistency comparison process can be performed synchronously, improving the real-time performance and efficiency of the test.

[0066] In this embodiment of the invention, the feedback video stream output by the relay device can be acquired in real time, and the feedback video stream can be decoded to obtain decoded video frames. For example, after decoding, a YUV format frame stream is obtained and cached in a memory buffer.

[0067] In some embodiments of the present invention, step S101 above may include the following sub-steps:

[0068] S1011, Preload template image.

[0069] In this embodiment of the invention, a template image is invoked and preloaded into memory.

[0070] S1012. Extract the Y component from the template image to obtain the first image.

[0071] In this embodiment of the invention, the Y component is extracted from the template image to obtain a first image characterizing the structural features of the template image. The Y component is an important part of the YUV color space, representing the image's luminance information, and is usually related to grayscale values. The U and V components represent chroma (hue and saturation), which in turn define two aspects of color: hue and saturation. The Y component of the image preserves the image's structural features to the greatest extent possible, such as shape, contour, texture, edges, and corners.

[0072] S1013. Scale the first image to a preset resolution to obtain the second image.

[0073] The first image is scaled to a preset resolution to obtain the second image. For example, the first image is scaled to a resolution of 640×360 to obtain the second image.

[0074] S1014. Extract the Y component from the video frames in the order of the video frames to obtain the third image.

[0075] For each video frame of the real-time acquired feedback video, the Y component is extracted from the video frames in order to obtain a third image characterizing the structural features of the video frame.

[0076] S1015. Scale the third image to a preset resolution to obtain the fourth image.

[0077] The third image is scaled to a preset resolution to obtain the fourth image. For example, the third image is scaled to a resolution of 640×360 to obtain the fourth image.

[0078] This invention extracts the Y component of an image, preserving its structural features for subsequent comparison while further reducing data processing volume and improving processing efficiency. Furthermore, uniformly scaling the images to a preset resolution serves two purposes: ensuring consistent comparison and reducing data processing volume, thus enhancing efficiency.

[0079] S1016. Calculate the structural similarity index of the template image and video frame based on the second image and the fourth image as the first similarity.

[0080] In this embodiment of the invention, the Structural Similarity Index (SSIM) of the template image and video frame is calculated based on the second and fourth images as the first similarity. SSIM is used to measure the similarity between two images, mainly considering three key features of the images: luminance, contrast, and structure.

[0081] S1017. When the first similarity is greater than the first preset value, determine that the video frame matches the template image.

[0082] When the first similarity is greater than a first preset value (e.g., 0.98), the video frame is determined to match the template image, indicating that the video frame is the template frame.

[0083] S102. When K consecutive video frames match the template image, the next frame of the K consecutive video frames is taken as the starting frame for consistency detection.

[0084] In the re-implementation of video, K consecutive template frames are typically inserted before the valid video data, where K is a positive integer greater than 1. When K consecutive video frames match the template image, the next frame after the K consecutive video frames is used as the starting frame for consistency detection, i.e., the starting frame for valid video frames.

[0085] S103. Starting from the initial frame, compare the video frames with the original video frame by frame in the order of the video frames to determine the consistency of the video frames with the original frames of the corresponding sequence numbers in the original video.

[0086] After determining the starting frame, starting from the starting frame, each video frame is compared with the original video frame by frame in sequence to determine the consistency between the video frame and the corresponding original frame in the original video. That is, the starting frame is compared with the starting frame of the original video, the second frame is compared with the second frame of the original video, and so on, to determine the consistency between each video frame and the corresponding original frame in the original video.

[0087] In some embodiments of the present invention, step S103 above includes the following sub-steps:

[0088] S1031. Starting from the initial frame, calculate the peak signal-to-noise ratio of the video frames and the corresponding original frames in the original video according to the order of the video frames.

[0089] In this embodiment of the invention, starting from the starting frame, the peak signal-to-noise ratio (PSNR) of the video frames and the corresponding original frames in the original video are calculated according to the order of the video frames. PSNR is one of the most classic objective quality evaluation indicators in the field of digital image processing. Its core idea is to measure the degree of distortion by calculating the mean square error (MSE) between the original image (the original frame in this invention) and the distorted image (the video frame in this invention), and then to quantify the evaluation by the ratio of the maximum signal power to the noise power.

[0090] S1032. Starting from the initial frame, calculate the structural similarity index between the video frame and the corresponding original frame in the original video according to the order of the video frames as the second similarity.

[0091] In this embodiment of the invention, starting from the starting frame, the structural similarity index (SSIM) between the video frame and the original frame with the corresponding sequence number in the original video is calculated as the second similarity, according to the order of the video frames.

[0092] S1033. When the peak signal-to-noise ratio is less than the second preset value and the second similarity is less than the third preset value, the consistency result between the video frame and the original frame with the corresponding sequence number is determined to be inconsistent.

[0093] When the peak signal-to-noise ratio is less than the second preset value (e.g., 30dB) and the second similarity is less than the third preset value (e.g., 0.85), the consistency result between the video frame and the original frame with the corresponding sequence number is determined to be inconsistent.

[0094] S1034. When the peak signal-to-noise ratio is greater than or equal to the second preset value, or the second similarity is greater than or equal to the third preset value, the consistency result between the video frame and the original frame with the corresponding sequence number is determined to be consistent.

[0095] When the peak signal-to-noise ratio is greater than or equal to the second preset value (e.g., 30dB), or the second similarity is greater than or equal to the third preset value (e.g., 0.85), the video frame is determined to be consistent with the original frame of the corresponding sequence number.

[0096] In this embodiment of the invention, when comparing the valid video frames of the re-entered video with the original frames of the original video, no preprocessing (e.g., resolution scaling, component extraction) is performed on the valid video frames and the original frames. The original resolution and color format are maintained, and the original information of the valid video frames and the original video frames is preserved to the maximum extent, so as to avoid the distortion caused by image preprocessing from affecting the consistency comparison results.

[0097] S104. When a video frame is detected that does not match the image consistency result of the original frame with the corresponding sequence number in the original video, the video frame is regarded as an abnormal frame.

[0098] When a video frame is detected whose image consistency result is inconsistent with the original frame with the corresponding sequence number in the original video, the video frame is designated as an abnormal frame. In this embodiment of the invention, the inconsistency between the video frame and the original frame with the corresponding sequence number in the original video may be due to an anomaly in the video frame itself (e.g., image freezing, distortion, etc.), or it may be because the video frame is a template frame. Therefore, in this embodiment of the invention, the video frame is marked as an abnormal frame.

[0099] S105. Compare the abnormal frame with the template image.

[0100] To further determine whether the abnormal frame is a template frame or a valid video frame with an abnormal image, the detected abnormal frame is compared with the template image. This comparison process can refer to the video frame and template image comparison process in the previous embodiments, and will not be repeated here.

[0101] S106. When an abnormal frame is matched with a template image, or when the last frame of the original video is compared, the consistency detection ends.

[0102] When an abnormal frame is matched with a template image, the abnormal frame is determined to be a template frame. Then all valid video frames have been compared and the consistency detection ends. Alternatively, the consistency detection ends when the last frame of the original video has been compared.

[0103] When an abnormal frame does not match the template image, it indicates that the abnormal frame is not a template frame. The abnormal frame is treated as an abnormal scene in the valid video, and the comparison continues with the next video frame and the next original frame until the template image is matched or the last frame of the original video is compared.

[0104] In some embodiments of the present invention, when an abnormal frame matching the template image is detected before the original video has been fully read, the frame before the abnormal frame is taken as the end frame of the consistency detection (i.e. the end frame of the valid video frame), and it is determined that the re-fed video has lost frames or that the re-fed video was injected into the template frame in advance, triggering a "data loss or re-fed interruption" alarm.

[0105] In some embodiments of the present invention, when the last frame of the original video is read and the current video frame does not match the template image, the next video frame is read. If the next video frame matches the template image, it is determined that the number of frames in the reloaded video is the same as that in the original video. If the next video frame does not match the template image, that is, the number of valid video frames in the reloaded video is greater than the number of frames in the original video, it is determined that there are redundant video frames in the reloaded video, triggering a "data overflow" alarm.

[0106] After the consistency test is completed, a consistency test report can be output, which includes both image consistency and frame rate consistency.

[0107] Image consistency: If any of the following conditions occur, it is judged as "image consistency not meeting the standard".

[0108] 1. Overall image quality does not meet requirements: If the average PSNR of all valid video frames is lower than a preset threshold (e.g., 30dB), or the average SSIM is lower than a threshold (e.g., 0.85), it indicates that there is widespread image distortion during the re-implantation process. For example, the mean of the peak signal-to-noise ratio and the mean of the second similarity of the valid video frames in the re-implanted video are calculated. Valid video frames are continuous video frames from the start frame to the end frame. When the mean of the peak signal-to-noise ratio is less than the first threshold (e.g., 30dB), or the mean of the second similarity is less than the second threshold (e.g., 0.85), it is determined that there is widespread image distortion in the re-implanted video.

[0109] 2. Localized image anomalies: There are abnormal frames with a single or multiple frames having a PSNR lower than the second preset value and an SSIM lower than the third preset value. These abnormal frames are confirmed as valid video frames by template frame identification (non-template frames). This indicates that localized image anomalies have occurred, such as image freezing, distortion, or content deviation.

[0110] Frame rate consistency: The following conditions are considered "frame rate inconsistency".

[0111] 1. If an abnormal frame matching the template image is detected before the original video has been fully read (i.e., a template frame is detected), it is determined that the re-uploaded video has lost frames or that the re-uploaded video was injected with the template frame in advance, meaning that the number of valid video frames in the re-uploaded video is less than the number of frames in the original video.

[0112] 2. If the original video has been read but no template frame has been detected, it is determined that there are redundant video frames in the reloaded video, and the number of valid video frames in the reloaded video is greater than the number of frames in the original video.

[0113] The re-implementation video consistency detection method provided by this invention compares each video frame of the real-time acquired re-implementation video with a template image in sequence. When K consecutive video frames match the template image, the next frame after the K consecutive video frames is used as the starting frame for consistency detection. Starting from the starting frame, the video frames are compared frame by frame with the original video in sequence to determine the consistency result between the video frame and the corresponding original frame in the original video. When a video frame is detected that does not match the consistency result of the corresponding original frame in the original video, it is designated as an abnormal frame. The abnormal frame is then compared with the template image. The consistency detection ends when the abnormal frame matches the template image or when the last frame of the original video has been compared. This invention, by acquiring video frames of the re-implemented video in real time and comparing them with template images to determine the valid video frames of the re-implemented video, and performing a frame-by-frame consistency comparison between the valid video frames and the video frames of the original video, can eliminate the interference of template frames before and after the valid video in the re-implemented video, thus improving the accuracy of simulation test results.

[0114] Figure 3 This is a schematic diagram of the structure of a video consistency detection device for reflow provided by the present invention, as shown below. Figure 3 As shown, the video consistency detection device includes:

[0115] The first comparison module 201 is used to compare each video frame of the real-time acquired re-entered video with the template image in the order of the video frames.

[0116] The starting frame determination module 202 is used to take the next frame of the K consecutive video frames as the starting frame for consistency detection when the K consecutive video frames match the template image, where K is a positive integer greater than 1.

[0117] The second comparison module 203 is used to compare the video frames with the original video frame by frame, starting from the starting frame and following the order of the video frames, to determine the consistency result of the video frames with the original frames with corresponding numbers in the original video.

[0118] The abnormal frame determination module 204 is used to identify the video frame as an abnormal frame when a video frame is detected that does not match the image consistency result of the original frame with the corresponding sequence number in the original video.

[0119] The third comparison module 205 is used to compare the abnormal frame with the template image;

[0120] The detection end module 206 is used to end the consistency detection when the abnormal frame matches the template image, or when the last frame of the original video has been compared.

[0121] In some embodiments of the present invention, the video consistency detection device further includes:

[0122] The feedback video stream acquisition module is used to acquire the feedback video stream output by the relay device in real time before comparing the video frames with the template image in the order of the video frames. The relay device is connected in series on the high-speed serial link between the feedback system and the vehicle domain controller. The relay device copies the video stream output by the feedback system into two video streams. One video stream is output to the vehicle domain controller, and the other video stream is used as the feedback video stream.

[0123] The decoding module is used to decode the re-fed video stream to obtain decoded video frames.

[0124] In some embodiments of the present invention, the first comparison module 201 includes:

[0125] The template image loading submodule is used to preload template images;

[0126] The first extraction submodule is used to extract the Y component from the template image to obtain a first image;

[0127] The first scaling submodule is used to scale the first image to a preset resolution to obtain the second image;

[0128] The second extraction submodule is used to extract the Y component from the video frames in the order of the video frames to obtain the third image;

[0129] The second scaling submodule is used to scale the third image to the preset resolution to obtain the fourth image;

[0130] The first similarity calculation submodule is used to calculate the structural similarity index between the template image and the video frame as the first similarity based on the second image and the fourth image;

[0131] The matching determination submodule is used to determine whether the video frame matches the template image when the first similarity is greater than a first preset value.

[0132] In some embodiments of the present invention, the second comparison module 203 includes:

[0133] The peak signal-to-noise ratio (PSNR) calculation submodule is used to calculate the PNR of the video frames and the corresponding original frames in the original video, starting from the starting frame and following the order of the video frames.

[0134] The second similarity calculation submodule is used to calculate the structural similarity index between the video frame and the original frame with the corresponding number in the original video, starting from the starting frame and in the order of the video frames, as the second similarity.

[0135] The first determination submodule is used to determine that the consistency result between the video frame and the original frame with the corresponding sequence number is inconsistent when the peak signal-to-noise ratio is less than a second preset value and the second similarity is less than a third preset value.

[0136] The second determination submodule is used to determine that the video frame and the original frame with the corresponding sequence number are consistent when the peak signal-to-noise ratio is greater than or equal to a second preset value, or the second similarity is greater than or equal to a third preset value.

[0137] In some embodiments of the present invention, the video consistency detection device further includes:

[0138] The repeated comparison module is used to treat the abnormal frame as an abnormal scene of the valid video when the abnormal frame does not match the template image, and continue to compare the next video frame and the next original frame until the template image is matched, or the last frame of the original video is compared.

[0139] In some embodiments of the present invention, the video consistency detection device further includes:

[0140] The end frame determination module is used to determine the frame before the abnormal frame as the end frame of consistency detection when an abnormal frame matching the template image is detected before the original video has been fully read, and to determine whether the re-fed video has lost frames or the re-fed video has been injected into the template frame in advance.

[0141] The next frame reading module is used to read the next video frame if the current video frame does not match the template image after reading the last frame of the original video.

[0142] The third determination submodule is used to determine that if the next video frame matches the template image, the frame number of the re-fed video is the same as that of the original video.

[0143] The fourth determination submodule is used to determine that there are redundant video frames in the re-fed video if the next video frame does not match the template image.

[0144] In some embodiments of the present invention, the video consistency detection device further includes:

[0145] The average value calculation module is used to calculate the average peak signal-to-noise ratio and the average second similarity of the effective video frames in the re-injected video after the consistency detection is completed. The effective video frames are continuous video frames from the start frame to the end frame.

[0146] The general distortion determination module is used to determine that the re-fed video has general image distortion when the mean of the peak signal-to-noise ratio is less than a first threshold or the mean of the second similarity is less than a second threshold.

[0147] The aforementioned reflow video consistency detection device can execute the reflow video consistency detection method provided in the foregoing embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the reflow video consistency detection method.

[0148] Figure 4 This is a schematic diagram of an electronic device provided by the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0149] like Figure 4 As shown, the electronic device includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer programs stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0150] Multiple components in the electronic device are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, optical disk, etc.; and a communication unit 19, such as a network card, modem, wireless transceiver, etc. The communication unit 19 allows the electronic device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0151] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the feedback video consistency detection method.

[0152] In some embodiments, the power-back video consistency detection method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on an electronic device via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the power-back video consistency detection method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the power-back video consistency detection method by any other suitable means (e.g., by means of firmware).

[0153] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0154] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0155] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0156] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0157] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0158] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0159] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the video backfeeding consistency detection method provided in any embodiment of this application.

[0160] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0161] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0162] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for detecting consistency in re-implemented video, characterized in that, include: For each video frame of the real-time acquired feedback video, the video frames are compared with the template image in the order of the video frames; When K consecutive video frames match the template image, the next frame of the K consecutive video frames is taken as the starting frame for consistency detection, where K is a positive integer greater than 1. Starting from the initial frame, the video frames are compared frame by frame with the original video in the order of the video frames to determine the consistency of the video frames with the original frames of the corresponding sequence number in the original video. When a video frame is detected that does not match the image consistency result of the original frame with the corresponding sequence number in the original video, the video frame is regarded as an abnormal frame. The abnormal frame is compared with the template image; The consistency detection ends when the abnormal frame matches the template image, or when the last frame of the original video has been compared.

2. The video consistency detection method according to claim 1, characterized in that, Before comparing the video frames with the template image in the order of the video frames, the method further includes: The relay device is connected in series on a high-speed serial link between the recharge system and the vehicle domain controller. The relay device copies the video stream output by the recharge system into two video streams. One video stream is output to the vehicle domain controller, and the other video stream is used as the recharge video stream. The re-fed video stream is decoded to obtain decoded video frames.

3. The video consistency detection method according to claim 1, characterized in that, The video frames are compared with the template image in the order of the video frames, including: Preload template images; Extract the Y component from the template image to obtain the first image; The first image is scaled to a preset resolution to obtain the second image; The Y component is extracted from the video frames in the order they appear to be to obtain the third image. The third image is scaled to the preset resolution to obtain the fourth image; The structural similarity index between the template image and the video frame is calculated based on the second image and the fourth image as the first similarity. When the first similarity is greater than a first preset value, the video frame is determined to match the template image.

4. The video consistency detection method according to any one of claims 1-3, characterized in that, Starting from the initial frame, the video frames are compared frame by frame with the original video in the order of the video frames to determine the consistency result of the video frames with the corresponding original frames in the original video, including: Starting from the initial frame, calculate the peak signal-to-noise ratio of the video frame and the corresponding original frame in the original video according to the order of the video frames; Starting from the initial frame, the structural similarity index between the video frame and the corresponding original frame in the original video is calculated as the second similarity, according to the order of the video frames. When the peak signal-to-noise ratio is less than a second preset value and the second similarity is less than a third preset value, the consistency result between the video frame and the original frame with the corresponding sequence number is determined to be inconsistent. When the peak signal-to-noise ratio is greater than or equal to the second preset value, or the second similarity is greater than or equal to the third preset value, the video frame is determined to be consistent with the original frame of the corresponding sequence number.

5. The video consistency detection method according to any one of claims 1-3, characterized in that, Also includes: When the abnormal frame does not match the template image, the abnormal frame is treated as an abnormal scene in the valid video, and the comparison between the next video frame and the next original frame continues until the template image is matched, or the last frame of the original video is compared.

6. The video consistency detection method according to any one of claims 1-3, characterized in that, Also includes: If an abnormal frame matching the template image is detected before the original video has been fully read, the frame before the abnormal frame is taken as the end frame of the consistency detection, and it is determined that the re-fed video has lost frames or that the re-fed video was injected into the template frame in advance. If the last frame of the original video is read and the current video frame does not match the template image, then continue reading the next video frame; If the next video frame matches the template image, it is determined that the number of frames in the re-fed video is the same as that in the original video; If the next video frame does not match the template image, it is determined that the re-fed video has redundant video frames.

7. The video consistency detection method according to claim 4, characterized in that, After the consistency check is completed, the following is also included: Calculate the mean of the peak signal-to-noise ratio and the mean of the second similarity of the effective video frames in the re-injected video, wherein the effective video frames are consecutive video frames from the start frame to the end frame; When the mean of the peak signal-to-noise ratio is less than the first threshold, or the mean of the second similarity is less than the second threshold, it is determined that the re-fed video has generalized image distortion.

8. A device for detecting video consistency during reflow, characterized in that, include: The first comparison module is used to compare each video frame of the real-time acquired re-entered video with the template image in the order of the video frames. The starting frame determination module is used to determine the next frame of the K consecutive video frames as the starting frame for consistency detection when the K consecutive video frames match the template image, where K is a positive integer greater than 1. The second comparison module is used to compare the video frames with the original video frame by frame, starting from the starting frame and following the order of the video frames, to determine the consistency result of the video frames with the original frames with corresponding numbers in the original video. The abnormal frame determination module is used to identify the video frame as an abnormal frame when a video frame is detected that does not match the image consistency result of the original frame with the corresponding sequence number in the original video. The third comparison module is used to compare the abnormal frame with the template image; The detection termination module is used to terminate the consistency detection when the abnormal frame matches the template image, or when the last frame of the original video has been compared.

9. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the video consistency detection method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the video consistency detection method as described in any one of claims 1-7.