Video evidence chain consistency verification method and system based on space-time metadata

By using a video evidence chain consistency verification method based on spatiotemporal metadata, the limitations of video data consistency assessment with the physical world are overcome. This method enables multi-dimensional cross-comparison of video content and external attributes, providing reliable and automated video evidence chain verification, and improving the credibility of evidence and review efficiency.

CN121170670APending Publication Date: 2025-12-19NANJING SPEED DISTRIBUTION INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511259912.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing technologies fail to effectively link video pixel content with its multidimensional attributes in video data processing, resulting in limitations in assessing the consistency between video data and the objective physical world. There is a lack of an effective mechanism to correlate and analyze the internal visual presentation of a video with its external recording attributes.

Method used

The video evidence chain consistency verification method based on spatiotemporal metadata acquires video files and decodes them into pixel streams and internal metadata. It generates an adaptive retrieval window by combining time and location tags, calls multi-source external data for spatiotemporal alignment, generates consistent data streams of illumination, weather, and device status, aggregates them into a multi-dimensional consistency vector based on frame timestamps, and then fuses them into a video consistency sequence based on credibility and feature stability, finally generating a verification report.

Benefits of technology

It enables multi-dimensional cross-comparison between video content and the objective physical world, providing quantifiable and reproducible data-driven verification, enhancing the reliability and persuasiveness of evidence in court, improving the efficiency of digital evidence review, identifying forged videos, and reducing human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170670A_ABST
    Figure CN121170670A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition and digital forensics, in particular to a video evidence chain consistency verification method and system based on space-time metadata, and the method comprises the steps: obtaining a pixel stream and internal metadata through decoding a video file, generating a self-adaptive retrieval window through combining time, a position label and a scene element, and carrying out the verification of the consistency of the video evidence chain. And calling multi-source external data in the window to carry out space-time alignment and weighted fusion to form structured external metadata. Illumination, weather and holding state consistency data streams are further extracted and aggregated into multi-dimensional consistency vectors under the global time reference, and a video consistency sequence is generated through dynamic weighted fusion. And the system analyzes the consistency sequence segment by segment, marks risk segments, calculates global credibility, and finally outputs a visual verification report. According to the method, multi-dimensional cross validation of the video and the objective physical world is realized, and the authenticity, the acquisition reliability and the examination efficiency of the digital evidence are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image recognition and digital forensics, in particular to a video evidence chain consistency verification method and system based on spatiotemporal metadata. BACKGROUND

[0002] In an information-based society, video data has become an indispensable evidence medium in public safety, judicial authentication, and traffic tracing scenarios. Its authenticity, integrity, and logical coherence are directly related to the effectiveness of analysis and conclusions. Existing technologies, focusing on computer vision and deep learning, identify objects and scenes in static frames through target detection and image segmentation, and capture changes through spatiotemporal modeling to understand behavior trajectories. Such methods can deeply interpret video content at the pixel level, significantly improving the efficiency and automation level of information retrieval, event detection, and behavior recognition.

[0003] However, in practical applications, existing technologies focus highly on the analysis of video pixel content, while the multi-dimensional attributes of video data as a complete record carrier are not adequately considered. Video files not only carry visual information, but also contain or associate rich contextual data, such as camera device information, precise geographic location, timestamp, file creation and modification history, etc.

[0004] The current technical paradigm generally treats the pixel content of a video and its accompanying metadata as two independent information domains for processing. Analysis of pixel content aims to understand the picture, while processing of metadata often stops at the simple level of reading, archiving, or indexing. This processing approach leads to a split between video content and its contextual attributes. Therefore, while existing technologies can finely analyze visual information within a single video segment, they have universal limitations in assessing the consistency of video data with the objective physical world, lacking an effective mechanism for correlating internal visual presentation of a video with external recording attributes.

[0005] To this end, a video evidence chain consistency verification method and system based on spatiotemporal metadata are proposed. SUMMARY

[0006] The present application aims to provide a video evidence chain consistency verification method and system based on spatiotemporal metadata.

[0007] To achieve the above-mentioned purpose, the present application provides the following technical solutions:

[0008] The method and system for checking the consistency of video evidence chain based on spatiotemporal metadata comprises: obtaining a video file to be checked, and separating the video file into a pixel stream and internal metadata after decoding; generating an adaptive retrieval window based on the time and location tags of the internal metadata, and combining scene elements identified by the pixel stream; calling multiple external data sources in parallel within the window to perform spatiotemporal alignment, establishing a traceability path, and weighting and fusing the path into structured external metadata according to the credibility and feature stability; extracting shadow trajectories from the pixel stream through illumination analysis, and generating shadow consistency data stream by combining the sun azimuth and elevation; obtaining weather features through weather identification, and generating weather consistency data stream by combining historical records; obtaining interframe displacement sequences through jitter detection, and generating mobile state consistency data stream by combining sensor information; aggregating the three types of consistency data streams into a multidimensional consistency vector under a global time reference according to the frame timestamps; weighting and fusing the multidimensional consistency vector into a video consistency sequence based on the source credibility and feature stability, and retaining the traceability index; determining the consistency level of the video consistency sequence segment by segment, and outputting a video evidence chain checking report.

[0009] Preferably, the step of generating an adaptive retrieval window comprises: parsing the internal metadata to obtain timestamps and location tags, and performing scene element identification on the pixel stream to extract environmental features, illumination conditions, and motion trajectories; dynamically adjusting the time interval length and spatial range size of the retrieval window according to the rate and direction parameters of the motion trajectory, wherein the length of the time interval is inversely proportional to the rate of the motion trajectory, and the size of the spatial range is proportional to the rate of the motion trajectory; mapping the adjusted time, spatial parameters, and environmental features into a standardized data structure together to form an adaptive retrieval window.

[0010] Preferably, the process of outputting structured external metadata comprises: accessing external databases through a data interface, wherein the external databases include satellite positioning databases, meteorological service databases, and sensor log databases; assigning initial weights to each traceability path according to the authority of the data source and the stability of the data link, wherein the authority is quantified according to a preset source rating table, the data link stability is quantified according to the real-time success rate and delay of data calling, and the initial weight is derived from the weighted sum of the quantified authority and data link stability; dynamically adjusting the weight by combining the volatility evaluation results of the data in the time sequence, wherein the adjustment method is to apply a multiplicative attenuation factor according to the volatility evaluation results on the basis of the initial weight, wherein the greater the volatility, the closer the attenuation factor is to zero; outputting a structured external metadata stream with unified dimensions and attached credibility reference parameters through a normalized weighted fusion algorithm.

[0011] Preferably, the step of generating the light consistency data stream comprises: under the condition of detecting that the scene has a single main light source, extracting the shadow area formed by static objects in the video frame by image brightness and gradient analysis and calculating the boundary track of the shadow area; geometrically comparing the geometric direction and length change of the shadow boundary track with the corresponding spatio-temporal solar azimuth and altitude angle parameters obtained from the external metadata; calculating the quantified light matching degree index according to the difference between the included angle of the two vector directions and the length change rate; sorting the light matching degree indexes calculated at each time point according to the corresponding time stamp to generate the time sequence and form the light consistency data stream.

[0012] Preferably, the step of generating the weather consistency data stream comprises: color and texture modeling of the sky area in the video frame to extract sky texture features including cloud shape and saturation; detecting precipitation particle tracks by the optical flow method combined with pixel intensity change; determining the visibility condition of the scene to be one of the qualitative categories of clear, hazy and foggy according to the scene's distant scene definition and contrast change; determining the qualitative state of wind effect and no wind by combining the swing mode of flags, trees and other flexible objects in the scene; combining the sky texture features, precipitation particle tracks, visibility qualitative categories and wind effect qualitative states into multi-dimensional weather features, and comparing the multi-dimensional weather features with historical weather records in the external metadata according to a preset matching rule, wherein the matching rule defines the corresponding relationship threshold between the qualitative categories in the video and the quantitative values of the external data; generating the weather consistency data stream according to the matching degree.

[0013] Preferably, the step of aggregating to form the multi-dimensional consistency vector comprises: setting a unified global time reference for the light consistency, weather consistency and machine state consistency data streams; resampling the three types of data streams with different update frequencies by using an interpolation method to make all frame time stamps have aligned data points; combining the three consistency data points aligned at each frame time stamp into a fixed-dimensional vector and embedding the corresponding traceability index identifier of the data points to form a multi-dimensional consistency vector data stream carrying light, weather and machine state consistency information.

[0014] Preferably, the dynamic weighted fusion process of generating the video consistency sequence comprises: configuring the source credibility and feature stability weight corresponding to each of the light, weather and machine state dimensions in the multi-dimensional consistency vector; the feature stability weight is dynamically updated according to the comprehensive calculation of the authority of the data source, the stability of the data link and the volatility of the data in the time sequence; the fusion process is to multiply each dimension value of the multi-dimensional consistency vector at each time by the corresponding weight value, then add all the products and normalize to fuse into a single scalar value; the scalar values at all times are arranged in time sequence to form the video consistency sequence.

[0015] Preferably, the step of generating the video evidence chain verification report comprises: calculating the consistency mean of each segment, and comparing it with the high and low thresholds dynamically set according to the statistical distribution of the entire sequence, and the segment below the low threshold is marked as a risk point; the consistency mean of each segment is taken as a data point, and a continuous consistency curve is drawn with time as the horizontal axis; the global reliability score is calculated by time-weighted averaging of the entire video consistency sequence, wherein the weight of the risk point marked segment is reduced according to a predetermined proportion; finally, the consistency curve, the global reliability score and the risk point reason of specific inconsistent dimension determined by backtracking the multi-dimensional consistency vector before fusion are integrated into the video evidence chain verification report and output.

[0016] The video evidence chain consistency verification system based on spatiotemporal metadata is used to perform any one of the video evidence chain consistency verification methods based on spatiotemporal metadata, and comprises:

[0017] A data acquisition and alignment module is configured to acquire a video file to be verified, and separate the video file into a pixel stream and internal metadata after decoding.

[0018] A time and location label based on the internal metadata is combined with a scene element identified by the pixel stream to generate an adaptive search window; a plurality of external data sources are called in parallel within the window to perform spatiotemporal alignment, and a traceability path is established and fused into structured external metadata according to the reliability and feature stability.

[0019] A feature analysis module is configured to extract a shadow trajectory from the pixel stream through illumination analysis, generate a shadow consistency data stream in combination with the sun azimuth and height, obtain weather features through weather identification, generate a weather consistency data stream in combination with historical records, and obtain a frame displacement sequence through jitter detection, and generate a holding device state consistency data stream in combination with sensor information.

[0020] A fusion determination module is configured to aggregate the three types of consistency data streams to generate a multi-dimensional consistency vector under a global time reference according to the frame time stamp; the multi-dimensional consistency vector is weighted and fused into a video consistency sequence based on the source reliability and feature stability, and a traceability index is retained; the consistency level of the video consistency sequence is determined segment by segment, and a video evidence chain verification report is output.

[0021] Compared with the prior art, the present application has the following advantages:

[0022] 1、The present application performs multidimensional cross comparison between the pixel content of the video and the objective physical world data from the outside and independent, changes the traditional mode relying on the subjective experience of the identification personnel and the detection of digital artifacts into a data-driven verification based on the principles of natural science, which is quantifiable and reproducible. This makes the verification conclusion no longer a vague inference, but an objective determination with a solid scientific basis, greatly enhancing the reliability, admissibility and court persuasiveness of the evidence.

[0023] 2、The present application provides a complete set of automatic processing procedures, from video decoding, data acquisition, multi-dimensional feature extraction to the final fusion determination and report generation, which minimizes human intervention. It can quickly process a large amount of video data, and automatically screen and highlight the fragments with spatiotemporal consistency risks, so that the identification personnel can be freed from tedious, time-consuming and error-prone manual information retrieval and comparison work, and focus their efforts on in-depth analysis of high-risk points, thereby improving the efficiency of digital evidence review by 100 times or even 1000 times.

[0024] 3、Unlike traditional forensics technology that only focuses on file encoding, compression traces and other digital levels, the present application goes deep into the situational authenticity level of video content. By checking whether the physical environment presented by the video picture conforms to the physical laws at the time and place it claims, the present application can effectively identify high-level fake videos that are seamless at the digital level but contradict the physical reality, providing a new and deeper means of counteracting false evidence. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 Figure 1 is a flowchart of the video evidence chain consistency verification method based on spatiotemporal metadata of the present application;

[0026] Figure 2 Figure 2 is a structural diagram of the video evidence chain consistency verification system based on spatiotemporal metadata of the present application. DETAILED DESCRIPTION

[0027] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0028] Embodiment one

[0029] As Figure 1The embodiment illustrates a video evidence chain consistency verification method based on spatiotemporal metadata in detail, and is specifically applied to public security video forensics, and comprises the following steps: decoding a video file to be verified to separate pixel stream and internal metadata; generating an adaptive search window according to time and location tags and scene elements of the pixel stream; performing spatiotemporal alignment on the window and establishing a traceability path by calling multiple source external data in parallel, and forming structured external metadata by weighted fusion according to source credibility and feature stability; performing light analysis on the pixel stream to extract shadow tracks and generate light consistency data stream by combining the sun azimuth and elevation, performing weather identification on the video frame to generate weather consistency data stream by combining historical records, and performing jitter detection on interframe displacement to generate holding machine state consistency data stream by combining sensor information; aggregating three types of consistency data streams to obtain a multi-dimensional consistency vector under a global time reference according to frame timestamps, and weighted fusion based on credibility and stability to form a video consistency sequence, and retaining a traceability index; determining the consistency level of the consistency sequence segment by segment, and outputting a video evidence chain verification report.

[0030] Further, the process of generating an adaptive search window first extracts timestamps and location tags contained in the video file by analyzing internal metadata. At the same time, the pixel stream is analyzed frame by frame. In a preferred embodiment, scene element recognition uses a pre-trained convolutional neural network such as the YOLO series model for target detection to identify geographical markers and dynamic targets. Sparse optical flow method is used to track the motion of stable feature points in the background to estimate the camera motion, and then Kalman filter is used to track the independent motion trajectory of the foreground dynamic target, so as to separate and calculate the speed and direction parameters of each target, providing the core basis for the dynamic adjustment of the window. By analyzing various lighting clues such as shadows, highlights and ambient light, the approximate direction and relative intensity of the dominant light source in the scene can also be estimated.

[0031] The timestamps and location tags are obtained by analyzing internal metadata, and the adaptive search window is generated by combining pixel stream frame-by-frame analysis. In scene element recognition, a pre-trained convolutional neural network is used to detect geographical markers and dynamic targets; in motion trajectory detection, sparse optical flow method is used to estimate the camera motion, and Kalman filter is used to track the independent trajectory of the foreground target to extract speed and direction parameters; at the same time, the direction and intensity of the dominant light source are estimated by combining shadow and highlight clues.

[0032] By jointly modeling target detection, optical flow, and Kalman filtering, the time interval is adaptively contracted and the spatial range is adaptively focused, significantly reducing the search range and external data calling overhead; by decoupling camera self-motion estimation and illumination cues, the robustness to fast motion and complex illumination scenes is improved, reducing false positives and false negatives; by standardizing the time, space, and environmental feature outputs, alignment and reuse with multi-source data are facilitated, improving the accuracy and efficiency of subsequent consistency checking; the overall scheme can be executed online, with good real-time performance and scalability.

[0033] On this basis, the time interval and spatial range of the search window are dynamically adjusted according to the change characteristics of the motion trajectory. In a preferred embodiment, this adjustment process is implemented through a multi-level rule mechanism. The multi-rule mechanism defines multiple motion states in advance, such as a stationary or extremely slow state with a speed of less than 5 pixels per second, a slow state with a speed between 5 and 30 pixels per second, a medium speed state with a speed between 30 and 100 pixels per second, and a fast state with a speed greater than 100 pixels per second. Each motion state is pre-configured with a specific time interval and spatial range corresponding to it. The specific correspondence is as shown in Table 1:

[0034] Table 1 Multi-level rule mechanism table

[0035] Motion state Time interval length (sec) Spatial range radius (m) Stationary or very slow 5 50 Slow 3 150 Medium 1.5 400 Fast 1 800

[0036] After completing the dynamic correction of the time interval and spatial range, these adjusted parameters are combined with the environmental features extracted in the early stage for further data standardization processing. Through a mapping mechanism, time information, spatial boundaries, environmental features, and illumination conditions are unified into a standardized data structure. In a specific embodiment, this standardized data structure is organized in JSON format. Its main fields include timestamp, which is an ISO 8601 format string; geographic boundary, which is a GeoJSON polygon format; motion vector, which contains an object with two floating point number fields of speed and direction; and environmental features, which contain a weather condition text label and an illumination estimation vector object. In this way, a unified and explicit data interface is provided for all subsequent analysis modules. The adaptive search window formed ultimately is a data structure that integrates multi-dimensional data, providing a precise reference framework for consistency checking of video evidence chains.

[0037] By analyzing internal metadata and pixel streams, combining target detection, optical flow tracking, and Kalman filtering, the time interval and spatial range of the search window are dynamically adjusted, and the motion parameters and environmental features are unified into a standardized data structure. The adaptive search window generated ultimately integrates time, space, and environmental multi-dimensional information, providing a precise and unified reference framework for video evidence chain consistency checking.

[0038] Further, the process of outputting structured external metadata includes: accessing external databases through a data interface, the external databases including a satellite positioning database, a meteorological service database, and a sensor log database; assigning an initial weight to each traceability path according to the authority of the data source and the stability of the data link, wherein the authority is quantified according to a preset source rating table, the data link stability is quantified according to the real-time success rate and delay of data calling, and the initial weight is derived from the weighted sum of the quantified authority and data link stability; dynamically adjusting the weight in combination with the fluctuation evaluation result of the data in the time sequence, the adjustment method being to apply a multiplicative attenuation factor to the initial weight according to the fluctuation evaluation result, wherein the greater the fluctuation, the closer the attenuation factor is to zero; and outputting a structured external metadata stream with unified dimensions and a credibility reference parameter through a normalized weighted fusion algorithm.

[0039] Further, the process of outputting structured external metadata first needs to establish a connection with multi-source external databases, which usually include satellite positioning databases, meteorological service databases, and sensor log databases.

[0040] For the satellite positioning database, high authority: including official global navigation satellite systems (such as GPS, Beidou); medium authority: API data from large, reputable commercial map service providers (such as Google Maps, Gaode Map); low authority: data from open source community map projects or third-party applications with low credibility;

[0041] For the meteorological service database, high authority: historical weather records from official agencies such as national meteorological bureaus or the World Meteorological Organization (WMO); medium authority: historical data from well-known commercial meteorological service companies; low authority: online weather aggregation services from small, regional weather stations;

[0042] For the sensor log database, high authority: log data from professionally calibrated, built-in high-precision sensors (such as industrial-grade gyroscopes, accelerometers); medium authority: built-in sensor logs from mainstream consumer electronics devices (such as well-known brand smartphones); low authority: data from uncalibrated, inexpensive external sensor modules or generated through software simulation;

[0043] After data calling is completed, an initial weight is assigned to the traceability path according to the properties of different data sources, and the weight mainly refers to the authority of the data source and the stability of the data link.

[0044] The data link stability (S) is quantified from the real-time success rate (R) and the average delay (T, unit: ms) of data calling, and the calculation formula is: The initial weight is weighted by the authority (A) and the stability (S) of the data link of the main reference data source; Specifically: W init = 0.6 S + 0.4 A; Wherein, W init represents the initial weight; min() represents taking the minimum value; For A, high authority takes 0.9; medium authority takes 0.6; low authority takes 0.4;

[0045] On this basis, a cross-modal volatility correlation analysis mechanism is also introduced to dynamically correct the weight. The specific formula is: W new = D W init ; Wherein, W new represents the adjusted weight; D represents the dynamic weight multiplier;

[0046] Specifically, the Bhattacharyya distance of the color histogram between consecutive video frames is calculated to quantify the visual change volatility, and a visual change score sequence is obtained. At the same time, the point-by-point change rate of the time series of external meteorological data such as wind speed is calculated, and an external data change score sequence is obtained. The logical rules of correlation comparison are as follows: 1. If the peak value appearance time of the visual change score and the external data change score is not more than, for example, 2 seconds within a time window of, for example, 10 seconds, it is determined that there is a strong positive correlation, and the weight of the external data is significantly increased (D takes 1.5) during this period. 2. If one score sequence changes dramatically while the other sequence remains stable, it is determined that there is a mismatch, and the weight is significantly reduced (D takes 0.3). 3. If both sequences remain stable for a long time, the weight remains at the baseline level (D takes 1).

[0047] A data score sequence is determined to be in a "stable" state if its standard deviation σ within a time window is lower than a preset low volatility threshold, which is 0.1 in this embodiment; When σ < 0.1, it is considered that the sequence remains stable; A data score sequence is determined to be in a "dramatic change" state if its standard deviation σ within a time window is higher than a preset high volatility threshold, which is 0.8 in this embodiment; That is, when σ > 0.8, it is considered that the sequence changes dramatically.

[0048] "Mismatch" is defined as a significant contradiction between visual change and external data change. In combination with the above quantified states, let the standard deviation of the visual change score sequence be σ vis , and the standard deviation of the external data change score sequence be σ ext ; The "mismatch" determination condition is that, within the same time window, one of the following logics is satisfied;

[0049] σ vis > 0.8 and σ ext < 0.1;

[0050] σ vis <0.1 and σ ext >0.8;

[0051] Finally, a normalized weighted fusion algorithm is used to integrate data from different sources and with different weights to form an external metadata stream with consistent dimensions, along with credibility reference parameters, providing reliable, stable and traceable data support for subsequent verification.

[0052] By connecting to external databases such as satellite positioning, meteorological services, and sensor logs, initial weights based on authority and link stability are set for the tracing path. These weights are then dynamically adjusted using cross-modal correlation analysis of visual change scores and external data change scores. Finally, a unified external metadata stream is generated through normalized weighted fusion, along with a credibility parameter, providing stable and traceable data support for subsequent verification.

[0053] Furthermore, the steps for generating the illumination consistency data stream include: under the condition that a single main light source is detected in the scene, extracting the shadow region formed by static objects in the video frame and calculating the boundary trajectory of the shadow region through image brightness and gradient analysis; geometrically comparing the geometric direction and length change of the shadow boundary trajectory with the corresponding spatiotemporal solar azimuth and elevation angle parameters obtained from external metadata; calculating a quantified illumination matching index based on the difference in the angle and length change rate between the two vector directions; and sorting the illumination matching index calculated at each time point according to the corresponding timestamp to generate the time series, forming the illumination consistency data stream.

[0054] Furthermore, the illumination matching degree index is calculated as follows:

[0055] Among them, M light Indicates the illumination matching degree value; θ diff ΔL represents the angle between the direction of the shadow observed in the video and the theoretical solar azimuth angle calculated from external metadata, with a value ranging from [0, 90] degrees; ratio This is the ratio of the rate of change of shadow length in the video to the theoretically calculated rate of change of shadow length.

[0056] Furthermore, the process of generating a consistent lighting data stream first identifies stable lighting cues formed by sunlight that can be analyzed, such as shadows cast by stationary objects or highlights on smooth object surfaces.

[0057] Subsequently, before performing the light cue analysis, the camera self-motion estimation module is first activated. In a preferred embodiment, the self-motion estimation module employs a method based on optical flow or feature points such as ORB features, tracks the stable feature points between frames, and calculates the camera's rotation, translation and scale parameters using the random sample consensus algorithm. Before extracting the displacement trajectories of the light cues, the estimated camera motion parameters are used to compensate the pixel coordinates, which maximally suppresses the trajectory changes introduced by the camera motion. Meanwhile, a camera motion compensation confidence score ranging from 0 to 1 is generated according to the internal state of the motion estimation algorithm, such as the number of successfully tracked feature points and the re-projection error.

[0058] After the motion compensation is completed, one or more reference objects and the light cues formed by them are locked by image segmentation and object tracking techniques. At the same time, the surface carrying the light cues is analyzed for local geometric features, such as the degree of perspective distortion of the surface texture, to estimate its confidence as an ideal plane, and a surface confidence score ranging from 0 to 1 is output. The method extracts and tracks the relative displacement trajectories of the key feature points of the light cues, such as the tips of the shadows, and matches their change trends with the theoretical motion trends of the sun calculated according to the external metadata.

[0059] Subsequently, a dynamic tolerance rating mechanism is used to convert the degree of fitting of the trend comparison into a quantifiable light matching degree indicator. First, the camera motion compensation confidence score and the surface confidence score are combined to calculate the overall uncertainty level of the current frame or segment. The overall uncertainty level U is calculated as follows: U = 1 - (0.7 · C motion + 0.3 · C surface ); where C motion and C surface represent the camera motion compensation confidence score and the surface confidence score, respectively.

[0060] This uncertainty level directly determines the tolerance Tol θ of the rating scale, for example, Tol θ = 5 + 10 · U; when the overall uncertainty level is low, for example, the camera is stable and the shadow-carrying surface is flat, a very strict rating standard is used, requiring a small pixel deviation to obtain a high matching degree. Conversely, when the overall uncertainty level is high, for example, the camera is shaking violently or the shadow-carrying surface is complex, the tolerance is automatically relaxed, allowing a larger pixel deviation to also obtain a reasonable matching degree. The resulting light consistency data stream can dynamically reflect the authenticity of the entire video under the lighting conditions.

[0061] The illumination trend comparison result is converted into a quantitative matching degree index through a dynamic tolerance rating mechanism; the overall uncertainty level is calculated by comprehensively considering the camera motion compensation confidence and the surface confidence, and the rating tolerance is adjusted according to the uncertainty level: strict standards are adopted under low uncertainty, and the tolerance is relaxed under high uncertainty.

[0062] The dynamic tolerance rating mechanism can adaptively adjust the comparison tolerance according to the differences in video acquisition environment and device state, provide strict judgment in the case of high-quality video, and avoid excessive punishment in complex or low-quality scenes. Thus, the robustness and accuracy of the illumination consistency evaluation are improved, and the generated matching degree index is more consistent with the real shooting conditions, providing reliable support for video evidence chain verification.

[0063] Through camera motion compensation, illumination clue extraction and surface confidence evaluation, the shadow or highlight trajectory is matched with the theoretical motion trend of the sun, and the illumination matching degree index is generated in combination with the dynamic tolerance rating. The finally output illumination consistency data stream can adapt to the environment and device conditions, accurately reflect the authenticity of the video in the illumination layer, and provide reliable basis for evidence chain verification.

[0064] Further, the step of generating weather consistency data stream comprises: color and texture modeling of the sky area in the video frame, extracting sky texture features including cloud shape and saturation; detecting precipitation particle trajectories by optical flow method combined with pixel intensity changes; determining the visibility condition of the scene as one of the qualitative categories of clear, hazy and foggy according to the scene's distant definition and contrast change; determining the qualitative state of wind effect and no wind by combining the swing mode of flags, trees and other flexible objects in the scene; combining the sky texture features, precipitation particle trajectories, visibility qualitative categories and wind effect qualitative states into multi-dimensional weather features, and comparing the multi-dimensional weather features with historical meteorological records in external metadata according to a preset matching rule, wherein the matching rule defines the corresponding relationship threshold between the qualitative categories in the video and the quantitative values in the external data; generating weather consistency data stream according to the matching compliance degree.

[0065] The matching rule is shown in Table 2 as follows:

[0066] Table 2 Matching rule table

[0067]

[0068]

[0069] The matching compliance degree is quantified by using a soft threshold function. For example, when the video is determined to be 'hazy' (rule 2-10km), and the external data visibility is V km, the compliance degree is

[0070]

[0071] Further, the process of generating weather consistency data stream first assesses macro weather conditions by analyzing the overall visual features of the video frames, including assessing atmospheric visibility, judging sky conditions, etc., and cross-verifying with external meteorological records.

[0072] On this basis, the method identifies precipitation events by detecting physical state changes of key surfaces in the scene, such as analyzing road reflectivity to determine whether there is a wet state. To enhance the robustness of wind effect judgment, the method uses an opportunistic multi-feature fusion strategy to search for motion patterns of vegetation, flexible objects or particle flow in the scene in turn. If none of the above features exist, it is determined that the current scene cannot extract effective wind effect features, and the weight of this dimension is set to neutral or zero.

[0073] To effectively compare video visual features with external data, a set of pre-set matching rules are used. Examples of these rules are as follows: regarding visibility, if the video analysis classifies the visibility as moderate blur, the rule requires the visibility value recorded in the external meteorological database to be within the range of 2 to 6 kilometers. Regarding precipitation, if the video analysis detects that the road surface has obvious wet reflection, the rule requires that the cumulative rainfall in the past one hour recorded in the external database be greater than 0.1 millimeters. Regarding sky conditions, if the video analysis classifies the sky as mostly covered by clouds, the rule requires that the cloud cover recorded in the external database should be higher than sixty percent. Finally, the multi-dimensional weather features will be compared with the historical records of the external meteorological database according to these matching rules. The degree of conformity after quantization is output as a weather consistency data stream.

[0074] By analyzing video frames to obtain multi-dimensional weather features such as visibility, precipitation and sky conditions, and combining with flexible object or surface state to identify wind effect, and then comparing with external meteorological records, the final output of the weather consistency data stream can quantify the degree of fit between the weather performance in the video and the real meteorological conditions, providing reliable evidence for the environmental level of the video evidence chain verification.

[0075] Further, the process of generating the consistency data stream of the holding state aims to verify whether the motion characteristics of the video picture are consistent with the physical motion recorded by the sensor. This process first parses and extracts the time series data of the gyroscope and accelerometer from the internal metadata of the video. If there is no sensor data, the generation of the consistency stream is aborted, and its weight is set to zero. If there is, the sensor data is integrated and the attitude is solved to construct a high-frequency camera six-degree-of-freedom motion trajectory recorded by the physical sensor. At the same time, the aforementioned camera self-motion estimation module is used to independently calculate a camera motion trajectory from the video pixel stream. The core of the holding state consistency is to compare the two motion trajectories. By using dynamic time warping and other sequence alignment algorithms, the similarity of the two trajectories in the rotation and translation components is calculated. If the two trajectories are highly consistent, it indicates that the visual motion is consistent with the physical motion, and a high consistency score is given. If there is a significant difference, for example, the pixel stream trajectory is much smoother than the sensor trajectory, it may mean that the video has been post-processed by digital stabilization, and a low consistency score is given. These scores are arranged in chronological order to form the holding state consistency data stream.

[0076] Further, the normalized distance D between the two trajectories (visual motion trajectory and physical motion trajectory) is calculated by the dynamic time warping (DTW) algorithm dtw , which is converted into a consistency score M by the following formula holding ;

[0077]

[0078] where T dist is a preset distance threshold, for example, 10% of the diagonal length of the scene.

[0079] The high consistency score represents the case where M holding > 0.8, the low consistency score represents the case where M holding < 0.3, and the intermediate consistency score is between the two.

[0080] By extracting the gyroscope and accelerometer data in the internal metadata of the video to generate the physical motion trajectory, and independently calculating the visual motion trajectory from the pixel stream, the similarity of the two is compared by dynamic time warping. When the trajectories are highly consistent, a high score is given, and when there is a significant difference, a low score is given, and finally a holding state consistency data stream arranged in chronological order is formed.

[0081] By extracting the gyroscope and accelerometer data to generate the physical motion trajectory, and independently calculating the visual motion trajectory from the pixel stream, the similarity of the two is compared by dynamic time warping algorithm. When the trajectories are highly consistent, a high score is given, and when there is a significant difference, a low score is given, and finally a holding state consistency data stream arranged in chronological order is formed.

[0082] Further, the step of aggregating to form the multi-dimensional consistency vector comprises: setting a unified global time reference for the three types of data streams of lighting consistency, weather consistency and holding state consistency; using an interpolation method to resample the three types of data streams with different update frequencies, so that there are aligned data points on all frame timestamps; combining the three consistency data points aligned on each frame timestamp into a fixed-dimensional vector, and embedding the traceability index corresponding to the data points to form a multi-dimensional consistency vector data stream carrying lighting, weather and holding state consistency information

[0083] Further, the process of aggregating to form the multi-dimensional consistency vector first establishes a unified global time reference. Since the update frequencies of the three types of consistency data streams are different, they are resampled by a specific interpolation method to achieve time alignment. In a preferred embodiment, for weather and lighting data streams which usually change continuously, linear interpolation is used for resampling. For holding state data streams which may have abrupt changes, a forward filling nearest neighbor interpolation method is used to retain their step change characteristics. After alignment, the consistency values at the same timestamp are combined into a fixed-dimensional vector structure with traceability index.

[0084] By establishing a unified time reference, and using linear interpolation for resampling of continuous change data streams and nearest neighbor interpolation for abrupt change data streams, time alignment of the three types of consistency data is achieved.

[0085] Further, the dynamic weighted fusion process of generating the video consistency sequence comprises: configuring the source credibility and feature stability weight for each of the three dimensions of lighting, weather and holding state in the multi-dimensional consistency vector; the feature stability weight is dynamically updated according to the comprehensive calculation of the authority of the data source, the stability of the data link and the volatility of the data in the time sequence; the fusion process is to multiply each dimension value of the multi-dimensional consistency vector at each time by its corresponding weight value, then add all the products and normalize them to fuse into a single scalar value; all the scalar values at all times are arranged in time sequence to form the video consistency sequence.

[0086] The feature stability weight is directly related to the maximum weight W of the external data it depends on final , for example, the stability weight W of the lighting consistency data stream stable_light is equal to the W final value of the solar position data it depends on, and the fusion process is to calculate the final consistency score value at time t where M i is the value of each consistency data stream at this time;

[0087] Further, the dynamic weighting fusion process of generating the video consistency sequence first assigns weights. The weight system consists of two parts: one is the source credibility weight, which is statically set according to the authority level of the external data provider; the other is the feature saliency weight, which is a dynamic weight. Specifically, the sharpness of the shadow in the light clue can be quantified by calculating the average gradient amplitude of its edge region. A gradient amplitude threshold is preset, when the calculated value is higher than the threshold, it is determined to be clear, and the feature saliency weight is set to a high value such as 1.2; if it is lower than the threshold, it is determined to be fuzzy, and the weight is set to a low value such as 0.5; if the shadow is blocked by a large area and cannot extract effective edges, the weight is set to 0. In this way, the visual quality is mapped to a specific weight adjustment value.

[0088] After completing the weight assignment, enter the fusion calculation stage. Multiply the three-dimensional values of the multi-dimensional consistency vector at a certain time by the corresponding weights one by one, then accumulate and summarize the results, and generate a single scalar value through normalization processing. This calculation process will continue to be executed in the full time sequence range of the video, all scalar values are arranged in timestamp order, and finally a continuous video consistency sequence is formed. The generated video consistency sequence integrates multi-dimensional features and improves credibility through a dynamic weight mechanism, providing a solid data foundation for subsequent segmentation judgment.

[0089] Through the double weight adjustment of source credibility and feature saliency, the multi-dimensional consistency vector is dynamically weighted and fused into a video consistency sequence. This sequence integrates light, weather and device state features while improving credibility, providing a reliable data foundation for subsequent segmentation judgment and verification report generation.

[0090] Further, the steps of generating a video evidence chain verification report include: calculating the consistency average of each segment and comparing it with the high and low thresholds dynamically set according to the statistical distribution of the entire sequence, and the segments below the low threshold are marked as risk points; plot a continuous consistency curve with the consistency average of each segment as a data point and time as the horizontal axis; calculate the global credibility score by time-weighted averaging the entire video consistency sequence, wherein the weight of the risk point marked segment is reduced according to the predetermined proportion (20% in this embodiment); finally, integrate the consistency curve, global credibility score and specific inconsistency dimension risk point reason determined by backtracking the multi-dimensional consistency vector before fusion into the video evidence chain verification report and output

[0091] Further, the process of generating the video evidence chain verification report first segments the video consistency sequence, calculates the consistency mean of each segment. Then, compare these means with the high and low thresholds dynamically set based on the distribution statistics of the entire sequence, mark potential risk points. Specifically, the low threshold is dynamically set to the tenth percentile value of the entire video consistency sequence, while the high threshold is set to the ninetieth percentile value. Any segment mean below the tenth percentile threshold will be marked as a potential risk point. After completing the risk segment marking, draw a continuous consistency curve to intuitively reflect the consistency trend. To quantify the overall credibility, the time-weighted average of the entire video consistency sequence is calculated, where the weight of the risk point is proportionally reduced, and finally a global credibility score is generated. Finally, combine the consistency curve, global credibility score and the multi-dimensional consistency vector obtained by backtracking to accurately locate the specific reasons for the occurrence of risk points. Based on the above analysis results, a complete video evidence chain verification report is formed, and output in a visual and numerical way, providing clear, traceable and interpretable basis for judicial evidence, compliance review and other applications.

[0092] By segmenting the consistency mean and setting dynamic thresholds to mark risk segments, combining the consistency curve with the time-weighted average to generate a global credibility score, and locating the risk causes by backtracking the multi-dimensional consistency vector. The final output verification report is presented in a visual and numerical form, providing clear, traceable and interpretable basis for judicial evidence and compliance review.

[0093] Embodiment two

[0094] Figure 2 The structure diagram of the video evidence chain consistency verification system based on spatiotemporal metadata of the application; the embodiment provides a video evidence chain consistency verification system based on spatiotemporal metadata, comprising:

[0095] The data acquisition and alignment module is used for acquiring the video file to be verified, and separating the decoded video file into a pixel stream and internal metadata;

[0096] The time and location labels based on internal metadata generate an adaptive search window in combination with the scene elements identified by the pixel stream; the window is used to call multiple source external data in parallel to perform spatiotemporal alignment, establish a traceability path, and fuse the path into structured external metadata according to the credibility and feature stability;

[0097] The feature analysis module is used for extracting shadow trajectories from the pixel stream through illumination analysis, generating shadow consistency data stream in combination with the sun azimuth and height; obtaining weather features through weather identification, generating weather consistency data stream in combination with historical records; obtaining interframe displacement sequences through jitter detection, generating holding state consistency data stream in combination with sensor information;

[0098] The fusion determination module aggregates the three types of consistent data streams into a multi-dimensional consistency vector under a global time reference according to the frame timestamps; fuses the video consistency sequence based on the source credibility and feature stability, and retains the traceability index; determines the consistency level for the video consistency sequence in segments, and outputs a video evidence chain verification report.

[0099] Further, the step of generating an adaptive retrieval window comprises: parsing the internal metadata to obtain the timestamp and position label, and performing scene element recognition on the pixel stream to extract environmental features, lighting conditions, and motion trajectories; dynamically adjusting the time interval length and spatial range size of the retrieval window according to the rate and direction parameters of the motion trajectory, wherein the length of the time interval is inversely proportional to the rate of the motion trajectory, and the size of the spatial range is proportional to the rate of the motion trajectory; mapping the adjusted time, spatial parameters and environmental features into a standardized data structure to form an adaptive retrieval window.

[0100] Further, the process of outputting structured external metadata comprises: accessing external databases through a data interface, the external databases including satellite positioning databases, weather service databases, and sensor log databases; assigning an initial weight to each traceability path according to the authority of the data source and the stability of the data link, wherein the authority is quantified according to a pre-set source rating table, the data link stability is quantified according to the real-time success rate and delay of data calling, and the initial weight is derived from the weighted sum of the quantified authority and data link stability; dynamically adjusting the weight by combining the volatility evaluation results of the data in the time sequence, the adjustment method being to apply a multiplicative decay factor according to the volatility evaluation results on the basis of the initial weight, wherein the greater the volatility, the closer the decay factor is to zero; outputting a structured external metadata stream with unified dimensions and attached credibility reference parameters through a normalized weighted fusion algorithm.

[0101] Further, the step of generating the lighting consistency data stream comprises: under the condition that the scene has a single main light source, extracting the shadow area formed by static objects in the video frame through image brightness and gradient analysis, and calculating the boundary trajectory of the shadow area; geometrically comparing the geometric direction and length change of the shadow boundary trajectory with the corresponding spatio-temporal solar azimuth and elevation angle parameters obtained from the external metadata; calculating a quantitative lighting matching degree index according to the difference between the included angle of the two vector directions and the length change rate; sorting the lighting matching degree indexes calculated at each time point according to the corresponding timestamps to generate the time sequence, forming the lighting consistency data stream.

[0102] Further, the step of generating the weather consistency data stream comprises: color and texture modeling of the sky region in the video frame, extracting sky texture features including cloud shape and saturation; detecting precipitation particle trajectories by optical flow method combined with pixel intensity changes; determining the visibility condition of the scene as one of the qualitative categories of clear, hazy and foggy according to the scene's distant scene definition and contrast changes; determining the qualitative state of wind effect and no wind by combining the waving patterns of flags, trees and other flexible objects in the scene; combining the sky texture features, precipitation particle trajectories, visibility qualitative categories and wind effect qualitative state into multi-dimensional weather features, and comparing the multi-dimensional weather features with historical meteorological records in external metadata according to a preset matching rule, wherein the matching rule defines the corresponding relationship threshold between the qualitative categories in the video and the quantitative values in the external data; generating the weather consistency data stream according to the matching degree.

[0103] Further, the step of aggregating the multi-dimensional consistency vector comprises: setting a unified global time reference for the three types of data streams of lighting consistency, weather consistency and machine state consistency; using an interpolation method to resample the three types of data streams with different update frequencies, so that there are aligned data points on all frame timestamps; combining the three consistency data points aligned on each frame timestamp into a fixed-dimensional vector, and embedding the corresponding traceability index of the data points to form a multi-dimensional consistency vector data stream carrying lighting, weather and machine state consistency information.

[0104] Further, the dynamic weighted fusion process of generating the video consistency sequence comprises: configuring the source credibility and feature stability weight of the lighting, weather and machine state dimensions in the multi-dimensional consistency vector respectively; the feature stability weight is dynamically updated according to the evaluation results of the authority of the data source, the stability of the data link and the volatility of the data in the time sequence; the fusion process is to multiply the value of each dimension of the multi-dimensional consistency vector at each time with its corresponding weight value, then add all the products and normalize them to fuse into a single scalar value; all the scalar values at all times are arranged in time sequence to form the video consistency sequence.

[0105] Further, the step of generating the video evidence chain verification report comprises: calculating the consistency mean of each segment and comparing it with the high and low thresholds dynamically set according to the statistical distribution of the entire sequence, and the segments below the low threshold are marked as risk points; drawing a continuous consistency curve with the consistency mean of each segment as the data point and time as the horizontal axis; calculating the global credibility score by time-weighted averaging of the entire video consistency sequence, wherein the weight of the risk point marked segment is reduced according to the predetermined proportion; finally, integrating the consistency curve, global credibility score and specific inconsistent dimension risk point reasons determined by tracing the multi-dimensional consistency vector before fusion into the video evidence chain verification report and outputting.

[0106] While embodiments of the application have been shown and described, it is to be understood that the application is not limited to the details of the embodiments described, since numerous changes, modifications, substitutions and variations can be made thereto without departing from the spirit and scope of the application as defined by the appended claims and their equivalents.

Claims

1. A method for checking the consistency of a video evidence chain based on spatiotemporal metadata, characterized in that, The method comprises the following steps: Obtain the video file to be verified, and separate the video file into a pixel stream and internal metadata after decoding; Generate an adaptive retrieval window based on the time and location tags of the internal metadata and the scene elements identified from the pixel stream; Call multiple external data sources in parallel within the window to perform spatio-temporal alignment, establish a traceability path, and fuse the external data into structured external metadata based on the credibility and feature stability; Extract shadow trajectories from the pixel stream through illumination analysis, and generate a shadow consistency data stream based on the sun's azimuth and altitude; Obtain weather features through weather identification, and generate a weather consistency data stream based on historical records; obtain a frame displacement sequence through jitter detection, and generate a device state consistency data stream based on sensor information; Aggregate the three types of consistency data streams into a multi-dimensional consistency vector based on frame timestamps under a global time reference; Fuse the consistency vector into a video consistency sequence based on the credibility of the source and the stability of the features, and retain the traceability index; determine the consistency level of the video consistency sequence segment by segment, and output a video evidence chain verification report.

2. The video evidence chain consistency verification method based on spatiotemporal metadata according to claim 1, characterized in that, The step of generating an adaptive retrieval window comprises the following steps: parsing the internal metadata to obtain timestamps and location tags, and identifying scene elements from the pixel stream to extract environmental features, lighting conditions, and motion trajectories; dynamically adjusting the time interval length and spatial range size of the retrieval window based on the speed and direction parameters of the motion trajectory, wherein the length of the time interval is inversely proportional to the speed of the motion trajectory, and the size of the spatial range is proportional to the speed of the motion trajectory; mapping the adjusted time, spatial parameters, and environmental features into a standardized data structure to form an adaptive retrieval window. 3.The video evidence chain consistency checking method based on spatiotemporal metadata according to claim 1, characterized in that, The process of outputting structured external metadata comprises the following steps: accessing external databases through a data interface, wherein the external databases include satellite positioning databases, meteorological service databases, and sensor log databases; assigning an initial weight to each traceability path based on the authority of the data source and the stability of the data link, wherein the authority is quantified based on a pre-set source rating table, the data link stability is quantified based on the real-time success rate and delay of data calling, and the initial weight is derived from the weighted sum of the quantified authority and data link stability; dynamically adjusting the weight based on the volatility evaluation results of the data in the time sequence, wherein the adjustment method is to apply a multiplicative attenuation factor based on the initial weight according to the volatility evaluation results, wherein the greater the volatility, the closer the attenuation factor to zero; output a structured external metadata stream with unified dimensions and attached credibility reference parameters through a normalized weighted fusion algorithm.

4. The video evidence chain consistency checking method based on spatio-temporal metadata according to claim 1, characterized in that, The step of generating the light consistency data stream comprises: under the condition of detecting that the scene has a single main light source, extracting a shadow area formed by a static object in the video frame by image brightness and gradient analysis, and calculating a boundary track of the shadow area; geometrically comparing the geometric direction and length change of the shadow boundary track with corresponding spatio-temporal solar azimuth and altitude angle parameters obtained from external metadata; calculating a quantified light matching degree index according to the difference between the included angle of the two vector directions and the length change rate; sorting the light matching degree indexes calculated at each time point according to the corresponding time stamps to generate the time sequence, thereby forming the light consistency data stream.

5. The method for video evidence chain consistency verification based on spatiotemporal metadata according to claim 1, characterized in that, The step of generating the weather consistency data stream comprises: color and texture modeling of the sky area in the video frame, extracting sky texture features including cloud shape and saturation; detecting precipitation particle tracks by the optical flow method combined with pixel intensity change; determining the visibility condition of the scene as one of the qualitative categories of clear, hazy and foggy according to the scene's distant definition and contrast change; determining the qualitative state of wind effect and no wind by combining the swing mode of flags, trees and other flexible objects in the scene; combining the sky texture features, precipitation particle tracks, visibility qualitative categories and wind effect qualitative states into multi-dimensional weather features, and comparing the multi-dimensional weather features with historical weather records in external metadata according to a preset matching rule, wherein the matching rule defines the corresponding relationship threshold between the qualitative categories in the video and the quantitative values of the external data; generating the weather consistency data stream according to the matching degree.

6. The method for video evidence chain consistency verification based on spatiotemporal metadata according to claim 1, characterized in that, The step of aggregating the multi-dimensional consistency vector comprises: setting a unified global time reference for the light consistency, weather consistency and machine state consistency data streams; using an interpolation method to resample the three types of data streams with different update frequencies, so that there are aligned data points on all frame time stamps; combining the three consistency data points aligned on each frame time stamp into a fixed-dimensional vector, and embedding the corresponding traceability index identifier of the data points, thereby forming a multi-dimensional consistency vector data stream carrying light, weather and machine state consistency information.

7. The method of claim 1, wherein, The dynamic weighted fusion process of generating the video consistency sequence comprises: configuring the source credibility and feature stability weight corresponding to each of the light, weather and machine state dimensions in the multi-dimensional consistency vector; the feature stability weight is dynamically updated according to the comprehensive calculation of the authority of the data source, the stability of the data link and the volatility of the data in the time sequence; the fusion process is to multiply each dimension value of the multi-dimensional consistency vector at each time by the corresponding weight value, then add all the products and normalize them to fuse into a single scalar value; the scalar values at all times are arranged in sequence to form the video consistency sequence.

8. The method for video evidence chain consistency verification based on spatio-temporal metadata according to claim 1, characterized in that, The steps of generating the video evidence chain verification report include: calculating the consistency mean of each segment and comparing it with the high and low thresholds dynamically set according to the statistical distribution of the entire sequence, and the segments below the low threshold are marked as risk points; the consistency mean of each segment is taken as a data point, and the time is taken as the horizontal axis to draw a continuous consistency curve; the global credibility score is calculated by time-weighted average of the entire video consistency sequence, wherein the weight of the risk point marked segment is reduced according to the predetermined proportion; finally, the consistency curve, the global credibility score and the risk point reason of specific inconsistent dimension by backtracking the multi-dimensional consistency vector before fusion are determined, and are integrated into the video evidence chain verification report and output.

9. A video evidence chain consistency checking system based on spatiotemporal metadata, characterized in that, The system is used for performing the video evidence chain consistency verification method based on spatiotemporal metadata as claimed in any one of claims 1 to 8, comprising: A data acquisition and alignment module is used for acquiring a video file to be verified, and separating the video file into a pixel stream and internal metadata after decoding; based on the time and position labels of the internal metadata, an adaptive search window is generated combining the scene elements identified by the pixel stream; multi-source external data is called in parallel within the window to perform spatiotemporal alignment, establish a traceability path and fuse into structured external metadata according to the credibility and feature stability; A feature analysis module is used for extracting shadow trajectories from the pixel stream through illumination analysis, generating shadow consistency data stream combining the sun azimuth and height; obtaining weather features through weather recognition, generating weather consistency data stream combining historical records; obtaining interframe displacement sequences through jitter detection, generating holding machine state consistency data stream combining sensor information; A fusion judgment module is used for aggregating three types of consistency data streams to generate a multi-dimensional consistency vector under a global time reference according to the frame timestamp; fusing into a video consistency sequence based on the source credibility and feature stability, and retaining the traceability index; determining the consistency level of the video consistency sequence segment by segment, and outputting the video evidence chain verification report.

Citation Information

Patent Citations

  • Aviation video stream target identification processing method and system

    CN119151984A

  • Business weight analysis method and device based on video auditing, equipment and medium

    CN120563060A

  • Assessing video stream quality

    US20210004600A1

  • Segmenting video stream frames

    US20210012114A1

  • Methods and systems for detecting video artifacts

    US9232118B1

Cited By

  • Enterprise video authentication consistency verification method and system based on multi-dimensional data fusion

    CN122049587A

  • Enterprise video authentication consistency verification method and system based on multi-dimensional data fusion

    CN122049587B