Video information monitoring method and device based on semantic skeleton, equipment and medium

By extracting multimodal feature vectors from video, audio, and metadata, generating semantic skeleton fingerprints, and constructing video trajectory maps, the problem of low accuracy in video monitoring in existing technologies is solved, and stable representation and source tracing of video content are achieved.

CN122490462APending Publication Date: 2026-07-31BEIJING ZHONGQING HUAYUN NEW MEDIA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZHONGQING HUAYUN NEW MEDIA TECHNOLOGY CO LTD
Filing Date
2026-04-24
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing video monitoring technologies struggle to effectively identify content created through secondary editing, resulting in low monitoring accuracy and an inability to achieve coordinated monitoring and tracing of multiple video data sources.

Method used

By extracting multimodal feature vectors from video, audio, and metadata, logical consistency verification is performed to generate semantic skeleton fingerprints. Combined with video account information, a video trajectory map is constructed to achieve intelligent monitoring across the entire chain.

Benefits of technology

It improves the overall accuracy of video monitoring, resists pixel-level changes such as occlusion and editing, and achieves accurate video correlation and source tracing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490462A_ABST
    Figure CN122490462A_ABST
Patent Text Reader

Abstract

This application provides a video information monitoring method, apparatus, device, and medium based on a semantic skeleton, relating to the field of video processing technology. The method involves extracting feature vectors from video data in the spatiotemporal dimension, calculating confidence levels for the video semantic feature vector, audio feature vector, and metadata feature vector, respectively; performing logical consistency verification on the video semantic feature vector, audio feature vector, and metadata feature vector; and, when the verification result is normal, dynamically assigning weights based on the verification result and the video confidence level, audio confidence level, and metadata confidence level to construct a semantic skeleton fingerprint; calculating the similarity between the semantic skeleton fingerprint and a preset video database, and associating videos based on the similarity values ​​to obtain associated videos; constructing a video trajectory map based on video account information, associated videos, and associated account information; and generating monitoring and control instructions for video data based on the video trajectory map, thereby improving the accuracy of video monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and particularly to a method, device, equipment, and medium for video information monitoring based on a semantic skeleton. Background Art

[0002] With the rapid development of the Internet and social media, online public opinion has become an important area of concern in social governance. Public opinion monitoring aims to timely discover and evaluate the evolution trend of public opinion by collecting, analyzing, and tracking a large amount of information on network platforms in real time, so as to provide decision-making support for relevant departments. Traditional public opinion monitoring mainly focuses on the capture and sentiment analysis of text information. However, with the rise of short video platforms and the rapid development of artificial intelligence technology, videos have gradually replaced pictures and texts as the main carrier of information dissemination. The popularization of AI-generated videos and the wide spread of a large number of video contents after secondary editing based on the original works have made the form of public opinion dissemination more complex and diverse. Such videos often contain sensitive words or sensitive pictures, which have had an adverse impact on the current video creation ecosystem and dissemination order, and also brought unprecedented challenges to public opinion monitoring. Therefore, the information identification and monitoring of video content have become a key link that urgently needs to be broken through in the public opinion monitoring system.

[0003] In related technical research, existing solutions mostly target specific scenarios and use convolutional neural network models for video recognition and monitoring. However, such neural network models mainly rely on convolutional layers to extract pixel features in video frames, focusing on distinguishing the visual differences between the main object and the background. When the video undergoes secondary creation, such as the picture being blocked, the background being replaced, being spliced by editing, or other visual elements being superimposed, the original pixel structure and background information change, resulting in the model being difficult to effectively identify whether the video is a derivative version of the original content. Rooted in this, the crux of this problem lies in that existing methods are often limited to the independent analysis of single-point video content, only focusing on the shallow pixel features of the picture itself, and failing to achieve the linkage monitoring between multi-source video data. When facing content variation, an effective traceability judgment cannot be formed, resulting in the overall accuracy of video monitoring being relatively low. Summary of the Invention

[0004] This application provides a method, device, equipment, and medium for video information monitoring based on a semantic skeleton, which can perform full-link video recognition and improve the overall accuracy of video monitoring.

[0005] The technical solution of the embodiment of this application is as follows: In the first aspect, the embodiment of this application provides a method for video information monitoring based on a semantic skeleton, and the method includes: Obtain video streams from each platform, and generate video data according to the video streams; Extract the video semantic feature vector, audio feature vector, and metadata feature vector in the spatiotemporal dimension of the video data, and calculate the confidence scores of the video semantic feature vector, the audio feature vector, and the metadata feature vector respectively to obtain the video confidence score, audio confidence score, and metadata confidence score. Logical consistency verification is performed on the video semantic feature vector, the audio feature vector, and the metadata feature vector to obtain the verification result. If the verification result is normal, dynamic weight allocation is performed based on the verification result and the video confidence, audio confidence, and metadata confidence to construct a semantic skeleton fingerprint. The semantic skeleton fingerprint is compared with a preset video database to obtain a similarity value. Based on the similarity value, the videos are associated to obtain the associated videos. Based on the video account information corresponding to the video data, the associated video and the associated account information corresponding to the associated video, a video trajectory map is constructed, and monitoring and control instructions for the video data are generated based on the video trajectory map.

[0006] In the above technical solution, a basic data source is constructed by accessing video streams from multiple platforms. Subsequently, video semantic feature vectors, audio feature vectors, and metadata feature vectors are extracted in the spatiotemporal dimensions of the video, and confidence levels are evaluated for each. Based on this, logical consistency verification is performed on the multimodal features. If the verification is successful, weights are dynamically allocated according to the confidence level to generate a unique semantic skeleton fingerprint. This semantic skeleton fingerprint can effectively resist pixel-level changes caused by secondary creations such as occlusion and editing, achieving a stable representation of the core semantics of the video. Next, accurate video association and source tracing are achieved through similarity calculation between the semantic skeleton fingerprint and the database. Finally, a video trajectory map is constructed by combining video account information and associated propagation paths, and targeted monitoring and control instructions are generated accordingly. This forms a full-link intelligent monitoring system from multi-source perception and semantic fingerprint construction to propagation source tracing, improving the overall accuracy of video monitoring.

[0007] In some embodiments of this application, the step of performing logical consistency verification on the video semantic feature vector, the audio feature vector, and the metadata feature vector to obtain the verification result includes: Extract visual depth features from the video semantic feature vector and acoustic reverberation features from the audio feature vector, and calculate the physical spatial coupling degree between the visual depth features and the acoustic reverberation features; Identify the visual action abrupt change peak of the video semantic feature vector on the time axis and the acoustic energy peak of the audio feature vector on the corresponding time axis, and calculate the temporal alignment error rate between the visual action abrupt change peak and the acoustic energy peak; Extract the environmental entity identifiers contained in the video semantic feature vector and / or the audio feature vector, and extract the device location information and timestamp from the metadata feature vector; calculate the spatiotemporal anchor point deviation value based on whether the location and time attributes represented by the environmental entity identifier match the device location information and timestamp. The physical space coupling degree, the temporal alignment error rate, and the spatiotemporal anchor point deviation value are input into a preset logical evaluation model, and a comprehensive consistency score is output. If the overall consistency score is greater than or equal to the consistency threshold, the verification result is determined to be normal; if it is less than the consistency threshold, the verification result is determined to be abnormal.

[0008] In some embodiments of this application, the step of dynamically assigning weights based on the verification results and the video confidence, audio confidence, and metadata confidence to construct a semantic skeleton fingerprint includes: The video data is divided into continuous time windows. Within each time window, the quantized value of the verification result is used as the basic adjustment coefficient and nonlinearly mapped to the video confidence, the audio confidence, and the metadata confidence, respectively, to generate the cross-modal weight matrix for the corresponding time window. For each time series window, the extreme mode is extracted from the cross-modal weight matrix, and the feature vector corresponding to the extreme mode is used as the main vector of the current window, and the feature vectors corresponding to the other modes are used as branch vectors. Using the main vector as the core node and the branch vector as the satellite node, the topological edge weights between the core node and each satellite node are calculated based on the weight difference in the cross-modal weight matrix. The core node, the satellite node, and the topological edge weights are input into a preset spatiotemporal graph neural network model to perform graph convolution feature aggregation, extracting the principal components with invariant topological structures as the semantic skeleton fingerprint.

[0009] In some embodiments of this application, the step of constructing a video trajectory map based on the video account information corresponding to the video data, the associated video, and the associated account information corresponding to the associated video includes: Construct a multi-dimensional feature tensor space, map the semantic skeleton fingerprint of the video data to the feature tensor space as the first-generation source anchor point, map the associated feature fingerprint of the associated video to the feature tensor space as the second-generation mutation anchor point, concatenate the first-generation source anchor point and the second-generation mutation anchor point according to the order, and initialize to generate a multi-order reference benchmark library. Acquire candidate videos on the distribution chain, synchronously map the candidate feature fingerprints of the candidate videos to the feature tensor space, and calculate the first feature distance between the candidate feature fingerprint and the first-generation source anchor point, and the second feature distance between the candidate feature fingerprint and each of the second-generation mutation anchor points based on the multi-order reference benchmark library. Based on the first feature distance and each of the second feature distances, it is determined whether the candidate video is a derived video. If the candidate video is a derived video, the candidate feature fingerprint is marked as a new mutation anchor point. The new mutation anchor point is inserted into the multi-order reference benchmark library to form a feature decay gradient. Using the video account information and associated account information corresponding to each order anchor point in the multi-order reference benchmark library as graph nodes, and the order concatenation relationship between nodes and the corresponding feature distance as weighted directed edges, a video trajectory graph of video evolution is constructed.

[0010] In some embodiments of this application, determining whether the candidate video is a derived video based on the first feature distance and each of the second feature distances includes: Select the minimum value among the second feature distances as the third feature distance; If the first feature distance is less than a preset source verification threshold, or if the first feature distance is greater than the source verification threshold and the third feature distance is less than a preset mutation verification threshold, the candidate video is determined to be a derived video.

[0011] In some embodiments of this application, generating monitoring and control instructions for the video data based on the video trajectory map includes: A preset multi-scene classification engine is obtained, and the video data in the video trajectory map is input into the multi-scene classification engine, wherein the multi-scene classification engine includes a copyright comparison sub-engine, a compliance semantic sub-engine, and an abnormal behavior topology sub-engine that are executed in parallel. If the copyright comparison sub-engine identifies a match with a degree greater than a preset match threshold, then a copyright infringement feature label is output. If the preset sensitive element library is hit in the compliance semantic sub-engine, the content violation feature tag and its corresponding severity coefficient quantification value will be output. If the abnormal behavior topology sub-engine determines that the state is malicious propagation based on the video trajectory map, then an abnormal account action tag is output. The monitoring and control instructions are generated based on the copyright infringement feature tags, the content violation feature tags, the severity coefficient quantification value, and the account abnormal behavior tags.

[0012] In some embodiments of this application, generating the monitoring and control instructions based on the copyright infringement feature tag, the content violation feature tag, the severity coefficient quantification value, and the account abnormal behavior tag includes: A comprehensive violation entropy is calculated based on the copyright infringement feature tags, the content violation feature tags, the severity coefficient quantification value, and the account abnormal behavior tags. When the overall violation entropy falls within a preset first threshold range, a preliminary warning instruction is generated. When the overall violation entropy falls within a preset second threshold range, a medium-level early warning instruction is generated. When the comprehensive violation entropy falls within a preset third threshold range, a severe warning instruction is generated, wherein the upper limit of the first threshold range is equal to the lower limit of the second threshold range, and the upper limit of the second threshold range is equal to the lower limit of the third threshold range.

[0013] Secondly, embodiments of this application provide a video information monitoring device based on a semantic skeleton, the device comprising: The video acquisition module is used to acquire video streams from various platforms and generate video data based on the video streams; The video processing module is used to extract the video semantic feature vector, audio feature vector, and metadata feature vector of the video data in the spatiotemporal dimension, and to calculate the confidence of the video semantic feature vector, the audio feature vector, and the metadata feature vector respectively to obtain the video confidence, audio confidence, and metadata confidence. The verification processing module is used to perform logical consistency verification on the video semantic feature vector, the audio feature vector, and the metadata feature vector to obtain a verification result. If the verification result is normal, dynamic weight allocation is performed based on the verification result and the video confidence, audio confidence, and metadata confidence to construct a semantic skeleton fingerprint. The association calculation module is used to calculate the similarity between the semantic skeleton fingerprint and a preset video database to obtain a similarity value, and to associate the video based on the similarity value to obtain the associated video; The monitoring and control module is used to construct a video trajectory map based on the video account information corresponding to the video data, the associated video and the associated account information corresponding to the associated video, and to generate monitoring and control instructions for the video data based on the video trajectory map.

[0014] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, a user interface, a communication bus, and a network interface. The processor, the memory, the user interface, and the network interface are respectively connected to the communication bus. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method described in any one of the first aspects.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed, perform the method described in any one of the methods provided in the first aspect above.

[0016] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. This method first establishes a basic data source by accessing video streams from multiple platforms. Then, it extracts video semantic feature vectors, audio feature vectors, and metadata feature vectors in the spatiotemporal dimensions, and evaluates their confidence levels. Based on this, it performs logical consistency verification on the multimodal features. If the verification is successful, weights are dynamically allocated according to the confidence level to generate a unique semantic skeleton fingerprint. This semantic skeleton fingerprint effectively resists pixel-level changes caused by secondary creations such as occlusion and editing, achieving a stable representation of the core semantics of the video. Next, it calculates the similarity between the semantic skeleton fingerprint and the database to achieve accurate video association and source tracing. Finally, it constructs a video trajectory map by combining video account information and associated propagation paths, and generates targeted monitoring and control instructions accordingly. This forms a full-link intelligent monitoring system from multi-source perception and semantic fingerprint construction to propagation source tracing, improving the overall accuracy of video monitoring. Therefore, it effectively solves the problem of low video monitoring accuracy caused by lack of source tracing and pixel recognition in related technologies. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a video information monitoring method based on a semantic skeleton provided in one embodiment of this application; Figure 2 This is a schematic diagram of the feature attenuation gradient of a video information monitoring method based on a semantic skeleton provided in one embodiment of this application; Figure 3 This is a schematic diagram of the structure of a video information monitoring device based on a semantic skeleton provided in one embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0018] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0019] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0020] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0021] Currently, video information monitoring employs automated monitoring solutions based on single-dimensional recognition. These solutions only identify video images (such as faces or sensitive scenes) or audio (such as sensitive words), failing to achieve audio-visual fusion or metadata-linked monitoring. Single-dimensional automated monitoring solutions cannot achieve the fusion and recognition of video images, audio, and metadata. When faced with blurred, occluded, edited, or noise-reduced videos, the accuracy rate is low, and false positives and false negatives are prone to occur.

[0022] Based on this, embodiments of this application provide a video information monitoring method, apparatus, electronic device, and readable storage medium based on a semantic skeleton. This video information monitoring method first constructs a basic data source by accessing video streams from multiple platforms. Then, it extracts video semantic feature vectors, audio feature vectors, and metadata feature vectors in the spatiotemporal dimensions of the video, and performs confidence assessments on each. On this basis, it performs logical consistency verification on the multimodal features. If the verification is normal, it dynamically assigns weights according to the confidence level to generate a unique semantic skeleton fingerprint. This semantic skeleton fingerprint can effectively resist pixel-level changes caused by secondary creations such as occlusion and editing, achieving a stable representation of the core semantics of the video. Next, it achieves accurate video association and source tracing through similarity calculation between the semantic skeleton fingerprint and the database. Finally, it constructs a video trajectory map by combining video account information and associated propagation paths, and generates targeted monitoring and control instructions accordingly, thereby forming a full-link intelligent monitoring system from multi-source perception and semantic fingerprint construction to propagation source tracing, improving the overall accuracy of video monitoring.

[0023] It should be noted that this video information monitoring method based on semantic skeleton is applicable to multiple scenarios such as short video platforms, live streaming platforms, film and television copyright monitoring, and online video compliance review, covering all dimensions of video content, metadata, and dissemination trajectory monitoring.

[0024] The technical solutions provided in the embodiments of this application will be further described below with reference to the accompanying drawings.

[0025] Reference Figure 1 , Figure 1 This is a flowchart illustrating the video information monitoring method based on semantic skeletons provided in this application embodiment. The video information monitoring method based on semantic skeletons is applied to a video information monitoring device based on semantic skeletons. The method is executed by a processor in an electronic device or a readable storage medium. The video information monitoring method based on semantic skeletons includes steps S100, S200, S300, S400, and S500.

[0026] Step S100: Obtain video streams from various platforms and generate video data based on the video streams.

[0027] Video streams are raw, unnormalized audio and video mixed messages transmitted in real time via Internet application layer protocols (such as RTMP, RTSP, or HTTP-based HLS and DASH protocols). Video data is a standardized multidimensional tensor sequence formed by loading various heterogeneous video streams into computer memory after low-level decoding, resampling, and spatial and temporal alignment.

[0028] Specifically, through a deployed distributed data acquisition cluster, the system calls the open APIs of short video platforms to monitor and retrieve the .m3u8 or .flv streaming media uniform resource locators (URLs) of the target videos in real time. Then, it uses dynamic link libraries to separate the audio and video tracks. Spatial normalization is then performed using a bicubic interpolation downsampling algorithm to uniformly scale all video frames to a fixed spatial resolution, eliminating differences in image size between different mobile phones. Temporal normalization is performed by frame extraction or interpolation to strictly lock the video frame rate at 25fps. For the audio track, it is resampled to a 16kHz sampling rate, mono PCM format audio stream. The processed data is stored in a memory buffer as an array to obtain the video data. By acquiring video streams and processing data from various platforms, the system overcomes the underlying data heterogeneity noise caused by cross-platform, cross-device, and cross-container format issues, ensuring that the dimensions of the subsequently received input tensors are completely consistent, avoiding misjudgments due to format differences, and providing support for subsequent data processing.

[0029] Step S200: Extract the video semantic feature vector, audio feature vector, and metadata feature vector of the video data in the spatiotemporal dimension. Calculate the confidence scores of the video semantic feature vector, audio feature vector, and metadata feature vector respectively to obtain the video confidence score, audio confidence score, and metadata confidence score.

[0030] The feature vector is a high-dimensional array that transforms unstructured images and sounds into computer-recognizable data. The confidence score is a number between 0 and 1 used to dynamically measure whether the extracted modality is clear, contains valid information, or is pure noise.

[0031] Specifically, deep neural network models are used to extract video semantic feature vectors, audio feature vectors, and metadata feature vectors in the spatiotemporal dimensions of video data. These deep neural network models include a pre-trained I3D (Inflated 3DConvNet) network, a VGGish network, and an encoder network, which are used for video semantic feature extraction and confidence calculation, audio feature extraction and confidence calculation, and metadata feature extraction and confidence calculation, respectively. For video semantic feature extraction and confidence calculation: 16 consecutive frames are input as a batch into the pre-trained I3D network, and the 1024-dimensional array output from the global pooling layer is extracted as the video semantic feature vector. Simultaneously, the Laplacian variance of the current 16 frames is calculated. If the variance is small, it indicates that the image is extremely blurry or a completely black screen; in this case, the video confidence is assigned a value close to 0. If the image has rich texture, a value close to 1.0 is assigned.

[0032] Audio Feature Extraction and Confidence Calculation: After converting PCM audio into a Log-Mel spectrogram, it is input into the VGGish network to extract a 128-dimensional audio feature vector. The signal-to-noise ratio (SNR) of the audio segment is calculated using the short-time energy method, and the SNR is mapped to the (0,1) interval using the Sigmoid function as the audio confidence score. For example, the confidence score for a completely silent or purely noise-based segment is 0.05. Specifically, converting PCM audio into a Log-Mel spectrogram involves audio loading and normalization. The normalized audio is divided into frames with a 25ms frame length and 10ms frame length. Each frame is multiplied by a Hamming window to reduce spectral leakage. Then, Fourier transform, Mel filter mapping, and logarithmic transform are performed on each frame to obtain the Log-Mel spectrogram. Metadata Feature Extraction and Confidence Calculation: Information is extracted from the video file header, including the release timestamp, device hash, and location coordinates. This information is then One-Hot encoded and concatenated into a 256-dimensional metadata feature vector. The metadata confidence score is calculated based on the completeness of the extracted fields. For example, if only 6 out of 10 fields are extracted, the confidence score is 0.6. By introducing a confidence evaluation mechanism, all extracted features are no longer blindly trusted, resulting in strong anti-interference capabilities. This allows for automatic filtering or reduction of interference from black screens, muted audio, and meaningless footage on subsequent monitoring.

[0033] Step S300: Logical consistency verification is performed on the video semantic feature vector, audio feature vector, and metadata feature vector to obtain the verification result. If the verification result is normal, dynamic weight allocation is performed based on the verification result and the video confidence, audio confidence, and metadata confidence to construct the semantic skeleton fingerprint.

[0034] In one embodiment, logical consistency verification is performed on the video semantic feature vector, audio feature vector, and metadata feature vector to obtain the verification result, including but not limited to the following steps: Step S310: Extract visual depth features from the video semantic feature vector and acoustic reverberation features from the audio feature vector, and calculate the physical spatial coupling degree between the visual depth features and the acoustic reverberation features.

[0035] Visual depth features are a quantitative assessment of the three-dimensional spatial volume and physical depth distance contained in a two-dimensional monocular planar image. Acoustic reverberation features refer to the time and energy distribution of sound waves as they propagate within a specific physical space, experiencing multiple reflections, scattering, and gradual attenuation due to walls, ceilings, and obstacles. Physical spatial coupling is a mathematical scoring mechanism where a value of 0 measures whether the image space and sound space are physically compatible.

[0036] Specifically, video frames are input into a monocular depth estimation model to calculate the normalized mean depth from the foreground to the background. A larger value indicates a more open space, such as an outdoor plaza. The monocular depth estimation model can be either a MiDaS model or a Depth-Anything model. The model outputs a two-dimensional depth map for each frame. Each pixel value in the depth map matrix represents the relative physical distance of that point from the lens. The average depth value is calculated for the central region of the depth map (excluding edge distortion pixels). The Sigmoid function is used to normalize the average depth value to the (0, 1) interval, obtaining the visual depth characteristics. Audio PCM waveform data for the corresponding time period is extracted, and the energy attenuation curve of the signal is calculated using the Schroeder integral method. Based on this, an acoustic algorithm is used to measure the RT60 index. The calculated RT60 seconds are normalized to the (0, 1) interval using a mapping function, yielding the acoustic reverberation characteristics. Visual depth features and acoustic reverberation features are input into a preset coupling calculation formula to obtain the spatial coupling degree. This formula can be calculated using Manhattan distance, expressed as C=exp(-λ|DA|), where C represents the spatial coupling degree, D represents the visual depth feature, A represents the acoustic reverberation feature, and λ represents the scale parameter, which can be set to 2.5 or 3. By calculating the absolute difference between the visual depth feature and the acoustic reverberation feature, a large difference indicates the result is close to 0, suggesting possible dubbing or alteration of ambient sound. This calculation provides a strong physical dimension for identifying AI deepfakes or artificial audio track replacement.

[0037] Step S320: Identify the visual action mutation peak of the video semantic feature vector on the time axis and the acoustic energy peak of the audio feature vector on the corresponding time axis, and calculate the temporal alignment error rate between the visual action mutation peak and the acoustic energy peak.

[0038] The timing alignment error rate refers to the delay difference between the time point when a violent action occurs in the picture and the time point when the sound explodes / high-pitched.

[0039] Specifically, a dense optical flow algorithm is used to calculate the sum of motion vector amplitudes of adjacent video frames, forming a motion intensity curve that changes over time. A peak-finding algorithm is used to find the first timestamp corresponding to the extreme point; for example, a heavy object falls at 3.2 seconds, identifying the visual motion abrupt change peak, which prepares for subsequent calculation of the timing alignment error rate. Short-duration coarse audio energy is extracted, and a peak-finding algorithm is applied to find the steeply rising pulse peak on the audio energy curve, recording the second timestamp corresponding to the acoustic peak; for example, a sound occurs at 3.6 seconds, identifying the acoustic energy peak. The time difference between the first and second timestamps is calculated, and this time difference is compared to a preset time delay threshold to obtain the time ratio. The time ratio is then minimized by a preset error threshold to obtain the timing alignment error rate. The preset time delay threshold can be 500ms, and the preset error threshold can be 1. Poorly copied or stretched videos often result in misalignment of the audio and video tracks. By performing peak comparison calculations, we can identify data that is out of sync with the audio and video, and capture "asynchronous audio and video degraded pseudo-original videos" caused by malicious cutting and splicing or improper screen recording.

[0040] Step S330: Extract environmental entity identifiers contained in the video semantic feature vector and / or audio feature vector; extract device location information and timestamps from the metadata feature vector; calculate the spatiotemporal anchor point deviation value based on whether the location and time attributes represented by the environmental entity identifiers match the device location information and timestamps.

[0041] Specifically, environmental entity identifiers refer to objects in the image with a clear geographical orientation or lighting conditions with a clear temporal orientation. The video image is scanned using a YOLO object detector and an OCR text recognition engine. If landmark features or road signs are detected, they are converted into semantic spatial coordinates using a built-in dictionary. The sunlight angle and color temperature of the video image are detected, or the text in the image is extracted and recorded as semantic temporal attributes. Environmental entity identifiers can represent location and time attributes. Device location information and timestamp information in the metadata feature vector are parsed. The spherical distance formula is used to calculate whether the location information and device location information match. Based on the geographical location information, the latitude and longitude deviation of the geographical location is normalized when the timestamp and time attributes are consistent, outputting a normalized spatiotemporal anchor point deviation value. This spatiotemporal anchor point deviation value ranges from 0 to 1; the larger the deviation, the closer the value is to 1. Absolute time zone ratios are used for virtual positioning recognition and video recognition in different regions.

[0042] Step S340: Input the physical space coupling degree, timing alignment error rate and spatiotemporal anchor point deviation value into the preset logical evaluation model, and output the comprehensive consistency score.

[0043] The logical evaluation model is a machine learning decision-maker trained with supervised learning, such as an XGBoost multi-level regression tree classifier, random forest, or gradient boosting tree. The physical spatial coupling degree obtained in step S310, the temporal alignment error rate obtained in step S320, and the spatiotemporal anchor point deviation value obtained in step S330 are combined to form the input feature space vector matrix. This feature space vector matrix is ​​then input into the pre-trained XGBoost multi-level regression tree classifier. After the leaf nodes of the classifier are calculated, the confidence scores are aggregated using the Sigmoid function, outputting a comprehensive consistency score within the continuous interval [0, 1]. The closer this comprehensive consistency score is to 1, the more reasonable the spatiotemporal logic. By fusing multi-dimensional feature vectors, the spatiotemporal logic of the video data can be judged more comprehensively, resulting in more accurate subsequent recognition results. The pre-trained XGBoost multi-level regression tree classifier is trained on a large amount of video data after the above processing, which increases the model's generalization ability; the specific training process is not detailed here.

[0044] In step S350, if the overall consistency score is greater than or equal to the consistency threshold, the verification result is determined to be normal; if it is less than the consistency threshold, the verification result is determined to be abnormal.

[0045] Specifically, the consistency threshold is the threshold coefficient recorded in the environment variable configuration file, which can be 0.65 or 0.7. If the overall consistency score is greater than or equal to the consistency threshold, it indicates that the video logic is basically compliant, and the verification result is considered normal. If the overall consistency score is less than the consistency threshold, the verification result is considered abnormal, the abnormal video is discarded, not included in fingerprint construction, and the publishing account is added to the high-risk review queue. A quantitative diversion valve is provided to prevent abnormal video data from entering the high-computing-power calculation stage, saving computing power.

[0046] In one embodiment, when the verification result is normal, a semantic skeleton fingerprint is constructed to extract semantically invariant information from the video data, thereby improving the accuracy of video recognition. Based on the verification result and video confidence, audio confidence, and metadata confidence, dynamic weights are assigned to construct the semantic skeleton fingerprint, including but not limited to the following steps: Step S360: Divide the video data into continuous time windows. Within each time window, use the quantized value of the verification result as the basic adjustment coefficient and perform nonlinear mapping with the video confidence, audio confidence and metadata confidence respectively to generate the cross-modal weight matrix of the corresponding time window.

[0047] Specifically, a time-series window is a segmentation of a long video along the time dimension. The cross-modal weight matrix is ​​a scoring table indicating whether the visuals, sound, or additional attributes are more important within the current time window.

[0048] The video data is divided into continuous windows with a 5-second step size. For example, a 15-second video can be divided into 3 windows. Within each time window, the quantified value of the verification result is the comprehensive consistency score. The comprehensive consistency score is used as the base adjustment coefficient. Within this window, the corresponding video confidence score, audio confidence score, and metadata confidence score are obtained. The softmax function is used to perform a non-linear mapping on the base adjustment coefficient, video confidence score, audio confidence score, and metadata confidence score. The specific mapping process is as follows: First, the base adjustment coefficient is multiplied by the video confidence score, audio confidence score, and metadata confidence score respectively, thereby adjusting the value of each confidence score. This is because, when the video data is normal, it often relies heavily on the confidence score. When assigning weights, the weight of clear parts is increased while the weight of blurry parts is decreased, resulting in misjudgment of noisy videos. By multiplying them first, the confidence scores can be balanced, thereby improving the accuracy of semantic extraction. The multiplied confidence values ​​are used to form the weights corresponding to the current time window according to the softmax function, forming the cross-modal weight matrix of the time window, which provides a foundation for generating an accurate semantic skeleton in the future.

[0049] Step S370: For each time series window, extract the extreme mode from the cross-modal weight matrix, take the feature vector corresponding to the extreme mode as the main vector of the current window, and take the feature vectors corresponding to the other modes as branch vectors.

[0050] Specifically, for each time series window, the extreme modality in the cross-modal weight matrix within that window is calculated. The extreme value can be retrieved using ArgMax(W_v, W_a, W_m), where W_v represents the visual modality, W_a represents the audio modality, and W_m represents the metadata modality. For example, if W_v is the maximum value, then the visual modality of the current window is the extreme modality. Based on the determined extreme modality, the 1024-dimensional visual semantic feature vector corresponding to the extreme modality of that time period is extracted and set as the central backbone vector. The 128-dimensional audio features and 256-dimensional metadata features are uniformly downsampled to the same dimensional space through a fully connected layer as branch vectors. Within each time series window, the backbone vector and branch vectors are determined, conforming to the evolutionary pattern of video. As time progresses, the visual modality dominates in the previous time series window, while the audio modality may dominate in the next. This process transforms the disordered multimodal data vectors into a tree-like structure with a clear hierarchical relationship, filtering out noise from low-confidence modalities within the window.

[0051] Step S380: Using the trunk vector as the core node and the branch vector as the satellite node, calculate the topological edge weights between the core node and each satellite node based on the weight difference in the cross-modal weight matrix.

[0052] Specifically, the adjacency matrix of the graph structure is obtained, where the core node connects to all satellite nodes, and satellite nodes are not connected to each other. The core node is represented as i, and the satellite nodes as j. The weight between the core node and each satellite node is represented as the absolute value of wi-wj. Then, the absolute difference of the values ​​in the weight matrix is ​​multiplied by the cosine angle of the vector in the reduced-dimensional space to obtain the edge weight, which is then written into the graph adjacency matrix. By using a mathematical form to characterize how sound, vision, and information are interdependent within a specific time period, isolated features of different dimensions are established in a relational network.

[0053] Step S390: Input the core node, satellite node and topological edge weights into the preset spatiotemporal graph neural network model to perform graph convolution feature aggregation, and extract the principal components with invariant topological structure as semantic skeleton fingerprints.

[0054] Specifically, graph convolutional feature aggregation utilizes neural network computation to selectively fuse information from satellite nodes into core nodes based on connection edges, forming the main information. The pre-defined spatiotemporal graph neural network model is a two-layer spatiotemporal graph convolutional network (ST-GCN). Core nodes, satellite nodes, and topological edge weights are input into the pre-defined spatiotemporal graph neural network model. The first layer extracts the internal distribution correlations of modalities within a single window; the second layer uses gated recurrent units to sequentially pass correlation features along the temporal window. A global average pooling layer compresses the final state into a dense feature tensor of fixed length 256 dimensions. Subsequently, a symbolic hashing algorithm is performed on the tensor (above 0 is assigned a value of 1, less than or equal to 0 is assigned a value of 0), outputting a 256-bit semantic skeleton fingerprint. Through this fingerprint extraction paradigm based on topological relationship aggregation, topological invariance is utilized to achieve a stable representation of the core semantics of the video. Despite the high-intensity modifications to the original video, such as picture-in-picture nesting, brute-force acceleration, mirroring, and background audio mixing, the skeleton fingerprint remains highly consistent and can be accurately identified because the spatiotemporal relational topology within the audio and video remains conserved. The preset spatiotemporal graph neural network model is a pre-trained model; the training process is not detailed here.

[0055] It should be noted that if the consistency verification is abnormal, it indicates that the video data has been tampered with or contains illegal content, and it should be discarded directly to avoid increasing computing power in the future.

[0056] Step S400: Calculate the similarity between the semantic skeleton fingerprint and the preset video database to obtain a similarity value. Based on the similarity value, associate the videos to obtain the associated videos.

[0057] The similarity value measures the distance between two high-dimensional vectors in the feature tensor space; the closer the distance, the higher the similarity. Specifically, a pre-stored video database contains semantic skeleton fingerprints of tens of millions of high-definition videos. Based on the Faiss dense vector retrieval engine, a hierarchical video database is constructed. Skeleton fingerprints are extracted from the video data according to the above format, or video frames are extracted using OpenCV, and features are extracted using neural networks to obtain skeleton fingerprints. An index is created for the skeleton fingerprints; a hierarchical index can be established for retrieval. The generated semantic skeleton fingerprint is input into the Faiss engine, and its cosine similarity to fingerprints in the database is calculated to obtain a similarity value. A similarity threshold is set, which can be set to 0.85. Skeleton fingerprints with similarity values ​​greater than the threshold are returned, resulting in associated videos. Multiple associated videos can be obtained. An approximate nearest neighbor vector retrieval architecture is used to find similar videos for subsequent video source tracing.

[0058] Step S500: Construct a video trajectory map based on the video account information corresponding to the video data, the associated videos, and the associated account information corresponding to the associated videos; and generate monitoring and control instructions for the video data based on the video trajectory map.

[0059] In one embodiment, a video trajectory map is constructed based on the video account information corresponding to the video data, the associated videos, and the associated account information corresponding to the associated videos, including but not limited to the following steps: Step S510: Construct a multi-dimensional feature tensor space, map the semantic skeleton fingerprint of the video data to the feature tensor space as the first-generation source anchor, map the associated feature fingerprint of the associated video to the feature tensor space as the second-generation mutation anchor, and concatenate the first-generation source anchor and the second-generation mutation anchor according to their order to initialize and generate a multi-order reference benchmark library.

[0060] Specifically, the multidimensional feature tensor space is a high-dimensional mathematical vector space representation used to transform various unstructured data of videos (such as image texture, audio spectrum, temporal changes, etc.) into machine-computable coordinate points. In this space, videos with more similar content are located closer together at their coordinate positions. The semantic skeleton fingerprint is a deep feature vector extracted after removing variable information from the video's surface (such as lighting, resolution scaling, and adding or removing black borders), representing the core event content of the video (such as character actions, specific scene sequences). The primary source anchor point represents the coordinates of the first absolutely original video in the representation space, which is the root of the entire evolutionary tree. The secondary mutation anchor point is the coordinate of the derived video generated after the original video has undergone secondary editing, filtering, and compression, which is equivalent to the branches and leaves of the evolutionary tree. The order cascade records the levels of the intergenerational inheritance relationship of videos (e.g., the first original video is level 0, the first repost is level 1, and the repost is edited again to level 2), and so on.

[0061] A multi-dimensional feature tensor space is constructed. When a confirmed original video is identified, its semantic skeleton fingerprint is extracted and its vector is mapped into this tensor space, designated as the initial source anchor point (order 0), and the mapped coordinates are recorded. The original video is directly associated with known related videos, and their associated feature fingerprints are extracted. The fingerprint extraction method can follow the same approach as the skeleton fingerprint extraction described above, and these are mapped into the feature tensor space as the second-generation mutation anchor points (order 1) of the original video data. The order 0 and order 1 anchor points are cascaded at the parent-child level, layer by layer, to form a multi-order reference benchmark library for subsequent comparisons.

[0062] Step S520: Obtain candidate videos on the distribution link, synchronously map the candidate feature fingerprints of the candidate videos to the feature tensor space, and calculate the first feature distance between the candidate feature fingerprint and the first generation source anchor point, and the second feature distance between the candidate feature fingerprint and each second generation mutation anchor point, based on the multi-order reference benchmark library.

[0063] Specifically, on each platform, application layer protocols are used to transmit messages in real time to obtain videos on the distribution link. After decoding, resampling, and spatial and temporal alignment, candidate videos are obtained. Candidate feature fingerprints are extracted from the candidate videos according to steps S200 and S300, and these fingerprints are synchronously mapped to the feature tensor space, converting them to the same tensor space to calculate the distances between various feature fingerprints in the feature tensor space. Based on the multi-level reference benchmark library already constructed in the feature tensor space, the first feature distance between the candidate feature fingerprint and the initial source anchor (the semantic skeleton fingerprint of the video data) is calculated, as well as the second feature distance between the feature fingerprint and each subsequent mutation anchor (the associated feature fingerprint of the associated video). Both the first and second feature distances use Euclidean distance; Hamming distance can also be used for feature distance calculation, which will not be elaborated here. Through these calculations, it is possible to determine whether the candidate video is closer to the original video or whether the associated video is more recent, selecting the most similar version to accurately determine whether the candidate video is an original video.

[0064] Step S530: Based on the first feature distance and each of the second feature distances, determine whether the candidate video is a derived video. If the candidate video is a derived video, mark the candidate feature fingerprint as a new mutation anchor point, and insert the new mutation anchor point into the multi-level reference benchmark library to form a feature decay gradient.

[0065] Specifically, the feature decay gradient means that as the distribution chain extends, the video undergoes multiple distributions, compressions, and cuts, and the differences between its features and the "original source anchor point" gradually increase in a regular manner. This directional feature decay trend along the propagation path is called the feature decay gradient.

[0066] In one embodiment, determining whether a candidate video is a derived video based on a first feature distance and various second feature distances includes, but is not limited to, the following steps: Step S531: Select the minimum value in the second feature distance as the third feature distance.

[0067] Step S532: If the first feature distance is less than the preset source verification threshold, or if the first feature distance is greater than the source verification threshold and the third feature distance is less than the preset mutation verification threshold, the candidate video is determined to be a derived video.

[0068] Specifically, the minimum value among multiple second feature distances is selected as the third feature distance. This represents the video version closest to or most similar to the current candidate video among known mutations, reducing the computational effort required to calculate each second feature distance. If the first feature distance falls below the preset source verification threshold, it indicates that the candidate video is extremely similar to the original source video, and the candidate video is determined to be a derived video. Alternatively, in the process of multiple overlays of videos, the continuously overlaid candidate videos become dissimilar to the original video data. In this case, the first feature distance exceeds the source verification threshold. Since the minimum third feature distance has been extracted, as long as the third feature distance between the candidate video and the associated video is less than the mutation verification threshold, even though the candidate video does not resemble the original video data, it is very similar to the associated data. Based on the "genetic transmission" characteristic, it can still be accurately determined to be a derived video. The source verification threshold can be set to 0.15, which indicates a derived video; the mutation verification threshold can be set to 0.08. Both are calculated using Euclidean distance, with smaller values ​​indicating closer similarity between the videos.

[0069] Traditional video fingerprinting techniques often only compare against the original video. Once the video has undergone more than three rounds of heavy watermarking, filtering, and frame extraction, the feature distance far exceeds the source verification threshold, leading to a "broken link in the tracing chain." By introducing a third feature distance and a mutation verification threshold, candidate videos are allowed to be compared with the next generation of variants. This constructs a continuous genetic mutation gradient, forming a video trajectory chain and a semantic skeleton fingerprint, enabling accurate identification regardless of how the original video data is forged.

[0070] like Figure 2As shown, when a candidate video is a derived video, it is not discarded but used as a new mutation anchor point. The anchor point is set to the order of the associated video plus one, and then its mapping is used to interpolate the multi-order reference base library to form a feature decay gradient. This ensures that subsequent candidate videos, even after multiple evolutions, can still be identified. With continuous operation, countless mutation branches will grow around the initial source anchor point. Because each time a small change is introduced into the candidate video, a "feature decay gradient" is formed, progressing from clear to blurry, from complete to incomplete. When N+2 hands of video appear in the future, there are closer anchor points for identification, thereby improving the accuracy of video recognition and monitoring.

[0071] In another embodiment, if the candidate video is not a derived video, the candidate video can be published, and control instructions for publishing the candidate video can be generated.

[0072] Step S540: Using the video account information and associated account information corresponding to each order anchor point in the multi-order reference benchmark library as graph nodes, and using the order concatenation relationship between nodes and the corresponding feature distance as weighted directed edges, a video trajectory graph of video evolution is constructed.

[0073] Specifically, after constructing the feature decay gradient, video tracking is achieved, moving beyond traditional simple video tagging. Based on the video account information corresponding to the video data and the associated account information corresponding to related videos, a video propagation trajectory graph is constructed according to the feature decay gradient. The video account information and associated account information are used as graph nodes. The end of the feature decay gradient is used as the order relationship between nodes, and the feature distance is used as the weight of the directed edges, forming a weighted propagation trajectory graph. For example, the video account information (e.g., who posted it) and its associated information (e.g., which channel it was downloaded from) corresponding to each order anchor point are used as entity nodes in the graph. The order concatenation relationship between anchor points (0th order points to 1st order, 1st order points to 2nd order) determines the time sequence and the order of infringement. The feature distance between them (the larger the distance, the more egregious the modification) is used as a weighted directed edge, thus drawing a clear video trajectory graph that tracks the complete propagation evolution and network topology, enabling subsequent control based on the video trajectory graph.

[0074] In one embodiment, generating monitoring and control instructions for video data based on video trajectory maps includes, but is not limited to, the following steps: Step S550: Obtain a preset multi-scene classification engine and input the video data in the video trajectory map into the multi-scene classification engine. The multi-scene classification engine includes a copyright comparison sub-engine, a compliance semantic sub-engine, and an abnormal behavior topology sub-engine that are executed in parallel.

[0075] Step S560: If the matching degree is found to be greater than the preset matching threshold in the copyright comparison sub-engine, then output the copyright infringement feature label.

[0076] Step S570: If the preset sensitive element library is hit in the compliance semantic sub-engine, the content violation feature tag and its corresponding severity coefficient quantification value are output.

[0077] Step S580: If the abnormal behavior topology sub-engine determines that the state is malicious propagation based on the video trajectory map, then output the account abnormal action tag.

[0078] Among them, the multi-scenario classification engine is the judgment center with a distributed parallel computing architecture, which can analyze from different dimensions. The multi-scenario classification engine includes a copyright comparison sub-engine, a compliance semantic sub-engine, and an abnormal behavior topology sub-engine that are executed in parallel. Different engines handle comparisons under different scenarios.

[0079] Specifically, the copyright comparison sub-engine extracts keyframes and audio spectrum from the video and searches through libraries of copyrighted films and television dramas or original content protection libraries. If the feature repetition rate between the candidate video and the protected library is found to be greater than a preset 95%, the sub-engine outputs a copyright infringement feature tag. For example, this copyright infringement feature tag could be "picture-in-picture copying" or "segmented explanation".

[0080] The compliance semantic sub-engine operates by using natural language processing and convolutional neural networks to perform semantic understanding and analysis on video images and words. If a preset sensitive element library is hit, the engine not only outputs a "content violation feature tag," but also outputs a severity coefficient quantification value based on the sensitive occupancy ratio and dwell time of the element. For example, the severity of a 0.5-second flash ad is 0.2, and the verification level of sensitive content promotion throughout the entire video is 0.9.

[0081] The abnormal behavior topology sub-engine analyzes the spatiotemporal relationships of nodes in the video trajectory graph. For example, if the graph shows: "Within a brief 10 minutes from 2:00 AM to 2:10 AM, 150 previously unconnected accounts with 0 followers (nodes) simultaneously distributed various derivative variants of the video to the entire network with the same order of mutation distance (edge ​​weight)." This sharing would be immediately identified as malicious propagation by a bot, resulting in an abnormal account action tag. The parallel execution of the copyright comparison sub-engine, compliance semantic sub-engine, and abnormal behavior topology sub-engine analyzes the video from different dimensions, improving the accuracy of end-to-end video monitoring. Through the parallel execution architecture of the three sub-engines in the multi-scene classification engine, copyright, compliance, and graph behavior analysis can be completed simultaneously using the same computing power in the initial stage of video uploading or distribution. This reduces the decision delay from tens of minutes to seconds when dealing with tens of millions of newly added short videos daily.

[0082] Step S590: Generate monitoring and control instructions based on copyright infringement feature tags, content violation feature tags, severity coefficient quantification values, and abnormal account action tags.

[0083] In one embodiment, monitoring and control instructions are generated based on copyright infringement feature tags, content violation feature tags, severity coefficient quantification values, and abnormal account action tags, including but not limited to the following steps: Step S591: Calculate the comprehensive violation entropy based on copyright infringement feature tags, content violation feature tags, severity coefficient quantification value, and abnormal account action tags.

[0084] The comprehensive violation entropy is a dimensionality-reduced calculation that integrates video infringement certainty, content harmfulness, and the severity of online troll dissemination. The more violation factors and the higher their weights, the higher the overall risk control entropy value. Specifically, copyright infringement feature tags, content violation feature tags, and account abnormal behavior tags are quantified according to preset rules to form copyright infringement scores, content violation scores, and account abnormality scores. The comprehensive violation entropy is obtained by multiplying the copyright infringement score by the first weight, adding the product of the content violation score, the second weight (severity / degree coefficient), and then adding the account abnormality score and the third weight. The first, second, and third weights are weight adjustment parameters, and the preset rules are the scores corresponding to the tags. These parameters are set by professionals based on statistical analysis of a large amount of historical data. The overall risk level is reflected through violation entropy calculation.

[0085] Step S592: If the comprehensive violation entropy is within the preset first threshold range, generate a primary warning instruction.

[0086] Step S593: If the comprehensive violation entropy is within the preset second threshold range, generate a medium-level warning instruction.

[0087] Step S594: When the comprehensive violation entropy is within a preset third threshold interval, a severe warning instruction is generated, wherein the upper limit of the first threshold interval is equal to the lower limit of the second threshold interval, and the upper limit of the second threshold interval is equal to the lower limit of the third threshold interval.

[0088] The upper limit of the first threshold interval is equal to the lower limit of the second threshold interval, and the upper limit of the second threshold interval is equal to the lower limit of the third threshold interval. For example, the first threshold interval is set to [0,30), the second threshold interval is set to [30,70), and the third threshold interval is set to [70,100].

[0089] Specifically, if the overall violation entropy is 25, falling into the first threshold range, a primary warning instruction is directly generated. The actions include: reducing the traffic priority of candidate videos (not recommending them to new users), silently removing relevant tags, and automatically placing them in the low-priority queue for manual review.

[0090] If the infringement is serious enough, and the overall violation entropy is 55, falling into the second threshold range, a medium-level warning instruction will be generated. The actions taken include: immediately removing the infringing video from the platform and disconnecting its link, sending an in-site warning letter to the uploader's account, deducting the account's credit score, and freezing any revenue generated from the video.

[0091] If a serious violation occurs, and a network of online trolls is detected in the video trajectory graph, resulting in a comprehensive violation entropy of 92, falling into the third threshold range, a severe warning is triggered. The actions include not only removing the video but also batch-blocking all associated physical accounts on that topology tree using the video trajectory graph and extracting their device fingerprint blacklist. By calculating continuous comprehensive violation entropy and ensuring the seamless equality of the upper and lower limits of the first, second, and third intervals, combined with a severity coefficient quantification value, automated, multi-level, flexible tiered management of demotion, rate limiting, account blocking, and source tracing alarms is achieved.

[0092] like Figure 3 As shown, this application embodiment provides a video information monitoring device 100 based on a semantic skeleton. This device 100 acquires video streams from various platforms through a video acquisition module 110 and generates video data based on the video streams. A video processing module 120 extracts video semantic feature vectors, audio feature vectors, and metadata feature vectors in the spatiotemporal dimension of the video data. Confidence scores are calculated for each of the video semantic feature vectors, audio feature vectors, and metadata feature vectors to obtain video confidence scores, audio confidence scores, and metadata confidence scores. A verification processing module 130 verifies the video semantic feature vectors, audio feature vectors, and metadata feature vectors. The data feature vectors are logically consistent to obtain the verification result. If the verification result is normal, dynamic weight allocation is performed based on the verification result and video confidence, audio confidence, and metadata confidence to construct a semantic skeleton fingerprint. The semantic skeleton fingerprint is then compared with the preset video database using the association calculation module 140 to obtain a similarity value. Based on the similarity value, videos are associated to obtain associated videos. The monitoring and control module 150 constructs a video trajectory map based on the video account information corresponding to the video data, the associated videos, and the associated account information corresponding to the associated videos. Based on the video trajectory map, monitoring and control instructions for the video data are generated.

[0093] It should be noted that the video acquisition module 110 is connected to the video processing module 120, the video processing module 120 is connected to the verification processing module 130, the verification processing module 130 is connected to the association calculation module 140, and the association calculation module 140 is connected to the monitoring and control module 150. The aforementioned video information monitoring method based on semantic skeletons is applied to a video information monitoring device 100 based on semantic skeletons. The device 100 constructs a basic data source by accessing video streams from multiple platforms. Subsequently, it extracts the video semantic feature vector, audio feature vector, and metadata feature vector in the spatiotemporal dimension of the video, and performs confidence assessments on each. Based on this, it performs logical consistency verification on the multimodal features. If the verification is normal, it dynamically allocates weights according to the confidence level to generate a unique semantic skeleton fingerprint. This semantic skeleton fingerprint can effectively resist pixel-level changes caused by secondary creations such as occlusion and editing, achieving a stable representation of the core semantics of the video. Then, it achieves accurate video association and source tracing through similarity calculation between the semantic skeleton fingerprint and the database. Finally, by combining video account information with associated propagation paths, a video trajectory map is constructed, and targeted monitoring and control instructions are generated accordingly. This forms a full-link intelligent monitoring system from multi-source perception and semantic fingerprint construction to propagation tracing, thereby improving the overall accuracy of video monitoring.

[0094] It should also be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0095] This application also discloses an electronic device. (See reference...) Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. The electronic device 500 may include: at least one processor 501, at least one network interface 504, a user interface 503, a memory 505, and at least one communication bus 502.

[0096] The communication bus 502 is used to enable communication between these components.

[0097] The user interface 503 may include a display screen and a camera. Optionally, the user interface 503 may also include a standard wired interface and a wireless interface.

[0098] The network interface 504 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0099] The processor 501 may include one or more processing cores. The processor 501 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 505, and by calling data stored in memory 505. Optionally, the processor 501 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array. The processor 501 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and Modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor 501.

[0100] The memory 505 may include random access memory (RAM) or read-only memory. Optionally, the memory 505 may include a non-transitory computer-readable storage medium. The memory 505 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 505 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 505 may also be at least one storage device located remotely from the aforementioned processor 501. (Refer to...) Figure 4 The memory 505, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a video information monitoring method based on a semantic skeleton.

[0101] exist Figure 4In the illustrated electronic device 500, the user interface 503 is mainly used to provide an input interface for the user and acquire user input data; while the processor 501 can be used to call an application program stored in the memory 505 for a video information monitoring method based on a semantic skeleton. When executed by one or more processors 501, the electronic device 500 performs one or more methods as described in the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0102] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0103] In the various embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.

[0104] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0105] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0106] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0107] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will readily conceive of those skilled in the art upon consideration of the specification and the disclosure of practical truths.

[0108] This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A method for video information monitoring based on semantic skeleton, characterized in that, The method includes: Acquire video streams from various platforms and generate video data based on the video streams; Extract the video semantic feature vector, audio feature vector, and metadata feature vector in the spatiotemporal dimension of the video data, and calculate the confidence scores of the video semantic feature vector, the audio feature vector, and the metadata feature vector respectively to obtain the video confidence score, audio confidence score, and metadata confidence score. Logical consistency verification is performed on the video semantic feature vector, the audio feature vector, and the metadata feature vector to obtain the verification result. If the verification result is normal, dynamic weight allocation is performed based on the verification result and the video confidence, audio confidence, and metadata confidence to construct a semantic skeleton fingerprint. The semantic skeleton fingerprint is compared with a preset video database to obtain a similarity value. Based on the similarity value, the videos are associated to obtain the associated videos. Based on the video account information corresponding to the video data, the associated video and the associated account information corresponding to the associated video, a video trajectory map is constructed, and monitoring and control instructions for the video data are generated based on the video trajectory map.

2. The method of claim 1, wherein, The logical consistency verification of the video semantic feature vector, the audio feature vector, and the metadata feature vector, to obtain the verification result, includes: Extract visual depth features from the video semantic feature vector and acoustic reverberation features from the audio feature vector, and calculate the physical spatial coupling degree between the visual depth features and the acoustic reverberation features; Identify the visual action abrupt change peak of the video semantic feature vector on the time axis and the acoustic energy peak of the audio feature vector on the corresponding time axis, and calculate the temporal alignment error rate between the visual action abrupt change peak and the acoustic energy peak; Extract the environmental entity identifiers contained in the video semantic feature vector and / or the audio feature vector, and extract the device location information and timestamp from the metadata feature vector; calculate the spatiotemporal anchor point deviation value based on whether the location and time attributes represented by the environmental entity identifier match the device location information and timestamp. The physical space coupling degree, the temporal alignment error rate, and the spatiotemporal anchor point deviation value are input into a preset logical evaluation model, and a comprehensive consistency score is output. If the overall consistency score is greater than or equal to the consistency threshold, the verification result is determined to be normal; if it is less than the consistency threshold, the verification result is determined to be abnormal.

3. The method of claim 2, wherein, The step of dynamically assigning weights based on the verification results and the video confidence, audio confidence, and metadata confidence to construct a semantic skeleton fingerprint includes: The video data is divided into continuous time windows. Within each time window, the quantized value of the verification result is used as the basic adjustment coefficient and nonlinearly mapped to the video confidence, the audio confidence, and the metadata confidence, respectively, to generate the cross-modal weight matrix for the corresponding time window. For each time series window, the extreme mode is extracted from the cross-modal weight matrix, and the feature vector corresponding to the extreme mode is used as the main vector of the current window, and the feature vectors corresponding to the other modes are used as branch vectors. Using the main vector as the core node and the branch vector as the satellite node, the topological edge weights between the core node and each satellite node are calculated based on the weight difference in the cross-modal weight matrix. The core node, the satellite node, and the topological edge weights are input into a preset spatiotemporal graph neural network model to perform graph convolution feature aggregation, extracting the principal components with invariant topological structures as the semantic skeleton fingerprint.

4. The method of claim 1, wherein, The construction of a video trajectory map based on the video account information corresponding to the video data, the associated video, and the associated account information corresponding to the associated video includes: Construct a multi-dimensional feature tensor space, map the semantic skeleton fingerprint of the video data to the feature tensor space as the first-generation source anchor point, map the associated feature fingerprint of the associated video to the feature tensor space as the second-generation mutation anchor point, concatenate the first-generation source anchor point and the second-generation mutation anchor point according to the order, and initialize to generate a multi-order reference benchmark library. Acquire candidate videos on the distribution chain, synchronously map the candidate feature fingerprints of the candidate videos to the feature tensor space, and calculate the first feature distance between the candidate feature fingerprint and the first-generation source anchor point, and the second feature distance between the candidate feature fingerprint and each of the second-generation mutation anchor points based on the multi-order reference benchmark library. Based on the first feature distance and each of the second feature distances, it is determined whether the candidate video is a derived video. If the candidate video is a derived video, the candidate feature fingerprint is marked as a new mutation anchor point. The new mutation anchor point is inserted into the multi-order reference benchmark library to form a feature decay gradient. Using the video account information and associated account information corresponding to each order anchor point in the multi-order reference benchmark library as graph nodes, and the order concatenation relationship between nodes and the corresponding feature distance as weighted directed edges, a video trajectory graph of video evolution is constructed.

5. The method of claim 4, wherein, The step of determining whether the candidate video is a derived video based on the first feature distance and each of the second feature distances includes: Select the minimum value among the second feature distances as the third feature distance; If the first feature distance is less than a preset source verification threshold, or if the first feature distance is greater than the source verification threshold and the third feature distance is less than a preset mutation verification threshold, the candidate video is determined to be a derived video.

6. The method according to claim 4, characterized in that, The step of generating monitoring and control instructions for the video data based on the video trajectory map includes: A preset multi-scene classification engine is obtained, and the video data in the video trajectory map is input into the multi-scene classification engine, wherein the multi-scene classification engine includes a copyright comparison sub-engine, a compliance semantic sub-engine, and an abnormal behavior topology sub-engine that are executed in parallel. If the copyright comparison sub-engine identifies a match with a degree greater than a preset match threshold, then a copyright infringement feature label is output. If the preset sensitive element library is hit in the compliance semantic sub-engine, the content violation feature tag and its corresponding severity coefficient quantification value will be output. If the abnormal behavior topology sub-engine determines that the state is malicious propagation based on the video trajectory map, then an abnormal account action tag is output. The monitoring and control instructions are generated based on the copyright infringement feature tags, the content violation feature tags, the severity coefficient quantification value, and the account abnormal behavior tags.

7. The method according to claim 6, characterized in that, The process of generating the monitoring and control instructions based on the copyright infringement feature tags, the content violation feature tags, the severity coefficient quantification value, and the account abnormal behavior tags includes: A comprehensive violation entropy is calculated based on the copyright infringement feature tags, the content violation feature tags, the severity coefficient quantification value, and the account abnormal behavior tags. When the overall violation entropy falls within a preset first threshold range, a preliminary warning instruction is generated. When the overall violation entropy falls within a preset second threshold range, a medium-level early warning instruction is generated. When the comprehensive violation entropy falls within a preset third threshold range, a severe warning instruction is generated, wherein the upper limit of the first threshold range is equal to the lower limit of the second threshold range, and the upper limit of the second threshold range is equal to the lower limit of the third threshold range.

8. A video information monitoring device based on semantic skeleton, characterized in that, The device includes: The video acquisition module is used to acquire video streams from various platforms and generate video data based on the video streams; The video processing module is used to extract the video semantic feature vector, audio feature vector, and metadata feature vector of the video data in the spatiotemporal dimension, and to calculate the confidence of the video semantic feature vector, the audio feature vector, and the metadata feature vector respectively to obtain the video confidence, audio confidence, and metadata confidence. The verification processing module is used to perform logical consistency verification on the video semantic feature vector, the audio feature vector, and the metadata feature vector to obtain a verification result. If the verification result is normal, dynamic weight allocation is performed based on the verification result and the video confidence, audio confidence, and metadata confidence to construct a semantic skeleton fingerprint. The association calculation module is used to calculate the similarity between the semantic skeleton fingerprint and a preset video database to obtain a similarity value, and to associate the video based on the similarity value to obtain the associated video; The monitoring and control module is used to construct a video trajectory map based on the video account information corresponding to the video data, the associated video and the associated account information corresponding to the associated video, and to generate monitoring and control instructions for the video data based on the video trajectory map.

9. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, a communication bus, and a network interface. The processor, the memory, the user interface, and the network interface are respectively connected to the communication bus. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.