New media data transmission management method and system
By extracting video frames and metadata packets from new media data streams and combining them with synchronization anomaly detection and cross-layer collaborative compensation mechanisms, the problem of insufficient synchronization processing in new media data transmission is solved, high-precision picture synchronization and transmission strategy optimization are achieved, and the transmission stability and intelligence level are improved.
Patent Information
- Application Number
- CN202511039625.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-03
AI Technical Summary
In the transmission of new media data, existing technologies lack the ability to synchronize video images and metadata, resulting in image delays, drift, or abnormal operation responses. Traditional methods are unable to repair synchronization errors in real time, affecting transmission stability and intelligence.
By acquiring real-time media data streams, extracting video frame sequences and their associated metadata packets, performing motion trajectory feature extraction and spatiotemporal parameter analysis, and using a synchronous anomaly detection model to calculate spatial offset and temporal deviation values, dynamic metadata misalignment judgment is achieved, and a cross-layer collaborative compensation mechanism is triggered to generate new media picture compensation data, dynamically adjusting parameter weights to optimize transmission management.
It significantly improves the synchronization accuracy and robustness of multi-source real-time media, enhances the system's self-recovery and adaptability, and is suitable for highly interactive scenarios such as monitoring, AR/VR, reducing latency and redundant data, and improving transmission efficiency and controllability.
Smart Images

Figure CN120751190A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data transmission, and in particular to a new media data transmission management method and system. Background Art
[0002] Early data transmission relied on the traditional TCP / IP protocol for file-level transmission, resulting in low transmission efficiency and real-time performance, making it difficult to meet the high-frequency, high-volume, and low-latency demands of multimedia data (such as video, audio, and images). With the advent of the mobile internet era and the development of 4G and 5G technologies, data transmission is moving towards high-bandwidth, low-latency options. Content delivery networks (CDNs), peer-to-peer (P2P) distribution, and streaming protocols (such as RTMP and HLS) are widely adopted in new media platforms, enabling distributed, efficient, and stable data transmission. In recent years, with the convergence of artificial intelligence, big data, and edge computing, the transmission and management of new media data has evolved from passive transmission to intelligent scheduling, real-time perception, and dynamic optimization. Simultaneously, technologies such as data security, encrypted transmission, and privacy protection have become research priorities to address the risks of attacks and information leakage during data transmission. However, in the current traditional real-time media transmission process, the synchronization processing capabilities of video images and metadata (such as user interaction, spatial location, etc.) are insufficient, which can easily lead to image delays, drift or abnormal operation responses. At the same time, when synchronization errors occur, traditional methods mostly rely on post-manual adjustments or delay processing, and cannot be repaired immediately in the transmission link, which leads to low stability and intelligence levels in new media data transmission management. Summary of the Invention
[0003] Based on this, it is necessary to provide a new media data transmission management method and system to solve at least one of the above technical problems.
[0004] To achieve the above object, a new media data transmission management method is provided, the method comprising the following steps:
[0005] Step S1: obtaining a real-time media data stream, which includes a video frame sequence and its associated metadata packet set, wherein the metadata packet set includes spatial coordinate data, timestamp data, and an operation type identifier;
[0006] Step S2: extracting motion trajectory features from the video frame sequence to generate a video trajectory feature vector, and performing spatiotemporal parameter analysis on the metadata packet set to generate a metadata spatiotemporal parameter set;
[0007] Step S3: Based on the video trajectory feature vector and the metadata spatiotemporal parameter set, the spatial offset and temporal deviation values are input into a preset synchronization anomaly detection model; the spatial offset and temporal deviation values are used to perform media image misalignment judgment on the real-time media data stream to obtain a dynamic metadata misalignment judgment result;
[0008] Step S4: If the dynamic metadata misalignment judgment result is true, cross-layer collaborative compensation is performed on the real-time media data stream to generate new media picture compensation data; the picture space alignment accuracy of the compensated real-time media data stream is collected through the new media picture compensation data, and the parameter weights of the new media picture compensation data are dynamically adjusted to perform new media data transmission management optimization operations.
[0009] This invention combines video trajectory feature vectors with spatiotemporal parameters of metadata and utilizes a synchronization anomaly detection model to effectively detect spatial and temporal offsets, achieving high-precision identification of media synchronization misalignment and significantly improving the synchronization accuracy of multi-source real-time media. Upon detecting metadata misalignment, a cross-layer collaborative compensation mechanism is automatically triggered. This mechanism not only generates compensation data but also dynamically collects spatial alignment accuracy and adaptively adjusts parameter weights to achieve real-time media image repair and transmission strategy optimization, enhancing the system's robustness and self-recovery capabilities. It supports the joint parsing and processing of video frames and their metadata (such as coordinates, timestamps, and operation types) in complex dynamic scenes, making it suitable for a variety of highly interactive scenarios such as surveillance, AR / VR, and remote collaboration, improving the system's adaptability and versatility. By dynamically adjusting the parameter weights of the compensation data, bandwidth control and data compression strategies are optimized, helping to reduce latency, reduce redundant data, and improve the overall transmission efficiency and controllability of media data streams. The logical logic of each step is clear and can be flexibly embedded in existing media processing or data synchronization systems. Both the synchronization anomaly detection model and compensation mechanism support model upgrades and policy customization, facilitating future iterations and intelligent expansion of the system. Therefore, the present invention improves the stability and intelligence level of new media data transmission management by dynamically distinguishing picture delays and introducing a cross-layer collaborative compensation mechanism.
[0010] Preferably, step S1 includes the following steps:
[0011] Step S11: continuously capturing data of the video frame input signal to obtain original video frame sequence data;
[0012] Step S12: performing frame-level data decoding on the original video frame sequence data to generate structured video frame data and extracting a frame sequence of structured time-frequency frame data to obtain a video frame sequence;
[0013] Step S13: using the video frame sequence to identify control signaling tags on the video frame input signal, generating a preliminary metadata segment set; performing field extraction on the metadata segment set, extracting spatial coordinate data, timestamp data, and operation type identifier, and generating an associated metadata package set;
[0014] Step S14: synchronously mapping the video frame sequence and the associated metadata packet set to obtain a real-time media data stream.
[0015] The present invention ensures the integrity of video data and the accuracy of time-frequency structure by continuously capturing and decoding the video frame input signal at the frame level, laying a structured foundation for subsequent processing. The control signaling tag recognition mechanism is introduced to effectively improve the recognition ability of operation behavior. Combined with the field extraction operation, a set of metadata packets containing spatial coordinates, timestamps and operation types is accurately generated, which improves the semantic integrity and practical usability of metadata. Through the synchronous mapping strategy, the structured video frame sequence and its corresponding metadata are coupled in the time and space dimensions to generate a real-time media data stream with strong consistency and time-space alignment, providing accurate input for subsequent analysis models. Accurate structured data stream input provides a more reliable basic input for the synchronization anomaly detection model, improves the model's perception of micro-time-space offsets, and enhances the system's fault tolerance and adaptability to synchronization anomalies. Based on modular data capture, structured decoding, signaling parsing and synchronization mapping mechanisms, the system has good scalability and is suitable for real-time media acquisition needs under multiple types of media sources, heterogeneous video formats or different transmission protocols.
[0016] Preferably, in step S13, using the video frame sequence to perform control signaling tag recognition on the video frame input signal includes:
[0017] Extracting frame header data of a video frame input signal using a video frame sequence to generate video frame synchronization index data;
[0018] Divide the video frame synchronization index data into content segments to generate signaling candidate area data;
[0019] Scan key bits of the signaling candidate area data to generate control signaling feature segment data;
[0020] Performing operation identification matching on the control signaling feature segment data to generate initial operation type identification data;
[0021] Perform frame sequence timing mapping processing on the initial operation type identification data to generate a preliminary metadata segment set.
[0022] This invention generates video frame synchronization index data by extracting frame header data. This allows the system to accurately locate the physical position and logical structure of each frame in the video stream, enhancing the deconstruction capabilities and synchronization accuracy of the video frame input signal. A content segmentation mechanism is introduced to effectively narrow the target area for signaling recognition and improve the efficiency of key bit scanning, making it particularly effective when processing complex video structures or high frame rates. Key bits within the candidate area are deeply scanned to accurately extract control signaling feature segments, ensuring that the system can quickly identify and classify implicit control operation signals (such as gesture control and interaction triggering). An operation identifier matching mechanism rapidly generates initial operation type identification data, providing structured semantic support for subsequent functions such as behavior prediction, operation response, or event annotation. Through frame sequence temporal mapping, the identified operation type data is sequentially mapped to the frame sequence, effectively generating a structurally continuous and logically complete set of preliminary metadata segments, ensuring high consistency between the metadata timeline and the frame. This method is highly adaptable and generalizable, adapting to video sources with various codec formats and signaling protocols, providing strong data support for widespread applications in multimedia interaction, intelligent monitoring, and automatic annotation systems.
[0023] Preferably, step S2 includes the following steps:
[0024] Step S21: analyzing the key frames of the video frame sequence and performing key frame screening to generate media video motion key frame data; continuously tracking the target position of the video motion key frame data to obtain trajectory point sequence data;
[0025] Step S22: performing trajectory coherence fitting on the trajectory point sequence data to generate a video trajectory feature vector;
[0026] Step S23: performing time stamp calculation on the metadata package set to generate synchronization timing parameter data; performing spatial position calculation on the spatial coordinates in the metadata package set to generate position mapping parameter data;
[0027] Step S24: jointly associate the synchronization timing parameter data and the position mapping parameter data to generate a metadata spatiotemporal parameter set.
[0028] Through keyframe screening and target position tracking mechanisms, the present invention effectively identifies keyframe regions with significant motion information in the video and extracts trajectory point sequence data, thereby providing high-quality input for behavioral trajectory modeling and improving the system's ability to accurately perceive entity behavior in dynamic images. Trajectory coherence fitting (S22) can perform temporal consistency analysis and structured encoding on complex trajectory data. The generated video trajectory feature vector has good spatiotemporal continuity and discriminability, providing robust feature expression for subsequent anomaly detection models and image alignment models. Timestamp solution and spatial position solution (S23) respectively perform deep modeling of the time dimension and space dimension to generate high-precision synchronized time series parameter data and position mapping parameter data, effectively avoiding spatiotemporal drift problems caused by time offset or coordinate errors. By jointly fusing the time and space parameters, a metadata spatiotemporal parameter set with a clear structure and high matching accuracy is generated, laying a solid data foundation for the system's subsequent alignment accuracy judgment and synchronization compensation operations. For multi-target, multi-perspective, or multi-timeline problems in the scene, this step has powerful trajectory fusion and metadata synchronization capabilities, wide adaptability, and high computational stability, significantly enhancing the system's generalization and stable operation capabilities in dynamic environments.
[0029] Preferably, step S3 includes the following steps:
[0030] Step S31: performing time series synchronization on the video trajectory feature vector and the metadata spatiotemporal parameter set to generate synchronized input feature data;
[0031] Step S32: inputting the synchronous input feature data into a preset synchronous anomaly detection model to perform spatial coordinate error calculation and timestamp offset analysis to generate spatial offset data and time deviation value data;
[0032] Step S33: performing position error threshold comparison on the spatial offset data to generate image position misalignment identification data; performing delay threshold comparison on the time deviation value data to generate time synchronization abnormality identification data;
[0033] Step S34: performing media picture misalignment determination based on the picture position misalignment identification data and the time synchronization anomaly identification data, and generating a dynamic metadata misalignment determination result.
[0034] By synchronizing the video trajectory feature vectors with the metadata spatiotemporal parameter sets in a time series, the present invention effectively generates synchronized input feature data in a unified format, ensuring strict temporal alignment of video behavior and associated metadata, significantly improving the accuracy and interpretability of subsequent error detection. Through a synchronized anomaly detection model (step S32), the system simultaneously calculates spatial coordinate errors and timestamp offsets, generating spatial offset data and temporal deviation data. This model supports multi-dimensional error modeling and exhibits excellent versatility and algorithmic portability. By setting threshold comparison mechanisms for spatial and temporal offsets, it rapidly outputs image position misalignment indicators and temporal synchronization anomaly indicators. This approach offers advantages such as low latency and high accuracy, making it suitable for rapid quality analysis of large-scale, real-time video data. By integrating and judging the two types of anomaly indicator data in step S34, the system outputs dynamic metadata misalignment judgment results in real time. These results are structured and time-series-based, helping to drive subsequent compensation, correction, and optimization actions, providing a precise basis for real-time media quality control. This step process can adapt to complex scenarios such as multiple targets, multiple trajectories, and multi-level metadata, and can still maintain a high level of error discrimination stability under dynamic changing conditions such as nonlinear motion, sudden occlusion, or perspective conversion.
[0035] Preferably, step S34 includes the following steps:
[0036] Step S341: performing spatial anomaly clustering on the image position misalignment identification data to generate spatial misalignment distribution data;
[0037] Step S342: performing time series focusing on the time synchronization anomaly identification data to generate time misalignment aggregation data;
[0038] Step S343: performing segmented co-frequency vibration superposition on the spatial misalignment distribution data and the temporal misalignment aggregation data to generate spatiotemporal coupling resonance feature data; performing dynamic recognition processing on the spatiotemporal coupling resonance feature data to generate media image misalignment status data;
[0039] Step S344: Performing logic rule judgment on the media image misalignment status data to generate a dynamic metadata misalignment judgment result.
[0040] The present invention performs spatial anomaly clustering and temporal focusing aggregation operations on the image position misalignment identification data and the time synchronization anomaly identification data respectively through steps S341 and S342, which can construct a spatial distribution model and a temporal evolution model of the misalignment behavior, realize the accurate structured expression of multi-source misalignment information, and provide high-quality input for downstream feature fusion. By superimposing the spatial misalignment and the temporal misalignment in segmented and co-frequency vibrations, spatiotemporal resonance feature data with coupling characteristics is generated, which can effectively identify the potential image anomaly emergence mode caused by synchronization mismatch, thereby significantly improving the detection capability of complex phenomena such as "intermittent misalignment" and "additive drift". By using a dynamic recognition algorithm to quickly process the spatiotemporal coupling resonance feature data, the system can generate media image misalignment status data in real time, so that the overall processing flow has low latency and high-frequency response characteristics, which is particularly suitable for the recognition tasks of non-stationary states such as abnormal drift and time difference rebound in dynamic media scenes. By introducing a logical rule judgment mechanism into the media picture misalignment status data in step S344, it is possible to output a structured dynamic metadata misalignment judgment result based on preset model rules or adaptive strategies, thereby providing an accurate judgment basis for subsequent picture compensation, transmission control, view switching and other operations.
[0041] Preferably, step S4 includes the following steps:
[0042] Step S41: Performing status judgment on the dynamic metadata inaccuracy judgment result. If the dynamic metadata inaccuracy judgment result is true, performing cross-layer spatial structure analysis on the real-time media data stream to generate cross-layer collaborative mapping relationship data;
[0043] Step S42: Compensate and distribute the real-time media data stream according to the cross-layer collaborative mapping relationship data to generate cross-layer compensation control instruction data; perform cross-layer collaborative compensation on the real-time media data stream to generate new media picture compensation data;
[0044] Step S43: extracting the spatial alignment accuracy of the new media picture compensation data to generate picture spatial alignment accuracy data; adaptively adjusting the parameter weights of the picture spatial alignment accuracy data to generate dynamically optimized parameter weight data;
[0045] Step S44: performing fusion correction processing on the new media picture compensation data and the dynamic optimization parameter weight data to generate a final compensated real-time media data stream to perform new media data transmission management optimization operations.
[0046] The present invention makes a state judgment on the dynamic metadata misalignment judgment result, and triggers a cross-layer spatial structure analysis mechanism when the judgment is true, so as to accurately establish a collaborative mapping relationship between media data and multi-level system components, thereby improving the structural relevance and response flexibility of the system in complex media scenarios. According to the cross-layer collaborative mapping relationship, compensation control instruction data is generated and used to drive the cross-layer compensation execution of real-time media data streams, significantly improving the spatial consistency and timing coordination during the media compensation process, and ensuring high-quality synchronization in multi-device, multi-terminal or multi-channel video interactions. Spatial alignment accuracy extraction is performed on the new media picture compensation data, which can quantitatively reflect the accuracy changes of the current compensation effect in the spatial distribution, and thus realize visual monitoring and index analysis of various spatial offsets, projection misalignments, picture drift and other problems. Using spatial alignment accuracy data, the compensation parameter weights are dynamically adjusted, and a parameter weight adaptive adjustment mechanism is constructed to realize real-time correction of parameter configuration according to changes in compensation effect, thereby enhancing the system's adaptive compensation capabilities in different scenarios and different misalignment characteristics. By fusing and correcting the compensation data with the optimization parameter weights, the resulting compensated real-time media data stream achieves optimal alignment in terms of spatial location, temporal sequence, and inter-frame consistency, fundamentally improving the transmission stability and output consistency of media data. This resulting compensated real-time media data stream can be used to drive downstream media data transmission management optimization operations, such as bandwidth scheduling, image quality control, and node synchronization management. This forms an intelligent closed-loop control system based on compensation accuracy feedback, significantly improving the overall system Quality of Service (QoS).
[0047] Preferably, step S42 includes the following steps:
[0048] Step S421: performing key frame boundary region analysis on the real-time media data stream according to the cross-layer collaborative mapping relationship data to obtain key frame boundary regions; extracting the key frame boundary regions to perform edge structure extraction to obtain frame boundary feature data;
[0049] Step S422: performing perturbation variation analysis on the metadata spatiotemporal parameter set using the frame boundary feature data to generate spatiotemporal perturbation marker data; performing inter-layer difference positioning on the frame boundary feature data and the spatiotemporal perturbation marker data to generate cross-layer error anchor data;
[0050] Step S423: performing cross-layer boundary guidance compensation on the cross-layer error anchoring data to generate cross-layer compensation control instruction data;
[0051] Step S424: superimpose and correct the cross-layer compensation control instruction data and the current real-time media data stream to generate new media picture compensation data; perform frame sequence consistency verification processing on the new media picture compensation data to generate new media picture compensation data.
[0052] By analyzing keyframe boundary regions in real-time media data streams, the present invention accurately identifies structural turning points within keyframes of video. This, combined with edge structure extraction, generates frame boundary feature data, effectively enhancing the ability to perceive changes in video structure hierarchies and providing a high-resolution structural foundation for subsequent compensation positioning. By combining the extracted frame boundary feature data with metadata spatiotemporal parameter sets for perturbation variation analysis, the system constructs spatiotemporal perturbation marker data reflecting different spatial hierarchies and temporal evolution trends, helping to identify potential interference factors and abnormal evolution paths before fine-grained compensation. Combining frame boundary features with spatiotemporal perturbation marker data, the system locates inter-layer differences and generates cross-layer error anchor data. This allows for the establishment of compensation anchor points at the intersection of spatial structure and temporal sequence, providing precise target points and correction directions for subsequent compensation control, significantly improving compensation efficiency and effectiveness. By designing a boundary-guided compensation strategy for the error anchor data and generating compensation control command data with cross-layer guidance and spatial directionality, the system effectively achieves coordinated error correction across multiple layers, enhancing the system's adaptability to heterogeneous devices or multi-protocol environments. After the new media compensation data is generated, a frame consistency verification mechanism is introduced to ensure that the compensation process does not disrupt the temporal continuity of the original video, effectively preventing frame skipping, frame loss, and frame misalignment, thereby ensuring the playback stability and user-perceived coherence of the final output media stream. By superimposing the compensation control command data with the current media data stream for correction, dynamically responsive new media compensation data is generated in real time, achieving a low-latency compensation process that simultaneously corrects and outputs data, meeting the requirements of scenarios such as real-time video analysis and interactive media transmission that require both timeliness and accuracy.
[0053] Preferably, step S423 includes the following steps:
[0054] Step S4231: performing boundary direction aggregation on the spatial offset information in the cross-layer error anchor data to generate boundary guidance vector data; performing inter-layer association path sorting on the boundary guidance vector data to generate hierarchical guidance path data;
[0055] Step S4232: Identify the compensation signal injection point on the hierarchical guidance path data to generate cross-layer signal injection position data;
[0056] Step S4233: Dynamically weighting the cross-layer signal injection position data and the spatial offset information to generate compensation response priority data;
[0057] Step S4234: Perform instruction parameter parsing on the compensation response priority data to generate cross-layer compensation control instruction data.
[0058] The present invention can automatically extract boundary guidance vector data with correction guidance by performing boundary direction aggregation on the spatial offset information in the cross-layer error anchoring data, and further sort out the inter-layer correlation paths to form hierarchical guidance path data with contextual structure logic, thereby significantly improving the spatial target directionality and controllability of cross-layer compensation. Using the cross-layer signal injection position data generated in step S4232, the system can identify the compensation signal injection points of the response paths of different levels based on the sorted paths, accurately control the entry position of the signal action, reduce invalid compensation and resource redundancy, and effectively support the precise delivery of compensation for structured multi-layer media models. The injection position data and the spatial offset information are dynamically weighted and fused. The generated compensation response priority data can flexibly adjust the execution priority of different compensation channels in the processing flow, improve the system's resource scheduling flexibility and the ability to immediately control the response to sudden errors, and ensure that the system can still output stably in high-concurrency scenarios. The injection position data and spatial offset information are dynamically weighted and fused. The generated compensation response priority data can flexibly adjust the execution priority of different compensation channels in the processing flow, improve the system's resource scheduling flexibility and the ability to immediately control the response to sudden errors, and ensure that the system can still output stably in high-concurrency scenarios.
[0059] In this specification, a new media data transmission management system is provided, which is used to execute the above-mentioned new media data transmission management method. The new media data transmission management system includes:
[0060] A media stream acquisition module is used to obtain a real-time media data stream, which includes a video frame sequence and its associated metadata packet set, wherein the metadata packet set includes spatial coordinate data, timestamp data and an operation type identifier;
[0061] A trajectory analysis module is used to extract motion trajectory features from the video frame sequence to generate a video trajectory feature vector, and simultaneously perform spatiotemporal parameter analysis on the metadata packet set to generate a metadata spatiotemporal parameter set;
[0062] The image misalignment detection module is used to calculate the spatial offset and temporal deviation values based on the video trajectory feature vector and metadata spatiotemporal parameter set input into a preset synchronization anomaly detection model; the spatial offset and temporal deviation values are used to detect media image misalignment in the real-time media data stream, and a dynamic metadata misalignment judgment result is obtained;
[0063] The cross-layer collaborative compensation module is used to perform cross-layer collaborative compensation on the real-time media data stream if the dynamic metadata inaccuracy judgment result is true, thereby generating new media picture compensation data; the picture space alignment accuracy of the compensated real-time media data stream is collected through the new media picture compensation data, and the parameter weights of the new media picture compensation data are dynamically adjusted to perform new media data transmission management optimization operations.
[0064] The beneficial effect of the present invention lies in the fact that the video frame sequence and its associated metadata package collection, acquired by the media stream acquisition module, encompass spatial coordinate data, timestamp data, and operation type identifiers. This provides a comprehensive and accurate source of original information for subsequent motion trajectory analysis and spatiotemporal mapping, thereby enhancing the overall data acquisition integrity and accuracy of the system from the source. The video trajectory feature vectors extracted by the trajectory analysis module not only accurately restore the dynamic path of moving targets in the media but also, through linkage with the spatiotemporal parameter set of the metadata, provide a strong foundation for image synchronization analysis and content behavior recognition. By inputting the video trajectory feature vectors and the spatiotemporal parameter set of the metadata into a preset synchronization anomaly detection model to calculate spatial and temporal deviations, it can effectively identify media image misalignment, delay, or drift caused by device latency, network fluctuations, or frame rate jitter, thereby achieving a more sensitive and reliable dynamic metadata misalignment detection mechanism. Once image misalignment is determined, the cross-layer collaborative compensation module triggers multi-level compensation processing for the real-time media data stream. This, combined with the collaborative mapping relationship between different hierarchical structures, enables rapid location and directional correction of offsets, frame errors, and other issues, effectively ensuring the spatial consistency and visual stability of the final image output. By extracting the spatial alignment accuracy of new media compensation data and combining it with a parameter weight adaptive adjustment mechanism, the system can optimize control parameters in real time based on the current transmission status and compensation effect, achieving precise compensation strategy updates, thereby improving the stability, accuracy, and resource utilization efficiency of the entire transmission link. Therefore, this invention improves the stability and intelligence of new media data transmission management through dynamic image delay determination and the introduction of a cross-layer collaborative compensation mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 A schematic diagram of the steps of a new media data transmission management method;
[0066] Figure 2 for Figure 1 Detailed implementation steps of step S2 in FIG.
[0067] Figure 3 for Figure 1 Detailed implementation steps of step S3 in FIG.
[0068] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0069] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. It is obvious that the embodiments described are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.
[0070] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller approaches.
[0071] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.
[0072] To achieve this, please refer to Figures 1 to 3 , a new media data transmission management method, the method comprising the following steps:
[0073] Step S1: obtaining a real-time media data stream, which includes a video frame sequence and its associated metadata packet set, wherein the metadata packet set includes spatial coordinate data, timestamp data, and an operation type identifier;
[0074] Step S2: extracting motion trajectory features from the video frame sequence to generate a video trajectory feature vector, and performing spatiotemporal parameter analysis on the metadata packet set to generate a metadata spatiotemporal parameter set;
[0075] Step S3: Based on the video trajectory feature vector and the metadata spatiotemporal parameter set, the spatial offset and temporal deviation values are input into a preset synchronization anomaly detection model; the spatial offset and temporal deviation values are used to perform media image misalignment judgment on the real-time media data stream to obtain a dynamic metadata misalignment judgment result;
[0076] Step S4: If the dynamic metadata misalignment judgment result is true, cross-layer collaborative compensation is performed on the real-time media data stream to generate new media picture compensation data; the picture space alignment accuracy of the compensated real-time media data stream is collected through the new media picture compensation data, and the parameter weights of the new media picture compensation data are dynamically adjusted to perform new media data transmission management optimization operations.
[0077] This invention combines video trajectory feature vectors with spatiotemporal parameters of metadata and utilizes a synchronization anomaly detection model to effectively detect spatial and temporal offsets, achieving high-precision identification of media synchronization misalignment and significantly improving the synchronization accuracy of multi-source real-time media. Upon detecting metadata misalignment, a cross-layer collaborative compensation mechanism is automatically triggered. This mechanism not only generates compensation data but also dynamically collects spatial alignment accuracy and adaptively adjusts parameter weights to achieve real-time media image repair and transmission strategy optimization, enhancing the system's robustness and self-recovery capabilities. It supports the joint parsing and processing of video frames and their metadata (such as coordinates, timestamps, and operation types) in complex dynamic scenes, making it suitable for a variety of highly interactive scenarios such as surveillance, AR / VR, and remote collaboration, improving the system's adaptability and versatility. By dynamically adjusting the parameter weights of the compensation data, bandwidth control and data compression strategies are optimized, helping to reduce latency, reduce redundant data, and improve the overall transmission efficiency and controllability of media data streams. The logical logic of each step is clear and can be flexibly embedded in existing media processing or data synchronization systems. Both the synchronization anomaly detection model and compensation mechanism support model upgrades and policy customization, facilitating future iterations and intelligent expansion of the system. Therefore, the present invention improves the stability and intelligence level of new media data transmission management by dynamically distinguishing picture delays and introducing a cross-layer collaborative compensation mechanism.
[0078] In the embodiment of the present invention, reference Figure 1 FIG. 1 is a flow chart of a new media data transmission management method according to the present invention. In this example, the new media data transmission management method includes the following steps:
[0079] Step S1: obtaining a real-time media data stream, which includes a video frame sequence and its associated metadata packet set, wherein the metadata packet set includes spatial coordinate data, timestamp data, and an operation type identifier;
[0080] Step S2: extracting motion trajectory features from the video frame sequence to generate a video trajectory feature vector, and performing spatiotemporal parameter analysis on the metadata packet set to generate a metadata spatiotemporal parameter set;
[0081] Step S3: Based on the video trajectory feature vector and the metadata spatiotemporal parameter set, the spatial offset and temporal deviation values are input into a preset synchronization anomaly detection model; the spatial offset and temporal deviation values are used to perform media image misalignment judgment on the real-time media data stream to obtain a dynamic metadata misalignment judgment result;
[0082] Step S4: If the dynamic metadata misalignment judgment result is true, cross-layer collaborative compensation is performed on the real-time media data stream to generate new media picture compensation data; the picture space alignment accuracy of the compensated real-time media data stream is collected through the new media picture compensation data, and the parameter weights of the new media picture compensation data are dynamically adjusted to perform new media data transmission management optimization operations.
[0083] In an embodiment of the present invention, a multimedia processing system receives a real-time media data stream from a front-end acquisition device (such as a smart camera or embedded video module). This data stream contains an encoded video frame sequence (supporting H.264 / H.265 encoding formats) and its corresponding metadata package set. Each metadata package includes: 3D spatial coordinate data (X, Y, Z, in pixels or millimeters); a frame-level timestamp (supporting millisecond accuracy); and an operation type identifier (including rotation, scaling, and displacement, represented in an enumerated encoding format). The video frame sequence and metadata package set undergo frame-level decoding and synchronous caching via a unified video decoding interface (supporting the FFmpeg framework) before entering the subsequent feature extraction module for processing. An integrated motion trajectory analysis engine (based on OpenCV and the TensorRT deep learning library optimization module) performs frame-by-frame motion trajectory detection on the video frame sequence. Keypoint tracking algorithms (such as KLT optical flow or DeepSORT) and a skeleton pose estimation model are used to extract the target trajectory in the video, generating a video trajectory feature vector containing the direction vector, velocity change rate, and trajectory curvature. Simultaneously, the system parses the metadata data packets, performs timestamp alignment and spatial coordinate normalization, and generates a standardized metadata spatiotemporal parameter set, including parameters such as trajectory start and end times, spatial activity range, and operation frequency. The video trajectory feature vector and metadata spatiotemporal parameter set are then fed into a pre-defined synchronization anomaly detection model. This model, which utilizes a dual-channel LSTM architecture and an attention mechanism fusion module, calculates the spatial offset (Δx, Δy, Δz) and temporal deviation (Δt) between the two in real time. The spatial offset reflects the difference between the spatial position indicated by the metadata and the actual motion path of the target in the image, while the temporal deviation measures the synchronization error of the data stream (in milliseconds). Based on the offset threshold strategy (for example, Δx or Δy exceeding 15 pixels, Δt exceeding 50 milliseconds), the system determines whether the current frame exhibits media misalignment and outputs a dynamic metadata misalignment determination result (a Boolean flag). If the determination in step S3 is true, the system proceeds to step S4, triggering a cross-layer collaborative compensation mechanism. This mechanism comprises an image alignment network (e.g., a pixel-level alignment network based on a U-Net architecture) and a metadata back-interpolation module. The system uses spatial offsets to inversely adjust the spatial mapping of key image regions. It also reconstructs and interpolates lost frames based on temporal deviations to generate new media compensation data. High-precision image overlap metrics (such as structural similarity (SSIM) and feature point alignment) are then used to assess the spatial alignment accuracy of the compensated video frames with the ideal metadata trajectory. Based on these assessments, the system dynamically optimizes compensation weighting parameters (for example, increasing temporal interpolation weights or reducing spatial transformation amplitudes), thereby optimizing new media data transmission management and ensuring the consistency and accuracy of image and metadata.Ultimately, the compensated real-time media data stream will be repackaged through an optimized encoding strategy and transmitted to the back-end processing system, achieving improved media synchronization quality and enhanced abnormal fault tolerance.
[0084] Preferably, step S1 includes the following steps:
[0085] Step S11: continuously capturing data of the video frame input signal to obtain original video frame sequence data;
[0086] Step S12: performing frame-level data decoding on the original video frame sequence data to generate structured video frame data and extracting a frame sequence of structured time-frequency frame data to obtain a video frame sequence;
[0087] Step S13: using the video frame sequence to identify control signaling tags on the video frame input signal, generating a preliminary metadata segment set; performing field extraction on the metadata segment set, extracting spatial coordinate data, timestamp data, and operation type identifier, and generating an associated metadata package set;
[0088] Step S14: synchronously mapping the video frame sequence and the associated metadata packet set to obtain a real-time media data stream.
[0089] In an embodiment of the present invention, a continuous data capture operation is performed on the video frame input signal based on a front-end acquisition module, and the acquisition module can be an industrial camera or an intelligent monitoring terminal with a high-definition image sensor. A buffer queue mechanism and a double buffer strategy are used to achieve high-speed sampling and inter-frame stability maintenance to obtain uninterrupted original video frame sequence data. The sequence data contains uncompressed or unencoded video frame signal data and supports frame-by-frame timing numbering. The captured original video frame sequence data is frame-level decoded, and the decoding module supports the H.264 / H.265 encoding standard and is adapted to mainstream decoding frameworks such as FFmpeg and GStreamer. After decoding, the system generates structured video frame data based on the intra-frame pixel matrix, timing index and coding hierarchy structure. On this basis, the system extracts the time-frequency domain features contained in the structured data to obtain a structured time-frequency frame data frame sequence with clear inter-frame timestamps, key frame marks and content summaries, forming a highly reliable video frame sequence for subsequent processing. The system uses the aforementioned video frame sequence to identify control signaling tags from the raw video frame input signal. It then generates a preliminary metadata segment set by identifying embedded operation instructions, frame type tags, or external additional signals (such as MPEG SEI information or RTMP stream metadata). The system then uses a field parsing engine to extract semantic-level fields from this set. Based on a pre-set field mapping table, the system extracts: spatial coordinate data (such as the coordinates of the target location in the video scene or camera pose information); timestamp data (the time of frame capture recorded in milliseconds); and operation type identifiers (used to identify operation events such as zooming, rotating, and view switching). Ultimately, a set of related metadata packets is constructed, ensuring a uniform field encoding format to support cross-module parsing. The generated video frame sequence and metadata packet set are input to the synchronization mapping module. This module maps each video frame to its corresponding metadata packet based on a frame-level timestamp comparison mechanism and content matching strategy. The synchronization mapping process considers inter-frame delay tolerance and metadata update intervals, maintaining spatiotemporal consistency between the two using a sliding window and resampling algorithm. The final output is a complete, structured, and time-ordered real-time media data stream, which contains each frame of picture data and its corresponding spatial and temporal feature metadata for use by subsequent media processing and anomaly detection modules.
[0090] Preferably, in step S13, using the video frame sequence to perform control signaling tag recognition on the video frame input signal includes:
[0091] Extracting frame header data of a video frame input signal using a video frame sequence to generate video frame synchronization index data;
[0092] Divide the video frame synchronization index data into content segments to generate signaling candidate area data;
[0093] Scan key bits of the signaling candidate area data to generate control signaling feature segment data;
[0094] Performing operation identification matching on the control signaling feature segment data to generate initial operation type identification data;
[0095] Perform frame sequence timing mapping processing on the initial operation type identification data to generate a preliminary metadata segment set.
[0096] In this embodiment of the present invention, frame header data extraction is performed on the video frame input signal using frame-level index information in the video frame sequence. This frame header data typically includes a frame type identifier (e.g., I-frame, P-frame, B-frame), a frame sequence number, a timestamp field, and a stream channel identifier. The extracted frame header data serves as the basis for intra-frame synchronization features and is used to generate video frame synchronization index data, ensuring that subsequent signaling markers can be accurately located within the frame structure. The generated video frame synchronization index data is then segmented into content segments. Based on the length and position offset information of the frame header fields, the index data is logically partitioned into areas such as a payload area, an extended parameter area, and a control information reserved area. Based on heuristic thresholds or protocol description files, the system extracts areas potentially containing control signaling and generates candidate signaling region data with defined spatial boundaries. Key bit scanning is performed on the candidate signaling region data. This scanning is based on a bitstream parsing model to identify specific bit patterns (e.g., "1010" edge features, control instruction start bits, special symbol markers, etc.). By matching and counting feature patterns within frames, bit segments with a high correlation with control signaling are extracted, generating control signaling feature segment data with structural integrity and trigger semantics. This control signaling feature segment data is then fed into the operation identifier matching module. Combined with the system's built-in operation instruction template library, the module identifies the operation type (such as zoom, translation, angle rotation, brightness adjustment, etc.) through bit sequence structure matching and semantic rule parsing. Initial operation type identification data corresponding to each signaling feature segment is generated. This process accelerates operation instruction classification efficiency through hash mapping and support vector structures. This initial operation type identification data undergoes frame-to-frame temporal mapping, establishing a temporal alignment between the operation type and the specific frame based on the timestamp and frame number in the video frame sequence. This process ensures that each operation type has a clear occurrence time and applicable frame range. Finally, the temporal mapping results are integrated to generate a complete preliminary metadata segment set, providing the data foundation for subsequent field extraction and metadata package generation.
[0097] As an example of the present invention, refer to Figure 2 As shown, in this example, step S2 includes:
[0098] Step S21: analyzing the key frames of the video frame sequence and performing key frame screening to generate media video motion key frame data; continuously tracking the target position of the video motion key frame data to obtain trajectory point sequence data;
[0099] Step S22: performing trajectory coherence fitting on the trajectory point sequence data to generate a video trajectory feature vector;
[0100] Step S23: performing time stamp calculation on the metadata package set to generate synchronization timing parameter data; performing spatial position calculation on the spatial coordinates in the metadata package set to generate position mapping parameter data;
[0101] Step S24: jointly associate the synchronization timing parameter data and the position mapping parameter data to generate a metadata spatiotemporal parameter set.
[0102] In this embodiment of the present invention, keyframe filtering is performed on the input video frame sequence. A detection algorithm based on the rate of change of image structure is used to identify frames with significant inter-frame motion changes. For example, SIFT / SURF feature differences, edge density changes, and grayscale histogram similarity are used for comprehensive evaluation. This allows the extraction of keyframe data representing the main trends in the video content. Based on the keyframe data, the system further performs continuous target position tracking on the moving targets contained therein. This tracking method can be based on a deep neural network (such as YOLO or DeepSORT) or a multi-target tracking model based on a Kalman filter. The position of each moving target in each keyframe is identified and numbered, thereby constructing a trajectory point sequence data representing the target's motion path. This obtained trajectory point sequence data is then subjected to trajectory continuity fitting using Bezier curve fitting, polynomial interpolation, or least-squares-based curve regression methods to improve the smoothness and continuity of the trajectory path and eliminate position jitter and discontinuities between keyframes. After fitting, the system extracts high-dimensional features from the continuous trajectory morphology, including path curvature changes, directional change rates, time span distribution, and trajectory density distribution. Using vectorized encoding, it constructs a video trajectory feature vector describing the overall video motion state, which serves as an input to the subsequent synchronization anomaly detection model. Timestamp resolution is performed on the previously generated metadata data set, extracting the timestamp field of each metadata record and normalizing it based on the frame clock frequency. Furthermore, an inter-frame differential time calculation strategy is incorporated to compensate for frame loss or out-of-order sequences, ultimately generating frame-level accurate synchronization timing parameter data. The system then performs spatial position resolution on the spatial coordinate information contained in the metadata data set. This process, based on a calibration model of the original coordinate system, accounts for factors such as device shooting angle, optical distortion, and spatial projection transformations. Using homography or perspective projection reconstruction, the two-dimensional coordinates are restored to the world coordinate system or standard image plane, generating position mapping parameter data for spatial alignment. After obtaining the synchronization timing parameter data and position mapping parameter data, the system jointly correlates them using unique frame identifiers and timestamps. This association process utilizes a multidimensional hash index structure and a windowed time alignment strategy to ensure a one-to-one correspondence between spatial and temporal data. Ultimately, the system structures the results of this combined processing into a set of metadata spatiotemporal parameters, including time-position pairs and operation type indexes. This serves as input for the media data synchronization detection and compensation module.
[0103] As an example of the present invention, refer to Figure 3 As shown, in this example, step S3 includes:
[0104] Step S31: performing time series synchronization on the video trajectory feature vector and the metadata spatiotemporal parameter set to generate synchronized input feature data;
[0105] Step S32: inputting the synchronous input feature data into a preset synchronous anomaly detection model to perform spatial coordinate error calculation and timestamp offset analysis to generate spatial offset data and time deviation value data;
[0106] Step S33: performing position error threshold comparison on the spatial offset data to generate image position misalignment identification data; performing delay threshold comparison on the time deviation value data to generate time synchronization abnormality identification data;
[0107] Step S34: performing media picture misalignment determination based on the picture position misalignment identification data and the time synchronization anomaly identification data, and generating a dynamic metadata misalignment determination result.
[0108] In this embodiment of the present invention, time series synchronization processing is performed using the previously generated video trajectory feature vectors and metadata spatiotemporal parameter sets as input. During this process, the system performs a one-to-one alignment based on the global timestamp sequence and frame synchronization index, aligns data sequences from different sources using the dynamic time warping (DTW) algorithm, and constructs synchronized input feature data that integrates frame timestamps, spatial position pairs, and motion features. The synchronized input feature data structure includes, but is not limited to, frame sequence number, frame timestamp, trajectory vector segment, corresponding spatial position, and operation type identifier, maintaining a unified time series alignment format to facilitate unified processing by subsequent models. This synchronized input feature data is fed into a pre-defined synchronization anomaly detection model. This model can be based on a multi-layer neural network (such as a CNN+LSTM architecture), an attention mechanism model, or a spatiotemporal modeling framework based on a graph neural network (GNN), capable of jointly analyzing video trajectory changes and metadata spatiotemporal anomalies. The model calculates the Euclidean distance between the predicted spatial position in the trajectory vector and the metadata position mapping value to obtain the spatial coordinate error value. Simultaneously, a sliding window analysis and hysteresis value extraction are performed on the timestamp field to obtain the timestamp offset value. Finally, a structured output is generated: spatial offset data reflects the degree of spatial mismatch between the trajectory path and the metadata location; and temporal deviation data reflects the degree of temporal synchronization delay between the trajectory features and the metadata. More specifically, the model construction process includes the following: input features are derived from the synchronized input feature data generated in step S31, primarily including: video frame ID: Frame_ID, frame timestamp: Timestamp, video trajectory feature vector: Trajectory_Vector (e.g., [x1, y1], [x2, y2], ..., [xn, yn]), spatial coordinate value: Spatial_Meta (e.g., [x_meta, y_meta] in metadata), and operation type identifier: Action_Type (which can be one-hot encoded). After unified formatting and normalization, these features form a time series tensor structure: Input_Tensor∈ℝ^(T×D), where T is the time series length (i.e., the number of frame windows) and D is the sum of feature dimensions per frame. The model employs a multi-layered spatiotemporal feature fusion network structure, including a bidirectional LSTM (Bi-LSTM) to extract temporal features from the input sequence. The output structure is: H_t∈ℝ^(T×H), where H is the hidden state dimension. This enhances the ability to model the temporal dependencies between trajectory motion and metadata. The trajectory vector and spatial metadata are combined to calculate the frame-level spatial residual: ΔS_t=||(x_t, y_t) - (x_meta_t, y_meta_t)||. This residual channel is then introduced into subsequent neural layers to enhance spatial anomaly perception.A convolution operation (e.g., 1D-CNN) is performed on the sequence of frame timestamp differences to detect sudden delays. Local statistical features with sliding windows are used as input for temporal anomaly detection. Spatial residuals, temporal offsets, and LSTM hidden states are fused. A fully connected neural network (MLP) and softmax are used to output the discrimination results: Output 1: spatial offset (Regression), Output 2: temporal deviation (Regression), and Output 3: misalignment indicator (Binary Classification). To train this model, a training set containing both normal and anomalous data is required. The main steps are as follows: Real video streams and metadata are collected from multiple scenarios, such as camera surveillance streams and human-computer interaction recordings. Simulated offsets and temporal delays are performed on some video frames to construct spatial / temporal misalignment samples. The misalignment type of each frame is labeled (no misalignment, positional misalignment, temporal misalignment, or mixed misalignment). Trajectory vectors and metadata fields are extracted to form a complete training input tensor and corresponding supervision label set. The model is trained using an end-to-end multi-task loss function, including: spatial offset loss: L_s = MSE(ΔS_pred, ΔS_true), temporal offset loss: L_t = MSE(ΔT_pred, ΔT_true), and classification discriminant loss: L_c = CrossEntropy(y_pred, y_true). The total loss function is a weighted combination: ;in 、 、 α is a hyperparameter that controls the weights of the three tasks and is typically set to a balanced value (e.g., α = β = 1, γ = 2). The training optimizer uses Adam with an adaptive learning rate and supports early stopping to prevent overfitting. After training, the model can be deployed as an embedded inference module or edge node detection service. Input: Real-time synchronized input feature data sequence; Output: Spatial offset value, temporal offset value, and a Boolean flag indicating whether the frame is misaligned. Spatial offset data is subjected to position error threshold comparison. Spatial offset values are compared frame by frame based on a set spatial accuracy threshold (e.g., pixel distance, millimeter-level units). Any value exceeding the threshold is marked as "misaligned," generating frame-level position misalignment identification data. Similarly, the system compares temporal offset data against a delay threshold. For example, a maximum allowable delay (e.g., Δt_max = 50ms) is set. Each frame's temporal offset is evaluated. Any value exceeding the limit is marked as "temporal synchronization anomaly," and corresponding temporal anomaly identification data is generated. After obtaining frame-level identification of position misalignment and temporal anomaly, the system performs media image misalignment detection. The discrimination logic can be based on Boolean combination rules, discriminant functions, or fusion models. For example, a joint discrimination rule can be set as follows: if the frame's image position misalignment flag = 1 and the time synchronization anomaly flag = 1, the frame is judged to be a "dynamic metadata misalignment frame." This is performed in batches on a frame-by-frame basis, ultimately outputting a set of results reflecting the synchronization accuracy of the entire media data stream, known as the dynamic metadata misalignment judgment result, which is used to trigger subsequent compensation mechanisms.
[0109] Preferably, step S34 includes the following steps:
[0110] Step S341: performing spatial anomaly clustering on the image position misalignment identification data to generate spatial misalignment distribution data;
[0111] Step S342: performing time series focusing on the time synchronization anomaly identification data to generate time misalignment aggregation data;
[0112] Step S343: performing segmented co-frequency vibration superposition on the spatial misalignment distribution data and the temporal misalignment aggregation data to generate spatiotemporal coupling resonance feature data; performing dynamic recognition processing on the spatiotemporal coupling resonance feature data to generate media image misalignment status data;
[0113] Step S344: Performing logic rule judgment on the media image misalignment status data to generate a dynamic metadata misalignment judgment result.
[0114] In this embodiment of the present invention, the system receives image position misalignment identification data output by the spatial offset analysis module, which identifies the areas of spatial misalignment detected in each frame. The system then clusters this data using a density clustering algorithm (such as DBSCAN) or a K-means-based spatial clustering algorithm. The clustering results are used to characterize spatial anomaly features such as the density distribution, concentrated areas, and contour boundaries of the spatial misalignment, thereby generating multi-scale, multi-region spatial misalignment distribution data for subsequent coupled analysis. The temporal synchronization anomaly identification data output by the temporal deviation value analysis module is extracted and subjected to focusing processing based on the frame time series. This processing uses a sliding window anomaly integration algorithm or a temporal mutation detection algorithm (such as CUSUM or SPOT) to dynamically aggregate the density of anomaly time points. The results obtained after focusing processing can reflect the frequently occurring segments of misalignment on the time axis, the mutation time nodes, and the stability evolution trend, ultimately generating temporal misalignment aggregated data with a clear clustering pattern. The spatial misalignment distribution data and temporal misalignment aggregation data are input into the coupling analysis module. A segmented co-frequency vibration superposition strategy is employed, using spatial cluster blocks and temporal focus windows as synchronization units to perform local frequency domain superposition analysis. Specifically, a bidirectional Fourier transform of the spatial-temporal co-spectrum is constructed to cross-match spatial frequencies and temporal periods. This process generates high-dimensional cross-feature spatiotemporal coupling resonance signature data, reflecting the synchronization and resonance strength of the misalignment in the spatiotemporal dimension. The system further performs dynamic recognition processing on this coupling data. Using a pre-trained deep neural network model (such as an LSTM-Transformer structure), it determines whether the coupling resonance has reached the image misalignment trigger threshold, ultimately generating stateful media image misalignment status data. Conditional judgment is performed based on the media image misalignment status data and a pre-set logical rule base. Logical rules may include: the spatial cluster density exceeds a set threshold; the temporal misalignment window frequency reaches a set standard; the spatiotemporal resonance activation value reaches the upper limit of the confidence interval; and the hit rate of multiple recognitions meets stability requirements. When the trigger conditions are met, the system determines that there are significant spatiotemporal synchronization anomalies in the current real-time media data stream, generates a dynamic metadata misalignment judgment result, and uses it as input to start the subsequent cross-layer collaborative compensation process.
[0115] Preferably, step S4 includes the following steps:
[0116] Step S41: Performing status judgment on the dynamic metadata inaccuracy judgment result. If the dynamic metadata inaccuracy judgment result is true, performing cross-layer spatial structure analysis on the real-time media data stream to generate cross-layer collaborative mapping relationship data;
[0117] Step S42: Compensate and distribute the real-time media data stream according to the cross-layer collaborative mapping relationship data to generate cross-layer compensation control instruction data; perform cross-layer collaborative compensation on the real-time media data stream to generate new media picture compensation data;
[0118] Step S43: extracting the spatial alignment accuracy of the new media picture compensation data to generate picture spatial alignment accuracy data; adaptively adjusting the parameter weights of the picture spatial alignment accuracy data to generate dynamically optimized parameter weight data;
[0119] Step S44: performing fusion correction processing on the new media picture compensation data and the dynamic optimization parameter weight data to generate a final compensated real-time media data stream to perform new media data transmission management optimization operations.
[0120] In this embodiment of the present invention, the compensation process is initiated upon receiving a true result from dynamic metadata misalignment. First, a multi-level spatial structure modeling analysis is performed on the real-time media data stream to extract its spatial topological relationships at the logical frame, pixel, and content semantic layers. Using a graph neural network (GNN) and a multi-scale spatial modeling algorithm, the position information, hierarchical features, and region boundaries within the media data frame structure are jointly modeled to generate a cross-layer mapping structure model. Combining the spatial coordinate dimensions in the metadata with the logical frame sequence relationship, cross-layer collaborative mapping relationship data is constructed for compensation calculations. Based on this generated cross-layer collaborative mapping relationship data, the system constructs a compensation path map and implements a dynamic control instruction scheduling mechanism. This mechanism assigns different spatial compensation strategies based on the severity and type of the misaligned area, including but not limited to: interpolation compensation using adjacent frame references; spatial feature remapping; and multi-scale affine transformation alignment. After executing the compensation instructions, the system generates control-level cross-layer compensation control instruction data, which drives the underlying compensation module to reconstruct the image and generate new media image compensation data with greater structural consistency. The generated new media image compensation data undergoes spatial alignment accuracy analysis. Using geometric consistency detection algorithms (such as phase correlation analysis and SIFT+RANSAC matching), metrics such as spatial offset, overlap, and edge misalignment rate are calculated for key areas before and after compensation, generating quantitative spatial alignment accuracy data. Based on this alignment accuracy data, the system then dynamically adjusts the spatial mapping parameters and adaptive filter parameters involved in the original compensation process using a gradient back-adjustment algorithm or a Bayesian weight optimization model. This generates dynamically optimized parameter weights that respond to scene changes, improving compensation accuracy and robustness. The generated new media image compensation data and dynamically optimized parameter weights are input into a fusion correction module. This module, using a dynamic fusion function and edge enhancement correction mechanism, compensates for local details and enhances texture consistency in areas with minor errors. After fusion correction, the module outputs a final compensated real-time media data stream with enhanced visual and structural consistency. This stream serves as the optimized target data stream and is then fed into the new media data transmission channel for subsequent management tasks such as transmission scheduling, network adaptation, and terminal rendering optimization.
[0121] Preferably, step S42 includes the following steps:
[0122] Step S421: performing key frame boundary region analysis on the real-time media data stream according to the cross-layer collaborative mapping relationship data to obtain key frame boundary regions; extracting the key frame boundary regions to perform edge structure extraction to obtain frame boundary feature data;
[0123] Step S422: performing perturbation variation analysis on the metadata spatiotemporal parameter set using the frame boundary feature data to generate spatiotemporal perturbation marker data; performing inter-layer difference positioning on the frame boundary feature data and the spatiotemporal perturbation marker data to generate cross-layer error anchor data;
[0124] Step S423: performing cross-layer boundary guidance compensation on the cross-layer error anchoring data to generate cross-layer compensation control instruction data;
[0125] Step S424: superimpose and correct the cross-layer compensation control instruction data and the current real-time media data stream to generate new media picture compensation data; perform frame sequence consistency verification processing on the new media picture compensation data to generate new media picture compensation data.
[0126] In this embodiment of the present invention, keyframes in a real-time media data stream are identified based on cross-layer collaborative mapping relationship data. Region boundary identification algorithms, such as those based on region growing and graph cuts, are then applied to the keyframe images to delineate keyframe boundary regions. Subsequently, a combination of edge operators (such as a Canny-Sobel joint filter) is used to perform multi-scale edge structure analysis on the boundary regions, extracting features such as boundary texture, corner distribution, and shape curvature. This generates frame boundary feature data describing the spatial structural differences of the keyframes. This frame boundary feature data is then combined with a set of metadata spatiotemporal parameters to perform a perturbation variation analysis model. This model uses a small-amplitude perturbation simulation mechanism and a temporal drift window to estimate perturbation offsets on timestamp and coordinate data, generating spatiotemporal perturbation marker data reflecting potential alignment errors. Based on this frame boundary feature data and the spatiotemporal perturbation marker data, the system performs a multi-layer fusion difference localization algorithm to compare the spatial consistency and temporal correlation between the structure and control layers, identifying regions with structural drift or alignment offsets. This generates quantifiable cross-layer error anchor data, which is used to calibrate the specific location, scope, and extent of compensation operations. For cross-layer error anchor data, the system utilizes a boundary-guided compensation strategy framework to construct a compensation path and parameter set. Combined with the current media data frame structure and temporal synchronization information, this system generates cross-layer compensation control command data. This command data, centered around multi-layer frame indices, carries key parameters such as spatial adjustment matrices, affine repair coefficients, and temporal interpolation factors. These serve as the basis for driving the compensation module to perform cross-layer mapping corrections, ensuring the coordinated optimization of structural consistency and spatiotemporal continuity. The cross-layer compensation control command data is then fused, overlaid, and spatially corrected with the real-time media data stream. Using algorithms such as image registration, inter-frame reconstruction, and texture mapping, local redrawing and alignment repair are performed on the compensation area, ultimately generating structurally complete new media image compensation data. To ensure the temporal consistency and multi-frame synchronization of the new media image compensation data, the system further incorporates a frame sequence consistency verification mechanism. This mechanism utilizes a sliding time window and inter-frame variability analysis to detect and correct frame-level jumps, frame errors, or temporal misalignments, ensuring that the output meets continuous playback and encoding transmission standards. The new media picture compensation data finally outputted will serve as the core content of subsequent optimization processing and data stream transmission, and will be used by step S43 to call and execute the spatial alignment accuracy extraction and parameter adaptive optimization tasks.
[0127] Preferably, step S423 includes the following steps:
[0128] Step S4231: performing boundary direction aggregation on the spatial offset information in the cross-layer error anchor data to generate boundary guidance vector data; performing inter-layer association path sorting on the boundary guidance vector data to generate hierarchical guidance path data;
[0129] Step S4232: Identify the compensation signal injection point on the hierarchical guidance path data to generate cross-layer signal injection position data;
[0130] Step S4233: Dynamically weighting the cross-layer signal injection position data and the spatial offset information to generate compensation response priority data;
[0131] Step S4234: Perform instruction parameter parsing on the compensation response priority data to generate cross-layer compensation control instruction data.
[0132] In this embodiment of the present invention, spatial offset information is first extracted from cross-layer error anchor data. This information typically includes spatial displacement data between different layers (such as video frames, control signal frames, and timestamp frames). Using a boundary direction aggregation algorithm, this spatial offset information is statistically and directionally analyzed, clustering individual error points according to their spatial distribution patterns to generate representative boundary guidance vector data. Next, based on this boundary guidance vector data, the system uses an inter-layer correlation path combing algorithm to analyze the spatial misalignment correlations between different layers and identify optimal paths for error correction at each layer. This process, through boundary correction curve fitting and path planning algorithms, ultimately generates temporally and spatially consistent hierarchical guidance path data for subsequent compensation path planning. The hierarchical guidance path data is then used to identify compensation signal injection points. This process, based on a dynamic signal injection strategy, identifies the most suitable injection points (e.g., the nodes most in need of correction in terms of time or space) through segmented analysis of path data. Further analysis of the error type (e.g., spatial offset, time delay, etc.) at each layer path point generates cross-layer signal injection location data, which are the key areas for cross-layer compensation. This step uses an adaptive signal positioning algorithm (such as the UWB positioning algorithm or the Kalman positioning algorithm) to self-adjust based on historical data and the accuracy of feedback control to ensure accurate injection point identification. Dynamic weight fusion is performed based on cross-layer signal injection location data and spatial offset information. During this process, the error value of the spatial offset information is combined with the priority data of the signal injection point. Compensation response priority data is dynamically generated based on a preset weighting algorithm (such as the weighted mean method or weighted nonlinear optimization method). This data reflects the weight of each layer in the compensation operation and the importance of the response. Based on this real-time priority data, the system also comprehensively evaluates factors such as the severity of the error and the urgency of the correction need to ensure that the most critical or impactful error areas are compensated first. The compensation response priority data is input into the command parsing module, which generates cross-layer compensation control command data with specific control parameters based on the priority data. The command data includes, but is not limited to, spatiotemporal adjustment parameters for each compensation area, specific compensation operation timing, and frame-level correction factors. The instruction generation is parsed by a rule-based instruction generation system, and the priority areas of the compensation response priority are mapped into a set of standardized control instructions, which are finally transmitted to modules at each level through the control platform to perform cross-layer compensation operations.
[0133] In this specification, a new media data transmission management system is provided, which is used to execute the above-mentioned new media data transmission management method. The new media data transmission management system includes:
[0134] A media stream acquisition module is used to obtain a real-time media data stream, which includes a video frame sequence and its associated metadata packet set, wherein the metadata packet set includes spatial coordinate data, timestamp data and an operation type identifier;
[0135] A trajectory analysis module is used to extract motion trajectory features from the video frame sequence to generate a video trajectory feature vector, and simultaneously perform spatiotemporal parameter analysis on the metadata packet set to generate a metadata spatiotemporal parameter set;
[0136] The image misalignment detection module is used to calculate the spatial offset and temporal deviation values based on the video trajectory feature vector and metadata spatiotemporal parameter set input into a preset synchronization anomaly detection model; the spatial offset and temporal deviation values are used to detect media image misalignment in the real-time media data stream, and a dynamic metadata misalignment judgment result is obtained;
[0137] The cross-layer collaborative compensation module is used to perform cross-layer collaborative compensation on the real-time media data stream if the dynamic metadata inaccuracy judgment result is true, thereby generating new media picture compensation data; the picture space alignment accuracy of the compensated real-time media data stream is collected through the new media picture compensation data, and the parameter weights of the new media picture compensation data are dynamically adjusted to perform new media data transmission management optimization operations.
[0138] The beneficial effect of the present invention lies in the fact that the video frame sequence and its associated metadata package collection, acquired by the media stream acquisition module, encompass spatial coordinate data, timestamp data, and operation type identifiers. This provides a comprehensive and accurate source of original information for subsequent motion trajectory analysis and spatiotemporal mapping, thereby enhancing the overall data acquisition integrity and accuracy of the system from the source. The video trajectory feature vectors extracted by the trajectory analysis module not only accurately restore the dynamic path of moving targets in the media but also, through linkage with the spatiotemporal parameter set of the metadata, provide a strong foundation for image synchronization analysis and content behavior recognition. By inputting the video trajectory feature vectors and the spatiotemporal parameter set of the metadata into a preset synchronization anomaly detection model to calculate spatial and temporal deviations, it can effectively identify media image misalignment, delay, or drift caused by device latency, network fluctuations, or frame rate jitter, thereby achieving a more sensitive and reliable dynamic metadata misalignment detection mechanism. Once image misalignment is determined, the cross-layer collaborative compensation module triggers multi-level compensation processing for the real-time media data stream. This, combined with the collaborative mapping relationship between different hierarchical structures, enables rapid location and directional correction of offsets, frame errors, and other issues, effectively ensuring the spatial consistency and visual stability of the final image output. By extracting the spatial alignment accuracy of new media compensation data and combining it with a parameter weight adaptive adjustment mechanism, the system can optimize control parameters in real time based on the current transmission status and compensation effect, achieving precise compensation strategy updates, thereby improving the stability, accuracy, and resource utilization efficiency of the entire transmission link. Therefore, this invention improves the stability and intelligence of new media data transmission management through dynamic image delay determination and the introduction of a cross-layer collaborative compensation mechanism.
[0139] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.
[0140] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A new media data transmission management method, characterized in that: The following steps are involved: Step S1: obtaining a real-time media data stream, which includes a video frame sequence and its associated metadata packet set, wherein the metadata packet set includes spatial coordinate data, timestamp data, and an operation type identifier; Step S2: extracting motion trajectory features from the video frame sequence to generate a video trajectory feature vector, and performing spatiotemporal parameter analysis on the metadata packet set to generate a metadata spatiotemporal parameter set; Step S3: Based on the video trajectory feature vector and the metadata spatiotemporal parameter set, the spatial offset and temporal deviation values are input into a preset synchronization anomaly detection model; the spatial offset and temporal deviation values are used to perform media image misalignment judgment on the real-time media data stream to obtain a dynamic metadata misalignment judgment result; Step S4: If the dynamic metadata misalignment judgment result is true, cross-layer collaborative compensation is performed on the real-time media data stream to generate new media picture compensation data; the picture space alignment accuracy of the compensated real-time media data stream is collected through the new media picture compensation data, and the parameter weights of the new media picture compensation data are dynamically adjusted to perform new media data transmission management optimization operations.
2. The new media data transmission management method according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: continuously capturing data of the video frame input signal to obtain original video frame sequence data; Step S12: performing frame-level data decoding on the original video frame sequence data to generate structured video frame data and extracting a frame sequence of structured time-frequency frame data to obtain a video frame sequence; Step S13: using the video frame sequence to identify control signaling tags on the video frame input signal, generating a preliminary metadata segment set; performing field extraction on the metadata segment set, extracting spatial coordinate data, timestamp data, and operation type identifier, and generating an associated metadata package set; Step S14: synchronously mapping the video frame sequence and the associated metadata packet set to obtain a real-time media data stream.
3. The new media data transmission management method according to claim 2, characterized in that: In step S13, the control signaling mark recognition of the video frame input signal using the video frame sequence includes: Extracting frame header data of a video frame input signal using a video frame sequence to generate video frame synchronization index data; Divide the video frame synchronization index data into content segments to generate signaling candidate area data; Scan key bits of the signaling candidate area data to generate control signaling feature segment data; Performing operation identification matching on the control signaling feature segment data to generate initial operation type identification data; Perform frame sequence timing mapping processing on the initial operation type identification data to generate a preliminary metadata segment set.
4. The new media data transmission management method according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: analyzing the key frames of the video frame sequence and performing key frame screening to generate media video motion key frame data; continuously tracking the target position of the video motion key frame data to obtain trajectory point sequence data; Step S22: performing trajectory coherence fitting on the trajectory point sequence data to generate a video trajectory feature vector; Step S23: performing time stamp calculation on the metadata package set to generate synchronization timing parameter data; performing spatial position calculation on the spatial coordinates in the metadata package set to generate position mapping parameter data; Step S24: jointly associate the synchronization timing parameter data and the position mapping parameter data to generate a metadata spatiotemporal parameter set.
5. The new media data transmission management method according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: performing time series synchronization on the video trajectory feature vector and the metadata spatiotemporal parameter set to generate synchronized input feature data; Step S32: inputting the synchronous input feature data into a preset synchronous anomaly detection model to perform spatial coordinate error calculation and timestamp offset analysis to generate spatial offset data and time deviation value data; Step S33: performing position error threshold comparison on the spatial offset data to generate image position misalignment identification data; performing delay threshold comparison on the time deviation value data to generate time synchronization abnormality identification data; Step S34: performing media picture misalignment determination based on the picture position misalignment identification data and the time synchronization anomaly identification data, and generating a dynamic metadata misalignment determination result.
6. The new media data transmission management method according to claim 4, characterized in that: Step S34 includes the following steps: Step S341: performing spatial anomaly clustering on the image position misalignment identification data to generate spatial misalignment distribution data; Step S342: performing time series focusing on the time synchronization anomaly identification data to generate time misalignment aggregation data; Step S343: performing segmented co-frequency vibration superposition on the spatial misalignment distribution data and the temporal misalignment aggregation data to generate spatiotemporal coupling resonance feature data; performing dynamic recognition processing on the spatiotemporal coupling resonance feature data to generate media image misalignment status data; Step S344: Performing logic rule judgment on the media image misalignment status data to generate a dynamic metadata misalignment judgment result.
7. The new media data transmission management method according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: Performing status judgment on the dynamic metadata inaccuracy judgment result. If the dynamic metadata inaccuracy judgment result is true, performing cross-layer spatial structure analysis on the real-time media data stream to generate cross-layer collaborative mapping relationship data; Step S42: Compensate and distribute the real-time media data stream according to the cross-layer collaborative mapping relationship data to generate cross-layer compensation control instruction data; perform cross-layer collaborative compensation on the real-time media data stream to generate new media picture compensation data; Step S43: extracting the spatial alignment accuracy of the new media picture compensation data to generate picture spatial alignment accuracy data; adaptively adjusting the parameter weights of the picture spatial alignment accuracy data to generate dynamically optimized parameter weight data; Step S44: performing fusion correction processing on the new media picture compensation data and the dynamic optimization parameter weight data to generate a final compensated real-time media data stream to perform new media data transmission management optimization operations.
8. The new media data transmission management method according to claim 7, characterized in that: Step S42 includes the following steps: Step S421: performing key frame boundary region analysis on the real-time media data stream according to the cross-layer collaborative mapping relationship data to obtain key frame boundary regions; extracting the key frame boundary regions to perform edge structure extraction to obtain frame boundary feature data; Step S422: performing perturbation variation analysis on the metadata spatiotemporal parameter set using the frame boundary feature data to generate spatiotemporal perturbation marker data; performing inter-layer difference positioning on the frame boundary feature data and the spatiotemporal perturbation marker data to generate cross-layer error anchor data; Step S423: performing cross-layer boundary guidance compensation on the cross-layer error anchoring data to generate cross-layer compensation control instruction data; Step S424: superimpose and correct the cross-layer compensation control instruction data and the current real-time media data stream to generate new media picture compensation data; perform frame sequence consistency verification processing on the new media picture compensation data to generate new media picture compensation data.
9. The new media data transmission management method according to claim 8, characterized in that: Step S423 includes the following steps: Step S4231: performing boundary direction aggregation on the spatial offset information in the cross-layer error anchor data to generate boundary guidance vector data; performing inter-layer association path sorting on the boundary guidance vector data to generate hierarchical guidance path data; Step S4232: Identify the compensation signal injection point on the hierarchical guidance path data to generate cross-layer signal injection position data; Step S4233: Dynamically weighting the cross-layer signal injection position data and the spatial offset information to generate compensation response priority data; Step S4234: Perform instruction parameter parsing on the compensation response priority data to generate cross-layer compensation control instruction data.
10. A new media data transmission management system, characterized in that: For executing the new media data transmission management method according to claim 1, the new media data transmission management system comprises: A media stream acquisition module is used to obtain a real-time media data stream, which includes a video frame sequence and its associated metadata packet set, wherein the metadata packet set includes spatial coordinate data, timestamp data and an operation type identifier; A trajectory analysis module is used to extract motion trajectory features from the video frame sequence to generate a video trajectory feature vector, and simultaneously perform spatiotemporal parameter analysis on the metadata packet set to generate a metadata spatiotemporal parameter set; The image misalignment detection module is used to calculate the spatial offset and temporal deviation values based on the video trajectory feature vector and metadata spatiotemporal parameter set input into a preset synchronization anomaly detection model; the spatial offset and temporal deviation values are used to detect media image misalignment in the real-time media data stream, and a dynamic metadata misalignment judgment result is obtained; The cross-layer collaborative compensation module is used to perform cross-layer collaborative compensation on the real-time media data stream if the dynamic metadata inaccuracy judgment result is true, thereby generating new media picture compensation data; the picture space alignment accuracy of the compensated real-time media data stream is collected through the new media picture compensation data, and the parameter weights of the new media picture compensation data are dynamically adjusted to perform new media data transmission management optimization operations.