Media data management method and system based on hybrid cloud architecture, and storage medium
Through the media data management method under the hybrid cloud architecture, the storage hierarchy and migration strategy are optimized by using the popularity prediction and bandwidth prediction models, which solves the problems of unreasonable storage resource allocation and low migration efficiency in the existing technology, and realizes efficient media data management and differentiated services.
Patent Information
- Application Number
- CN202511032970.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-10-03
AI Technical Summary
Existing media data management technologies have problems in large-scale data processing, such as hot data allocation errors, network resource waste, inefficient migration, storage space waste, and low retrieval efficiency. They cannot meet the differentiated service needs of different types of media data.
Adopting a media data management method under a hybrid cloud architecture, through a hybrid cloud media data heat prediction network, a bidirectional inter-cloud bandwidth prediction model, and genetic algorithm optimization technology, we achieve refined hierarchical storage, dynamic migration, and resource scheduling of media data. Combined with distributed index construction and time-aware scheduling, we optimize storage tier allocation and cross-cloud migration strategies.
It improves the accuracy of media data access prediction, optimizes storage resource allocation, reduces network congestion, enables fast retrieval and differentiated services, and improves overall management efficiency and resource utilization.
Smart Images

Figure CN120751180A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a media data management method, system, and storage medium based on a hybrid cloud architecture. Background Art
[0002] Existing media data management technologies are mainly based on traditional centralized storage architectures, which allocate media files to different storage devices through static tiered storage strategies. Existing technologies usually use a simple classification method based on file access frequency to divide media data into two tiers: hot data and cold data. Hot data is stored in high-performance SSD devices, and cold data is stored in lower-cost mechanical hard drives. In terms of data migration, existing technologies mainly rely on preset migration strategies and fixed time scheduling, and execute data migration tasks within a specific time period through batch processing. For index management of media content, existing technologies use traditional indexing methods based on file paths, and locate and retrieve content through the directory structure and file names of the file system. In terms of resource scheduling, existing technologies usually use static load balancing strategies, which allocate computing and storage tasks to different server nodes according to pre-configured rules. They lack the ability to dynamically adapt to changes in media data access patterns and network conditions.
[0003] However, existing technologies have significant shortcomings when processing large-scale media data: First, traditional static tiered storage cannot accurately predict changes in the access popularity of media data, resulting in frequently accessed data being incorrectly allocated to low-performance storage, affecting the user access experience; second, existing data migration strategies lack the ability to perceive dynamic changes in network bandwidth, and still perform large-scale data migration when the network is congested, resulting in wasted network resources and low migration efficiency; third, file path-based indexing methods cannot effectively handle duplicate files with the same content but different file names, resulting in wasted storage space and reduced retrieval efficiency; finally, static resource scheduling strategies cannot intelligently allocate resources based on the time-sensitive characteristics of media data. Real-time streaming data and archived data are treated equally, and cannot meet the differentiated service requirements of different types of media data. Summary of the Invention
[0004] The present application provides a media data management method, system, and storage medium based on a hybrid cloud architecture, which are used to dynamically optimize the storage tier allocation and cross-cloud migration strategy of media data by intelligently predicting media data access popularity and network bandwidth changes in a hybrid cloud architecture environment, thereby solving the problems of low data migration efficiency and unreasonable storage resource allocation in the existing technology.
[0005] In a first aspect, the present application provides a media data management method based on a hybrid cloud architecture, the media data management method based on the hybrid cloud architecture comprising: Perform heat prediction on media data through the hybrid cloud media data heat prediction network to obtain media feature vectors and heat prediction values; Performing storage hierarchical division processing on the media data according to the media feature vector and the heat prediction value to obtain a four-level storage allocation scheme; Perform cross-cloud migration processing on the four-level storage allocation solution using a bidirectional inter-cloud bandwidth prediction model to obtain a data migration execution plan; Performing distributed index construction processing on the media data according to the data migration execution plan to obtain a media content index table; Based on the media content index table, time-series-aware scheduling processing is performed on the media data to obtain a dynamic resource allocation result.
[0006] Optionally, performing heat prediction processing on the media data by using the hybrid cloud media data heat prediction network to obtain the media feature vector and the heat prediction value includes: The media data is input into the 3D convolutional neural network for convolution operation processing to obtain raw feature data including video frame sequence features, audio spectrum features and metadata features; Inputting the original feature data into a long short-term memory network for time series analysis and processing to obtain a historical access pattern sequence and time dependency relationship data of the media data; Performing feature splicing processing on the original feature data and the time dependency relationship data to obtain a fused feature vector; Inputting the fused feature vector into a fully connected layer for weight calculation and activation function processing to obtain a 128-dimensional media feature vector; Softmax normalization is performed on the media feature vector to obtain a popularity prediction value representing the access probability distribution in the next 7 days.
[0007] Optionally, performing storage hierarchical division processing on the media data according to the media feature vector and the heat prediction value to obtain a four-level storage allocation scheme includes: Performing data sensitivity calculation based on the media feature vector to obtain a sensitivity score value ranging from 0 to 1; Calculating access delay requirements based on the predicted popularity value and the sensitivity score value to obtain a delay tolerance parameter for the media data; Performing a weighted sum operation on the heat prediction value, the sensitivity score value, the delay tolerance parameter, and the storage cost coefficient to obtain a storage tier score value; Comparing the storage tier score with four preset threshold intervals to obtain classification identifiers of the hot data layer, the warm data layer, the cold data layer, and the ice data layer; The storage location mapping process is performed on the media data according to the classification identifier to obtain a four-level storage allocation plan including a target storage level, a migration priority, and an expected migration time.
[0008] Optionally, performing cross-cloud migration processing on the four-level storage allocation solution using a bidirectional inter-cloud bandwidth prediction model to obtain a data migration execution plan includes: The current network throughput, delay time, packet loss rate and historical bandwidth data are input into the bidirectional LSTM network for sequence encoding processing to obtain the network status feature sequence; Performing forward and reverse time series analysis on the network status feature sequence to obtain a bandwidth availability prediction value and confidence interval within the next 2 hours; Performing genetic algorithm optimization processing based on the migration priorities and bandwidth availability prediction values in the four-level storage allocation scheme to obtain the optimal time scheduling sequence for cross-cloud data migration; The media data is divided into blocks based on the optimal time scheduling sequence to obtain a 256KB data block sequence and parallel transmission path allocation; The data block sequence and the parallel transmission path are allocated to perform migration time calculation processing to obtain a data migration execution plan including a migration start time, a transmission path, and a completion time.
[0009] Optionally, performing genetic algorithm optimization processing according to the migration priority and bandwidth availability prediction value in the four-level storage allocation scheme to obtain an optimal time scheduling sequence for cross-cloud data migration includes: Encoding and arranging the media data in the four-level storage allocation scheme according to the migration priority to obtain a chromosome sequence including the migration order of the hot data layer, the warm data layer, the cold data layer, and the ice data layer; Performing fitness function calculation on the chromosome sequence based on the bandwidth availability prediction value to obtain a fitness score that integrates migration time and storage cost; Performing a high-fitness individual selection process based on the fitness score to obtain excellent chromosome pairs for reproducing the next generation; Performing crossover recombination and site mutation processing on the media data migration sequence in the excellent chromosome pair to obtain a new migration sequence combination; The new migration sequence combination is screened and processed based on the fitness score to obtain the optimal time scheduling sequence for cross-cloud data migration of media data from the private cloud to each storage layer of the public cloud.
[0010] Optionally, performing distributed index construction processing on the media data according to the data migration execution plan to obtain a media content index table includes: Performing content feature hash calculation processing based on the media data in the data migration execution plan to obtain a 64-bit media content fingerprint including a visual feature hash value, an audio feature hash value, and a content signature; Input the media content fingerprint into the consistent hash ring for node positioning processing to obtain the distributed storage location of the media data in the private cloud and public cloud storage nodes; Encapsulating the access rights and metadata information of the media data according to the distributed storage location to obtain an index entry including a file hash value, a storage path, and access control information; The index entries are transmitted across clouds according to a synchronous update mechanism to obtain index copy data maintained in the private cloud and the public cloud respectively; Update log recording and version control processing are performed on the index copy data to obtain a media content index table that supports fast retrieval and duplicate file detection.
[0011] Optionally, performing time-series-aware scheduling processing on the media data based on the media content index table to obtain a dynamic resource allocation result includes: Performing access time characteristic analysis on the media data according to the storage location information in the media content index table to obtain a time sensitivity classification including real-time streaming media data and batch archived data; Based on the time sensitivity classification, the resource utilization of the private cloud and the public cloud is monitored and calculated to obtain the current load status and available resource capacity of each storage node; Prioritizing the time sensitivity classification and available resource capacity to obtain a scheduling strategy for allocating real-time streaming media data to private cloud nodes in a priority manner; Performing predictive analysis on the historical access patterns and service periodicity characteristics of media data according to the scheduling strategy to obtain a resource reservation plan and a data preloading schedule; Based on the resource reservation plan and the data preloading schedule, cross-cloud resource dynamic allocation processing is performed to obtain a dynamic resource allocation result including a resource allocation path, execution time and load balancing result.
[0012] In a second aspect, the present application provides a media data management system based on a hybrid cloud architecture, the media data management system based on the hybrid cloud architecture comprising: The prediction module is used to perform heat prediction processing on media data through the hybrid cloud media data heat prediction network to obtain media feature vectors and heat prediction values; a partitioning module, configured to perform storage hierarchical partitioning processing on the media data according to the media feature vector and the heat prediction value, and obtain a four-level storage allocation scheme; A migration module, configured to perform cross-cloud migration processing on the four-level storage allocation solution using a bidirectional inter-cloud bandwidth prediction model to obtain a data migration execution plan; A construction module, configured to perform distributed index construction processing on the media data according to the data migration execution plan to obtain a media content index table; The scheduling module is used to perform time-series-aware scheduling processing on the media data based on the media content index table to obtain a dynamic resource allocation result.
[0013] In a third aspect, a media data management device based on a hybrid cloud architecture is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the media data management device based on the hybrid cloud architecture executes the above-mentioned media data management method based on the hybrid cloud architecture.
[0014] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, which, when executed on a computer, enables the computer to execute the above-mentioned media data management method based on the hybrid cloud architecture.
[0015] In the technical solution provided by this application, through the technical features of the hybrid cloud media data heat prediction network, combined with the synergy of 3D convolutional neural networks and long short-term memory networks, it is possible to simultaneously extract the spatial and temporal features of media data, generate high-dimensional media feature vectors and accurate heat prediction values, and significantly enhance the accuracy of predicting media content access patterns compared to traditional simple statistical methods based on access frequency. The technical features of the four-level storage allocation scheme achieve refined hierarchical management of media data by comprehensively considering multi-dimensional factors such as heat prediction values, sensitivity scores, delay tolerance, and storage costs, avoiding the limitations of the simple dichotomy of hot data and cold data in the prior art. The technical features of the bidirectional inter-cloud bandwidth prediction model adopt a bidirectional LSTM network structure, which can simultaneously analyze the forward and backward time dependencies of the network status, accurately predict the bandwidth availability in the future time period, and combine the cross-cloud migration processing optimized by genetic algorithms to effectively solve the network congestion problem caused by improper selection of data migration timing in the prior art. The technical features of distributed index construction achieve fast retrieval and duplicate file detection based on content features by calculating the 64-bit fingerprint of media content and consistent hash ring positioning, overcoming the low retrieval efficiency and storage redundancy problems of traditional file path indexing methods.
[0016] The deep learning algorithm features of the hybrid cloud media data popularity prediction network can adaptively learn the access patterns of different types of media content, providing personalized storage optimization strategies for scenarios such as streaming platforms, online education, and corporate training. The timing analysis algorithm features of the bidirectional inter-cloud bandwidth prediction model fully utilize the periodic and bursty characteristics of network traffic. In network-sensitive applications such as video conferencing and live streaming services, it can intelligently select the optimal data migration timing to avoid resource conflicts during peak business hours. The technical features of time-aware scheduling processing achieve differentiated service quality assurance in multimedia content distribution networks by distinguishing the time sensitivity of real-time streaming data and batch archived data, ensuring that critical business data receives priority resource allocation. The technical features of genetic algorithm optimization demonstrate excellent global optimization capabilities when handling large-scale media data migration scheduling problems. It is particularly suitable for application scenarios such as media asset management and digital libraries that require frequent data reorganization. Through multi-generational evolution, it finds the optimal time scheduling sequence, significantly improving the overall efficiency and resource utilization of media data management in hybrid cloud environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 This is a schematic diagram of an embodiment of a media data management method based on a hybrid cloud architecture in an embodiment of the present application; Figure 2 This is a schematic diagram of an embodiment of a media data management system based on a hybrid cloud architecture in an embodiment of the present application; Figure 3 This is a schematic block diagram of the structure of a media data management device based on a hybrid cloud architecture in an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The embodiments of the present application provide a media data management method, system and storage medium based on a hybrid cloud architecture. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0020] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In one embodiment of the present application, a method for managing media data based on a hybrid cloud architecture includes: Step S101: Performing heat prediction processing on media data through the hybrid cloud media data heat prediction network to obtain media feature vectors and heat prediction values; Step S102: performing storage hierarchical division processing on the media data according to the media feature vector and the popularity prediction value to obtain a four-level storage allocation scheme; Step S103: Perform cross-cloud migration processing on the four-level storage allocation solution using a bidirectional inter-cloud bandwidth prediction model to obtain a data migration execution plan; Step S104: Perform distributed index construction processing on the media data according to the data migration execution plan to obtain a media content index table; Step S105: Perform time-series-aware scheduling processing on the media data based on the media content index table to obtain a dynamic resource allocation result.
[0021] It is understandable that the execution subject of this application can be a media data management system based on a hybrid cloud architecture, or a terminal or a server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.
[0022] Specifically, when processing data through the hybrid cloud media data popularity prediction network, the input media data is first fed into a 3D convolutional neural network. This network uses three-dimensional convolution kernels to extract spatial and temporal features from the continuous frame sequence of the video file. It also extracts spectral features of the audio signal and file metadata. The 3D convolutional neural network uses a sliding window mechanism to perform convolution operations on the width, height, and time dimensions of the video frames, generating raw feature data containing motion and texture information. This raw feature data is then fed into a long-short-term memory network (LSTM). This network uses a gating mechanism to control the inflow and outflow of information. The forget gate determines which historical information is discarded, the input gate determines which new information is retained, and the output gate determines which information is output. The LSTM network analyzes time-series data such as the media data's historical access frequency, user click behavior, and download counts to generate a historical access pattern sequence and temporal dependency data. The raw feature data and temporal dependency data are then subjected to feature concatenation. The spatial feature vector from the 3D convolutional neural network is concatenated with the temporal feature vector from the LSTM network, dimensionally, to form a fused feature vector.
[0023] During the storage tiering phase, data sensitivity scores are calculated based on media feature vectors. These features, such as access rights identifiers, user permission levels, and content classification tags, are analyzed and mapped to a numerical range of 0 to 1. Higher data sensitivity scores indicate a greater need for storage in a private cloud environment. Access latency requirements are calculated by weighting the predicted popularity value with the sensitivity score. The predicted popularity value reflects the frequency of media data access, while the sensitivity score reflects the security requirements of the data. These two factors are combined to generate a latency tolerance parameter, which determines the access speed requirements for the media data. When calculating the storage tier score, the predicted popularity value, sensitivity score, latency tolerance parameter, and storage cost coefficient are multiplied by their corresponding weighting factors and then added together. The weighting factors are dynamically adjusted based on business needs. The calculated result is then compared with four preset threshold ranges. Media data with a score between 0.8 and 1.0 is assigned to the hot data tier, between 0.6 and 0.8 to the warm data tier, between 0.3 and 0.6 to the cold data tier, and between 0 and 0.3 to the cold data tier.
[0024] During the cross-cloud migration phase, the bidirectional inter-cloud bandwidth prediction model processes network status data using a bidirectional LSTM network consisting of forward and backward LSTM units. The forward units process network throughput, latency, and packet loss rate data in chronological order, while the backward units process the same data in reverse chronological order. The results from both directions are combined to generate a network status feature sequence. Bandwidth availability predictions are derived by analyzing the periodic patterns and sudden fluctuations in the network status feature sequence. The predictions include available bandwidth values and corresponding confidence intervals for each time period within the next two hours. A genetic algorithm optimization process encodes media data in the four-level storage allocation scheme into a chromosome sequence based on migration priority. Each gene bit represents a migration time node for a media file, and the chromosome sequence represents the overall migration schedule. A fitness function calculates a fitness score for each chromosome based on the bandwidth availability predictions. The score takes into account migration time and storage cost. Chromosomes with higher fitness scores represent more optimal migration plans.
[0025] The distributed index construction process achieves fast retrieval by calculating media content fingerprints. Visual feature hash values are generated by extracting the color histogram and edge features of video keyframes. Audio feature hash values are generated by analyzing the Mel-frequency cepstral coefficients of the audio signal. Content signatures are generated by combining metadata such as file size, duration, and encoding format. These three hash values are combined into a 64-bit media content fingerprint through an XOR operation. The consistent hash ring maps media content fingerprints to specific locations in a ring space and determines the distribution of media data in private and public cloud storage nodes based on the hash values. The hash ring design ensures uniform data distribution and minimizes data migration when nodes change. Index entries encapsulate file hash values, storage paths, access control information, and metadata information. The cross-cloud synchronization update mechanism ensures consistency of index data in private and public clouds by comparing version timestamps.
[0026] Time-aware scheduling processing analyzes the access time characteristics of media data based on the storage location information in the media content index table. Real-time streaming data has low latency requirements and high concurrent access characteristics, while batch archived data has high latency tolerance and low access frequency characteristics. Resource utilization monitoring calculates the current load status and remaining available resource capacity of each node by collecting CPU usage, memory occupancy, and network bandwidth occupancy data of each storage node. Priority matching processing matches the time-sensitivity classification results with the available resource capacity. Real-time streaming data is preferentially allocated to the high-performance storage nodes of the private cloud, and batch archived data is allocated to the low-cost storage nodes of the public cloud. The resource reservation plan predicts future peak resource demand periods based on the periodic laws in historical access patterns. The data preloading schedule arranges the migration of hot media data from the public cloud to the private cloud cache before the access peak.
[0027] In a specific embodiment, the process of executing step S101 may specifically include the following steps: The media data is input into the 3D convolutional neural network for convolution operation processing to obtain raw feature data including video frame sequence features, audio spectrum features and metadata features; The original feature data is input into the long short-term memory network for time series analysis and processing to obtain the historical access pattern sequence and time dependency relationship data of the media data; Perform feature splicing processing on the original feature data and time dependency data to obtain the fused feature vector; The fused feature vector is input into the fully connected layer for weight calculation and activation function processing to obtain a 128-dimensional media feature vector; Softmax normalization is performed on the media feature vector to obtain the popularity prediction value representing the access probability distribution in the next 7 days.
[0028] Specifically, the media data is input into a 3D convolutional neural network for feature extraction. The 3D convolutional neural network is a deep learning network structure that specializes in processing three-dimensional data. Its convolution kernel slides across the width, height, and time dimensions of the video data. Unlike traditional two-dimensional convolution, the 3D convolution kernel has a depth dimension and can capture both spatial and temporal features. During the convolution operation, the network analyzes the input video file as a frame sequence. Each convolution kernel slides across consecutive video frames to detect motion patterns and texture changes. At the same time, the audio signal is analyzed in the frequency domain, and the time domain audio signal is converted into spectral data through a fast Fourier transform. The frequency distribution characteristics and energy distribution characteristics of the audio are extracted. The network also reads the metadata information of the media file, including technical parameters such as file format, encoding parameters, resolution, and frame rate, and integrates these different types of feature data into the original feature data.
[0029] The raw feature data is then fed into a long-short-term memory (LSTM) network for time series analysis. The LSTM network is a specialized recurrent neural network structure specifically designed to process sequential data and address long-term dependencies. The network incorporates three key gating mechanisms. The forget gate uses a sigmoid activation function to determine which information to discard from the cell state. The input gate also uses a sigmoid function to determine which new information to store in the cell state. The output gate controls which parts of the cell state are output. During the time series analysis process, the network receives historical access records for media data, including time series data such as daily visits, user dwell time, and download frequency. The gating mechanism filters and retains important historical access pattern information, identifying periodic access patterns and sudden access events in the data. This generates a historical access pattern sequence and analyzes the correlations between different time points to form temporal dependency data. This data reflects the changing popularity of media content over time. The feature splicing processing stage merges the original feature data from the 3D convolutional neural network with the time dependency data from the long short-term memory network in terms of dimensions. The splicing processing uses vector connection to arrange and combine the spatial feature vectors and the temporal feature vectors in a specific order to form a higher-dimensional fusion feature vector. This vector also contains the visual features, auditory features, metadata features, and historical access pattern features of the media content. The splicing process needs to ensure that the feature data from different sources maintains consistency in numerical range and dimension, and the various feature values are mapped to the same numerical range through normalization.
[0030] The fused feature vector is then input into the fully connected layer for weight calculation and activation function processing. The fully connected layer is a basic layer structure in the neural network. Each neuron in the layer is connected to all neurons in the previous layer. The weight calculation process multiplies the input feature vector by the weight matrix through matrix multiplication, and then adds the bias term to obtain the result after linear transformation. The activation function processing uses ReLU or other nonlinear functions to perform nonlinear mapping on the linear transformation result to enhance the network's expressive power. After layer-by-layer processing of multiple fully connected layers, the dimension of the feature vector is gradually compressed and refined, and finally a 128-dimensional media feature vector is output. This vector is a highly abstract and compressed representation of the original media data.
[0031] Softmax normalization converts the 128-dimensional media feature vector into a popularity prediction value in the form of a probability distribution. The Softmax function is a commonly used probabilistic normalization function. Its calculation process uses each element in the feature vector as the input of an exponential function, calculates the exponential value, and then divides it by the sum of all exponential values. The result is a probability distribution where the sum of all elements equals 1, and each element represents the probability of the media data being accessed within the corresponding time period. Normalization maps the media feature vector into an access probability distribution for the next seven days. The first element represents the access probability for the first day, the second element for the second day, and so on. The calculation of the popularity prediction value takes into account the historical access patterns, content characteristics, and temporal periodicity of the media content.
[0032] In a specific embodiment, the process of executing step S102 may specifically include the following steps: Calculate the data sensitivity based on the media feature vector to obtain a sensitivity score value ranging from 0 to 1; Calculate the access delay requirement based on the popularity prediction value and sensitivity score value to obtain the delay tolerance parameter of the media data; Perform a weighted sum operation on the heat prediction value, sensitivity score value, delay tolerance parameter, and storage cost coefficient to obtain the storage tier score value; Compare the storage tier score with four preset thresholds to obtain the classification identification of the hot data layer, warm data layer, cold data layer, and ice data layer; The storage location mapping process of the media data is performed according to the classification identifier to obtain a four-level storage allocation plan including the target storage level, migration priority and expected migration time.
[0033] Specifically, based on the analysis of 128-dimensional media feature vectors, data sensitivity refers to the degree to which media content requires security for the storage environment. The calculation process identifies sensitivity indicators by parsing specific dimensions in the media feature vectors. Different dimensions in the feature vector correspond to different media attributes, including content classification identification, user permission requirements, access control levels and other information. The calculation process determines the sensitivity level by looking up the dimension values related to access rights in the feature vector and combining them with the content label information. The sensitivity score is calculated using a weighted average method, which assigns different weights to factors such as access rights level, content classification level, and user group permissions. The results are normalized after weighted summation and mapped to a numerical range of 0 to 1, where 0 represents completely public content, 1 represents highly confidential content, and intermediate values represent different degrees of sensitivity requirements.
[0034] The access delay requirement calculation process correlates the heat prediction value with the sensitivity score value. The heat prediction value reflects how frequently the media data is accessed, and the sensitivity score value reflects the security requirements of the data. The combination of the two determines the access speed requirements of the media data. The calculation process uses an inverse proportional relationship. When the heat prediction value is high, users expect faster access responses, so the delay tolerance is low. When the sensitivity score is high, the data needs to be stored in a secure private cloud environment, and security must be guaranteed even if the delay is slightly higher. The delay tolerance parameter is calculated by taking the inverse of the heat prediction value and multiplying it with the sensitivity score value. The value is then corrected by an adjustment factor to obtain a numerical parameter that reflects the media data's tolerance for access delay. The higher the parameter value, the greater the data's tolerance for delay, and the lower the value, the stricter the delay requirement.
[0035] A weighted summation operation combines four key factors: the predicted popularity value, the sensitivity score, the latency tolerance parameter, and the storage cost coefficient. The storage cost coefficient is a pre-set parameter that reflects the cost differences between different storage tiers. Private cloud storage is more expensive but more secure, while public cloud storage is less expensive but less secure. The weighting of the weighted summation operation is dynamically adjusted based on the business needs of the hybrid cloud architecture. Generally, the predicted popularity value is given a higher weight because access frequency directly affects the user experience. The sensitivity score is given a lower weight because security is a fundamental requirement. Latency tolerance and storage cost are relatively lower weights. The operation multiplies each of the four parameters by the corresponding weight coefficient and then adds the results to obtain a comprehensive storage tier score, which comprehensively reflects the storage requirements of media data in the hybrid cloud environment.
[0036] The numerical comparison process compares and analyzes the storage tier score against four preset threshold intervals. These threshold intervals correspond to the classification criteria for the hot, warm, cold, and ice data tiers. These threshold intervals are fixed numerical ranges pre-set based on the storage characteristics and business requirements of the hybrid cloud architecture. The hot data tier corresponds to a score range of 0.8 to 1.0, which is suitable for frequently accessed media data with strict latency requirements. The warm data tier corresponds to a score range of 0.6 to 0.8, which is suitable for moderately accessed media data. The cold data tier corresponds to a score range of 0.3 to 0.6, which is suitable for infrequently accessed media data. The ice data tier corresponds to a score range of 0 to 0.3, which is suitable for rarely accessed archived media data. The comparison process determines the storage tier to which the media data should be assigned by determining the range within which the storage tier score falls. The comparison result generates a corresponding classification identifier, which contains the tier name and corresponding storage environment information.
[0037] The storage location mapping process allocates media data to specific storage locations based on classification identifiers. The mapping process needs to consider the distribution of storage resources in the private cloud and public cloud in the hybrid cloud architecture. The hot data tier is mapped to the private cloud's high-performance SSD storage array, the warm data tier is mapped to the private cloud's mechanical hard disk storage area, the cold data tier is mapped to the public cloud's standard storage service, and the ice data tier is mapped to the public cloud's archive storage service. The mapping process also needs to determine the migration priority, which is sorted according to the storage tier score value. The higher the score, the higher the migration priority of the media data, and the more urgent the migration schedule. The calculation of the expected migration time takes into account factors such as the current storage load, network bandwidth conditions, and data file size. The expected time point for migration completion is determined by evaluating the time required for data transmission.
[0038] In a specific embodiment, the process of executing step S103 may specifically include the following steps: The current network throughput, delay time, packet loss rate and historical bandwidth data are input into the bidirectional LSTM network for sequence encoding processing to obtain the network status feature sequence; Perform forward and reverse time series analysis on the network status feature sequence to obtain the bandwidth availability prediction value and confidence interval within the next 2 hours; Based on the migration priorities and bandwidth availability predictions in the four-level storage allocation scheme, a genetic algorithm optimization process is performed to obtain the optimal time scheduling sequence for cross-cloud data migration. Based on the optimal time scheduling sequence, media data is divided into blocks and processed to obtain a 256KB data block sequence and parallel transmission path allocation; The data block sequence and the parallel transmission path are allocated to perform migration time calculation processing to obtain a data migration execution plan including the migration start time, transmission path and completion time.
[0039] Specifically, the data processing process of the bidirectional inter-cloud bandwidth prediction model first inputs the current network throughput, delay time, packet loss rate and historical bandwidth data into the bidirectional LSTM network for sequence encoding. Network throughput refers to the amount of data transmitted through the network connection per unit time, usually measured in Mbps. Delay time refers to the time required for a data packet to travel from the sender to the receiver, usually measured in milliseconds. Packet loss rate refers to the proportion of data packets lost during transmission to the total number of data packets. Historical bandwidth data contains time series records of network bandwidth usage over a period of time. The bidirectional LSTM network is a special recurrent neural network structure that contains two processing directions: forward LSTM units and backward LSTM units. The forward unit processes the input sequence in chronological order from the past to the present, and the backward unit processes the same input sequence in reverse chronological order from the present to the past. The sequence encoding process forms an input sequence by arranging the network performance parameters in chronological order. The forward LSTM unit gradually processes the network status data at each time point through a gating mechanism to extract the forward dependency in the time series. The backward LSTM unit starts reverse processing from the end of the sequence to extract the backward dependency in the time series. The processing results of the two directions are fused by splicing or summing to generate a network status feature sequence containing bidirectional time dependency information. The forward and reverse time series analysis of network status feature sequences utilizes the output of a bidirectional LSTM network for in-depth analysis. Forward time series analysis, based on the hidden state sequence of the forward LSTM unit, identifies trends and cyclical patterns in network bandwidth changes. The analysis predicts future bandwidth changes by detecting upward, downward, and plateauing trends in the feature sequence. Reverse time series analysis, based on the hidden state sequence of the backward LSTM unit, traces back from the current time point to identify historical factors and unexpected events that have impacted the current network status. The time series analysis combines the forward and reverse analysis results and generates a bandwidth availability forecast for the next two hours through a weighted fusion mechanism. The forecast represents the available bandwidth capacity of the network connection within each time period. The confidence interval, derived by calculating the statistical distribution of the prediction error, reflects the reliability of the prediction result. Narrower confidence intervals indicate more accurate predictions, while wider confidence intervals indicate greater uncertainty.
[0040] The genetic algorithm optimization process comprehensively analyzes the migration priorities and bandwidth availability predictions within the four-level storage allocation scheme. Migration priorities are ranked based on the importance and urgency of media data, while bandwidth availability predictions provide information about network transmission capacity at different time periods. A genetic algorithm, an optimization algorithm that mimics biological evolution, searches for the optimal solution through selection, crossover, and mutation. The optimization process first encodes media data migration tasks into a chromosome sequence, where each gene bit represents the migration schedule for a media file. The chromosome length is equal to the number of media files to be migrated. An initial population is formed by randomly generating multiple different migration schedules. The fitness function comprehensively considers migration priority and bandwidth availability. High-priority media data should be migrated during time periods with sufficient bandwidth. Fitness is calculated by multiplying the migration priority by the bandwidth availability in the corresponding time period. Higher fitness values indicate better migration plans. The selection operation selects individuals with good fitness as parents based on their fitness values. The crossover operation swaps gene bits between parent individuals to generate new individuals. The mutation operation randomly changes the values of certain gene bits within individuals to increase population diversity. After multiple generations of evolution, the optimal time schedule for cross-cloud data migration is obtained.
[0041] Chunking segments media data files based on an optimal time-scheduled sequence. Chunking refers to the technique of breaking large files into multiple smaller blocks for parallel transmission. Each data block is set to 256KB in size, a balance between network transmission efficiency and error recovery costs. Blocks that are too small increase transmission overhead, while blocks that are too large incur high retransmission costs in the event of transmission errors. Chunking segments media data into consecutive 256KB blocks according to the file's byte order. The size of the last block is determined by the number of remaining bytes in the file. Each block is assigned a sequence number and a checksum. The sequence number is used for file reassembly at the receiving end, while the checksum is used to detect transmission errors. Parallel transmission path allocation determines the transmission path for each data block based on the network topology and bandwidth distribution. Path allocation utilizes a load balancing strategy, assigning data blocks to different network links for simultaneous transmission. The allocation process considers each path's bandwidth capacity, latency characteristics, and reliability, prioritizing the allocation of important data blocks to high-quality transmission paths.
[0042] The migration time calculation process comprehensively analyzes the data block sequence and parallel transmission path allocation information to calculate the transmission time of each data block and the overall migration completion time. The transmission time of a single data block is obtained by dividing the data block size by the available bandwidth of the allocation path. The transmission time calculation also needs to consider factors such as network delay, protocol overhead, and retransmission time. The migration start time is determined according to the arrangement in the optimal time scheduling sequence to ensure that the migration task is started during a time period with sufficient network bandwidth. The transmission path information records the specific network connection and relay nodes used by each data block. The completion time is determined by the maximum value of the transmission time of all data blocks, because the slowest data block in parallel transmission determines the overall completion time. The data migration execution plan integrates key information such as migration start time, transmission path, and completion time to form a detailed migration task schedule. The plan also includes error handling strategies and progress monitoring mechanisms to ensure the controllability and traceability of the migration process.
[0043] In a specific embodiment, the process of performing the genetic algorithm optimization process according to the migration priority and bandwidth availability prediction value in the four-level storage allocation scheme may specifically include the following steps: The media data in the four-level storage allocation scheme is encoded and arranged according to the migration priority to obtain a chromosome sequence including the migration order of the hot data layer, the warm data layer, the cold data layer and the ice data layer; The fitness function of the chromosome sequence is calculated based on the bandwidth availability prediction value to obtain a fitness score that combines migration time and storage cost. According to the fitness score, high-fitness individuals are selected to obtain excellent chromosome pairs for reproducing the next generation; Perform crossover recombination and site mutation processing on the media data migration sequence in the excellent chromosome pair to obtain a new migration sequence combination; The new migration sequence combinations are screened according to the fitness score to obtain the optimal time scheduling sequence for cross-cloud data migration of media data from private cloud to public cloud storage layers.
[0044] Specifically, the encoding permutation process optimized by the genetic algorithm serializes and encodes the media data in the four-level storage allocation scheme according to migration priority. The encoding process uses integer encoding, assigning each media file a unique number. The order of the numbers reflects the order of migration. Media files in the hot data layer, due to their high access frequency and sensitivity to latency, are given the highest migration priority and occupy the front position in the chromosome sequence. Media files in the warm data layer have medium priority and are ranked after the hot data in the sequence. Media files in the cold data layer and the ice data layer have decreasing priority, occupying the middle and end positions in the sequence, respectively. The chromosome sequence is the data structure used to represent the solution in the genetic algorithm. Each chromosome represents a time schedule for media data migration. Each gene bit in the sequence corresponds to a media file, and the value of the gene bit represents the migration time period of the file. The length of the sequence is equal to the total number of media files to be migrated. The permutation process ensures that media files within the same storage tier are sorted according to factors such as file size and access frequency, forming a complete chromosome sequence that contains the migration order of the four tiers.
[0045] The fitness function calculation process evaluates each chromosome sequence based on the predicted bandwidth availability. The fitness function is a mathematical function used in genetic algorithms to evaluate the quality of solutions. The calculation comprehensively considers two key factors: migration time and storage cost. Migration time is calculated by dividing the size of each media file by the predicted bandwidth availability for the corresponding time period. The result reflects the length of time required to complete the file migration. Storage cost is calculated by weighting the price differences and data storage capacity of different storage tiers. Private cloud storage is more expensive but offers better performance and security, while public cloud storage is cheaper but compromises on access speed. The fitness score is calculated by weighting the inverse of the migration time and the inverse of the storage cost. The weighting coefficient is adjusted based on business needs. Solutions with shorter migration times receive higher time scores, while solutions with lower storage costs receive higher cost scores. The weighted sum of these two scores constitutes the chromosome's fitness score, with higher scores indicating a better migration solution.
[0046] The high-fitness individual selection process uses a roulette wheel selection mechanism to select excellent chromosomes from the current population for reproduction. Roulette wheel selection is a common selection strategy in genetic algorithms. The selection probability is proportional to the individual's fitness score; individuals with higher fitness have a greater probability of being selected. The selection process first calculates the sum of the fitness scores of all individuals in the population. Then, the selection probability of each individual is calculated. The selection probability is equal to the individual's fitness score divided by the sum. A random number is then generated and compared with the cumulative probability to determine the selected individual. The selection operation is repeated until a sufficient number of parent individuals are obtained. Excellent chromosome pairs are formed by randomly pairing the selected individuals. Each pair of two individuals is used in the subsequent crossover operation. The pairing process ensures that each individual has the opportunity to exchange genes with other excellent individuals. The resulting chromosome pairs contain the best gene combinations in the current population.
[0047] Crossover recombination and site mutation are the core operations for generating new individuals in genetic algorithms. Crossover simulates the genetic recombination phenomenon during biological reproduction. This process randomly selects a crossover point between pairs of excellent chromosomes and swaps the gene segments before and after the crossover point to generate new offspring individuals. The crossover operation uses a single-point crossover method, randomly selecting a location in the chromosome sequence as the crossover point. The gene segment before the crossover point of the parent individual is combined with the gene segment after the crossover point of the mother individual. Simultaneously, the gene segment before the crossover point of the mother individual is combined with the gene segment after the crossover point of the parent individual to form two new offspring individuals. Site mutation randomly changes the values of certain gene positions in the offspring individuals with a small probability. Mutation replaces the original migration schedule with an alternative time period by randomly selecting gene positions in the chromosome sequence. The mutation probability is typically set low to avoid disrupting the combination of excellent genes. Mutation increases the genetic diversity of the population and prevents the algorithm from prematurely converging to a local optimum.
[0048] The survival of the fittest screening process compares the fitness of newly generated migration sequence combinations with that of the original population. This screening process employs an elite retention strategy combined with a fitness ranking mechanism. Elite retention ensures that the individuals with the highest fitness in the current population advance directly to the next generation, preventing the loss of high-performing genes during evolution. Fitness ranking ranks all individuals from high to low according to their fitness scores, selecting the top-ranked individuals to form the new population. The screening process sets a fixed population size; individuals exceeding this size are eliminated, and the remaining individuals form the evolutionary foundation for the next generation. This screening process is repeated for multiple generations until fitness converges or a preset number of iterations is reached. Convergence is determined by monitoring the change in fitness of the best individuals over several consecutive generations. Convergence is considered achieved when the change is below a set threshold. The optimal time schedule for cross-cloud data migration is determined by the best individuals retained. This sequence details the specific time schedule for migrating each media file from the private cloud to each storage tier in the public cloud. This sequence takes into account the temporal variations in network bandwidth and the cost differences between different storage tiers, ensuring that the migration process meets performance requirements while minimizing overall cost.
[0049] In a specific embodiment, the process of executing step S104 may specifically include the following steps: Perform content feature hash calculation based on the media data in the data migration execution plan to obtain a 64-bit media content fingerprint including visual feature hash values, audio feature hash values, and content signatures; Input the media content fingerprint into the consistent hash ring for node positioning processing to obtain the distributed storage location of the media data in the private cloud and public cloud storage nodes; Encapsulate the access rights and metadata information of media data according to the distributed storage location to obtain an index entry containing the file hash value, storage path and access control information; The index entries are transferred across clouds according to the synchronous update mechanism to obtain index copy data maintained in the private cloud and the public cloud respectively; Update log records and version control are performed on the index copy data to obtain a media content index table that supports fast retrieval and duplicate file detection.
[0050] Specifically, multi-dimensional feature extraction is performed based on the media data in the data migration execution plan. Hash calculation is a mathematical function that converts input data of arbitrary length into an output value of fixed length. Media content fingerprints generate unique identifiers by combining multiple hash algorithms. The calculation of visual feature hash values first extracts key frames from the video file. Key frames are representative static images in the video. The extraction process selects frames according to fixed time intervals or scene change thresholds. Then, each key frame is subjected to color histogram analysis and edge detection. The color histogram records the distribution of various colors in the image, and edge detection identifies the outline and shape features of objects in the image. These visual feature data are input into the perceptual hash algorithm for compression encoding. The perceptual hash algorithm can generate feature codes that produce similar hash values for images with similar content but different formats. The calculation of audio feature hash values involves analyzing the audio signal in the media file to extract the audio's spectral and time-domain features. Spectral features are converted into frequency-domain representations using a Fast Fourier Transform (FFT), identifying the primary frequency components and energy distribution within the audio. Time-domain features include parameters such as amplitude variation, zero-crossing rate, and short-term energy, reflecting the temporal characteristics and dynamic changes of the audio signal. The audio feature data is encoded using a locality-sensitive hashing algorithm to generate a hash value that reflects the similarity of the audio content. Content signatures are generated based on the metadata of the media file, including technical parameters such as file size, creation time, modification time, encoding format, resolution, and frame rate. This metadata is concatenated according to a predefined format, and a fixed-length content signature is generated using cryptographic hash functions such as MD5 or SHA-256. Finally, the visual feature hash value, audio feature hash value, and content signature are combined using an exclusive-OR operation to generate a 64-bit media content fingerprint.
[0051] The consistent hashing ring node location process maps media content fingerprints to specific locations in a circular address space. A consistent hashing ring is a commonly used data distribution algorithm in distributed systems. The ring structure represents the hash space as a connected ring, with each point on the ring corresponding to a hash value. Node location first uses a hash function to map storage nodes in the private and public clouds to different locations on the ring. Each storage node calculates a hash value based on its identifying information, such as its IP address and port number, to determine its location on the ring. The media content fingerprint is then mapped to a specific location on the ring using the same hash function. The location process begins at the location of the media content fingerprint on the ring and searches for the nearest storage node in a clockwise direction. This node becomes the primary storage location for the media data. To enhance data reliability, the location algorithm also selects one or more nodes in the next clockwise direction as backup storage locations. The determination of distributed storage locations takes into account load balancing and fault tolerance. If a storage node fails or is overloaded, the consistent hashing ring can redistribute data to other nodes, minimizing the scope of data migration. The location results contain detailed information about the distribution of media data across storage nodes in the private and public clouds.
[0052] Encapsulation structures information related to media data based on distributed storage locations. Encapsulation refers to the process of packaging and organizing different types of data into a predefined format, creating a standardized data structure. Encapsulation of access rights information includes security-related data such as user authentication, role-based access control, and access policies. User authentication records which users or user groups have access to the media file. Role-based access control defines the file operation permissions of different roles, such as read-only, read-write, and delete. Access policies specify access conditions such as time limits, IP address restrictions, and device restrictions. Encapsulation of metadata information integrates the technical and business attributes of the media file. Technical attributes include information such as file format, encoding parameters, and quality level. Business attributes include content management information such as file title, description, tags, and classification. The encapsulation process stores this information in a structured format using standard formats such as JSON or XML. Index entry generation combines the file hash value, storage path, and access control information to form an index record. The file hash value serves as the index's primary key, uniquely identifying the media file. The storage path records the file's specific address in the distributed storage location. Access control information provides the basis for security verification. Index entries use a fixed data structure format to facilitate subsequent query and management operations.
[0053] Cross-cloud transmission uses a synchronous update mechanism to ensure consistency of index data between the private and public clouds. Synchronous updates are a technical mechanism that automatically propagates changes to the index data on one end. Transmission begins by establishing a secure communication channel between the private and public clouds. This channel uses the TLS encryption protocol to protect data transmission security and prevent interception or tampering of index information during transmission. Then, an incremental synchronization strategy is employed to reduce network transmission overhead. Incremental synchronization only transmits changed index entries, rather than the entire index data. Change detection is achieved by comparing index entry version numbers or timestamps. When an index entry is updated, the transmission mechanism sends the updated entry to the peer for synchronization. The transmission process utilizes message queues to ensure data reliability. Message queues cache pending index update requests, maintaining data integrity during network outages or temporary peer unavailability. Index replica data is maintained by establishing separate index storage areas in the private and public clouds, each of which maintains media content index information. Synchronization of replica data utilizes an eventual consistency strategy, allowing for temporary data inconsistencies while ensuring eventual convergence to a consistent state.
[0054] Update logging and version control processes track and manage changes to index replica data in detail. The update log is a mechanism for recording the history of index data changes. Every creation, modification, or deletion of an index entry is recorded in the log. Log records contain information such as the operation time, operation type, operation object, and operation result. The operation time uses a high-precision timestamp to ensure chronological accuracy. The operation type distinguishes between addition, modification, and deletion operations. The operation object records the identifier of the index entry being operated on, and the operation result records whether the operation was successful. Logging uses an append-only write method, allowing only new records to be added to the end of the log without modifying the previous record, ensuring log integrity and traceability. Version control assigns a version number to each index entry. The version number is an ascending numeric sequence and is automatically incremented with each change to the index entry. Version control supports rollback of index data. When an index error is discovered or a restoration to a previous state is required, the version number can be used to quickly locate and restore to a specific version. Version information is stored alongside the index entry, forming a complete record of historical changes. The construction of the media content index table integrates the functions of fast retrieval and duplicate file detection. Fast retrieval supports file search operations with O(1) time complexity by establishing an index structure based on hash values. Duplicate file detection identifies duplicate files with the same content but different file names by comparing media content fingerprints. The detection results help optimize storage space usage and reduce redundant data.
[0055] In a specific embodiment, the process of executing step S105 may specifically include the following steps: Analyze and process the access time characteristics of media data based on the storage location information in the media content index table to obtain a time sensitivity classification including real-time streaming data and batch archived data; Monitor and calculate the resource utilization of private and public clouds based on time sensitivity classification to obtain the current load status and available resource capacity of each storage node; Prioritize the time sensitivity classification and available resource capacity to obtain a scheduling strategy that prioritizes the allocation of real-time streaming data to private cloud nodes. Perform predictive analysis on the historical access patterns and business periodicity characteristics of media data based on the scheduling strategy to obtain resource reservation plans and data preloading schedules; Based on the resource reservation plan and data preloading schedule, cross-cloud resource dynamic allocation processing is performed to obtain dynamic resource allocation results including resource allocation path, execution time and load balancing results.
[0056] Specifically, the access time characteristics analysis of time-aware scheduling performs in-depth analysis of media data based on the storage location information in the media content index table. Access time characteristics refer to the access patterns and response requirements of media data over different time periods. The analysis process identifies the time sensitivity level of the data by parsing historical access logs, user behavior data, and business scenario identifiers recorded in the index table. Real-time streaming media data refers to media content that requires continuous transmission and is extremely sensitive to latency, including live video streams, online meeting recordings, and real-time surveillance video. This type of data is characterized by high access frequency, strict transmission latency requirements, and users expecting instant responses. The analysis process identifies real-time streaming data by examining the access frequency field, latency tolerance parameters, and user priority identifiers in the index table. Batch archived data refers to historical media content with lower access frequency and less stringent latency requirements, including backup files, historical records, and long-term archived videos. This type of data is characterized by longer access intervals, a tolerance for higher access latency, and is primarily used for storage and backup purposes. The analysis process identifies batch archived data by counting the last access time, access interval, and data importance level in the index table. The generation of time sensitivity classification adopts a multi-dimensional evaluation mechanism, comprehensively considering factors such as data access frequency, delay tolerance, business importance and user level. The classification algorithm performs weighted calculations on these factors to generate a time sensitivity score. According to the scoring threshold, media data is divided into three levels of high sensitivity, medium sensitivity and low sensitivity. High sensitivity corresponds to real-time streaming data, low sensitivity corresponds to batch archived data, and medium sensitivity corresponds to regular media data in between.
[0057] Resource utilization monitoring and calculation processing performs real-time monitoring of storage nodes in private and public clouds based on time-sensitivity classification. Resource utilization refers to the ratio of currently used resources of a storage node to total resource capacity. The monitoring scope includes key performance indicators such as CPU usage, memory occupancy, disk I / O throughput, and network bandwidth occupancy. Monitoring and calculation regularly collect resource usage data through monitoring agents deployed on each storage node. The collection interval is usually set to seconds or minutes to ensure the real-time nature of the data. The collected raw data includes instantaneous values and cumulative values. Instantaneous values reflect the resource status at the current moment, and cumulative values reflect resource usage trends over a period of time. The current load status is calculated by performing sliding window averaging on the collected data to eliminate the impact of instantaneous fluctuations on the evaluation results. The size of the sliding window is adjusted according to business needs and data characteristics, and is generally set to a time window of 5 to 15 minutes. The calculation results generate a status report reflecting the real-time load level of each storage node. Available resource capacity is calculated by subtracting current usage from total resource capacity. The calculation process also needs to consider the resource reservation mechanism, reserving a certain proportion of resource capacity for critical businesses to ensure sufficient resource buffer in emergency situations. Available resource capacity data is updated in real time and synchronized to the resource scheduling decision center.
[0058] Priority matching correlates the time-sensitivity classification results with available resource capacity information. The matching algorithm employs a multi-level priority strategy. First, a baseline priority is determined based on the time-sensitivity level, with high-sensitivity data receiving the highest priority and low-sensitivity data receiving the lowest. Priorities are then adjusted based on resource capacity. When resources are sufficient, the original priority is maintained; when resources are limited, the priority of some data is appropriately lowered. The scheduling strategy prioritizes real-time streaming data to private cloud nodes based on both security and performance considerations. Private cloud nodes typically offer higher security levels and more stable network connections, meeting the stringent requirements of real-time streaming data. The scheduling strategy adopts the principle of proximity, prioritizing private cloud nodes with the lowest network latency. When private cloud node resources are insufficient, the scheduling strategy selects public cloud nodes with better performance as a backup. The matching results generate a detailed scheduling strategy table, which records the target storage environment, backup options, and switching conditions for each type of media data. The scheduling strategy supports dynamic adjustment, automatically updating allocation decisions based on changes in resource availability.
[0059] Predictive analysis and processing deeply mines historical access patterns and service cyclical characteristics of media data based on scheduling strategies. Historical access pattern analysis identifies regular patterns in data access by analyzing historical data access frequency, access time distribution, and user access behavior over a period of time. The analysis utilizes time series analysis to chronologically arrange historical access data. Key features of access patterns are extracted through trend analysis, cycle detection, and anomaly identification. Service cyclical characteristics are identified based on cyclical patterns within business scenarios, such as differences in access between weekdays and weekends, load variations between daytime and nighttime, and seasonal service peaks. The identification process decomposes historical data into cyclical components at different levels, such as daily, weekly, and monthly cycles, and analyzes access intensity and resource demand fluctuations within each cycle. Resource reservation plans are developed based on the results of predictive analysis, anticipating peaks and troughs in resource demand over the next period. Sufficient resource capacity is reserved before peak periods arrive, releasing excess resources for other tasks during trough periods. The reservation plan utilizes a dynamic adjustment mechanism, making real-time adjustments based on deviations between actual access and the predicted results. The data preloading schedule is arranged based on the predicted peak access period, and hot media data is migrated from low-performance storage to high-performance storage in advance. Preloading operations are usually scheduled during the access off-peak period to avoid resource competition with normal business access. The schedule is formulated by comprehensively considering data transmission time, storage switching time and business continuity requirements.
[0060] Dynamic cross-cloud resource allocation coordinates and schedules resources in a multi-cloud environment based on resource reservation plans and data preloading schedules. Dynamic allocation refers to the flexible allocation of computing and storage resources between private and public clouds based on real-time resource demand and availability. Resource allocation paths are determined using a multi-path selection algorithm that comprehensively considers network latency, bandwidth capacity, security level, and cost factors to select the optimal transmission and storage path for different types of media data. Path selection supports dynamic switching, automatically switching to a backup path when the primary path fails or performance degrades. Execution time is calculated based on the resource allocation path and current network conditions to estimate the start, execution, and completion times of resource allocation tasks. This time calculation takes into account resource preparation time, data transmission time, and configuration deployment time. The results are used to generate a detailed execution schedule to guide the implementation of resource allocation tasks. Load balancing results are generated by monitoring the load distribution of each storage node and using a load balancing algorithm to adjust the resource allocation plan to ensure that the load level of each node remains within a reasonable range. Load balancing considers differences in node processing power and business importance, allocating more critical tasks to high-performance nodes and routine tasks to standard nodes. The dynamic resource allocation results integrate key information such as resource allocation path, execution time and load balancing results to form a resource scheduling plan. The plan supports real-time monitoring and dynamic adjustment, and optimizes the allocation strategy according to the actual situation during the execution process.
[0061] The above describes the media data management method based on the hybrid cloud architecture in the embodiment of the present application. The following describes the media data management system based on the hybrid cloud architecture in the embodiment of the present application. Figure 2 In one embodiment of the present application, a media data management system based on a hybrid cloud architecture includes: The prediction module is used to perform heat prediction processing on media data through the hybrid cloud media data heat prediction network to obtain media feature vectors and heat prediction values; a partitioning module, configured to perform storage hierarchical partitioning processing on the media data according to the media feature vector and the heat prediction value, and obtain a four-level storage allocation scheme; A migration module, configured to perform cross-cloud migration processing on the four-level storage allocation solution using a bidirectional inter-cloud bandwidth prediction model to obtain a data migration execution plan; A construction module, configured to perform distributed index construction processing on the media data according to the data migration execution plan to obtain a media content index table; The scheduling module is used to perform time-series-aware scheduling processing on the media data based on the media content index table to obtain a dynamic resource allocation result.
[0062] above Figure 2The media data management system based on the hybrid cloud architecture in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The media data management device based on the hybrid cloud architecture in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0063] Reference Figure 3 In an embodiment of the present invention, a media data management device based on a hybrid cloud architecture is further provided. The media data management device based on a hybrid cloud architecture may be a server, and its internal structure may be as follows: Figure 3 As shown. The media data management device based on the hybrid cloud architecture includes a processor, a memory, a display screen, an input device, a network interface and a database connected via a system bus. The computer-designed processor is used to provide computing and control capabilities. The memory of the media data management device based on the hybrid cloud architecture includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the media data management device based on the hybrid cloud architecture is used to store the corresponding data in this embodiment. The network interface of the media data management device based on the hybrid cloud architecture is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above method is implemented.
[0064] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the media data management device based on the hybrid cloud architecture to which the solution of the present invention is applied.
[0065] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the media data management method based on a hybrid cloud architecture.
[0066] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0067] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a media data management device based on a hybrid cloud architecture (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0068] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A media data management method based on hybrid cloud architecture, characterized in that: The method comprises: Perform heat prediction on media data through the hybrid cloud media data heat prediction network to obtain media feature vectors and heat prediction values; Performing storage hierarchical division processing on the media data according to the media feature vector and the heat prediction value to obtain a four-level storage allocation scheme; Perform cross-cloud migration processing on the four-level storage allocation solution using a bidirectional inter-cloud bandwidth prediction model to obtain a data migration execution plan; Performing distributed index construction processing on the media data according to the data migration execution plan to obtain a media content index table; Based on the media content index table, time-series-aware scheduling processing is performed on the media data to obtain a dynamic resource allocation result.
2. The media data management method based on hybrid cloud architecture according to claim 1, characterized in that: The method of performing heat prediction processing on the media data through the hybrid cloud media data heat prediction network to obtain the media feature vector and heat prediction value includes: The media data is input into the 3D convolutional neural network for convolution operation processing to obtain raw feature data including video frame sequence features, audio spectrum features and metadata features; Inputting the original feature data into a long short-term memory network for time series analysis and processing to obtain a historical access pattern sequence and time dependency relationship data of the media data; Performing feature splicing processing on the original feature data and the time dependency relationship data to obtain a fused feature vector; Inputting the fused feature vector into a fully connected layer for weight calculation and activation function processing to obtain a 128-dimensional media feature vector; Softmax normalization is performed on the media feature vector to obtain a popularity prediction value representing the access probability distribution in the next 7 days.
3. The media data management method based on hybrid cloud architecture according to claim 1, characterized in that: The media data is divided into storage levels according to the media feature vector and the heat prediction value to obtain a four-level storage allocation scheme, including: Performing data sensitivity calculation based on the media feature vector to obtain a sensitivity score value ranging from 0 to 1; Calculating access delay requirements based on the predicted popularity value and the sensitivity score value to obtain a delay tolerance parameter for the media data; Performing a weighted sum operation on the heat prediction value, the sensitivity score value, the delay tolerance parameter, and the storage cost coefficient to obtain a storage tier score value; Comparing the storage tier score with four preset threshold intervals to obtain classification identifiers of the hot data layer, the warm data layer, the cold data layer, and the ice data layer; The storage location mapping process is performed on the media data according to the classification identifier to obtain a four-level storage allocation plan including a target storage level, a migration priority, and an expected migration time.
4. The media data management method based on hybrid cloud architecture according to claim 1, characterized in that: The cross-cloud migration process of the four-level storage allocation solution is performed using a bidirectional inter-cloud bandwidth prediction model to obtain a data migration execution plan, including: The current network throughput, delay time, packet loss rate and historical bandwidth data are input into the bidirectional LSTM network for sequence encoding processing to obtain the network status feature sequence; Performing forward and reverse time series analysis on the network status feature sequence to obtain a bandwidth availability prediction value and confidence interval within the next 2 hours; Performing genetic algorithm optimization processing based on the migration priorities and bandwidth availability prediction values in the four-level storage allocation scheme to obtain the optimal time scheduling sequence for cross-cloud data migration; The media data is divided into blocks based on the optimal time scheduling sequence to obtain a 256KB data block sequence and parallel transmission path allocation; The data block sequence and the parallel transmission path are allocated to perform migration time calculation processing to obtain a data migration execution plan including a migration start time, a transmission path, and a completion time.
5. The media data management method based on hybrid cloud architecture according to claim 4, characterized in that: The genetic algorithm optimization process is performed based on the migration priority and bandwidth availability prediction value in the four-level storage allocation scheme to obtain the optimal time scheduling sequence for cross-cloud data migration, including: Encoding and arranging the media data in the four-level storage allocation scheme according to the migration priority to obtain a chromosome sequence including the migration order of the hot data layer, the warm data layer, the cold data layer, and the ice data layer; Performing fitness function calculation on the chromosome sequence based on the bandwidth availability prediction value to obtain a fitness score that integrates migration time and storage cost; Performing a high-fitness individual selection process based on the fitness score to obtain excellent chromosome pairs for reproducing the next generation; Performing crossover recombination and site mutation processing on the media data migration sequence in the excellent chromosome pair to obtain a new migration sequence combination; The new migration sequence combination is screened and processed based on the fitness score to obtain the optimal time scheduling sequence for cross-cloud data migration of media data from the private cloud to each storage layer of the public cloud.
6. The media data management method based on hybrid cloud architecture according to claim 1, characterized in that: The step of performing distributed index construction processing on the media data according to the data migration execution plan to obtain a media content index table includes: Performing content feature hash calculation processing based on the media data in the data migration execution plan to obtain a 64-bit media content fingerprint including a visual feature hash value, an audio feature hash value, and a content signature; Input the media content fingerprint into the consistent hash ring for node positioning processing to obtain the distributed storage location of the media data in the private cloud and public cloud storage nodes; Encapsulating the access rights and metadata information of the media data according to the distributed storage location to obtain an index entry including a file hash value, a storage path, and access control information; The index entries are transmitted across clouds according to a synchronous update mechanism to obtain index copy data maintained in the private cloud and the public cloud respectively; Update log recording and version control processing are performed on the index copy data to obtain a media content index table that supports fast retrieval and duplicate file detection.
7. The media data management method based on hybrid cloud architecture according to claim 1, characterized in that: The performing time-series-aware scheduling processing on the media data based on the media content index table to obtain a dynamic resource allocation result includes: Performing access time characteristic analysis on the media data according to the storage location information in the media content index table to obtain a time sensitivity classification including real-time streaming media data and batch archived data; Based on the time sensitivity classification, the resource utilization of the private cloud and the public cloud is monitored and calculated to obtain the current load status and available resource capacity of each storage node; Prioritizing the time sensitivity classification and available resource capacity to obtain a scheduling strategy for allocating real-time streaming media data to private cloud nodes in a priority manner; Performing predictive analysis on the historical access patterns and service periodicity characteristics of media data according to the scheduling strategy to obtain a resource reservation plan and a data preloading schedule; Based on the resource reservation plan and the data preloading schedule, cross-cloud resource dynamic allocation processing is performed to obtain a dynamic resource allocation result including a resource allocation path, execution time and load balancing result.
8. A media data management system based on hybrid cloud architecture, characterized in that: A method for managing media data based on a hybrid cloud architecture according to any one of claims 1 to 7, wherein the media data management system based on a hybrid cloud architecture comprises: The prediction module is used to perform heat prediction processing on media data through the hybrid cloud media data heat prediction network to obtain media feature vectors and heat prediction values; a partitioning module, configured to perform storage hierarchical partitioning processing on the media data according to the media feature vector and the heat prediction value, and obtain a four-level storage allocation scheme; A migration module, configured to perform cross-cloud migration processing on the four-level storage allocation solution using a bidirectional inter-cloud bandwidth prediction model to obtain a data migration execution plan; A construction module, configured to perform distributed index construction processing on the media data according to the data migration execution plan to obtain a media content index table; The scheduling module is used to perform time-series-aware scheduling processing on the media data based on the media content index table to obtain a dynamic resource allocation result.
9. A media data management device based on a hybrid cloud architecture, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the media data management method based on the hybrid cloud architecture according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor is enabled to execute the media data management method based on a hybrid cloud architecture according to any one of claims 1 to 7.