Intelligent media distribution system integrated with multi-dimensional state perception
By integrating a multi-dimensional state-aware intelligent media distribution system and utilizing deep reinforcement learning and edge computing, the system addresses the adaptability issues of media distribution systems to changes in user terminal performance and network conditions. This enables adaptive transmission and personalized recommendations, reduces latency and energy consumption, protects user privacy, and improves user experience and resource utilization efficiency.
Patent Information
- Application Number
- CN202511110366.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-12-09
AI Technical Summary
Existing media distribution systems cannot adapt to drastic changes in user terminal performance and network conditions, resulting in playback stuttering, excessive power consumption, or resource waste. At the same time, they cannot capture users' immediate needs in real time, and the centralized cloud content generation model brings high latency and data privacy risks.
The system employs an intelligent media distribution system that integrates multi-dimensional state awareness, including a device status monitoring module, a content feature analysis module, a distribution decision reasoning module, and a distribution and transmission execution module. It utilizes deep reinforcement learning algorithms to adjust media parameters in real time, and combines edge computing and cloud collaborative computing to achieve adaptive transmission and personalized recommendations. Furthermore, it protects user privacy through a decentralized privacy protection mechanism.
It achieves adaptive transmission of media content, eliminates playback stuttering and latency, reduces terminal power consumption, provides accurate and personalized content recommendations, improves recommendation accuracy and user experience, protects user privacy, and optimizes resource utilization efficiency.
Smart Images

Figure CN121099082A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of media content distribution technology, specifically relating to an intelligent media distribution system that integrates multi-dimensional state perception. Background Technology
[0002] With the rapid development of mobile internet and the widespread adoption of smart devices, digital media content consumption has become an important part of users' daily lives. Statistics show that global mobile video traffic accounts for over 70% of mobile data traffic, and this figure is projected to reach 80% by 2025. However, media distribution in the mobile internet environment also faces unprecedented challenges.
[0003] Existing media distribution systems mainly face the following technical challenges:
[0004] First, the traditional "one-size-fits-all" distribution strategy cannot adapt to the drastic changes in user terminal performance, network conditions, and real-time environmental scenarios. That is, when users use different media display devices at different times and locations, their device computing power (such as CPU utilization, GPU load, and NPU availability), network status (such as WiFi, 4G, or 5G network status), battery level, and ambient light conditions vary significantly. However, existing media distribution systems still generally adopt fixed cloud strategies, which can easily lead to problems such as playback stuttering, excessive power consumption, or resource waste.
[0005] Secondly, existing media content recommendation mechanisms are primarily based on outdated historical behavioral data, failing to capture users' immediate contextual needs. For example, users may prefer short videos during their commute, while at home they may prefer longer videos; users need subtitles in noisy environments, but may prefer pure audio in quiet environments. Traditional media content recommendation systems struggle to perceive these dynamically changing environmental factors, resulting in recommended content that is disconnected from the user's current situation.
[0006] Furthermore, the centralized cloud-based content generation model introduces high latency and data privacy risks. All media processing and personalized content generation tasks rely on cloud servers, increasing network latency and potentially leading to the leakage of sensitive user data. This is particularly problematic in scenarios requiring real-time responses, such as personalized caption generation and content summarization, where the cloud processing model struggles to meet low-latency requirements.
[0007] Therefore, there is an urgent need for an intelligent media distribution system that can integrate multi-dimensional state perception, achieve cloud-based collaborative computing, and protect user privacy in order to solve the above-mentioned technical challenges and provide more efficient, accurate, and secure personalized media distribution services. Summary of the Invention
[0008] The purpose of this invention is to provide an intelligent media distribution system that integrates multi-dimensional state awareness, in order to solve the problem that existing media distribution strategies cannot adapt to drastic changes in user terminal performance and network conditions.
[0009] To achieve the above objectives, the present invention adopts the following technical solution:
[0010] This invention provides an intelligent media distribution system that integrates multi-dimensional state perception, including a device state monitoring module, a content feature analysis module, a distribution decision reasoning module, and a distribution and transmission execution module;
[0011] The device status monitoring module is used to collect multi-dimensional device status characteristics of the media display terminal in real time. The multi-dimensional device status characteristics include device computing power indicators, network status parameters, battery power and / or remaining available storage space.
[0012] The content feature analysis module is used to perform feature extraction processing on the media content to be transmitted to obtain multi-dimensional content complexity features of the media content to be transmitted, wherein the multi-dimensional content complexity features include video feature parameters, audio feature parameters, content complexity score and / or file size;
[0013] The distribution decision reasoning module is communicatively connected to the device status monitoring module and the content feature analysis module, respectively. It is used to import the multi-dimensional device status features and the multi-dimensional content complexity features into the media distribution decision model pre-trained based on the deep reinforcement learning algorithm, and output the distribution and transmission scheme of the media content to be transmitted. The distribution and transmission scheme includes encoding format decision results, resolution decision results, bitrate decision results, transmission protocol decision results and / or computation task allocation decision results.
[0014] The distribution and transmission execution module is communicatively connected to the distribution decision reasoning module and is used to perform corresponding operations according to the distribution and transmission scheme so as to adaptively transmit the media content to be transmitted to the media display terminal.
[0015] Based on the above-mentioned invention, a novel scheme for optimizing distribution and transmission strategies by fusing multi-dimensional state perception results using deep reinforcement learning is provided. This scheme includes a device state monitoring module, a content feature analysis module, a distribution decision reasoning module, and a distribution and transmission execution module. The device state monitoring module collects multi-dimensional device state features of the media display terminal in real time. The content feature analysis module extracts multi-dimensional content complexity features of the media content to be transmitted. The distribution decision reasoning module imports the multi-dimensional device state features and multi-dimensional content complexity features into a media distribution decision model pre-trained using a deep reinforcement learning algorithm, outputting a distribution and transmission scheme, which is then executed by the distribution and transmission execution module. By using multi-dimensional state perception and dynamically adjusting media parameters based on real-time network conditions and device performance for adaptive media content transmission, playback stuttering and latency can be eliminated, terminal power consumption reduced, and precise personalized content highly matched to the current situation can be provided, achieving an immersive and smooth experience, facilitating practical application and promotion.
[0016] In one possible design, multi-dimensional device status characteristics of the media display terminal are collected in real time, including:
[0017] CPU utilization, used as an indicator of device computing power, is obtained by reading the / proc / stat file in real time.
[0018] And / or, obtain the GPU load as an indicator of device computing power in real time through the OpenGL ES extension interface;
[0019] And / or, the availability of NPUs, used as a metric for device computing power, can be detected in real time via the NNAPI interface;
[0020] And / or, obtain bandwidth and / or latency information in real time as network status parameters through the NetworkCapabilities function;
[0021] And / or, the battery level can be monitored in real time via the BatteryManager API interface.
[0022] In one possible design, feature extraction processing is performed on the media content to be transmitted to obtain multi-dimensional content complexity features of the media content to be transmitted, including:
[0023] Use the FFmpeg library to parse the media content to be transmitted to obtain the video resolution, frame rate, encoding format and / or bit rate as video feature parameters;
[0024] And / or, by reading the header information of the media content to be transmitted to obtain the audio encoding format, sampling rate and / or number of channels used as audio feature parameters;
[0025] And / or, the content complexity score S of the transmitted media content is calculated according to the following formula. complexity :
[0026]
[0027] In the formula, W represents the width of the video frame of the transmitted media content, H represents the height of the video frame of the transmitted media content, FPS represents the number of frames per second of the video frame of the transmitted media content, and R... bit The video bitrate of the transmitted media content is represented by DBM, the computing power index of the media display terminal is represented by NBW, and the bandwidth of the media display terminal is represented by NBW.
[0028] In one possible design, the intelligent media distribution system also includes a behavioral event collection module, a scene data acquisition module, an association strength inference module, and a recommendation decision fusion module;
[0029] The behavior event collection module is used to collect explicit behavior operation events of the user to which the media display terminal belongs, wherein the explicit behavior operation events include click events, search events, favorite events and / or rating events;
[0030] The scene data acquisition module is used to collect implicit scene data of the environment in which the media display terminal is located. The implicit scene data includes geographical location, time of day, ambient brightness, ambient noise level and / or terminal motion status.
[0031] The association strength inference module is communicatively connected to the behavior event collection module and the scene data acquisition module. It is used to construct a dynamic heterogeneous graph containing user nodes, scene nodes, and content nodes using the DA-GCN graph neural network model, based on the event data of the explicit behavior operation events, the implicit scene data, and the attribute data of each media content to be recommended. Then, using the graph attention mechanism formula of the DA-GCN graph neural network model, Attention(u,i,s)=softmax(LeakyReLU(W[hu||hi||hs])), it calculates the first association strength between each pair of two nodes and the second association strength between each group of three nodes. Here, the two nodes refer to the user node and the content node. The three nodes refer to user nodes, scene nodes, and content nodes. The user node is used to represent the current interests of the user based on the event data. The scene node is used to represent the dynamic environment state based on the implicit scene data in real time. The content node is used to represent the attribute characteristics of the media content to be recommended based on the attribute data. hu represents the embedding vector based on the event data, hi represents the embedding vector based on the attribute data, hs represents the embedding vector based on the implicit scene data, || represents the concatenation of two embedding vectors, W[] represents the decoder, LeakyReLU() represents the leaky linear rectified function, and softmax() represents the normalized exponential function.
[0032] The recommendation decision fusion module is communicatively connected to the association strength inference module. It is used to first calculate the corresponding comprehensive recommendation score for each media content to be recommended based on the corresponding first association strength and second association strength. Then, it arranges each media content to be recommended in descending order of comprehensive recommendation score to obtain a media content queue. Finally, it selects the top N media content to be recommended from the media content queue to obtain the TopN media content recommendation list, where N represents a positive integer.
[0033] In one possible design, the top N media content items to be recommended are selected from the media content queue to obtain a TopN media content recommendation list, including:
[0034] The order of media content in the media content queue is optimized and adjusted by applying a diversity optimization algorithm to obtain a new media content queue.
[0035] The top N media content to be recommended are selected from the new media content queue to obtain the TopN media content recommendation list, where N represents a positive integer.
[0036] In one possible design, the intelligent media distribution system further includes a task complexity assessment module, a task allocation decision module, and a distributed task execution module that are sequentially connected in communication.
[0037] The task complexity assessment module is used to analyze and obtain the corresponding task complexity score for each of the multiple generation tasks of the media content to be transmitted, wherein the multiple generation tasks include video rendering tasks, subtitle generation tasks and / or content summarization tasks.
[0038] The task allocation decision module is used to allocate the corresponding task to a task execution node corresponding to the task complexity score interval if, based on the multiple task complexity score intervals that correspond one-to-one with the multiple task execution nodes, the corresponding task is found to be located in a certain task complexity score interval among the multiple task complexity score intervals. The multiple task execution nodes include cloud, edge computer and / or the media display terminal.
[0039] The distributed task execution module is used to allocate each generation task to a corresponding task execution node according to the allocation strategy of each generation task, so as to execute the multiple generation tasks of the media content to be transmitted.
[0040] In one possible design, the intelligent media distribution system further includes a model distillation deployment module, which is used to distill the cloud-based BERT-Large teacher model into a terminal TinyBERT student model, and deploy the terminal TinyBERT student model to a terminal node to perform the generation task assigned to the terminal node, wherein the terminal node is an edge computer or the media display terminal.
[0041] In one possible design, the intelligent media distribution system also includes a decentralized privacy protection module, which is used to desensitize user data through a federated learning framework combined with differential privacy and homomorphic encryption technologies.
[0042] In one possible design, the decentralized privacy protection module includes a model federation synchronization unit and a data-sensitive classification unit and a differential privacy noise addition unit that are connected in communication.
[0043] The model federation synchronization unit is used to periodically perform federated aggregation of models on the cloud and all terminal nodes in the following manner: the terminal nodes calculate the gradient vector of their local model and upload it to the cloud using homomorphic encryption; the cloud aggregates and updates the global model using the FedAvg algorithm based on the gradient vectors from all terminal nodes, and then distributes the updated model parameters of the global model to each terminal node; the terminal nodes use the updated model parameters to update their local model.
[0044] The data sensitivity classification unit is used to automatically perform sensitivity classification processing on the user data of the user to which the media display terminal belongs, and obtain the sensitivity level of the user data;
[0045] The differential privacy noise-adding unit is used to add Laplace noise to user data that is sensitive data to protect privacy, add location offset to user data that is both sensitive data and location data to protect privacy, and add counting noise to user data that is both sensitive data and behavioral statistics data to protect privacy, based on the sensitivity level of the user data.
[0046] In one possible design, homomorphic encryption is used to protect the gradient vector when it is uploaded to the cloud. This includes: first quantizing the gradient vector into integer form, then using the SEAL homomorphic encryption library and employing the BFV homomorphic encryption scheme and key to encrypt the quantization result, and finally uploading the encrypted result to the cloud so that the cloud can perform gradient vector aggregation calculation in the encrypted domain and complete the global model update without decryption.
[0047] And / or, the aggregation is performed securely by multiple participants using the following secret sharing protocol: the gradient vector is divided into multiple parts and the multiple parts are distributed one-to-one to multiple aggregation nodes; the aggregation nodes use the Shamir threshold secret sharing protocol and achieve gradient recovery through polynomial interpolation; and the identities of the participants and data integrity are verified through zero-knowledge proof.
[0048] The beneficial effects of the above scheme are:
[0049] (1) This invention provides a new scheme for optimizing distribution and transmission strategies based on the fusion of multi-dimensional state perception results through deep reinforcement learning. The scheme includes a device state monitoring module, a content feature analysis module, a distribution decision reasoning module, and a distribution and transmission execution module. The device state monitoring module is used to collect multi-dimensional device state features of the media display terminal in real time. The content feature analysis module is used to extract multi-dimensional content complexity features of the media content to be transmitted. The distribution decision reasoning module is used to import the multi-dimensional device state features and multi-dimensional content complexity features into a media distribution decision model pre-trained based on a deep reinforcement learning algorithm, output the distribution and transmission scheme, and hand it over to the distribution and transmission execution module for execution. Thus, by using multi-dimensional state perception and dynamically adjusting media parameters according to real-time network conditions and device performance to adaptively transmit media content, playback stuttering and delay can be eliminated, terminal power consumption can be reduced, and accurate personalized content that is highly matched with the current situation can be provided to achieve an immersive and smooth experience.
[0050] (2) By building a new system that integrates explicit user operations and implicit scene data to construct dynamic user profiles and achieve scene-aware media content accurate recommendation, it can ensure that the content obtained by users is highly relevant to their current needs and environment based on real-time multi-dimensional state awareness dynamic recommendation, thereby achieving the goal of greatly improving recommendation accuracy.
[0051] (3) By building a cloud-based collaborative distributed media content generation system, the latency of generation tasks such as real-time subtitle generation and content summary generation can be reduced, resource utilization can be improved, unnecessary cloud service overhead can be reduced, and the overall system response speed can be improved.
[0052] (4) The embedding vector of scene nodes can capture the user's current scene status and potential needs in real time (such as short entertainment on the way to work or in-depth movie watching in a coffee shop), further ensuring the accuracy of recommendation decision results and improving user experience, so as to achieve the goal of "pushing appropriate content in appropriate scenes".
[0053] (5) Based on the privacy protection mechanism of federated learning, it can ensure that sensitive user data does not leave the terminal device without affecting the recommendation effect. Through decentralized architecture, federated learning and data classification and desensitization mechanisms, it can effectively protect user privacy and reduce data compliance risks.
[0054] (6) Through the intelligent computing task offloading mechanism, compared with the traditional cloud processing mode, the dynamic optimization scheduling and utilization of end-edge-cloud computing resources can be realized, reducing cloud computing load, reducing network bandwidth consumption, improving edge node efficiency, and thus optimizing resource utilization efficiency.
[0055] (7) By sensing multi-dimensional information such as the status of user terminal devices, network environment and usage scenarios in real time, and combining cloud collaborative computing and edge intelligence technology, adaptive media content transmission, personalized recommendation and privacy protection can be achieved, solving the technical limitations of traditional media distribution systems in dynamic mobile environments, and providing users with a smoother, more accurate and secure personalized media consumption experience. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a schematic diagram of the structure of an intelligent media distribution system that integrates multi-dimensional state perception, provided in an embodiment of the present invention.
[0058] Figure 2 This is a flowchart illustrating the process of multi-dimensional context-adaptive transmission provided in an embodiment of the present invention.
[0059] Figure 3 This is an example diagram of an architecture for intelligent scene fusion recommendation provided in an embodiment of the present invention.
[0060] Figure 4 This is an example diagram of the architecture for cloud-based collaborative content generation provided in an embodiment of the present invention.
[0061] Figure 5 This is an example diagram of an architecture for decentralized privacy protection provided in an embodiment of the present invention. Detailed Implementation
[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and descriptions of the embodiments or the prior art. Obviously, the following description of the structure of the accompanying drawings is only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these embodiments without creative effort. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention.
[0063] It should be understood that although the terms "first" and "second", etc., may be used herein to describe various objects, these objects should not be limited by these terms. These terms are only used to distinguish one object from another. For example, the first object may be referred to as the second object, and similarly, the second object may be referred to as the first object, without departing from the scope of the exemplary embodiments of the invention.
[0064] It should be understood that the term "and / or" that may appear in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, or A and B exist simultaneously. Another example is A, B and / or C, which can mean that any one of A, B, and C or any combination thereof exists. The term " / and" that may appear in this document describes another relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone or A and B exist simultaneously. In addition, the character " / " that may appear in this document generally indicates that the related objects before and after it are in an "or" relationship.
[0065] Example
[0066] like Figure 1 As shown, the intelligent media distribution system provided in this embodiment, which integrates multi-dimensional state perception, includes, but is not limited to, a device state monitoring module, a content feature analysis module, a distribution decision reasoning module, and a distribution transmission execution module.
[0067] The device status monitoring module is used to collect multi-dimensional device status characteristics of the media display terminal in real time. These multi-dimensional device status characteristics include, but are not limited to, device computing power indicators, network status parameters, battery level, and / or remaining available storage space. The media display terminal is an electronic device owned by the user that can play and display media content (such as video files), such as a smartphone or tablet. Specifically, the device computing power indicators include, but are not limited to, CPU (Central Processing Unit) utilization, GPU (Graphics Processing Unit) load, and NPU (Neural Network Processing Unit) availability. The network status parameters include, but are not limited to, bandwidth, latency information, packet loss rate, and connection type. The multi-dimensional device status characteristics can be monitored in real time through Android / iOS native APIs (Application Programming Interfaces). Specifically, the multi-dimensional device status characteristics of the media display terminal are collected in real time, including but not limited to: CPU utilization as a device computing power indicator by reading the / proc / stat file in real time; and / or GPU load as a device computing power indicator by querying the OpenGL ES extension interface in real time; and / or NPU availability as a device computing power indicator by detecting the NNAPI interface in real time; and / or bandwidth and / or latency information as network status parameters by obtaining the NetworkCapabilities function in real time; and / or battery level by monitoring the BatteryManager API interface in real time; and so on.The aforementioned ` / proc / stat` file (a virtual file in the Linux system that provides real-time data on system status and various statistics), OpenGL ES (OpenGL for Embedded Systems, a subset of the OpenGL 3D graphics API designed for embedded devices such as mobile phones, PDAs, and game consoles) extension interface, NNAPI (Android Neural Networks API, an Android C API designed to run computationally intensive machine learning operations on mobile devices) interface, NetworkCapabilities (which can be understood as identifiers of network capabilities, similar to the Capability of Call, and more like a utility class) function, and BatteryManager API (a JavaScript API used to access device battery status information) interface are all existing terms and will not be elaborated upon here. Furthermore, the monitoring frequency of these multi-dimensional device status characteristics can be set to once per second, and the corresponding data can be stored in a local sliding window.
[0068] The content feature analysis module is used to perform feature extraction processing on the media content to be transmitted to obtain multi-dimensional content complexity features of the media content to be transmitted. These multi-dimensional content complexity features include, but are not limited to, video feature parameters, audio feature parameters, content complexity score, and / or file size. The media content to be transmitted is the media content to be transmitted from the cloud to the media display terminal, such as a video file. Specifically, the video feature parameters include, but are not limited to, video resolution, frame rate, encoding format, and / or bitrate, and the audio feature parameters include, but are not limited to, audio encoding format, sampling rate, and / or number of channels. Specifically, feature extraction processing is performed on the media content to be transmitted to obtain multi-dimensional content complexity features of the media content to be transmitted, including but not limited to: using the FFmpeg library to parse the media content to be transmitted to obtain video resolution, frame rate, encoding format, and / or bit rate, etc., used as video feature parameters; and / or, by reading the file header information of the media content to be transmitted to obtain audio encoding format, sampling rate, and / or number of channels, etc., used as audio feature parameters; and / or, calculating the content complexity score S of the transmitted media content according to the following formula. complexity :
[0069]
[0070] In the formula, W represents the width of the video frame of the transmitted media content, H represents the height of the video frame of the transmitted media content, FPS represents the number of frames per second of the video frame of the transmitted media content, and R...bit The video bitrate of the transmitted media content is represented by DBM, the computing power of the media display terminal is represented by NBW, and the bandwidth of the media display terminal is represented by NBW. The aforementioned FFmpeg (an open-source computer program that can record, convert, and stream digital audio and video) library and file header information are all existing terms and will not be elaborated upon here.
[0071] The distribution decision reasoning module is communicatively connected to the device status monitoring module and the content feature analysis module, respectively. It is used to import the multi-dimensional device status features and the multi-dimensional content complexity features into a media distribution decision model pre-trained based on a deep reinforcement learning algorithm, and output a distribution and transmission scheme for the media content to be transmitted. This distribution and transmission scheme includes, but is not limited to, encoding format decision results, resolution decision results, bitrate decision results, transmission protocol decision results, and / or computational task allocation decision results. The deep reinforcement learning algorithm (DRL) is an existing artificial intelligence method that combines the perception capabilities of deep learning with the decision-making mechanism of reinforcement learning. It achieves optimal policy learning for complex tasks through agent-environment interaction. Deep reinforcement learning can learn better policy representations to solve some problems in optimization models, such as large-scale, complex constraints, and multi-coupling problems (in practical optimization problems, there are often problems such as large model size, complex constraints, and mutual coupling between multiple variables, which pose significant challenges to optimization solutions. Deep reinforcement learning, as a powerful learning method, can learn optimal policies through interaction with the environment, thus providing a possibility for solving these problems). Therefore, a deep reinforcement learning model (i.e., the media distribution decision model) for outputting media distribution and transmission schemes can be conventionally trained based on model training algorithms such as PPO (Proximal Policy Optimization, a policy optimization method in reinforcement learning, whose core idea is to avoid instability during training by limiting the magnitude of policy updates). This allows for the output of the optimal distribution and transmission scheme for the media content to be transmitted.Specifically, the encoding format decision result refers to the specific encoding format chosen from H.264, H.265, and AVI (Audio Video Interleaved) (i.e., when network bandwidth is insufficient, the bitrate can be reduced by switching encoding formats); the resolution decision result refers to the specific resolution chosen from 720P, 1080P, and 4K; the transmission protocol decision result refers to the specific transmission protocol chosen from HTTP (Hypertext Transfer Protocol) and QUIC (QuickUDP Internet Connections, an experimental transport layer network transmission protocol) (i.e., when network latency is high, the QUIC protocol can be dynamically enabled to improve transmission efficiency); the computing task allocation decision result refers to whether to offload high-load tasks such as video transcoding or special effects rendering to edge node devices (i.e., when device computing power is insufficient, these high-load tasks can be offloaded to edge node devices); and so on. Furthermore, the media distribution decision model can be trained in the cloud and then deployed on terminal devices (such as the media display terminal) via TensorFlowLite (an open-source deep learning framework that runs TensorFlow models on the device) for real-time inference.
[0072] The distribution and transmission execution module, communicatively connected to the distribution decision and reasoning module, is used to perform corresponding operations according to the distribution and transmission scheme, so as to adaptively transmit the media content to be transmitted to the media display terminal. For example, according to the distribution and transmission scheme, if transcoding is required, the transcoding task is sent to the FFmpeg service on the edge node; if the QUIC protocol is selected, Chrome's QUIC is used to establish the connection; if the resolution needs to be reduced, a video stream of the corresponding resolution is requested; and so on. Furthermore, the entire process is transparent to the user and will switch smoothly without interrupting the media content display.
[0073] Based on the detailed description of the above intelligent media distribution system and Figure 2As illustrated in the example of multi-dimensional context-adaptive transmission, this embodiment provides a novel scheme for optimizing distribution and transmission strategies based on deep reinforcement learning and the fusion of multi-dimensional state perception results. This scheme includes a device state monitoring module, a content feature analysis module, a distribution decision reasoning module, and a distribution and transmission execution module. The device state monitoring module collects multi-dimensional device state features of the media display terminal in real time. The content feature analysis module extracts multi-dimensional content complexity features of the media content to be transmitted. The distribution decision reasoning module imports the multi-dimensional device state features and multi-dimensional content complexity features into a media distribution decision model pre-trained based on a deep reinforcement learning algorithm, outputs a distribution and transmission scheme, and executes it through the distribution and transmission execution module. Thus, by using multi-dimensional state perception and dynamically adjusting media parameters according to real-time network conditions and device performance for adaptive media content transmission, playback stuttering and latency can be eliminated, terminal power consumption reduced, and precise personalized content highly matched to the current situation can be provided, achieving an immersive and smooth experience, facilitating practical application and promotion.
[0074] Preferably, the intelligent media distribution system also includes, but is not limited to, a behavior event collection module, a scene data acquisition module, an association strength inference module, and a recommendation decision fusion module.
[0075] The behavioral event collection module is used to collect explicit behavioral operation events of the users belonging to the media display terminal. These explicit behavioral operation events include, but are not limited to, click events, search events, favorite events, and / or rating events. These explicit behavioral operation events reflect the user's recent / current interests. Specifically, the event data for click events includes, but is not limited to, click timestamps, clicked content identifiers, and dwell time after clicking; the event data for search events includes, but is not limited to, query terms, clicked content of search results, and search depth; the event data for favorite events includes, but is not limited to, favorite time and unfavorited time; and the rating events include, but are not limited to, rating values and rating time. The aforementioned click events, search events, favorite events, and rating events can be automatically recorded by the terminal device during runtime and obtained through regular access history records. Furthermore, the event data for explicit behavioral operation events can be pre-processed locally by the terminal device and then uploaded to the cloud for in-depth analysis.
[0076] The scene data acquisition module is used to collect implicit scene data of the environment in which the media display terminal is located. This implicit scene data includes, but is not limited to, geographical location, time of day, ambient brightness, ambient noise level, and / or terminal motion status. The implicit scene data reflects the recent / current environment of the device. With user authorization, the geographical location can be obtained through GPS (Global Positioning System) and WiFi positioning technology; the time of day can be obtained through system time and time zone information analysis; the ambient brightness can be detected through a light intensity sensor; the ambient noise level can be detected through a microphone (only volume is analyzed, no content is recorded); and the terminal motion status can be determined through an accelerometer and gyroscope, etc. Furthermore, the implicit scene data can be processed locally on the terminal device in real time to obtain its feature vector, and only this feature vector needs to be uploaded to the cloud for in-depth analysis.
[0077] The association strength inference module is communicatively connected to the behavior event collection module and the scene data acquisition module. It is used to construct a dynamic heterogeneous graph containing user nodes, scene nodes, and content nodes using the DA-GCN graph neural network model, based on the event data of the explicit behavior operation events, the implicit scene data, and the attribute data of each media content to be recommended. Then, using the graph attention mechanism formula of the DA-GCN graph neural network model, Attention(u,i,s)=softmax(LeakyReLU(W[hu||hi||hs])), it calculates the first association strength between each pair of two nodes and the second association strength between each group of three nodes. Here, the two nodes refer to user nodes and content nodes. The three nodes refer to the user node, scene node, and content node. The user node is used to represent the current interests of the user based on the event data. The scene node is used to represent the dynamic environment state based on the implicit scene data in real time. The content node is used to represent the attribute characteristics of the media content to be recommended based on the attribute data. hu represents the embedding vector based on the event data, hi represents the embedding vector based on the attribute data, hs represents the embedding vector based on the implicit scene data, || represents concatenating two embedding vectors, W[] represents the decoder, LeakyReLU() represents the leaky linear rectified function, which is an improved ReLU (Rectified Linear Unit) activation function, and softmax() represents the normalized exponential function. The DA-GCN graph neural network model (Dynamic Attention Graph Convolutional Network) is an existing deep learning model that combines dynamic graph convolutional networks with attention mechanisms. It is mainly used to process spatiotemporal data association modeling and dynamic feature extraction. Therefore, the dynamic heterogeneous graph can be conventionally constructed based on the event data, the implicit scene data, and the attribute data, and the first association strength A1 and the second association strength can be conventionally calculated based on the graph attention mechanism formula.
[0078] The recommendation decision fusion module, communicatively connected to the association strength inference module, first calculates a comprehensive recommendation score for each media content to be recommended based on its corresponding first and second association strengths. Then, it arranges the media content to be recommended in descending order of their comprehensive recommendation scores, forming a media content queue. Finally, it selects the top N media content to be recommended from the media content queue to obtain a TopN media content recommendation list, where N represents a positive integer. For example, the two weight values used in the aforementioned weighted calculation can be: a weight of 0.7 corresponding to the first association strength and a weight of 0.3 corresponding to the second association strength. The TopN media content recommendation list is pushed to the media display terminal for output display, allowing users to easily select the media content to be transmitted. Furthermore, to ensure the richness of the recommended content, preferably, the top N media content to be recommended is selected from the media content queue to obtain a Top N media content recommendation list. This includes, but is not limited to: applying a diversity optimization algorithm to optimize and adjust the order of media content in the media content queue to obtain a new media content queue; selecting the top N media content to be recommended from the new media content queue to obtain a Top N media content recommendation list, where N represents a positive integer. The diversity optimization algorithm is an existing algorithm used to improve the diversity of the recommendation system and balance the breadth and depth of user interests. Common algorithms include dynamic diversity parameter adjustment, Monte Carlo tree search, DPP (Determinantal Point Process) algorithm, and MMR (Maximal Marginal Relevance) algorithm.
[0079] Based on the aforementioned behavioral event collection module, scene data acquisition module, association strength inference module, and recommendation decision fusion module, a new system can be built that integrates explicit user operations and implicit scene data to construct dynamic user profiles and achieve scene-aware, accurate media content recommendation (i.e., obtaining...). Figure 3 The technical architecture shown here enables intelligent fusion and recommendation based on scenarios. This allows for dynamic recommendations based on real-time multi-dimensional state perception, ensuring that the content received by users is highly relevant to their current needs and environment, thereby significantly improving the accuracy of recommendations.
[0080] Preferably, the intelligent media distribution system further includes, but is not limited to, a task complexity assessment module, a task allocation decision module, and a distributed task execution module that are sequentially connected in communication.
[0081] The task complexity assessment module is used to analyze and obtain a corresponding task complexity score for each of the multiple generation tasks of the media content to be transmitted. The multiple generation tasks include, but are not limited to, video rendering tasks, subtitle generation tasks, and / or content summarization tasks. The specific analysis process for the aforementioned task complexity score is based on existing technologies: for example, for the video rendering task, the corresponding task complexity score can be calculated based on resolution, frame rate, and effects complexity; for the subtitle generation task, the corresponding task complexity score can be evaluated based on audio duration and language complexity; and for the content summarization task, the corresponding task complexity score can be evaluated based on text length and semantic complexity.
[0082] The task allocation decision module is used to, for each generated task, if a task complexity score is found to fall within a certain task complexity score interval based on multiple task complexity score intervals corresponding to multiple task execution nodes, then the corresponding task is allocated to a task execution node corresponding to that task complexity score interval. The multiple task execution nodes include a cloud platform, an edge computer, and / or the media display terminal. The multiple task complexity score intervals do not overlap. For example, when the multiple task execution nodes include a cloud platform, an edge computer, and the media display terminal, the task complexity score interval corresponding to the cloud platform can be set to 0.8 or higher, the task complexity score interval corresponding to the edge computer can be set to [0.3, 0.8], and the task complexity score interval corresponding to the media display terminal can be set to below 0.3.
[0083] The distributed task execution module is used to allocate each generation task to a corresponding task execution node according to the allocation strategy of each generation task, so as to complete the multiple generation tasks of the media content to be transmitted. For example, an NVIDIA V100 GPU cluster can be used in the cloud to process video rendering tasks and call FFmpeg and OpenCV libraries for video processing; the integrated GPU of Intel NUC can be used in the edge computer to process image enhancement and audio noise reduction tasks; and an NPU can be used in the media display device to accelerate the running of a lightweight model for real-time subtitle generation. In addition, each task execution node can also coordinate tasks and synchronize results through a message queue (Redis).
[0084] Based on the aforementioned task complexity assessment module, task allocation decision module, and distributed task execution module, a cloud-based collaborative distributed media content generation system can be constructed (i.e., to obtain...). Figure 4The technical architecture shown can be used for cloud-based collaborative content generation, thereby reducing the latency of generation tasks such as real-time caption generation and content summary generation, improving resource utilization, reducing unnecessary cloud service overhead, and improving the overall system response speed.
[0085] Further preferably, the intelligent media distribution system also includes, but is not limited to, a model distillation deployment module, wherein the model distillation deployment module is used to distill the cloud-based BERT-Large teacher model into a terminal TinyBERT student model, and deploy the terminal TinyBERT student model to a terminal node to execute the generation task assigned to the terminal node, wherein the terminal node is an edge computer or the media display terminal. The cloud-based BERT-Large teacher model is an existing large language model deployed on a cloud server for real-time caption generation and content summarization generation; the specific process of the aforementioned distillation is as follows: using the knowledge distillation loss function L=α×LCE+(1-α)×LKD, where LCE represents cross-entropy loss, LKD represents knowledge distillation loss, and α is set to 0.3, the size of the model after distillation can be compressed from 400MB to 50MB, thereby increasing the inference speed of the terminal TinyBERT student model by 5 times, which is conducive to its deployment on the terminal node to achieve real-time caption generation and content summarization generation. Thus, knowledge transfer can also be achieved through model distillation technology.
[0086] Preferably, the intelligent media distribution system further includes a decentralized privacy protection module, which is used to de-identify user data through a federated learning framework combined with differential privacy and homomorphic encryption technologies. Specifically, it employs a federated learning framework and differential privacy technology to achieve continuous model optimization while protecting user data privacy: on the one hand, it performs sensitivity analysis on user data, and for sensitive data (such as biometrics and location information), it performs de-identification and feature extraction locally, uploading only the feature vectors with differential privacy noise enhancement; while for non-sensitive data (such as operation frequency), it can be uploaded to the cloud after standardization. On the other hand, it uses a federated learning framework to ensure that the user-side model is trained only locally, and the gradient parameters are securely uploaded to the federated aggregation server through homomorphic encryption technology, so as to use a secure multi-party computation protocol (such as using the Shamir threshold privacy sharing protocol) for global model updates, ensuring that the original data does not leave the user's device throughout the process. Examples of the technical architecture of the decentralized privacy protection module include... Figure 5 As shown.
[0087] Specifically, the decentralized privacy protection module includes, but is not limited to, a model federation synchronization unit, a data-sensitive classification unit and a differential privacy noise-adding unit that are connected in communication.
[0088] The model federation synchronization unit is used to periodically perform federated aggregation of models on the cloud and all terminal nodes in the following manner: Terminal nodes calculate the gradient vector of their local model and upload it to the cloud using homomorphic encryption; the cloud aggregates and updates the global model using the FedAvg algorithm based on the gradient vectors from all terminal nodes, and then distributes the updated model parameters to each terminal node; the terminal nodes update their local model using the updated model parameters. The aforementioned period may be, but is not limited to, 24 hours; the aforementioned model may include, but is not limited to, the media distribution decision model and / or the DA-GCN graph neural network model, etc. To ensure the security of the computation process (even though encryption overhead increases communication by approximately 2 times), specifically, the gradient vector is uploaded to the cloud using homomorphic encryption, including but not limited to: first quantizing the gradient vector into integer form, then encrypting the quantization result using the SEAL homomorphic encryption library and employing the BFV homomorphic encryption scheme and key, and finally uploading the encrypted result to the cloud so that the cloud can perform gradient vector aggregation calculation in the encrypted domain and complete the global model update without decryption. The aforementioned SEAL (Simple Encrypted Arithmetic Library, an open-source library developed by Microsoft to provide developers with implementations of homomorphic encryption algorithms), BFV (Brakerski-Fan-Vercauteren, a method for implementing fully homomorphic encryption), and FedAvg (Federated Averaging, one of the most classic algorithms in federated learning, which trains a model on a local device and uploads parameters to a server for weighted average aggregation to achieve global model updates) algorithms are all existing technologies and will not be elaborated upon here. To ensure that no single aggregation node can obtain complete user data during the aggregation process, thereby enhancing security, the aggregation process employs a secure aggregation method using a secret sharing protocol among multiple participants: the gradient vector is divided into multiple parts, and these parts are distributed one-to-one to multiple aggregation nodes; the aggregation nodes use the Shamir threshold secret sharing protocol (a threshold secret sharing algorithm based on polynomial interpolation, proposed by Adi Shamir in 1979) and achieve gradient recovery through polynomial interpolation; zero-knowledge proofs are used to verify the identities of the participants and the integrity of the data. Therefore, based on the privacy protection mechanism of federated learning, sensitive user data can be ensured not to leave the terminal device without affecting the recommendation effect. Through decentralized architecture and federated learning, user privacy can be effectively protected and data compliance risks reduced. Furthermore, version control can ensure model consistency and automatically roll back abnormal nodes to a stable version.
[0089] The data sensitivity classification unit is used to automatically classify the user data of the users belonging to the media display terminal to obtain the sensitivity level of the user data. The user data specifically includes, but is not limited to, biometric data (such as fingerprints and facial data, which can be classified as high sensitivity level), precise location information (which can be classified as high sensitivity level), address book (which can be classified as high sensitivity level), device identifier (which can be classified as medium sensitivity level), coarse location (which can be classified as high sensitivity level), operation statistics (which can be classified as low sensitivity level), and content preferences (which can be classified as low sensitivity level), etc.; different processing strategies are adopted for different sensitivity levels.
[0090] The differential privacy noise-adding unit is used to add Laplace noise to user data that is sensitive data to protect privacy, add location offset to user data that is both sensitive data and location data to protect privacy, and add counting noise to user data that is both sensitive data and behavioral statistics data to protect privacy, based on the sensitivity level of the user data. In the aforementioned process of adding Laplace noise, the noise magnitude is determined according to the privacy budget ε, where ε is 0.1 and the noise standard deviation is... During the aforementioned process of adding position offset, the offset distance follows an exponential distribution; during the aforementioned process of adding counting noise, the counting noise can be made to satisfy the (ε,δ) differential privacy condition, where δ=10. 6 Therefore, the aforementioned data classification and anonymization mechanism can further and effectively protect user privacy and reduce data compliance risks.
[0091] In summary, the intelligent media distribution system provided in this embodiment has the following technical effects:
[0092] (1) This embodiment provides a new scheme for optimizing the distribution and transmission strategy based on the fusion of multi-dimensional state perception results through deep reinforcement learning. It includes a device state monitoring module, a content feature analysis module, a distribution decision reasoning module, and a distribution and transmission execution module. The device state monitoring module is used to collect multi-dimensional device state features of the media display terminal in real time. The content feature analysis module is used to extract multi-dimensional content complexity features of the media content to be transmitted. The distribution decision reasoning module is used to import the multi-dimensional device state features and multi-dimensional content complexity features into the media distribution decision model pre-trained based on the deep reinforcement learning algorithm, output the distribution and transmission scheme, and hand it over to the distribution and transmission execution module for execution. Thus, by using multi-dimensional state perception and dynamically adjusting media parameters according to real-time network conditions and device performance to adaptively transmit media content, playback stuttering and delay can be eliminated, terminal power consumption can be reduced, and accurate personalized content that is highly matched with the current situation can be provided to achieve an immersive and smooth experience.
[0093] (2) By building a new system that integrates explicit user operations and implicit scene data to construct dynamic user profiles and achieve scene-aware media content accurate recommendation, it can ensure that the content obtained by users is highly relevant to their current needs and environment based on real-time multi-dimensional state awareness dynamic recommendation, thereby achieving the goal of greatly improving recommendation accuracy.
[0094] (3) By building a cloud-based collaborative distributed media content generation system, the latency of generation tasks such as real-time subtitle generation and content summary generation can be reduced, resource utilization can be improved, unnecessary cloud service overhead can be reduced, and the overall system response speed can be improved.
[0095] (4) The embedding vector of scene nodes can capture the user's current scene status and potential needs in real time (such as short entertainment on the way to work or in-depth movie watching in a coffee shop), further ensuring the accuracy of recommendation decision results and improving user experience, so as to achieve the goal of "pushing appropriate content in appropriate scenes".
[0096] (5) Based on the privacy protection mechanism of federated learning, it can ensure that sensitive user data does not leave the terminal device without affecting the recommendation effect. Through decentralized architecture, federated learning and data classification and desensitization mechanisms, it can effectively protect user privacy and reduce data compliance risks.
[0097] (6) Through the intelligent computing task offloading mechanism, compared with the traditional cloud processing mode, the dynamic optimization scheduling and utilization of end-edge-cloud computing resources can be realized, reducing cloud computing load, reducing network bandwidth consumption, improving edge node efficiency, and thus optimizing resource utilization efficiency.
[0098] (7) By sensing multi-dimensional information such as the status of user terminal devices, network environment and usage scenarios in real time, and combining cloud collaborative computing and edge intelligence technology, adaptive media content transmission, personalized recommendation and privacy protection can be achieved, solving the technical limitations of traditional media distribution systems in dynamic mobile environments, and providing users with a smoother, more accurate and secure personalized media consumption experience.
[0099] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An intelligent media distribution system integrating multi-dimensional state perception, characterized in that, It includes a device status monitoring module, a content feature analysis module, a distribution decision reasoning module, and a distribution and transmission execution module; The device status monitoring module is used to collect multi-dimensional device status characteristics of the media display terminal in real time. The multi-dimensional device status characteristics include device computing power indicators, network status parameters, battery power and / or remaining available storage space. The content feature analysis module is used to perform feature extraction processing on the media content to be transmitted to obtain multi-dimensional content complexity features of the media content to be transmitted, wherein the multi-dimensional content complexity features include video feature parameters, audio feature parameters, content complexity score and / or file size; The distribution decision reasoning module is communicatively connected to the device status monitoring module and the content feature analysis module, respectively. It is used to import the multi-dimensional device status features and the multi-dimensional content complexity features into the media distribution decision model pre-trained based on the deep reinforcement learning algorithm, and output the distribution and transmission scheme of the media content to be transmitted. The distribution and transmission scheme includes encoding format decision results, resolution decision results, bitrate decision results, transmission protocol decision results and / or computation task allocation decision results. The distribution and transmission execution module is communicatively connected to the distribution decision reasoning module and is used to perform corresponding operations according to the distribution and transmission scheme so as to adaptively transmit the media content to be transmitted to the media display terminal.
2. The intelligent media distribution system as described in claim 1, characterized in that, Real-time collection of multi-dimensional device status characteristics of media display terminals, including: CPU utilization, used as an indicator of device computing power, is obtained by reading the / proc / stat file in real time. And / or, obtain the GPU load as an indicator of device computing power in real time through the OpenGL ES extension interface; And / or, the availability of NPUs, used as a metric for device computing power, can be detected in real time via the NNAPI interface; And / or, obtain bandwidth and / or latency information in real time as network status parameters through the NetworkCapabilities function; And / or, the battery level can be monitored in real time via the BatteryManager API interface.
3. The intelligent media distribution system as described in claim 1, characterized in that, Feature extraction processing is performed on the content of the medium to be transmitted to obtain multi-dimensional content complexity features of the content, including: Use the FFmpeg library to parse the media content to be transmitted to obtain the video resolution, frame rate, encoding format and / or bit rate as video feature parameters; And / or, by reading the header information of the media content to be transmitted to obtain the audio encoding format, sampling rate and / or number of channels used as audio feature parameters; And / or, the content complexity score S of the transmitted media content is calculated according to the following formula. complexity : In the formula, W represents the width of the video frame of the transmitted media content, H represents the height of the video frame of the transmitted media content, FPS represents the number of frames per second of the video frame of the transmitted media content, and R... bit The video bitrate of the transmitted media content is represented by DBM, the computing power index of the media display terminal is represented by NBW, and the bandwidth of the media display terminal is represented by NBW.
4. The intelligent media distribution system as described in claim 1, characterized in that, The intelligent media distribution system also includes a behavior event collection module, a scene data acquisition module, an association strength inference module, and a recommendation decision fusion module; The behavior event collection module is used to collect explicit behavior operation events of the user to which the media display terminal belongs, wherein the explicit behavior operation events include click events, search events, favorite events and / or rating events; The scene data acquisition module is used to collect implicit scene data of the environment in which the media display terminal is located. The implicit scene data includes geographical location, time of day, ambient brightness, ambient noise level and / or terminal motion status. The association strength inference module is communicatively connected to the behavior event collection module and the scene data acquisition module. It is used to construct a dynamic heterogeneous graph containing user nodes, scene nodes, and content nodes using the DA-GCN graph neural network model, based on the event data of the explicit behavior operation events, the implicit scene data, and the attribute data of each media content to be recommended. Then, using the graph attention mechanism formula of the DA-GCN graph neural network model, Attention(u,i,s)=softmax(LeakyReLU(W[hu||hi||hs])), it calculates the first association strength between each pair of two nodes and the second association strength between each group of three nodes. Here, the two nodes refer to the user node and the content node. The three nodes refer to user nodes, scene nodes, and content nodes. The user node is used to represent the current interests of the user based on the event data. The scene node is used to represent the dynamic environment state based on the implicit scene data in real time. The content node is used to represent the attribute characteristics of the media content to be recommended based on the attribute data. hu represents the embedding vector based on the event data, hi represents the embedding vector based on the attribute data, hs represents the embedding vector based on the implicit scene data, || represents the concatenation of two embedding vectors, W[] represents the decoder, LeakyReLU() represents the leaky linear rectified function, and softmax() represents the normalized exponential function. The recommendation decision fusion module is communicatively connected to the association strength inference module. It is used to first calculate the corresponding comprehensive recommendation score for each media content to be recommended based on the corresponding first association strength and second association strength. Then, it arranges each media content to be recommended in descending order of comprehensive recommendation score to obtain a media content queue. Finally, it selects the top N media content to be recommended from the media content queue to obtain the TopN media content recommendation list, where N represents a positive integer.
5. The intelligent media distribution system as described in claim 4, characterized in that, The top N media content to be recommended are selected from the media content queue to obtain a TopN media content recommendation list, including: The order of media content in the media content queue is optimized and adjusted by applying a diversity optimization algorithm to obtain a new media content queue. The top N media content to be recommended are selected from the new media content queue to obtain the TopN media content recommendation list, where N represents a positive integer.
6. The intelligent media distribution system as described in claim 1, characterized in that, The intelligent media distribution system also includes a task complexity assessment module, a task allocation decision module, and a distributed task execution module that are sequentially connected in communication. The task complexity assessment module is used to analyze and obtain the corresponding task complexity score for each of the multiple generation tasks of the media content to be transmitted, wherein the multiple generation tasks include video rendering tasks, subtitle generation tasks and / or content summarization tasks. The task allocation decision module is used to allocate the corresponding task to a task execution node corresponding to the task complexity score interval if, based on the multiple task complexity score intervals that correspond one-to-one with the multiple task execution nodes, the corresponding task is found to be located in a certain task complexity score interval among the multiple task complexity score intervals. The multiple task execution nodes include cloud, edge computer and / or the media display terminal. The distributed task execution module is used to allocate each generation task to a corresponding task execution node according to the allocation strategy of each generation task, so as to execute the multiple generation tasks of the media content to be transmitted.
7. The intelligent media distribution system as described in claim 6, characterized in that, The intelligent media distribution system also includes a model distillation deployment module, which is used to distill the cloud-based BERT-Large teacher model into a terminal TinyBERT student model, and deploy the terminal TinyBERT student model to a terminal node to perform the generation task assigned to the terminal node, wherein the terminal node is an edge computer or the media display terminal.
8. The intelligent media distribution system as described in claim 1, characterized in that, The intelligent media distribution system also includes a decentralized privacy protection module, which is used to desensitize user data through a federated learning framework combined with differential privacy technology and homomorphic encryption technology.
9. The intelligent media distribution system as described in claim 8, characterized in that, The decentralized privacy protection module includes a model federation synchronization unit, a data-sensitive classification unit and a differential privacy noise addition unit that are connected in communication. The model federation synchronization unit is used to periodically perform federated aggregation of models in the cloud and all terminal nodes in the following manner: the terminal nodes calculate the gradient vector of the local model and upload the gradient vector to the cloud using homomorphic encryption; the cloud aggregates and updates the global model using the FedAvg algorithm based on the gradient vectors from all terminal nodes, and then distributes the updated model parameters of the global model to each terminal node. The terminal node updates the local model using the updated model parameters; The data sensitivity classification unit is used to automatically perform sensitivity classification processing on the user data of the user to which the media display terminal belongs, and obtain the sensitivity level of the user data; The differential privacy noise-adding unit is used to add Laplace noise to user data that is sensitive data to protect privacy, add location offset to user data that is both sensitive data and location data to protect privacy, and add counting noise to user data that is both sensitive data and behavioral statistics data to protect privacy, based on the sensitivity level of the user data.
10. The intelligent media distribution system as described in claim 9, characterized in that, The gradient vector is protected and uploaded to the cloud using homomorphic encryption, which includes: first quantizing the gradient vector into integer form, then using the SEAL homomorphic encryption library and employing the BFV homomorphic encryption scheme and key to encrypt the quantization result, and finally uploading the encrypted result to the cloud so that the cloud can perform gradient vector aggregation calculation in the encrypted domain and complete the global model update without decryption. And / or, the aggregation is performed securely by multiple participants using the following secret sharing protocol: the gradient vector is divided into multiple parts and the multiple parts are distributed one-to-one to multiple aggregation nodes; the aggregation nodes use the Shamir threshold secret sharing protocol and achieve gradient recovery through polynomial interpolation; and the identities of the participants and data integrity are verified through zero-knowledge proof.