Video efficient production method and system based on cloud edge collaboration
Through the efficient video production method based on cloud-edge collaboration, the problems of material management and personalized needs in traditional video production are solved, and efficient, personalized and collaborative management of video production are achieved.
Patent Information
- Application Number
- CN202510216708.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
In the traditional video production process, the materials are stored scatteredly and the lack of a unified management mechanism makes it time-consuming to find specific materials and difficult to meet the personalized needs of users. Especially for non-professional users, the threshold for complex processes and professional tools is high.
Efficient video production method based on cloud-edge collaboration is adopted, and video production is efficient and personalized through data fusion acquisition and processing, cloud-edge collaborative transmission and cache, task disassembly and collaborative processing, content optimization and distribution.
This method can significantly improve the efficiency and quality of video production, provide personalized creative recommendations, support large-scale team collaboration, optimize video management and distribution, and meet users' diverse needs.
Smart Images

Figure CN120075480A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video production, and particularly to an efficient video production method and system based on cloud-edge collaboration. Background Art
[0002] Video production is a process of re-editing, integrating, and arranging pictures, videos, and background music to generate a new video file. It is not only a synthesis of original materials but also a reprocessing of the original materials.
[0003] In the traditional video production process, materials are scattered and stored in different devices or storage media, lacking a unified and effective management mechanism. When specific materials need to be searched for, it often takes a lot of time and effort. At the same time, traditional video production mainly relies on manual experience and manual operations, which are difficult to meet the increasingly diverse personalized needs of users. For some non-professional users, the complex video production process and the threshold of professional tools are relatively high.
[0004] Therefore, it is necessary to propose an efficient video production method and system based on cloud-edge collaboration to solve the above problems. Summary of the Invention
[0005] The main purpose of the present invention is to provide an efficient video production method and system based on cloud-edge collaboration, which can effectively solve the problems in the background art.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0007] An efficient video production method based on cloud-edge collaboration includes the following operating steps:
[0008] S1: Data fusion acquisition and processing. Video is acquired through a data acquisition device, and at the same time, audio, light, and temperature data of the shooting location are synchronously collected and integrated.
[0009] S2: Cloud-edge collaborative transmission and caching. The acquired data is first transmitted to the nearest CDN node, and the CDN node intelligently selects the optimal path to transmit the data to the cloud according to the network conditions and cloud load.
[0010] S3: Task decomposition and collaborative processing. The overall video production process is decomposed, and appropriate creators are selected for collaborative processing according to the decomposed process. At the same time, quality control is carried out on the finished products of the creators.
[0011] S4: Content optimization and distribution. Basic behavior data of users is collected, and at the same time, a user portrait is constructed, the characteristics of users are deeply mined, and content is recommended according to the behavior preference characteristics of users.
[0012] Preferably, in the S1, the following operating steps are specifically included:
[0013] S101: Data acquisition strategy. Before shooting, use a deep learning model to classify the shooting scene, adjust the data sampling frequency according to the importance of each data source in different scenarios, continuously analyze the quality of the collected data during the data acquisition process, and when it is found that the video picture is blurred, the audio is noisy, or the environmental data is abnormal, dynamically adjust the sampling parameters in real time;
[0014] S102: Distributed acquisition. Connect multiple shooting devices on site and sensors for data acquisition into a distributed acquisition network through wireless ad hoc network technology. Each device serves as a node in the network and undertakes different data acquisition tasks according to its own location and functional characteristics. The node devices exchange the collected data and status information in real time. When a device detects a certain key event, it sends a trigger signal to other devices through the network to coordinate other devices to perform multi-angle and multi-modal data acquisition simultaneously to ensure comprehensive capture of key scenarios;
[0015] S103: Data fusion algorithm, including temporal dimension fusion, spatial dimension fusion, and feature fusion. Among them, temporal dimension fusion: Build a time series model for different modal data, and achieve accurate fusion in the temporal dimension by analyzing the change trends and correlations of the data over time;
[0016] Spatial dimension fusion: Consider the spatial location information of the data acquisition devices and perform spatial fusion on the multi-modal data collected at different locations;
[0017] Feature fusion: Use a multi-scale convolutional neural network to extract features from video, audio, and environmental data, and perform weighted fusion on the extracted features of different modalities and different scales through an attention mechanism;
[0018] S104: Perform preliminary processing based on S101 - 103. Collect domain knowledge related to the shooting scene, including video types and geographical information, and build a knowledge graph based on this. The knowledge graph contains various entities that appear in the video and the relationships between them; During the preliminary processing process, use the knowledge graph to perform semantic annotation on the collected data, and through semantic enhancement processing, provide information for subsequent video analysis and understanding;
[0019] Establish a multi-modal data quality evaluation index system, including indicators such as video clarity, color restoration, audio loudness, sound quality, and environmental data accuracy;
[0020] According to the evaluation results of the multi-modal data quality evaluation index system, automatically adjust the parameters and strategies of the preliminary processing. When the video clarity does not meet the standard, adopt a more complex image enhancement algorithm; when the audio loudness is inappropriate, perform volume adjustment and equalization processing, and record the parameters and effects of each processing at the same time. Continuously optimize the processing strategy through machine learning algorithms to improve the processing efficiency and quality.
[0021] Preferably, in S2, a multi-level caching mechanism is set on the data acquisition device and the CDN node. When the video data is requested multiple times, the CDN node caches the popular segments, and the next request can directly obtain from the cache to reduce the transmission time; the data acquisition device also caches the recently shot data for immediate calling and viewing to ensure the continuity of the shooting work when the network is unstable.
[0022] Preferably, in S3, it specifically includes the following operation steps:
[0023] S301: In-depth decomposition based on content understanding, use advanced deep learning technology to comprehensively analyze the video content. For the video, adopt an architecture combining convolutional neural network and recurrent neural network. First, extract the visual features of the video through the convolutional neural network, including scenes, objects, and human actions, and then use the recurrent neural network to analyze the temporal information of the video to judge the development and rhythm changes of the plot. For the audio, use a deep autoencoder model for feature extraction to identify the speech content, music style, and sound effect type in the video;
[0024] S302: Consider the adaptability decomposition of skills and resources, establish a creator skill and resource database, record the fields that each creator is good at, the tools mastered, and the material resources owned. When decomposing tasks, consider the factors that the creator is good at based on the creator skill and resource database to match the tasks with the creator's capabilities and resources;
[0025] S303: Intelligent matching and recommendation, build an intelligent matching and recommendation system based on big data and artificial intelligence, which is used to comprehensively consider the creator's skill level, the quality of historical works, the time record of completing tasks, and user evaluations, and accurately match the most suitable creator for each decomposed task. At the same time, according to the creator's preferences and the direction of ability improvement, recommend tasks that meet their development needs to them to improve the creator's participation and satisfaction;
[0026] S304: Collaborative processing, adopt a distributed file system and message queue technology to achieve real-time data synchronization and sharing among multiple creators. When a creator updates the video material and task progress, immediately send the update information to other relevant creators through the message queue, and at the same time use the distributed file system to ensure data consistency and availability;
[0027] The blockchain technology is introduced to record the entire process of task allocation, execution and completion. Each task has a unique identification and detailed operation record on the blockchain to ensure the transparency and traceability of the task. At the same time, the smart contract function based on the blockchain can automatically execute the delivery and payment process of the task.
[0028] S305: Quality control and feedback, establish a multi-level quality review system, including creator self-review, peer review and professional team final review;
[0029] During the task execution process, a real-time feedback mechanism is established, and creators can feedback the problems and difficulties encountered to the project managers and other relevant personnel at any time. The project managers will provide timely guidance and support. At the same time, based on the review results and user feedback, the task processing process and quality standards are iteratively optimized to continuously improve the overall work quality and efficiency.
[0030] Preferably, S303 also includes a dynamic reward mechanism for setting different reward standards according to the difficulty, urgency and importance of the task. Creators who complete tasks on time and with high quality are given additional bonuses, honorary titles and priority opportunities to undertake subsequent high-quality tasks. At the same time, a points system is introduced, and creators can obtain corresponding points when completing tasks. Points can be used to exchange learning materials, use rights of advanced tools, and participate in training activities organized by the platform, which is used to motivate creators to continuously improve their skills and work quality.
[0031] Preferably, the step S4 specifically includes the following steps:
[0032] S401: Multi-dimensional data collection and integration, collect users' basic behavior data, including browsing time, search keywords, and viewing completion, and access third-party data, including cooperation with social media platforms to obtain users' public interests and hobbies, social relationship chains; and cooperate with e-commerce platforms to understand users' consumption levels and purchased product categories, integrate the above data, and build a user data set to provide sufficient information for subsequent analysis;
[0033] S402: Construction of user portraits. The integrated data is processed by a deep neural network. Users with similar behaviors and attributes are first classified into one category through a clustering algorithm. Then, the neural network model is used to deeply mine the characteristics of each category of users.
[0034] S403: Content optimization: using collaborative filtering algorithms to recommend personalized content to each user based on the user's behavior and the preferences of similar users, and combining content-based recommendation algorithms to make recommendations based on the matching degree between the content attributes and the user's profile;
[0035] Dynamically adjust content based on real-time data feedback;
[0036] S404: Content distribution. Analyze the user characteristics and traffic patterns of different content distribution platforms, accurately deliver different types of content to appropriate platforms, and optimize the titles and tags of the content according to the platform's algorithm rules to increase the exposure rate of the content on the platform;
[0037] Through time series analysis and machine learning models, predict the traffic in different time periods and regions, and adjust the content distribution strategy in advance according to the prediction results to reasonably allocate traffic resources.
[0038] Preferably, in S403, according to the real-time data feedback, dynamically adjust the content by analyzing the user's interests, using the click-through rate CTR and the watch completion rate WCR indicators. Let the number of displays of a certain video segment be N sh ow , the number of clicks be N click , the number of views be N complete , then the formula for the click-through rate CTR is:
[0039]
[0040] The formula for the watch completion rate WCR is:
[0041]
[0042] By analyzing the CTR and WCR of different segments, the degree of user interest in each segment can be evaluated, and the content can be optimized accordingly.
[0043] An efficient video production system based on cloud-edge collaboration includes a data acquisition module, a data processing module, a caching and prefetching module, a multi-device collaborative management module, a blockchain application module, a data storage and management module, a creative recommendation module, and a quality evaluation and feedback module. The data acquisition module uses an ultra-high-definition panoramic camera and an environmental perception sensor to shoot videos and collect environmental data at the shooting site;
[0044] The data processing module is used to perform semantic segmentation and object detection on the video data collected by the data acquisition module to identify people, animals, and objects in the video and add semantic tags to them. The semantic tags are used for quick retrieval and precise editing in subsequent video production;
[0045] The caching and prefetching module, based on historical data and real-time traffic prediction, uses an intelligent caching algorithm to cache popular video segments and common materials to edge computing nodes. When multiple data acquisition modules request the same data simultaneously, it can be directly obtained from the cache to reduce the access pressure on the cloud server;
[0046] The multi-device collaborative management module is used to support the collaborative work among multiple data acquisition modules. Through a distributed task scheduling algorithm, it decomposes complex video production tasks and assigns them to different devices for parallel processing;
[0047] The blockchain application module realizes the secure sharing and trusted transactions of data based on blockchain technology. It establishes a blockchain ledger at the edge computing node, records the data contributions and processing results of each data acquisition module, and ensures the authenticity and integrity of the data;
[0048] The data storage and management module is used to store and manage the data for video production. By building a data lake, it stores and manages the data uniformly. Through the metadata management and data governance functions of the data lake, it is used to achieve data integration, cleaning, and annotation;
[0049] The creative recommendation module uses deep learning models and natural language processing technologies to provide users with creative recommendations and inspiration;
[0050] The quality assessment and feedback module is used to evaluate the quality of the video in real time during the user's creation process, including picture quality, audio quality, and content rationality, and provides feedback and improvement suggestions to the user in a timely manner to help the user improve the video production level.
[0051] Preferably, the cache and prefetch module analyzes the user's operation habits and video production process, predicts the data that may be needed in advance, and actively prefetches it from the cloud server to the edge computing node.
[0052] Preferably, the creative recommendation module specifically includes introducing knowledge graph technology to conduct in-depth correlation analysis on the collected text, image, video, and audio data;
[0053] Adopting a combination of transfer learning and reinforcement learning, transfer learning uses a model pre-trained on large-scale general data to quickly adapt to specific tasks in the video creation field, reducing the training time and data volume requirements. Reinforcement learning continuously optimizes the recommendation strategy according to the user's real-time feedback.
[0054] Compared with the prior art, the present invention provides a method and system for efficient video production based on cloud-edge collaboration, having the following beneficial effects:
[0055] 1. The method and system for efficient video production based on cloud-edge collaboration can provide users with personalized and accurate creative recommendations, can recommend unique materials according to the user's past creation styles and input themes, can effectively stimulate the user's creative inspiration, and can also provide a collaboration platform for various participants in video production, enabling easy organization of large-scale team collaboration and effectively improving the efficiency of video production.
[0056] 2. The video efficient production method and system based on cloud-edge collaboration can not only focus on the video production process, but also manage videos, including material collection, production, storage, distribution, and subsequent updates. It can record detailed data of each video, facilitating users to query and call at any time. In all aspects of video production, it can monitor and evaluate the video quality and give optimization suggestions for problems occurring in video production.
[0057] 3. The video efficient production method and system based on cloud-edge collaboration can comprehensively analyze video content through in-depth decomposition based on content understanding, and can break down tasks meticulously, making each subtask more targeted. At the same time, the decomposition strategy considering the adaptability of creators' skills and resources, relying on the creators' skills and resource database, can achieve precise matching of tasks and creators, greatly improving the efficiency and quality of task completion.
[0058] 4. The video efficient production method and system based on cloud-edge collaboration can, through the analysis of a large amount of user data, including behavioral data such as users' browsing history, search records, likes and comments, deeply understand users' preferences and needs for different types of content, and can real-time monitor information such as hot topics and popular trends on the network, helping content creators promptly capture the focus of current audiences' attention and integrating hot elements into the content, which can effectively improve the timeliness and attractiveness of video content. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is the flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] To make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with the specific embodiments.
[0061] Embodiment 1:
[0062] As Figure 1 shown, the video efficient production method based on cloud-edge collaboration includes the following operation steps:
[0063] S1: Data fusion collection and processing. Video is collected through data collection devices, and at the same time, audio, light, and temperature data at the shooting location are synchronously collected and integrated.
[0064] Specifically, it includes the following operation steps:
[0065] S101: Data acquisition strategy. Before shooting, a deep learning model is used to classify the shooting scene, and the data sampling frequency is adjusted according to the importance of each data source in different scenarios. For example, in an outdoor natural scene, environmental sound and light data are crucial for creating the atmosphere; while in an indoor stage scene, the synchronization and quality of video and audio are more important. According to the scene classification result, the data sampling frequency of various sensors is automatically adjusted. During the data acquisition process, the quality of the already acquired data is continuously analyzed. When it is found that the video picture is blurred, the audio is noisy, or the environmental data is abnormal, the sampling parameters are adjusted in real time and dynamically. When it is detected that the video picture is too dark, the sampling frequency of the light sensor is increased to obtain more accurate light information for subsequent processing; if the audio is distorted, the sampling points of the microphone are increased to improve the audio quality.
[0066] S102: Distributed acquisition. Multiple shooting devices on-site and sensors for data acquisition are connected into a distributed acquisition network through wireless ad-hoc networking technology. Each device acts as a node in the network and undertakes different data acquisition tasks according to its own location and functional characteristics. The devices located in front of the stage are mainly responsible for collecting the video and audio from the front, while the devices distributed around focus on collecting environmental sound and light data. The data and status information collected by each node device are exchanged in real time. When a device detects a certain key event, it sends a trigger signal to other devices through the network to coordinate other devices to perform multi-angle and multi-modal data acquisition simultaneously, so as to ensure the comprehensive capture of key scenes.
[0067] S103: Data fusion algorithm, including time dimension fusion, space dimension fusion, and feature fusion. Among them, time dimension fusion: A time series model is constructed for data of different modalities. By analyzing the change trends and correlations of the data over time, accurate fusion in the time dimension is achieved. In a dance performance video, the video frames of the dancer's movements are synchronously fused with the music rhythm and on-site audio data at the corresponding moments, so that the audience can feel a more real on-site atmosphere when watching the video.
[0068] Space dimension fusion: Considering the spatial position information of the data acquisition devices, multi-modal data collected from different positions are fused in space. At a large event site, by combining the video pictures of multiple cameras and the audio data of microphones distributed at different positions, a three-dimensional audio-visual scene is constructed through a spatial mapping algorithm, providing a basis for subsequent immersive experiences.
[0069] Feature fusion: A multi-scale convolutional neural network is used to extract features from video, audio, and environmental data, and the features of different modalities and different scales extracted are weighted and fused through an attention mechanism.
[0070] S104: Based on S101 - 103, perform preliminary processing to collect domain knowledge related to the shooting scene, including video type and geographical information, and construct a knowledge graph based on this. The knowledge graph contains entities that appear in various videos and the relationships between them; during the preliminary processing, use the knowledge graph to perform semantic annotation on the collected data, and through semantic enhancement processing, provide information for subsequent video analysis and understanding;
[0071] Establish a multi - modal data quality evaluation index system, including indicators such as video clarity, color restoration, audio loudness, audio quality, and the accuracy of environmental data;
[0072] According to the evaluation results of the multi - modal data quality evaluation index system, automatically adjust the parameters and strategies of the preliminary processing. When the video clarity does not meet the standard, adopt a more complex image enhancement algorithm; when the audio loudness is inappropriate, perform volume adjustment and equalization processing, and record the parameters and effects of each processing at the same time. Continuously optimize the processing strategy through machine learning algorithms to improve processing efficiency and quality.
[0073] S2: Cloud - edge collaborative transmission and caching. First, transmit the collected data to the nearest CDN node. The CDN node intelligently selects the optimal path to transmit the data to the cloud according to the network conditions and cloud load;
[0074] Set up a multi - level caching mechanism on the data acquisition device and the CDN node. When the video data is requested multiple times, the CDN node caches the popular segments, and the next request can directly obtain from the cache to reduce the transmission time; the data acquisition device also caches the recently shot data for immediate access and viewing, ensuring the continuity of the shooting work when the network is unstable.
[0075] S3: Task decomposition and collaborative processing. Decompose the overall production process of the video, select appropriate creators for collaborative processing according to the decomposed process, and at the same time perform quality control on the finished products of the creators;
[0076] Specifically, it includes the following operation steps:
[0077] S301: In-depth disassembly based on content understanding. Use advanced deep learning techniques to comprehensively analyze video content. For videos, adopt an architecture that combines a convolutional neural network and a recurrent neural network. First, use the convolutional neural network to extract visual features of the video, including scenes, objects, and human actions. Then, use the recurrent neural network to analyze the temporal information of the video to judge the development of the plot and the change of rhythm. For audio, use a deep autoencoder model for feature extraction to identify the speech content, music style, and sound effect type in the video. According to the analysis results, break down the video production tasks in detail, including editing, special effects, and subtitles. At the same time, the editing task can be broken down into multiple small clip edits according to the plot paragraphs. The special effects task is subdivided according to the special effect type and the object of action. The subtitle task can be divided into text extraction, translation, proofreading, and timeline matching.
[0078] S302: Consider the adaptation disassembly of skills and resources. Establish a database of creators' skills and resources, record the fields that each creator is good at, the tools they master, and the material resources they own. When disassembling tasks, consider the factors that the creator is good at based on the database of creators' skills and resources to match the tasks with the creator's capabilities and resources.
[0079] S303: Intelligent matching and recommendation. Build an intelligent matching and recommendation system based on big data and artificial intelligence. Comprehensively consider the creator's skill level, the quality of historical works, the time record of task completion, and user evaluations to accurately match the most suitable creator for each disassembled task. At the same time, according to the creator's preferences and the direction of ability improvement, recommend tasks that meet their development needs to them to improve the creator's participation and satisfaction. It also includes a dynamic reward mechanism, which sets different reward criteria according to the difficulty, urgency, and importance of the task. For creators who complete tasks on time and with high quality, give additional bonuses, honorary titles, and the opportunity to undertake subsequent high-quality tasks preferentially. At the same time, introduce an integral system. Creators can obtain corresponding points for completing tasks, and the points can be used to exchange for learning materials, the right to use advanced tools, and participation in training activities organized by the platform to encourage creators to continuously improve their skills and work quality.
[0080] S304: Collaborative processing. Adopt a distributed file system and message queue technology to achieve real-time data synchronization and sharing among multiple creators. When a creator updates video materials and task progress, immediately send the update information to other relevant creators through the message queue. At the same time, use the distributed file system to ensure data consistency and availability. When the editor trims the video clip, other creators responsible for special effects and subtitles can immediately obtain the updated video materials to avoid work conflicts caused by data inconsistency.
[0081] The blockchain technology is introduced to record the entire process of task allocation, execution and completion. Each task has a unique identification and detailed operation record on the blockchain to ensure the transparency and traceability of the task. At the same time, the smart contract function based on the blockchain can automatically execute the delivery and payment process of the task.
[0082] S305: Quality control and feedback. A multi-level quality review system is established, including self-review by creators, peer review, and final review by professional teams. After completing a task, creators first conduct self-review to ensure that the task meets basic requirements. Then, the task is submitted to peers for mutual review, and peers can provide opinions and suggestions from different perspectives. Finally, a professional team conducts a final review to comprehensively evaluate and control the quality of the task.
[0083] During the task execution process, a real-time feedback mechanism is established, and creators can feedback the problems and difficulties encountered to the project managers and other relevant personnel at any time. The project managers will provide timely guidance and support. At the same time, based on the review results and user feedback, the task processing process and quality standards are iteratively optimized to continuously improve the overall work quality and efficiency.
[0084] S4: Content optimization and distribution: collect basic user behavior data, build user portraits, conduct in-depth mining of user characteristics, and recommend content based on user behavior and preference characteristics;
[0085] The specific steps include the following:
[0086] S401: Multi-dimensional data collection and integration, collect users' basic behavior data, including browsing time, search keywords, and viewing completion, and access third-party data, including cooperation with social media platforms to obtain users' public interests and hobbies, social relationship chains; and cooperate with e-commerce platforms to understand users' consumption levels and purchased product categories, integrate the above data, and build a user data set to provide sufficient information for subsequent analysis;
[0087] S402: Construction of user portraits. The integrated data is processed by a deep neural network. Users with similar behaviors and attributes are first classified into one category through a clustering algorithm. Then, the neural network model is used to deeply mine the characteristics of each category of users, and the preferences of a certain category of users for specific subject content in a specific time period and on a specific device are analyzed, with the accuracy down to the specific style, actor and other details, to establish a comprehensive user portrait.
[0088] S403: Content optimization: using collaborative filtering algorithms to recommend personalized content to each user based on the user's behavior and the preferences of similar users, and combining content-based recommendation algorithms to make recommendations based on the matching degree between the content attributes and the user's profile;
[0089] Dynamically adjust the content based on real-time data feedback. During video playback, analyze the user's real-time viewing behavior, such as the time points of fast-forwarding, pausing, and repeated viewing, to understand the user's interest level in different segments. When it is found that most users frequently fast-forward in a certain segment, consider streamlining or optimizing this segment, such as shortening the duration or adding special effects, to improve the attractiveness of the content;
[0090] Dynamically adjust the content based on real-time data feedback. Specifically, analyze the user's interests, and use the click-through rate (CTR) and watch completion rate (WCR) metrics. Assume the number of impressions of a certain video segment is N sh ow , the number of clicks is N click , and the number of views is N complete , then the formula for the click-through rate (CTR) is:
[0091]
[0092] The formula for the watch completion rate (WCR) is:
[0093]
[0094] By analyzing the CTR and WCR of different segments, the user's interest level in each segment can be evaluated, and then the content can be optimized accordingly;
[0095] S404: Content distribution. Analyze the user characteristics and traffic patterns of different content distribution platforms. Users of short video platforms tend to prefer easy, interesting, and fast-paced content; users of long video platforms have a higher demand for in-depth and high-quality content. According to these characteristics, accurately deliver different types of content to the appropriate platforms, and at the same time, optimize the title and tags of the content according to the platform's algorithm rules to improve the exposure rate of the content on the platform;
[0096] Through time series analysis and machine learning models, predict the traffic in different time periods and regions. According to the prediction results, adjust the content distribution strategy in advance, reasonably allocate traffic resources. When a large event is about to be held in a certain region, predict that the traffic of relevant content in this region will increase significantly, and then push relevant high-quality content in advance to meet the user's needs and avoid service instability caused by traffic overload.
[0097] Embodiment 2:
[0098] An efficient video production system based on cloud-edge collaboration includes a data acquisition module, a data processing module, a caching and prefetching module, a multi-device collaborative management module, a blockchain application module, a data storage and management module, a creative recommendation module, and a quality assessment and feedback module. The data acquisition module uses an ultra-high-definition panoramic camera and an environmental perception sensor to shoot videos and collect environmental data at the shooting site;
[0099] The data processing module is used to perform semantic segmentation and object detection on the video data collected by the data acquisition module, to identify the people, animals, and objects in the video, and to add semantic tags to them. The semantic tags are used for quick retrieval and precise editing in subsequent video production;
[0100] The cache and prefetch module, based on historical data and real-time traffic prediction, adopts an intelligent caching algorithm to cache popular video segments and commonly used materials to the edge computing nodes. When multiple data acquisition modules request the same data simultaneously, they can directly obtain it from the cache, which is used to reduce the access pressure on the cloud server. The cache and prefetch module analyzes the user's operation habits and video production process, predicts the data that may be needed in advance, and actively prefetches it from the cloud server to the edge computing nodes;
[0101] The multi-device collaborative management module is used to support the collaborative work between multiple data acquisition modules. Through a distributed task scheduling algorithm, it decomposes complex video production tasks and assigns them to different devices for parallel processing;
[0102] The blockchain application module realizes the secure sharing and trusted transactions of data based on blockchain technology. It establishes a blockchain ledger at the edge computing nodes to record the data contributions and processing results of each data acquisition module, ensuring the authenticity and integrity of the data;
[0103] The data storage and management module is used to store and manage the data for video production. By constructing a data lake, it stores and manages the data uniformly. Through the metadata management and data governance functions of the data lake, it is used to realize the integration, cleaning, and annotation of the data;
[0104] The creative recommendation module uses deep learning models and natural language processing technologies to provide users with creative recommendations and inspiration. The creative recommendation module specifically includes introducing knowledge graph technology to perform in-depth correlation analysis on the collected text, image, video, and audio data. When dealing with travel-related creation needs, it correlates destination attractions, local special foods, and folk culture information with each other through the knowledge graph to construct a comprehensive travel knowledge system. When the user inputs a travel-related theme, the creative recommendation module can provide richer and logically related creative recommendations based on the knowledge graph;
[0105] Adopt a combination of transfer learning and reinforcement learning. Transfer learning utilizes a model pre-trained on large-scale general data to quickly adapt to specific tasks in the field of video creation, aiming to reduce the training time and data volume requirements. Reinforcement learning continuously optimizes the recommendation strategy based on the user's real-time feedback. When the user frequently uses a certain creative element recommended to complete a video work and obtains a high evaluation, the recommendation weight for similar creative elements is enhanced; when a certain recommendation is not adopted by the user, the recommendation strategy is adjusted to reduce the frequency of similar video recommendations.
[0106] The quality assessment and feedback module is used to evaluate the quality of the video in real time during the user's creation process, including picture quality, audio quality, and content rationality, and provide feedback and improvement suggestions to the user in a timely manner to help the user improve the video production level.
[0107] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An efficient video production method based on cloud-edge collaboration, characterized by: The steps include: S1: Data fusion acquisition and processing: Video is collected through data acquisition equipment, while audio, lighting and temperature data of the shooting location are collected synchronously, and these data are integrated; S2: Cloud-edge collaborative transmission and caching: the collected data is first transmitted to the nearest CDN node. The CDN node intelligently selects the optimal path to transmit the data to the cloud based on the network status and cloud load. S3: Task decomposition and collaborative processing: decomposing the overall video production process, selecting appropriate creators for collaborative processing based on the decomposition process, and performing quality control on the creators’ finished products; S4: Content optimization and distribution, collect users' basic behavior data, build user portraits, conduct in-depth mining of user characteristics, and recommend content based on user behavior and preference characteristics.
2. The efficient video production method based on cloud-edge collaboration according to claim 1 is characterized in that: The S1 specifically includes the following steps: S101: Data collection strategy: Before shooting, use a deep learning model to classify the shooting scenes, adjust the data sampling frequency according to the importance of each data source in different scenes, and continuously analyze the quality of the collected data during the data collection process. If the video is blurred, the audio is noisy, or the environmental data is abnormal, the sampling parameters are adjusted dynamically in real time. S102: Distributed collection: multiple shooting devices and sensors for collecting data on site are connected into a distributed collection network through wireless ad hoc networking technology. Each device acts as a node in the network and undertakes different data collection tasks according to its own location and functional characteristics. The node devices exchange the collected data and status information in real time. When a device finds a key event, it sends a trigger signal to other devices through the network, and cooperates with other devices to collect data from multiple angles and modes at the same time to ensure that key scenes are fully captured. S103: Data fusion algorithm, including time dimension fusion, space dimension fusion, and feature fusion. Time dimension fusion: build a time series model for data of different modalities, and achieve accurate fusion of the time dimension by analyzing the change trend and correlation of the data over time; Spatial dimension fusion: Considering the spatial location information of the data acquisition device, the multimodal data collected at different locations are spatially fused; Feature fusion: Use multi-scale convolutional neural networks to extract features from video, audio, and environmental data, and perform weighted fusion of the extracted features of different modalities and scales through an attention mechanism; S104: Perform preliminary processing based on S101-103 to collect domain knowledge related to the shooting scene, including video type and geographic information, and build a knowledge graph based on this. The knowledge graph contains entities appearing in various videos and the relationships between them; In the initial processing, the collected data is semantically annotated using knowledge graphs and processed through semantic enhancement to provide information for subsequent video analysis and understanding; Establish a multimodal data quality assessment indicator system, including indicators of video clarity, color reproduction, audio loudness, sound quality, and accuracy of environmental data; According to the evaluation results of the multimodal data quality assessment index system, the parameters and strategies of the initial processing are automatically adjusted. When the video clarity does not meet the standard, a more complex image enhancement algorithm is used. When the audio loudness is inappropriate, volume adjustment and equalization processing are performed. At the same time, the parameters and effects of each processing are recorded, and the processing strategy is continuously optimized through machine learning algorithms to improve processing efficiency and quality.
3. The efficient video production method based on cloud-edge collaboration according to claim 1 is characterized in that: In S2, a multi-level cache mechanism is set on the data acquisition device and the CDN node. When the video data is requested multiple times, the CDN node caches the popular segments, and the next request can be directly obtained from the cache to reduce the transmission time; The data acquisition device also caches recently captured data for easy access and viewing at any time, ensuring the continuity of shooting when the network is unstable.
4. The method for efficient video production based on cloud-edge collaboration according to claim 1 is characterized in that: The S3 specifically includes the following steps: S301: Deep analysis based on content understanding uses advanced deep learning technology to conduct a comprehensive analysis of video content. For video, a convolutional neural network and a recurrent neural network are used to extract visual features of the video, including scenes, objects, and character actions. A recurrent neural network is then used to analyze the timing information of the video to determine the development of the plot and rhythm changes. For audio, a deep autoencoder model is used to extract features and identify the voice content, music style, and sound effect type in the video. S302: Consider the adaptability of skills and resources, establish a creator skills and resources database, record the areas of expertise, tools mastered, and material resources owned by each creator, and when decomposing tasks, consider the factors that the creator is good at based on the creator skills and resources database to match the tasks with the creator's abilities and resources; S303: Intelligent matching and recommendation: building an intelligent matching and recommendation system based on big data and artificial intelligence, which comprehensively considers the skill level, quality of historical works, time record of completing tasks, and user evaluation of creators, and accurately matches the most suitable creator for each disassembled task. At the same time, based on the creators' preferences and ability improvement directions, tasks that meet their development needs are recommended to them, so as to improve the participation and satisfaction of creators; S304: Collaborative processing, using distributed file systems and message queue technology to achieve real-time data synchronization and sharing among multiple creators. When a creator updates video materials or task progress, the updated information is immediately sent to other related creators through the message queue. At the same time, the distributed file system is used to ensure data consistency and availability. The blockchain technology is introduced to record the entire process of task allocation, execution and completion. Each task has a unique identification and detailed operation record on the blockchain to ensure the transparency and traceability of the task. At the same time, the smart contract function based on the blockchain can automatically execute the delivery and payment process of the task. S305: Quality control and feedback, establish a multi-level quality review system, including creator self-review, peer review and professional team final review; During the task execution process, a real-time feedback mechanism is established, and creators can feedback the problems and difficulties encountered to the project managers and other relevant personnel at any time. The project managers will provide timely guidance and support. At the same time, based on the review results and user feedback, the task processing process and quality standards are iteratively optimized to continuously improve the overall work quality and efficiency.
5. The method for efficient video production based on cloud-edge collaboration according to claim 4 is characterized in that: The S303 also includes a dynamic reward mechanism for setting different reward standards based on the difficulty, urgency and importance of the task. Creators who complete tasks on time and with high quality are given additional bonuses, honorary titles and priority opportunities to undertake subsequent high-quality tasks. At the same time, a points system is introduced, whereby creators can obtain corresponding points for completing tasks. Points can be used to redeem learning materials, access to advanced tools, and participation in training activities organized by the platform, to motivate creators to continuously improve their skills and work quality.
6. The efficient video production method based on cloud-edge collaboration according to claim 1 is characterized in that: The S4 specifically includes the following steps: S401: Multi-dimensional data collection and integration, collecting basic user behavior data, including browsing time, search keywords, and viewing completion, while accessing third-party data, including cooperation with social media platforms to obtain users' public interests and hobbies, and social relationship chains; Cooperate with e-commerce platforms to understand users' consumption levels and the categories of goods they purchase, integrate the above data, and build a user data set to provide sufficient information for subsequent analysis; S402: Construction of user portraits. The integrated data is processed by a deep neural network. Users with similar behaviors and attributes are first classified into one category through a clustering algorithm. Then, the neural network model is used to deeply mine the characteristics of each category of users. S403: Content optimization: using collaborative filtering algorithms to recommend personalized content to each user based on the user's behavior and the preferences of similar users, and combining content-based recommendation algorithms to make recommendations based on the matching degree between the content attributes and the user's profile; Dynamically adjust content based on real-time data feedback; S404: Content distribution: analyzing the user characteristics and traffic patterns of different content distribution platforms, accurately delivering different types of content to appropriate platforms, and optimizing the titles and tags of content according to the platform algorithm rules to increase the exposure of the content on the platform; Through time series analysis and machine learning models, traffic in different time periods and regions can be predicted. Based on the prediction results, content distribution strategies can be adjusted in advance and traffic resources can be allocated reasonably.
7. The method for efficient video production based on cloud-edge collaboration according to claim 6 is characterized in that: In S403, the content is dynamically adjusted according to the real-time data feedback. Specifically, the user's interests are analyzed, and the click-through rate CTR and the viewing completion rate WCR indicators are used. Suppose the number of times a certain video clip is displayed is N. show , the number of clicks is N click , the number of views is N complete , then the formula for click-through rate CTR is: The formula for viewing completion rate WCR is: By analyzing the CTR and WCR of different segments, we can evaluate the user's interest in each segment and optimize the content accordingly.
8. An efficient video production system based on cloud-edge collaboration, adopting the efficient video production system based on cloud-edge collaboration as described in any one of claims 1-7, including a data acquisition module, a data processing module, a cache and pre-fetching module, a multi-device collaborative management module, a blockchain application module, a data storage and management module, a creative recommendation module, and a quality assessment and feedback module, characterized in that: The data acquisition module captures video and collects environmental data of the shooting scene based on an ultra-high-definition panoramic camera and an environmental perception sensor; The data processing module is used to perform semantic segmentation and target detection on the video data collected by the data acquisition module, to identify people, animals, and objects in the video, and to add semantic tags to them. The semantic tags are used for rapid retrieval and accurate editing in subsequent video production; The cache and pre-fetch module uses an intelligent cache algorithm based on historical data and real-time traffic prediction to cache popular video clips and commonly used materials to edge computing nodes. When multiple data acquisition modules request the same data at the same time, they can directly obtain it from the cache to reduce the access pressure on the cloud server; The multi-device collaborative management module is used to support the collaborative work between multiple data acquisition modules, and decomposes and distributes complex video production tasks to different devices for parallel processing through a distributed task scheduling algorithm; The blockchain application module realizes secure data sharing and trusted transactions based on blockchain technology, establishes a blockchain account book on the edge computing node, records the data contribution and processing results of each data collection module, and ensures the authenticity and integrity of the data; The data storage and management module is used to store and manage the data of video production. By building a data lake, the data is uniformly stored and managed. The metadata management and data governance functions of the data lake are used to achieve data integration, cleaning and annotation. The creative recommendation module uses deep learning models and natural language processing technology to provide users with creative recommendations and inspiration; The quality assessment and feedback module is used to evaluate the quality of the video in real time during the user's creation process, including picture quality, audio quality, and content rationality, and to provide users with feedback and improvement suggestions in a timely manner to help users improve their video production level.
9. The efficient video production system based on cloud-edge collaboration according to claim 8 is characterized in that: The cache and pre-fetch module predicts the data that may be needed in advance by analyzing the user's operating habits and video production process, and actively pre-fetches the data from the cloud server to the edge computing node.
10. The efficient video production system based on cloud-edge collaboration according to claim 8 is characterized in that: The creative recommendation module specifically includes introducing knowledge graph technology to perform in-depth correlation analysis on the collected text, image, video, and audio data; It adopts a combination of transfer learning and reinforcement learning. Transfer learning uses models pre-trained on large-scale general data to quickly adapt to specific tasks in the field of video creation, thereby reducing training time and data requirements. Reinforcement learning continuously optimizes recommendation strategies based on real-time feedback from users.
Citation Information
Patent Citations
Multi-source multi-protocol video fusion transmission switching system and method
CN110519641A
Video data processing method, device, system and equipment based on cloud edge collaboration
CN113783944A
Cloud-edge collaborative video processing method and device, medium and equipment
CN115174848A
Real-time video intelligent processing method and system based on cloud edge collaboration
CN116310935A
End-side cloud collaboration method and system based on visual scene perception
CN117319402A
Cited By
Multi-terminal-oriented distributed audio and video adaptive production system
CN120301991A
Distributed audio and video adaptive production system for multiple terminals
CN120301991B
Cross-platform remote networking method and system
CN120710887A
Video production control method and system based on cloud platform
CN120916024A
A video production control method and system based on a cloud platform
CN120916024B