Cross-platform convergence media intelligent creation method, server, medium and product
The server obtains video resources in the integrated media center, builds video association maps and learns user editing habits, solves the problem of inefficient in integrated media content creation, and realizes efficient unified recall and intelligent combination of cross-platform materials.
Patent Information
- Application Number
- CN202510620023.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-22
AI Technical Summary
In the process of creating existing integrated media content, creators need to switch and download materials between multiple platforms, resulting in inefficiency and spend a lot of time on material combination and editing adjustments.
Obtain user login information through the server, authenticate access to video resources in multiple media centers, extract the content theme and emotional characteristics of the video, build a video association map, and learn user editing operation habits, perform intelligent optimization, and generate creative combination solutions.
It realizes unified retrieval of cross-platform video materials, improves the efficiency of integrated media content creation, reduces the repetitive work of material acquisition and editing adjustment, and generates creative solutions for content coherence and emotional continuity.
Smart Images

Figure CN120358393A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to a cross-platform media convergence intelligent creation method, server, medium, and product. Background Art
[0002] With the rapid development of Internet technology and the rise of self-media platforms, the demand for the creation of media convergence content such as short videos and live broadcasts has increased sharply. Creators need to publish content on multiple platforms to expand their influence, which promotes the development of media convergence content creation towards specialization and high efficiency.
[0003] Currently, the creation of media convergence content is mainly carried out through video editing software. Creators download materials from various media platforms and import the materials into the video editing software, and select and splice video clips through the editing tools provided by the video editing software.
[0004] Due to the scattered sources of video materials, creators need to switch between multiple media platforms and download materials, which consumes a lot of time. When creating videos, creators often need to repeatedly adjust and preview to find a suitable combination of materials. This purely manual exploration creation process is inefficient. Summary of the Invention
[0005] This application provides a cross-platform media convergence intelligent creation method, server, medium, and product for improving the efficiency of video creation.
[0006] In a first aspect, the present application provides a cross-platform integrated media intelligent creation method, which is applied to a server. The method includes: obtaining user login information, and sending a user authentication request to one or more integrated media centers based on the user login information, where the user authentication request is used to retrieve accessible videos from the integrated media centers; receiving the accessible videos returned by the integrated media centers, and determining one or more videos to be edited from the accessible videos; extracting the content theme and emotional features of the videos to be edited to generate video feature vectors; constructing a video association graph according to the video feature vectors, where the video association graph is used to represent the content association degree and emotional association degree between different videos to be edited; performing combined matching on the videos to be edited based on the video association graph, and outputting multiple creative combination schemes, where the creative combination schemes are used to represent the arrangement order, connection method, and visual settings of the videos to be edited; in response to a selection instruction of the user for a target creative combination scheme, obtaining the user's editing operation sequence, where the target creative combination scheme is any one of the creative combination schemes, and the editing operation sequence is used to represent the adjustment process of the user for the target creative combination scheme; extracting the user's creative intention features according to the editing operation sequence, where the creative intention features include video editing features, transition operation features, and audio-visual configuration features; performing intelligent optimization on the target creative combination scheme based on the creative intention features to generate a media file, and the intelligent optimization includes adjusting video connection points, optimizing transition special effects, and audio-visual matching.
[0007] By adopting the above technical solution, the server obtains the user login information and authenticates and accesses the video resources of multiple integrated media centers, realizing the unified retrieval of cross-platform video materials. The server extracts the content theme and emotional features of each video to be edited, determines the content association degree and emotional association degree between different videos to be edited, and constructs a video association graph to provide a basis for subsequent intelligent combination. At the same time, the server will learn the user's editing operation habits, extract the user's creative intention features in aspects such as video editing, transition operation, and audio-visual configuration, and perform intelligent optimization on the target creative combination scheme accordingly. This method significantly improves the efficiency of integrated media content creation and reduces the repetitive work of creators in material acquisition and editing adjustment.
[0008] Combined with some embodiments of the first aspect, in some embodiments, extracting the content theme and emotional features of the videos to be edited to generate video feature vectors specifically includes: dividing the videos to be edited into multiple video segments based on a preset video segment length; processing the video segments through a preset feature extraction network to obtain multi-modal features including picture features and audio features; identifying the content theme of the video segments based on the multi-modal features by using a theme classification model; determining the emotional features of the video segments by using an emotional recognition model based on the multi-modal features; and performing feature fusion on the content theme and emotional features to generate video feature vectors.
[0009] By adopting the above technical solution, the server divides the video to be edited into multiple video segments based on a preset video segment length, and uses a preset feature extraction network to determine the picture features and audio features of the video segments, so as to analyze the video content more precisely and comprehensively. The server uses a theme classification model and an emotion recognition model to identify the content theme and emotion features of the video segments respectively, and then fuses these features to generate a video feature vector. This feature extraction method based on deep learning can accurately capture the content and emotion of the video to be edited, and provide a more reliable feature representation for subsequent video correlation analysis and intelligent combination.
[0010] Combined with some embodiments of the first aspect, in some embodiments, according to the video feature vector, a video correlation graph is constructed. The video correlation graph is used to represent the content correlation degree and emotion correlation degree between different videos to be edited, and specifically includes: calculating the cosine similarity of the content themes of any two videos to be edited to construct a content correlation degree matrix of the videos to be edited; calculating the emotion distance of the emotion features of any two videos to be edited to construct an emotion correlation degree matrix of the videos to be edited; performing weighted fusion on the content correlation degree matrix and the emotion correlation degree matrix to obtain a comprehensive correlation degree matrix; and determining the video correlation graph based on the comprehensive correlation degree matrix.
[0011] By adopting the above technical solution, the server calculates the cosine similarity of the content themes of any two videos to be edited and the emotion distance of the emotion features, and constructs a content correlation degree matrix and an emotion correlation degree matrix respectively. Then, the server performs weighted fusion on the two matrices to obtain a comprehensive correlation degree matrix, and finally determines the video correlation graph. This multi-dimensional correlation analysis method not only considers the similarity of videos at the content level, but also pays attention to the correlation at the emotion level, enabling the server to more comprehensively understand the correlation relationship between different videos to be edited. Based on the video correlation graph for video combination matching, creative solutions that are more reasonable in terms of content coherence and emotion continuity can be generated, improving the quality of the automatically generated solutions.
[0012] In some embodiments in combination with some embodiments of the first aspect, according to the editing operation sequence, the creative intention features of the user are extracted. The creative intention features include video editing features, transition operation features, and audio-visual configuration features, specifically including: based on the cut point setting operations in the editing operation sequence, calculating the duration interval and picture change between adjacent cut points to construct the user's video editing features; according to the transition effect configuration operations in the editing operation sequence, counting the usage patterns of transition effects under different combinations of scene types to determine the user's transition operation features, where the usage patterns of transition effects include the probability of selecting the effect type and the configuration of effect parameters; according to the audio adjustment operations in the editing operation sequence, extracting the audio adjustment timing features and the audio-visual synchronization features to determine the user's audio-visual configuration features, and the audio-visual synchronization features are used to characterize the time correlation between the picture switching and the audio change.
[0013] By adopting the above technical solutions, the server analyzes the user's editing operation sequence, extracts multi-dimensional creative intention features from it to deeply understand the user's creative habits and preferences, and then provides more personalized creation suggestions that better meet the user's personalized needs when intelligently optimizing the target creative combination plan in the follow-up, significantly improving the personalization degree of the creation effect.
[0014] In some embodiments in combination with some embodiments of the first aspect, after the step of obtaining the user login information and sending a user authentication request to one or more media convergence centers based on the user login information, the method further includes: obtaining the video resource address sent by the media convergence center; based on the preset web page parsing rules, extracting video information from the video resource address to obtain video metadata including the video title, video description, video duration, release time, and video type; storing the video metadata in the video resource index database.
[0015] By adopting the above technical solutions, after the user authentication, the server obtains the video resource address sent by the media convergence center, and based on the preset web page parsing rules, extracts video information from the video resource address to obtain video metadata. The server stores the video metadata in the video resource index database to establish a structured video resource management system, which facilitates the user to quickly retrieve and filter the required materials.
[0016] In some embodiments in combination with some embodiments of the first aspect, after the step of receiving the accessible video returned by the media convergence center and determining one or more videos to be edited from the accessible video, the method further includes: in response to the user's selection instruction for the target video template, obtaining the target video template; receiving the video title and text description input by the user; based on the preset configuration of the target video template, performing style rendering on the video title and text description to obtain the output video.
[0017] By adopting the above technical solutions, users can freely select a target video template and input a video title and a text description, and the server automatically completes style rendering based on the preset configuration of the target video template. This templated creation method greatly simplifies the user's creation process and is especially suitable for quickly creating a series of video content with a unified style. At the same time, the preset configuration of the target video template ensures the professionalism and consistency of the created finished product, enabling users lacking professional design experience to quickly create beautiful and standardized video works.
[0018] In combination with some embodiments of the first aspect, in some embodiments, after the step of intelligently optimizing a target creative combination plan based on creative intention features to generate a media file, the method further includes: obtaining video specification parameters of multiple publishing platforms, where the video specification parameters include video resolution, aspect ratio, and bitrate requirements; and adjusting the media file based on the video specification parameters to generate a final media file adapted to different publishing platforms.
[0019] By adopting the above technical solutions, the server obtains the video specification parameters of different publishing platforms and automatically adjusts the generated media file accordingly, realizing the function of adapting to multiple platforms with one creation. This intelligent adaptation mechanism eliminates the cumbersome work of users manually adjusting parameters for different platforms, ensures that the video can obtain the best playback effect on each platform, and improves the efficiency of multi-platform content publishing.
[0020] In a second aspect, an embodiment of the present application provides a server, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the server to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0021] In a third aspect, an embodiment of the present application provides a computer program product containing instructions, which, when the computer program product runs on a server, causes the server to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions, which, when the instructions run on a server, cause the server to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0023] Understandably, the server provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the method provided in the embodiments of the present application. Therefore, the beneficial effects they can achieve can refer to the beneficial effects in the corresponding method, which will not be elaborated here.
[0024] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By adopting the above technical solution, the server obtains the user login information and authenticates the access to the video resources of multiple media convergence centers, realizing the unified retrieval of cross-platform video materials. The server extracts the content theme and emotional characteristics of each video to be edited, determines the content correlation degree and emotional correlation degree between different videos to be edited, and constructs a video correlation graph, providing a basis for subsequent intelligent combination. At the same time, the server will learn the user's editing operation habits, extract the creative intention characteristics of the user in video editing, transition operation, audio-visual configuration, etc., and intelligently optimize the target creative combination plan accordingly. This method significantly improves the efficiency of media convergence content creation and reduces the repetitive work of creators in material acquisition and editing adjustment.
[0025] 2. By adopting the above technical solution, the server divides the video to be edited into multiple video segments based on a preset video segment length, and uses a preset feature extraction network to determine the video frame features and audio features of the video segments, so as to analyze the video content more finely and comprehensively. The server uses a theme classification model and an emotion recognition model to identify the content theme and emotional characteristics of the video segments respectively, and then fuses these features to generate a video feature vector. This feature extraction method based on deep learning can accurately capture the content and emotion of the video to be edited, providing a more reliable feature representation for subsequent video correlation analysis and intelligent combination.
[0026] 3. By adopting the above technical solution, the server calculates the cosine similarity of the content themes and the emotional distance of the emotional characteristics of any two videos to be edited, and constructs a content correlation matrix and an emotional correlation matrix respectively. Then, the server performs weighted fusion on the two matrices to obtain a comprehensive correlation matrix, and finally determines the video correlation graph. This multi-dimensional correlation analysis method not only considers the similarity of videos at the content level, but also pays attention to the relevance at the emotional level, enabling the server to more comprehensively understand the correlation relationship between different videos to be edited. Based on the video correlation graph for video combination matching, a creative plan that is more reasonable in terms of content coherence and emotional continuity can be generated, improving the quality of the automatically generated plan. Description of the Drawings
[0027] Figure 1 is a flowchart of a cross-platform media convergence intelligent creation method in the embodiments of the present application; Figure 2 It is another process schematic diagram of the cross-platform media integration intelligent creation method in the embodiments of the present application; Figure 3 It is a schematic structural diagram of an entity device of the server in the embodiments of the present application. Detailed implementation manners
[0028] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application, the singular forms "a", "an", "above-mentioned", "the", and "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations including one or more of the listed items.
[0029] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.
[0030] Next, in combination with the above scenarios, the process of the method provided in this embodiment will be described. Please refer to Figure 1 , which is a process schematic diagram of the cross-platform media integration intelligent creation method in the embodiments of the present application.
[0031] S101. Obtain user login information, and send a user authentication request to one or more media integration centers based on the user login information. The user authentication request is used to retrieve accessible videos from the media integration centers; Among them, the user login information refers to the identity verification information of the user on the media integration creation platform, including username, password, token, etc.; the media integration center refers to a content platform with video resources, such as major video websites, news platforms, etc.; the user authentication request is used to verify the user's access permission and obtain authorization from the media integration center; the accessible video refers to the video content to which the user has legal access rights.
[0032] Specifically, when a user needs to create cross-platform integrated media content, the server first receives the user login information, then packages the user login information into a user authentication request and sends it to multiple integrated media centers specified by the user. Each user authentication request contains necessary user identity information and access permission declarations. The server communicates securely with each integrated media center through a unified authentication interface to ensure that the user can legally access the video resources of these integrated media centers. This centralized authentication method avoids the trouble of the user having to log in to multiple integrated media centers separately.
[0033] S102. Receive the accessible videos returned by the integrated media center, and determine one or more videos to be edited from the accessible videos; Among them, the accessible video refers to the video content that the user has the right to access after authentication; the video to be edited refers to the video material selected from the accessible videos that needs to be creatively processed.
[0034] Specifically, the server receives the accessible videos returned by each integrated media center. The server uniformly organizes and displays these accessible videos to facilitate user browsing and selection. The user can select one or more of them as the videos to be edited according to the creation requirements. The server will record the user's selection and prepare for subsequent editing processing. This centralized display and selection method enables the user to more efficiently manage the video resources from different integrated media centers.
[0035] S103. Extract the content theme and emotional characteristics of the video to be edited to generate a video feature vector; Among them, the content theme is used to represent the core content category of the video to be edited, such as news, entertainment, education, etc.; the emotional characteristic is used to represent the emotional tendency conveyed by the video to be edited, such as positive, negative, etc.; the video feature vector is used to represent the content characteristics and emotional characteristics of the video to be edited in numerical form.
[0036] Specifically, the server deeply analyzes each video to be edited: first, the video to be edited is segmented into video clips of appropriate length, and then a pre-trained deep learning model is used to extract the visual features and audio features of each video clip. The server identifies the category attributes of the video clips through a theme classification model and analyzes the emotional color of the video clips through an emotion recognition model. The server integrates the content theme and emotional characteristics of the video to be edited into a multi-dimensional feature vector to represent the semantic information of the video to be edited. The feature extraction process uses a multi-modal analysis method, considering both the picture content and the sound features, so that the server's understanding of the video content is more comprehensive.
[0037] Suppose the server analyzes a 3-minute food cooking video: (1) Video segmentation processing (divided into 12 segments every 15 seconds, each segment representing a different cooking stage): Introduction of ingredients at the beginning (15 seconds); Preparation and cutting process (30 seconds); Actual cooking process (75 seconds); Plating and garnishing (45 seconds); Presentation of the finished product (15 seconds); (2) Multi-modal feature extraction: Visual features: Ingredient features: Prawns, green peppers, minced garlic, etc.; Scene features: Kitchen environment, stove, etc.; Action features: Actions such as cutting vegetables and stir-frying; Audio features: Background music: Lively background score; Ambient sounds: Sound of cutting vegetables, sound of stir-frying; Commentary sound: Step-by-step explanation; (3) Theme classification results: Main category: Food making (95% confidence); Sub-category: Seafood cuisine (88% confidence); Scene category: Home cooking (82% confidence); (4) Emotional feature analysis: Beginning: Sense of anticipation (when introducing ingredients); Middle: Sense of concentration (during the cooking process); End: Sense of pleasure (when presenting the finished product); (5) Feature fusion results: Finally, a multi-dimensional video feature vector is generated, including: Content dimension: Probability distribution of food, tutorial, and lifestyle categories; Emotional dimension: Intensity distribution of various emotions; Visual dimension: Feature distribution of visual elements; Sound dimension: Distribution of audio features.
[0038] Optionally, generally, to extract the content theme and emotional features of the video to be edited to generate a video feature vector can be achieved in the following ways, which are not limited here: Based on the preset video segment length, divide the video to be edited into multiple video segments; Process the video segments through a preset feature extraction network to obtain multi-modal features including visual features and audio features; Based on the multi-modal features, use a theme classification model to identify the content theme of the video segments; Based on the multi-modal features, use an emotion recognition model to determine the emotional features of the video segments; Fuse the content theme and emotional features to generate a video feature vector.
[0039] The steps to build a theme classification model based on deep learning are as follows: First, the server collects multiple videos and their corresponding content theme tags. The way to obtain the videos is not limited here. The content theme tags corresponding to the videos are usually manually labeled. The server stores the multiple videos and their corresponding content theme tags in the dataset D, and the format of each piece of data is (video, content theme tag). Among them, the video is the input feature for model training, and the content theme tag is the output feature for model training.
[0040] Then, the server constructs a recurrent neural network based on LSTM, which includes an input layer, two LSTM hidden layers, a fully connected layer, and an output layer. The input layer inputs the videos. The output layer outputs the content theme tags.
[0041] Next, the server uses the Adam optimizer, sets the learning rate to 0.001, and the training batch size to 32. It can also be set according to the actual situation and is not limited here. 80% of the data in the historical data is divided into the training set, and 20% is divided into the validation set. Train for 100 epochs, and save the model with the highest accuracy on the validation set. It can also be set according to the actual situation and is not limited here. An epoch is the process of the entire training dataset passing through the neural network once. In machine learning and deep learning, an epoch is a unit used to measure the number of times the entire training set is repeatedly learned. Specifically, when the neural network completes a forward calculation and a backward propagation process, that is, all data has been processed by the network once, this completes one epoch. The server uses binary cross-entropy as the loss function and adopts Early Stopping to prevent overfitting. When the value of the loss function exceeds the preset function threshold, it is determined that the model training is completed, and the theme classification model is obtained. Early Stopping is a technique in deep learning and machine learning to prevent model overfitting. It decides when to stop training by monitoring the performance of the model on the validation set.
[0042] Finally, the server inputs the input features in the validation set into the theme classification model, then obtains the predicted output of the theme classification model, compares the predicted output of the theme classification model with the actual output features in the validation set, and uses some performance metrics such as accuracy, precision, recall, F1 score, mean squared error (MSE), etc. to evaluate the performance of the theme classification model. According to the performance of the theme classification model on the validation set, adjust the parameters of the theme classification model, including adjusting the learning rate, changing the model complexity (such as increasing or decreasing the number of layers or nodes in the neural network), modifying the regularization strength, etc. This process may require multiple iterations, and each time it is adjusted based on the previous learning results to optimize the theme classification model.
[0043] The method for constructing the emotion recognition model is similar and will not be elaborated here.
[0044] S104. Construct a video association graph based on the video feature vectors, where the video association graph is used to represent the content association degree and emotion association degree between different videos to be edited; Among them, the video association graph is used to represent the content association degree and emotion association degree between different videos to be edited; the content association degree is used to quantify the similarity degree of different videos to be edited in terms of content themes; the emotion association degree is used to measure the consistency degree of different videos to be edited in terms of emotion expression.
[0045] Specifically, at the content level, the server calculates the cosine similarity of the content themes of any two videos to be edited to construct a content association degree matrix of the videos to be edited. At the emotion level, the server calculates the emotion distance of the emotion features of any two videos to be edited to construct an emotion association degree matrix of the videos to be edited. The server organizes these association relationships into a graph structure, where the nodes represent the videos to be edited and the edges represent the association strength between the videos to be edited. This graph structure intuitively shows the semantic connections between the videos to be edited and provides an important reference for subsequent intelligent editing. The server will consider the weighted balance of content and emotion when constructing the video association graph to ensure that the generated works have both content coherence and emotional fluency.
[0046] Suppose there are 4 videos: V1: Jogging at the seaside during sunset (calm and pleasant); V2: City marathon (passionate); V3: Seaside yoga (peaceful and serene); V4: Urban extreme parkour (exciting); (1) Content theme vector representation (feature dimensions: running, yoga, extreme sports, seaside scene, urban scene): V1: [0.8, 0.1, 0.1, 0.9, 0.1]; V2: [0.9, 0.0, 0.2, 0.1, 0.9]; V3: [0.1, 0.9, 0.0, 0.9, 0.1]; V4: [0.3, 0.0, 0.9, 0.1, 0.9]; Calculate the content association degree matrix (cosine similarity): V1 V2 V3 V4 V1 1.0 0.6 0.7 0.2 V2 0.6 1.0 0.1 0.7 V3 0.7 0.1 1.0 0.1 V4 0.2 0.7 0.1 1.0 (2) Emotion feature vector representation (feature dimensions: pleasantness, intensity, calmness): V1: [0.8, 0.3, 0.7]; V2: [0.7, 0.9, 0.2]; V3: [0.6, 0.1, 0.9]; V4: [0.8, 0.9, 0.1]; Calculate the emotional correlation matrix (similarity after Euclidean distance normalization): (3) Comprehensive correlation matrix (weights: content 0.6, emotion 0.4): (4) Construct a video association graph based on the comprehensive correlation: Association Strength Division Video Association Map Representation Strong Association: 0.7 - 1.0 Strong Association: ==== Moderately Strong Association: 0.5 - 0.7 Moderately Strong Association: === Weak Association: 0.1 - 0.5 Weak Association: == V3 ======= V1 ======= V2 ===== V4; (5) Analysis results: Strong correlation group: V1 - V3: Seaside sports group (0.74); V2 - V4: Urban sports group (0.78); Medium - strong correlation group: V1 - V2: Running theme group (0.56); Weak correlation group: V3 - V2 / V4: Static vs dynamic sports (0.14 - 0.18); (6) Application scenario examples: Video recommendation: After watching V1: Give priority to recommending V3 (similar scenes and emotions); After watching V2: Give priority to recommending V4 (similar scenes and intensities); Intelligent editing: V1 and V3 can be edited into a "seaside sports collection"; V2 and V4 can be edited into an "urban sports collection"; Content organization: Scene dimension: Seaside group (V1, V3), urban group (V2, V4); Exercise intensity: High - intensity group (V2, V4), medium - low intensity group (V1, V3).
[0047] Optionally, generally, according to the video feature vectors, a video association graph is constructed. The video association graph is used to represent the content association degree and emotional association degree between different videos to be edited, which can be achieved in the following ways (not limited here): calculate the cosine similarity of the content themes of any two videos to be edited to construct a content association degree matrix of the videos to be edited; calculate the emotional distance of the emotional features of any two videos to be edited to construct an emotional association degree matrix of the videos to be edited; perform weighted fusion on the content association degree matrix and the emotional association degree matrix to obtain a comprehensive association degree matrix; and determine the video association graph based on the comprehensive association degree matrix.
[0048] S105. Based on the video association graph, perform combined matching on the videos to be edited and output multiple creative combination plans, where the creative combination plans are used to represent the arrangement order, connection method, and visual settings of the videos to be edited; Among them, the creative combination plan refers to the overall plan for intelligent sorting and effect setting of the videos to be edited; the arrangement order refers to the playing order of different videos to be edited; the connection method refers to the transition method between the videos to be edited; and the visual settings refer to the setting parameters of visual elements including filters, special effects, subtitles, etc.
[0049] Specifically, the server determines the optimal sorting sequence between the videos to be edited based on the association strength of each video to be edited in the video association graph, thereby generating multiple feasible video combination plans, ensuring that adjacent videos to be edited can naturally transition in terms of content and emotion. The server determines appropriate transition effects for each pair of videos to be edited through a preset algorithm and configures corresponding visual processing parameters. The preset algorithm will consider factors such as the scene type and rhythm change of the videos to be edited and automatically recommend multiple creative combination plans with different styles. Each creative combination plan includes a complete editing timeline, transition effect settings, and visual effect parameters, enabling the creator to quickly preview different creative effects.
[0050] S106. In response to the user's selection instruction for the target creative combination plan, obtain the user's editing operation sequence, where the target creative combination plan is any one of the creative combination plans, and the editing operation sequence is used to represent the adjustment process of the user for the target creative combination plan; Among them, the target creative combination plan refers to the creative combination plan selected by the user for adjustment; the editing operation sequence refers to a series of modification operations performed by the user, including clip point adjustment, special effect modification, parameter setting, etc.
[0051] Specifically, after the user selects a target creative combination plan as the base plan from multiple creative combination plans, the server starts recording all the user's editing operations. These editing operations include adjusting the time length of the video to be edited, modifying the transition effects, adjusting the audio parameters, etc. The server will record in detail the type, parameter changes, and timestamp of each editing operation to form a complete editing operation sequence. This editing operation sequence not only reflects the specific modification content of the user but also embodies the user's creative preferences and style characteristics. By analyzing the editing operation sequence, the server can better understand the user's creative intention.
[0052] S107. Extract the user's creative intention features according to the editing operation sequence. The creative intention features include video editing features, transition operation features, and audio-visual configuration features. Among them, the creative intention feature refers to the creative preference feature extracted from the user's editing behavior; the video editing feature is used to represent the user's habit of dividing video segments and controlling the duration; the transition operation feature is used to represent the user's preference for the effect selection when switching between different scenes; the audio-visual configuration feature is used to represent the user's way of processing the coordination between audio and video.
[0053] Specifically, the server conducts in-depth analysis on the recorded editing operation sequence. In terms of video editing, the server statistics the distribution of the segment lengths set by the user, the position characteristics of the cut points, etc. In terms of transition operations, the server analyzes the types of transition effects and parameter settings preferred by the user under different scene types. In terms of audio-visual configuration, the server extracts the timing characteristics of the user's audio adjustment and the synchronization relationship between audio changes and video switching. Through these multi-dimensional analyses, the server constructs the user's personalized creative intention features. Optionally, generally, according to the editing operation sequence, extracting the user's creative intention features, which include video editing features, transition operation features, and audio-visual configuration features, can be achieved in the following ways (not limited here): Based on the cut point setting operations in the editing operation sequence, calculate the duration interval and video change between adjacent cut points to construct the user's video editing features; According to the transition effect configuration operations in the editing operation sequence, statistics the usage patterns of transition effects under different combinations of scene types to determine the user's transition operation features. The usage patterns of transition effects include the probability of effect type selection and the configuration of effect parameters; According to the audio adjustment operations in the editing operation sequence, extract the audio adjustment timing characteristics and audio-visual synchronization characteristics to determine the user's audio-visual configuration features. The audio-visual synchronization characteristics are used to characterize the time correlation between video switching and audio changes.
[0054] The following lists a specific example of the analysis of the editing operation sequence of a travel Vlog creator: 1. Video editing features: (1) Segment duration rule: Opening segment: 5 - 6 seconds (showing the iconic building of the destination); Character introduction: 8 - 10 seconds (self - introduction and itinerary preview); Scenery display: 12 - 15 seconds (slow - motion display of scenic spots); Food segment: 6 - 8 seconds (display of each dish); Transition shot: 2 - 3 seconds (walking on the street, means of transportation); Ending segment: 4 - 5 seconds (ending with sunset or night view); (2) Cutting - point features: At scene changes: 40% (switching between different scenic spots); At the completion of actions: 35% (end of the character's actions); At pauses in speech: 25% (natural pauses in the commentary); (3) Characteristics of picture changes: Gradual progression from panoramic to close - up, rhythmic alternation between dynamic and static, and natural transition between bright and dark scenes.
[0055] 2. Transition operation features: (1) Between scenic spots: Fade - in and fade - out effect (accounting for 60%); Push - pull effect (accounting for 30%); Other effects (accounting for 10%); (2) Food display: Blind - effect (accounting for 50%); Dissolve effect (accounting for 40%); Other effects (accounting for 10%); (3) Character activities: Smooth transition (accounting for 70%); Blur transition (accounting for 20%); Other effects (accounting for 10%); (4) Transition parameter settings: Fade - in and fade - out: 1.5 - second duration; Push - pull effect: 2 - second duration, medium speed; Blind - effect: 1 - second duration, 12 stripes; Blur transition: 1 - second duration, 30% blur degree.
[0056] 3. Audio - visual configuration features: (1) Background music processing: Opening music: 2 - second fade - in, energetic and lively; Scenery display: volume 70%, soothing style; Food segment: volume 60%, lively rhythm; Character commentary: Background sound is reduced to 30%; Ending music: Fades out in 3 seconds, with a warm style; (2) Audio adjustment rules: Voice priority: Commentary volume is 90%; Ambient sound: 40% of the original sound is retained; Music coordination: Switches according to the scene style; (3) Audio-visual synchronization features: The scenic spot switching is synchronized with the music rhythm; The food display is precisely coordinated with the sound effects; The commentary is naturally connected with the picture.
[0057] 4. Summary of the creative intention characteristics: (1) Video style: Prefers a natural and smooth narrative rhythm; Values the integrity display of scenic spots; Focuses on conveying the human experience; (2) Editing habits: Takes 12 - 15 seconds as the main content unit; Is good at using transitional shots to connect scenes; Maintains the continuity of the picture composition; (3) Sound processing: Values the clarity of the commentary; Supplements with music emotion rendering; Focuses on the harmony and unity of audio and video.
[0058] S108. Based on the creative intention characteristics, the target creative combination plan is intelligently optimized to generate a media file. The intelligent optimization includes adjusting the video connection points, optimizing the transition effects, and the audio-visual matching.
[0059] Among them, intelligent optimization refers to the process of automatically improving the video effect according to the user's creative intention characteristics; the video connection point refers to the connection between adjacent videos to be edited; the transition effect refers to the transition effect between the videos to be edited; the audio-visual matching is used to handle the coordination relationship between the audio and the picture.
[0060] Specifically, the server systematically optimizes the target creative combination plan according to the extracted creative intention characteristics. First, the server fine-tunes the connection time points of the videos to be edited according to the user's editing habits, making the rhythm more in line with the user's style. Then, based on the user's transition operation characteristics, the server selects the most suitable special effect type and parameters for each scene transition. Next, the server optimizes the audio transition effect according to the user's audio-visual configuration characteristics to ensure audio-visual synchronization. Finally, the server integrates all the optimized settings to generate the final media file. This intelligent optimization based on the user's intention not only retains the advantages of the original plan but also incorporates the user's personalized style.
[0061] Suppose the target creative combination plan is a city travel video: Video to be edited 1: City landmarks (20 seconds); Video to be edited 2: Characteristic streets (30 seconds); Video to be edited 3: Local cuisine (25 seconds); Video to be edited 4: Night scene (20 seconds); Continuing with the example in step S107, the server optimizes the target creative combination plan based on the extracted creative intention features: 1. Apply video editing features: Clip Duration Pattern Opening: 5 - 6 seconds Scenery Display: 12 - 15 seconds Food Clip: 6 - 8 seconds Transition Shot: 2 - 3 seconds End Clip: 4 - 5 seconds Optimization and adjustment: Optimization of Video to be edited 1 (City landmarks): Original: 20 seconds → Optimized: 15 seconds (Opening iconic building: 6 seconds, Landmark detail display: 9 seconds); Optimization of Video to be edited 2 (Characteristic streets): Original: 30 seconds → Optimized: 24 seconds (Divided into 2 display units: Street panorama: 13 seconds, Street culture: 11 seconds); Optimization of Video to be edited 3 (Local cuisine): Original: 25 seconds → Optimized: 18 seconds (Each dish: 6 - 8 seconds, A total of 3 special dishes are shown); Optimization of Video to be edited 4 (Night scene): Original: 20 seconds → Optimized: 14 seconds (Night scene panorama: 9 seconds, Ending scene: 5 seconds).
[0062] 2. Apply transition operation features: Scene Transition Pattern Between Scenic Spots: Fade In and Out (60%) Food Display: Blinds Effect (50%) Character Activities: Smooth Transition (70%) Optimization process: Scene 1 → 2: Use a 1.5 - second fade - in and fade - out; Natural transition from landmarks to streets; Scene 2 → 3: Use a 1 - second blinds effect; Transition from the street to the food scene; Scene 3 → 4: Use a 2 - second push - pull effect; Time transition from food to night scene.
[0063] 3. Apply audio - visual configuration features: Background Music Processing Audio Adjustment Pattern Opening: Fade In for 2 seconds Voice Priority: 90% Scenery Display: Volume 70% Ambient Sound: 40% Food Clip: Volume 60% Music Matches the Scene Ending: Fade Out for 3 seconds Optimization implementation: Audio level processing: Landmark commentary: Volume 90%; Street ambient sound: Keep 40%; Food commentary: Volume 85%; Night scene background music: volume 70%; Music style matching: Opening: grand and expansive; Street: brisk pace; Food: warm and energetic; Night scene: soft and soothing; Audio and video synchronization: The transitions should be coordinated with the rhythm of the music; The commentary corresponds precisely to the images.
[0064] By adopting the above technical solution, the server obtains the user's login information and authenticates access to the video resources of multiple integrated media centers, realizing the unified retrieval of cross-platform video materials. The server extracts the content theme and emotional characteristics of each video to be edited, determines the content relevance and emotional relevance between different videos to be edited, and constructs a video association map to provide a basis for subsequent intelligent combination. At the same time, the server will learn the user's editing operation habits, extract the user's creative intention characteristics in video editing, transition operations, audio and video configuration, etc., and intelligently optimize the target creative combination plan accordingly. This method significantly improves the efficiency of integrated media content creation and reduces the repetitive work of creators in material acquisition and editing adjustments.
[0065] The following is a more detailed description of the process of the method provided by this implementation. Figure 2 , which is another flow chart of the cross-platform integrated media intelligent creation method in the embodiment of the present application.
[0066] S201, obtaining user login information, and sending a user authentication request to one or more integrated media centers based on the user login information, where the user authentication request is used to retrieve accessible videos from the integrated media center; For details, please refer to step S101, which is not limited here.
[0067] S202, obtaining the video resource address sent by the integrated media center; The video resource address refers to a network link used to locate and access specific video content, including a URL address of a video file, a video playback page address, or a video streaming media address. The video resource address may be in the form of a network link such as "http: / / example.com / videos / 12345".
[0068] Specifically, first, the server receives the authentication response information returned by each media convergence center, and parses the video resource address from the authentication response information. The server normalizes these video resource addresses to ensure that the address format is unified and valid.
[0069] S203. Extract video information from the video resource address based on a preset web page parsing rule to obtain video metadata including a video title, a video description, a video duration, a release time, and a video type. Among them, the preset web page parsing rule refers to a set of structured rules for extracting video information from different platform pages, including parsing methods such as DOM selectors and regular expressions; video metadata refers to structured data describing the basic attributes and content characteristics of a video; the video title refers to the name or theme of the video; the video description refers to the text description of the video content; the video duration refers to the playing duration of the video; the release time refers to the date and time when the video was first released; the video type refers to the classification label of the video content, such as news, entertainment, education, etc.
[0070] Specifically, the server calls the corresponding preset web page parsing rule according to the page structure characteristics of different media convergence platforms. First, the server accesses the page pointed to by the video resource address to obtain the page source code. Then, the server uses a DOM selector to locate the page elements containing the target information and extracts the text content such as the video title and video description. For the video duration, the server parses the duration attribute or duration label in the player component. The release time is usually extracted from the article release information or video metadata. The video type is identified according to the page classification label or video tag. The server performs format unification and validity verification on the extracted target information to ensure the accuracy and integrity of the video metadata.
[0071] S204. Store the video metadata in the video resource index database. Among them, the video resource index database is a structured database system for centrally storing and managing video metadata; storage refers to the process of writing the processed video metadata into the video resource index database.
[0072] Specifically, first, the server normalizes the video metadata to unify the data format and encoding method. Then, the server writes the processed video metadata into the video resource index database according to a predefined database schema. At the same time, the server creates indexes for important fields, such as the video title, release time, video type, etc., to improve the subsequent query efficiency. The server also records management information such as the update time and data source of the video metadata, and regularly performs data consistency checks and redundant data cleaning to ensure that the information in the video resource index database always remains accurate and up-to-date.
[0073] S205. Receive the accessible videos returned by the media convergence center and determine one or more videos to be edited from the accessible videos. Specifically, refer to step S102, which is not limited here.
[0074] S206. In response to the user's selection instruction for the target video template, obtain the target video template; Among them, the target video template refers to a pre-designed video production template, which includes predefined visual styles, layout structures, animation effects, and interaction methods; the selection instruction refers to the template selection operation performed by the user on the creation interface; obtaining refers to the process in which the server reads and loads the selected template. For example, the target video template can be a production template with a specific style such as a news report template, a product display template, or a travel vlog template.
[0075] Specifically, the server monitors the user's template selection operation on the creation interface. When it detects that the user clicks or confirms a certain video template, it triggers the template acquisition process. The server reads the complete configuration information of the target video template from the template library, including the visual style definition, animation parameter settings, layout rule configuration, etc. of the target video template. The server will verify the integrity and availability of the target video template to ensure that all necessary resources and configurations can be correctly loaded. At the same time, the server will pre-load the common materials and effect components in the target video template to improve the response speed of subsequent rendering.
[0076] S207. Receive the video title and text description input by the user; Among them, the video title refers to the main title text defined by the user for creating the video, which is used to summarize the video theme or content; the text description refers to the supplementary description text of the video content, which can include subtitles, chapter descriptions, or caption content; receiving refers to the process in which the server obtains and processes the user's text input. For example, the video title can be "Spring Travel Record in 2024", and the text description can be "Record the wonderful time of the cherry blossom season".
[0077] Specifically, the server provides input areas for the video title and text description on the creation interface and receives the user's text input in real time. The server preprocesses the video title and text description input by the user, including removing extra spaces, unifying line breaks, checking special characters, etc. The server will also perform text length verification to ensure that the input text meets the display limit. For text containing multiple languages, the server will perform character encoding conversion to ensure correct display during subsequent rendering.
[0078] S208. Based on the preset configuration of the target video template, perform style rendering on the video title and text description to obtain the output video; Among them, the preset configuration refers to the text style rules defined in the target video template, including parameters such as font, size, color, and animation effects; style rendering refers to the process of visually processing the video title and text description according to the preset configuration; the output video refers to the final video file after rendering. For example, the preset configuration can specify that the video title uses bold font size 36 with a fade-in animation, and the explanatory text uses regular script font size 24 with a fade-in effect.
[0079] Specifically, the server reads the preset configuration of the text style from the target video template, including font settings, position layouts, animation parameters, etc. of different text elements. The server performs rendering processing on the video title and text description according to the preset configuration, calculates the precise position and timeline of the text, and generates corresponding visual effects. If the text content is long, the server will automatically adjust the layout or perform paging processing. The server also processes the easing effect and timing control of the text animation to ensure that the transition of the text appearance and disappearance is natural and smooth. Finally, the server synthesizes the rendered text effect with the video content and outputs the final video file. During the rendering process, the server will perform performance optimization, use a caching mechanism to improve the rendering speed, and ensure that the quality of the output video meets the requirements.
[0080] S209. Extract the content theme and emotional characteristics of the video to be edited to generate a video feature vector; Specifically, reference can be made to step S103, which is not limited here.
[0081] S210. Construct a video association graph based on the video feature vector. The video association graph is used to represent the content association degree and emotional association degree between different videos to be edited; Specifically, reference can be made to step S104, which is not limited here.
[0082] S211. Based on the video association graph, perform combination matching on the videos to be edited and output multiple creative combination plans. The creative combination plans are used to represent the arrangement order, connection method, and visual settings of the videos to be edited; Specifically, reference can be made to step S105, which is not limited here.
[0083] S212. In response to the user's selection instruction for the target creative combination plan, obtain the user's editing operation sequence. The target creative combination plan is any one of the creative combination plans, and the editing operation sequence is used to represent the adjustment process of the user for the target creative combination plan; Specifically, reference can be made to step S106, which is not limited here.
[0084] S213. Extract the user's creative intention characteristics according to the editing operation sequence. The creative intention characteristics include video editing characteristics, transition operation characteristics, and audio-visual configuration characteristics; Specifically, refer to step S107, which is not limited here.
[0085] S214. Based on the creative intention features, intelligently optimize the target creative combination plan to generate a media file. The intelligent optimization includes adjusting video connection points, optimizing transition effects, and audio-visual matching. Specifically, refer to step S108, which is not limited here.
[0086] S215. Obtain the video specification parameters of multiple publishing platforms. The video specification parameters include video resolution, aspect ratio, and bitrate requirements. Among them, the publishing platform refers to various network platforms used for distributing and playing video content, such as short video platforms, social media platforms, and video websites. The video specification parameters refer to the technical specification requirements put forward by different publishing platforms for uploaded videos. The video resolution refers to the pixel size of the video frame, such as 1920x1080. The aspect ratio refers to the width-to-height ratio of the video frame, such as 16:9, 4:3, etc. The bitrate requirement refers to the limit of the video data transmission rate, usually in bits per second (bps). For example, a certain publishing platform may require the video resolution to be no less than 720p, the aspect ratio to be 16:9, and the bitrate to be between 2 - 8 Mbps.
[0087] Specifically, the server establishes connections with each publishing platform and obtains the latest video specification parameters through the API interface or configuration file. The server will parse and organize these video specification parameters and establish a parameter table containing the video specification requirements of each publishing platform. For different types of video content, the server will consider the differentiated requirements of the publishing platforms. For example, short video platforms may prefer a vertical screen with an aspect ratio of 9:16, while traditional video websites may require a horizontal screen with an aspect ratio of 16:9. The server will also verify the validity of these parameters to ensure that the obtained specification requirements are up-to-date and executable.
[0088] S216. Based on the video specification parameters, adjust the media file to generate the final media file adapted to different publishing platforms.
[0089] Among them, the media file refers to the original video content generated by the server in the early stage. Adjustment refers to the modification and optimization of the technical parameters of the media file. Adaptation means customized processing according to the requirements of different publishing platforms. The final media file refers to the video file that can be directly published after adaptation. For example, adjust a 1080p horizontal video to a 720p vertical video suitable for mobile phone vertical screen viewing.
[0090] Specifically, the server analyzes the specification differences of each publishing platform and formulates a video conversion plan for each publishing platform. For video resolution adjustment, the server will select an appropriate scaling algorithm to meet the size requirements while ensuring the picture quality. For aspect ratio adjustment, the server may need to crop the picture or add background filling to ensure that the core content can be fully displayed in different aspect ratios. In terms of bitrate control, the server will select appropriate encoding parameters according to the publishing platform's limitations to find a balance between file size and picture quality. The server will also consider the audio requirements of different publishing platforms and perform corresponding resampling or compression processing on the audio stream. During the generation process, the server will process multiple platform versions in parallel and improve the conversion efficiency through task queue management. Finally, the server will perform quality checks on the generated files of each version to ensure that they all meet the specification requirements of the corresponding publishing platform and prepare for subsequent file uploads.
[0091] The server in the embodiment of the present invention application will be described from the perspective of hardware processing. Please refer to Figure 3 , which is a schematic structural diagram of an entity device of the server in the embodiment of the present application.
[0092] It should be noted that Figure 3 the structure of the server shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0093] As Figure 3 shown, the server includes a CPU 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory ROM 302 or the program loaded from the storage section 308 into the random access memory RAM 303, such as executing the method described in the above embodiment. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other through a bus 304. The I / O interface 305 is also connected to the bus 304.
[0094] The following components are connected to the I / O interface 305: an input section 306 including an audio input device, a button switch, etc.; an output section 307 including a liquid crystal display (LCD), an audio output device, an indicator light, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read from it can be installed into the storage section 308 as needed.
[0095] Specifically, according to an embodiment of the present invention, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the CPU 301, various functions defined in the present invention are performed.
[0096] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings.
[0098] Specifically, the server in this embodiment includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, it implements the cross-platform integrated media intelligent creation method provided in the above embodiment.
[0099] On the other hand, the present invention also provides a computer-readable storage medium, which may be included in the server described in the above embodiment; or it may exist separately without being assembled into the server. The above storage medium carries one or more computer programs. When the above one or more computer programs are executed by a processor of the server, the server is enabled to implement the cross-platform integrated media intelligent creation method provided in the above embodiment.
[0100] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.
[0101] As used in the above embodiments, depending on the context, the term "when..." may be interpreted to mean "if...", or "after...", or "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" may be interpreted to mean "if determining...", or "in response to determining...", or "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".
[0102] Those of ordinary skill in the art can understand all or part of the processes in the methods of the above embodiments. These processes can be completed by relevant hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. The foregoing storage medium includes: various media such as ROM or random access memory RAM, magnetic disk, or optical disk that can store program codes.
Claims
1. A cross-platform integrated media intelligent creation method, characterized in that, Applied to a server, the method includes: Obtain user login information, and send a user authentication request to one or more media convergence centers based on the user login information, where the user authentication request is used to retrieve accessible videos from the media convergence centers; Receive the accessible videos returned by the media convergence centers, and determine one or more videos to be edited from the accessible videos; Extract the content theme and emotional features of the videos to be edited to generate video feature vectors; Construct a video association graph according to the video feature vectors, where the video association graph is used to represent the content association degree and emotional association degree between different videos to be edited; Based on the video association graph, perform combination matching on the videos to be edited, and output multiple creative combination schemes, where the creative combination schemes are used to represent the arrangement order, connection method, and visual settings of the videos to be edited; In response to a selection instruction of a user for a target creative combination scheme, obtain the user's editing operation sequence, where the target creative combination scheme is any one of the creative combination schemes, and the editing operation sequence is used to represent the adjustment process of the user for the target creative combination scheme; Extract the user's creative intention features according to the editing operation sequence, where the creative intention features include video editing features, transition operation features, and audio-visual configuration features; Based on the creative intention features, perform intelligent optimization on the target creative combination scheme to generate a media file, where the intelligent optimization includes adjusting video connection points, optimizing transition special effects, and audio-visual matching.
2. The method according to claim 1, characterized in that, The extracting the content theme and emotional features of the videos to be edited to generate video feature vectors specifically includes: Divide the videos to be edited into multiple video segments based on a preset video segment length; Process the video segments through a preset feature extraction network to obtain multi-modal features including picture features and audio features; Based on the multi-modal features, use a theme classification model to identify the content theme of the video segments; Based on the multi-modal features, use an emotion recognition model to determine the emotional features of the video segments; Fuse the content theme and the emotional features to generate the video feature vectors.
3. The method according to claim 1, characterized in that, The constructing a video association graph according to the video feature vectors, where the video association graph is used to represent the content association degree and emotional association degree between different videos to be edited, specifically includes: Calculate the cosine similarity of the content themes of any two videos to be edited to construct a content association degree matrix of the videos to be edited; Calculate the emotional distance of the emotional features of any two videos to be edited to construct an emotional association degree matrix of the videos to be edited; Perform weighted fusion on the content association degree matrix and the emotional association degree matrix to obtain a comprehensive association degree matrix; Based on the comprehensive association degree matrix, determine the video association graph.
4. The method according to claim 1, wherein The extracting the user's creative intention features according to the editing operation sequence, where the creative intention features include video editing features, transition operation features, and audio-visual configuration features, specifically includes: Based on the cut point setting operations in the editing operation sequence, calculate the duration intervals and picture changes between adjacent cut points to construct the video editing features of the user; According to the transition effect configuration operations in the editing operation sequence, count the transition effect usage patterns under different combinations of scene types to determine the user's transition operation features, where the transition effect usage patterns include the probability of selecting the effect type and the configuration of effect parameters; According to the audio adjustment operations in the editing operation sequence, extract the audio adjustment timing features and the audio-visual synchronization features to determine the user's audio-visual configuration features, and the audio-visual synchronization features are used to characterize the time correlation between the picture switching and the audio change.
5. The method according to claim 1, wherein After the step of obtaining the user login information and sending a user authentication request to one or more media convergence centers based on the user login information, the method further includes: Obtain the video resource address sent by the media convergence center; Based on the preset web page parsing rules, extract video information from the video resource address to obtain video metadata including the video title, video description, video duration, release time, and video type; Store the video metadata in the video resource index database.
6. The method according to claim 1, characterized in that, After the step of receiving the accessible video returned by the media convergence center and determining one or more videos to be edited from the accessible video, the method further includes: In response to the user's selection instruction for the target video template, obtain the target video template; Receive the video title and text description input by the user; Based on the preset configuration of the target video template, perform style rendering on the video title and text description to obtain the output video.
7. The method according to claim 1, wherein After the step of intelligently optimizing the target creative combination plan based on the creative intention features to generate a media file, the method further includes: Obtain the video specification parameters of multiple publishing platforms, where the video specification parameters include video resolution, picture ratio, and bit rate requirements; Based on the video specification parameters, adjust the media file to generate the final media file adapted to different publishing platforms.
8. A server, characterized in that, The server includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the server to execute the method according to any one of claims 1-7.
9. A computer-readable storage medium, comprising instructions, characterized in that, When the instruction runs on the server, cause the server to execute the method according to any one of claims 1-7.
10. A computer program product, characterized in that, When the computer program product runs on the server, cause the server to execute the method according to any one of claims 1-7.