Short play delivery optimization method and system based on artificial intelligence, and storage medium
By combining cross-modal feature fusion and deep temporal networks with multi-agent reinforcement learning and conditional generative adversarial networks, the problem of inaccurate user interest characterization in short drama delivery is solved, enabling intelligent optimization of delivery strategies and automatic generation of creative materials, thereby improving delivery effectiveness and resource allocation efficiency.
Patent Information
- Application Number
- CN202510738505.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Existing methods for placing short dramas rely on human experience and simple data analysis, which cannot accurately capture the deep connection between users and content, fail to reflect changes in user interests in a timely manner, and lack a deep understanding of the market competition environment, resulting in insufficient targeting accuracy and suboptimal resource allocation.
A cross-modal feature fusion network is used to deeply integrate the visual, auditory and user interaction data of short dramas. A deep temporal network is used to model the evolution of user interests. A multi-agent reinforcement learning network is used to model the delivery environment and predict strategies. A conditional generative adversarial network is used to generate creative materials. Finally, a distributed computing framework is used to analyze the delivery effect.
It enabled precise capture of user interests, improved campaign performance, optimized resource allocation, enhanced the differentiated expression of creative materials, and improved real-time monitoring of campaign performance and resource utilization efficiency.
Smart Images

Figure CN120786103B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and particularly relates to a short drama delivery optimization method and system based on artificial intelligence and a storage medium. BACKGROUND
[0002] With the rapid development of short videos and short drama contents, the demand of users for high-quality contents is continuously increasing. The existing short drama delivery method mainly relies on artificial experience and simple data analysis, and the delivery strategy is usually decided based on the basic portrait features and historical viewing data of users. In the content distribution process, fixed delivery rules and preset target audience groups are usually adopted, and there is a lack of dynamic perception of changes in user interests and real-time optimization of delivery effects. At the same time, the traditional short drama creative optimization method mainly relies on artificial editing and experience summary, and cannot fully utilize multi-modal data features for intelligent content adaptation.
[0003] However, this traditional delivery method has obvious deficiencies. First, a single data dimension cannot accurately capture the deep association between users and contents, resulting in insufficient delivery accuracy. Second, the static user portrait construction method cannot timely reflect the dynamic changes of user interests, affecting the delivery effect. Third, the mechanical delivery strategy optimization lacks deep perception of the market competition environment, making it difficult to achieve optimal allocation of resources. In addition, the optimization process of creative materials lacks consideration of the user emotional resonance dimension, affecting the content reach effect. SUMMARY
[0004] The present application provides a short drama delivery optimization method and system based on artificial intelligence and a storage medium, which solves the technical problem of inaccurate user interest description in the traditional delivery method by introducing multi-modal feature fusion and deep time series modeling technology. The method uses a cross-modal feature fusion network to deeply integrate the visual, auditory and user interaction data of short dramas, combines a deep time series network to model the evolution law of user interests, and realizes accurate capture of user interest features, significantly improving the delivery effect.
[0005] In a first aspect, the application provides an artificial intelligence-based short play delivery optimization method, which comprises: collecting short play picture data, audio data and user interaction data through a multi-modal data acquisition unit, performing feature extraction and fusion operation on the short play picture data, audio data and user interaction data by using a cross-modal feature fusion network to generate a short play fusion feature vector; based on the short play fusion feature vector, modeling and analyzing user viewing behavior sequences by using a deep time sequence network to obtain user interest transfer features, and constructing user portrait data according to the user interest transfer features; according to the user portrait data and the short play fusion feature vector, performing delivery environment modeling and delivery action prediction by using a multi-agent reinforcement learning network to obtain short play delivery strategy data; using the short play delivery strategy data to segment and reorganize short play materials, and generating multiple versions of creative materials by using a conditional generative adversarial network to form a short play creative material library; for the materials in the short play creative material library, using a distributed computing framework to analyze delivery effect, obtaining delivery effect data, and generating a delivery resource configuration scheme according to the delivery effect data.
[0006] In a second aspect, the application provides an artificial intelligence-based short play delivery optimization system, which comprises:
[0007] An acquisition module is configured to collect short play picture data, audio data and user interaction data through a multi-modal data acquisition unit, perform feature extraction and fusion operation on the short play picture data, audio data and user interaction data by using a cross-modal feature fusion network to generate a short play fusion feature vector;
[0008] A modeling module is configured to, based on the short play fusion feature vector, model and analyze user viewing behavior sequences by using a deep time sequence network to obtain user interest transfer features, and construct user portrait data according to the user interest transfer features;
[0009] A prediction module is configured to, according to the user portrait data and the short play fusion feature vector, perform delivery environment modeling and delivery action prediction by using a multi-agent reinforcement learning network to obtain short play delivery strategy data;
[0010] A reorganization module is configured to use the short play delivery strategy data to segment and reorganize short play materials, and generate multiple versions of creative materials by using a conditional generative adversarial network to form a short play creative material library;
[0011] An analysis module is configured to, for the materials in the short play creative material library, use a distributed computing framework to analyze delivery effect, obtain delivery effect data, and generate a delivery resource configuration scheme according to the delivery effect data.
[0012] The third aspect of the present application provides a computer device, comprising a memory and at least one processor, the memory storing instructions; the at least one processor calling the instructions in the memory to make the computer device execute the short play delivery optimization method based on artificial intelligence.
[0013] The fourth aspect of the present application provides a computer readable storage medium, the computer readable storage medium storing instructions, when running on a computer, making the computer execute the short play delivery optimization method based on artificial intelligence.
[0014] In the technical scheme provided in the present application, the short play picture data, audio data and user interaction data are collected by the multi-modal data acquisition unit, the feature extraction and fusion operation are performed by using the cross-modal feature fusion network, the short play content features are captured in all directions, and the problem of insufficient feature expression caused by single data dimension in the traditional method is effectively overcome; the user viewing behavior sequence is modeled and analyzed by using the deep time sequence network, the user interest transfer features are obtained and the user portrait data are constructed, the dynamic evolution law of the user interest is accurately described, and the timeliness and accuracy of the user portrait are improved; the delivery environment modeling and delivery action prediction are performed based on the multi-agent reinforcement learning network, the precise perception and strategy optimization of the complex delivery environment are realized, and the adaptability of the delivery decision is significantly improved; the short play materials are intelligently segmented and reorganized by using the conditional generative adversarial network, a plurality of versions of creative materials are generated, and the differentiation expression ability of the content is enhanced; the delivery effect analysis is performed by using the distributed computing framework, and the efficient configuration of the delivery resources is realized. The main innovations of the present application at the algorithm level are as follows: the cross-modal feature fusion network realizes the deep integration of multi-modal data by using the attention mechanism, and extracts more expressive content features; the deep time sequence network introduces the long short-term memory unit, and effectively models the time sequence dependence relationship of the user interest; the multi-agent reinforcement learning network adopts a hierarchical architecture design, and improves the generalization ability of the delivery strategy; the conditional generative adversarial network combines the prior knowledge in the field, and optimizes the generation quality of the creative materials. These algorithm innovations are deeply combined with specific application scenarios, and play an important role in the short play content distribution process: the precise identification of the user interest, the intelligent optimization of the delivery strategy, the automatic generation of the creative materials, and the real-time monitoring of the delivery effect are realized, and finally the goal of improving the short play delivery effect is achieved. Through actual application verification, the present application significantly improves the touch accuracy of the short play content, optimizes the use efficiency of the delivery resources, enhances the expressiveness of the creative materials, and provides effective technical support for the intelligent distribution of the short play content. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor based on these drawings.
[0016] Figure 1 An embodiment schematic diagram of the short play delivery optimization method based on artificial intelligence in the embodiments of the present application;
[0017] Figure 2 A user behavior schematic diagram of the user viewing behavior sequence in the embodiments of the present application;
[0018] Figure 3 An embodiment schematic diagram of the short play delivery optimization system based on artificial intelligence in the embodiments of the present application;
[0019] Figure 4 An embodiment schematic block diagram of the computer device in the embodiments of the present application. DETAILED DESCRIPTION
[0020] The embodiments of the present application provide a short play delivery optimization method and system based on artificial intelligence and a storage medium. The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0021] For the convenience of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 An embodiment of the short play delivery optimization method based on artificial intelligence in the embodiments of the present application includes:
[0022] In step S101, short play picture data, audio data and user interaction data are collected by a multi-modal data acquisition unit, and a cross-modal feature fusion network is used to perform feature extraction and fusion operation on the short play picture data, audio data and user interaction data to generate a short play fusion feature vector;
[0023] In step S102, based on the short drama fusion feature vector, a user viewing behavior sequence is modeled and analyzed by using a deep time sequence network to obtain a user interest transfer feature, and a user portrait data is constructed according to the user interest transfer feature.
[0024] In step S103, according to the user portrait data and the short drama fusion feature vector, a multi-agent reinforcement learning network is used to model a delivery environment and predict a delivery action to obtain short drama delivery strategy data.
[0025] In step S104, the short drama delivery strategy data is used to segment and reorganize short drama materials, and a conditional generative adversarial network is used to generate multiple versions of creative materials to form a short drama creative material library.
[0026] In step S105, for the materials in the short drama creative material library, a distributed computing framework is used to analyze the delivery effect to obtain delivery effect data, and a delivery resource configuration scheme is generated according to the delivery effect data.
[0027] It can be understood that the execution subject of the present application can be a short drama delivery optimization system based on artificial intelligence, and can also be a terminal or a server, which is not limited here. The server is taken as an example for description in the embodiments of the present application.
[0028] Specifically, in the multi-modal data acquisition stage, the multi-modal data acquisition unit is used to collect data of short drama content in all directions. The multi-modal data acquisition unit includes a video processing module, an audio analysis module and a user behavior tracking module, which are respectively responsible for collecting short drama picture data, audio data and user interaction data. For short drama picture data, the system extracts key frames at fixed time intervals (such as every second), analyzes the scene of each key frame, and extracts visual features such as color distribution, texture features and scene semantics. For audio data, the system divides the audio stream into several audio segments and extracts acoustic features such as tone, rhythm and emotion. For user interaction data, the system records user behaviors such as likes, comments and shares and their time stamps. Through a cross-modal feature fusion network, the three types of feature data are time-aligned and feature-fused to generate a unified short drama fusion feature vector.
[0029] In the user behavior analysis stage, the system analyzes the user's viewing behavior sequence based on the short drama fusion feature vector using a deep time series network. The deep time series network uses an attention mechanism to capture the time-dependent relationship of the user's viewing behavior. The system divides the user's viewing history into time windows and encodes the viewing behavior in each time window to extract the user's short-term and long-term interest features. By analyzing the user's migration patterns between different types of short dramas, the system generates user interest migration features to represent the dynamic change of the user's interest. Based on the user interest migration features, the system constructs a multi-dimensional user portrait data, including the user's basic attributes, interest preferences, and activity levels. In the delivery strategy generation stage, the system generates delivery strategies based on the user portrait data and short drama fusion feature vector using a multi-agent reinforcement learning network. The system models the delivery environment as a state space of multi-agent interaction, with each agent representing a delivery channel. Based on historical delivery data, the system trains the agents to learn the optimal delivery strategy, including delivery time, target audience, and budget allocation. The system generates specific short drama delivery strategy data through the delivery action prediction module.
[0030] In the creative material generation stage, the system intelligently segments and reorganizes the original short drama material based on the short drama delivery strategy data. The system identifies key scenes and highlights in the short drama and then reorganizes them based on different creative themes to generate multiple versions of creative materials. Specifically, the system uses a conditional generative adversarial network to generate creative variations that meet the target audience's preferences by taking the delivery strategy as a conditional input. All generated creative materials are stored in the short drama creative material library for subsequent delivery use. In the delivery effect analysis stage, the system uses a distributed computing framework to process large-scale delivery data. The system groups the delivery data of creative materials by delivery channel and calculates the effect indicators of each channel, such as exposure, clicks, and conversions. By calculating the performance indicators and delivery cost ratio of each channel, the system evaluates the delivery effect of different channels. Based on the delivery effect data, the system generates the optimal delivery resource allocation scheme.
[0031] For example, for a popular short drama work, the system extracts its visual features (such as scene transitions, character expressions, and picture composition), audio features (such as music mood, dialogue content), and user interaction features (such as comment sentiment, and forwarding time distribution). Through deep time series network analysis, it is found that the user's viewing behavior has obvious time period characteristics, and the user groups in different time periods have significant preference differences. Based on these analysis results, the system generates targeted delivery strategies, such as focusing on delivering emotional short drama content during the lunch and evening prime time. The system also intelligently edits the original short drama based on the characteristics of different target groups to generate multiple versions highlighting different points. Through continuous effect monitoring and optimization, the system can dynamically adjust resource allocation to improve overall delivery effect.
[0032] In the embodiment of the present application, short play picture data, audio data and user interaction data are collected by a multi-modal data acquisition unit, feature extraction and fusion operations are performed by a cross-modal feature fusion network, and the short play content features are captured in all directions, effectively overcoming the problem of insufficient feature expression caused by single data dimension in traditional methods; user interest migration features are obtained and user portrait data are constructed by modeling and analyzing user viewing behavior sequences through a deep time sequence network, accurately depicting the dynamic evolution law of user interest and improving the timeliness and accuracy of the user portrait; the adaptive of the delivery decision is significantly improved by precisely perceiving the complex delivery environment and optimizing the strategy based on the multi-agent reinforcement learning network for delivery environment modeling and delivery action prediction; the creative material is segmented and reorganized intelligently by the conditional generative adversarial network to generate multiple versions of creative material, enhancing the differentiated expression ability of the content; and the distributed computing framework is used for delivery effect analysis to realize efficient allocation of delivery resources. The main innovations of the present application at the algorithm level are as follows: the cross-modal feature fusion network realizes the deep integration of multi-modal data through the attention mechanism, and extracts more expressive content features; the deep time sequence network introduces a long short-term memory unit to effectively model the time sequence dependence of user interest; the multi-agent reinforcement learning network adopts a hierarchical architecture design to improve the generalization ability of the delivery strategy; and the conditional generative adversarial network combines domain prior knowledge to optimize the generation quality of creative material. These algorithm innovations are deeply combined with specific application scenarios, and play an important role in the short play content distribution process: accurate identification of user interest, intelligent optimization of delivery strategy, automatic generation of creative material, and real-time monitoring of delivery effect, ultimately achieving the goal of improving the delivery effect of short plays. Through actual application verification, the present application significantly improves the accuracy of short play content reach, optimizes the use efficiency of delivery resources, enhances the expressiveness of creative material, and provides effective technical support for the intelligent distribution of short play content.
[0033] In a specific embodiment, the process of step S101 can specifically include the following steps:
[0034] (1) The short play picture data is segmented and sampled according to a time window to obtain a set of sampling segments, and multi-level feature decomposition is performed on each sampling segment in the set of sampling segments to extract color distribution features, texture features and motion trajectory features. The color distribution features, texture features and motion trajectory features are spliced and combined to obtain a visual feature matrix;
[0035] (2) The audio data is divided into multiple overlapping audio frames, and a Fourier transform is performed on each audio frame to obtain a spectrum graph. The pitch feature, rhythm feature and emotion feature are extracted from the spectrum graph, and the pitch feature, rhythm feature and emotion feature are combined in time sequence to obtain an audio feature matrix;
[0036] (3) The user interaction data is sorted according to the timestamp, the interaction behavior sequence is constructed, the click density, the stay time and the interaction frequency in the interaction behavior sequence are calculated, and the click density, the stay time and the interaction frequency are combined into a user behavior feature matrix;
[0037] (4) The time sequence correlation between the visual feature matrix, the audio feature matrix and the user behavior feature matrix is calculated by using the cross-modal feature fusion network, a time sequence correlation matrix is generated, and the three feature matrices are resampled and interpolated based on the time sequence correlation matrix to obtain aligned feature data;
[0038] (5) The information entropy of each modal feature in the aligned feature data is calculated to determine the modal fusion weight, the aligned feature data is weighted and summed by using the modal fusion weight, and normalized to generate a short drama fusion feature vector.
[0039] Specifically, in the short drama picture data processing stage, a time window segmentation sampling method is adopted. The window length is set to 3 seconds, and the adjacent windows overlap by 1 second. This overlapping design ensures the continuity of feature extraction. For the video sequence in each window, the sampling formula is:
[0040]
[0041] Wherein, is the sampling result at time t, N is the number of sampling points in the window (for example, for a 30fps video, a 3-second window corresponds to 90 sampling points), is the weight coefficient of the i-th sampling point, F(t) is the original video frame sequence, is the sampling interval. Taking an action scene as an example, when a rapid action change is detected, the weight coefficient of the corresponding position will be given a larger value to highlight the features of key action frames. For the frame sequence obtained by sampling, color distribution features (histogram of HSV color space), texture features (gray level co-occurrence matrix statistics) and motion trajectory features (optical flow field vector) are extracted, and these features are combined into a visual feature matrix.
[0042] In the audio processing link, the audio data is first divided into frames according to a window length of 25ms and a frame shift of 10ms. Short-time Fourier transform is applied to each audio frame:
[0043]
[0044] Here represents the time-frequency spectrum, is the Hamming window function, and x(t) is the audio signal, is the angular frequency. Through this formula, the system converts the time-domain signal into a time-frequency representation, facilitating the extraction of audio features. For example, when processing a scene of emotional dialogue, the system extracts pitch variations (fundamental frequency trajectory), rhythm features (energy envelope), and emotional features (harmonic structure) by analyzing the spectrogram, which are organized into an audio feature matrix.
[0045] For the processing of user interaction data, various types of interaction behaviors are sorted and feature extracted based on timestamps. Within each analysis window (usually 10 seconds), the cumulative value of interaction features is calculated:
[0046]
[0047] where is the comprehensive interaction index at time t, is the click density (times / second), is the average dwell time (seconds), F(t) is the interaction frequency (times / minute), , , are the weight coefficients of each feature, respectively. For example, for a dramatic climax segment, when the click density reaches 10 times / second, the average dwell time exceeds 30 seconds, and the interaction frequency reaches 60 times / minute, the system will identify this as a key moment of high user engagement.
[0048] In the cross-modal feature fusion stage, the system establishes a feature alignment mechanism by calculating the temporal correlation between the three feature matrices. All features are resampled to a unified time scale (such as 100ms as the basic unit), and then linear interpolation is used to fill in missing values. For a typical short drama scene, such as an emotional outburst point, visual features may exhibit dramatic changes in expressions and actions, audio features may exhibit significant changes in volume and tone, and user interactions may exhibit dense comments and likes. The system ensures the precise correspondence of these features in the time dimension through feature alignment. By calculating the information entropy of each modal feature, the weights of different features in the fusion process are determined. For different types of scenes, these weights will be dynamically adjusted. For example, in action scenes dominated by visual effects, the weight of visual features will be correspondingly increased; in emotional scenes dominated by dialogue, audio features will be given greater weight. This dynamic weight adjustment mechanism ensures the flexibility and accuracy of feature fusion.
[0049] In a specific embodiment, the process of performing step S102 can specifically include the following steps:
[0050] (1) Construct a viewing feature sequence according to the short drama fusion feature vector, and divide the viewing feature sequence into time windows to obtain viewing window data;
[0051] (2) sequentially encode the user viewing behavior sequence in each time window in the viewing window data to generate a behavior encoding sequence;
[0052] (3) perform time sequence correlation analysis on the behavior encoding sequence by using a deep time sequence network to extract viewing duration, switching frequency and interaction intensity, and generate a behavior feature matrix;
[0053] (4) calculate the user interest change rate based on the behavior feature matrix, and perform time sequence accumulation on the user interest change rate to obtain a user interest migration feature;
[0054] (5) cluster and group the user interest migration feature according to interest themes to construct an interest category table;
[0055] According to the weight distribution of each category in the interest category table, user portrait data is generated.
[0056] Specifically, as shown in Figure 2 , it is a user behavior schematic diagram of the user viewing behavior sequence in the embodiment of the present application. When constructing the viewing feature sequence based on the short drama fusion feature vector, the time window division method is adopted, and the length of each window is set to 5 minutes, and the adjacent windows overlap by 1 minute. For each viewing sequence, the following feature extraction formula is adopted:
[0057]
[0058] Among them, represents the viewing behavior feature at time t, is the number of sampling points in the window, is the weight coefficient of the kth sampling point, Q(t) is the original viewing behavior sequence, is the sampling time interval. For example, when processing a continuous drama segment, if the user repeatedly watches at a certain time point, the weight coefficient of the corresponding position will be assigned a larger value, highlighting the behavior feature of the time point.
[0059] For the generation of the user behavior encoding sequence, the sequence encoding function is adopted:
[0060]
[0061] Here is the behavior encoding value at position r, is the play state (1 represents playing and 0 represents pausing), is the viewing time (unit: seconds), is the number of jumps, , , are the weight coefficients of each behavior feature. Taking a typical scenario as an example, when a user is watching a plot, if the behaviors of frequent pausing and repeatedly watching a certain segment occur, these features will be encoded into higher behavior feature values.
[0062] In calculating the user interest change rate, the interest dynamic change function is adopted:
[0063]
[0064] wherein is the interest change rate at time t, is the content type change (the number of conversions between different types), is the viewing speed change (the number of adjustment times of the playback speed), is the interaction intensity (the frequency of behaviors such as likes and comments), , , are the corresponding weight coefficients. In actual application, for example, when a user switches from a light plot to a suspense plot, the system will detect a significant change in the content type, and at the same time, in combination with the user's viewing speed adjustment and interaction behavior, calculate the trend of interest change.
[0065] Based on the calculated interest change rate, the system performs time series accumulation to obtain the user interest migration feature. The evolution process of the user interest over time is recorded. For example, if a user initially prefers light and funny content, but gradually starts to focus on suspense plots, this interest migration will be accurately captured and quantified by the system.
[0066] In the interest clustering grouping phase, the system clusters the user interest migration feature according to the content theme. For example, the user's interest is divided into suspense reasoning, urban sentiment, light and funny, etc. For each category, the system calculates its weight value, reflecting the user's preference for the content of that category. For example, if the user's viewing time of the suspense category accounts for 60%, and the interaction behavior accounts for 70%, the suspense category will be given a higher weight. According to the weight distribution of the interest categories, the system generates complete user portrait data. This portrait contains information on the user's interest preferences, viewing habits, and interaction patterns, among other dimensions. In this way, the system can accurately depict the user's interest features.
[0067] In a specific embodiment, the process of performing step S103 can specifically include the following steps:
[0068] (1) Layering the user portrait data according to the interest dimension, constructing a user layered matrix, and mapping the short drama fusion feature vector into the user layered matrix to generate an interest matching degree distribution;
[0069] (2) Expand the interest matching degree distribution in the spatiotemporal dimensions, establish a delivery status space table, and extract the delivery time period, delivery area and delivery channel from the delivery status space table to form delivery environment characteristics;
[0070] (3) Input the characteristics of the delivery environment into the multi-agent reinforcement learning network to model the competitive relationship of each delivery channel and obtain the channel competition matrix;
[0071] (4) Calculate the reward value of the placement action based on the channel competition matrix, and use the reward value as the optimization target to construct the action value table;
[0072] (5) Select the optimal combination of deployment actions from the action value table and generate a deployment decision sequence;
[0073] (6) Match the decision sequence with the constraints to obtain short drama delivery strategy data.
[0074] Specifically, in the user segmentation stage, user profile data is segmented according to different interest dimensions. The user interest matching degree is calculated using the following formula:
[0075]
[0076] in Indicates interest levels The matching degree value, The number of dimensions of interest features. Let G(x) be the weight coefficient of the p-th dimension, and G(x) be the interest feature function. The dimension interval is used. For example, for a young user group, if their viewing history shows that 70% of their viewing time is spent watching suspenseful short dramas and their interaction frequency is 5 times per hour, then the weight coefficient of the suspense category will be increased accordingly, which will be reflected in the interest matching degree distribution.
[0077] For modeling the characteristics of the deployment environment, the system adopts a spatiotemporal state transition function:
[0078]
[0079] here For spacetime points Environmental characteristic values, It is a time-related characteristic (such as peak period, trough period). Spatial characteristics (such as regional attributes). It is characterized by spatiotemporal coupling. , , corresponding weight coefficient. In actual application, for example, 8-10 pm in a certain city is the peak time for watching short videos, and the time feature weight of this time period will be increased, and combined with the user activity data in this region, the delivery environment features are formed.
[0080] The modeling of the channel competition relationship adopts a reward calculation formula:
[0081]
[0082] wherein R(a, b) is the competition relationship value between channels a and b, V(a, b) is the traffic competition degree, C(a, b) is the cost ratio, F(a, b) is the conversion efficiency ratio, 、 、 corresponding weight coefficient. For example, between two mainstream short video platforms, if the unit delivery cost of platform A is 0.8 times that of platform B, but the conversion rate is 1.2 times that of platform B, these factors will be comprehensively calculated through the formula to obtain the competition relationship value between channels. When generating the delivery action value table, the system comprehensively considers the competition relationship and delivery effect of each channel. For example, when it is found that the delivery effect of a certain channel in a certain period of time is significantly better than that of other channels, the system will increase the delivery weight of the channel in this period of time. In specific operation, the system will record the historical performance of each delivery action, including click rate, conversion rate, return on investment, etc., and update the action value table accordingly.
[0083] In the generation process of the delivery decision sequence, the system will select the optimal combination of delivery actions from the action value table. Multiple factors are considered, such as the active period of the target audience, the traffic distribution of the channel, the delivery situation of the competitors, etc. For example, for young user groups, the system will increase the delivery intensity in the evening rest period and preferentially select short video platforms with high user activity. The system will match the generated delivery decision sequence with actual delivery constraint conditions. These constraint conditions include budget limitations, delivery time period limitations, material specification requirements, etc. Through this matching process, it is ensured that the final generated short drama delivery strategy data not only meets the optimization target, but also meets the actual execution conditions. For example, when the daily budget is 10000 yuan, the system will reasonably allocate the budget according to the expected effect of each period of time to ensure that key periods and channels obtain sufficient resource input.
[0084] In a specific embodiment, the process of performing step S104 can specifically include the following steps:
[0085] (1) converting the short drama delivery strategy data into a scene description matrix, and performing key frame extraction on the short drama materials based on the scene description matrix to obtain a scene segment set;
[0086] (2) Perform semantic analysis on the scene segment set to generate scene semantic vectors, and establish a segment association graph based on the scene semantic vectors;
[0087] (3) Reorganize and sort the scene segment set according to the scene connection relationship in the segment association graph to generate a reorganized scene sequence;
[0088] (4) Input the reorganized scene sequence into a conditional generative adversarial network to generate multiple initial creative materials of different theme styles;
[0089] (5) Perform creative scoring on the initial creative materials, extract creative theme tags and style feature tags, and construct a creative material feature table;
[0090] (6) Classify and organize the creative materials in the creative material feature table according to the theme tags and style feature tags to generate a short play creative material library.
[0091] Specifically, it is converted into a scene description matrix. This conversion process takes into account information in multiple dimensions such as delivery time, target audience, and expected effect. For example, a late-night delivery strategy targeting young people will be converted into scene description information containing time characteristics, population characteristics, and effect targets. Based on this scene description matrix, key frames are extracted from the original short play materials, and the extraction rules include scene cut points, expression change points, and action climax points. Each key frame has a timestamp and a scene type label, forming a scene segment set. For the extracted scene segment set, the next step is semantic analysis. The semantic analysis process includes identifying the main elements in the scene (such as characters, actions, and environment), emotional tone, plot development, and other content. Through these analysis results, a scene semantic vector is generated, which contains the core semantic information of the scene. Based on these semantic vectors, an association graph between segments is established, which reflects the logical relationship between different scene segments. For example, in an emotional short play, scenes with similar emotional tones will have a strong association, and emotional turning points will be marked as key nodes.
[0092] According to the established segment association graph, the scene segments are reorganized and sorted. The reorganization process needs to consider the continuity between scenes, the emotional progression relationship, the plot development rule and other factors. By analyzing the scene connection relationship in the association graph, the best scene combination order is determined to generate a reorganized scene sequence. This reorganization process is not a simple linear arrangement, but an intelligent scene arrangement according to creative needs. The reorganized scene sequence is input into a conditional generative adversarial network to generate multiple initial creative materials of different theme styles. During the generation process, the style features of the generated materials are controlled by adjusting the conditional parameters. For example, for different target audience groups, creative versions of different styles such as energetic, warm, and suspenseful are generated. Each generated creative material retains the core content of the original scene while having unique style features.
[0093] The generated initial creative material is scored for creativity. The scoring dimensions include visual appeal, emotional appeal, information transmission effect, and other aspects. At the same time, the theme tag (such as "youth", "inspiring", "warm" and the like) and the style feature tag (such as "bright", "warm", "tense" and the like) of each creative material are extracted. These scores and tag information are sorted into a creative material feature table, which records the detailed feature description of each creative material. The materials in the creative material feature table are classified and sorted. The classification standard is mainly based on the theme tag and the style feature tag, and the materials with similar features are classified into the same category. This classification method facilitates the quick call of creative materials that meet the needs in actual placement. The sorted materials are stored in the short drama creative material library, forming a systematic creative resource library.
[0094] Taking a typical short drama placement optimization case as an example: the original material of an emotional short drama is 3 minutes long, and 12 key scene segments are extracted through scene analysis. After semantic analysis, core plot nodes such as "first meeting", "complications", and "reconciliation" are identified. Based on the semantic correlation between the segments, a complete scene correlation network is constructed. Based on this network, multiple creative versions are generated for different placement scenarios: the version for young people highlights the youth and vitality elements, focusing on the interaction scenes of the characters; the version for mature people focuses on emotional resonance, highlighting the inner play of the characters. These different versions of creative materials are sorted into the creative material library after scoring and tag extraction. When placement in a specific scenario is needed, the most suitable creative version is retrieved from the material library to achieve precise placement.
[0095] In a specific embodiment, the process of performing step S105 can specifically include the following steps:
[0096] (1) Group the materials in the short drama creative material library according to the placement channels, construct a channel placement record table, and from the channel placement record table, count the exposure, click volume and conversion volume to generate an effect statistics matrix;
[0097] (2) Calculate the click conversion ratio and placement cost ratio of each placement channel based on the effect statistics matrix, and combine the click conversion ratio and placement cost ratio to form a channel efficiency index;
[0098] (3) Time series expansion is performed on the channel efficiency index to construct an efficiency trend chart, and the placement time period weight is extracted from the efficiency trend chart;
[0099] (4) Cross operation is performed on the placement time period weight and the channel efficiency index to obtain a space-time placement score table;
[0100] (5) According to the space-time release scoring table, the priority of each release channel is sorted, and the release effect data is generated;
[0101] (6) Based on the release effect data, the space-time dimension distribution of the release resource is obtained, and the release resource configuration scheme is obtained.
[0102] Specifically, the 24-hour release data of six typical release channels (channels AA, BB, CC, DD, EE, FF) is grouped and arranged. The key indicators of channel AA in the evening peak period (20:00-21:00) are: total exposure 2.2 million times, effective click 132,000 times, deep conversion 26,400 times, benchmark cost 9,000 yuan, quality interaction 46,200 times, and benchmark interaction 132,000 times. The data of channel BB at the same time shows that the total exposure is 180,000 times, the effective click is 90,000 times, the deep conversion is 18,000 times, the benchmark cost is 7,000 yuan, the quality interaction is 27,000 times, and the benchmark interaction is 90,000 times. The data of channel CC is: total exposure 150,000 times, effective click 60,000 times, deep conversion 12,000 times, benchmark cost 5,500 yuan, quality interaction 18,000 times, and benchmark interaction 60,000 times.
[0103] The channel efficiency evaluation adopts a multi-level calculation model:
[0104]
[0105] Among them, is the efficiency index value of channel v, is the effective click (times / hour), is the total exposure (times / hour), is the deep conversion (times / hour), is the total click (times / hour), is the conversion revenue (yuan / hour), is the benchmark cost (yuan / hour), is the quality interaction (times / hour), is the benchmark interaction (times / hour), and are balance coefficients (values 0.4 and 0.6).
[0106] Taking channel AA as an example, the specific data is substituted: , indicating that the effective click rate is 6%; reflects the deep conversion effect; indicates the logarithmic measure of revenue and cost; reflects the influence of quality interaction proportion.
[0107] The time series analysis introduces a progressive evaluation function:
[0108]
[0109] wherein is the performance value of the channel v at the time period i, is the peak flow (times / minute), is the average flow (times / minute), is the actual income (yuan / hour), is the expected income (yuan / hour), is the cumulative interaction (times), is the target interaction (times), , , is the dynamic weight coefficient (0.3, 0.4, 0.3 respectively).
[0110] Specific to the data of channel AA: peak flow 3600 times / minute, average flow 1285 times / minute, get ; actual income 22500 yuan / hour, expected income 17300 yuan / hour, calculated ; cumulative interaction 125000 times, target interaction 100000 times, get .
[0111] Cross-dimension feature fusion model:
[0112]
[0113] wherein is the comprehensive score of the space-time point is the real-time conversion (times / minute), is the benchmark conversion (times / minute), is the user stay duration (seconds), is the average stay duration (seconds), is the sharing propagation (times), is the exposure (times), , , is the feature weight (value 0.35, 0.35, 0.3). The specific performance of channel AA is: real-time conversion 29 times / minute, benchmark conversion 20 times / minute, get
[0114] ; user stay duration 96 seconds, average stay duration 60 seconds, calculate ; sharing propagation 110000 times, exposure 2200000 times, get .
[0115] Based on the above comprehensive score results, the resource allocation of the six channels is optimized. The specific allocation scheme of the daily budget of 200,000 yuan is as follows: channel AA obtains 60,000 yuan of investment budget, which is mainly used for investment in the evening peak period (20:00-22:00) because it shows the highest efficiency index and conversion effect in this period; channel BB is allocated 45,000 yuan, focusing on covering the evening and weekend periods, taking advantage of the active users in these periods; channel CC obtains 35,000 yuan, which is specifically invested in its advantage performance period; the remaining 60,000 yuan is allocated to channels DD, EE and FF according to the efficiency proportion of each channel.
[0116] In a specific embodiment, the process of performing step S106 can specifically include the following steps:
[0117] (1) Normalize the score data in the space-time investment score table, construct a standardized score matrix, and extract the time period performance value of each channel from the standardized score matrix;
[0118] (2) Perform aggregation operation on the time period performance value according to the channel dimension to obtain the channel performance index, and construct a channel ranking table based on the channel performance index;
[0119] (3) Calculate the historical investment yield rate for each channel in the channel ranking table to generate a yield distribution curve, and extract the investment advantage coefficient from the yield distribution curve;
[0120] (4) Weight and combine the investment advantage coefficient and the channel performance index to form a channel comprehensive score;
[0121] (5) Sort the channel comprehensive score in descending order to generate a channel priority list;
[0122] (6) Combine the channel priority list with the time period performance value to generate investment effect data.
[0123] Specifically, the space-time investment score data of the five investment channels (channels KK, LL, MM, NN and PP) is standardized. The original score data of channel KK in the peak period (20:00-21:00) is as follows: original score 92 points, actual conversion quantity 2800 times / hour, target conversion quantity 2000 times / hour, effective interaction quantity 4200 times / hour, and benchmark interaction quantity 3000 times / hour. These data are standardized to construct a standardized score matrix.
[0124] The normalization calculation uses a comprehensive evaluation model:
[0125]
[0126] where H(x) is the standardized score value of channel x, I(x) is the original score value, and are the maximum and minimum values of the score, respectively, is the actual conversion amount (times / hour), is the target conversion amount (times / hour), is the effective interaction amount (times / hour), is the benchmark interaction amount (times / hour), and are weight coefficients (values 0.6 and 0.4). Taking channel KK as an example, the actual data is substituted: original score 92 points, actual conversion amount 2800 times / hour, target conversion amount 2000 times / hour, effective interaction amount 4200 times / hour, and benchmark interaction amount 3000 times / hour. After standardization processing, the performance values of each period are extracted from the score matrix. The performance data of channel KK at different periods shows: early peak (8:00-10:00) standardized score 0.75, midday peak (12:00-14:00) score 0.82, and late peak (20:00-22:00) score 0.95. These period performance values are further aggregated according to the channel dimension to obtain the overall performance index of the channel.
[0127] The channel performance index calculation introduces a multi-level evaluation model:
[0128]
[0129] wherein is the performance index of channel x in period t, is the peak flow (times / minute), P(t,x) is the average flow (times / minute), Q(t,x) is the user stay duration (seconds), R(t,x) is the benchmark stay duration (seconds), S(t,x) is the actual revenue (yuan), and T(t,x) is the target revenue (yuan), , , are weight coefficients (values 0.3, 0.4, and 0.3, respectively). The specific data of channel KK at the peak period: peak flow 4000 times / minute, average flow 1500 times / minute; user stay duration 120 seconds, benchmark stay duration 80 seconds; actual revenue 25000 yuan, and target revenue 20000 yuan. Based on the calculated performance index, a channel ranking table is constructed. The ranking data shows: the comprehensive performance index of channel KK is the highest, reaching 0.92; channel LL is 0.85; and channel MM is 0.78. At the same time, the historical delivery data of each channel in the past 30 days is analyzed, the delivery yield is calculated, and a yield distribution curve is generated. The historical yield data of channel KK shows: the average daily yield is 185%, and the fluctuation range is controlled within 15%; the yield at the peak period can reach as high as 230%.
[0130] The comprehensive score system adopts a dynamic weight model:
[0131]
[0132] wherein is a comprehensive score of channel x in time period t, is a historical performance index, is a current revenue (yuan), is a historical average revenue (yuan), is a user activity (times / hour), and V(t, x) is a benchmark activity (times / hour), , , is a dynamic weight coefficient (sum is 1). The data of channel KK shows that the historical performance index is 0.92, the current revenue is 28000 yuan, the historical average revenue is 22000 yuan, the user activity is 8500 times / hour, and the benchmark activity is 6000 times / hour. After extracting the delivery advantage coefficient from the revenue distribution curve, the channel performance index is weighted and fused to form the final comprehensive score of the channel. The advantage coefficient of channel KK is 1.28, combined with the performance index 0.92, the final comprehensive score reaches 94 points. The comprehensive scores of each channel are sorted in descending order to generate a priority list: channel KK > LL > MM > NN > PP.
[0133] The priority list is combined with the time period performance value to generate complete delivery effect data. The specific delivery scheme is: channel KK obtains a budget of 54000 yuan, focuses on covering the late peak period, and the expected ROI is 220%; channel LL is allocated 45000 yuan, mainly attacking the noon and evening periods, and the expected ROI is 180%; the remaining funds are allocated to other channels according to the score ratio. The one-week running data shows that after the implementation of the scheme, the overall delivery effect is improved by 35%, and the ROI of each channel exceeds the expected target.
[0134] The short drama delivery optimization method based on artificial intelligence in the embodiments of the application is described above, and the short drama delivery optimization system based on artificial intelligence in the embodiments of the application is described below. Please refer to Figure 3 The short drama delivery optimization system based on artificial intelligence in the embodiments of the application includes one embodiment:
[0135] The acquisition module is configured to acquire short drama picture data, audio data and user interaction data through a multi-modal data acquisition unit, and perform feature extraction and fusion operation on the short drama picture data, audio data and user interaction data by using a cross-modal feature fusion network to generate a short drama fusion feature vector.
[0136] The modeling module is configured to model and analyze the user viewing behavior sequence based on the short drama fusion feature vector by using a deep time sequence network, obtain a user interest transfer feature, and construct user portrait data according to the user interest transfer feature.
[0137] The prediction module is configured to model and predict the delivery action by using a multi-agent reinforcement learning network according to the user portrait data and the short drama fusion feature vector, and obtain short drama delivery strategy data.
[0138] The reorganization module is configured to segment and reorganize the short drama material by using the short drama delivery strategy data, generate a plurality of versions of creative material by using a conditional generative adversarial network, and form a short drama creative material library.
[0139] The analysis module is configured to analyze the delivery effect of the material in the short drama creative material library by using a distributed computing framework, obtain delivery effect data, and generate a delivery resource configuration scheme according to the delivery effect data.
[0140] Through the cooperation of the above-mentioned components, the short drama picture data, audio data and user interaction data are collected through the multi-modal data acquisition unit, the cross-modal feature fusion network is used for feature extraction and fusion operation, the short drama content features are captured in all directions, and the problem of insufficient feature expression caused by single data dimension in the traditional method is effectively overcome; the user interest transfer features are obtained and the user portrait data are constructed by modeling and analyzing the user viewing behavior sequence through the deep time sequence network, the dynamic evolution law of the user interest is accurately described, and the timeliness and accuracy of the user portrait are improved; the multi-agent reinforcement learning network is used for modeling the delivery environment and predicting the delivery action, the complex delivery environment is accurately perceived and the strategy is optimized, and the adaptability of the delivery decision is significantly improved; the short drama materials are intelligently segmented and reorganized through the conditional generative adversarial network, and multiple versions of creative materials are generated, and the differentiated expression ability of the content is enhanced; the distributed computing framework is used for delivery effect analysis, and efficient allocation of delivery resources is realized. The main innovations of the present application at the algorithm level are as follows: the cross-modal feature fusion network realizes the deep integration of multi-modal data through the attention mechanism, and extracts more expressive content features; the deep time sequence network introduces the long short-term memory unit, and effectively models the time sequence dependence relationship of the user interest; the multi-agent reinforcement learning network adopts a hierarchical architecture design, which improves the generalization ability of the delivery strategy; the conditional generative adversarial network combines with the prior knowledge in the field, and optimizes the generation quality of the creative materials. These algorithm innovations and specific application scenarios are deeply combined, and play an important role in the short drama content distribution process: accurate identification of user interest, intelligent optimization of delivery strategy, automatic generation of creative materials, and real-time monitoring of delivery effect, ultimately achieving the goal of improving the delivery effect of short dramas. Through actual application verification, the present application significantly improves the accuracy of short drama content reach, optimizes the use efficiency of delivery resources, enhances the expressiveness of creative materials, and provides effective technical support for the intelligent distribution of short drama content.
[0141] Referring to Figure 4 In the embodiment of the present application, a computer device is also provided, which can be a server, and the internal structure thereof can be as shown in Figure 4 The computer device comprises a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is used to store the corresponding data in the embodiment. The network interface of the computer device is used to communicate with the external terminal through network connection. The computer program is executed by the processor to implement the above-mentioned method.
[0142] Those skilled in the art can understand that, Figure 4 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied.
[0143] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the method. It can be understood that the computer readable storage medium in the embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0144] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiment methods can be included. Any reference to memory, storage, database or other medium provided by the present application and used in the embodiment can include non-volatile and / or volatile memory. The non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. The volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM, etc.
[0145] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-mentioned system, system and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0146] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0147] The above-described embodiments are merely used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still make modifications to the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A short drama delivery optimization method based on artificial intelligence, characterized in that, The AI-based short drama delivery optimization method includes: The short drama's visual data, audio data, and user interaction data are collected by a multimodal data acquisition unit. A cross-modal feature fusion network is then used to extract and fuse the features of the short drama's visual data, audio data, and user interaction data to generate a short drama fusion feature vector. Based on the short drama fusion feature vector, a deep temporal network is used to model and analyze the user viewing behavior sequence to obtain user interest migration features, and user profile data is constructed based on the user interest migration features. Based on the user profile data and the short drama fusion feature vector, the short drama delivery strategy data is obtained by modeling the delivery environment and predicting delivery actions through a multi-agent reinforcement learning network. The short drama delivery strategy data is used to segment and reorganize the short drama materials, and multiple versions of creative materials are generated through a conditional generative adversarial network to form a short drama creative material library; For the materials in the short drama creative material library, a distributed computing framework is used to analyze the delivery effect, obtain delivery effect data, and generate a delivery resource configuration plan based on the delivery effect data. This includes: grouping the materials in the short drama creative material library according to delivery channels, constructing a channel delivery record table, and calculating exposure, clicks, and conversions from the channel delivery record table to generate an effect statistics matrix; calculating the click-through rate to conversion rate and delivery cost ratio for each delivery channel based on the effect statistics matrix, and combining the click-through rate to conversion rate and delivery cost ratio to form a channel performance index; performing time series expansion on the channel performance index to construct an performance trend chart, and extracting the delivery time period weights from the performance trend chart; performing cross-operation between the delivery time period weights and the channel performance index to obtain a spatiotemporal delivery scoring table; prioritizing each delivery channel according to the spatiotemporal delivery scoring table to generate the delivery effect data; and allocating delivery resources according to the spatiotemporal dimensions based on the delivery effect data to obtain the delivery resource configuration plan.
2. The short drama delivery optimization method based on artificial intelligence according to claim 1, characterized in that, The process involves collecting short drama video data, audio data, and user interaction data through a multimodal data acquisition unit, and then using a cross-modal feature fusion network to extract and fuse features from these data to generate a short drama fusion feature vector. This includes: The short drama video data is sampled in segments according to time windows to obtain a set of sampled segments. Multi-level feature decomposition is performed on each sampled segment in the set of sampled segments to extract color distribution features, texture features and motion trajectory features. The color distribution features, texture features and motion trajectory features are spliced and combined to obtain a visual feature matrix. The audio data is divided into multiple overlapping audio frames. A Fourier transform is performed on each audio frame to obtain a spectrogram. Pitch features, rhythm features, and emotional features are extracted from the spectrogram. The pitch features, rhythm features, and emotional features are combined in time sequence to obtain an audio feature matrix. The user interaction data is sorted according to the timestamp to construct an interaction behavior sequence. The click density, dwell time and interaction frequency in the interaction behavior sequence are calculated, and the click density, dwell time and interaction frequency are combined into a user behavior feature matrix. The temporal correlation between the visual feature matrix, the audio feature matrix, and the user behavior feature matrix is calculated using the cross-modal feature fusion network to generate a temporal correlation matrix. Based on the temporal correlation matrix, the three feature matrices are resampled and interpolated to obtain aligned feature data. By calculating the information entropy of each modal feature in the alignment feature data, the modal fusion weights are determined. The alignment feature data is then weighted and summed using the modal fusion weights and normalized to generate the short drama fusion feature vector.
3. The short drama delivery optimization method based on artificial intelligence according to claim 1, characterized in that, The process involves using a deep temporal network to model and analyze user viewing behavior sequences based on the fused feature vectors of the short dramas, obtaining user interest migration features, and constructing user profile data based on these features, including: Based on the fusion feature vector of the short drama, a viewing feature sequence is constructed, and the viewing feature sequence is divided into time windows to obtain viewing window data; The user viewing behavior sequence within each time window in the viewing window data is sequentially encoded to generate a behavior encoding sequence; The deep temporal network is used to perform temporal correlation analysis on the behavior encoding sequence to extract viewing duration, switching frequency and interaction intensity, and generate a behavior feature matrix. The user interest change rate is calculated based on the behavioral feature matrix, and the user interest change rate is accumulated over time to obtain the user interest migration feature. The user interest migration features are clustered and grouped according to interest topics to construct an interest category table; The user profile data is generated based on the weight distribution of each category in the interest category table.
4. The short drama delivery optimization method based on artificial intelligence according to claim 1, characterized in that, The step involves modeling the delivery environment and predicting delivery actions using a multi-agent reinforcement learning network based on the user profile data and the fused feature vector of the short drama, to obtain short drama delivery strategy data, including: The user profile data is stratified according to the interest dimension to construct a user stratification matrix, and the short drama fusion feature vector is mapped to the user stratification matrix to generate an interest matching degree distribution; The interest matching degree distribution is expanded in a spatiotemporal dimension to establish a delivery state space table, and the delivery time period, delivery area and delivery channel are extracted from the delivery state space table to form delivery environment characteristics; The delivery environment features are input into the multi-agent reinforcement learning network to model the competitive relationship of each delivery channel and obtain the channel competition matrix. The reward value of the campaign is calculated based on the channel competition matrix, and the reward value is used as the optimization target to construct the campaign value table. Select the optimal combination of actions from the action value table to generate an action decision sequence; match the action decision sequence with the action constraints to obtain the short drama action strategy data.
5. The short drama delivery optimization method based on artificial intelligence according to claim 1, characterized in that, The process involves segmenting and reorganizing short drama materials using the short drama delivery strategy data, and generating multiple versions of creative materials through a conditional generative adversarial network to form a short drama creative material library, including: The short drama delivery strategy data is converted into a scene description matrix, and keyframes are extracted from the short drama material based on the scene description matrix to obtain a set of scene segments. Semantic analysis is performed on the set of scene fragments to generate scene semantic vectors, and a fragment association graph is established based on the scene semantic vectors; Based on the scene connection relationships in the fragment association graph, the scene fragment set is reorganized and sorted to generate a reorganized scene sequence; The recombined scene sequence is input into the conditional generative adversarial network to generate multiple initial creative materials with different themes and styles; The initial creative materials are scored creatively, creative theme tags and style feature tags are extracted, and a creative material feature table is constructed. The creative materials in the creative material feature table are categorized and organized according to theme tags and style feature tags to generate the short drama creative material library.
6. The short drama delivery optimization method based on artificial intelligence according to claim 1, characterized in that, The step of prioritizing each delivery channel according to the spatiotemporal delivery scoring table and generating the delivery performance data includes: The scoring data in the spatiotemporal delivery scoring table is normalized to construct a standardized scoring matrix, and the time period performance values of each channel are extracted from the standardized scoring matrix. The time period performance values are aggregated and calculated according to the channel dimension to obtain the channel performance index, and a channel ranking table is constructed based on the channel performance index; the historical ad placement rate of return is calculated for each channel in the channel ranking table, and a revenue distribution curve is generated; the ad placement advantage coefficient is extracted from the revenue distribution curve. The placement advantage coefficient and the channel performance index are weighted and combined to form a comprehensive channel score; The channel comprehensive scores are sorted in descending order to generate a channel priority list; the channel priority list is combined with the time period performance value to generate the campaign performance data.
7. An AI-based short drama delivery optimization system, used to implement the AI-based short drama delivery optimization method as described in any one of claims 1 to 6, characterized in that, The AI-based short drama delivery optimization system includes: The acquisition module is used to acquire short drama screen data, audio data and user interaction data through a multimodal data acquisition unit, and to perform feature extraction and fusion operations on the short drama screen data, audio data and user interaction data using a cross-modal feature fusion network to generate a short drama fusion feature vector; The modeling module is used to model and analyze the user's viewing behavior sequence using a deep temporal network based on the fusion feature vector of the short drama, obtain user interest migration features, and construct user profile data based on the user interest migration features. The prediction module is used to model the delivery environment and predict delivery actions through a multi-agent reinforcement learning network based on the user profile data and the fusion feature vector of the short drama, so as to obtain short drama delivery strategy data. The reorganization module is used to segment and reorganize the short drama materials using the short drama delivery strategy data, and generate multiple versions of creative materials through a conditional generative adversarial network to form a short drama creative material library. The analysis module is used to analyze the performance of materials in the short drama creative material library using a distributed computing framework, obtain performance data, and generate a resource allocation plan based on the performance data. This includes: grouping the materials in the short drama creative material library according to the distribution channels, constructing a channel distribution record table, and calculating exposure, clicks, and conversions from the channel distribution record table to generate a performance statistics matrix; calculating the click-through rate (CTR) and cost-per-view (CPS) ratio for each distribution channel based on the performance statistics matrix, and combining the CTR and CPS ratio to form a channel performance index; performing time-series expansion on the channel performance index to construct a performance trend chart, and extracting the distribution time period weights from the performance trend chart; performing cross-operation between the distribution time period weights and the channel performance index to obtain a spatiotemporal distribution scoring table; prioritizing each distribution channel according to the spatiotemporal distribution scoring table to generate the performance data; and allocating distribution resources according to the spatiotemporal dimensions based on the performance data to obtain the resource allocation plan.
8. A computer device, characterized in that, The device includes a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that the processor, when executing the computer program, implements the artificial intelligence-based short drama delivery optimization method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, the computer program causing a processor, when executed by a processor, to perform the artificial intelligence-based short drama delivery optimization method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent delivery parameter optimization method and system based on data delivery feedback effect
CN120031612A