Intelligent projection algorithm optimization method and system applied to short video projection flow platform
By categorizing content pools by theme and user behavior characteristics on short video streaming platforms, collecting real-time interactive data to generate content adaptation coefficients, and dynamically adjusting the delivery strategy, the problem of inaccurate delivery has been solved, and the efficiency and effectiveness of delivery have been improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FEIKE WANGHONG (HANGZHOU) TECH CO LTD
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-08
AI Technical Summary
Existing short video delivery algorithms lack precise matching in determining the content to be delivered and the target user group, and cannot optimize the delivery strategy in a timely manner based on user feedback, resulting in poor delivery performance and serious waste of resources.
By determining the initial content pool categorized by theme and the reach of the target user group, collecting real-time interaction data, generating content adaptation coefficients, and dynamically adjusting the delivery priority and coverage dimensions, a stable delivery strategy is formed.
It achieves precise matching between the content delivered and user needs, improves delivery efficiency and effectiveness, and adapts to dynamic adjustments based on market changes and user demands.
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] This invention relates to the field of internet advertising technology, and more specifically, to an intelligent advertising algorithm optimization method and system applied to short video streaming platforms. Background Technology
[0002] With the booming development of the short video industry, short video streaming platforms have become a key channel for advertisers and content creators to promote their products and content. However, existing short video streaming algorithms have many shortcomings.
[0003] Currently, most streaming algorithms tend to use a rather crude approach when determining content and target user groups. The initial content pool is typically built by simply listing various short video materials, lacking detailed categorization of themes and precise content tagging, making it difficult to accurately match target users. When defining the reach of the target user group, they only consider basic user attributes such as age and gender, failing to delve deeply into user behavior characteristics and geographic distribution features, thus unable to accurately grasp the users' true needs and preferences.
[0004] Meanwhile, existing ad delivery algorithms lack dynamic adjustment mechanisms. During the campaign, they cannot collect real-time interaction data from the target user group regarding the content, making it impossible to optimize the delivery strategy based on actual user feedback. This makes it difficult to achieve optimal campaign results, hindering advertisers and content creators from achieving efficient promotional goals and wasting significant resources and costs. Summary of the Invention
[0005] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide an intelligent delivery algorithm optimization method applied to a short video streaming platform, the method comprising: Determine the initial content pool for short video streaming and the reach of the target user group. The initial content pool includes short video materials categorized by theme and corresponding content tags. The reach of the target user group includes user behavior characteristics and geographic distribution characteristics. Collect real-time interaction data of the target user group on the initial content pool. The real-time interaction data includes the user's browsing time, interaction operations and jump paths on the short video materials. A content adaptation coefficient is generated based on the real-time interactive data. The content adaptation coefficient is used to reflect the matching relationship between short video materials and target user groups. The delivery priority of the initial content pool and the coverage dimension of the target user group are adjusted according to the content adaptation coefficient. The process involves repeatedly collecting real-time interactive data, generating content adaptation coefficients, and adjusting delivery parameters until a stable delivery strategy is formed and applied to the continuous delivery process.
[0006] In another aspect, embodiments of the present invention also provide an intelligent delivery algorithm optimization system for short video streaming platforms, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the machine-readable storage medium to implement the above-mentioned method.
[0007] Based on the above, this embodiment of the invention determines an initial content pool containing short video materials categorized by theme and corresponding content tags, as well as a target user group reach range covering user behavior characteristics and geographical distribution characteristics. It collects real-time interaction data of the target user group on the initial content pool, enabling timely capture of genuine user feedback on the short video materials. The content adaptation coefficient generated based on this accurately reflects the matching degree between the materials and the target user group. Adjusting the placement priority of the initial content pool and the coverage dimension of the target user group reach range according to the content adaptation coefficient achieves dynamic optimization of the placement strategy, making the placed content more aligned with user needs and improving the accuracy and effectiveness of placement. By iteratively executing the operations of collecting data, generating coefficients, and adjusting parameters until a stable placement strategy is formed and applied to the continuous streaming process, it can adapt to dynamic adjustments in market changes and user needs, effectively improving the efficiency and effectiveness of short video streaming. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of the execution flow of the intelligent delivery algorithm optimization method for short video delivery platforms provided in this embodiment of the invention.
[0009] Figure 2 This is a schematic diagram of exemplary hardware and software components of an intelligent delivery algorithm optimization system for short video streaming platforms provided in an embodiment of the present invention. Detailed Implementation
[0010] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating an intelligent delivery algorithm optimization method for short video streaming platforms, provided by an embodiment of the present invention. The following is a detailed description of this intelligent delivery algorithm optimization method for short video streaming platforms.
[0011] Step S110: Determine the initial content pool for short video streaming and the reach of the target user group. The initial content pool includes short video materials categorized by theme and corresponding content tags. The reach of the target user group includes user behavior characteristics and geographic distribution characteristics.
[0012] Before determining the initial content pool, it is necessary to connect to the short video distribution platform's material repository. This involves using API calls to obtain metadata for all short video materials on the platform, including basic information such as material ID, upload time, format, and duration. Based on a pre-defined set of keywords for distribution direction, text matching is performed on the title and description fields in the metadata to filter out short video materials with relevance exceeding a threshold as a candidate set.
[0013] Determining the reach of the target user group requires accessing the platform's user data center to extract users' basic attributes and behavioral logs. Basic attributes include non-private information submitted during registration, while behavioral logs contain records of user actions within the platform. Analysis of this data identifies user groups with shared behavioral patterns, and combined with geographic data, preliminary reach boundaries are defined.
[0014] Step S111: Select short video materials that match the theme from the content library of the short video streaming platform, and divide the short video materials into multiple theme subsets according to the theme classification criteria. Each theme subset contains at least one short video material.
[0015] The content library's indexing system is built upon a thesaurus, with each thesaurus corresponding to multiple synonyms and related terms. After inputting the core keywords of the streaming topic, the indexing system returns a list of all associated short video clip IDs. For the clips in the list, keyframes and audio text are extracted using the content parsing module and compared with features in the topic classification criteria. Clips meeting the set conformity threshold are categorized into the corresponding topic subset.
[0016] The thematic classification standard adopts a hierarchical structure, with multiple sub-themes under a primary theme, each with clearly defined characteristic parameters. For example, the sub-theme "Ball Games" under the primary theme "Sports" has characteristic parameters including the frequency of ball-like objects appearing in the image and the number of times related sports terms appear. By comparing these characteristic parameters, accurate classification of the materials is achieved.
[0017] Step S112: Add content tags to each short video material. The content tags include the core elements, expression methods, and emotional tendencies of the material.
[0018] The content parsing module is invoked to perform full parsing of the short video footage. Core element extraction is achieved by processing video frames using an object recognition model. The model outputs the categories and confidence levels of objects appearing in the scene, and objects with confidence levels higher than a threshold are selected as core elements. Simultaneously, audio is converted into text through speech recognition, and keywords are extracted to supplement the core elements.
[0019] The determination of the expression mode is based on the technical characteristics of the video, including frame rate changes, transition effects, and camera movement trajectories. These features are extracted using video analysis tools and matched against a pre-defined expression mode feature library to determine the corresponding expression mode category. Emotional tendency is then assessed using a text sentiment analysis model that processes the audio-to-text transcription and subtitles, outputting results indicating positive, neutral, calm, or lively sentiments.
[0020] Step S1121: Share the complete content of each short video material and extract the core elements appearing in the short video material. The core elements include at least one identifiable key object, which includes people, objects, scenes, and events.
[0021] The complete data stream of short video footage is transmitted to the parsing server via a content distribution interface. The parsing server uses a segmentation processing mechanism to divide the video into multiple segments of fixed duration. For each segment, a pre-trained object detection model is used for frame-level processing, and the model outputs the category, location coordinates, and duration of appearance of each key object.
[0022] Person recognition is achieved through facial feature extraction algorithms. Feature values are calculated for faces appearing in the image and compared with a known person database. Unknown persons are tagged as "unknown person + frequency of occurrence." Object recognition covers common item categories, determining object names through feature matching. Scene recognition combines geographic information tagging and image element combination analysis; for example, the simultaneous appearance of "beach," "waves," and "sun umbrella" indicates a "beach scene." Event extraction uses temporal sequence analysis to combine consecutively occurring related actions into event descriptions; for example, "an object being lifted and then moved" can be described as a "moving event."
[0023] Step S1122: Analyze the shooting techniques, editing style, and narrative structure of the short video material to determine the expression mode of the short video material. The expression mode includes documentary, animation, interview, demonstration, and plot type description.
[0024] Shooting technique analysis is achieved using a lens parameter extraction tool. This tool retrieves parameters such as focal length, aperture, and shutter speed from video metadata and combines this with image stability data to determine whether the shot is static or moving, and whether it's a long or short shot. Editing style is determined by matching transition effects against a database, analyzing the frequency of different transition types, and combining this with the speed of scene transitions to identify the rhythmic characteristics of the edit.
[0025] Narrative structure analysis employs text summarization algorithms to process the video's narration and dialogue text, extracting timelines, character relationships, and the sequence of events to construct a narrative graph. This narrative graph is then compared to pre-defined structural templates for documentary, animation, and other expressive styles; the template with the highest match corresponds to the expressive style of the video.
[0026] Step S1123: Determine the emotional tendency conveyed by the short video material through the color tone, background music, and language style elements in the short video material. The emotional tendency description includes qualitative expressions such as positive, neutral, calm, and lively.
[0027] Image tone analysis uses color extraction tools to obtain the RGB value distribution of each frame, calculates the average saturation and brightness of the tone, and generates a tone curve. The curve is then compared with tone models of positive (high saturation, high brightness) and calm (low saturation, medium brightness) emotional tendencies to obtain the tone matching degree.
[0028] Background music analysis uses an audio feature extraction module to obtain parameters such as rhythm, melody, and volume. Music with fast rhythms and large volume variations tends to be lively, while music with slow rhythms and stable volume tends to be calm. Language style analysis uses a text tone analysis tool to process speech-to-text, extracting the frequency of interjections and modal particles, and combining this with sentence length to determine the emotional tone of the language.
[0029] Based on the analysis results of the overall color tone, background music, and language style, a weighted fusion method was used to determine the final emotional tendency. The weights of each element were obtained by training a machine learning model based on historical labeled data.
[0030] Step S1124: Combine the core element, the expression method, and the emotional tendency description into content tags according to a preset format. Each content tag consists of multiple descriptive items, which are separated by delimiters.
[0031] The default format uses the structure "core element | expression method | sentiment tendency". Multiple descriptive items within the core element are separated by commas, and the expression method and sentiment tendency are each a separate descriptive item. For example, if the core element contains "basketball, court, audience", the expression method is "documentary", and the sentiment tendency is "positive", then the content tags would be "basketball, court, audience | documentary | positive".
[0032] The choice of delimiters is based on text processing compatibility. Commas are used to separate descriptive items of the same type, while vertical bars are used to separate different types, ensuring accurate splitting of content in subsequent retrieval and analysis. After tags are generated, their format is checked for correctness using a validation tool; tags that do not conform to the format will be returned for regeneration.
[0033] Step S1125: Add a weight identifier to each descriptive item in the content tag. The weight identifier reflects the prominence of the descriptive item in the short video material. The prominence is determined based on the frequency, duration and influence of the descriptive item.
[0034] Frequency calculation measures the number of times a key object corresponding to a given descriptor appears in the video; the more times, the higher the base weight. Duration statistics calculate the proportion of the total duration of the key object's continuous appearance in the video to the total video duration; the higher the proportion, the higher the duration weight. Influence is assessed by evaluating the key object's contribution to the video's theme expression. This contribution is calculated using a theme association model; the higher the association, the higher the influence weight.
[0035] The weighting is based on a percentage system and is calculated using the formula "weight = frequency weight × 0.3 + duration weight × 0.4 + influence weight × 0.3". The weighting of each description item is appended to the description item and enclosed in parentheses, such as "basketball (85), court (60), audience (30) | documentary (90) | positive (75)".
[0036] Step S1126: Bind the content tags with added weights to the short video materials and store them in the attribute information of the short video materials to establish a one-to-one correspondence between the content tags and the short video materials.
[0037] By calling the attribute information editing function of short video clips through the material management interface, content tags are written into the "tag field" of the attribute information. In the database, a unique index is created for each short video clip ID, and the tag information and the clip ID are stored as key-value pairs. At the same time, an inverted index of tags is built, with each tag item corresponding to a list of all clip IDs containing that tag, ensuring that the corresponding short video clip can be quickly retrieved by tag.
[0038] During the binding process, a transaction processing mechanism is employed to ensure the synchronized storage of tag information and materials, preventing data inconsistencies. After binding is complete, the correspondence between tags and materials is verified through a validation interface; only after successful verification is the binding confirmed as successful.
[0039] Step S113: Integrate the short video materials with added content tags and the corresponding theme subsets into an initial content pool, and establish an index directory for the initial content pool. The index directory contains the correspondence between theme subsets and short video materials and the retrieval path of content tags.
[0040] The integration process is implemented through a directory management system. The system creates a directory tree according to the hierarchical structure of thematic subsets, with each thematic subset corresponding to a directory node. The short video clip IDs with added tags are categorized by thematic subset and stored under the corresponding directory nodes. Simultaneously, metadata such as the storage path, tag information, and format of each clip is recorded to form a clip list.
[0041] The index directory is constructed using a B+ tree structure, with leaf nodes storing material IDs and corresponding tag index pointers. The retrieval path for content tags is generated using a path planning algorithm. This algorithm determines the optimal retrieval path from the root directory to the target material based on the hierarchical relationship and relevance of the tags. The index directory is updated periodically; an automatic index update mechanism is triggered when new materials are added or material tags are modified.
[0042] Step S114: Extract user registration information and historical behavior records from the user database of the short video streaming platform. The user registration information includes the basic information filled in by the user, and the historical behavior records include the themes and content tags of the short video materials that the user has previously viewed and interacted with.
[0043] The system connects to the user database via a data access interface, which employs an encrypted transmission protocol to ensure secure data transmission. The extracted user registration information does not contain any privacy-sensitive data, but only basic information that users have voluntarily disclosed, such as their interests and occupation. During the extraction process, the data is anonymized to remove information that could potentially identify the user.
[0044] The extraction of historical behavior records covers all user operation logs within a preset time period, including the ID of the video viewed, the start and end times of the viewing, the type of interaction (comment, share, favorite, etc.), and the corresponding video topic and tags. These records are grouped by user ID and stored as user behavior sequences, with each sequence containing a timestamp and behavior details.
[0045] Step S115: Based on the geographic field in the user registration information and the location association data in the historical behavior records, classify the geographic distribution characteristics of the user. The geographic distribution characteristics include the administrative division information of the user's permanent residence area and active area.
[0046] The geographic field in user registration information is analyzed by parsing the text to extract administrative region codes. Location-related data from historical behavior records includes the geographic region corresponding to the IP address used when a user posted content, and the frequency of browsing videos tagged with geographic regions. This data is then input into a geographic analysis model, which uses clustering algorithms to group users into different geographic clusters.
[0047] The determination of a user's permanent residence area is based on the percentage of time a user spends actively in each area. The area with the highest percentage and a duration exceeding a threshold is designated as the permanent residence area. Active areas are determined by the frequency of user interaction across different areas; areas with an interaction frequency above the average are marked as active areas. Geographical distribution characteristics are stored in the form of administrative region codes and names, including information at the provincial, municipal, and district levels.
[0048] Step S116: Analyze the frequency of attention and interaction depth of short video materials with different themes and content tags in the user's historical behavior records, and extract user behavior features, including the user's preferred theme type, content tag combination and interaction method.
[0049] Attention frequency is calculated for each topic and content tag, showing the ratio of the number of times a user views related videos to the total number of views within a unit of time. Interaction depth is measured by an interaction intensity index, which is calculated by combining factors such as the weight of interaction type (e.g., comments have a higher weight than likes), the length and quality of the interactive content, etc.
[0050] The co-occurrence of content tags is analyzed using association rule mining algorithms to identify frequently occurring tag combinations. The frequency of different interaction methods is statistically analyzed, and the most frequent interaction methods are identified as user-preferred interaction methods. These analytical results are then integrated to form a user behavior feature vector.
[0051] Step S1161: Filter out all short video materials that the user has viewed from the user's historical behavior records, and extract the theme and content tags of all the short video materials.
[0052] Iterate through the user's browsing history logs, extract the video ID for each log entry, and then query the media database for the corresponding theme and content tags using the video ID. Aggregate all video themes and tags for the same user, remove duplicates, and form the user's browsing theme set and tag set.
[0053] For deleted or expired video IDs, they are marked as "invalid records" and logged, and are not included in the analysis. During the extraction process, a relationship matrix is established between users and topics / tags, where each matrix element represents the number of times a user has viewed a particular topic or tag.
[0054] Step S1162: Count the number of times users view short video materials for each topic and the total viewing time within a preset period, and calculate the attention frequency of each topic. The attention frequency is the proportion of the number of times the topic is viewed to the total number of times all topics are viewed.
[0055] The preset period can be set according to actual needs, such as 7 days, 30 days, etc. Within the period, the number of views and total viewing time are counted by topic. The total viewing time is the sum of the viewing time of all videos under that topic. The frequency of attention is calculated by division, that is, the number of views for a single topic is divided by the total number of views for all topics, and the result is rounded to two decimal places.
[0056] After calculation, the attention frequencies are normalized to ensure that the sum of the attention frequencies for all topics is 1. The results are then stored in a user behavior feature database and associated with the user ID for easy subsequent querying and analysis.
[0057] Step S1163: Analyze user interaction behavior on short video materials for each content tag. The interaction behavior includes commenting, sharing, and collecting. Calculate the number of interactions and interaction quality for each content tag. The interaction quality is comprehensively evaluated based on factors such as comment length and sharing scope.
[0058] The number of interactions is the sum of the number of comments, shares, and favorites (dimensionless integer). The interaction quality is calculated by weighting and combining the comment length index and the sharing scope index. The comment length index is the ratio of the average number of characters per comment to the average comment length on the platform (dimensionless), and the sharing scope index is the ratio of the number of shares to external platforms to the total number of shares (dimensionless).
[0059] The weights of the two are 0.6 and 0.4 respectively. After concatenation, a two-dimensional vector is formed by [comment length index × 0.6, sharing range index × 0.4]. Then, the interaction quality value in the range of 0-1 is obtained by normalization.
[0060] Step S1164: Sort by attention frequency and select the top N topics as the user's preferred topic types.
[0061] All topics are sorted by frequency of attention in descending order, and the top N topics are selected. The value of N can be adjusted according to the diversity of user behavior, and is usually set to cover more than 70% of the topics with the highest user attention frequency. The selected topic types are stored in list format, associated with user IDs, and serve as an important component of user behavior characteristics.
[0062] During the sorting process, if topics with the same frequency of attention appear, they will be sorted by total browsing time, with topics with longer browsing time being selected first.
[0063] Step S1165: Perform combination analysis on content tags that have more than a set number of user interactions and have a higher than a set quality of interaction, identify content tag combinations that appear at least M times simultaneously, and determine the content tag combinations as user-preferred content tag combinations.
[0064] Content tags that meet both interaction frequency and quality criteria are selected, and a tag co-occurrence matrix is constructed. Matrix elements represent the number of times two tags appear simultaneously in a user's viewed video. When the co-occurrence count reaches M, the two tags are considered a combination. For combinations of multiple tags, a frequent itemset mining algorithm is used for identification, with the minimum support set to M.
[0065] The identified content tag combinations are sorted by frequency of occurrence, and the combinations with higher frequency are selected as the content tag combinations preferred by users and stored in the form of sets, with each set containing two or more tags.
[0066] Step S1166: Statistically analyze the distribution characteristics of the interaction methods used by users during the interaction process to determine the interaction methods preferred by users.
[0067] The system tracks the frequency of user interactions such as commenting, sharing, and saving, calculating the percentage of each method's usage relative to the total interactions. The two interaction methods with the highest percentages are identified as the user's preferred interaction methods. If the usage frequency of a particular interaction method is significantly higher than other methods (e.g., exceeding 50%), only that method is selected.
[0068] The distribution characteristics are stored in the form of pie chart data, reflecting the usage ratio of different interaction methods.
[0069] Step S1167: Integrate the user's preferred topic type, content tag combination, and interaction method into user behavior characteristics to form a description file of user behavior characteristics.
[0070] The user's preferred topic types, content tag combinations, interaction methods, and corresponding distribution characteristics are integrated and a description file is generated in JSON format. The file contains the user ID, feature generation time, each feature item, and its detailed data. The description file is stored in a user feature database and a timed update mechanism is set up to ensure the timeliness of the features.
[0071] During the generation process, feature items are standardized to unify data formats and naming conventions, facilitating cross-system calls and analysis.
[0072] Step S117: Integrate user behavior characteristics and geographic distribution characteristics to define the reach of the target user group and form a description document of the reach. The description document includes the correspondence between user behavior characteristics and geographic distribution characteristics and a set of identifiers of the covered users.
[0073] A correlation analysis model is used to match user behavior characteristics with geographic distribution characteristics. The model calculates the similarity of behavioral characteristics among user groups in different regions and merges user groups with similarity scores exceeding a threshold. Based on the merging results, the boundaries of the reach area are defined, including both geographic scope and behavioral characteristic range.
[0074] The description document is in XML format and includes geographic codes, behavioral characteristic parameters, and a list of user identifiers. The user identifier list is a collection of user IDs that meet the reach criteria, retrieved through database queries and updated periodically. The description document is stored in the reach management system and serves as a crucial basis for traffic delivery.
[0075] Step S120: Collect real-time interaction data of the target user group on the initial content pool. The real-time interaction data includes the user's browsing time, interaction operations and jump paths for short video materials.
[0076] A data collection script is embedded in the short video playback page. The script captures user actions through an event listener mechanism. When a user opens a video, the script records the start timestamp; when the user closes the video or navigates to a different page, it records the end timestamp. The difference between the two timestamps is the viewing duration.
[0077] Interactive operations are implemented by binding click events. When a user clicks buttons such as comment, share, or favorite, the script records the operation type, timestamp, and corresponding video ID. Navigation paths are captured by listening for changes in page URLs, recording information such as the navigation time from the current video page to other pages and the target page URL. The collected data is transmitted to the data processing center in real time, using a streaming protocol to ensure data continuity and integrity.
[0078] Step S121: Embed a data collection module in the content display interface of the short video streaming platform. The data collection module is used to record the operation behavior of the target user group on the short video materials in the initial content pool.
[0079] The data acquisition module is integrated into the front-end code of the content display interface as an SDK. The module consists of three core units: event capture, data encapsulation, and transmission control. The event capture unit listens for user actions by binding DOM events, including event types such as click, scroll, and keydown. Each event type corresponds to a preset operation behavior mapping table. For example, when the click event triggers the comment button, it is automatically mapped to "comment operation".
[0080] The data encapsulation unit formats the captured event data according to a preset JSON structure, including fields such as event type, trigger timestamp, user identifier, video clip ID, and operation location coordinates. For privacy-sensitive data, such as the user identifier, a hash algorithm is used for one-way encryption to ensure that the original information cannot be reverse-engineered.
[0081] The transmission control unit employs an adaptive transmission strategy, dynamically adjusting the transmission frequency and data packet size based on network conditions. When the network is readily available, it uses real-time transmission mode, sending each event data record immediately upon generation. When the network is congested, it automatically switches to batch transmission mode, merging multiple event data records into a single data packet and sending it at preset intervals. Simultaneously, a breakpoint resumption mechanism is enabled during transmission; if transmission is interrupted, it resumes sending the unsuccessfully transmitted data from the point of interruption after reconnection.
[0082] Step S122: When a user in the target user group opens the short video material, the data collection module starts timing until the user closes the short video material or jumps to other content, and records the corresponding time period as the user's browsing time of the short video material.
[0083] Upon detecting that a user has opened a short video clip, the data acquisition module immediately calls the system clock interface to obtain the current timestamp as the starting point for timing. This timestamp is then associated with the video clip ID and the user identifier and stored in the local cache. During the timing process, the module periodically polls the page status to determine if the video is playing. If playback is detected to be paused, the timing is paused; when playback resumes, the timing resumes from the paused point.
[0084] When a user closes the video or navigates to other content, the module again calls the system clock interface to obtain the current timestamp as the timing endpoint. The browsing duration is calculated by subtracting the starting timestamp from the endpoint timestamp, accurate to the millisecond. After calculation, the browsing duration and related information are packaged and sent to the data processing center via the transmission control unit, and the corresponding timing data in the local cache is cleared.
[0085] Step S123: Monitor the user's actions while browsing short video materials, and record the time, duration and associated object of each action as interactive action data. The actions include clicking, swiping, commenting, sharing and collecting.
[0086] Click detection is achieved by binding the click event of the button element. Each time a click is triggered, the time of the click, the button's DOM path (used to locate the specific operation object), and the mouse coordinates at the time of the click are recorded. For comment operations, additional events capturing changes in the comment input box are captured, recording the start and submission times of the comment, calculating the duration, and anonymizing the comment content to remove potentially sensitive information such as phone numbers and email addresses.
[0087] The swipe operation is implemented by listening to the touchmove or mousemove events, recording the start and end times of the swipe, calculating the duration, and simultaneously collecting the coordinate change sequence during the swipe process. The swipe direction and distance are analyzed through these coordinate changes. When a sharing operation is triggered, the identifier of the target sharing platform, the generation time of the sharing link, and the completion status of the sharing operation (success or failure) are recorded.
[0088] The "favorite" action is linked to the state change of the "favorite" button. When the button switches from "not favorited" to "favorited," the time of the favorite is recorded. If the state switches again later, the time of the unfavorite is recorded. Both are included in the interaction data. All action records include the corresponding video material ID and user identifier to ensure data traceability.
[0089] Step S124: Track the user's switching order between different short video materials in the initial content pool, record the triggering method and page path of the user jumping from one short video material to another, and form jump path data.
[0090] A unique identifier parameter, such as "videoId=xxx", is embedded in the URL of the short video playback page. This parameter changes as the user switches between different short videos. The data acquisition module captures these URL changes by listening to the `hashchange` or `popstate` events of the `window` object, extracts the `videoId` before and after the change, and determines the switching order.
[0091] The trigger method is determined based on user actions before the URL change. If the user clicks on a video thumbnail in the recommendation list before the change, the trigger method is marked as "recommendation list click"; if the user performs a left or right swipe, it is marked as "swipe to switch"; if the user clicks on an associated link within the video, it is marked as "internal link jump".
[0092] The page paths traversed are recorded by tracking all page addresses during the URL changes, arranged chronologically, with each page address associated with a corresponding dwell time (the time difference between entering and leaving the page). For example, if a user navigates from video A to the homepage and then from the homepage to video B, the page path is recorded as "Video A → Homepage → Video B," with the dwell time on the homepage marked for each. After the navigation path data is generated, it is stored in association with the user's identifier, forming the user's browsing history profile.
[0093] Step S1241: Set a tracking identifier in the page jump node of the short video streaming platform. The tracking identifier is used to identify the user's operation of jumping from the current page to other pages.
[0094] Page navigation nodes include clickable elements such as navigation menus, recommended video lists, related video links, and back buttons. Add a custom attribute "data-trace-id" to the HTML tags of these elements as a tracking identifier. The attribute value is a unique identifier for the node, such as "nav-home", "recommend-list-1", "related-video-5", etc.
[0095] The mapping relationship between tracking identifiers and target pages is stored in a configuration file. The configuration file uses key-value pairs, where the key is the tracking identifier and the value is the URL template of the target page. When a user clicks a node with a tracking identifier, the data acquisition module reads the identifier and queries the corresponding target page URL through the configuration file to determine the target location for redirection. Simultaneously, the module records the tracking identifier in the redirection path data for subsequent analysis of the redirection source.
[0096] Step S1242: When a user performs a jump operation on a short video material page in the initial content pool, the tracking identifier records the identifier of the current short video material, the identifier of the target short video material to jump to, and the operation type triggered by the jump, including clicking a recommended link and swiping to switch.
[0097] The current short video clip is identified by the `videoId` parameter value carried in the page URL, and the target short video clip is identified by the `videoId` parameter value in the URL of the page after the jump. The operation type triggered by the jump is determined by analyzing the correlation between user behavior and tracking identifiers: if an element with "data-trace-id=recommend-video-x" is clicked before the jump, the operation type is marked as "clicking a recommended link"; if no click operation is detected, but a swipe event is detected and the page URL changes, it is marked as "swiping to switch"; if an element with "data-trace-id=related-video-y" is clicked, it is marked as "clicking a related link".
[0098] This information is written to the redirection operation log in real time. The log entries include fields such as operation timestamp, user ID, current video ID, target video ID, and operation type.
[0099] Step S1243: Record the intermediate pages visited by the user during the redirection process. The intermediate pages include the platform homepage, category pages, and search results pages. Record the access order and dwell time of the intermediate pages in chronological order.
[0100] When a user navigates from the current short video page to the target short video page, if they pass through other pages, the data collection module records the access information of the intermediate pages by listening for page load events (load event) and page unload events (unload event). When the load event is triggered, the page URL and entry timestamp are recorded; when the unload event is triggered, the exit timestamp is recorded. The difference between the two is the duration of the user's stay on that intermediate page.
[0101] The type of intermediate page is identified by its URL path. For example, " / home" corresponds to the platform homepage, " / category / sports" corresponds to the sports category page, and " / search?keyword=xxx" corresponds to the search results page. The URLs, entry times, exit times, and dwell times of the intermediate pages are arranged in chronological order to form an intermediate page access sequence, which is used as part of the redirection path data.
[0102] Step S1244: Integrate the current short video material identifier, jump trigger method, intermediate page access order and target short video material identifier according to the timeline to form a single jump path record.
[0103] Sort by timestamp, the current short video clip identifier (jump start point), jump trigger method (e.g., "click recommended link"), intermediate page access sequence (if any), and target short video clip identifier (jump end point) are arranged in sequence to form a complete jump path record. The record is in JSON array format, and each element contains the event type (e.g., "start point", "trigger method", "intermediate page", "end point") and the corresponding data content.
[0104] For example, a redirect path record might look like this: [{"type":"Starting point","value":"video123"},{"type":"Trigger method","value":"Click the recommended link"},{"type":"Intermediate page","value":[{"url":" / home","stayTime":5000}]},{"type":"End point","value":"video456"}]. This record fully reflects the user's entire process of jumping from video123 to video456.
[0105] Step S1245: Associate multiple jump path records of the same user within the same time period, analyze the continuity and correlation of the jump paths, and identify the user's browsing patterns in the initial content pool.
[0106] Multiple redirection path records for the same user are associated through user identifiers, and the same time period refers to a preset time window, such as within one hour. The path matching algorithm analyzes whether the endpoint and starting point of adjacent redirection path records are consistent. If they are consistent, they are determined to be continuous paths and merged into a longer browsing sequence.
[0107] Association analysis is achieved by calculating the frequency of different video identifiers appearing together in the navigation path; video combinations with frequencies exceeding a threshold are considered related. Simultaneously, the distribution of navigation trigger methods is statistically analyzed to identify user-preferred navigation methods. Through these analyses, user browsing patterns are summarized, such as "a preference for jumping from sports videos to fitness videos via recommended links" and "a tendency to continuously swipe to switch between videos on the same topic."
[0108] Step S1246: Summarize all jump path records into jump path data. The jump path data includes user identifiers, a set of jump path records, and path analysis conclusions. The path analysis conclusions are used to describe the user's habitual patterns of switching short video materials in the initial content pool.
[0109] The redirection path data is aggregated using user identifiers as keys. The value corresponding to each key contains a set of all redirection path records for that user and path analysis conclusions. The path analysis conclusions are generated based on the path coherence and relevance analysis results, and use natural language to describe the user's habit patterns, such as "Users often browse food-themed videos by swiping, and on average, after every 3 videos, they will jump to kitchenware recommendation videos via recommended links."
[0110] The redirect path data is stored in a dedicated path database, which adopts a document-oriented database structure and supports querying and statistical analysis by user ID, time range, video theme, and other dimensions.
[0111] Step S130: Generate a content adaptation coefficient based on the real-time interactive data. The content adaptation coefficient is used to reflect the matching relationship between short video materials and target user groups.
[0112] Real-time interaction data of the target user group is extracted from the real-time database and grouped by short video clip ID. Each group contains all viewing durations, interaction data, and navigation path data for that clip. Multi-dimensional analysis is performed on each group: the average viewing duration as a percentage of the total video duration is calculated; the total number of interactions and average interaction quality are statistically analyzed; and the jump-away ratio (the proportion of times users jump from this video to other videos out of the total number of views) is analyzed.
[0113] These analysis results are used as input parameters and substituted into the content fit coefficient calculation model. The model calculates the content fit coefficient for each short video clip through weighted summation. The higher the coefficient value, the stronger the match between the clip and the target user group. After calculation, the content fit coefficient is associated with the video clip ID and stored, and a coefficient distribution report is generated to reflect the distribution of fit levels among different clips.
[0114] Step S131: Extract the browsing duration, interactive operation data and jump path data corresponding to each short video material from the real-time interactive data, group them according to the user identifier in the target user group, and obtain the interaction record of each user for each short video material.
[0115] Real-time interaction data of the target user group is retrieved from the real-time database through a data query interface. The search criteria are the set of identifiers that identify users belonging to the target user group. The search results are grouped according to the combination of "user identifier + video material ID". Each group forms an interaction record, which includes all browsing time records, interaction operation record list and jump path records of the user for that video.
[0116] For browsing time, the total browsing time, average browsing time, and longest browsing time for the user are calculated. For interactive actions, the number of times each type of action is performed and the total interaction quality score are recorded. For navigation paths, the number of times the user navigates from this video to other videos and the topic distribution of the navigation targets are recorded. This data is integrated into the interaction log to form the basic data for the user-video interaction matrix.
[0117] Step S132: Analyze the ratio of a single user's browsing time for a single short video clip to the total duration of that short video clip, and calculate the user's attention index for that short video clip by combining the number and type of user's interactive operations on that short video clip.
[0118] The percentage of viewing time is calculated by dividing the actual viewing time of a single user on the short video by the total duration of the short video. The resulting percentage ranges from 0 to 1, and it reflects the degree to which the user has watched the video content completely.
[0119] For interactive actions, corresponding weight coefficients are assigned to different types of interactive actions. For example, the weight coefficient for commenting is higher than that for liking, and the weight coefficient for sharing is higher than that for saving. The weight coefficients are determined through correlation analysis of historical interactive data and video conversion effects, and the weight coefficients of all types of interactive actions are normalized to ensure that their values are within the range of 0 to 1.
[0120] The number of various interactive operations performed by the user on the short video material is counted. The number of each interactive operation is multiplied by the corresponding weight coefficient and summed to obtain the weighted sum of interactive operations. Then, this sum is divided by the average of the weighted sums of the user's interactive operations on all short video materials within the same time period to obtain the standardized value of the interactive operation. The standardized value of the interactive operation is controlled within the range of 0 to 1.
[0121] The attention index is generated by weighted concatenation of the browsing time percentage and the standardized value of interaction. Specifically, the two values are scaled according to preset weights (e.g., browsing time percentage weight 0.6, interaction standardization value weight 0.4) and then combined to form a two-dimensional vector. The two elements in the two-dimensional vector retain their original dimensional attributes and are used for subsequent adaptation value calculation.
[0122] Step S133: Based on the user's jump path data in the initial content pool, analyze the reasons why the user jumps from one short video material to another, determine whether the jump is caused by the current short video material not meeting the user's expectations, and count the proportion of each short video material that the user actively jumps away from.
[0123] The reason for redirection is analyzed by comparing the differences in themes and content tags between the current video and the target video: if the theme similarity between the target video and the current video is below a threshold, and the user's viewing time in the current video is below average, the reason for redirection is determined to be "the current video does not meet expectations"; if the theme similarity between the target video and the current video is above a threshold, it is determined to be "the user continues browsing on the same theme".
[0124] The number of times users actively leave the video is considered as "the current video does not meet expectations." The total number of views for the video is the total number of times users open the video (regardless of whether they finish watching it). The percentage of users actively leaving the video is calculated as "number of active leaves / total number of views." The higher the percentage, the more likely the video is to be abandoned by users because it does not meet their expectations.
[0125] Step S134: Combine the attention index and the proportion of users who are redirected to another site to calculate the fit value of a single short video material for a single user.
[0126] The attention index and the proportion of users who were redirected away are weighted and concatenated to form a two-dimensional adaptation vector. The vector format is [attention index × 0.6, (1 - proportion of users redirected away) × 0.4]. Through vector normalization, the two-dimensional vector is converted into a single adaptation value (range 0-1). The normalization process uses the min-max method, i.e., adaptation value = (vector magnitude - minimum magnitude) / (maximum magnitude - minimum magnitude). The higher the adaptation value, the better the user's match with the material.
[0127] The adaptation value also ranges from 0 to 1, with a higher value indicating better compatibility between the user and the video. After calculation, the adaptation value is associated with the user identifier and video material ID and stored as the basis for subsequent summary calculations.
[0128] Step S135: Summarize the adaptation values of all users in the target user group for the same short video material, and take the comprehensive result as the content adaptation coefficient between the short video material and the target user group.
[0129] For the same short video clip, we collect the fit values of all users in the target user group and calculate the average of these fit values as a preliminary result. At the same time, we consider the representativeness weight of each user in the target group; for example, active users have a higher weight than occasionally active users. The weight values are set based on the user's historical interaction frequency.
[0130] The content adaptation coefficient is a "weighted average," calculated by multiplying each user's adaptation value by its corresponding weight, summing the results, and then dividing by the total weights. The coefficient ranges from 0 to 1, with values above 0.8 indicating high adaptation, 0.5-0.8 indicating medium adaptation, and below 0.5 indicating low adaptation. After calculation, a coefficient list is generated, containing the video clip ID and its corresponding content adaptation coefficient.
[0131] Step S136: Establish a description rule for the content adaptation coefficient. The description rule includes a description of the matching correlation degree corresponding to the numerical range of the content adaptation coefficient. The description of the matching correlation degree is used to explain the fit between the short video material and the target user group reflected by the content adaptation coefficient.
[0132] The description rules adopt a hierarchical definition approach, dividing the content matching coefficient into multiple consecutive numerical ranges, each range corresponding to a specific degree of matching correlation. For example: 0.8-1.0 corresponds to "high matching, the target user group shows strong interest in the material, high browsing completion rate, active interaction, and very few jumps due to not meeting expectations"; 0.5-0.8 corresponds to "medium matching, the target user group has some interest in the material, moderate browsing completion rate, some interaction, and occasionally jumps due to not meeting expectations"; 0-0.5 corresponds to "low matching, the target user group has low interest in the material, low browsing completion rate, few interactions, and frequently jumps due to not meeting expectations".
[0133] The description rules are stored in the rule database and can be dynamically adjusted based on actual traffic delivery results and user feedback to ensure that the rules accurately reflect the correspondence between coefficients and matching degree.
[0134] Step S137: Store the content adaptation coefficient of each short video material and its corresponding description rules in the coefficient database, and establish an association index between the content adaptation coefficient and the short video material and the target user group.
[0135] The coefficient database adopts a relational database structure and includes fields such as video material ID, target user group identifier, content adaptation coefficient, and description rule ID. The description rule ID is associated with a description rule entry in the rule database, and can be used to query the corresponding matching degree description.
[0136] The associated index adopts a composite index structure, using "target user group identifier + video material ID" as the index key to improve the efficiency of querying coefficients by group and material. Simultaneously, an index is created using content adaptation coefficients as the sorting key, supporting quick queries of video material lists within specific coefficient ranges. A transaction mechanism is enabled during data storage to ensure consistency between coefficient data and associated information.
[0137] Step S140: Adjust the delivery priority of the initial content pool and the coverage dimension of the target user group's reach based on the content adaptation coefficient.
[0138] The content adaptation coefficients of all short video materials are retrieved from the coefficient database. The materials are then sorted from highest to lowest coefficient value, and the sorting results serve as the direct basis for adjusting the delivery priority. For materials in the first half of the delivery priority list, their display probability in the recommendation list is increased; for materials in the second half, their display probability is reduced or delivery is paused.
[0139] Simultaneously, the content tags and target user characteristics of highly compatible creatives are analyzed, and common features are extracted as new coverage dimensions to supplement the reach, while existing coverage dimensions that do not match the characteristics of highly compatible creatives are removed. After adjustments are made, a priority list and an updated reach description document are generated as parameters for the next round of ad delivery.
[0140] Step S141: Extract the content adaptation coefficient of each short video material from the coefficient database, sort the short video materials in the initial content pool according to the size of the content adaptation coefficient, and use the sorting result as the basis for adjusting the delivery priority.
[0141] A dataset consisting of "video material ID + content adaptation coefficient" is extracted from the coefficient database using SQL queries. After extraction, a data cleaning script is used to remove outliers, including data with empty content adaptation coefficients and values exceeding a preset reasonable range. Once cleaned, the dataset is imported into a sorting engine, which uses a quicksort algorithm to sort the data in descending order based on the content adaptation coefficient.
[0142] During the sorting process, if two short video clips have the same content fit coefficient, their historical conversion rates are further compared, with the clip with the higher conversion rate ranked higher. After the sorting results are generated, they are synchronized to the placement priority index table. The index table assigns priority weights to each video clip according to its ranking position. The weight value increases with the ranking position, and the weight difference between adjacent positions remains at a fixed ratio.
[0143] Step S142: The short video materials in the first half of the ranking results are designated as the priority delivery group, and the short video materials in the second half of the ranking results are designated as the secondary priority delivery group. The short video materials in the priority delivery group will receive more display opportunities during the delivery process.
[0144] Extract the total number of materials from the sorting results, and determine the boundary between the first and second halves using integer division. If the total number is odd, the first half contains (total number + 1) / 2 materials, and the second half contains (total number - 1) / 2 materials.
[0145] Set display quota parameters for the priority and secondary priority ad groups respectively. The quota parameter for the priority ad group should be 2-3 times that of the secondary priority ad group, with the specific multiple dynamically adjusted based on the ad spend budget. The display quota parameters are linked to the ad spend scheduling system. When allocating display resources, the scheduling system allocates exposure opportunities to the two groups of creatives according to the quota ratio, with creatives from the priority ad group occupying a higher position in the user's recommendation list.
[0146] Step S143: Perform aggregate analysis on the content tags of short video materials in the priority delivery group, extract the content tag combination that appears more frequently than the set frequency, use the content tag combination as a supplementary screening condition for the initial content pool, select new short video materials that match the content tag combination from the content library of the short video delivery platform, and add them to the corresponding theme subset of the initial content pool.
[0147] The tag parsing tool is used to break down the content tags of all materials in the priority delivery group, resulting in a list of individual tags and the combination relationships between tags. A frequency statistics algorithm is used to calculate the frequency of each tag combination, setting the frequency to 30%-50% of the total number of materials in the priority delivery group; the specific value is adjusted according to the diversity requirements of the materials.
[0148] The tag combinations that meet the frequency criteria are converted into a filter expression, using the format "tag A AND tag BAND...". This filter expression is then passed to the content library's search interface, which returns a list of new content IDs that meet the criteria. These new content items must not have appeared in the initial content pool.
[0149] New materials are subject to theme matching verification. The theme classification model is used to determine the theme subset to which they belong. After the verification is successful, the material ID and associated information are written to the corresponding directory in the initial content pool, and the version number of the index directory is updated.
[0150] Step S144: Analyze the interaction records of the short video materials of the priority delivery group in the target user group, extract the common features of the user behavior characteristics and geographical distribution characteristics of the target user group, and use the common features as the core coverage dimension of the reach.
[0151] The system extracts all user interaction records for the prioritized ad group from the real-time interaction database, groups them by user ID, and calculates the probability of occurrence of each behavioral feature (such as interaction method and browsing duration distribution) in each group. The user behavioral features are then aggregated using a clustering algorithm, and the feature values corresponding to the cluster centers are the common behavioral features.
[0152] The commonalities of regional distribution characteristics are analyzed using a heatmap generation tool. The tool maps the user's regional code to geographic coordinates, counts the user density at each coordinate point, and forms a dense area where the density is higher than a threshold. The administrative division information of the dense area is the common regional characteristics.
[0153] Common behavioral features and common regional features are integrated into core coverage dimensions. The dimensions are stored in key-value pairs, where the key is the feature type (such as "interaction method" or "resident area") and the value is the specific feature content.
[0154] Step S145: Compare the original coverage dimensions and core coverage dimensions in the reach of the target user group, retain the overlapping coverage dimensions, expand the newly added feature descriptions in the core coverage dimensions, and form the adjusted coverage dimensions.
[0155] The original coverage dimensions and the core coverage dimensions are compared at the field level using a difference comparison algorithm. Dimensions with the same fields and identical content are marked as "overlapping", dimensions with the same fields but different content are marked as "to be updated", and fields unique to the core coverage dimensions are marked as "new".
[0156] For dimensions marked "to be updated," replace the existing content with the content of the core coverage dimension; for "new" dimensions, add them directly to the coverage dimension set; "overlapping" dimensions remain unchanged. After the adjustment is complete, generate a coverage dimension change list, which records the operation type (retain, update, add) and specific content for each dimension.
[0157] Step S146: Based on the adjusted coverage dimensions, select newly added users who meet the criteria from the user database of the short video streaming platform and include them in the reach of the target user group, while removing the original users who do not meet the adjusted coverage dimensions.
[0158] Step S1461: Decompose the adjusted coverage dimension into multiple filtering conditions. Each filtering condition corresponds to a feature description in the coverage dimension. The filtering conditions include specific parameters of user behavior features and administrative division range of geographical distribution features.
[0159] The condition parser converts the key-value pairs covering dimensions into structured filtering conditions. For example, "interaction method = comment" is converted into "user_behavior.interaction_type='comment'", and "resident region = province A, city B" is converted into "user_region.resident_areaIN('province A, city B')".
[0160] For numerical feature parameters (such as browsing time thresholds), the filtering criteria use range expressions (such as "browsing time >= 30 seconds"); for categorical feature parameters (such as topic preferences), the filtering criteria use inclusion expressions (such as "preferred_topicIN('technology', 'sports')"). All filtering criteria are combined to form a filtering condition set, and the condition sets are linked together using the logical operator "AND".
[0161] Step S1462: Extract feature information of users who are not included in the reach of the target user group from the user database of the short video streaming platform. The feature information includes the user behavior characteristics and geographical distribution characteristics of the users.
[0162] By executing the "NOTIN" statement through the user database query interface, user IDs already within the reach range are excluded, and the feature information of the remaining users is extracted. The feature information is indexed by user ID and contains the association data between the user behavior feature table and the geographic feature table. During the extraction process, sensitive fields (such as specific addresses) are blurred (only the municipal administrative unit level is retained).
[0163] The extracted feature information is stored in a temporary data table with the same structure as the feature table of users within the reach range, which facilitates subsequent comparison operations.
[0164] Step S1463: Compare the feature information of users not included in the reach scope with the filtering conditions. Users who meet all the filtering conditions are identified as new users, and the unique code of the new users is recorded.
[0165] The rule engine applies a set of filtering conditions to each user feature in the temporary data table. Users who meet all conditions are marked as "compliant," and those who do not are marked as "non-compliant." The rule engine uses a pipelined processing model, where the next condition is only validated after the previous one has passed, thus improving filtering efficiency.
[0166] For users marked as "matching", extract their unique code (such as user ID) to generate a list of new users. The list includes user ID and feature matching score (the higher the score, the better the match with the screening criteria).
[0167] Step S1464: Extract the current user behavior characteristics and geographic distribution characteristics of each user from the original user identifier set in the reach of the target user group, and compare them with the adjusted coverage dimension filtering conditions.
[0168] By linking the user identifier set to the user feature database, the latest feature information of each user is obtained (ensuring that the data has been updated within the last 24 hours). The same rule engine as in step S1463 is used for condition comparison, and users who do not meet any of the filtering conditions are marked as "to be removed".
[0169] Step S1465: If the characteristics of an existing user do not meet any of the screening criteria, then the user is identified as a user to be removed, and the unique code of the user to be removed is recorded.
[0170] A second verification is performed on users to be removed to check whether their behavior characteristics in the past 7 days have fluctuated temporarily (such as failing to meet the conditions due to accidental factors). If the probability of fluctuation exceeds the set threshold, the removal is temporarily suspended; otherwise, they are included in the list of users to be removed.
[0171] Step S1466: In the user identifier set of the target user group's reach range, add the unique code of the new user and delete the unique code of the user to be removed to form an updated user identifier set.
[0172] The set operation tool performs the operation of "existing set - set to be removed + new set". The operation process adopts a transaction mechanism to ensure the atomicity of addition and deletion and avoid data inconsistency. The updated user ID set is stored in a distributed cache with a 10-minute expiration time to ensure real-time access for subsequent accesses.
[0173] Step S1467: Based on the updated set of user identifiers, extract the corresponding user information from the user database and update the description document of the reach of the target user group.
[0174] Extract the basic user information (non-sensitive fields) corresponding to the updated user identifier set and replace the original user information section in the description document. Simultaneously, update the document's version number and update timestamp. The version number should be in the format "major version number.minor version number," with the minor version number incrementing by 1 after each update.
[0175] After the update is complete, the document verification tool is used to check the format integrity. Once the verification is successful, the document is synchronized to the main document library of the reach management system.
[0176] Step S147: Record the results of the initial content pool's delivery priority adjustment and the results of the target user group's reach coverage dimension adjustment to form an adjustment log. The adjustment log includes a comparison of parameters before and after the adjustment and the basis for the adjustment.
[0177] The adjustment logs are in a structured format. Each log entry includes the adjustment time, adjustment type (priority adjustment / coverage dimension adjustment), parameter values before adjustment, parameter values after adjustment, and adjustment basis (such as changes in content adaptation coefficients or changes in user characteristics).
[0178] Log data is stored in an audit database, which uses a time-series partitioned table, storing data by day and retaining the most recent 90 days of log data. Simultaneously, the log data is synchronized to a visualization monitoring platform, which displays parameter adjustment trends using line charts, facilitating operations personnel to trace the adjustment process.
[0179] Step S150: Repeatedly execute the operations of collecting real-time interactive data, generating content adaptation coefficients, and adjusting delivery parameters until a stable delivery strategy is formed and applied to the continuous delivery process.
[0180] The cyclic process is scheduled by the workflow engine, which sets the trigger and termination conditions for the cycle. Before each cycle begins, the workflow engine checks whether the current feed status is normal (e.g., no system failure, traffic not exceeded). If it is normal, the cyclic process is triggered; otherwise, it enters a waiting state and retryes every 5 minutes.
[0181] During the loop, the three stages of real-time interactive data acquisition, content adaptation coefficient generation, and deployment parameter adjustment are executed sequentially. The next stage is only started after the previous stage is completed and the result is verified, forming a closed-loop control. The execution result of each stage is recorded in the loop status table, which includes the stage name, start time, end time, status (success / failure), and reason for failure (if any).
[0182] For example, step S151: Set the start conditions for cyclic execution, wherein the start conditions are the initial content pool and the reach of the target user group are determined for the first time and the streaming begins.
[0183] The startup conditions are implemented through a status monitor. The monitor monitors the status of the index directory of the initial content pool and the status of the reach description document in real time. When both are marked as "effective" and the "stream status" field of the streaming scheduling system is "running", a loop startup signal is triggered.
[0184] The start signal is sent to the workflow engine's start interface, which records the start time and initial parameters (such as the content pool version and user-scope version for the first loop) as the baseline data for the loop.
[0185] Step S152: At the beginning of each loop, trigger the data acquisition module to collect real-time interaction data of the target user group on the current initial content pool, generate a real-time interaction data snapshot according to the collection rules, and store it.
[0186] At the start of each loop, the workflow engine sends a collection instruction to the data acquisition module. This instruction includes the batch number of the current loop, the set of target user identifiers, and the list of content pool asset IDs. The data acquisition module filters the data according to the instruction range, collecting only the interaction behavior of the specified users with the specified assets.
[0187] Real-time interactive data snapshots adopt an incremental acquisition mode, which only includes new interaction records compared to the previous snapshot. The snapshot file is named in the format of "cycle batch number_acquisition timestamp.snapshot" and stored in a distributed file system in the Parquet format to improve subsequent reading efficiency.
[0188] Step S153: Based on the newly collected real-time interactive data, recalculate the content adaptation coefficient between each short video material and the target user group, and update the content adaptation coefficient value in the coefficient database.
[0189] The content adaptation coefficient calculation module is invoked. The module loads newly collected real-time interactive data snapshots, calculates the adaptation value of a single user for a single material by user group, and then summarizes them into the content adaptation coefficient of the material through a weighted average algorithm (the weight is the representative weight of the user in the target group).
[0190] After the calculation is completed, the coefficient database is updated through a database transaction. The old value is backed up to the historical coefficient table, and the new value is written to the current coefficient field. The update operation is accompanied by a timestamp, marked as "cycle batch number_update time".
[0191] Step S154: Based on the updated content adaptation coefficient, repeat the operation of adjusting the delivery priority of the initial content pool and the coverage dimension of the target user group's reach to generate new adjustment results.
[0192] The adjustment process reuses the logic of steps S141 to S147, but updates the input parameters to the latest content adaptation coefficients and user interaction data. During the adjustment process, all intermediate results are marked with the current cycle batch number for easy comparison with historical adjustment results.
[0193] The new adjustment results are stored in the adjustment result database, which records core data such as the deployment priority list, coverage dimension list, and user identifier set for each cycle.
[0194] Step S155: Compare the results of this adjustment with the results of the previous adjustment, analyze the magnitude of changes in the delivery priority and the changes in the coverage dimensions, and determine whether the adjustment results tend to stabilize.
[0195] The magnitude of changes in ad placement priority is calculated using a similarity algorithm. The algorithm converts the list of creative IDs from two sorting processes into a sequence, calculates the edit distance of the sequence, and the ratio of the edit distance to the total length represents the magnitude of the change. Changes in coverage dimensions are statistically analyzed using a difference comparison tool, which tracks the number of newly added, deleted, and modified dimensions and the scale of users involved.
[0196] Set stability criteria: the change in campaign priority is less than 10% for a continuous period, and the proportion of users whose coverage dimensions have changed is less than 5% for a continuous period. When both criteria are met, the adjustment results are considered to be stable.
[0197] Step S156: If the change in the adjustment result remains within a preset stable range in K consecutive cycles, and there are no new or deleted core features in the coverage dimension, then a stable delivery strategy is determined to be formed.
[0198] The K value is set according to the traffic delivery scenario, usually 3-5 times, with a time interval of 24 hours between each cycle (to ensure sufficient data to reflect user behavior trends). Core features refer to dimensions that play a decisive role in the reach (such as topic preferences and core regions), which are determined by a feature importance assessment model. The model outputs an importance score for each dimension, and the top 5 dimensions with the highest scores are considered core features.
[0199] After K consecutive cycles meet the stability criteria, the workflow engine generates a "stability strategy confirmation signal". The signal includes data such as the average adjustment range and a list of core features within the stability period.
[0200] Step S157: Solidify the stable delivery strategy into a delivery execution plan, which includes the final delivery priority ranking of the initial content pool, the final coverage dimension of the target user group's reach, and the application rules of the content adaptation coefficient.
[0201] The execution plan for ad delivery is generated using a standardized template, which includes the following sections: effective time of the plan, content pool configuration (material ID, priority weight, display quota), user scope configuration (user identifier set, coverage dimension parameters), adaptation coefficient application rules (such as the mapping relationship between coefficient threshold and display frequency), and exception handling mechanism (such as replacement rules when materials become invalid).
[0202] After the scheme is generated, its integrity is ensured by digital signature. The signing key is kept by the system administrator, and the signature information is attached to the end of the scheme.
[0203] Step S158: Synchronize the streaming execution plan to the streaming execution module of the short video streaming platform, replace the original temporary streaming parameters and apply them to the subsequent continuous streaming process, and continuously monitor the streaming effect to make corresponding fine adjustments.
[0204] The traffic delivery execution module receives the traffic delivery execution plan through an interface, parses the configuration parameters in the plan, and updates the internal delivery rule engine. The replacement process adopts a canary release strategy, first switching 10% of the traffic to the new plan, observing for 2 hours without any anomalies, and then gradually expanding to 100% of the traffic.
[0205] Continuous monitoring is achieved through the performance analysis module, which provides real-time statistics on metrics such as ad impressions, clicks, and conversion rates. When these metrics are compared with the expected metrics of the plan, a fine-tuning mechanism is triggered if the deviation exceeds 15%. Fine-tuning includes minor adjustments to ad priority weights and local optimizations of the user scope. The results of the fine-tuning do not change the core framework of the plan.
[0206] Figure 2 The illustration shows exemplary hardware and software components of an intelligent video streaming algorithm optimization system 100 for short video streaming platforms, which can implement the ideas of this application, according to some embodiments of this application. For example, processor 120 can be used in the intelligent video streaming algorithm optimization system 100 for short video streaming platforms and to perform the functions in this application.
[0207] The intelligent ad-hoc algorithm optimization system 100 applied to short video streaming platforms can be a general-purpose server or a special-purpose server; both can be used to implement the intelligent ad-hoc algorithm optimization method for short video streaming platforms described in this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the load.
[0208] For example, the intelligent video streaming algorithm optimization system 100 applied to a short video streaming platform may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and various forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the intelligent video streaming algorithm optimization system 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The methods of this application can be implemented according to these program instructions. The intelligent video streaming algorithm optimization system 100 also includes an I / O interface 150 between the computer and other input / output devices.
[0209] For ease of explanation, only one processor is described in the intelligent ad-hoc algorithm optimization system 100 applied to a short video streaming platform. However, it should be noted that the intelligent ad-hoc algorithm optimization system 100 applied to a short video streaming platform in this application may also include multiple processors. Therefore, the steps performed by one processor described in this application may also be performed jointly by multiple processors or individually. For example, if the processor of the intelligent ad-hoc algorithm optimization system 100 applied to a short video streaming platform executes steps A and B, it should be understood that steps A and B may also be performed jointly by two different processors or individually by one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.
[0210] Furthermore, this embodiment of the invention also provides a readable storage medium, wherein computer-executable instructions are preset in the readable storage medium, and when the processor executes the computer-executable instructions, the above-mentioned intelligent delivery algorithm optimization method applied to the short video streaming platform is implemented.
[0211] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A method for optimizing intelligent delivery algorithms applied to short video streaming platforms, characterized in that, The method includes: Determine the initial content pool for short video streaming and the reach of the target user group. The initial content pool includes short video materials categorized by theme and corresponding content tags. The reach of the target user group includes user behavior characteristics and geographic distribution characteristics. Collect real-time interaction data of the target user group on the initial content pool. The real-time interaction data includes the user's browsing time, interaction operations and jump paths on the short video materials. A content adaptation coefficient is generated based on the real-time interactive data. The content adaptation coefficient is used to reflect the matching relationship between short video materials and target user groups. The delivery priority of the initial content pool and the coverage dimension of the target user group are adjusted according to the content adaptation coefficient. The process involves repeatedly collecting real-time interactive data, generating content adaptation coefficients, and adjusting delivery parameters until a stable delivery strategy is formed and applied to the continuous delivery process.
2. The intelligent delivery algorithm optimization method applied to short video streaming platforms according to claim 1, characterized in that, The determination of the initial content pool and the reach of the target user group for short video streaming includes: Short video materials that match the theme are selected from the content library of the short video streaming platform. The short video materials are divided into multiple theme subsets according to the theme classification criteria. Each theme subset contains at least one short video material. Add content tags to each short video clip, and the content tags include the core elements involved in the clip, the expression methods, and the emotional tendencies of the clip. The short video materials with added content tags and their corresponding theme subsets are integrated into an initial content pool, and an index directory for the initial content pool is established. The index directory contains the correspondence between theme subsets and short video materials and the retrieval path for content tags. User registration information and historical behavior records are extracted from the user database of the short video streaming platform. The user registration information includes the basic information filled in by the user, and the historical behavior records include the themes and content tags of the short video materials that the user has viewed and interacted with in the past. Based on the geographic field in the user registration information and the location association data in the historical behavior records, the geographic distribution characteristics of users are divided. The geographic distribution characteristics include the administrative division information of the user's permanent residence area and active area. Analyze the frequency of attention and interaction depth of short video materials with different themes and content tags in the user's historical behavior records to extract user behavior features, which include the user's preferred theme type, content tag combination and interaction method; By associating and integrating user behavior characteristics and geographic distribution characteristics, the reach of the target user group is defined, and a description document of the reach is formed. The description document includes the correspondence between user behavior characteristics and geographic distribution characteristics and a set of identifiers of the covered users.
3. The intelligent delivery algorithm optimization method applied to short video streaming platforms according to claim 2, characterized in that, The process of adding content tags to each short video clip includes: Share the complete content of each short video clip, extract the core elements appearing in the short video clip, the core elements include at least one identifiable key object, the key object includes people, objects, scenes, and events; Analyze the shooting techniques, editing style, and narrative structure of the short video materials to determine their expression methods, which include documentary, animation, interview, demonstration, and narrative description. The emotional tendency conveyed by the short video material is determined by the color tone, background music, and language style elements in the short video material. The emotional tendency description includes qualitative expressions such as positive, neutral, calm, and lively. The core elements, the expression methods, and the emotional tendency descriptions are combined into content tags according to a preset format. Each content tag consists of multiple descriptive items, which are separated by delimiters. Add a weight identifier to each descriptive item in the content tag. The weight identifier reflects the prominence of the descriptive item in the short video material. The prominence is determined based on the frequency, duration and influence of the descriptive item. The content tags with added weights are bound to the short video clips and stored in the attribute information of the short video clips, establishing a one-to-one correspondence between the content tags and the short video clips.
4. The intelligent delivery algorithm optimization method applied to short video streaming platforms according to claim 2, characterized in that, The analysis examines the frequency and depth of user engagement with short video content on different themes and tags in the user's historical behavior records to extract user behavior characteristics, including: Filter out all short video materials that the user has viewed from the user's historical behavior records, and extract the theme and content tags of all the short video materials; The system counts the number of times users view short video materials for each topic and the total viewing time within a preset period, and calculates the attention frequency for each topic. The attention frequency is the proportion of the number of times a topic is viewed to the total number of times all topics are viewed. Analyze user interaction behavior with short video materials for each content tag, including comments, sharing, and collection. Calculate the number of interactions and interaction quality for each content tag, with interaction quality comprehensively evaluated based on factors such as comment length and sharing scope. Based on the frequency of attention, select the top N topics with the highest frequency of attention as the topic types preferred by users; Perform combination analysis on content tags that have more than a set number of user interactions and have a higher than a set quality of interaction, identify content tag combinations that appear at least M times simultaneously, and determine the content tag combinations that the user prefers. Analyze the distribution characteristics of the interaction methods used by users during the interaction process to determine the interaction methods preferred by users; The user's preferred topic types, content tag combinations, and interaction methods are integrated into user behavior characteristics, forming a description file of user behavior characteristics.
5. The intelligent delivery algorithm optimization method applied to short video streaming platforms according to claim 1, characterized in that, The collection of real-time interaction data between the target user group and the initial content pool includes: A data collection module is embedded in the content display interface of the short video streaming platform. The data collection module is used to record the operation behavior of the target user group on the short video materials in the initial content pool. When a user in the target user group opens the short video material, the data collection module starts timing until the user closes the short video material or jumps to other content, and records the corresponding time period as the user's browsing time of the short video material; Monitor user actions while browsing short video content, and record the time, duration, and associated objects of each action as interactive action data. The actions include clicking, swiping, commenting, sharing, and saving. Track the user's switching order between different short video clips in the initial content pool, record the triggering method and the page path the user takes when jumping from one short video clip to another, and form jump path data; The data acquisition module aggregates browsing time, interactive operation data, and jump path data into real-time interactive data snapshots at preset time intervals. Each real-time interactive data snapshot contains all operation records of the target user group on the initial content pool within that time interval. Add a time identifier and a user identifier to each real-time interactive data snapshot. The time identifier includes the start time and end time of data collection, and the user identifier includes a unique code of the user in the target user group that generated the operation. Real-time interactive data snapshots with added time and user identifiers are stored in a real-time database, and an index is established to correspond the real-time interactive data with short video materials and target user groups in the initial content pool.
6. The intelligent delivery algorithm optimization method applied to short video streaming platforms according to claim 5, characterized in that, The tracking method monitors the user's switching order between different short video clips in the initial content pool, records the triggering method and page path traversed when the user jumps from one short video clip to another, forming jump path data, including: Set a tracking identifier at the page jump node of the short video streaming platform. The tracking identifier is used to identify the user's operation of jumping from the current page to other pages. When a user performs a jump operation on a short video material page in the initial content pool, the tracking identifier records the identifier of the current short video material, the identifier of the target short video material, and the type of operation triggered by the jump, including clicking a recommended link and swiping to switch. Record the intermediate pages that the user visits during the navigation process. The intermediate pages include the platform homepage, category pages, and search results pages. Record the access order and dwell time of the intermediate pages in chronological order. The current short video material identifier, jump trigger method, intermediate page access order and target short video material identifier are integrated in timeline to form a single jump path record; By associating multiple redirection path records of the same user within the same time period, the continuity and correlation of the redirection paths can be analyzed to identify the user's browsing patterns in the initial content pool. All redirection path records are aggregated into redirection path data, which includes user identifiers, a set of redirection path records, and path analysis conclusions. The path analysis conclusions are used to describe the user's habitual patterns of switching short video materials in the initial content pool.
7. The intelligent delivery algorithm optimization method applied to short video streaming platforms according to claim 1, characterized in that, The generation of content adaptation coefficients based on the real-time interactive data includes: Extract the browsing duration, interactive operation data, and jump path data corresponding to each short video material from the real-time interactive data, group them according to the user identifier in the target user group, and obtain the interaction record of each user for each short video material; Analyze the ratio of a single user's viewing time for a single short video clip to the total duration of that short video clip, and combine this with the number and type of user interactions on that short video clip to calculate the user's attention index for that short video clip; Based on the user's jump path data in the initial content pool, analyze the reasons why users jump from one short video clip to another, determine whether the jump is due to the current short video clip not meeting the user's expectations, and count the proportion of each short video clip that users actively jump away from. By combining the attention index and the proportion of users who are redirected and leave, the suitability value of a single short video material for a single user is calculated. The compatibility values of all users in the target user group for the same short video material are aggregated, and the overall result is taken as the content compatibility coefficient between the short video material and the target user group. A description rule is established for the content adaptation coefficient. The description rule includes a description of the matching correlation degree corresponding to the numerical range of the content adaptation coefficient. The description of the matching correlation degree is used to explain the fit between the short video material and the target user group reflected by the content adaptation coefficient. The content adaptation coefficient of each short video material and its corresponding description rules are stored in the coefficient database, and an index is established to associate the content adaptation coefficient with the short video material and the target user group.
8. The intelligent delivery algorithm optimization method applied to short video streaming platforms according to claim 1, characterized in that, The step of adjusting the delivery priority of the initial content pool and the coverage dimension of the reach of the target user group based on the content adaptation coefficient includes: Extract the content adaptation coefficient of each short video material from the coefficient database, sort the short video materials in the initial content pool according to the size of the content adaptation coefficient, and use the sorting result as the basis for adjusting the delivery priority; The short video materials in the first half of the ranking results are designated as the priority delivery group, and the short video materials in the second half of the ranking results are designated as the secondary priority delivery group. The short video materials in the priority delivery group receive more display opportunities during the delivery process. The content tags of short video materials in the priority delivery group are aggregated and analyzed to extract content tag combinations that appear more frequently than a set frequency. These content tag combinations are used as supplementary filtering conditions for the initial content pool. New short video materials that match the content tag combinations are selected from the content library of the short video delivery platform and added to the corresponding theme subset of the initial content pool. Analyze the interaction records of short video materials of the priority delivery group among the target user group, extract the common features of user behavior characteristics and geographical distribution characteristics of the target user group, and use the common features as the core coverage dimension of the reach; By comparing the original coverage dimensions and core coverage dimensions in the reach of the target user group, overlapping coverage dimensions are retained, and the newly added feature descriptions in the core coverage dimensions are expanded to form the adjusted coverage dimensions. Based on the adjusted coverage dimensions, newly added eligible users are selected from the user database of the short video streaming platform and included in the reach of the target user group, while existing users who do not meet the adjusted coverage dimensions are removed. Record the results of adjusting the initial content pool's delivery priority and the coverage dimension of the target user group's reach, forming an adjustment log. The adjustment log includes a comparison of parameters before and after the adjustment and the basis for the adjustment.
9. The intelligent delivery algorithm optimization method applied to a short video streaming platform according to claim 8, characterized in that, The process involves selecting newly added eligible users from the user database of the short video streaming platform based on the adjusted coverage dimensions, including them in the reach of the target user group, and removing existing users who no longer meet the adjusted coverage dimensions. This includes: The adjusted coverage dimension is decomposed into multiple filtering conditions. Each filtering condition corresponds to a feature description in the coverage dimension. The filtering conditions include specific parameters of user behavior features and administrative division range of geographical distribution features. Extract feature information of users who are not included in the reach of the target user group from the user database of the short video streaming platform. The feature information includes the user behavior characteristics and geographical distribution characteristics of the users. The feature information of users not included in the reach scope is compared with the filtering conditions. Users who meet all the filtering conditions are identified as new users, and the unique code of the new users is recorded. Extract the current user behavior characteristics and geographic distribution characteristics of each user from the original set of user identifiers within the reach of the target user group, and compare them with the adjusted screening criteria of the coverage dimensions. If the characteristics of an existing user do not meet any of the screening criteria, the user is identified as a user to be removed, and the unique code of the user to be removed is recorded. In the user identifier set within the reach of the target user group, add the unique code of the new user and delete the unique code of the user to be removed to form an updated user identifier set; Based on the updated set of user identifiers, extract the corresponding user information from the user database and update the description document of the reach of the target user group.
10. A smart delivery algorithm optimization system for short video streaming platforms, characterized in that, The device includes a processor and a memory, the memory being connected to the processor. The memory is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the memory to implement the intelligent delivery algorithm optimization method for short video streaming platforms as described in any one of claims 1-9.