Audio and video high-precision pushing intelligent algorithm matching method and system

CN122838656APending Publication Date: 2026-09-29BEIJING LONGSHI HAIZHI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610986283.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

现有算法多聚焦于单次推荐的局部匹配,忽略从历史转移模式中提取的可通行路径,使得推荐结果难以有效引导用户沿其潜在兴趣链条扩展探索,加剧了推荐内容的同质化问题,限制了推送的多样性

Benefits of technology

[0055]本发明获取用户交互行为数据流和音视频内容特征数据集,通过时序解析提取行为模式向量并进行多模态融合生成内容表征向量,能够全面捕捉用户行为动态与内容语义的深层关联,显著提升推送匹配的精准度。基于行为模式向量中的消费切换序列构建演化图谱,可直观反映用户在不同内容类型间的转移规律,为后续路径推荐提供结构化支撑。在行为模式向量中识别中断交互特征并生成排斥分布图,将累积频次超过抑制阈值的类型标记为规避节点,从而在演化图谱中执行绕行路径搜索。这一机制有效避免向用户推送已明确反感的类型,降低误推率与用户流失风险,同时通过提取未经过规避节点的连通路径,实现内容推荐的屏蔽化与安全化。沿绕行路径从对应内容类型中匹配内容表征向量,生成扩展候选集,再综合演化图谱、排斥分布图与行为模式向量计算综合匹配度,确保推荐结果既契合用户潜在兴趣,又规避负面体验,显著提升推送的适应性与用户满意度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838656A_ABST
    Figure CN122838656A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of audio and video pushing, and more particularly to an audio and video high-precision pushing intelligent algorithm matching method and system. User behavior data stream and audio and video content feature dataset are obtained, behavior mode vectors are extracted through time sequence analysis, content representation vectors are generated in combination with multi-modal fusion, an evolution graph is constructed based on consumption switching sequences, an exclusion distribution graph is generated by identifying interrupt interaction features, exclusion types exceeding a threshold value are taken as avoidance nodes for graph traversal search in the evolution graph to generate an expansion candidate set, and audio and video content is sorted and pushed after comprehensive calculation of matching degrees. The method improves pushing precision and user satisfaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio and video push technology, and in particular to a method and system for intelligent algorithm matching for high-precision audio and video push. Background Technology

[0002] In the field of audio and video content recommendation, current conventional approaches mainly rely on building interest models based on users' historical behavior data and employing algorithms such as collaborative filtering, content filtering, or deep neural networks to achieve personalized recommendations. Collaborative filtering methods analyze the similarity of user group behavior to uncover common preferences for recommendations; content filtering matches audio and video tags, keywords, and other metadata with users' historical preferences; deep learning methods further utilize time-series models such as recurrent neural networks and self-attention mechanisms to capture short-term fluctuations in users' interests and integrate multimodal content features to improve representation accuracy. These approaches are widely used in scenarios such as video platforms, music streaming, and live streaming services, and often iteratively optimize recommendation strategies through real-time feedback.

[0003] Interrupted interactions during user viewing (such as quickly skipping, closing early, switching content, etc.) often reflect a clear aversion to that type of content. However, traditional methods usually treat such behaviors as noise or simply reduce their weight, lacking a systematic modeling mechanism to accumulate and analyze the frequency of rejection. This leads to the frequent repetition of recommended content of types that users have clearly rejected, resulting in a decline in user experience.

[0004] User interests are not isolated but rather exhibit regular consumption patterns across different content types (e.g., shifting from short videos to sports events and then to documentaries). Existing algorithms often focus on localized matching in a single recommendation, neglecting viable paths extracted from historical transition patterns. This makes it difficult for recommendations to effectively guide users to explore along their potential interest chains, exacerbating the homogenization of recommended content and limiting the diversity of push notifications. Therefore, current technologies still need improvement in refining the handling of user rejection signals and the evolutionary relationship between content types. Summary of the Invention

[0005] This invention provides a method and system for intelligent algorithm matching for high-precision audio and video push, which can solve the problems in the prior art.

[0006] A first aspect of this invention provides a method for intelligent algorithm matching for high-precision audio and video push, comprising:

[0007] Acquire user interaction behavior data streams and audio / video content feature datasets;

[0008] The user interaction behavior data stream is analyzed in time series to extract behavior pattern vectors. At the same time, the audio and video content feature dataset is fused in multimodal mode to generate content representation vectors.

[0009] Based on the consumption switching sequence of users between different content types in the behavioral pattern vector, an evolutionary map reflecting the content type transfer relationship is constructed.

[0010] Identify user interruption interaction features in behavior pattern vectors, mark the content types corresponding to interruption interaction features as exclusion types, count the cumulative frequency of each exclusion type, and generate an exclusion distribution map.

[0011] The exclusion type with a cumulative frequency exceeding the preset suppression threshold in the exclusion distribution map is taken as the avoidance node. A graph traversal search starting from the non-avoidance node is performed in the evolution graph. Connected paths that do not pass through the avoidance node in the traversal path are extracted as bypass paths. Audio and video content matching the content representation vector is extracted from the corresponding content type along the bypass path to generate an extended candidate set.

[0012] The comprehensive matching degree of audio and video content in the extended candidate set is calculated based on the evolutionary map, rejection distribution map and behavior pattern vector, and audio and video content is pushed to users according to the ranking result of the comprehensive matching degree.

[0013] The user interaction behavior data stream is subjected to time-series analysis to extract behavior pattern vectors. Simultaneously, multimodal fusion processing is performed on the audio and video content feature dataset to generate content representation vectors, including:

[0014] The interaction events in the user interaction behavior data stream are sorted by timestamp to construct a time-series interaction sequence, and a fixed step size is set to segment the time-series interaction sequence by a sliding window.

[0015] The interaction type, interaction duration, and interaction interval within the window are extracted as behavioral feature dimensions. After normalizing the behavioral feature dimensions, the statistical distribution characteristics within the window and the rate of change of behavioral intensity between adjacent windows are calculated. The statistical distribution characteristics and the rate of change of behavioral intensity are concatenated and encoded to generate a behavioral pattern vector.

[0016] Visual representations are generated by extracting body contours and texture information from visual frame sequences in audio and video content feature datasets, and audio representations are generated by extracting frequency components and energy distribution from audio waveform sequences through short-time Fourier transform.

[0017] The system detects whether there is an alignment deviation between visual and audio representations in the time dimension. When an alignment deviation exists, it calculates the sliding cross-correlation function of visual and audio representations in the time dimension, extracts the time offset corresponding to the global maximum value of the sliding cross-correlation function, and then translates the visual representation according to the time offset and sums it with the audio representation element by element to generate a content representation vector.

[0018] Based on the user consumption switching sequences between different content types in the behavioral pattern vector, an evolutionary map reflecting the content type transfer relationship is constructed, including:

[0019] Extract user consumption switching sequences between different content types from the behavior pattern vector, identify content type labels of adjacent consumption behaviors in the consumption switching sequence, and combine adjacent content type labels into ordered pairs as content type switching pairs according to the chronological order.

[0020] The frequency of occurrence of content type switching pairs in the consumption switching sequence is statistically analyzed, and a content type transition matrix is ​​constructed with content type as the row and column index and occurrence frequency as the matrix element.

[0021] The transition probability of each content type to other content types is calculated by row normalization of the content type transition matrix. The eigenvector of the content type transition matrix is ​​calculated, and the eigenvector corresponding to the largest eigenvalue is extracted as the content type importance vector.

[0022] Calculate the Shannon entropy for the transition probability distribution of each row in the content type transition matrix, multiply the Shannon entropy by the corresponding element value in the content type importance vector to generate a comprehensive stability index, and sort and label the node level of the content types according to the comprehensive stability index.

[0023] An evolutionary graph reflecting the content type transfer relationship is constructed using content type as nodes, transfer probability as directed edge weights, and node level as node attributes.

[0024] Identify user interruption interaction features in the behavior pattern vector, label the content types corresponding to the interruption interaction features as exclusion types, count the cumulative frequency of each exclusion type, and generate an exclusion distribution map, including:

[0025] Extract the interaction duration and interaction completion fields from the behavior pattern vector. Group and standardize the interaction duration field according to the content type label to obtain the standardized duration. Construct a feature space with the standardized duration and interaction completion field and perform isolation tree anomaly detection to obtain anomaly score.

[0026] Interaction events with abnormal scores exceeding a preset abnormal threshold are marked as interrupted interaction features, and the corresponding content type tags are extracted and marked as exclusion types;

[0027] The cumulative frequency is constructed by counting the number of occurrences of the exclusion type in the behavior pattern vector. The cumulative frequency is normalized to generate the exclusion strength. The distribution uniformity of the cumulative frequency is evaluated to obtain the concentration index.

[0028] In the consumption switching sequence, the location of the rejection type is retrieved and a location index set is constructed. For each location in the location index set, the content type label within a fixed step size is extracted to construct a pre-label matrix. Principal component analysis is performed on the pre-label matrix to extract the dominant feature vector. The loading coefficient of each content type label in the dominant feature vector is calculated as the inducing factor intensity. The content type label with the largest inducing factor intensity is selected as the rejection inducing factor type.

[0029] A rejection distribution map is generated using rejection type as the horizontal axis, rejection intensity as the vertical axis, rejection cause type as the node label, and concentration index as the transparency parameter.

[0030] Principal component analysis is performed on the pre-label matrix to extract the dominant feature vector. The loading coefficients of each content type label in the dominant feature vector are calculated as the inducing factor strength, including:

[0031] The pre-label matrix is ​​converted into a numerical matrix by one-hot encoding. Singular value decomposition is performed on the numerical matrix to obtain the left singular vector matrix and the singular value sequence. The maximum singular value is selected from the singular value sequence, and the column vector corresponding to the maximum singular value in the left singular vector matrix is ​​extracted as the dominant feature vector.

[0032] Extract the component values ​​corresponding to each content type label from the dominant feature vector as the pattern contribution, and count the number of times each content type label appears in the position before the exclusion type in the preceding label matrix as the neighbor frequency.

[0033] Normalization is performed on both the pattern contribution and the neighbor frequency. The normalized pattern contribution and the normalized neighbor frequency are then weighted and fused to obtain the loading coefficient of each content type tag in the dominant feature vector as the inducing force intensity.

[0034] In the evolutionary graph, a graph traversal search starting from non-avoided nodes is performed. Connected paths that do not pass through avoided nodes are extracted as bypass paths. Audio and video content matching the content representation vector is extracted from the corresponding content type along the bypass paths to generate an extended candidate set, including:

[0035] In the evolutionary graph, retrieve all reachable paths that start from the causative node corresponding to the causative type of the causative type of the causative type in the causative distribution map and do not pass through the avoidance node. Extract the paths in the reachable paths where the transition probability of the subsequent node is greater than the transition probability of the preceding node as the bypass path.

[0036] For the extracted detour path, the shortest path length between each node in the detour path and the avoidance node is calculated. Based on the shortest path length, an isolation metric is generated to characterize the degree of isolation between the node and the avoidance node. The transition probability of each edge in the detour path is weighted and fused with the isolation metric of the corresponding node to generate a path stability index.

[0037] Select the detour path with the highest path stability index, extract the content type corresponding to each node in the detour path, retrieve the audio and video content that matches the content representation vector in the content type, and generate an extended candidate set according to the order of the nodes to which the audio and video content belongs in the detour path.

[0038] Based on evolutionary maps, rejection distribution maps, and behavioral pattern vectors, a comprehensive matching degree is calculated for the audio and video content in the expanded candidate set. Audio and video content is then pushed to users based on the ranking of the comprehensive matching degree, including:

[0039] Identify the content type of each audio and video content in the extended candidate set, calculate the shortest path between the node corresponding to the content type and the node corresponding to the current interactive content type in the evolution graph, extract the transition probability of each edge in the shortest path and calculate the product of the transition probabilities as a path reachability metric.

[0040] Query the rejection strength corresponding to the content type in the rejection distribution map;

[0041] Query the frequency of users' historical interactions with the content type in the behavior pattern vector, and generate preference weights based on the historical interaction frequency;

[0042] The path reachability metric and preference weight are weighted and fused to generate an initial matching score. The initial matching score is adjusted based on the repulsion strength to generate a comprehensive matching score for each audio and video content in the extended candidate set. The audio and video content in the extended candidate set is sorted from largest to smallest according to the comprehensive matching score, and audio and video content is pushed to the user according to the sorting results.

[0043] A second aspect of this invention provides an intelligent algorithm matching system for high-precision audio and video push, comprising:

[0044] The data acquisition unit is used to acquire user interaction behavior data streams and audio / video content feature datasets.

[0045] The feature extraction unit is used to perform time-series analysis on user interaction behavior data streams, extract behavior pattern vectors, and perform multimodal fusion processing on audio and video content feature datasets to generate content representation vectors.

[0046] The graph construction unit is used to construct an evolutionary graph reflecting the content type transfer relationship based on the user's consumption switching sequence between different content types in the behavior pattern vector;

[0047] The exclusion analysis unit is used to identify the user's interrupted interaction features in the behavior pattern vector, mark the content type corresponding to the interrupted interaction feature as the exclusion type, count the cumulative frequency of each exclusion type, and generate an exclusion distribution map.

[0048] The candidate generation unit is used to take the rejection type with a cumulative frequency exceeding a preset suppression threshold in the rejection distribution map as the avoidance node, perform a graph traversal search starting from the non-avoidance node in the evolution graph, extract the connected path in the traversal path that does not pass through the avoidance node as the bypass path, and extract the audio and video content that matches the content representation vector from the corresponding content type along the bypass path to generate an extended candidate set.

[0049] The recommendation push unit is used to calculate the comprehensive matching degree of audio and video content in the extended candidate set based on the evolutionary map, rejection distribution map and behavior pattern vector, and push audio and video content to users according to the comprehensive matching degree ranking results.

[0050] A third aspect of the present invention provides an electronic device, comprising:

[0051] processor;

[0052] Memory used to store processor-executable instructions;

[0053] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0054] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0055] This invention acquires user interaction behavior data streams and audio / video content feature datasets. Through temporal analysis, it extracts behavior pattern vectors and performs multimodal fusion to generate content representation vectors. This comprehensively captures the deep correlation between user behavior dynamics and content semantics, significantly improving the accuracy of push matching. An evolutionary graph is constructed based on the consumption switching sequences in the behavior pattern vectors, intuitively reflecting the user's transition patterns between different content types and providing structured support for subsequent path recommendations. Interruption interaction features are identified in the behavior pattern vectors, and an exclusion distribution map is generated. Types with accumulated frequencies exceeding the suppression threshold are marked as avoidance nodes, thereby performing detour path search in the evolutionary graph. This mechanism effectively avoids pushing explicitly disliked types to users, reducing false positives and user churn risk. Simultaneously, by extracting connected paths that do not pass through avoidance nodes, it achieves content recommendation shielding and security. Content representation vectors are matched from corresponding content types along the detour paths to generate an expanded candidate set. Finally, the evolutionary graph, exclusion distribution map, and behavior pattern vectors are combined to calculate the comprehensive matching degree, ensuring that the recommendation results both match the user's potential interests and avoid negative experiences, significantly improving the adaptability of push notifications and user satisfaction. Attached Figure Description

[0056] Figure 1 A flowchart illustrating the intelligent algorithm matching method for high-precision audio and video push;

[0057] Figure 2 The flowchart shows the process of extracting detour paths and generating extended audio and video candidate sets based on evolutionary graphs. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0059] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0060] Figure 1 This is a flowchart illustrating the intelligent algorithm matching method for high-precision audio and video push according to an embodiment of the present invention.

[0061] The intelligent algorithm matching methods for high-precision audio and video push include:

[0062] Acquire user interaction behavior data streams and audio / video content feature datasets;

[0063] The user interaction behavior data stream is analyzed in time series to extract behavior pattern vectors. At the same time, the audio and video content feature dataset is fused in multimodal mode to generate content representation vectors.

[0064] Based on the consumption switching sequence of users between different content types in the behavioral pattern vector, an evolutionary map reflecting the content type transfer relationship is constructed.

[0065] Identify user interruption interaction features in behavior pattern vectors, mark the content types corresponding to interruption interaction features as exclusion types, count the cumulative frequency of each exclusion type, and generate an exclusion distribution map.

[0066] The exclusion type with a cumulative frequency exceeding the preset suppression threshold in the exclusion distribution map is taken as the avoidance node. A graph traversal search starting from the non-avoidance node is performed in the evolution graph. Connected paths that do not pass through the avoidance node in the traversal path are extracted as bypass paths. Audio and video content matching the content representation vector is extracted from the corresponding content type along the bypass path to generate an extended candidate set.

[0067] The comprehensive matching degree of audio and video content in the extended candidate set is calculated based on the evolutionary map, rejection distribution map and behavior pattern vector, and audio and video content is pushed to users according to the ranking result of the comprehensive matching degree.

[0068] The user interaction behavior data stream is subjected to time-series analysis to extract behavior pattern vectors. Simultaneously, multimodal fusion processing is performed on the audio and video content feature dataset to generate content representation vectors, including:

[0069] The interaction events in the user interaction behavior data stream are sorted by timestamp to construct a time-series interaction sequence, and a fixed step size is set to segment the time-series interaction sequence by a sliding window.

[0070] The interaction type, interaction duration, and interaction interval within the window are extracted as behavioral feature dimensions. After normalizing the behavioral feature dimensions, the statistical distribution characteristics within the window and the rate of change of behavioral intensity between adjacent windows are calculated. The statistical distribution characteristics and the rate of change of behavioral intensity are concatenated and encoded to generate a behavioral pattern vector.

[0071] Visual representations are generated by extracting body contours and texture information from visual frame sequences in audio and video content feature datasets, and audio representations are generated by extracting frequency components and energy distribution from audio waveform sequences through short-time Fourier transform.

[0072] The system detects whether there is an alignment deviation between visual and audio representations in the time dimension. When an alignment deviation exists, it calculates the sliding cross-correlation function of visual and audio representations in the time dimension, extracts the time offset corresponding to the global maximum value of the sliding cross-correlation function, and then translates the visual representation according to the time offset and sums it with the audio representation element by element to generate a content representation vector.

[0073] The user interaction data stream contains various interaction events generated by users on the platform, such as clicks, play, pause, fast forward, favorites, and comments. Each event record is accompanied by precise timestamp information. Arranging these interaction events in ascending order of timestamps yields a complete time-series interaction sequence, which accurately reconstructs the user's behavioral trajectory along the timeline. To capture local behavioral patterns from continuous time-series interaction sequences, a fixed step size is set. The time-series interaction sequence is segmented using a sliding window, with each window covering a length of [length missing]. The time interval, with adjacent windows separated by a step size By moving backward, the entire sequence is divided into several independently analyzable behavioral segments while maintaining temporal continuity.

[0074] Within each sliding window, three behavioral feature dimensions are extracted: interaction type, interaction duration, and interaction interval. Interaction type reflects the types of operations performed by the user within that window, which can be quantified using one-hot encoding or category frequency statistics. Interaction duration represents the time a user invests in a single interaction, directly reflecting the user's level of attention to the content. Interaction interval describes the time difference between two adjacent interaction events, indirectly reflecting the user's browsing rhythm and engagement continuity. Because the dimensions and numerical ranges of the three feature dimensions differ significantly, min-max normalization is performed on each dimension separately to map the original values ​​to... The interval is used to eliminate the interference of inconsistent dimensions on subsequent calculations.

[0075] After normalization, the statistical distribution characteristics of each behavioral feature dimension within the window are calculated, including descriptive statistics such as mean, variance, skewness, and kurtosis, to characterize the overall distribution of behavior within the window. Simultaneously, the rate of change in behavioral intensity between adjacent windows is calculated. Behavioral intensity can be defined as the weighted cumulative value of all interaction durations within the window. Let the i-th... The behavior intensity of each window is The rate of change of behavioral intensity between adjacent windows for This rate of change can capture the fluctuations in user interest over time, such as the turning point where a user suddenly accelerates consumption or noticeably stagnates on a certain piece of content. The statistical distribution characteristic vector and the behavior intensity change rate vector are concatenated in dimensional order and mapped to a fixed-dimensional latent space through a linear encoding layer, ultimately generating a behavior pattern vector. This vector simultaneously carries static distribution information within the window and dynamic evolution information across windows.

[0076] For the audio and video content feature dataset, content features are extracted from both visual and audio modalities. For the visual frame sequence, keyframes are extracted at a fixed frame rate. Edge detection operators are performed on each frame to extract geometric shape descriptors of object contours. Simultaneously, local binary mode or Gabor filters are used to extract texture information. The contour descriptors and texture features are concatenated and compressed to a fixed dimension using a convolutional coding network to generate frame-by-frame visual feature vectors, which are then aggregated along the time axis to form a visual representation. For the audio waveform sequence, a short-time Fourier transform is performed on the original waveform signal to convert the time-domain signal into a time-frequency domain representation. The frequency component distribution and energy distribution corresponding to each time frame are extracted to form a spectrogram matrix. The spectrogram matrix is ​​then weighted and integrated along the frequency axis to extract frequency band features with concentrated energy, which are then encoded using a fully connected layer to generate an audio representation. Both visual and audio representations use time frames as the basic unit and have the same time axis resolution, laying the foundation for subsequent alignment processing.

[0077] Before fusing visual and audio representations, it is necessary to detect any alignment discrepancies between them in the temporal dimension. Alignment discrepancy refers to the temporal misalignment between the visual frame sequence and the audio waveform sequence during acquisition, encoding, or transmission. Without correction, direct fusion will introduce noise and reduce the semantic accuracy of the content representation. The detection method involves calculating the frame-by-frame similarity between the visual and audio representation sequences in the temporal dimension. If the peak of the similarity curve deviates from the zero-latency position, an alignment discrepancy is determined to exist.

[0078] When alignment deviation is detected, the visual representation sequence is calculated. With audio representation sequence Sliding cross-correlation function between ,in This is the time offset. Indicates at offset The degree of correlation between the two sequences at time points. The calculation method involves sliding frame by frame along the time axis with a step size, and calculating the mean of the inner product of the feature vectors of corresponding frames in the two sequences at each offset position. Extraction The time offset corresponding to obtaining the global maximum value among all offset positions. This offset is the estimated actual time delay of the visual representation relative to the audio representation. Subsequently, the visual representation sequence is shifted entirely along the time axis. Each time unit allows the visual and audio frames to be realigned in the time dimension.

[0079] After alignment correction, the visual representation after translation is... With audio representation Perform an element-wise weighted summation operation, and let the visual modality weight be... Audio modal weights are And satisfy The fused content representation vector The Frame representation Weight and It can adaptively adjust based on content type; for example, it can assign higher priority to content that is primarily visual (such as landscape videos). Values ​​are assigned higher to content that is primarily audio information (such as music videos). Values. For each frame along the timeline. Average pooling is performed to ultimately generate a single content representation vector. This vector integrates semantic information of audio and video content across multiple dimensions, including visual texture, contour structure, frequency components, and energy distribution, providing a high-quality feature foundation for accurate matching of subsequent content with user preferences.

[0080] Behavioral pattern vectors and content representation vectors together constitute the two-sided inputs of the recommendation algorithm. The former encodes the user's dynamic behavioral preferences, while the latter encodes the multimodal semantic features of audio and video content. In the subsequent matching stage, the two achieve accurate matching between users and content through similarity calculation, thereby supporting the goal of high accuracy in the overall push process.

[0081] Based on the user consumption switching sequences between different content types in the behavioral pattern vector, an evolutionary map reflecting the content type transfer relationship is constructed, including:

[0082] Extract user consumption switching sequences between different content types from the behavior pattern vector, identify content type labels of adjacent consumption behaviors in the consumption switching sequence, and combine adjacent content type labels into ordered pairs as content type switching pairs according to the chronological order.

[0083] The frequency of occurrence of content type switching pairs in the consumption switching sequence is statistically analyzed, and a content type transition matrix is ​​constructed with content type as the row and column index and occurrence frequency as the matrix element.

[0084] The transition probability of each content type to other content types is calculated by row normalization of the content type transition matrix. The eigenvector of the content type transition matrix is ​​calculated, and the eigenvector corresponding to the largest eigenvalue is extracted as the content type importance vector.

[0085] Calculate the Shannon entropy for the transition probability distribution of each row in the content type transition matrix, multiply the Shannon entropy by the corresponding element value in the content type importance vector to generate a comprehensive stability index, and sort and label the node level of the content types according to the comprehensive stability index.

[0086] An evolutionary graph reflecting the content type transfer relationship is constructed using content type as nodes, transfer probability as directed edge weights, and node level as node attributes.

[0087] When extracting user consumption switching sequences between different content types from the behavior pattern vector, each interaction record in the behavior pattern vector needs to be tagged with a type. Each time a user completes the consumption of one piece of content and switches to another, the content type tag of the current content and the content type tag of the next piece of content are recorded. These two adjacent content type tags are then combined into ordered pairs according to their chronological order, forming a content type switching pair. For example, if a user first consumes a video of type "Technology News" and then switches to content of type "Lifestyle & Entertainment," the corresponding content type switching pair is recorded as (Technology News → Lifestyle & Entertainment). By traversing the entire consumption switching sequence, a set of ordered pairs covering all of the user's historical switching behaviors can be obtained. This set completely depicts the user's migration trajectory along the content type dimension.

[0088] After obtaining all content type switching pairs, frequency statistics are performed on these ordered pairs to count the number of times each switching pair appears in the consumption switching sequence. A content type transition matrix is ​​constructed using all appearing content types as row and column indices, and the frequency of the corresponding switching pair as matrix elements. Let the total number of content types be... Then the transition matrix is A square matrix, the first element in the matrix is... Line 1 Column elements Indicates that the user starts from the first Switching to the first content type The historical frequency of each content type. For switching directions that have never occurred, the corresponding matrix element has a value of 0. The sparsity of this matrix directly reflects the diversity of user interest migration: if the non-zero elements in the matrix are relatively concentrated, it indicates that the user's switching behavior has obvious regularity; if the non-zero elements are relatively evenly distributed, it indicates that the user's content consumption preferences are relatively broad.

[0089] The content type transition matrix is ​​normalized row by row by dividing each element of the row by the sum of the elements in that row, thus converting the frequency values ​​of each row into probability values. After normalization, the matrix... Line 1 Column elements This indicates that the user has completed the first consumption. After this type of content, the next consumption... The transition probabilities of different content types are given, and the sum of all elements in each row is 1. Based on this, the eigenvalues ​​and eigenvectors of the normalized content type transition matrix are calculated. Let the largest eigenvalue in the set of eigenvalues ​​of the matrix be denoted as . Its corresponding feature vector is ,Will This serves as a content type importance vector. Each component of this feature vector reflects the core position of each content type in the overall transfer structure: the larger the component value, the more pivotal the content type is in the user's consumption switching behavior, and the key node in the user's interest migration.

[0090] Calculate the Shannon entropy for the transition probability distribution of each row in the normalized content type transition matrix. For the ... The content type and its corresponding Shannon entropy Calculated by iterating through all non-zero probability elements in the row. ,calculate Among them, The agreed contribution value for this item is 0. The size reflects the degree of dispersion of user switching behavior when starting from this content type: The larger the value, the more evenly distributed the switching direction from this type is, and the higher the uncertainty of users' subsequent consumption behavior in this type; The smaller the value, the more concentrated the user's switching behavior is on a few target types after starting from this type, indicating a strong regularity.

[0091] Shannon entropy for each content type Content type importance vector Corresponding components Multiply by each other to obtain the overall stability index for this content type. . The study comprehensively considered the importance of content types within the overall transfer structure and the dispersion of their switching behavior: if a content type has high importance but its switching direction is dispersed (i.e., high entropy), its overall stability index is high, meaning that this type needs to be assigned a higher level in the graph to reflect its complex pivotal characteristics; conversely, if a content type has low importance and its switching direction is concentrated, its overall stability index is low, and its corresponding node level is relatively low. Based on the overall stability index of all content types... Content types are sorted in descending order, and then divided into several levels based on the sorting results. Content types with higher overall stability indices are marked as higher-level nodes, while those with lower indices are marked as lower-level nodes. The node levels can be divided using methods such as equally spaced segments or natural breakpoint classification to ensure that the level division effectively distinguishes the structural differences in user consumption behavior among different content types.

[0092] Using all occurrences of content type as nodes, and normalized transition probabilities... Using the weights of directed edges and the hierarchy determined by the comprehensive stability index ranking results corresponding to each content type as node attributes, an evolutionary graph reflecting the content type transfer relationship is constructed. Each directed edge in the graph originates from the source node (content type). ) points to the target node (content type) ), edge weight The transition strength between the two types of content was quantified. Node hierarchy attributes provide structured priority information for subsequent graph traversal searches, enabling priority expansion along higher-level nodes during detour path extraction, thus more accurately capturing potential user interest migration directions. When the transition probability... When the edge weight falls below a certain set pruning threshold, the corresponding directed edges can be pruned to reduce the complexity of the graph and filter out the interference of occasional switching behaviors on the graph structure. The final generated evolutionary graph completely preserves the user's migration patterns at the content type level in a structured form, providing a reliable graph structure foundation for subsequent avoidance node identification, detour path search, and expansion candidate set generation.

[0093] Identify user interruption interaction features in the behavior pattern vector, label the content types corresponding to the interruption interaction features as exclusion types, count the cumulative frequency of each exclusion type, and generate an exclusion distribution map, including:

[0094] Extract the interaction duration and interaction completion fields from the behavior pattern vector. Group and standardize the interaction duration field according to the content type label to obtain the standardized duration. Construct a feature space with the standardized duration and interaction completion field and perform isolation tree anomaly detection to obtain anomaly score.

[0095] Interaction events with abnormal scores exceeding a preset abnormal threshold are marked as interrupted interaction features, and the corresponding content type tags are extracted and marked as exclusion types;

[0096] The cumulative frequency is constructed by counting the number of occurrences of the exclusion type in the behavior pattern vector. The cumulative frequency is normalized to generate the exclusion strength. The distribution uniformity of the cumulative frequency is evaluated to obtain the concentration index.

[0097] In the consumption switching sequence, the location of the rejection type is retrieved and a location index set is constructed. For each location in the location index set, the content type label within a fixed step size is extracted to construct a pre-label matrix. Principal component analysis is performed on the pre-label matrix to extract the dominant feature vector. The loading coefficient of each content type label in the dominant feature vector is calculated as the inducing factor intensity. The content type label with the largest inducing factor intensity is selected as the rejection inducing factor type.

[0098] A rejection distribution map is generated using rejection type as the horizontal axis, rejection intensity as the vertical axis, rejection cause type as the node label, and concentration index as the transparency parameter.

[0099] Extracting the interaction duration and interaction completion fields from the behavior pattern vector is a fundamental data preparation step for identifying interrupted interaction features. The interaction duration field records the actual time a user spends on a specific piece of content, while the interaction completion field records the percentage of that content consumed, such as the percentage of video playback progress relative to the total duration, or the percentage of audio listened to. Because the duration distribution varies significantly across different content types—short videos and long documentaries differ drastically in length—directly comparing raw interaction durations would introduce systematic bias. Therefore, the interaction duration field is grouped by content type label, and standardization is performed within each content type group to convert the raw duration into a relative duration level within that type, resulting in a standardized duration. Standardized durations eliminate the impact of inconsistent duration benchmarks between different content types, enabling subsequent anomaly detection to perform cross-type comparisons on a uniform scale.

[0100] A two-dimensional feature space is constructed using standardized duration and interaction completion fields, within which isolation tree anomaly detection is performed. The core idea of ​​the isolation tree is to randomly partition the feature space, making sample points with shorter path lengths easier to isolate and corresponding to higher anomaly scores. Normal interaction behavior typically shows a positive correlation between standardized duration and interaction completion, meaning that content with longer consumption time often has higher completion. Interrupted interaction behavior, on the other hand, is characterized by the coexistence of low completion and relatively normal duration, or extremely short duration and extremely low completion. These sample points are located in sparse regions in the feature space, and the isolation tree can separate them with a shorter average path length, assigning them higher anomaly scores. Interaction events with anomaly scores exceeding a preset anomaly threshold are extracted and marked as interrupted interaction features. Simultaneously, the content type labels corresponding to these interaction events are extracted and marked as exclusion types. The preset anomaly threshold can be dynamically determined based on the quantile distribution of historical data; for example, the 90th percentile of the anomaly score distribution can be used as the threshold to ensure that only events that statistically significantly deviate from the normal interaction pattern are included in the interrupted interaction feature set.

[0101] The cumulative frequency is constructed by counting the number of times each rejection type appears in the behavior pattern vector. The cumulative frequency directly reflects the historical accumulation of users' rejection behavior towards that type of content; a higher frequency indicates that the user's rejection behavior towards that type is more persistent and stable. The cumulative frequency of each rejection type is normalized and mapped to the interval between 0 and 1 to obtain the rejection strength. Rejection strength is a relative expression of the cumulative frequency, making the rejection levels comparable across different users or time periods. Simultaneously, the distribution uniformity of the cumulative frequency is assessed, and a concentration index is calculated. The concentration index measures whether users' interrupted interaction behavior is concentrated on a few rejection types or evenly distributed across multiple types. A high concentration index indicates that users' rejection of specific content types is highly targeted; a low concentration index indicates that the rejection behavior is more dispersed, possibly reflecting a decline in overall user activity rather than active avoidance of specific content types. Concentration indicators can be calculated using the Gini coefficient or normalized entropy value based on frequency distribution. The closer the Gini coefficient is to 1, the higher the concentration. Similarly, the closer the normalized entropy value is to 0, the higher the concentration. Both can be used as quantitative implementation schemes for concentration indicators.

[0102] The system retrieves the locations where rejection types occur within the consumption switching sequence, constructing a location index set. The consumption switching sequence records the user's switching trajectory between different content types in chronological order. The location of a rejection type within the switching sequence indicates that the user consumed content of that type at that moment, triggering an interruption of interaction. For each location in the location index set, content type tags within a fixed step size are extracted. The fixed step size can be set to 3 to 5 switching steps, and the specific value can be adjusted according to the average length and sparsity of the user behavior sequence. These preceding content type tags are summarized to construct a preceding tag matrix. Each row of the matrix corresponds to a rejection event, and each column corresponds to a content type tag (represented by one-hot encoding or frequency encoding) appearing at a certain location within a fixed step size before the event. The preceding tag matrix captures the user's content consumption path pattern before the rejection behavior occurs, which may implicitly suggest that certain content types have a continuous inducing effect on rejection behavior.

[0103] Principal component analysis (PCA) is performed on the pre-label matrix to extract the dominant eigenvector (VIV) that explains the largest proportion of variance. The VVIV reflects the most significant direction of change in the pre-label matrix, and the loading coefficients of each dimension represent the contribution of the corresponding content type label to that principal component. Content type labels with larger absolute loading coefficients indicate a highly consistent pattern of occurrence in the consumption path before the rejection behavior, meaning that the consumption of this type of content often precedes the rejection behavior, exhibiting strong causal characteristics. The loading coefficients of each content type label in the dominant eigenvector are used as causal strength, and the content type label with the highest causal strength is marked as the rejection causal type. Identifying rejection causal types has significant practical value: it reveals the pre-triggered patterns of user rejection behavior, providing additional avoidance references when planning detour paths in the evolutionary graph, and preventing recommended paths from passing through content type sequences that could induce rejection.

[0104] A rejection distribution map is generated using rejection type as the horizontal axis, rejection strength as the vertical axis, rejection trigger type as node labels, and concentration index as the transparency parameter. In the rejection distribution map, each rejection type corresponds to a discrete coordinate position on the horizontal axis, and its height on the vertical axis directly reflects the rejection strength of that type. The rejection trigger type labeled on the node indicates which preceding content type is most likely to induce interruption of interaction for that rejection type. The concentration index is visually encoded through the transparency parameter—rejection type nodes with higher concentration have lower transparency (darker color), and nodes with lower concentration have higher transparency (lighter color), allowing the graph to visually distinguish between targeted rejection behavior and scattered rejection behavior. The rejection distribution map serves as input for subsequent avoidance node determination and detour path search. Its compact encoding of multi-dimensional information allows the recommendation algorithm to simultaneously refer to information from three dimensions—rejection strength, rejection concentration, and trigger association—when processing graph traversal, thereby implementing a more refined content avoidance strategy during the candidate set generation stage.

[0105] Principal component analysis is performed on the pre-label matrix to extract the dominant feature vector. The loading coefficients of each content type label in the dominant feature vector are calculated as the inducing factor strength, including:

[0106] The pre-label matrix is ​​converted into a numerical matrix by one-hot encoding. Singular value decomposition is performed on the numerical matrix to obtain the left singular vector matrix and the singular value sequence. The maximum singular value is selected from the singular value sequence, and the column vector corresponding to the maximum singular value in the left singular vector matrix is ​​extracted as the dominant feature vector.

[0107] Extract the component values ​​corresponding to each content type label from the dominant feature vector as the pattern contribution, and count the number of times each content type label appears in the position before the exclusion type in the preceding label matrix as the neighbor frequency.

[0108] Normalization is performed on both the pattern contribution and the neighbor frequency. The normalized pattern contribution and the normalized neighbor frequency are then weighted and fused to obtain the loading coefficient of each content type tag in the dominant feature vector as the inducing force intensity.

[0109] The pre-interaction tag matrix stores the sequence of content type tags consumed by the user in the steps prior to each interruption event. These tags exist in the form of strings or category codes and cannot be directly used in numerical calculations. Therefore, a one-hot encoding transformation is required to convert the pre-interaction tag matrix, mapping each content type tag to a binary vector that takes a value of 1 only in the corresponding dimension and 0 in the other dimensions. After the transformation, each row in the pre-interaction tag matrix corresponds to the expanded tag sequence before an interruption event, and each column corresponds to whether a specific content type tag appears at that position, thus completely converting the original category information into a numerical matrix form. This process ensures that there are no implicit size relationships between different content type tags, avoiding the misinterpretation of category codes as ordered numerical values.

[0110] After completing the one-hot encoding transformation, singular value decomposition is performed on the resulting numerical matrix. Let the numerical matrix be... Singular value decomposition decomposes it into a left singular vector matrix. Singular value diagonal matrix and right singular vector matrix The product of, i.e. .in, The column vectors reflect the distribution structure of the samples (i.e., each interruption event) in the principal component space. The diagonal elements are the singular values, arranged in descending order. The magnitude of each singular value represents the proportion of data variance explained by the corresponding principal component direction. The largest singular value is selected from the singular value sequence. (Right now (first element on the diagonal), extract Zhongyu The corresponding first column vector is the dominant feature vector. . It captures the most explanatory linear structure direction in the pre-label matrix, representing the dominant pattern that recurs in the pre-sequence of all interrupted interaction events.

[0111] Dominant feature vector The dimension is consistent with the number of rows in the preceding tag matrix (i.e., the number of interruption interaction events), not with the number of content type tags. Therefore, it cannot be directly derived from... The contribution of each content type tag is read by column index. To obtain the projection intensity of each content type tag in the dominant direction, the right singular vector matrix is ​​needed. Zhongyu The corresponding first column vector . The dimension of is consistent with the feature dimension after one-hot encoding (i.e., the total number of content type tags), where the ... Each component Indicates the first The loading of a content type tag in the direction of the dominant feature reflects the degree of contribution of that tag to the dominant pattern, and is called the pattern contribution. The larger the absolute value of the pattern contribution, the stronger the regularity and the more significant the structure of the content type tag in the preceding sequence that triggers the interruption behavior.

[0112] While extracting the contribution of patterns, it is also necessary to count the frequency of each content type tag appearing in the position preceding the exclusion type from the original records of the preceding tag matrix, i.e., the proximity frequency. Specifically, for each preceding tag sequence, check the content type tag corresponding to the last position of the sequence (i.e., the most recent content consumption record immediately adjacent to the time of the interrupted interaction event). If the tag belongs to a certain content type... If the neighbor frequency count is incremented by 1, then the neighbor frequency count for that type is incremented by 1. The neighbor frequency directly measures the statistical frequency of a content type being closely adjacent to the exclusionary behavior in time. It is a direct counting indicator based on local temporal location, which complements the information captured from the global structural perspective of pattern contribution.

[0113] After obtaining the pattern contribution and nearest neighbor frequency, normalization is performed on both sets of values ​​to map their respective value ranges to a unified numerical scale, thereby eliminating the impact of dimensional differences on subsequent weighted fusion. The normalization method uses min-max normalization, assuming a certain content type... The original value of the pattern contribution is The normalized model contribution is The calculation method is as follows Similarly, let content type The original value of the nearest neighbor frequency is The normalized neighbor frequency is The calculation method is as follows After normalization, both sets of values ​​are at... Within the interval, the conditions for direct integration are met.

[0114] The normalized pattern contribution and the normalized neighbor frequency are weighted and fused to calculate the loading coefficient of each content type tag, i.e., the incentive strength. The fusion calculation formula is as follows: ,in The fusion weight for the contribution of the pattern, with a value range of [value range missing]. , The fusion weights are for neighboring frequencies. The value of can be adjusted according to the data size and the sparsity of interrupt events: when the number of interrupt events is large and the preceding sequence structure is highly regular, it should be appropriately increased. To enhance the weighting of principal component analysis results; when the number of interruption events is small and the statistical sample is limited, the weighting should be appropriately reduced. This allows the direct counting indicator of proximity frequency to play a greater role, ensuring the robustness of the causation intensity estimation.

[0115] Intensity of triggers The physical meaning is: content type In the consumption sequence before users exhibit rejection behavior, there are significant repetitive patterns at the global structural level and a high correlation with rejection events at the temporal proximity level. Therefore, these patterns are identified as potential factors that induce rejection behavior. The higher the value, the stronger the contribution of the content type to triggering rejection behavior. In subsequent push filtering and path planning, content types with high incentive intensity will not only be marked as precursor nodes of the rejection type itself, but will also be subject to additional weight penalties in the detour path search of the evolution graph, thereby reducing the probability of the push path passing through this type of content type and further reducing the possibility of users encountering rejection experiences.

[0116] Through the complete process described above, from one-hot encoding, singular value decomposition, pattern contribution extraction, neighbor frequency statistics to weighted fusion, a quantified incentive strength value is generated for each content type tag, forming an incentive strength distribution vector. This vector, together with the repulsion distribution map, constitutes a two-layer characterization of user repulsion behavior: the repulsion distribution map describes which content types users show a clear avoidance tendency, while the incentive strength distribution vector further reveals which content types act as pre-emptive triggers for repulsion behavior in the consumption sequence, providing a more refined basis for path avoidance decisions in the evolutionary graph.

[0117] like Figure 2 As shown, Figure 2 This is a flowchart illustrating the process of detour path extraction and audio / video extended candidate set generation based on evolutionary graphs in an embodiment of the present invention.

[0118] In the evolutionary graph, a graph traversal search starting from non-avoided nodes is performed. Connected paths that do not pass through avoided nodes are extracted as bypass paths. Audio and video content matching the content representation vector is extracted from the corresponding content type along the bypass paths to generate an extended candidate set, including:

[0119] In the evolutionary graph, retrieve all reachable paths that start from the causative node corresponding to the causative type of the causative type of the causative type in the causative distribution map and do not pass through the avoidance node. Extract the paths in the reachable paths where the transition probability of the subsequent node is greater than the transition probability of the preceding node as the bypass path.

[0120] For the extracted detour path, the shortest path length between each node in the detour path and the avoidance node is calculated. Based on the shortest path length, an isolation metric is generated to characterize the degree of isolation between the node and the avoidance node. The transition probability of each edge in the detour path is weighted and fused with the isolation metric of the corresponding node to generate a path stability index.

[0121] Select the detour path with the highest path stability index, extract the content type corresponding to each node in the detour path, retrieve the audio and video content that matches the content representation vector in the content type, and generate an extended candidate set according to the order of the nodes to which the audio and video content belongs in the detour path.

[0122] Before generating the expanded candidate set, it's necessary to define the origin of the detour path. The rejection distribution graph records the cumulative frequency of each rejection type and its corresponding rejection trigger type. The rejection trigger type reflects the preceding content type that triggers the user to interrupt interaction; that is, the content category in which the user immediately interrupts their behavior after consuming a certain type of content. In the evolutionary graph, each content type corresponds to a node, and the node corresponding to the rejection trigger type is called the trigger node. Graph traversal search starts from the trigger node, rather than from the avoidance node, because the trigger node itself does not trigger user rejection but is the preceding context of rejection behavior, possessing a certain degree of rationality in content consumption. Starting from the trigger node allows us to explore content migration paths that are still acceptable to the user when approaching the rejection boundary.

[0123] Starting from the trigger node, a depth-first or breadth-first graph traversal search is performed in the evolutionary graph. During the traversal, all path branches that pass through avoidance nodes are strictly filtered out, retaining only reachable paths that do not touch the avoidance nodes at all. Among all reachable paths, paths that satisfy the condition of monotonically increasing transition probabilities are further selected, that is, the transition probability of subsequent nodes in the path is strictly greater than the transition probability of preceding nodes. Let the i-th node in the path be... The node to the first The transition probability of each node is Then the path that satisfies the condition must satisfy... ,in The total number of path nodes. This selection criterion ensures that the detour path shows a trend of gradually increasing user interest along the evolution direction, avoiding the introduction of content type nodes with low switching probability and weak user willingness to switch, thereby ensuring that the content in the expanded candidate set has a high probability of user acceptance.

[0124] The set of paths that meet the above screening criteria constitutes the candidate detour path set. For each path in the candidate detour path set, it is necessary to further evaluate the overall stability of the path to prevent potential user rejection risks caused by some nodes in the path that, although not marked as avoidance nodes, are too close to avoidance nodes in the content type space. To this end, the shortest path length between each node in the path and all avoidance nodes is calculated. Let the path be the path with the shortest path length between each node and all avoidance nodes. Each node is The set of nodes to avoid is Then the node shortest isolation distance Defined as arrive The minimum shortest path length of all avoided nodes, i.e. ,in This represents the shortest path length between two nodes in the evolutionary graph.

[0125] Based on the shortest isolation distance, generate the node's isolation metric. The isolation metric uses a monotonically increasing mapping function to convert the shortest isolation distance into a normalized representation of the isolation level. A larger value indicates that the node is farther away from all avoided nodes in the graph structure. A higher isolation metric value means a lower correlation between the content type corresponding to the node and the content that the user avoids, and stronger security. Specifically, ,in To control the positive real hyperparameter of the decay rate, adaptive calibration can be performed based on the average node spacing of the evolution map. When this occurs, it indicates that the node itself is an avoidance node. ;along with Increase A value close to 1 indicates that the node is completely isolated from the avoidance area.

[0126] After obtaining the isolation metric values ​​of each node, the transition probabilities of each edge in the path are weighted and fused with the corresponding isolation metric values ​​of the nodes to calculate the path stability index. Let the path contain the first... The transition probability of the edge is The corresponding starting node isolation metric value is The fusion weight is (The range of values ​​is) The path stability index is then defined as the mean of the weighted fusion values ​​of all edges in the path. .in Used to adjust the relative weights of transition probability and isolation metric in path stability assessment, when When the value is large, path stability depends more on the strength of the transition probability; when When the path size is small, path stability focuses more on the degree of isolation between the path and avoidance nodes. In practical applications, The system can be dynamically adjusted based on the user's historical rejection intensity, and the intensity can be appropriately reduced for users with frequent rejection behaviors. This is to more conservatively avoid the risk of potential rejection.

[0127] Calculate the path stability index for each path in the candidate detour path set. Select The longest path is taken as the final detour path. This path has both a high probability of content type transfer and a strong degree of isolation of avoidance nodes in the evolutionary graph, which can meet the user's content consumption migration pattern while avoiding user-rejected areas to the greatest extent.

[0128] After determining the final detour path, the content type corresponding to each node is extracted sequentially according to the order of nodes in the path. For each content type, the content representation vector is retrieved from the audio and video content feature dataset. The audio and video content with the highest matching degree. The matching process uses cosine similarity to calculate the feature vector of the candidate content. The similarity between nodes is used to filter content items whose similarity exceeds a preset matching threshold and include them in the candidate pool. The content type corresponding to each node is retrieved independently, and the retrieval results are arranged according to the order of the nodes in the detour path. The content corresponding to the node near the starting point of the path (the side of the inducing node) is arranged first, and the content corresponding to the node near the ending point of the path is arranged last, thus generating an expanded candidate set with a path order structure.

[0129] This path-ordered extended candidate set is crucial in recommendation sequences. Content types near the path's starting point are more closely related to the user's current consumption context, reducing the user's unfamiliarity with the recommended content; content types near the path's ending point represent content directions with higher transition probabilities and stronger user acceptance in the evolutionary graph. The path-ordered candidate set can be weighted in the subsequent comprehensive matching degree calculation stage, incorporating path location information. This ensures that the recommendation results reflect both the gradual migration logic of content and effectively avoid content types that users are already aware of and dislike, achieving highly accurate personalized audio and video content delivery.

[0130] Based on evolutionary maps, rejection distribution maps, and behavioral pattern vectors, a comprehensive matching degree is calculated for the audio and video content in the expanded candidate set. Audio and video content is then pushed to users based on the ranking of the comprehensive matching degree, including:

[0131] Identify the content type of each audio and video content in the extended candidate set, calculate the shortest path between the node corresponding to the content type and the node corresponding to the current interactive content type in the evolution graph, extract the transition probability of each edge in the shortest path and calculate the product of the transition probabilities as a path reachability metric.

[0132] Query the rejection strength corresponding to the content type in the rejection distribution map;

[0133] Query the frequency of users' historical interactions with the content type in the behavior pattern vector, and generate preference weights based on the historical interaction frequency;

[0134] The path reachability metric and preference weight are weighted and fused to generate an initial matching score. The initial matching score is adjusted based on the repulsion strength to generate a comprehensive matching score for each audio and video content in the extended candidate set. The audio and video content in the extended candidate set is sorted from largest to smallest according to the comprehensive matching score, and audio and video content is pushed to the user according to the sorting results.

[0135] After obtaining the expanded candidate set, a comprehensive matching degree quantification calculation needs to be performed on each audio and video content to achieve accurate ranking and push. The calculation process involves three main information threads: the content type transfer relationship depicted by the evolutionary graph, the user rejection tendency recorded by the rejection distribution map, and the historical interaction preferences carried by the behavioral pattern vector. Each of the three threads generates an independent metric, which is ultimately merged into a unified comprehensive matching degree score.

[0136] For each audio / video content in the expanded candidate set, its content type is first identified, and the corresponding node for that content type is located in the evolutionary graph. Simultaneously, the node corresponding to the content type the user is currently interacting with is determined as the starting node. In the evolutionary graph, nodes represent content types, and directed edges carry transition probability attributes, reflecting the user's historical tendency to switch between two content types. Based on the edge weight information in the graph, a shortest path algorithm (such as Dijkstra's algorithm, using the negative logarithm of the transition probability as the edge cost) is used to calculate the shortest path from the starting node to the target content type node. The transition probabilities corresponding to each directed edge on this shortest path are extracted, and these transition probabilities are multiplied sequentially; the resulting product is the path reachability metric for that audio / video content. A higher path reachability metric indicates a stronger historical tendency for the user to naturally transition from their current consumption state to that content type, and a more sufficient graph basis for recommending that content. Let the set of directed edges traversed by the shortest path contain a total of... edge, the first The transition probability corresponding to each edge is Then the path reachability metric for: When the shortest path does not exist (i.e., the target node and the starting node are unreachable in the evolutionary graph), then... A value of 0 indicates that the content type has no transfer path support in the current graph, and the subsequent comprehensive matching degree will be suppressed accordingly.

[0137] After obtaining the path reachability metric, the process then moves to the exclusion distribution map to query the exclusion strength. The exclusion distribution map is indexed by content type, storing the cumulative interruption frequency and normalized exclusion strength value for each content type. For each audio / video content in the extended candidate set, its corresponding exclusion strength is retrieved from the exclusion distribution map and denoted as . If a content type does not appear in the exclusion distribution map, its exclusion strength is... A value of 0 indicates that the user has not historically exhibited any observable aversion to this type. Aversion strength. The range of values ​​is constrained to within after normalization. Within the range, the higher the value, the stronger the user's aversion to that type of content.

[0138] The behavior pattern vector records the historical interaction frequency of users with each content type, reflecting the depth and breadth of user consumption across different content types. For each audio / video content in the expanded candidate set, the raw value of the user's historical interaction frequency for that type is extracted from the behavior pattern vector and denoted as... To eliminate the dimensional imbalance caused by the difference in absolute frequency between different content types, the historical interaction frequency of all candidate content types is normalized by min-max normalization. The normalized historical interaction frequency is denoted as [missing information]. The calculation method is as follows: ,in and These represent the maximum and minimum historical interaction frequencies of all candidate content types in the expanded candidate set, respectively. (Normalized) As a preference weight, its value range is constrained by... Range. The higher the preference weight, the deeper the user's historical consumption accumulation of this content type, and the higher the degree of alignment between the recommended content of this type and the user's existing preferences.

[0139] In obtaining path reachability metrics and preference weights Then, the two are weighted and merged to generate a preliminary matching score. Let the fusion weight of the path reachability metric be... The fusion coefficient corresponding to the preference weight is And satisfy The initial matching degree The calculation method is as follows: fusion weight and The specific value can be adjusted according to the actual business scenario. When sufficient user behavior data has been accumulated, it can be appropriately increased. The weighting of historical preferences makes them play a stronger dominant role in recommendation results; when the user is a new user or historical data is sparse, the weighting can be appropriately increased. The proportion of [something] is inferred more from the content type transfer relationships in the evolutionary map.

[0140] Preliminary match This reflects a comprehensive level of content type accessibility and historical preferences, but it does not yet consider users' aversion tendencies. To incorporate the strength of aversion into the final scoring system, a rating based on the strength of aversion will be implemented. The initial matching score is adjusted with a suppressive effect to generate a comprehensive matching score. The adjustment method uses a multiplicative attenuation approach, specifically as follows: The physical meaning of this formula is: when the repulsion strength When it approaches 0, the overall matching degree is... Preliminary matching degree Almost equal, the rejection adjustment has a negligible impact on the score; when the rejection strength... When the overall matching degree approaches 1, the overall matching degree is... A value significantly lowered to near 0 indicates that while the content type may have a certain score in terms of graph accessibility or historical preference, its recommendation priority is significantly reduced due to strong user aversion. This design ensures that the negative user signal in the aversion distribution map effectively suppresses the recommendation results, preventing content types that users clearly dislike from being pushed to the user interface.

[0141] A comprehensive matching score was obtained for all audio and video content in the expanded candidate set. After calculation, according to The expanded candidate set is sorted from largest to smallest, with higher-ranked audio and video content receiving higher recommendation priority. During push notifications, several audio and video pieces are selected sequentially from the sorted results and presented to the user in ranking order. If there is a comprehensive matching degree within the expanded candidate set... For content with a value of 0 or a very small value, a cutoff threshold can be set to filter candidate content below this threshold from the final push list, ensuring the quality of the pushed content. The entire calculation and sorting process is executed dynamically each time a user triggers a recommendation request, ensuring that the push results can reflect the user's latest interaction status and changes in content preferences in real time.

[0142] A second aspect of this invention provides an intelligent algorithm matching system for high-precision audio and video push, comprising:

[0143] The data acquisition unit is used to acquire user interaction behavior data streams and audio / video content feature datasets.

[0144] The feature extraction unit is used to perform time-series analysis on user interaction behavior data streams, extract behavior pattern vectors, and perform multimodal fusion processing on audio and video content feature datasets to generate content representation vectors.

[0145] The graph construction unit is used to construct an evolutionary graph reflecting the content type transfer relationship based on the user's consumption switching sequence between different content types in the behavior pattern vector;

[0146] The exclusion analysis unit is used to identify the user's interrupted interaction features in the behavior pattern vector, mark the content type corresponding to the interrupted interaction feature as the exclusion type, count the cumulative frequency of each exclusion type, and generate an exclusion distribution map.

[0147] The candidate generation unit is used to take the rejection type with a cumulative frequency exceeding a preset suppression threshold in the rejection distribution map as the avoidance node, perform a graph traversal search starting from the non-avoidance node in the evolution graph, extract the connected path in the traversal path that does not pass through the avoidance node as the bypass path, and extract the audio and video content that matches the content representation vector from the corresponding content type along the bypass path to generate an extended candidate set.

[0148] The recommendation push unit is used to calculate the comprehensive matching degree of audio and video content in the extended candidate set based on the evolutionary map, rejection distribution map and behavior pattern vector, and push audio and video content to users according to the comprehensive matching degree ranking results.

[0149] A third aspect of the present invention provides an electronic device, comprising:

[0150] processor;

[0151] Memory used to store processor-executable instructions;

[0152] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0153] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0154] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for intelligent algorithm matching for high-precision audio and video push, characterized in that, include: Acquire user interaction behavior data streams and audio / video content feature datasets; The user interaction behavior data stream is analyzed in time series to extract behavior pattern vectors. At the same time, the audio and video content feature dataset is fused in multimodal mode to generate content representation vectors. Based on the consumption switching sequence of users between different content types in the behavioral pattern vector, an evolutionary map reflecting the content type transfer relationship is constructed. Identify user interruption interaction features in behavior pattern vectors, mark the content types corresponding to interruption interaction features as exclusion types, count the cumulative frequency of each exclusion type, and generate an exclusion distribution map. The exclusion type with a cumulative frequency exceeding the preset suppression threshold in the exclusion distribution map is taken as the avoidance node. A graph traversal search starting from the non-avoidance node is performed in the evolution graph. Connected paths that do not pass through the avoidance node in the traversal path are extracted as bypass paths. Audio and video content matching the content representation vector is extracted from the corresponding content type along the bypass path to generate an extended candidate set. The comprehensive matching degree of audio and video content in the extended candidate set is calculated based on the evolutionary map, rejection distribution map and behavior pattern vector, and audio and video content is pushed to users according to the ranking result of the comprehensive matching degree.

2. The method according to claim 1, characterized in that, The user interaction behavior data stream is subjected to time-series analysis to extract behavior pattern vectors. Simultaneously, multimodal fusion processing is performed on the audio and video content feature dataset to generate content representation vectors, including: The interaction events in the user interaction behavior data stream are sorted by timestamp to construct a time-series interaction sequence, and a fixed step size is set to segment the time-series interaction sequence by a sliding window. The interaction type, interaction duration, and interaction interval within the window are extracted as behavioral feature dimensions. After normalizing the behavioral feature dimensions, the statistical distribution characteristics within the window and the rate of change of behavioral intensity between adjacent windows are calculated. The statistical distribution characteristics and the rate of change of behavioral intensity are concatenated and encoded to generate a behavioral pattern vector. Visual representations are generated by extracting body contours and texture information from visual frame sequences in audio and video content feature datasets, and audio representations are generated by extracting frequency components and energy distribution from audio waveform sequences through short-time Fourier transform. The system detects whether there is an alignment deviation between visual and audio representations in the time dimension. When an alignment deviation exists, it calculates the sliding cross-correlation function of visual and audio representations in the time dimension, extracts the time offset corresponding to the global maximum value of the sliding cross-correlation function, and then translates the visual representation according to the time offset and sums it with the audio representation element by element to generate a content representation vector.

3. The method according to claim 1, characterized in that, Based on the user consumption switching sequences between different content types in the behavioral pattern vector, an evolutionary map reflecting the content type transfer relationship is constructed, including: Extract user consumption switching sequences between different content types from the behavior pattern vector, identify content type labels of adjacent consumption behaviors in the consumption switching sequence, and combine adjacent content type labels into ordered pairs as content type switching pairs according to the chronological order. The frequency of occurrence of content type switching pairs in the consumption switching sequence is statistically analyzed, and a content type transition matrix is ​​constructed with content type as the row and column index and occurrence frequency as the matrix element. The transition probability of each content type to other content types is calculated by row normalization of the content type transition matrix. The eigenvector of the content type transition matrix is ​​calculated, and the eigenvector corresponding to the largest eigenvalue is extracted as the content type importance vector. Calculate the Shannon entropy for the transition probability distribution of each row in the content type transition matrix, multiply the Shannon entropy by the corresponding element value in the content type importance vector to generate a comprehensive stability index, and sort and label the node level of the content types according to the comprehensive stability index. An evolutionary graph reflecting the content type transfer relationship is constructed using content type as nodes, transfer probability as directed edge weights, and node level as node attributes.

4. The method according to claim 1, characterized in that, Identify user interruption interaction features in the behavior pattern vector, label the content types corresponding to the interruption interaction features as exclusion types, count the cumulative frequency of each exclusion type, and generate an exclusion distribution map, including: Extract the interaction duration and interaction completion fields from the behavior pattern vector. Group and standardize the interaction duration field according to the content type label to obtain the standardized duration. Construct a feature space with the standardized duration and interaction completion field and perform isolation tree anomaly detection to obtain anomaly score. Interaction events with abnormal scores exceeding a preset abnormal threshold are marked as interrupted interaction features, and the corresponding content type tags are extracted and marked as exclusion types; The cumulative frequency is constructed by counting the number of occurrences of the exclusion type in the behavior pattern vector. The cumulative frequency is normalized to generate the exclusion strength. The distribution uniformity of the cumulative frequency is evaluated to obtain the concentration index. In the consumption switching sequence, the location of the rejection type is retrieved and a location index set is constructed. For each location in the location index set, the content type label within a fixed step size is extracted to construct a pre-label matrix. Principal component analysis is performed on the pre-label matrix to extract the dominant feature vector. The loading coefficient of each content type label in the dominant feature vector is calculated as the inducing factor strength. The content type label with the largest inducing factor strength is selected as the rejection inducing factor type. A rejection distribution map is generated using rejection type as the horizontal axis, rejection intensity as the vertical axis, rejection cause type as the node label, and concentration index as the transparency parameter.

5. The method according to claim 4, characterized in that, Principal component analysis is performed on the pre-label matrix to extract the dominant feature vector. The loading coefficients of each content type label in the dominant feature vector are calculated as the inducing factor strength, including: The pre-label matrix is ​​converted into a numerical matrix by one-hot encoding. Singular value decomposition is performed on the numerical matrix to obtain the left singular vector matrix and the singular value sequence. The maximum singular value is selected from the singular value sequence, and the column vector corresponding to the maximum singular value in the left singular vector matrix is ​​extracted as the dominant feature vector. Extract the component values ​​corresponding to each content type label from the dominant feature vector as the pattern contribution, and count the number of times each content type label appears in the position before the exclusion type in the preceding label matrix as the neighbor frequency. Normalization is performed on both the pattern contribution and the neighbor frequency. The normalized pattern contribution and the normalized neighbor frequency are then weighted and fused to obtain the loading coefficient of each content type tag in the dominant feature vector as the inducing force intensity.

6. The method according to claim 1, characterized in that, In the evolutionary graph, a graph traversal search starting from non-avoided nodes is performed. Connected paths that do not pass through avoided nodes are extracted as bypass paths. Audio and video content matching the content representation vector is extracted from the corresponding content type along the bypass paths to generate an extended candidate set, including: In the evolutionary graph, retrieve all reachable paths that start from the causative node corresponding to the causative type of the causative type of the causative type in the causative distribution map and do not pass through the avoidance node. Extract the paths in the reachable paths where the transition probability of the subsequent node is greater than the transition probability of the preceding node as the bypass path. For the extracted detour path, the shortest path length between each node in the detour path and the avoidance node is calculated. Based on the shortest path length, an isolation metric is generated to characterize the degree of isolation between the node and the avoidance node. The transition probability of each edge in the detour path is weighted and fused with the isolation metric of the corresponding node to generate a path stability index. Select the detour path with the highest path stability index, extract the content type corresponding to each node in the detour path, retrieve the audio and video content that matches the content representation vector in the content type, and generate an extended candidate set according to the order of the nodes to which the audio and video content belongs in the detour path.

7. The method according to claim 1, characterized in that, Based on evolutionary maps, rejection distribution maps, and behavioral pattern vectors, a comprehensive matching degree is calculated for the audio and video content in the expanded candidate set. Audio and video content is then pushed to users based on the ranking of the comprehensive matching degree, including: Identify the content type of each audio and video content in the extended candidate set, calculate the shortest path between the node corresponding to the content type and the node corresponding to the current interactive content type in the evolution graph, extract the transition probability of each edge in the shortest path and calculate the product of the transition probabilities as a path reachability metric. Query the rejection strength corresponding to the content type in the rejection distribution map; Query the frequency of users' historical interactions with the content type in the behavior pattern vector, and generate preference weights based on the historical interaction frequency; The path reachability metric and preference weight are weighted and fused to generate an initial matching score. The initial matching score is adjusted based on the repulsion strength to generate a comprehensive matching score for each audio and video content in the extended candidate set. The audio and video content in the extended candidate set is sorted from largest to smallest according to the comprehensive matching score, and audio and video content is pushed to the user according to the sorting results.

8. A high-precision audio and video push intelligent algorithm matching system, used to implement the method as described in any one of claims 1-7, characterized in that, include: The data acquisition unit is used to acquire user interaction behavior data streams and audio / video content feature datasets. The feature extraction unit is used to perform time-series analysis on user interaction behavior data streams, extract behavior pattern vectors, and perform multimodal fusion processing on audio and video content feature datasets to generate content representation vectors. The graph construction unit is used to construct an evolutionary graph reflecting the content type transfer relationship based on the user's consumption switching sequence between different content types in the behavior pattern vector; The exclusion analysis unit is used to identify the user's interrupted interaction features in the behavior pattern vector, mark the content type corresponding to the interrupted interaction feature as the exclusion type, count the cumulative frequency of each exclusion type, and generate an exclusion distribution map. The candidate generation unit is used to take the rejection type with a cumulative frequency exceeding a preset suppression threshold in the rejection distribution map as the avoidance node, perform a graph traversal search starting from the non-avoidance node in the evolution graph, extract the connected path in the traversal path that does not pass through the avoidance node as the bypass path, and extract the audio and video content that matches the content representation vector from the corresponding content type along the bypass path to generate an extended candidate set. The recommendation push unit is used to calculate the comprehensive matching degree of audio and video content in the extended candidate set based on the evolutionary map, rejection distribution map and behavior pattern vector, and push audio and video content to users according to the comprehensive matching degree ranking results.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.