Video quality determination method and server
By integrating video descriptions, user behavior, and comprehensive evaluation data, and performing encoding and correlation analysis to generate fusion features, the problem of inaccurate video quality assessment in existing technologies is solved, achieving accurate and comprehensive video quality assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 青岛聚看云科技有限公司
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies are insufficient for accurate and comprehensive assessment of video quality, and cannot balance the objective attributes of the video with the subjective perception of the user, resulting in biased assessment results.
By acquiring video description data, user behavior data, and comprehensive evaluation data, the data is encoded into feature vectors, and correlation analysis and feature fusion are performed to generate fused features, which are then used to ultimately score the quality.
It achieves accurate and comprehensive evaluation of video quality, taking into account both the objective attributes of the video and the subjective perception of the user, making the score more realistic.
Smart Images

Figure CN121985152A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method and server for determining video quality. Background Technology
[0002] With the rapid development of internet technology, multimedia encoding technology, and mobile terminal devices, video has become one of the main carriers for people to obtain information, engage in entertainment and social interaction, and disseminate knowledge, and is widely used in various scenarios. At the same time, how to accurately assess the quality of video content has become a crucial issue in content distribution, recommendation, and the control of low-quality content.
[0003] Existing technologies have proposed various video quality assessment schemes. These include assessments based on the video's objective characteristics, such as extracting inherent parameters like resolution, frame rate, bitrate, signal-to-noise ratio, and color saturation, and then using a pre-defined scoring model to quantify these parameters and obtain a video quality score; or assessments based on subjective user feedback, such as collecting user feedback likes, comments, and ratings, and directly using the user feedback results as the basis for video quality evaluation. However, these existing technologies struggle to achieve accurate and comprehensive video quality assessment. Summary of the Invention
[0004] This application provides a video quality determination method and server, which enables the final quality score to be more accurate and realistic, truly reflecting the actual quality level of the video, and achieving accurate and comprehensive evaluation of video quality.
[0005] In a first aspect, some embodiments provide a video quality determination method, including:
[0006] Acquire multi-source data of the video to be processed; wherein, the multi-source data includes video description data, user behavior data, and comprehensive evaluation data; the comprehensive evaluation data is determined based on the evaluation data of the video to be processed on different platforms;
[0007] The multi-source data is encoded to obtain video features corresponding to the video description data, behavioral features corresponding to the user behavior data, and evaluation features corresponding to the comprehensive evaluation data.
[0008] A correlation analysis is performed on the video features, the behavioral features, and the evaluation features to obtain correlation information between different features;
[0009] Based on the correlation information between different features, the video features, the behavioral features, and the evaluation features are fused to obtain fused features;
[0010] The fusion features are analyzed to obtain a quality score for the video to be processed.
[0011] In the above embodiments, by integrating multi-source data such as video description data, user behavior data, and comprehensive evaluation data of the video to be processed, the limitations of evaluation based on a single data dimension are overcome, enabling a comprehensive consideration of video quality and avoiding evaluation bias caused by incomplete data. Furthermore, by encoding and extracting multi-dimensional features and conducting correlation analysis, the inherent relationships between different features are fully explored, making feature fusion more targeted and improving data utilization efficiency. In addition, this embodiment obtains a quality score based on the fusion of multi-dimensional features according to the correlation information between different features, taking into account both the objective attributes of the video and the subjective perception of users, and integrating evaluation results from multiple platforms, making the final quality score more accurate and realistic, truly reflecting the actual quality level of the video, and achieving a precise and comprehensive evaluation of video quality.
[0012] Secondly, some embodiments also provide a video quality determination apparatus, comprising:
[0013] The data acquisition module is used to acquire multi-source data of the video to be processed; wherein, the multi-source data includes video description data, user behavior data, and comprehensive evaluation data; the comprehensive evaluation data is determined based on the evaluation data of the video to be processed on different platforms;
[0014] The data encoding module is used to encode the multi-source data to obtain video features corresponding to the video description data, behavioral features corresponding to the user behavior data, and evaluation features corresponding to the comprehensive evaluation data.
[0015] The information determination module is used to perform correlation analysis on the video features, the behavioral features, and the evaluation features to obtain correlation information between different features;
[0016] The feature fusion module is used to fuse the video features, the behavioral features, and the evaluation features based on the correlation information between different features to obtain fused features;
[0017] The quality scoring module is used to analyze the fusion features and obtain a quality score for the video to be processed.
[0018] Thirdly, some embodiments also provide a server, including:
[0019] At least one processor, and configured as follows:
[0020] Acquire multi-source data of the video to be processed; wherein, the multi-source data includes video description data, user behavior data, and comprehensive evaluation data; the comprehensive evaluation data is determined based on the evaluation data of the video to be processed on different platforms;
[0021] The multi-source data is encoded to obtain video features corresponding to the video description data, behavioral features corresponding to the user behavior data, and evaluation features corresponding to the comprehensive evaluation data.
[0022] A correlation analysis is performed on the video features, the behavioral features, and the evaluation features to obtain correlation information between different features;
[0023] Based on the correlation information between different features, the video features, the behavioral features, and the evaluation features are fused to obtain fused features;
[0024] The fusion features are analyzed to obtain a quality score for the video to be processed.
[0025] Fourthly, some embodiments also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the methods provided in some embodiments of the first aspect.
[0026] Fifthly, some embodiments also provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the methods provided in some embodiments of the first aspect. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A flowchart illustrating a video quality determination method provided in some embodiments of this application;
[0029] Figure 2 Flowcharts of the video recommendation process provided for some embodiments of this application;
[0030] Figure 3 This is a flowchart illustrating the process of determining multi-source data for a video to be processed, provided in some embodiments of this application.
[0031] Figure 4 This application provides a flowchart illustrating the process of determining video description data for a video to be processed, based on some embodiments of the present application.
[0032] Figure 5 A flowchart illustrating the process of determining multimodal features of a video segment according to some embodiments of this application;
[0033] Figure 6 This application provides a flowchart illustrating the process of determining comprehensive evaluation data in some embodiments.
[0034] Figure 7 A flowchart illustrating the process of determining user behavior data for a video to be processed, provided in some embodiments of this application;
[0035] Figure 8 A schematic diagram illustrating the determination of correlation information between different features according to some embodiments of this application;
[0036] Figure 9 A flowchart illustrating the process of determining fusion features provided for some embodiments of this application;
[0037] Figure 10 A flowchart illustrating the process of determining the quality score of a video to be processed, provided for other embodiments of this application;
[0038] Figure 11 A flowchart illustrating a video quality determination method provided in some embodiments of this application;
[0039] Figure 12 This is a schematic diagram of the structure of a video quality determination device provided in some embodiments of this application;
[0040] Figure 13 This is an internal structural diagram of a computer device provided in some embodiments of this application. Detailed Implementation
[0041] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.
[0042] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0043] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.
[0044] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0045] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0046] Traditional video quality assessment primarily relies on methods such as manual review and scoring, objective metrics based on video features, and feedback metrics based on user behavior. However, manual review and scoring are highly subjective, costly, and inefficient; objective metrics based on video features only reflect technical quality and fail to capture the quality at the content comprehension level; and feedback metrics based on user behavior are significantly influenced by external factors such as recommendation strategies and timing, failing to comprehensively reflect the intrinsic quality of the content. In recent years, large language models have demonstrated powerful semantic understanding and reasoning capabilities, enabling them to abstractly understand the plot logic, information density, emotional expression, and value orientation of video content. However, how to combine these comprehension capabilities with multi-dimensional objective data to achieve quantitative assessment of video content quality remains a challenge.
[0047] Based on this, in some embodiments, a video quality determination method is provided. This video quality determination method can be implemented by a server.
[0048] In an exemplary embodiment, taking the application of the video quality determination method to a processor in a server as an example, the server includes at least one processor for performing video-related tasks, such as determining video quality, filtering and recommending videos; further, the server may also include a communication device. The communication device can be communicatively connected to a terminal, and this communication connection can be a wired communication connection or a wireless communication connection. Optionally, at least one processor is connected to the communication device.
[0049] Among them, the terminal is a device that can interact with the user. It can be integrated into the server and is part of the server, such as the server's input and output devices, so that the server's communication device and the terminal achieve internal communication within the server; or it can be other electronic devices that are independent of the server, including but not limited to various personal computers, laptops, smartphones, tablets, display devices (such as smart TVs), etc., so that the server's communication device and the terminal achieve cross-device communication.
[0050] like Figure 1 As shown, the video quality determination and generation method applied to the processor in a server includes the following steps:
[0051] S101, acquire multi-source data of the video to be processed.
[0052] In some embodiments, multi-source data may include at least video description data, user behavior data, and comprehensive evaluation data. The comprehensive evaluation data is determined based on evaluation data of the video to be processed across different platforms.
[0053] The videos to be processed can be long videos (such as movies, TV series, documentaries, etc.) or short videos. These videos can be played on various video playback platforms for users to watch.
[0054] Video description data is raw data used to characterize the attributes and content features of the video itself. It is the basic static information of the video and can generally be obtained without user intervention.
[0055] User behavior data is dynamic data formed after users interact with the video to be processed on various platforms, reflecting their actual participation and behavioral preferences. For example, interaction behaviors can include clicks, interactive behaviors, transactional behaviors, playback behaviors, and feedback behaviors. Click behaviors include clicking the video's cover, opening the details page, and clicking related recommendations; interactive behaviors include sending bullet comments, commenting, favorites, sharing, adjusting playback speed, switching resolution, recommending, watching together, and sharing on social media platforms; transactional behaviors include paid clicks, purchases, and membership subscriptions; playback behaviors include playback duration, completion rate, pause / resume times, and dragging positions; feedback includes ratings, likes, dislikes, reports, and problem feedback.
[0056] The comprehensive evaluation data is a unified evaluation data obtained by integrating the original evaluation data of the video to be processed on different dissemination or publishing platforms, taking into account the comprehensiveness of evaluations from multiple platforms.
[0057] For the video to be processed, three types of raw data are collected: video description data, user behavior data, and comprehensive evaluation data. The multi-source data can be directly obtained raw data or data obtained after preprocessing the raw data. In one optional embodiment, the multi-source data is directly obtained raw data. For the video description data, the original static information of the video to be processed is collected, including basic attributes (such as duration, resolution, frame rate, release time, and genre category), content features (such as audio parameters, visual style, keywords, tags, plot summary, and shot transition frequency), and release attributes (such as the qualifications of the releasing account and the basic category of the releasing platform).
[0058] For user behavior data, a real-time processing engine can be integrated to collect data on core user behaviors such as clicks, plays, interactions, payments, and feedback related to the video to be processed, through logs from various video playback platforms. For comprehensive evaluation data, a distributed crawler architecture can be used to crawl raw evaluation data from various platforms, including but not limited to professional ratings (such as overall rating, number of ratings, and distribution of each star rating), professional comments (review text, number of likes, and usefulness polls), and ranking data (such as the ranking changes and trend data of the video to be processed on various platforms). Then, the data from multiple platforms is weighted and integrated according to platform weights to generate unified comprehensive evaluation data. The platform weights can be set based on the platform's traffic scale, user accuracy, and the authenticity of the evaluations.
[0059] In another optional embodiment, the multi-source data is the data obtained after preprocessing the original data. For the original video description data, a large language model can be used to perform semantic analysis on the original video description data beforehand, obtaining processed data as the video description data. For the original user behavior data, anomaly detection algorithms can be used to detect anomalies in the acquired original user behavior data, marking and removing abnormal data that clearly does not conform to normal user behavior; furthermore, feature extraction can be performed on the original user behavior data after removing the abnormal data to obtain the original user behavior data. For the original comprehensive evaluation data, the differentiated rating systems of different platforms can be unified into a single standard, eliminating the incomparability of evaluation data caused by different platform rating rules, thus obtaining the original comprehensive evaluation data.
[0060] S102, Encode the multi-source data to obtain the video features corresponding to the video description data, the behavioral features corresponding to the user behavior data, and the evaluation features corresponding to the comprehensive evaluation data.
[0061] Among them, video features are feature vectors obtained by transforming video description data; behavioral features are feature vectors obtained by transforming user behavior data encoding; and evaluation features are feature vectors obtained by transforming comprehensive evaluation data.
[0062] Optionally, the multi-source data acquired in the step can be transformed into quantifiable and analyzable structured features, resulting in video features corresponding to video description data, behavioral features corresponding to user behavior data, and evaluation features corresponding to comprehensive evaluation data.
[0063] In one optional embodiment, to address the attribute differences among video description data, user behavior data, and comprehensive evaluation data, an adaptive encoding method is employed to transform the raw data into standardized feature vectors, yielding video features, behavioral features, and evaluation features respectively. For example, the video description data may include different data types such as numerical, textual, and visual data. Methods such as Min-Max normalization and Z-Score normalization can be used to map the numerical video description data (e.g., resolution, frame rate, bitrate) to a unified range, preserving the original data patterns and forming a basic video feature vector. For textual data such as video titles, descriptions, and tags, natural language processing techniques can be used to transform the text into fixed-dimensional text embedding vectors, serving as video content features reflecting the video's content attributes. For visual data such as video image clarity and scene complexity, computer vision techniques can be used to extract visual features from video frames, concatenating them with numerical and textual features to form a complete video feature vector. Finally, the features obtained after encoding different data types are fused into a single video feature.
[0064] For user behavior data, statistical indicators can be calculated from user interaction data such as play counts, completion rates, and likes. These calculated indicators are then standardized to form statistical behavioral features, reflecting the user's engagement and preference for the videos being processed. For playback behavior data such as viewing duration, fast-forwarding, rewinding times, and pause points, temporal coding techniques can be used to transform the behavioral sequences into fixed-dimensional sequence feature vectors, capturing the temporal patterns and preference details of user behavior. Finally, the features obtained from encoding different types of books are merged into a single behavioral feature.
[0065] For comprehensive evaluation data, weighted fusion and normalization processing can be used to map the comprehensive score data after integration of multiple platforms to the interval [0,10] or [0,1] to form basic evaluation features. At the same time, the variance and standard deviation of the scores of multiple platforms are calculated to reflect the consistency of the evaluation and serve as supplementary evaluation features.
[0066] In another optional embodiment, corresponding feature encoders can be constructed for the video description data, user behavior data, and comprehensive evaluation data, respectively. The corresponding feature encoders are then used to encode the video description data, user behavior data, and comprehensive evaluation data to obtain the video features corresponding to the video description data, the behavioral features corresponding to the user behavior data, and the evaluation features corresponding to the comprehensive evaluation data.
[0067] S103, perform correlation analysis on video features, behavioral features and evaluation features to obtain correlation information between different features.
[0068] The correlation information between different features includes the correlation information between video features and behavioral features, the correlation information between video features and evaluation features, and the correlation information between behavioral features and evaluation features. Optionally, to ensure the richness of the correlation information, the correlation information between different features may also include the correlation information between video features, behavioral features, and evaluation features.
[0069] Optionally, the inherent relationships among video features, behavioral features, and evaluation features can be explored to clarify the degree of influence between different features, providing a basis for subsequent feature fusion, avoiding feature redundancy caused by blind fusion, and improving the effectiveness of fused features. For example, correlation analysis methods include, but are not limited to, linear correlation analysis methods (such as Pearson correlation coefficient analysis, Spearman rank correlation coefficient analysis, etc.), nonlinear correlation analysis methods (such as mutual information analysis methods, decision tree or random forest correlation analysis methods, etc.), and multi-feature comprehensive correlation analysis methods (association rule mining algorithms, attention mechanism correlation analysis methods, etc.). The specific correlation analysis method used can be selected according to specific needs and actual conditions, and is not limited here.
[0070] S104. Based on the correlation information between different features, video features, behavioral features, and evaluation features are fused to obtain fused features.
[0071] In one optional embodiment, for any given feature, invalid sub-dimension features with extremely low correlation to the other two types of features can be removed, reducing the computational load of fusion and improving feature effectiveness. Based on the correlation information with the other two features, a fusion weight is assigned to the feature, ensuring that the sum of the fusion weights is 1; the higher the correlation, the larger the fusion weight. For example, the fusion weight of the feature can be determined based on the ratio of the average correlation of the feature with the other two types of features to the sum of the average correlations of all features. Finally, the three types of features are multiplied by their corresponding fusion weights, and then integrated into a unified fusion feature vector through concatenation or superposition.
[0072] In another optional embodiment, based on the correlation information between different features obtained above, the correlation weight between each feature and video quality can be calculated. The higher the correlation, the greater the weight. Then, the video features, behavioral features, and evaluation features are multiplied by their respective weights, and then summed and concatenated to obtain a fused feature vector.
[0073] S105, analyze the fusion features to obtain the quality score of the video to be processed.
[0074] Among them, the quality score of the video to be processed is a numerical result that can directly represent the overall quality level of the video to be processed.
[0075] In one optional embodiment, based on the fused feature vector, an adapted analysis model is used to perform quantitative analysis on it, and finally outputs a standardized quality score of the video to be processed. The score range can be set according to the actual scenario. For example, the score range of the quality score can be 1-10 points or 1-100 points, which is not limited here.
[0076] For example, an appropriate analysis model can be selected to analyze the fused features based on the complexity of the application scenario. For instance, a traditional statistical model can be used for simple scenarios, while a machine learning or deep learning model can be used for complex scenarios. The fused features are input into the trained analysis model, which then outputs the original score of the video to be processed through quantization calculation. The original score output by the analysis model is then mapped to a preset score range to obtain the final standardized quality score.
[0077] Optionally, application scenarios include, but are not limited to, video recommendation, video management, and video rating. Specifically, for video recommendation, quality ratings can be used as the core weighting indicator. Videos with high quality ratings receive increased traffic and recommendation priority, while videos with low quality ratings receive less recommendation or are relegated to secondary traffic pools. This achieves precise reach of high-quality content, improving platform content quality and user viewing experience. For example, the video recommendation process generally includes recall, sorting, and reordering. Figure 2 The diagram shows a flowchart of the video recommendation process provided in this application embodiment. The content library includes all videos to be processed, serving as recommended content. During the recall process, quality control can be achieved at the recall layer through a multi-source recall fusion mechanism based on the quality score of the videos to be processed, ensuring that the videos entering the candidate pool meet a basic quality threshold. Simultaneously, a threshold is set; videos with quality scores below the threshold directly enter the manual review process. If the manual review is passed, the video is sent to the candidate pool; if the manual review fails, the video is either optimized or directly filtered. Videos with quality scores greater than or equal to the threshold are directly sent to the candidate pool. A random sampling algorithm based on quality scores is involved, constructing a quality-weighted recall mechanism that dynamically adjusts parameters to balance the quality and freshness of the videos to be processed. A dual-retrieval mechanism of "interest vector-quality vector" is constructed, along with a collaborative recall mechanism. Based on traditional user interest vectors, a quality feature vector is added to achieve synergy between interest matching and quality screening. During the ranking process, a quality-interest dual-objective optimization model can be used at the ranking layer, using the quality score of the videos to be processed as a constraint and optimization objective to achieve a balance between long-term user value and short-term conversion. During the rearrangement process, the diversity of recommendation results can be optimized through a quality diversity rearrangement algorithm at the rearrangement layer.
[0078] In the above embodiments, multi-source data such as video description data, user behavior data, and comprehensive evaluation data of the video to be processed are integrated to overcome the limitations of evaluation based on a single data dimension, achieve a comprehensive consideration of video quality, and avoid evaluation bias caused by one-sided data. By encoding and extracting multi-dimensional features and conducting correlation analysis, the inherent relationships between different features are fully explored, making feature fusion more targeted and improving data utilization efficiency. Based on the correlation information between different features, a quality score is obtained after fusing multi-dimensional features, taking into account both the objective attributes of the video and the subjective perception of users, and integrating evaluation results from multiple platforms, making the final quality score more accurate and realistic, and truly reflecting the actual quality level of the video, thus achieving an accurate and comprehensive evaluation of video quality.
[0079] Based on the above embodiments, in an exemplary embodiment, the process of determining the multi-source data of the video to be processed, in S101, using multi-source data as the data obtained after data preprocessing of the original data, is further refined. Optionally, such as Figure 3 As shown, it includes the following steps:
[0080] S301: Obtain the video to be processed, user log data for the video to be processed, and evaluation data of the video to be processed on different platforms.
[0081] The user log data consists of raw logs of user interactions with the videos to be processed, recorded by various video playback platforms. This data is either time-series or fragmented structured data. The evaluation data for the videos to be processed across different platforms consists of raw evaluation data for the videos on different platforms, including quantified ratings, text reviews, and tagged feedback.
[0082] Optionally, targeted data collection can be performed on the videos to be processed, user log data, and evaluation data from various platforms. For example, a direct collection method using the video playback platform API (Application Programming Interface) can be used to obtain the videos to be processed, user log data, and evaluation data from various platforms through open and fixed interfaces. Alternatively, a targeted collection method using web crawlers can be employed to obtain the videos to be processed, user log data, and evaluation data from various platforms.
[0083] S302, Perform semantic analysis on the video to be processed to obtain video description data of the video to be processed.
[0084] Optionally, multi-dimensional semantic analysis techniques can be used to process the video to obtain video description data that reflects the core content attributes of the video, solving the problem that video content is difficult to directly use for subsequent feature extraction. For example, image data, audio data, and text data can be extracted from the video to be processed. Furthermore, semantic analysis is performed on the image data, audio data, and text data respectively to mine the core semantic information of the video. The core semantic information obtained from the semantic analysis is then structured, integrated, and quantified to form video description data in a unified format that can be directly used for subsequent encoding.
[0085] S303, extract features from user log data to obtain user behavior data for the video to be processed.
[0086] One optional embodiment transforms user log data into quantifiable, dimensional, and analyzable user behavior data. Optionally, user behavior can be categorized by type, such as click behavior, interaction behavior, playback behavior, transaction behavior, and feedback (at least two of these). Different user behavior types are then analyzed to extract metrics. For example, for click behavior, metrics include, but are not limited to, cover click-through rate, detail page open rate, and related recommendation click distribution; for interaction behavior, metrics include, but are not limited to, sending bullet comments, commenting, favorites, sharing, speed adjustment, resolution switching, recommendations, watching, collaborative watching, and sharing on social media platforms; for playback behavior, metrics include, but are not limited to, playback duration, completion rate, pause / resume count, and drag operation position; for transaction behavior, metrics include, but are not limited to, transaction clicks, transaction conversion rate, membership activation path, and transaction amount; and for feedback, metrics include, but are not limited to, ratings, likes, dislikes, reports, and problem feedback content.
[0087] One optional implementation can extract the core metric values based on two fundamental dimensions: playback behavior and interaction behavior. Without complex calculations, user behavior data can be directly generated statistically. Alternatively, metric values can be extracted based on three core dimensions: playback behavior, interaction behavior, and feedback behavior. Basic data calculations, such as rate values and average values, can be performed to generate multi-dimensional user behavior data, eliminating raw logs with no analytical value. Furthermore, metric values can be extracted from multiple dimensions, including click behavior, interaction behavior, playback behavior, transaction behavior, and feedback, and complex calculations and analyses can be performed to generate comprehensive, fine-grained user behavior data, uncovering deeper patterns in user behavior.
[0088] S304 integrates the evaluation data of the video to be processed on different platforms to obtain comprehensive evaluation data.
[0089] Optionally, since evaluation data from different platforms differ in standards, formats, and weights, multi-platform evaluation data can be integrated into unified, standardized, and objective comprehensive evaluation data to avoid the bias of evaluations from a single platform. This can be achieved by processing the evaluation data from different platforms before weighting and merging them.
[0090] For example, weighted fusion can employ a simple weighted fusion algorithm, which calculates the data by weighting it according to preset fixed weights to generate comprehensive evaluation data. Considering that evaluation data from different platforms includes both quantitative scores and textual evaluations, a weighted fusion method combining quantitative and sentiment dimensions can also be used. This involves weighting and fusing the quantitative scores and textual evaluations from each platform separately, and then integrating the fusion results of the two dimensions according to preset weights to generate comprehensive evaluation data in two dimensions.
[0091] In this embodiment, semantic analysis is performed on the video to be processed to extract video description data, accurately mining the core content attributes of the video and avoiding the problem that the raw content data alone cannot be directly used for subsequent feature analysis; user behavior data is obtained by targeted feature extraction from user log data, improving the effectiveness of user behavior data and more accurately reflecting the user's actual interaction feedback and acceptance of the video; cross-platform evaluation data is fused to obtain comprehensive evaluation data, realizing the standardization and unification of different platform rating systems and evaluation forms, making the evaluation dimension data more comprehensive and objective.
[0092] Based on the above embodiments, in an exemplary embodiment, the video description data includes video type, emotion data, and narrative rhythm data; accordingly, the process of determining the video description data of the video to be processed in S302 is further refined. Optionally, such as Figure 4 As shown, it includes the following steps:
[0093] S401, the video to be processed is segmented to obtain at least two video segments.
[0094] Among them, a video segment is an independent video unit obtained by splitting the complete video to be processed.
[0095] In one optional embodiment, the long video to be processed can be divided into video segments of preset duration according to scene boundaries, with each video segment having the same duration. The preset time period can be 5 minutes or 10 minutes. For example, a keyframe sampling strategy can be used to segment the video to be processed, thereby reducing computational complexity while preserving narrative integrity.
[0096] Another optional embodiment can use video shots as the core segmentation basis. Shot transition points are determined by detecting inter-frame abrupt changes in the video frame, and segmentation is completed at these transition points. Each video segment corresponds to an independent shot or a group of consecutive shots. For example, during shot transitions, significant abrupt changes occur in the pixel differences, color histograms, and edge features of adjacent video frames. Therefore, the threshold for these abrupt changes can be determined by quantifying these indicators. For instance, the difference in pixel grayscale values between adjacent frames can be calculated; if the difference exceeds a preset threshold, a shot transition point is identified. Similarly, the similarity of the color histograms between adjacent frames can be calculated; if the similarity is below a preset threshold, a shot transition point is identified. Edge features between adjacent frames can also be extracted; if the edge feature matching degree is below a preset threshold, a shot transition point is identified. After extracting the shot transition points, the video is cut at these points to generate at least two video segments, each based on a shot. This ensures that the resulting video segments conform to the visual creation logic of the video, the segmentation results are consistent with the rhythm of the video frame presentation, and feature extraction avoids cross-shot confusion.
[0097] After all segmentation methods are completed, video segments can be numbered and labeled with attributes (such as video segment 1 / 2 / 3, corresponding time interval, number of shots) to facilitate subsequent feature extraction and correlation analysis.
[0098] S402: For each video segment, feature extraction is performed on the video segment to obtain the multimodal features of the video segment.
[0099] Multimodal features are a set of multi-dimensional structured features extracted from video segments, representing a digital, multi-dimensional representation of the video segment's content. Multimodal features include camera motion features, speech content features, and semantic features. Camera motion features characterize the motion attributes of camera shots and editing within the video segment, reflecting the visual rhythm and dynamics of the video. Speech content features characterize the acoustic properties and content information of the audio in the video segment, reflecting the intonation, speech rate, and core speech content. Semantic features characterize the core semantic information of the visual content and text within the video segment, reflecting the video segment's theme and core expression.
[0100] Optionally, for each independent video segment, structured and quantifiable features can be extracted from three modalities: camera motion, audio content, and semantics, generating a multimodal feature vector for each video segment. The feature vectors of all video segments constitute the multimodal features of the complete video. Features can be extracted independently for each video segment, standardized for features of the same modality, and unified in multimodal feature dimensions, ensuring the computability and comparability of the features.
[0101] S403 determines the video type and sentiment data of the video to be processed based on the multimodal features of different video segments.
[0102] Among them, sentiment data quantification represents the overall sentiment tendency and intensity of the video to be processed, such as sentiment orientation (positive, negative, neutral, etc.). Video type represents the theme expressed by the video to be processed.
[0103] Optionally, based on the multimodal features of all video segments obtained, the video type (such as food, technology, entertainment, drama, etc.) and the sentiment data (such as overall sentiment tendency, sentiment intensity, and sentiment distribution of each video segment) of the video to be processed can be obtained through feature aggregation, model determination, and rule matching.
[0104] One optional implementation classifies videos by content theme (e.g., food, technology, entertainment, scenery, education), presentation style (e.g., short videos, vlogs, tutorials, dramas, reviews), and dissemination attributes (e.g., humor, healing, suspense, inspirational). It supports single or multiple type tags, such as "food + tutorial." Type matching rules can be formulated based on the semantic features of video segments. The video type of a segment is determined by matching these rules, and then aggregated into an overall video type, without the need for model training. Specifically, a rule base needs to be pre-defined, setting core keywords or feature rules for each video type. Furthermore, the semantic features of each video segment are matched against the rule base; the rule with the highest matching degree corresponds to the video segment type. The duration percentage of each type in all video segments is calculated, and the top 1-2 types with the highest percentages are identified as the overall video type.
[0105] Another optional implementation involves first aggregating the multimodal features of all video segments into a single overall video feature vector, and then using a pre-trained classification model to determine the video type, supporting multi-type label output. The classification model can be a multimodal semantic model. For example, weights are assigned to the multimodal features of each video segment based on its duration, with longer segments receiving larger weights. The multimodal features of all video segments are then weighted and summed to generate the overall multimodal feature vector. This overall feature vector is input into the pre-trained classification model, which outputs the video's type label and matching probability. The top 1-2 probabilities are selected as the final type.
[0106] Optionally, for the sentiment data of the video to be processed, a sentiment classification model can be trained based on the multimodal features of video segments. This first determines the sentiment tendency and intensity of each video segment, and then aggregates them into overall video sentiment data, taking into account multimodal information and achieving higher accuracy. For example, the multimodal features of each video segment are input into a pre-trained sentiment classification model, and the model outputs the sentiment tendency and intensity values of the video segments as sentiment data. Further, the overall sentiment intensity value is calculated by weighting the video segment duration, and the overall sentiment tendency is determined based on the intensity value range.
[0107] S404: Extract the narrative structure of the video to be processed, and determine the narrative rhythm data of the video to be processed based on the narrative structure.
[0108] The narrative structure refers to the organization and presentation logic of the video content, representing the arrangement of the video plot and content. The narrative structure includes information such as the duration of each of the four stages: introduction, development, climax, and conclusion. Narrative rhythm data includes the structural distribution and rhythm score.
[0109] Optionally, narrative structure is fundamental, and narrative rhythm is the quantitative expression of that structure; different narrative structures correspond to different narrative rhythms. The narrative rhythm can be precisely quantified by combining the duration, feature intensity, and content density of different narrative stages in the video to be processed.
[0110] In one optional embodiment, the narrative structure can be divided into four stages: introduction, development, climax, and conclusion. The introduction is the opening stage, such as introducing the theme, providing background, and setting up suspense—the initial setup for the video content. The development stage is the content unfolding stage, such as advancing the plot, explaining details, and reinforcing the setup—the core exposition of the video content. The climax stage is the content turning point or climax stage, such as plot twists, reaching an emotional climax, resolving suspense, and presenting the core result—the core turning point of the video narrative. The conclusion stage is the ending stage, such as summarizing key points, elevating the theme, explaining the ending, and giving advice—the final conclusion of the video content. The time boundaries of the four stages can be extracted, and the duration of each stage can be calculated as the structural distribution, while simultaneously calculating a rhythm score.
[0111] The time boundaries refer to the start and end timestamps of each stage. For example, the "start" stage (00:00-00:15) has time boundaries of 00:00 and 00:15. The time percentage is the ratio of the actual duration of each stage to the total duration of the video to be processed, reflecting the allocation of narrative length in each stage.
[0112] Optionally, core feature thresholds can be set for the four stages of introduction, development, transition, and conclusion. For example, the intensity of camera movement and emotional intensity in the "transition" stage should reach their peak. By using feature mutation detection, threshold matching, or model judgment, time nodes that conform to the characteristics of each stage are found as the time boundaries of the four stages. The duration of each stage is calculated based on the time boundaries, and then compared with the total duration to obtain the percentage result. Furthermore, the narrative rhythm of the four stages is quantitatively scored from three dimensions: rhythm within each stage, connection between stages, and overall narrative fit. The higher the score, the more reasonable the rhythm and the more it fits the video type and narrative structure.
[0113] The above embodiments reduce the analysis granularity to the camera or content logic unit by segmenting the video, and then perform refined multimodal feature extraction on each video segment. This can capture the fine-grained content features of the video, avoid feature ambiguity and information loss caused by overall analysis, and restore the actual content and expression of the video to the greatest extent. By extracting features from three modalities—camera movement, audio content, and semantics—the three core dimensions of video—visual, auditory, and content—it can effectively capture the composite content features of the video. The narrative structure and narrative rhythm of the video are extracted and quantitatively analyzed in a structured manner, realizing the quantification of video narrative features.
[0114] Based on the above embodiments, in an exemplary embodiment, the process of determining the multimodal features of the video segment in S402 is further refined. Optionally, such as Figure 5 As shown, it includes the following steps:
[0115] S501 acquires image data, audio data, and text data from a video segment.
[0116] Among them, image data is frame-level visual data extracted from video segments, which are the images in the video segments; audio data is pure audio stream data extracted from video segments; and text data is structured information converted from video segments, including speech-to-text and subtitle text.
[0117] Optionally, for each video segment, three types of raw data—image data, audio data, and text data—are selectively extracted from that video segment to ensure data integrity, temporal alignment, and consistency.
[0118] S502 extracts features from the image data to obtain the camera motion features of the video segment.
[0119] Optionally, inter-frame feature analysis and motion detection algorithms can be used to extract quantitative indicators, classification features, and high-dimensional vector features that characterize the motion of video clips. The extracted features should be able to reflect the shooting motion, editing motion, and dynamics of the footage.
[0120] In one optional embodiment, the pixel difference or histogram difference between adjacent frames in the image data can be calculated to obtain the inter-frame difference value; the inter-frame difference value can be matched with the rule base to determine the camera motion type, and a classification feature can be assigned to each type; the number of camera switches within the video segment can be counted as the basic quantitative feature, and finally the camera motion type and the number of camera switches can be output as camera motion features.
[0121] S503 converts the audio data into a Mel spectrogram and extracts features from the Mel spectrogram to obtain the speech content features of the video segment.
[0122] Among them, the Mel spectrogram is a time-frequency domain visualization image that transforms audio waveform data through the Mel scale, preserving the core acoustic features of the audio such as frequency, duration, and energy.
[0123] Optionally, the audio data can be converted into a Mel spectrogram first, and then quantitative indicators and high-dimensional acoustic feature vectors representing speech content can be extracted from the Mel spectrogram. The features should be able to reflect the core acoustic attributes of speech such as speech rate, pitch, volume, energy, and frequency distribution.
[0124] S504 performs word segmentation on the text data and extracts features from the segmented text data to obtain the semantic features of the video segment.
[0125] Optionally, the text data can be segmented into words, with the segmentation tailored to the characteristics of the video domain; a second cleaning process should follow after segmentation. For example, the segmentation process could involve using a general Chinese word segmentation tool to segment the text data, loading a general stop word library to filter out meaningless words; alternatively, a pre-trained model could be used for contextual semantic segmentation to maximize segmentation accuracy, loading a customized hybrid stop word library for fine-tuning, and labeling the segmentation results with word weights to reflect word importance, resulting in a weighted word sequence.
[0126] Optionally, a pre-trained general word vector model is used to transform the segmented word sequence into a set of word vectors. The word vector set is then fused into a text vector using the mean or weighted mean (by word weight). The text vectors are then standardized to generate semantic features.
[0127] The above embodiments design dedicated processing and feature extraction paths for the three core modalities of image, audio, and text, avoiding feature ambiguity caused by multimodal mixed processing and improving the accuracy and targeting of single-modal feature extraction; from raw data acquisition to modal preprocessing and then to feature extraction, noise, errors and redundancy in the original data are reduced to the greatest extent, ensuring the reliability and accuracy of the obtained multimodal features and avoiding feature bias caused by low-quality data.
[0128] Based on the above embodiments, in an exemplary embodiment, the process of determining the comprehensive evaluation data in S304 is further refined. Optionally, such as Figure 6 As shown, it includes the following steps:
[0129] S601 determines the deviation factor and platform weight for different platforms based on the platform attribute information of different platforms.
[0130] The platform attribute information for each platform is structured data representing its core characteristics, encompassing four main categories: user profiles, content positioning, evaluation mechanisms, and data characteristics. The deviation factor for each platform quantifies the coefficient of systematic bias in its evaluation data, used to correct for differences in evaluation standards between platforms, ensuring comparability of the corrected data. The platform weight for each platform reflects the credibility, representativeness, and relevance of its evaluation data; a higher weight indicates a greater impact on the overall evaluation result.
[0131] Optionally, based on the platform attribute information of each platform, corresponding methods can be used to calibrate the deviation factor and platform weight respectively, to ensure that the deviation factor can accurately offset the systematic deviation of the platform, and to ensure that the platform weight can reflect the credibility and representativeness of the platform evaluation data.
[0132] In one optional embodiment, a batch of cross-platform videos with the same source (consistent content) can be selected, and the average evaluation score of each platform can be calculated. Based on the average score of the entire platform, the deviation factor is calculated as the average score of the entire platform / the average score of the platform. When there are no videos with the same source, the platform's evaluation severity attribute can be directly used as the deviation factor.
[0133] Another optional embodiment can integrate platform attribute features, video type features, and evaluation index features to construct a nonlinear model, such as a random forest or neural network model; using manually labeled objective evaluations as labels, the model is trained to learn the mapping relationship from single-platform data to objective data, and the mapping function output by the model is the bias factor.
[0134] In one optional embodiment, a weight judgment matrix can be constructed based on platform attributes, such as data authenticity, matching degree, and professionalism, and the importance of each attribute can be compared pairwise. The eigenvectors of the judgment matrix are calculated and normalized to obtain the weights of each platform, ensuring that the sum of the weights is 1.
[0135] Another optional embodiment can be based on the attribute characteristics of the platform, and use the entropy weight method to calculate the objective weight of each attribute. The smaller the entropy value, the higher the attribute discrimination and the greater the weight. The attribute characteristics of each platform are weighted and summed, and the weight is the attribute entropy weight to obtain the platform comprehensive score. The comprehensive score is normalized and used as the final platform weight.
[0136] S602, for any platform, corrects the evaluation data of the video to be processed on the platform according to the platform's deviation factor, and obtains the platform's standardized evaluation data.
[0137] Among them, the standardized evaluation data of the video to be processed on each platform is the objective data obtained by correcting the evaluation data of that platform for the deviation factor, thereby eliminating the systematic bias of the platform.
[0138] Optionally, for any platform, the evaluation data of the video to be processed on that platform can be preprocessed, such as outlier extraction, metric normalization, and data completion. Furthermore, different data correction methods can be used to correct the evaluation data of the video to be processed on that platform.
[0139] In one alternative embodiment, the preprocessed raw evaluation data can be multiplied by the platform's deviation factor to obtain the platform's standardized evaluation data.
[0140] S603 weights and merges standardized evaluation data from different platforms based on their respective platform weights to obtain comprehensive evaluation data.
[0141] Optionally, the standardized evaluation data of each platform can be weighted and merged according to the determined platform weights to generate an evaluation result that reflects the overall performance of the video across all platforms.
[0142] In one optional embodiment, the standardized evaluation data for each platform may include the values of multi-dimensional indicators. First, the indicator values of the multi-dimensional indicators for a single platform are manually assigned weights, and the comprehensive score of the single platform is calculated. Further, the comprehensive scores of each platform are weighted and summed using the platform weights to obtain the final comprehensive evaluation data.
[0143] Another alternative implementation involves first constructing a fusion model, such as an attention-based neural network, and inputting standardized data and platform weights from each platform. The fusion model then dynamically adjusts the weight ratios of each platform and indicator through an attention mechanism to adapt to the video characteristics, ultimately outputting comprehensive evaluation data.
[0144] The above embodiments effectively offset the systematic biases caused by differences in evaluation standards, user habits, and mechanisms across different platforms by using attribute-driven bias factor calibration for different platforms. This makes cross-platform evaluation data not only incomparable but also directly comparable. By quantifying weights based on platform attributes, the biases of subjective human assignment are avoided, ensuring that the comprehensive evaluation results are more consistent with the actual performance of the video and improving the credibility of the results.
[0145] Based on the above embodiments, in an exemplary embodiment, user behavior data includes behavior distribution characteristics, playback rhythm characteristics, and interaction depth characteristics; based on this, the process of determining user behavior data for the video to be processed in S303 is further refined. Optionally, such as Figure 7 As shown, it includes the following steps:
[0146] S701, for each first data point in the user log data, determine the frequency of the behavior corresponding to the first data point.
[0147] The first data refers to the core behavioral raw data directly collected from user log data. Different first data include at least two of the following: click behavior data, interaction behavior data, and transaction behavior data. The behavior frequency corresponding to the first data is the quantified number of times the first data is collected according to a specified time or behavior dimension, reflecting the frequency and intensity of a certain type of user behavior.
[0148] Optionally, the first data that meets the requirements can be precisely selected from user log data, and behavior frequency statistics can be completed at a unified time granularity and user granularity to ensure the consistency and comparability of frequency data. In addition, the statistical time granularity and user granularity of all first data must be consistent, and abnormal data in the logs should be removed to avoid interfering with the frequency statistics results.
[0149] For example, for click behavior data, raw log records such as content clicks, page clicks, button clicks, and search clicks can be obtained, and duplicate clicks and invalid clicks within a short period of time can be removed. For interaction behavior data, raw log records such as comments, likes, favorites, reposts, follows, and bullet screen comments can be obtained, and meaningless interactions (such as blank comments and pure symbol likes) and malicious interactions (such as spamming comments) can be removed. For transaction behavior data, raw log records such as order placement, payment, refund, repeat purchase, and product click conversion can be obtained, and test transactions and invalid transactions with a 100% refund rate can be removed, retaining the actual completed transaction behaviors.
[0150] Furthermore, the first data can be directly counted at a fixed time granularity and a single user dimension to obtain the basic behavior frequency; only the number of occurrences is counted, without considering the intensity or scale of the behavior, such as only the number of orders placed in a transaction, without considering the amount; the frequency is simply deduplicated, such as if the same user triggers the same behavior multiple times on the same day, only once is counted, and finally the behavior frequency corresponding to the first data is obtained.
[0151] S702, determine the information entropy based on the frequency of behavior corresponding to different first data, and determine the behavior distribution characteristics based on the information entropy.
[0152] Among them, the behavior distribution feature is a characteristic that represents the diversity and concentration of user behavior. The higher the entropy value, the more uniform the behavior distribution and the stronger the diversity.
[0153] Optionally, information entropy can be calculated based on the frequency of behaviors corresponding to different primary data. Indicators derived from this information entropy can then be used to quantify the diversity, concentration, and evenness of user behavior distribution, forming behavioral distribution characteristics that reflect user behavior preferences. This requires first normalizing the frequency of behaviors with different dimensions to the same interval, then calculating the information entropy, and finally deriving multi-dimensional distribution characteristics from the entropy value to avoid the limitations of a single entropy value.
[0154] In one optional embodiment, an information entropy calculation model can be invoked, and the frequency of behaviors corresponding to different first data can be input into the information entropy calculation model to obtain the information entropy output by the information entropy calculation model. Furthermore, the information entropy can be directly and intentionally used as the sole characteristic of the behavior distribution.
[0155] In another optional embodiment, information entropy can be calculated based on the frequency of behaviors corresponding to different first data according to the basic formula, and derived indicators, namely concentration, evenness, and dominance, can be calculated based on the probability distribution. The values of the derived indicators are all normalized to between 0 and 1, and then concatenated with the information entropy to form the behavioral distribution features.
[0156] S703 extracts features from the playback behavior data in the user log data to obtain playback rhythm features.
[0157] Among them, playback behavior data refers to time-series data related to content playback from user log data, such as playback duration, number of pauses, fast forwards, and rewinds, playback completion rate, and continuous playback interval. Playback rhythm characteristics are quantitative features that characterize the rhythm of user content playback, reflecting the patterns of speed, continuity, and completeness of user playback.
[0158] Optionally, quantitative features that characterize the user's content playback rhythm can be extracted from the playback behavior data in user logs, covering four dimensions: playback continuity, completeness, speed, and interactivity, reflecting the user's acceptance of the content and viewing habits. It is necessary to align the playback behavior data with the content timeline, design features for the core dimensions of playback rhythm, avoid redundant features, and standardize all features.
[0159] In one optional embodiment, the timestamp of the playback behavior in the playback behavior data is aligned with the original duration of the video to be processed, and unified as a relative playback time, such as playing to 50% of the video to be processed. Abnormal data with a playback duration of less than 1 second and a playback completion rate of more than 100% are removed, and the correct playback indicators are padded with zeros. For records of interrupted playback, the interruption position and duration are marked.
[0160] Furthermore, from the preprocessed playback behavior data described above, a predetermined number of indicators are extracted and their values are directly statistically analyzed. These indicators include, but are not limited to, playback completion rate, average playback speed, percentage of longest single playback duration, and number of pauses. All indicators are normalized to the 0-1 range and then concatenated to form the basic playback rhythm characteristics.
[0161] S704. Extract features from at least one of the interaction behavior data, transaction behavior data, and feedback data in the user log data to obtain interaction depth features.
[0162] Interaction depth features are comprehensive characteristics that characterize the depth of user participation, focusing on the quality and depth of behavior. Feedback data represents users' subjective feelings about the videos they process, such as ratings, comments, complaints, likes, and favorites.
[0163] Optionally, at least one type of data can be selected from interactive behavior data, transaction behavior data, and feedback data to extract interaction depth features that are different from the basic frequency, representing the user's depth of participation and value contribution to the video being processed.
[0164] Filter target data (at least one of the following: interaction behavior data, transaction behavior data, and feedback data) from user log data, and remove invalid, malicious, and test data; perform basic labeling on the data, such as interaction type, transaction amount tiers, and preliminary judgment of feedback sentiment; perform preliminary normalization on indicators of different dimensions to avoid subsequent calculation biases.
[0165] In one optional implementation, a single data type (such as interaction behavior data only) can be selected, and basic frequency and 2-3 quality indicators can be extracted, weighted, summed, and then normalized. Wherein, if the single data type is interaction behavior data, the quality indicators include, but are not limited to, the number of comments, the average comment length, the like conversion rate, and the favorite conversion rate. If the single data type is transaction behavior data, the quality indicators include, but are not limited to, the number of orders, the average order value, and the repurchase rate. If the single data type is feedback data, the quality indicators include, but are not limited to, the average rating, the number of valid feedback comments, and the percentage of positive feedback. Finally, the weighted result is directly used as the interaction depth feature.
[0166] Another optional implementation involves selecting two or more types of data (such as interaction behavior data + transaction behavior data or interaction behavior data + feedback data), extracting 3-4 quality indicators for each type of data, and removing the basic frequency; using the entropy weight method to calculate the objective weight of each quality indicator, performing weighted fusion on each type of data to obtain the deep sub-features of a single type of data; and then weighted fusion again on the deep sub-features of all single-type data to splice them into interaction deep features.
[0167] Another optional implementation can integrate three types of data: interactive behavior data, transaction behavior data, and feedback data, and perform refined quantification on each type of data. For example, for interactive behavior data, BERT (Bidirectional Encoder Representations from Transformers) can be used to perform semantic sentiment analysis of comments, interaction time-series features, and importance scoring of interactive content; for transaction behavior data, customer lifetime value, transaction frequency time-series, and category repurchase rate can be calculated; for feedback data, deep semantic analysis, feedback sentiment intensity, and demand mining matching degree can be performed. Eight to ten deep indicators are extracted from each type of data to form a high-dimensional indicator matrix; redundant indicators are eliminated through feature screening, and the normalized indicators are used as the final deep interaction features.
[0168] The above embodiments quantify the distribution patterns of behavior through information entropy and extract playback rhythm and interaction depth features in a targeted manner, transforming scattered user log data into structured, computable, and comparable features. The three types of core features extracted respectively cover the user's behavioral preference patterns, content consumption habits, and participation value depth, comprehensively representing the user's real behavioral characteristics from different dimensions, making the features more consistent with the user's actual level of participation.
[0169] Based on the above embodiments, in an exemplary embodiment, the process of determining the correlation information between different features in S103 is further refined. Optionally, such as Figure 8 As shown, it includes the following steps:
[0170] S801 concatenates video features, behavioral features, and evaluation features to obtain a feature sequence.
[0171] Among them, the feature sequence is a high-dimensional feature vector formed by directly concatenating the three types of features—video features, behavioral features, and evaluation features—after aligning them according to a specified dimension.
[0172] Optionally, video features, behavioral features, and evaluation features can be cleaned, standardized, and dimension-unified to eliminate differences in units and data noise, ensuring that the three types of features can be directly concatenated. All features are mapped to the same numerical range, outliers are removed, and dimension unification is achieved through dimensionality upgrades or downsizing, preserving the original feature information to the greatest extent possible.
[0173] In one optional embodiment, the original dimensions of video features, behavioral features, and evaluation features can be calculated, and a target alignment dimension can be determined. For example, the target alignment dimension can be the maximum dimension of the video features, behavioral features, and evaluation features. For features whose original dimensions are smaller than the target alignment dimension, zero-padding is used to increase the dimensionality. For features whose original dimensions are larger than the target alignment dimension, the first N dimensions are truncated, where N is the target alignment dimension. Finally, the aligned video features, behavioral features, and evaluation features are concatenated into a high-dimensional feature sequence in a specified order. The concatenation process ensures that the feature order is fixed and the dimensions are continuous, supporting subsequent linear projection and attention calculations.
[0174] S802 employs three independent linear projection layers to linearly map the feature sequences, thereby obtaining the query vector, key vector, and value vector of the multi-head attention module.
[0175] The three independent linear projection layers are fully connected linear layers used to map the feature sequence into query vectors (Q), key vectors (K), and value vectors (V) of uniform dimension. The multi-head attention module is the core module containing at least 3-4 independent attention heads, each capturing the correlation information between two different features. The query vector, key vector, and value vector are the three core vectors for multi-head attention computation, obtained by mapping the feature sequence through the three independent linear projection layers, with a dimension equal to the number of attention heads multiplied by the feature hidden dimension.
[0176] Optionally, three structurally independent linear projection layers with non-shared parameters are pre-constructed to linearly map the feature sequence into a query vector, a key vector, and a value vector, respectively. This ensures that the dimensions of the query vector, key vector, and value vector are uniform and adaptable to the computational requirements of the multi-head attention module. The linear projection layers only contain weight matrices and bias terms, without non-linear activation functions, to avoid disrupting the linear correlation of features. The parameters of the three linear projection layers are completely independent, and the output vector dimension is designed according to "number of attention heads × hidden dimension".
[0177] The layer structure of the linear projection layer is a fully connected linear layer, which can be expressed by the formula:
[0178]
[0179] in, Given the input feature sequence, This is the projection weight matrix; For bias terms; This is the output vector.
[0180] The weight matrices of the three linear projection layers corresponding to the query vector, key vector, and value vector. , , With bias term , , Completely independent, each feature map is updated separately during training to avoid interference between them. Let the number of attention heads in the multi-head attention module be... ( ≥3), hidden dimensions of a single attention head Then the dimension of the output vector of the linear projection layer is .
[0181] Construct three trainable linear projection layers with randomly initialized weights and biases, which are updated along with the model during training; set a uniform hidden dimension for all attention heads. Projection layer according to Dimensional output: The projected key and value vectors are scaled to avoid attention score saturation due to excessively large hidden dimensions.
[0182] After obtaining the query vector, key vector, and value vector, a multi-head attention module with at least three attention heads is constructed. Each attention head is assigned a dedicated task to capture two types of feature correlations, which are bound one-to-one with no overlap. The correlation information between features is accurately quantified through attention scores to achieve targeted mining of feature correlations.
[0183] Since the input consists of three types of features—video features, behavioral features, and evaluation features—it can form three feature pairs: video feature-behavioral feature, video feature-evaluation feature, and behavioral feature-evaluation feature. Therefore, the number of attention heads is greater than or equal to 3. Each attention head uniquely corresponds to a feature pair, and the task allocation is fixed as follows: Attention head 1: captures the correlation information between "video features and behavioral features"; Attention head 2: captures the correlation information between "video features and evaluation features"; Attention head 3: captures the correlation information between "behavioral features and evaluation features". If the number of attention heads is greater than 3, additional attention heads can be added to capture the correlation between the three types of features—"video features, behavioral features, and evaluation features"—to improve the feature fusion depth.
[0184] S803 captures correlation information based on the query vector, key vector, and value vector to obtain correlation information between different features.
[0185] Among them, the correlation information between different features is the correlation strength between two types of features quantified by attention scores by each attention head. The higher the correlation value, the stronger the mutual influence and correlation between features.
[0186] Optionally, each attention head is assigned its own value based on the query vector, key vector, and value vector. , and Each attention head is based on its own , and Attention scores, i.e., the correlation information between features, are calculated by scaling the dot product attention.
[0187] The above embodiments bind feature pairs to tasks one-to-one with at least three attention heads. Each attention head specifically captures the correlation between a set of features, avoiding interference between the correlation information of different feature pairs, greatly improving the accuracy and targeting of the correlation between features, and effectively uncovering the potential correlation patterns of different feature pairs.
[0188] Based on the above embodiments, in an exemplary embodiment, the process of determining the fusion features in S104 is further refined. Optionally, such as Figure 9 As shown, it includes the following steps:
[0189] S901, based on the correlation information between different features, update the video features, behavioral features, and evaluation features to obtain new video features, new behavioral features, and new evaluation features that are integrated with the corresponding correlation information.
[0190] Optionally, the correlation information between different features can be standardized, dimension matched, and matrix reconstructed to eliminate data noise, unify the units, and match the original feature dimensions, ensuring that the correlation information can be directly combined with the original features to complete feature updates.
[0191] Furthermore, for each type of feature—video features, behavioral features, and evaluation features—the correlation information between it and the other two types of features is fused in a targeted manner. Corresponding methods are then used to update the features, generating new video, behavioral, and evaluation features, thereby improving the representational power of individual features. Each type of feature only fuses correlation information directly related to it; for example, video features only fuse the correlation between video features and behavioral features, and video features and evaluation features, without introducing irrelevant information. The dimensions of the new features are completely consistent with the original features, ensuring the feasibility of subsequent weighted fusion. The correlation information is integrated into the original features as an enhancing factor, rather than replacing them.
[0192] S902, based on the fusion weights corresponding to the new video features, the new behavioral features, and the new evaluation features, the new video features, the new behavioral features, and the new evaluation features are weighted and fused to obtain the fused features.
[0193] Each fusion weight is used to characterize the degree of contribution of the corresponding feature to the fused feature.
[0194] In one optional embodiment, the fusion weights corresponding to different new features are first calculated. For example, the fusion weights of the three types of new features can be calculated using a corresponding method based on the correlation information between features and the information richness of the new features. The weight directly represents the degree of contribution of each new feature to the final fused feature, ensuring that the weight allocation is objective, personalized, and aligned with business needs, avoiding the bias of subjective human assignment. Furthermore, the fusion weight is positively correlated with the total correlation of the feature (the sum of the correlation between the feature and the other two types of features) and positively correlated with the information entropy (information richness) of the new feature.
[0195] In another optional embodiment, the fusion weights corresponding to new video features, new behavioral features, and new evaluation features can be determined based on the gating fusion unit.
[0196] Optionally, after obtaining the fusion weights corresponding to the new video features, the new behavioral features, and the new evaluation features, each new feature can be multiplied element-wise with its corresponding fusion weight to achieve feature weighting. The weighted new features are then summed to obtain the fused features.
[0197] The above embodiments integrate the correlation information between features into each type of original feature to generate new features. The new features retain the original core information and integrate the correlation information of other features, realizing the information complementarity of the three types of features: video features, behavioral features, and evaluation features. This qualitatively improves the representation ability of single features and lays a high-quality feature foundation for subsequent fusion.
[0198] Based on the above embodiments, in an exemplary embodiment, the process of determining the quality score of the video to be processed in S105 is further refined. Optionally, such as Figure 10 As shown, it includes the following steps:
[0199] S1001, analyze the fusion features to obtain the dimensional scores of the video to be processed under different evaluation dimensions.
[0200] In one optional embodiment, the evaluation dimensions are core dimensions characterizing video quality. For example, evaluation dimensions include, but are not limited to, content quality, user experience, technical quality, commercial value, and social value. Dimension scores reflect the specific performance of the video under processing in different dimensions. Furthermore, multiple indicators can be subdivided under different evaluation dimensions. For example, for the content quality dimension, indicators include, but are not limited to, narrative completeness, character development, and degree of innovation; for the user experience dimension, indicators include, but are not limited to, completion rate, interactive engagement, and emotional feedback; for the technical quality dimension, indicators include, but are not limited to, image clarity, audio fidelity, and editing smoothness; for the commercial value dimension, indicators include, but are not limited to, advertising monetization capability, paid conversion rate, and long-tail value; for the social value dimension, indicators include, but are not limited to, information accuracy, value orientation, and cultural dissemination power. Based on this, the fusion features can be analyzed to obtain the indicator values for each indicator under different evaluation dimensions, and based on the indicator values, the dimension scores for each evaluation dimension can be calculated.
[0201] For example, the fusion features are decomposed and quantified in multiple dimensions. For different preset evaluation dimensions, the fusion features are mapped to quantified scores using corresponding methods to obtain a dimension score for each dimension, accurately representing the video's specific performance in each dimension. Optionally, for any evaluation dimension, all fusion features related to the corresponding evaluation dimension can be extracted, and the entropy weight method can be used to calculate the weight of each fusion feature. This weight represents the degree of contribution of the feature to the evaluation dimension. The weighted sum of all indicator-related features under the evaluation dimension is then performed to obtain the dimension score for that evaluation dimension.
[0202] S1002, determine the dynamic weights of different evaluation dimensions based on the video type of the video to be processed.
[0203] The dynamic weight of each evaluation dimension is used to characterize the degree of contribution of that evaluation dimension to the video type.
[0204] Optionally, an initial weight range can be pre-assigned to each evaluation dimension based on its contribution to video quality. For example, the initial weight range for content quality could be set to 15%-30%, user experience to 20%-35%, technical quality to 15%-25%, commercial value to 15%-30%, and social value to 5%-25%. Furthermore, the specific video type to be processed is first defined. Then, based on the core requirements of the video type and the initial weight range of each evaluation dimension, the dynamic weight of each dimension is dynamically adjusted to ensure that the weight accurately reflects the contribution of different dimensions to the quality of that type of video. For example, if the video segment to be processed is a documentary, the dynamic weight of the social value dimension can be automatically increased to 25%; if the video segment to be processed is a film, the dynamic weight of the content quality dimension can be increased to 30%.
[0205] S1003, based on the dynamic weights of different evaluation dimensions, performs weighted fusion of the dimensional scores of different dimensions to obtain the quality score of the video to be processed.
[0206] Optionally, the dimensional scores of different evaluation dimensions are multiplied by their corresponding dynamic weights, and a preliminary quality score is obtained by weighted summation. After optimization and calibration, the final video quality score (0-10 points) is obtained, realizing the quantitative representation of the overall video quality.
[0207] In the above embodiments, weights for different evaluation dimensions are dynamically allocated according to the specific type of video to be processed, so that the weights accurately match the core needs of the video type. The dynamic weights determine the contribution of each dimension's score, and the weighted summation achieves the complementary integration of multi-dimensional scores. Subsequent optimization and calibration eliminate scoring biases, ensuring that the final quality score can objectively and accurately reflect the overall quality of the video.
[0208] Based on the above embodiments, in an exemplary embodiment, such as Figure 11 As shown, this video quality determination method is applied to a recommendation scenario for videos to be processed. In this scenario, the server stores many videos, each of which can serve as a candidate video for recommendation. The server pre-determines the quality score for each candidate video. In this scenario, the process of determining the quality score for each candidate video includes the following steps:
[0209] S1101, Obtain candidate videos, user log data for candidate videos, and evaluation data of candidate videos on different platforms.
[0210] S1102, Perform semantic analysis on the candidate videos to obtain video description data for the candidate videos.
[0211] S1103, extract features from user log data to obtain user behavior data for candidate videos.
[0212] S1104, the evaluation data of candidate videos on different platforms are fused to obtain comprehensive evaluation data.
[0213] The video description data includes video type, sentiment data, and narrative rhythm data.
[0214] Optionally, the candidate video is segmented to obtain at least two video segments; for each video segment, features are extracted to obtain multimodal features of the video segment; wherein, the multimodal features include camera motion features, speech content features and semantic features.
[0215] Optionally, image data, audio data, and text data of the video segment are acquired; feature extraction is performed on the image data to obtain the camera motion features of the video segment; the audio data is converted into a Mel spectrogram, and feature extraction is performed on the Mel spectrogram to obtain the speech content features of the video segment; word segmentation is performed on the text data, and feature extraction is performed on the segmented text data to obtain the semantic features of the video segment.
[0216] User behavior data includes behavior distribution characteristics, playback rhythm characteristics, and interaction depth characteristics.
[0217] Optionally, for each first data point in the user log data, the frequency of the corresponding behavior is determined based on the first data point; wherein, different first data points include at least two of click behavior data, interaction behavior data, and transaction behavior data; the information entropy is determined based on the behavior frequency corresponding to different first data points, and the behavior distribution characteristics are determined based on the information entropy; the playback behavior data in the user log data is subjected to feature extraction to obtain playback rhythm features; and at least one of the interaction behavior data, transaction behavior data, and feedback data in the user log data is subjected to feature extraction to obtain interaction depth features.
[0218] Optionally, based on the platform attribute information of different platforms, the deviation factor and platform weight of different platforms are determined; for any platform, the evaluation data of candidate videos under the platform are corrected according to the platform's deviation factor to obtain the platform's standardized evaluation data; based on the platform weight of different platforms, the standardized evaluation data of different platforms are weighted and fused to obtain comprehensive evaluation data.
[0219] S1105, encode the video description data, user behavior data and comprehensive evaluation data respectively to obtain the video features corresponding to the video description data, the behavioral features corresponding to the user behavior data and the evaluation features corresponding to the comprehensive evaluation data.
[0220] S1106, video features, behavioral features and evaluation features are concatenated to obtain a feature sequence.
[0221] S1107 employs three independent linear projection layers to linearly map the feature sequences, thereby obtaining the query vector, key vector, and value vector of the multi-head attention module.
[0222] S1108: Based on the query vector, key vector, and value vector, capture the correlation information to obtain the correlation information between different features.
[0223] The multi-head attention module includes at least three attention heads, and each attention head is used to capture the correlation information between two different features.
[0224] S1109, based on the correlation information between different features, update the video features, behavioral features, and evaluation features to obtain new video features, new behavioral features, and new evaluation features that are integrated with the corresponding correlation information.
[0225] S1110, Based on the fusion weights corresponding to the new video features, the new behavioral features, and the new evaluation features, the new video features, the new behavioral features, and the new evaluation features are weighted and fused to obtain the fused features.
[0226] Each fusion weight is used to characterize the degree of contribution of the corresponding feature to the fused feature.
[0227] S1111, Analyze the fusion features to obtain the quality score of the candidate video.
[0228] After the server determines the quality score of all candidate videos, it stores the candidate videos and their corresponding quality scores. Based on this:
[0229] S1112, the terminal sends a video recommendation request to the server.
[0230] If the quality score is greater than or equal to the threshold, the recommendation priority of the candidate video segments is determined based on the quality score.
[0231] S1113, In response to the video recommendation request, the server selects candidate videos with a quality score greater than or equal to the score threshold from among the candidate videos as the target recommended videos.
[0232] S1114, the server sends the target recommended video to the terminal.
[0233] S1115, the terminal displays the target recommended video.
[0234] The specific implementation methods of S1101-S1115 are the same as those in the above embodiments, and will not be repeated here.
[0235] Based on the same inventive concept, some embodiments also provide a video quality determination apparatus for implementing the video quality determination method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more video quality determination apparatus embodiments provided below can be found in the limitations of the video quality determination method described above, and will not be repeated here.
[0236] In one exemplary embodiment, such as Figure 12 As shown, a video quality determination device is provided, including: a data acquisition module 1201, a data encoding module 1202, an information determination module 1203, a feature fusion module 1204, and a quality scoring module 1205. Wherein:
[0237] The data acquisition module 1201 is used to acquire multi-source data of the video to be processed; the multi-source data includes video description data, user behavior data and comprehensive evaluation data; the comprehensive evaluation data is determined based on the evaluation data of the video to be processed on different platforms.
[0238] The data encoding module 1202 is used to encode multi-source data to obtain video features corresponding to video description data, behavioral features corresponding to user behavior data, and evaluation features corresponding to comprehensive evaluation data.
[0239] The information determination module 1203 is used to perform correlation analysis on video features, behavioral features and evaluation features to obtain correlation information between different features.
[0240] The feature fusion module 1204 is used to fuse video features, behavioral features and evaluation features based on the correlation information between different features to obtain fused features.
[0241] The quality scoring module 1205 is used to analyze the fusion features and obtain the quality score of the video to be processed.
[0242] In one exemplary embodiment, the information determination module 1203 is used to:
[0243] Video features, behavioral features, and evaluation features are concatenated to obtain a feature sequence. Three independent linear projection layers are used to linearly map the feature sequence to obtain the query vector, key vector, and value vector of the multi-head attention module. Based on the query vector, key vector, and value vector, correlation information is captured to obtain the correlation information between different features. The multi-head attention module includes at least three attention heads, and each attention head is used to capture the correlation information between two different features.
[0244] In one exemplary embodiment, the feature fusion module 1204 is used to:
[0245] Based on the correlation information between different features, the video features, behavioral features, and evaluation features are updated to obtain new video features, new behavioral features, and new evaluation features that are fused with corresponding correlation information. Based on the fusion weights corresponding to the new video features, the new behavioral features, and the new evaluation features, the new video features, new behavioral features, and new evaluation features are weighted and fused to obtain fused features. Each fusion weight is used to characterize the degree of contribution of the corresponding feature to the fused features.
[0246] In one exemplary embodiment, the data acquisition module 1201 includes:
[0247] The first acquisition unit is used to acquire the video to be processed, user log data for the video to be processed, and evaluation data of the video to be processed on different platforms.
[0248] The second acquisition unit is used to perform semantic analysis on the video to be processed and obtain video description data of the video to be processed.
[0249] The third acquisition unit is used to extract features from user log data to obtain user behavior data for the video to be processed.
[0250] The fourth acquisition unit is used to fuse the evaluation data of the video to be processed on different platforms to obtain comprehensive evaluation data.
[0251] In one exemplary embodiment, the video description data includes video type, emotion data, and narrative rhythm data; the second acquisition unit further includes:
[0252] The segmentation subunit is used to segment the video to be processed, resulting in at least two video segments.
[0253] The extraction sub-unit is used to extract features from each video segment to obtain the multimodal features of the video segment; among which, the multimodal features include camera motion features, speech content features and semantic features.
[0254] The sub-units are used to determine the video type and sentiment data of the video to be processed based on the multimodal features of different video segments; the narrative structure of the video to be processed is extracted, and the narrative rhythm data of the video to be processed is determined based on the narrative structure.
[0255] In one exemplary embodiment, the extraction subunit is used for:
[0256] Acquire image, audio, and text data from video segments;
[0257] Feature extraction is performed on the image data to obtain the camera motion features of the video segment;
[0258] The audio data is converted into a Mel spectrogram, and features are extracted from the Mel spectrogram to obtain the speech content features of the video segment.
[0259] The text data is segmented into words, and features are extracted from the segmented text data to obtain the semantic features of the video segment.
[0260] In one exemplary embodiment, the third acquisition unit is used for:
[0261] For each first data point in the user log data, determine the frequency of the corresponding behavior based on the first data point; wherein, different first data points include at least two of click behavior data, interaction behavior data, and transaction behavior data;
[0262] Based on the frequency of behavior corresponding to different first data, determine the information entropy, and determine the behavior distribution characteristics based on the information entropy;
[0263] Feature extraction is performed on playback behavior data in user log data to obtain playback rhythm features;
[0264] Feature extraction is performed on at least one of the interaction behavior data, transaction behavior data, and feedback data in the user log data to obtain interaction depth features.
[0265] In one exemplary embodiment, the fourth acquisition unit is used to:
[0266] Based on the platform attribute information of different platforms, determine the deviation factor and platform weight of different platforms;
[0267] For any given platform, the evaluation data of the video to be processed on the platform is corrected according to the platform's deviation factor to obtain the platform's standardized evaluation data.
[0268] Based on the platform weights of different platforms, the standardized evaluation data from different platforms are weighted and merged to obtain comprehensive evaluation data.
[0269] In one exemplary embodiment, the quality scoring module 1205 is used to:
[0270] Analyzing the fusion features yields dimensional scores for the video under different evaluation dimensions; and...
[0271] Based on the video type of the video to be processed, determine the dynamic weights of different evaluation dimensions;
[0272] Based on the dynamic weights of different evaluation dimensions, the dimensional scores of different dimensions are weighted and fused to obtain the quality score of the video to be processed.
[0273] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 13 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores PCB product-related data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements a video quality determination method.
[0274] Those skilled in the art will understand that Figure 13 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0275] In one alternative embodiment, Figure 13 The computer device shown can be the aforementioned server.
[0276] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0277] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0278] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0279] It should be noted that the user information (including but not limited to user behavior data) and data (including but not limited to data used for analysis, stored data, and displayed data) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0280] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0281] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0282] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for determining video quality, characterized in that, The method includes: Acquire multi-source data of the video to be processed; wherein, the multi-source data includes video description data, user behavior data, and comprehensive evaluation data; the comprehensive evaluation data is determined based on the evaluation data of the video to be processed on different platforms; The multi-source data is encoded to obtain video features corresponding to the video description data, behavioral features corresponding to the user behavior data, and evaluation features corresponding to the comprehensive evaluation data. A correlation analysis is performed on the video features, the behavioral features, and the evaluation features to obtain correlation information between different features; Based on the correlation information between different features, the video features, the behavioral features, and the evaluation features are fused to obtain fused features; The fusion features are analyzed to obtain a quality score for the video to be processed.
2. The method according to claim 1, characterized in that, The step of performing correlation analysis on the video features, the behavioral features, and the evaluation features to obtain correlation information between different features includes: The video features, behavioral features, and evaluation features are concatenated to obtain a feature sequence; Three independent linear projection layers are used to linearly map the feature sequence to obtain the query vector, key vector, and value vector of the multi-head attention module. Based on the query vector, the key vector, and the value vector, correlation information is captured to obtain correlation information between different features; wherein, the multi-head attention module includes at least three attention heads, and each attention head is used to capture correlation information between two different features.
3. The method according to claim 1, characterized in that, The step of fusing the video features, the behavioral features, and the evaluation features based on the correlation information between different features to obtain fused features includes: Based on the correlation information between different features, the video features, the behavioral features, and the evaluation features are updated to obtain new video features, new behavioral features, and new evaluation features that are fused with corresponding correlation information; Based on the fusion weights corresponding to the new video features, the new behavioral features, and the new evaluation features, a weighted fusion is performed on the new video features, the new behavioral features, and the new evaluation features to obtain fused features; wherein, each fusion weight is used to characterize the degree of contribution of the corresponding feature to the fused features.
4. The method according to any one of claims 1-3, characterized in that, The acquisition of multi-source data for the video to be processed includes: Acquire the video to be processed, user log data for the video to be processed, and evaluation data of the video to be processed on different platforms; Perform semantic analysis on the video to be processed to obtain video description data of the video to be processed; Feature extraction is performed on the user log data to obtain user behavior data for the video to be processed; The evaluation data of the video to be processed on different platforms are fused to obtain comprehensive evaluation data.
5. The method according to claim 4, characterized in that, The video description data includes video type, sentiment data, and narrative rhythm data; the semantic analysis of the video to be processed to obtain the video description data includes: The video to be processed is segmented to obtain at least two video segments; For each video segment, feature extraction is performed on the video segment to obtain the multimodal features of the video segment; wherein, the multimodal features include camera motion features, speech content features, and semantic features; Based on the multimodal features of different video segments, the video type and sentiment data of the video to be processed are determined; Extract the narrative structure of the video to be processed, and determine the narrative rhythm data of the video to be processed based on the narrative structure.
6. The method according to claim 5, characterized in that, The step of extracting features from the video segment to obtain the multimodal features of the video segment includes: Acquire the image data, audio data, and text data of the video segment; Feature extraction is performed on the image data to obtain the camera motion features of the video segment; The audio data is converted into a Mel spectrogram, and features are extracted from the Mel spectrogram to obtain the speech content features of the video segment. The text data is segmented into words, and features are extracted from the segmented text data to obtain the semantic features of the video segment.
7. The method according to claim 4, characterized in that, The process of fusing evaluation data of the video to be processed on different platforms to obtain comprehensive evaluation data includes: Based on the platform attribute information of different platforms, determine the deviation factor and platform weight of different platforms; For any platform, the evaluation data of the video to be processed under the platform is corrected according to the deviation factor of the platform to obtain the standardized evaluation data of the platform. Based on the platform weights of different platforms, the standardized evaluation data from different platforms are weighted and merged to obtain comprehensive evaluation data.
8. The method according to claim 4, characterized in that, The user behavior data includes behavior distribution features, playback rhythm features, and interaction depth features; the feature extraction of the user log sequence data to obtain the user behavior data of the video to be processed includes: For each first data point in the user log data, the frequency of the behavior corresponding to the first data point is determined based on the first data point; wherein, different first data points include at least two of click behavior data, interaction behavior data, and transaction behavior data; Based on the frequency of behavior corresponding to different first data, the information entropy is determined, and the behavior distribution characteristics are determined based on the information entropy; Feature extraction is performed on the playback behavior data in the user log data to obtain playback rhythm features; Feature extraction is performed on at least one of the interaction behavior data, transaction behavior data, and feedback data in the user log data to obtain interaction depth features.
9. The method according to any one of claims 1-3, characterized in that, The analysis of the fused features to obtain a quality score for the video to be processed includes: The fusion features are analyzed to obtain the dimensional scores of the video to be processed under different evaluation dimensions; and, Based on the video type of the video to be processed, determine the dynamic weights of different evaluation dimensions; Based on the dynamic weights of different evaluation dimensions, the dimensional scores of different dimensions are weighted and fused to obtain the quality score of the video to be processed.
10. A server, characterized in that, include: At least one processor, and configured as follows: Acquire multi-source data of the video to be processed; wherein, the multi-source data includes video description data, user behavior data, and comprehensive evaluation data; the comprehensive evaluation data is determined based on the evaluation data of the video to be processed on different platforms; The multi-source data is encoded to obtain video features corresponding to the video description data, behavioral features corresponding to the user behavior data, and evaluation features corresponding to the comprehensive evaluation data. A correlation analysis is performed on the video features, the behavioral features, and the evaluation features to obtain correlation information between different features; Based on the correlation information between different features, the video features, the behavioral features, and the evaluation features are fused to obtain fused features; The fusion features are analyzed to obtain a quality score for the video to be processed.