Sample screening method and device, equipment and storage medium
Through joint modeling of multi-dimensional features, dynamically adjusting the bucket threshold, the problem of inaccurate sample division in the graphics and dynamic resource recommendation systems in the existing technology is solved, and fine-grained reading time estimates and sample type recognition is realized, which improves the training data quality and model generalization performance of the recommendation system.
Patent Information
- Application Number
- CN202510479738.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-01
AI Technical Summary
In the recommendation system of graphics and dynamic resources, the method of dividing positive and negative samples by a single feature leads to rigidity of the bucket threshold, which cannot accurately reflect the user's reading intention, affecting the scientificity of the training samples and the generalization performance of the model.
The multi-dimensional feature joint modeling method is adopted to build a deep learning model through the number of words and pictures of text, dynamically adjust the bucket function and threshold boundaries, and combine text and picture sensitivity to achieve fine-grained reading time estimates and sample type division.
It improves the scientificity of the training samples and the generalization ability of the recommendation system, and improves the accuracy of the modeling of the finished target and the accuracy of user behavior prediction.
Smart Images

Figure CN120408303A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technologies, and particularly to the fields of deep learning and model training technologies. Background Art
[0002] The aggregation and recall modules in a recommendation system often design and optimize strategies for specific scenarios, while the ranking stage more relies on technologies with wide applicability such as machine learning and deep learning to better integrate the characteristics of different scenarios and improve the accuracy of ranking. Especially in the rough ranking link, as a key link in the recommendation funnel, it needs to efficiently and accurately score and rank a large amount of resources. This requires a large number of accurate positive and negative training samples to improve the accuracy of ranking. Therefore, the positive and negative division of training samples greatly affects the accuracy and generalization performance of the model. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device, and storage medium for sample screening.
[0004] According to one aspect of the present disclosure, there is provided a method for sample screening, including:
[0005] Determining an estimated playing duration of the target resource according to the number of text words and the number of pictures of the target resource;
[0006] Adjusting a reference threshold of a corresponding bucket of the target resource according to the text sensitivity and picture sensitivity of the target resource to obtain a division threshold;
[0007] Determining a sample type of the target resource according to the estimated playing duration and the division threshold; wherein the sample type includes positive samples and negative samples.
[0008] According to another aspect of the present disclosure, there is provided a sample screening apparatus, including:
[0009] An estimation module for determining an estimated playing duration of the target resource according to the number of text words and the number of pictures of the target resource;
[0010] A threshold determination module for adjusting a reference threshold of a corresponding bucket of the target resource according to the text sensitivity and picture sensitivity of the target resource to obtain a division threshold;
[0011] A screening module for determining a sample type of the target resource according to the estimated playing duration and the division threshold; wherein the sample type includes positive samples and negative samples.
[0012] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.
[0016] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0017] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements any method in the embodiments of the present disclosure when executed by a processor.
[0018] Fine-grained reading duration prediction can be achieved according to the present disclosure, ensuring the scientificity of training sample division.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0021] Figure 1 is a flowchart of a sample screening method provided according to an embodiment of the present disclosure;
[0022] Figure 2 is a schematic structural diagram of a sample screening device provided according to an embodiment of the present disclosure;
[0023] Figure 3 is a block diagram of an electronic device for implementing the sample screening method in the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0025] In the related art, the current rough ranking framework defines the completion target for picture - text and dynamic resources by dividing the appropriate number of words through a single - sample feature, and fitting a curve function after statistically calculating the corresponding reading duration. Finally, the corresponding estimated reading duration for different word - count buckets is obtained according to the fitting function, and this is used as the basis for dividing training positive and negative samples in turn.
[0026] The current rough ranking framework has a defect of single - feature in the modeling method for the completion target of picture - text and dynamic resources. It divides positive and negative sample features by artificially defining the word - count bucket threshold and relies on statistically calculating the historical reading duration to fit a static function to achieve the estimated duration mapping. This method does not integrate diverse features and context - dynamic associations, resulting in rigid bucket thresholds, cumulative estimation biases, and a semantic gap between the basis for dividing positive and negative training samples and the true reading intentions of users.
[0027] To at least partially solve one or more of the above - mentioned problems and other potential problems, embodiments of the present disclosure provide a sample screening method. Using the technical solutions of the embodiments of the present disclosure, an optimization strategy for the global rough - ranking picture - text and dynamic completion target based on multi - dimensional feature joint modeling can be implemented. By integrating the text - word - count feature and the number - of - pictures feature within the resource, a deep learning model is constructed to automatically adaptively divide the bucket function and the dynamic threshold boundary according to the length of the resource, and achieve fine - grained reading - duration estimation. While ensuring the scientific nature of training sample division, this method significantly improves the generalization ability of the completion - target modeling and can be applied to optimize the interactive behavior prediction system in scenarios such as dynamic - resource recommendation.
[0028] Figure 1 is a schematic flowchart of a sample screening method according to an embodiment of the present disclosure. As Figure 1 shown, the method at least includes the following steps:
[0029] S110. Determine the estimated playback duration of the target resource according to the text word count and the number of pictures of the target resource.
[0030] In the embodiments of the present disclosure, the target resource can be a picture - text resource or a dynamic resource. A picture - text resource (Text - Image Content) is a multi - modal content form with text and pictures as core elements, usually in a static or light - interaction form. The text, which is its core feature, provides detailed descriptions, and the pictures enhance the intuitive perception. It is common to have text paragraphs interspersed with pictures. Such resources usually include:
[0031] Social media: Weibo picture - text, Xiaohongshu notes, Instagram posts
[0032] Information platforms: news articles with pictures, Zhihu columns
[0033] E - commerce: product detail pages (text description + product pictures).
[0034] Dynamic resources are content forms centered around real-time updates, interactive feedback, or dynamic elements, usually containing time series or user behavior responses. Their core feature is that the content value decays over time (such as news hotspots). User behavior (likes, comments) affects content dissemination. Such resources typically include:
[0035] Short video platforms: the dynamic recommendation streams of Douyin and Kuaishou
[0036] Social dynamics: WeChat Moments, Facebook dynamics (including text, pictures, videos, likes, and comments)
[0037] Real-time information: stock market quotes, live sports event bullet screens.
[0038] The common point between graphic resources and dynamic resources is that both have text and usually pictures. The number of text words and the number of pictures have a decisive impact on the playback duration of the resource. For the number of text words, an increase in the number of text words can provide more complete information (such as tutorial content) and extend the user's reading duration. However, when the number of words exceeds the user's patience threshold (such as >2000 words on mobile devices), it may lead to dropping out midway. For example, when the number of words increases from 100 to 500, the increase in reading duration is significant; but when it increases from 2000 to 2500, the increase slows down or even becomes negative. For different application scenarios, the optimal word count range varies significantly. For example: social media: 300 - 800 words, in-depth analysis: 1500 - 3000 words.
[0039] For the pictures in the resources, pictures can quickly capture the user's attention, reduce the bounce rate, simplify the understanding of complex concepts through charts (such as flowcharts, infographics), and extend the stay time. High-quality pictures (such as humanistic photography) stimulate the user's emotional investment and enhance content stickiness. However, too many pictures (such as >10) lead to distraction of attention and will instead reduce the effective reading duration. There may be a non-linear relationship between pictures and text, and the picture-text ratio needs to be balanced. For example: for e-commerce product detail pages, 1 - 2 pictures (product pictures + scene pictures) are paired with every 200 words, and for tutorial content, 1 schematic diagram (such as code screenshots + effect diagrams) is paired with each step.
[0040] The influence of the number of text words and the number of pictures on the playback duration shows non-linear and scenario-dependent characteristics, and the limitations of existing methods need to be solved through multi-modal fusion and dynamic bucketing.
[0041] Using a pre-trained multi-modal regression model, taking the number of text words L and the number of pictures I of the target resource as input features, and based on the non-linear mapping relationship between the number of words, the number of pictures, and the playback duration in historical behavior data, the estimated playback duration T of the target resource is output pred .
[0042] S120. Adjust the reference threshold of the corresponding bucket of the target resource according to the text sensitivity and picture sensitivity of the target resource to obtain the division threshold.
[0043] In the embodiments of the present disclosure, the preset static bucket threshold is weighted and corrected according to the text sensitivity and picture sensitivity of the target resource to generate a dynamic division threshold.
[0044] The text sensitivity and picture sensitivity reflect the contribution intensity of text and pictures in the current resource to the duration.
[0045] S130. Determine the sample type of the target resource according to the estimated playback duration and the division threshold. Among them, the sample type includes positive samples and negative samples.
[0046] In the embodiments of the present disclosure, the estimated playback duration T of the target resource pred is compared with the dynamic threshold θ dynamic as follows:
[0047] If T pred ≥θ dynamic , it is marked as a positive sample (the user may finish playing);
[0048] If T pred <θ dynamic , it is marked as a negative sample (the user may jump out midway).
[0049] According to the solution of the embodiments of the present disclosure, the semantic deviation of the static bucket is eliminated through the dynamic threshold, so that the division of positive and negative samples is more in line with the true intention of the user. Provide high-quality training data for subsequent sorting models (such as CTR prediction) and improve the prediction accuracy of the completion rate.
[0050] In a possible implementation manner, step S110 determines the estimated playback duration of the target resource according to the text word count and picture count of the target resource, and further includes the steps of:
[0051] Input the text word count and picture count of the target resource into a pre-trained playback duration prediction model to obtain the estimated playback duration.
[0052] Among them, the playback duration prediction model is a polynomial regression model, and the parameters of the playback duration prediction model are trained according to the text word count, the number of pictures, and the actual playback duration of multiple sample resources.
[0053] In the embodiments of the present disclosure, a playback duration prediction model can be trained using multiple sample resources in a historical resource library. The sample resources are resources that have been recommended to users or have been played by users, and the sample resources include the number of text words, the number of pictures, and the actual playback duration. The parameters of a playback duration prediction model can be determined based on the number of text words, the number of pictures, and the actual playback duration. Thus, for a new resource, the estimated playback duration can be calculated by inputting the number of text words and the number of pictures.
[0054] The playback duration prediction model can be a polynomial regression model. For example:
[0055]
[0056] where γ0~γ5 are parameters to be trained, L i is the number of text words of sample resource i, and I i is the number of pictures of sample resource i.
[0057] According to the solution of the embodiments of the present disclosure, a playback duration prediction model is trained through the historical data of sample resources, so that the playback duration of the target resource can be predicted based on the historical data through the playback duration prediction model.
[0058] In a possible implementation manner, step S120 adjusts the reference threshold of the corresponding bucket of the target resource according to the text sensitivity and picture sensitivity of the target resource to obtain a division threshold, and further includes the steps of:
[0059] S121. Determine the text bucket and picture bucket corresponding to the target resource according to the number of text words and the number of pictures included in the target resource.
[0060] In the embodiments of the present disclosure, the target resource is classified into a text bucket according to the number of text words. Multiple text buckets can be divided according to a preset word count interval, and each bucket corresponds to a word count interval (such as 0 - 500 words, 501 - 1000 words, >1000 words).
[0061] The target resource is classified into a picture bucket according to the number of pictures. Multiple picture buckets can be divided according to a preset picture quantity interval, and each bucket corresponds to a picture quantity interval (such as 0 pictures, 1 - 2 pictures, 3 - 5 pictures, >5 pictures).
[0062] Example: If the target resource contains 1200 words + 4 pictures, it falls into the >1000 - word text bucket and the 3 - 5 - picture bucket.
[0063] S122. Determine the first reference threshold of the text bucket and the second reference threshold of the picture bucket.
[0064] The first reference threshold (text) can be obtained through offline statistics based on historical data. For example, the average playback duration of each word count bucket is statistically calculated (such as the threshold for the bucket with more than 1000 words is 60 seconds).
[0065] The second reference threshold (picture) can also be obtained through offline statistics based on historical data. For example, the average playback duration of each picture bucket is statistically calculated (such as the threshold for the 3 - 5 picture bucket is 45 seconds).
[0066] S123. Determine the text sensitivity and picture sensitivity of the target resource.
[0067] Text sensitivity (ω L ): It is calculated through the partial derivative of the regression model. For example, to determine the impact of each 1 - unit increase in word count on the duration.
[0068] Picture sensitivity (ω I ): Similarly calculated
[0069] S124. Dynamically adjust the first reference threshold and the second reference threshold according to the text sensitivity and the picture sensitivity to obtain the division threshold.
[0070] In one example, the dynamic threshold adjustment formula can be: θ 动态 = θ 参考 ·(1 + ω)
[0071] If ω L = 0.1 (high text sensitivity), then the text threshold is increased by 10% (such as 60 seconds → 66 seconds).
[0072] If ω I = -0.05 (low picture sensitivity), then the picture threshold is decreased by 5% (such as 45 seconds → 42.75 seconds).
[0073] The final division threshold: Take the weighted average of the adjusted thresholds of the text and the picture (such as (66 + 42.75) / 2 = 54.38 seconds)
[0074] According to the solution of the embodiment of the present disclosure, when the sample is more sensitive to the text word count (ω L (x i ) > ω I (x i ))), the dynamically generated division threshold is more biased towards L * ; conversely, it is more dependent on I * .
[0075] In one possible implementation, step S123 of determining the text sensitivity and picture sensitivity of the target resource further includes the steps:[[]]
[0076] S123a. Calculate the partial derivative of the text word count of the target resource according to the playback duration prediction model to obtain the text sensitivity.
[0077] S123b. Calculate the partial derivative of the number of pictures of the target resource according to the playback duration prediction model to obtain the picture sensitivity.
[0078] In the embodiments of the present disclosure, in order to better explore the dependencies of the model on different features, it is necessary to calculate the partial derivatives of the number-of-pictures feature and the number-of-words feature:
[0079] For L i The partial derivative is:
[0080] For I i The partial derivative is:
[0081] Calculate the "sensitivity" or marginal impact of each feature on the playback duration prediction, that is, how much the prediction result will be affected if the text or the number of pictures changes slightly. ω L (x i ) represents how fast the predicted playback duration changes when the text word count changes, ω I (x i ) represents the impact on the predicted playback duration when the number of pictures changes.
[0082] For the regression model constituted by the above formula (1), the partial derivative calculation formula is:
[0083] ω L =γ1 + 2γ3L i +γ5I i
[0084] ω I =γ2 + 2γ4I i +γ5L i
[0085] Taking this regression model as an example, assume the following parameters for the online new resource:
[0086] Word count L i = 3000 (bucket 2000 - 4000 words, L * = 120s)
[0087] Number of pictures I i = 5 (bucket 5 - 6 pictures, I * = 90s)
[0088] The parameters of the regression model are γ1 = 0.02, γ2 = 0.01, γ3 = 0.0001, γ4 = 0.0002, γ5 = 0.005
[0089] Calculate the sensitivity:
[0090] ω L = 0.02 + 2 × 0.0001 × 3000 + 0.005 × 5 = 0.02 + 0.6 + 0.025 = 0.645
[0091] ω I = 0.01 + 2 × 0.0002 × 5 + 0.005 × 3000 = 0.01 + 0.002 + 15 = 15.012
[0092] In this example, ω L (x i ) < ω I (x i ), the resource is more sensitive to the number of pictures, and the threshold should be more biased towards or dependent on I * .
[0093] According to the solution of the embodiment of the present disclosure, the rigidity of static bucketing is avoided, and the differences in resource characteristics are adapted. The "one-size-fits-all" problem of traditional manual bucketing is avoided. For example, the threshold is relaxed for high-sensitivity resources (such as tutorial graphics), and the threshold is tightened for low-sensitivity resources (such as entertainment dynamics).
[0094] In a possible implementation manner, S124 dynamically adjusts the first reference threshold and the second reference threshold according to the text sensitivity and the picture sensitivity to obtain a partitioning threshold, further including the steps of:
[0095] S124-1. Map the text sensitivity and the picture sensitivity to the probability space through an exponential function to obtain a normalized text weight and a picture weight.
[0096] S124-2. Adjust the first reference threshold according to the text weight.
[0097] S124-3. Adjust the second reference threshold according to the picture weight.
[0098] S124-4. Obtain a partitioning threshold according to the adjusted first reference threshold and the adjusted second reference threshold.
[0099] In the embodiment of the present disclosure, the text weight and the picture weight can be determined according to the text sensitivity and the picture sensitivity. The corresponding reference threshold is adjusted based on the weight.
[0100] In one example, the calculation process of the partitioning threshold is as follows:
[0101] Bucket matching: Divide the new resource into the corresponding L and I buckets according to the number of words and the number of pictures (such as 2000 - 4000 words, 5 - 6 pictures).
[0102] Obtain the bucketing threshold: Read the L corresponding to the bucket from the offline statistical results * and I * (such as 120s and 90s).
[0103] Prediction of playing duration: Calculate using the regression model trained offline
[0104] Sensitivity calculation: Substitute L i and I i into the fixed partial derivative formula to calculate ω L and ω I .
[0105] Weight allocation: Normalize the sensitivity through Softmax:
[0106]
[0107] Dynamic division threshold generation:
[0108] θ i =α i L * +β i I *
[0109] Assume α i =0.645 / (0.645 + 15.012) ≈ 0.041, β i =15.012 / (0.645 + 15.012) ≈ 0.959.
[0110] The dynamically generated division threshold: θ i =0.041×120 + 0.959×90 ≈ 4.92 + 86.31 = 91.23s.
[0111] According to the solution of the embodiment of the present disclosure, the weight parameters are obtained by normalizing the sensitivity through Softmax, which can be directly regarded as the relative contribution ratio of the duration of text and pictures. The picture sensitivity and text sensitivity may have large numerical differences due to different units. Softmax eliminates the influence of dimensions through exponential transformation and normalization. The exponential function exp(ω) of Softmax amplifies the weights of high-sensitivity features and suppresses low-sensitivity features. The sensitivity may be negative, but the Softmax output is always positive, which conforms to the physical meaning of weights.
[0112] In a possible implementation manner, it further includes the steps of:
[0113] S210. Respectively according to the number of text words and the number of pictures, classify multiple sample resources in the historical resource library into one of multiple text buckets and one of multiple picture buckets.
[0114] S220. For multiple text buckets and multiple image buckets, sort them according to the actual playing duration of the sample resources in each bucket.
[0115] S230. Determine the first reference threshold for multiple text buckets and the second reference threshold for multiple image buckets according to the actual playing duration of the sample resources at the preset quantiles in the sorting result.
[0116] In the embodiments of the present disclosure, the sample types may include neutral samples in addition to positive samples and negative samples. According to different preset quantiles, the first reference threshold and the second reference threshold may be positive sample thresholds or negative sample thresholds.
[0117] When offline statistically calculating the sample thresholds of each bucket in step S122, the sample resources in each bucket can be sorted according to their actual playing durations. The resources at the top 20% quantile of the playing durations in each bucket are used as positive samples. The actual playing duration of the sample resources at the preset quantile is the positive sample threshold or the negative sample threshold. For example: the playing duration of the samples at the top 20% quantile of the playing durations from high to low in each bucket is used as the positive sample threshold, which is the positive sample threshold for the text bucket, and which is the positive sample threshold for the image bucket. The playing duration of the samples at the bottom 20% quantile of the playing durations in each bucket is used as the negative sample threshold, such as The preset quantile of 20% is only an example.
[0118] According to the solution of the embodiments of the present disclosure, the positive and negative sample demarcation points are defined based on the actual playing duration distribution. The positive and negative samples are distinguished, improving the accuracy and interpretability of sample screening.
[0119] In a possible implementation manner, step S130 further includes the steps of determining the sample type of the target resource according to the estimated playing duration and the division threshold:
[0120] When the estimated playing duration is greater than the positive sample division threshold, determine the sample type of the target resource as a positive sample.
[0121] In a possible implementation manner, step S130 further includes the steps of determining the sample type of the target resource according to the estimated playing duration and the division threshold:
[0122] When the estimated playing duration is less than the negative sample division threshold, determine the sample type of the target resource as a negative sample.
[0123] In the embodiments of the present disclosure, the positive and negative sample determination rules can be changed to:
[0124] Positive samples:
[0125] Negative samples:
[0126] Neutral samples:
[0127] According to the solution of the embodiment of the present disclosure, the existence of neutral samples reduces noise, improves the distinguishability between positive and negative samples, and adapts to the feature sensitivity of different resources.
[0128] In a possible implementation, it further includes the steps of:
[0129] Obtain multiple sample resources from the historical resource library.
[0130] Determine the parameters of the playback duration prediction model according to the text word count, the number of pictures, and the actual playback duration of the multiple sample resources.
[0131] Among them, the playback duration prediction model is a polynomial regression model, and the parameters further include the steps of: intercept term parameter, word count linear term parameter, picture count linear term parameter, word count logarithmic term parameter, picture count logarithmic term parameter, interaction term coefficient.
[0132] In the embodiment of the present disclosure, a polynomial regression model is used for modeling, and the playback duration T is predicted through the text word count L, the number of pictures I, and their non-linear transformations. The model is as follows:
[0133]
[0134] Among them, α0, α1……,α5 are regression parameters to be estimated, L i is the text word count of sample i, and I i is the number of pictures of sample i. Since it is difficult to estimate online user behavior with offline samples, this model is used to capture non-linear effects through logarithmic transformation and interaction terms while retaining the linear relationship to fit the relationship between the playback duration and the features.
[0135] α1L i ,α2I i respectively represent the direct linear effects of the text word count and the number of pictures on the playback duration
[0136] α3log(1+L i ),α4log(1+I i ) uses the logarithmic function to process the number of texts and pictures. This is because when the number of words or pictures is large, the impact of adding one more word or picture may weaken.
[0137] α5L i I i is the interaction term, indicating the impact on the playback duration under the combined action of text and pictures.
[0138] This model combines linear terms, logarithmic terms, and interaction terms. The meanings of each parameter are as follows:
[0139] 1. Intercept term α0
[0140] Mathematical meaning: When the number of words (L i ) and the number of images (I i ) are both zero, it is the predicted baseline playback duration. It represents the basic attractiveness of the content, such as the default duration that attracts users to watch based solely on the title or cover.
[0141] Example: If α0 = 15 seconds, it means that even without text and images, users may stay for 15 seconds due to the title.
[0142] 2. Linear term coefficients α1 and α2
[0143] α1 (linear term of the number of words):
[0144] Mathematical meaning: The fixed gain in playback duration for each additional word (independent effect unrelated to the logarithmic term). It reflects the direct positive / negative effect of the number of words on the playback duration.
[0145] Example: α1 = 0.05 → For each additional word, the playback duration increases by 0.05 seconds.
[0146] α2 (linear term of the number of images):
[0147] Mathematical meaning: The fixed gain in playback duration for each additional image. It measures the independent value of the number of images, such as visual attractiveness or information supplementation.
[0148] Example: α2 = 2 → For each additional image, the playback duration increases by 2 seconds.
[0149] 3. Logarithmic term coefficients α3 and α4
[0150] α3log(1 + L i )(linear term of the number of words):
[0151] Mathematical meaning: The adjustment amount of the natural logarithm of the number of words to the playback duration, capturing the diminishing marginal effect. The more words there are, the gradually decreasing gain of each additional word to the duration, which conforms to the phenomenon of user reading fatigue or information overload.
[0152] Example: α3 = 10 → When the number of words increases from 100 to 1000, the logarithmic term gain is 10 (ln(1001) - ln(101)) ≈ 23 seconds, but the marginal gain for each newly added word decreases from 10 / (1 + 100) ≈ 0.099 seconds to 10 / (1 + 1000) ≈ 0.01 seconds.
[0153] α4log(1 + I i )(linear term of the number of images):
[0154] Mathematical meaning: The amount of adjustment of the natural logarithm of the number of pictures to the playing duration. When the number of pictures increases, the gain per single picture gradually decreases, possibly due to visual fatigue or information redundancy.
[0155] Example: α4 = 5 → When the number of pictures increases from 2 to 5, the gain of the logarithmic term is 5(ln(6)-ln(3)) ≈ 5×0.693 ≈ 3.475 seconds, but the marginal gain per additional picture decreases from 5 / (1 + 2) ≈ 1.67 seconds to 5 / (1 + 5) ≈ 0.83 seconds.
[0156] 4. Interaction term coefficient α5
[0157] Mathematical meaning: The amount of adjustment of the product of the number of words and the number of pictures to the playing duration, measuring the synergy effect.
[0158] α5 > 0: The text and pictures complement each other (e.g., the text explanation enhances the understanding of the pictures), further increasing the duration.
[0159] α5 < 0: The text and pictures conflict (e.g., content overload), resulting in a decrease in the duration.
[0160] Example: α5 = 0.001 → If L i = 2000 words, I i = 5 pictures, the gain of the interaction term is 0.001×2000×5 = 10 seconds, indicating that the combination of text and pictures is reasonable.
[0161] Summary of model parameter characteristics
[0162]
[0163] In an example where a combination of parameters affects the prediction, assume the following content parameters:
[0164] Number of words L i = 1500 words, number of pictures I i = 4 pictures
[0165] Model parameters: α0 = 15, α1 = 0.05, α2 = 2, α3 = 10, α4 = 5, α5 = 0.001
[0166] Calculation of playing duration:
[0167]
[0168] Analysis of key influences:
[0169] The linear term of the number of words (+75 seconds) dominates, and the logarithmic term (+73.13 seconds) further amplifies the gain, but the marginal effect decreases (the gain of the 1000th word is higher than that of the 1500th word).
[0170] The contribution of the logarithm term of the number of pictures (+8.05 seconds) is limited, indicating that the number of pictures is approaching saturation.
[0171] The interaction term (+6 seconds) shows a weak synergistic effect between pictures and text, and the combination of pictures and text needs to be optimized.
[0172] According to the solution of the embodiment of the present disclosure, through the combination of the linear term, the logarithm term and the interaction term, the logarithm term captures the attenuation effect, avoids the deviation of the simple linear model, and balances the direct effect, the attenuation effect and the synergistic effect.
[0173] The whole process is divided into two stages: offline training and online application:
[0174] 1. In the offline training stage, it is used to construct the training set of the ranking model
[0175] Input: Historical resource library (including actual play duration, number of text words, number of pictures).
[0176] Objective: Screen high-quality positive and negative samples for training the ranking model.
[0177] Implementation steps:
[0178] Step 1: Static bucketing (bucketing by the number of text words and the number of pictures).
[0179] Step 2: Train a polynomial regression model to predict the play duration.
[0180] Step 3: Calculate the picture sensitivity and text sensitivity of each resource.
[0181] Step 4: Generate dynamic threshold positive sample threshold and negative sample threshold (based on bucketing statistics and sensitivity weights).
[0182] Step 5: Compare the predicted value with the dynamic threshold to label positive and negative samples.
[0183] In one example, the actual play duration of resource A is 180s (relatively high because it is recommended to the home page), but the model predicts T A = 120s (reflecting its true value of picture and text quality). The actual play duration of resource B is 60s (underestimated due to insufficient exposure), but the model predicts T B = 100s. Using the predicted value can avoid the interference of the recommendation strategy and screen out purer positive and negative samples.
[0184] 2. Online application stage (processing new resources)
[0185] Input: New resource (without play duration, only the number of text words and the number of pictures).
[0186] Objective: Predict its play duration and determine whether it should be used as a candidate positive sample or negative sample.
[0187] Implementation steps:
[0188] Step 1: Divide the new resources into the corresponding L and I buckets.
[0189] Step 2: Use the regression model trained offline to calculate
[0190] Step 3: Calculate the picture sensitivity and text sensitivity of the resource.
[0191] Step 4: Based on the bucket threshold L * and I * and the sensitivity weight to generate the dynamic threshold θ i .
[0192] Step 5: If mark as a positive sample; if mark as a negative sample; if mark as a neutral sample.
[0193] According to the solution of the embodiment of the present disclosure, the dynamic threshold is fused by bucket statistics and sensitivity weights to achieve a balance between bucket standardization and personalized adjustment. After the regression model training and offline statistics are completed, it can be applied to the positive and negative sample division of new resources. The new resources predict the playback duration and classify through the same process to ensure consistency with the training set and guarantee the generalization of the model.
[0194] Figure 2 is a schematic structural diagram of a sample screening device provided according to an embodiment of the present disclosure. As Figure 2 shown, the device includes:
[0195] An estimation module 201, configured to determine the estimated playback duration of the target resource according to the number of text words and the number of pictures of the target resource.
[0196] A threshold determination module 202, configured to adjust the reference threshold of the corresponding bucket of the target resource according to the text sensitivity and picture sensitivity of the target resource to obtain a division threshold.
[0197] A screening module 203, configured to determine the sample type of the target resource according to the estimated playback duration and the division threshold. Among them, the sample type includes positive samples and negative samples.
[0198] In a possible implementation manner, the estimation module 201 is configured to:
[0199] Input the number of text words and the number of pictures of the target resource into the playback duration prediction model to obtain the estimated playback duration.
[0200] Among them, the playback duration prediction model is a polynomial regression model.
[0201] In a possible implementation, the threshold determination module 202 is configured to:
[0202] Determine a text bucket and a picture bucket corresponding to the target resource according to the number of words and the number of pictures included in the target resource.
[0203] Determine a first reference threshold for the text bucket and a second reference threshold for the picture bucket.
[0204] Determine the text sensitivity and the picture sensitivity of the target resource.
[0205] Dynamically adjust the first reference threshold and the second reference threshold according to the text sensitivity and the picture sensitivity to obtain a division threshold.
[0206] In a possible implementation, the threshold determination module 202 is configured to:
[0207] Perform a partial derivative calculation on the number of words of the target resource according to the playback duration prediction model to obtain the text sensitivity.
[0208] Perform a partial derivative calculation on the number of pictures of the target resource according to the playback duration prediction model to obtain the picture sensitivity.
[0209] In a possible implementation, the threshold determination module 202 is configured to:
[0210] Map the text sensitivity and the picture sensitivity to the probability space through an exponential function to obtain a normalized text weight and a picture weight.
[0211] Adjust the first reference threshold according to the text weight.
[0212] Adjust the second reference threshold according to the picture weight.
[0213] Obtain a division threshold according to the adjusted first reference threshold and the adjusted second reference threshold.
[0214] In a possible implementation, the device further includes an offline statistics module, configured to:
[0215] Divide multiple sample resources in the historical resource library into one of multiple text buckets and one of multiple picture buckets respectively according to the number of words and the number of pictures.
[0216] For multiple text buckets and multiple picture buckets, sort according to the actual playback duration of the sample resources in each bucket.
[0217] Determine a first reference threshold for multiple text buckets and a second reference threshold for multiple picture buckets according to the actual playback duration of the sample resources at a preset quantile in the sorting result.
[0218] In a possible implementation, the screening module 203 is configured to:
[0219] When the estimated playing duration is greater than the positive sample division threshold, determine the sample type of the target resource as a positive sample.
[0220] In a possible implementation, the screening module 203 is configured to:
[0221] When the estimated playing duration is less than the negative sample division threshold, determine the sample type of the target resource as a negative sample.
[0222] In a possible implementation, it further includes a module training module, which is configured to:
[0223] Obtain multiple sample resources from the historical resource library.
[0224] Determine the parameters of the playing duration prediction model according to the text word count, number of pictures, and actual playing duration of the multiple sample resources.
[0225] Wherein, the playing duration prediction model is a polynomial regression model, and the parameters include: intercept term parameter, word count linear term parameter, picture count linear term parameter, word count logarithmic term parameter, picture count logarithmic term parameter, interaction term coefficient.
[0226] For the specific functions and examples of each module and sub-module of the device in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.
[0227] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0228] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0229] Figure 3 FIG. shows a schematic block diagram of an exemplary electronic device 300 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0230] Such as Figure 3As shown, device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the device 300 can also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0231] Multiple components in the device 300 are connected to the I / O interface 305, including: an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a disk, an optical disc, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the device 300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0232] The computing unit 301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 301 executes the various methods and processes described above, such as the sample screening method. For example, in some embodiments, the sample screening method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of the sample screening method described above can be executed. Alternatively, in other embodiments, the computing unit 301 can be configured to execute the sample screening method in any other appropriate way (e.g., by means of firmware).
[0233] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0234] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0235] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0236] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0237] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0238] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.
[0239] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0240] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for sample screening, comprising: Determining an estimated playback duration of the target resource according to the number of words and the number of pictures of the target resource; Adjusting a reference threshold of a corresponding bucket of the target resource according to the text sensitivity and the picture sensitivity of the target resource to obtain a division threshold; Determining a sample type of the target resource according to the estimated playback duration and the division threshold; wherein, the sample type includes positive samples and negative samples.
2. The method according to claim 1, wherein Determining an estimated playback duration of the target resource according to the number of words and the number of pictures of the target resource, comprising: Inputting the number of words and the number of pictures of the target resource into a pre-trained playback duration prediction model to obtain an estimated playback duration; Wherein, the playback duration prediction model is a polynomial regression model, and the parameters of the playback duration prediction model are trained according to the number of words, the number of pictures and the actual playback duration of multiple sample resources.
3. The method according to claim 1, wherein, Adjusting a reference threshold of a corresponding bucket of the target resource according to the text sensitivity and the picture sensitivity of the target resource to obtain a division threshold, comprising: Determining a text bucket and a picture bucket corresponding to the target resource according to the number of words and the number of pictures included in the target resource; Determining a first reference threshold of the text bucket and a second reference threshold of the picture bucket; Determining the text sensitivity and the picture sensitivity of the target resource; Dynamically adjusting the first reference threshold and the second reference threshold according to the text sensitivity and the picture sensitivity to obtain a division threshold.
4. The method according to claim 3, wherein, Determining the text sensitivity and the picture sensitivity of the target resource, comprising: Performing a partial derivative calculation on the number of words of the target resource according to the playback duration prediction model to obtain text sensitivity; Performing a partial derivative calculation on the number of pictures of the target resource according to the playback duration prediction model to obtain picture sensitivity.
5. The method according to claim 3 or 4, wherein Dynamically adjusting the first reference threshold and the second reference threshold according to the text sensitivity and the picture sensitivity to obtain a division threshold, comprising: Mapping the text sensitivity and the picture sensitivity to a probability space through an exponential function to obtain a normalized text weight and a picture weight; Adjusting the first reference threshold according to the text weight; Adjusting the second reference threshold according to the picture weight; Obtaining a division threshold according to the adjusted first reference threshold and the adjusted second reference threshold.
6. The method according to claim 3, further comprising: Dividing multiple sample resources in a historical resource library into one of multiple text buckets and one of multiple picture buckets respectively according to the number of words and the number of pictures; Sorting the multiple text buckets and the multiple picture buckets according to the actual playback duration of the sample resources in each bucket; Determining a first reference threshold of the multiple text buckets and a second reference threshold of the multiple picture buckets according to the actual playback duration of the sample resources at a preset quantile in the sorting result.
7. The method according to claim 1, wherein, The determining the sample type of the target resource according to the estimated playback duration and the division threshold includes: In the case where the predicted playing duration is greater than the positive sample division threshold, determine the sample type of the target resource as a positive sample.
8. The method according to claim 1, wherein, The determining the sample type of the target resource according to the predicted playing duration and the division threshold includes: In the case where the predicted playing duration is less than the negative sample division threshold, determine the sample type of the target resource as a negative sample.
9. The method according to claim 2, further comprising: Obtain a plurality of sample resources from the historical resource library; Determine the parameters of the playing duration prediction model according to the text word count, the number of pictures, and the actual playing duration of the plurality of sample resources; Wherein, the playing duration prediction model is a polynomial regression model, and the parameters include: an intercept term parameter, a word count linear term parameter, a picture count linear term parameter, a word count logarithmic term parameter, a picture count logarithmic term parameter, and an interaction term coefficient.
10. A sample screening device, comprising: An estimation module, configured to determine the predicted playing duration of the target resource according to the text word count and the number of pictures of the target resource; A threshold determination module, configured to adjust the reference threshold of the corresponding bucket of the target resource according to the text sensitivity and the picture sensitivity of the target resource to obtain a division threshold; A screening module, configured to determine the sample type of the target resource according to the predicted playing duration and the division threshold; wherein, the sample type includes positive samples and negative samples.
11. An electronic device, comprising: [[ID= 12. A non-transitory computer-readable storage medium storing computer instructions, wherein,