Method and device for scoring and ranking based on business object features, electronic device, storage medium

By constructing a ranking model based on the characteristics of business objects and learning the online performance relationships, the problem of high bandwidth consumption in online control experiments was solved. This enabled the ranking of business objects before online experiments, improving decision-making efficiency and accuracy.

CN122153371APending Publication Date: 2026-06-05GUANGZHOU HUYA TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU HUYA TECH CO LTD
Filing Date
2026-03-05
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

In existing technologies, online control experiments require a large amount of real traffic, have long experimental cycles, and cannot verify a large number of business objects to be tested one by one, resulting in slow iteration speed of key decisions and high cost of decision trial and error.

Method used

By constructing a ranking model based on business object features, and using pre-trained samples to learn the online performance relationships in the feature space of business objects, the model extracts the features of the objects to be evaluated and ranks them according to the predicted scores, thus achieving performance prediction without online experiments.

Benefits of technology

It significantly reduces the cost of decision-making trial and error, improves the speed of decision-making iteration, and enables accurate prediction of the superiority and inferiority of business objects before they go live, thereby reducing traffic consumption and experimental cycles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122153371A_ABST
    Figure CN122153371A_ABST
Patent Text Reader

Abstract

The application provides a scoring and ranking method and device based on business object features, an electronic device and a storage medium. The method comprises: obtaining a pre-trained ranking model; wherein the ranking model is trained by a pre-constructed sample pair, and the training target is to make the predicted score of the sample feature with a high business index value higher than the predicted score of the sample feature with a low business index value in the same sample pair; obtaining different versions of a to-be-evaluated business object, and extracting the features corresponding to each version of the to-be-evaluated business object; inputting the features under each version into the ranking model to obtain the predicted scores corresponding to the features under each version; and ranking the different versions of the to-be-evaluated business object according to the predicted scores. The application can automatically and accurately predict the online performance of the business object in the future, and reduce the decision-making trial and error cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of content recommendation technology, and more specifically, to a rating and ranking method, apparatus, electronic device, and storage medium based on business object characteristics and a ranking model training method. Background Technology

[0002] In internet products and services, numerous key decisions (such as content recommendation, advertising, UI design, and marketing strategies) are systematic choices that directly impact user experience and business results. Currently, judging the merits of these key decisions generally relies on experimental methods such as online controlled trials for final verification. This involves deploying different versions of test materials to real user traffic and observing and comparing their performance differences in key performance indicators such as click-through rate, conversion rate, viewing time, retention rate, or revenue to determine which version is superior.

[0003] However, current controlled experiments and their shortcomings are as follows: online controlled experiments require a large amount of real traffic, and sufficient sample size must be reserved before the experiment starts to ensure statistical significance. The experimental cycle usually lasts for several days or even weeks, which severely limits the iteration speed of key decisions. At the same time, faced with a massive number of business objects to be evaluated, such as tens of thousands of advertising creatives, tens of thousands of video cover candidate images, or hundreds of UI layout variations, online controlled experiments cannot conduct tests on all the materials to be tested one by one. They can only sample a very small percentage of them for verification, and the vast majority of materials never get any empirical evaluation opportunity.

[0004] Therefore, the technical problem that needs to be solved is how to provide a universal and automated sorting method that can accurately predict the future online performance of business objects before they go live. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a scoring and ranking method, apparatus, electronic device, and storage medium based on business object characteristics, which can automatically and accurately predict the future online performance of business objects and reduce the cost of decision-making trial and error.

[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, the present invention provides a scoring and ranking method based on business object features. The method includes: obtaining a pre-trained ranking model; wherein the ranking model is trained by pre-constructed sample pairs, and the training objective is to make the predicted score of the sample features with high business indicator values ​​in the sample pairs higher than the predicted score of the sample features with low business indicator values; obtaining different versions of the business objects to be evaluated, and extracting features corresponding to each version of the business objects to be evaluated; inputting the features under each version into the ranking model to obtain the predicted score corresponding to the features under each version; and ranking the different versions of the business objects to be evaluated according to the predicted scores.

[0007] Secondly, the present invention provides a ranking model training method, the method comprising: constructing a training dataset; wherein the training dataset contains multiple sample pairs; each sample pair consists of features corresponding to two different versions of the same business object, the two different versions having different business indicator values ​​in online experiments; inputting the sample pairs into a machine learning model to obtain feature prediction scores for each feature and calculating score differences; when the score difference is less than or equal to a preset boundary value, calculating a loss value based on a preset ranking loss function, and updating the parameters of the machine learning model according to the loss value before continuing to process the next sample pair; stopping training after all the sample pairs have been processed, thereby obtaining a ranking model.

[0008] Thirdly, the present invention provides a scoring and ranking method apparatus based on business object features, comprising: an acquisition module for acquiring a pre-trained ranking model; wherein the ranking model is trained by pre-constructed sample pairs, and the training objective is to make the predicted score of the sample features with high business indicator values ​​in the sample pairs higher than the predicted score of the sample features with low business indicator values; the acquisition module is further configured to acquire different versions of the business objects to be evaluated; an extraction module is configured to extract features of the business objects to be evaluated under each version; a prediction module is configured to input the features under each version into the ranking model to obtain the predicted score corresponding to the content features under each version; and a ranking module is configured to rank the different versions of the business objects to be evaluated according to the predicted scores.

[0009] Fourthly, the present invention provides a ranking model training apparatus, comprising: a construction module for constructing a training dataset; wherein the training dataset contains multiple sample pairs; each sample pair consists of features corresponding to two different versions of the same business object, the two different versions having different business indicator values ​​in online experiments; a training module for inputting the sample pairs into a machine learning model, obtaining feature prediction scores for each feature and calculating score differences; when the score difference is less than or equal to a preset boundary value, calculating a loss value based on a preset ranking loss function, updating the parameters of the machine learning model according to the loss value, and continuing to process the next sample pair; stopping training after all the sample pairs have been processed, thus obtaining a ranking model. Fifthly, the present invention provides an electronic device, comprising a processor and a memory, the memory storing machine-executable instructions executable by the processor, the processor executing the machine-executable instructions to implement the method described in any of the foregoing embodiments.

[0010] In a sixth aspect, the present invention provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any of the foregoing embodiments.

[0011] The present invention provides a scoring and ranking method, apparatus, electronic device, and storage medium based on business object features for ranking and ranking model training. First, a pre-trained ranking model is obtained. The training objective of this model is to output higher predicted scores for samples with high business indicator values. This allows the model to learn the implicit order relationships in the business object feature space that are strongly correlated with real online performance, ensuring the model has the ability to discriminate against unseen business objects in the future. Second, for different versions of the business object to be evaluated, features corresponding to each version are extracted. Then, the features of each version are input into the ranking model to obtain predicted scores. Since the model has been repeatedly optimized during the training phase using massive sample pairs to determine the relationship between high scores for high business indicator values ​​and low scores for low business indicator values, the scores between different versions of the business object can support the ranking of different versions. Finally, different versions are ranked according to the predicted scores, thus achieving the effect of predicting the ranking of different versions of business objects based solely on features without undergoing online experiments, significantly reducing the cost of decision-making trial and error.

[0012] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 A schematic flowchart illustrating the scoring and ranking method based on business object characteristics provided in an embodiment of the present invention; Figure 2 A sorting learning system provided in an embodiment of the present invention; Figure 3 A schematic flowchart illustrating the ranking model training method provided in an embodiment of the present invention; Figure 4 A functional block diagram of a scoring and ranking device based on business object characteristics provided in an embodiment of the present invention; Figure 5 This is a functional block diagram of the sorting model training device provided in an embodiment of the present invention; Figure 6 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0016] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0017] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0018] In internet products (such as live streaming, short video, and shopping applications), technical staff make many critical decisions every day. These include deciding which video to recommend to users, which creative image to use for ads, and which UI version to use for the new app homepage. These decisions must be validated based on real user feedback. The most common method is online controlled experiments, which involve randomly dividing users into groups, each seeing different versions (e.g., cover A vs. cover B), and then tracking which group clicks more and stays longer. Ideally, this statistical data should be used to determine which version to choose.

[0019] However, in practical applications, online controlled experiments require a large amount of real traffic and have a long experimental cycle, resulting in slow decision-making iteration. For a massive number of business objects to be tested (such as advertising creatives and recommendation content), it is impossible to test each one individually. Furthermore, the results of online experiments are often isolated and one-off, and are difficult to systematically learn from and use to guide future decisions.

[0020] To address the aforementioned technical issues, embodiments of the present invention can combine existing online test results to provide an automated scoring and ranking method based on business object characteristics, enabling devices to determine which version performs better even before the business object is launched.

[0021] Please see Figure 1 , Figure 1 The schematic flowchart illustrates a scoring and ranking method based on business object characteristics provided in an embodiment of the present invention. The method may include steps S101 to S104, as described below: S101: Obtain a pre-trained ranking model; The ranking model in this embodiment of the invention is trained by pre-constructed sample pairs, and the training objective is to make the predicted score of the sample features with high business indicator values ​​in the sample pairs higher than the predicted score of the sample features with low business indicator values. S102: Obtain different versions of the business object to be evaluated, and extract the features corresponding to each version of the business object to be evaluated; In the embodiments of this invention, the business object can also be referred to as "material," which refers to any digital content unit or strategy configuration unit in Internet product services that can participate in online control experiments as an independent variable, including but not limited to: images, video frames, text fragments, audio fragments, UI elements (such as buttons and pop-ups), recommendation strategies, advertising creatives, algorithm parameter configurations, etc.

[0022] S103: Input the features of each version into the ranking model to obtain the prediction score corresponding to the features of each version; S104: Sort the different versions of the business objects to be evaluated according to the predicted scores.

[0023] Unlike existing technologies, this invention provides a pre-trained ranking model. The training objective of this model is to output higher predicted scores for samples with high business indicator values. This allows the model to learn the implicit order relationships in the feature space of business objects that are strongly correlated with real online performance, ensuring that the model has the ability to discriminate against unseen business objects in the future. Secondly, for different versions of the business objects to be evaluated, features corresponding to each version are extracted. Then, the features of each version are input into the ranking model to obtain predicted scores. Since the model has been repeatedly optimized during the training phase using massive sample pairs to determine the relationship between high scores for high business indicator values ​​and low scores for low business indicator values, the scores between different versions of business objects can support the ranking of different versions. Finally, the different versions are ranked according to the predicted scores, thus achieving the effect of predicting the ranking of different versions of business objects based solely on features without undergoing online experiments, significantly reducing the cost of decision-making trial and error.

[0024] Next, the embodiments of the present invention will be described in conjunction with the relevant accompanying drawings. Figure 1 The sorting process is explained in detail.

[0025] In one embodiment of the present invention, before performing step S101 "obtaining the pre-trained ranking model", the ranking model can be constructed and persisted in any of the following ways: In one implementation, the ranking model can be trained end-to-end locally by the device that performs the ranking method described above (hereinafter referred to as "this device"). This device integrates a sample construction module and a model training module, which can continuously accumulate training samples based on historical online experimental feedback and periodically update model parameters. The trained ranking model can be persistently stored in the local storage unit of this device in the form of weight files, graph structures, or serialized model objects.

[0026] In another implementation, the ranking model can also be independently trained by an external system and provided to this device for deployment. For example, it can be pre-trained by an independent machine learning platform, model repository, or other training tools, and the model-related files can be exported. The device executing the above ranking method can directly load the model.

[0027] It should be understood that, regardless of the implementation method described above, when the device has a sorting requirement, that is, when it receives feature representations of different versions of a new business object (such as a cover to be selected, an advertising creative, etc.), the sorting model can be directly called without performing time-consuming operations such as model retraining, fine-tuning, or online learning. It can directly receive the feature representations of the new business object and output the predicted score.

[0028] Next, this embodiment of the invention will take the example that the ranking model can be trained locally by the device that performs the ranking method described above, to introduce the training process of the ranking model.

[0029] In this embodiment of the invention, the ranking model is pre-trained based on sample pairs constructed from different business objects. These sample pairs are constructed by the device based on experimental results data of different versions of the business objects. The training objective of the model is to make the feature prediction score of the sample with the higher business indicator value in the sample pair higher than the feature prediction score of the sample with the lower business indicator value. In this way, if the feature of the business object to be evaluated falls into the high-score region in the space, it is assigned a high score; if it falls into the low-score region, it is assigned a low score. This enables the generalization of the relative superiority and inferiority relationships of different versions of any new business object.

[0030] Based on the above considerations, this device can first construct sample pairs during the training of the ranking model. This can be understood as follows: for different versions of the same business object, this embodiment of the invention can automatically filter out version pairs with sufficiently significant performance differences from the actual business performance data tested in online experiments, and transform them into sample pairs that the model can learn, thereby enabling the ranking model to achieve its training objective. Specifically, the method for constructing sample pairs is shown in steps a1 to a3, as explained below: Step a1: Obtain the business indicator values ​​corresponding to different versions of each business object; In this embodiment of the invention, the business metric value is obtained by conducting online experiments on different versions of the same business object and statistically analyzing the experimental results. The business metric value can be, but is not limited to, any quantifiable key performance indicator such as click-through rate, conversion rate, viewing time, retention rate, and revenue. When there are multiple key performance indicators, the final business metric value is the aggregate value of these multiple key performance indicators.

[0031] As an optional implementation method, business metric values ​​can be obtained by: distributing different versions of each business object to target users; collecting behavioral data for each target user; and determining the business metric value for each version based on the behavioral data of the target users corresponding to each version.

[0032] This can be understood as follows: The device first acquires multiple versions of each business object (such as different covers of the same video, different copy of the same advertisement, and different variations of the same UI design), and conducts online experiments for each version. During the experiment, these different versions are randomly distributed to a group of target users with similar characteristics, that is, a homogeneous user group. Then, the system continuously collects behavioral data generated by users after seeing or using the version, such as whether they click, whether they place an order, and how long they stay. Then, the system calculates the business indicator value corresponding to each version based on these raw behavioral data.

[0033] Step a2: When the difference between the business indicator values ​​of any two target versions of each business object is greater than or equal to the preset significant difference value, obtain the features corresponding to the two target versions; In this embodiment of the invention, the device compares the business indicator values ​​corresponding to any two versions under the same business object. If the numerical difference between them is greater than or equal to a pre-set significant difference value, the system considers that the two versions do indeed have a reliable difference in actual business performance. Then, it extracts the features corresponding to each of the two versions to construct a training sample pair with a clear superior-inferior relationship. This is not done simply by constructing a training sample pair if the business indicator value of one version is greater than that of the other.

[0034] It is understandable that the significant difference value in step a2 above is designed by the embodiments of the present invention to address the inherent characteristics of online experimental data. Considering that online A / B experiments are affected by multiple factors such as the randomness of real user behavior, traffic distribution disturbances, and occasional abnormal events, if sample pairs are constructed solely based on the fact that the business indicator value of one version is greater than that of another version, it is very easy to misjudge random fluctuations as real experimental results, leading the model to learn a large number of false superior-inferiority relationships. Introducing a significant difference value is equivalent to introducing a random fluctuation range. The system only constructs sample pairs when the difference in business indicator values ​​stably exceeds the random fluctuation range, so that each training sample represents a real intervention effect that can withstand statistical inference testing. Furthermore, constructing sample pairs solely based on the fact that one version's business metric value is greater than another's is essentially a weak signal of superiority or inferiority. If the model continuously learns these weak signals during training, it will overemphasize subtle feature noise (such as a pixel's color on a cover or a stop word in text), causing drastic fluctuations in performance in unseen scenarios such as new videos, new user groups, and new device types. Sample pairs filtered by significant difference values, however, represent reliable patterns with clear superiority or inferiority relationships. The model thus focuses on learning core patterns of superiority and inferiority that consistently emerge under different experimental conditions (such as "high-contrast dynamic covers are generally better among younger users"), significantly enhancing its generalization and predictive ability for a massive number of new business objects.

[0035] Optionally, during feature extraction, features can be extracted from different versions of the object to be evaluated. These features are multimodal representations of the business object itself, such as the visual feature vector extracted from the cover image, the semantic features extracted from the text fragment, etc. Of course, the extracted features can also include contextual information related to the business object (e.g., display location, user profile, time period tags, etc.). That is, the features ultimately used in this embodiment of the invention can be only the features of the business object itself, or only the contextual information related to the business object, or both the features of the business object itself and the contextual information. This embodiment of the invention does not impose any limitations here.

[0036] Step a3: The features of the two target versions constitute a sample pair.

[0037] In this embodiment of the invention, features from two target versions with significantly different real-world business performance are combined to form a sample pair with a clear indication of superiority or inferiority. It can be seen that for the same business object, the sample pair consists of features corresponding to two different versions of that business object, whose business metric values ​​differ in online experiments. For ease of description, the features of the version with the higher business metric value are labeled "Better_Item," and the features of the version with the lower business metric value are labeled "Worse_Item." This sample pair will then be used as training data to drive the model to learn how to predict relative superiority or inferiority relationships based on features.

[0038] To better understand the above process, consider this example: Suppose a business object has two versions, ItemA and ItemB. Online experiments yield the corresponding business metric values ​​for ItemA and ItemB, such as (ItemA, KPI(A)) and (ItemB, KPI(B)). If KPI(A) is statistically significantly better than KPI(B), i.e., KPI(A) > KPI(B) + confidence_interval), where confidence_interval is a preset significant difference value, then a sample pair is generated, such as (Better_Item, Worse_Item), where Better_Item is the feature representation of ItemA, and Worse_Item is the feature representation of ItemB.

[0039] As can be seen, this embodiment of the invention uses online A / B testing results to automatically generate ranked learning sample pairs. Its core lies in determining the superiority or inferiority relationship between different versions of business objects and constructing training samples based solely on verifiable business performance generated by real users under identical conditions. This method eliminates the subjectivity of manual annotation, ensuring that each set of samples reflects empirically verified causal knowledge. Furthermore, this method ensures that online experimental results are no longer isolated reports; the cognitive experience regarding which version of the business object is superior can be systematically learned. The model trained in this way can not only accurately predict the online performance of new business objects but also systematically distill fragmented experimental conclusions into reusable decision-making capabilities.

[0040] Next, based on the sample pairs constructed above, a training dataset can be obtained. This training dataset contains a large number (millions) of sample pairs generated based on online test results. This training dataset can be used as shown in steps b1 to b4, as explained below: Step b1: Construct the training dataset; Step b2: Input the sample pairs into the machine learning model to obtain the feature prediction score for each feature and calculate the score difference; The machine learning model selected in this embodiment of the invention can be a Siamese network, that is, a multilayer perceptron with two branches sharing the same set of parameters, processing the two features in the input sample pair respectively, and outputting a score for each, denoted as . and , These are the scores corresponding to the features of the versions with higher business metric values ​​in the sample pair. This refers to the score corresponding to the feature in the version with the lower business metric value. During this process, the model calculates the difference between these two scores, i.e. .

[0041] Step b3: When the score difference is less than or equal to the preset boundary value, calculate the loss value based on the preset ranking loss function, update the parameters of the machine learning model according to the loss value, and continue to process the next sample pair; It is understandable that when the score difference is less than or equal to the preset boundary value... (For example If the score difference is 0.1, it indicates that the model cannot yet reliably distinguish between good and bad samples. In this case, the system calculates the predicted loss value based on the preset ranking loss function and adjusts the internal parameters of the model accordingly to make the score of the next sample pair closer to the ideal state. When the score difference is greater than the boundary value, it means that the model has met the basic discrimination ability requirement, and the system continues to process the next sample pair.

[0042] In this embodiment of the invention, the loss value is calculated and the model parameters are updated only when the score difference is less than or equal to a preset boundary value; if only the model is required to output a higher score for the "better version" (i.e., > The model might try to minimize loss by infinitely amplifying score differences, causing all scores to detach from actual business scales. For example, assigning a high-quality cover with a 0.2% higher click-through rate a score of 1000, while assigning a cover with only a 0.1% lower click-through rate a score of 1000. 999 points. While this extreme output satisfies the ranking relationship, it completely renders the score meaningless for horizontal comparison, making it unsuitable for a unified evaluation of entirely new business objects. Introducing margin forces the model to require that the predicted difference between two sample features must at least reach this boundary value, ensuring that every ranking decision is based on signals with substantial business differences, rather than minor noise or numerical fluctuations.

[0043] Optionally, the values ​​of the above boundary values ​​can be derived from the analysis of relevant personnel's online experiment results, business sensitivity, and historical data distribution, and can be dynamically adjusted according to the quality of the online experiment system.

[0044] The present invention introduces preset boundary values ​​during the training process of the ranking model. Its fundamental motivation is to accurately transform the sample basis of "significant superiority-inferiority relationship" on which the online A / B experiment relies into mathematical constraints in the model learning process, thereby ensuring that the trained ranking prediction model not only has ranking ability, but also has business credibility, robust stability, interpretability and sustainable evolution engineering practicality.

[0045] Optionally, the sorting loss function can be the boundary sorting loss function, whose mathematical expression is: Alternatively, it could be a triplet loss function.

[0046] Step b4: Stop training after all sample pairs have been processed, and obtain the ranking model.

[0047] As can be seen, the core of the training method provided in this embodiment of the invention lies in having the model assign a score to each of the different versions of material features from the same business object (such as the same video) each time it receives such a pair. The model ensures that the better-performing version scores higher, and the difference is not less than a pre-set minimum gap, i.e., a boundary value. This allows the machine learning model to learn to compare and automatically understand which business objects perform better in online experiments, rather than directly predicting specific values ​​such as the absolute click-through rate or conversion rate of a particular business object. This reduces sensitivity to the quality of labeled data and aligns with the comparison-based verification logic of online experiments, significantly improving the model's learning efficiency and business adaptability. In one embodiment of the invention, the device can also obtain new sample pairs corresponding to each business object; and update the ranking model with these new sample pairs according to a preset update cycle. This can be understood as the entire training process not being run only once, but rather periodically and continuously, for example, once a day, using all newly generated sample pairs from the previous day's online tests to update the model, thereby continuously enhancing the ranking model's capabilities over time.

[0048] After obtaining an updated and more capable ranking model based on the above training method, in the actual application process of this embodiment of the invention, the device first obtains and deploys the ranking model, that is, executes step S101, and then continues to execute steps S102 to S104 on the object to be evaluated.

[0049] Specifically, in step S102, when different versions of the business object to be evaluated are obtained, features corresponding to each version of the business object to be evaluated are extracted. It is understood that the business object to be evaluated can be a business object that has already participated in online testing, or it can be a new business object. Next, features are extracted from the different versions of the business object to be evaluated; and / or, contextual information related to the business object to be evaluated is used as features. In step S103, the features extracted in step S102 are input into a trained ranking model, and the model outputs a predicted score for each version. Then, in step S104, the different versions of the business object to be evaluated are ranked according to the predicted scores.

[0050] Optionally, the sorting method can be based on the scores from highest to lowest or from lowest to highest; no specific restriction is imposed here.

[0051] By using the steps S101 to S104 above, business objects can be automatically scored and sorted before going online, and the sorting results can be used as the basis for optimization decisions.

[0052] In one embodiment of the present invention, the ranking results of different versions of the business objects to be evaluated can be used for online decision-making and to assist in control experiments. The following implementation method is shown: In one implementation, the system can select the business object version with the highest predicted score as the target recommendation object based on the ranking results and push it to the target user. This approach skips the traditional control experiment stage, directly enabling deployment decisions and significantly shortening the product iteration cycle.

[0053] In another implementation, the system can also select the two business object versions with the highest predicted scores as candidate test objects based on the ranking results. This reduces the scale of the control experiment and significantly improves the experimental throughput and resource reuse rate.

[0054] To facilitate a comprehensive understanding of the scoring and ranking method based on business object characteristics provided in this embodiment of the invention, please refer to [link to relevant documentation]. Figure 2 , Figure 2 The ranking learning system provided in this embodiment of the invention mainly consists of four parts: an online experiment module, a sample generation module, a model training module, and an offline prediction module.

[0055] like Figure 2 As shown, the online experiment module is mainly responsible for executing the standard online experiment process. For each version of the business object, it outputs one or more predefined business indicator values, that is, it completes the task of determining the business indicator values ​​of different versions of the same business object in this embodiment of the invention. The sample generation module, based on the experimental conclusions obtained from the online experiment module, that is, the business indicator values ​​corresponding to different versions of the same business object, constructs the sample pairs required for ranking learning, and completes the task of constructing the training dataset.

[0056] Next, after obtaining the set of sample pairs output by the sample generation module, the model training module trains a ranking model capable of predicting the potential performance of business objects. This model periodically (e.g., daily / weekly) absorbs the latest online experimental data to update itself, outputting an updated and more powerful ranking prediction model, thus completing the task of obtaining the ranking model. The offline prediction module can conduct "virtual experiments" before the business objects to be evaluated go live. This involves generating feature representations of one or more business objects to be evaluated, and then using the latest version of the ranking model to output an ordered list of business objects. The order of these objects represents the model's predicted ranking of their future online performance. This ranking result can be directly used to select the best from multiple candidates, perform preliminary screening of massive amounts of business objects, or serve as input features for recommendation systems or advertising systems, etc.

[0057] To more clearly illustrate the working process of the ranking learning system provided in the embodiments of the present invention, the following will take the specific application scenario of "intelligently selecting video cover images" as an example to describe the implementation of the present invention in detail. It should be emphasized that this embodiment is only one of the many possible applications of the present invention and should not be construed as limiting the scope of protection of the present invention in any way.

[0058] Assuming the application scenario is as follows: on a video platform, selecting a high-quality cover image for each video is crucial for attracting user clicks and increasing viewership. Platforms typically prepare multiple candidate covers for the same video (e.g., different frames extracted from the video, images created by designers, etc.). The goal of this embodiment is to automatically predict and select the cover image most likely to receive the highest click-through rate from multiple candidate covers using the method of this invention.

[0059] First, the online experiment module is primarily responsible for obtaining multiple candidate cover images for the same video. For example, for video V1, the candidate covers are Cover_A, Cover_B, and Cover_C. Then, an A / B / C control experiment is performed. When video V1 needs to be shown to users, the system randomly selects one from Cover_A, Cover_B, and Cover_C for display. Then, based on the collected user behavior data, the click-through rate (CTR) of the cover is calculated. That is, CTR = (number of times the cover was clicked) / (total number of times the cover was displayed). Finally, after a period of experimentation (e.g., after 100,000 displays), the final CTR for each cover is output. For example: (Cover_A, CTR=5.2%), (Cover_B, CTR=3.1%), and (Cover_C, CTR=4.5%).

[0060] Then, the sample generation module can extract high-dimensional visual feature vectors for each cover image (Cover_A, Cover_B, and Cover_C) using a pre-trained image encoding model (such as ViT or ResNet), resulting in Feature_A, Feature_B, and Feature_C. Then, the three covers are compared pairwise: (A,B), (A,C), and (B,C), by comparing their CTR values: if CTR(A) > CTR(B), a sample pair (Better_Item=Feature_A, Worse_Item=Feature_B) is generated; if CTR(A) > CTR(C), a sample pair (Better_Item=Feature_A, Worse_Item=Feature_C) is generated; if CTR(C) > CTR(B), a sample pair (Better_Item=Feature_C, Worse_Item=Feature_B) is generated. Finally, three comparison sample pairs containing the visual features of the covers and having a clear superior-inferior relationship are output. These sample pairs are stored in a continuously expanding training dataset.

[0061] Next, the model training module, based on a training dataset containing a large number (in the millions) of the aforementioned sample pairs (generated from all historical control experiments), selects a neural network model. This model can be a "Siamese network" that takes two feature vectors as input and outputs scores through a shared-weights MLP (Multilayer Perceptron). Training is performed according to a given training objective, for example, using ... as the loss function and setting margin=0.1. For each sample pair (Feature_Better, Feature_Worse), the model calculates Score_Better and Score_Worse respectively. The loss function drives the model to optimize, making Score_Better as large as possible by at least 0.1 compared to Score_Worse. This training task is executed daily to learn new knowledge from all online experiments of the previous day, ultimately resulting in an updated ranking model that can be used for cover click-through rate prediction.

[0062] Finally, after receiving the new video V_new, the offline prediction module automatically generates five candidate covers (New_1, New_2, New_3, New_4, New_5) and extracts the visual feature vectors (F_1, F_2, F_3, F_4, F_5) for each cover. Then, the latest version of the ranking model is loaded, and each feature vector is input into the model to obtain its prediction score: Score_1=0.87, Score_2=0.54, Score_3=0.95, Score_4=0.61, Score_5=0.72. The scores are then ranked as follows: Score_3>Score_1>Score_5>Score_4>Score_2. This allows for the rapid selection of the optimal cover without online testing.

[0063] Based on the sorting results, the following decision-making applications can be made: Application 1: The system directly selects New_3, which has the highest score, as the default cover for the video and uploads it to the entire platform, saving the cost of a control experiment.

[0064] Application 2: The system selects the two highest-scoring New_3 and New_1 and performs small-scale A / B testing only on them to significantly reduce the scale of the experiment while ensuring the effectiveness.

[0065] As demonstrated in this embodiment, the scoring and ranking method based on business object characteristics provided by this invention can be successfully applied to video cover selection scenarios. By constructing an automated learning loop from online experimentation to offline prediction, it significantly improves the efficiency and effectiveness of cover selection. This idea can also be extended to any field that requires optimization through online experimentation, such as advertising creativity, recommendation rationale, and UI design.

[0066] In one embodiment of the present invention, as previously described, the ranking model in this embodiment can be trained by an external system and then loaded and used by this device. Therefore, this embodiment also provides a ranking model training method, the execution entity of which can be an external system (i.e., a device not executing the ranking method provided in this embodiment). Please see [link to previous text]. Figure 3 , Figure 3 A schematic flowchart of the ranking model training method provided in this embodiment of the invention includes steps S301 to S304: S301: Construct a training dataset; the training dataset contains multiple sample pairs; each sample pair consists of features corresponding to two different versions of the same business object, and the business indicator values ​​of the two different versions differ in the online experiment; S302: Input the sample pairs into the machine learning model to obtain the feature prediction score for each feature and calculate the score difference; S303: When the score difference is less than or equal to the preset boundary value, calculate the loss value based on the preset ranking loss function, update the parameters of the machine learning model according to the loss value, and continue to process the next sample pair. S304: Training stops after all sample pairs have been processed, and the ranking model is obtained.

[0067] Unlike existing technologies, this invention uses online experimental results for training sample modeling. This ensures that only sample pairs exhibiting stable and observable differences in performance within the online experimental environment are converted into training samples. This effectively avoids noise interference and improves the accuracy and generalization ability of the model. Based on the constructed sample pairs, the model assigns a score to each feature in each sample pair for each attempt, ensuring that the better-performing feature receives a higher score, and the difference is no less than a pre-set minimum gap. This method allows the machine learning model to learn by comparison and automatically understand which business objects perform better in online experiments, rather than directly predicting specific values ​​such as the absolute click-through rate or conversion rate of a particular business object. This reduces sensitivity to the quality of labeled data and aligns with the comparison-based verification logic of online experiments, significantly improving the model's learning efficiency and business adaptability. It should be noted that the training process of the ranking model using an external system described above is the same as the ranking model training process performed by this device as described earlier, and will not be repeated here.

[0068] Based on and Figure 1 Using the same inventive concept, and in order to perform the corresponding steps in the above embodiments and various possible methods, an implementation of a scoring and ranking device 40 based on business object characteristics is given below. Please refer to... Figure 4 , Figure 4 The present invention provides a functional block diagram of a scoring and ranking device based on business object features. The scoring and ranking device 40 based on business object features includes: an acquisition module 401, an extraction module 402, a prediction module 403, and a ranking module 404.

[0069] The acquisition module 401 is used to acquire a pre-trained ranking model; the ranking model is trained by pre-constructed sample pairs, and the training objective is to make the predicted score of the sample features with high business indicator values ​​in the sample pairs higher than the predicted score of the sample features with low business indicator values. The acquisition module 401 is also used to obtain different versions of the business objects to be evaluated; Extraction module 402 is used to extract features corresponding to each version of the business object to be evaluated; Prediction module 403 is used to input the features of each version into the ranking model to obtain the prediction score corresponding to the content features of each version; The sorting module 404 is used to sort different versions of the business objects to be evaluated based on the predicted scores.

[0070] It is understandable that the acquisition module 401, extraction module 402, prediction module 403, and sorting module 404 can be executed collaboratively. Figure 1 Each step in the process is to achieve the corresponding technical effect.

[0071] It should be noted that the scoring and ranking device 40 based on business object characteristics provided in this embodiment of the invention can be specific hardware on the device or software or firmware installed on the device. The implementation principle and technical effects of the scoring and ranking device 40 based on business object characteristics provided in this embodiment of the invention are the same as those in the foregoing method embodiments. For the sake of brevity, any parts not mentioned in the device embodiments can be referred to the corresponding content in the foregoing method embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can all be referred to the corresponding processes in the above method embodiments, and will not be repeated here.

[0072] Based on and Figure 3 The same inventive concept, in order to carry out the above embodiments Figure 3 The following describes the implementation of a rating and ranking model training device 50 based on business object characteristics, outlining the corresponding steps in each possible approach. Please refer to [link to relevant documentation]. Figure 5 , Figure 5 The following is a functional block diagram of the sorting model training device provided in an embodiment of the present invention. The sorting model training device 50 includes: a construction module 501 and a training module 502. Module 501 is used to construct a training dataset; wherein the training dataset contains multiple sample pairs; each sample pair consists of features corresponding to two different versions of the same business object, and the business indicator values ​​of the two different versions differ in the online experiment; The training module 502 is used to input sample pairs into the machine learning model, obtain the feature prediction score of each feature and calculate the score difference; when the score difference is less than or equal to a preset boundary value, calculate the loss value based on the preset ranking loss function, update the parameters of the machine learning model according to the loss value and continue to process the next sample pair; when all the sample pairs have been processed, training stops and the ranking model is obtained.

[0073] It is understandable that the construction module 501 and the training module 502 can be executed collaboratively. Figure 3 Each step in the process is to achieve the corresponding technical effect.

[0074] Optionally, the above Figure 4 and Figure 5Modules can be stored in the form of software or firmware. Figure 6 The memory shown is either stored in or embedded in the operating system (OS) of the electronic device 60, and can be used by... Figure 6 The processor in the system executes the data. Specifically, when the electronic device 60 is used only to execute the scoring and ranking method based on business object characteristics provided in the embodiments of the present invention, it can be used to store data. Figure 4 The functional modules of the device shown; when the electronic device 60 is only used to execute the ranking model-based training method provided in the embodiments of the present invention, it can be used to store... Figure 5 The functional modules of the device shown; when the electronic device 60 can be used to execute both the scoring and ranking method based on business object characteristics provided in the embodiments of the present invention and the ranking model training method provided in the embodiments of the present invention, it can be used to simultaneously store... Figure 4 and Figure 5 The device shown has functional modules; at the same time, the data, program code, etc. required to execute the above modules can be stored in the memory.

[0075] Please see Figure 6 , Figure 6 The diagram illustrates the structure of an electronic device according to an embodiment of the present invention, including a memory 601, a processor 602, and a communication interface 603. The memory 601, processor 602, and communication interface 603 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0076] Optionally, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0077] In this embodiment of the invention, the processor 602 may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in this embodiment of the invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in this embodiment of the invention can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor. The software modules may be located in the memory 601, and the processor 602 reads the program instructions from the memory 601 and, in conjunction with its hardware, completes the steps of the aforementioned methods.

[0078] In this embodiment of the invention, the memory 601 can be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or it can be volatile memory, such as RAM. The memory can also be any other medium capable of carrying or storing desired executable program code having an instruction or data structure form and accessible by a computer, but is not limited thereto. The memory in this embodiment of the invention can also be a circuit or any other device capable of implementing a storage function for storing instructions and / or data.

[0079] The memory 601 can be used to store software programs and modules, such as the instructions / modules of the related devices provided in the embodiments of the present invention. These can be stored in the memory 601 in the form of software or firmware, or embedded in the operating system (OS) of the electronic device 60. The processor 602 executes various functional applications and data processing by executing the software programs and modules stored in the memory 601. The communication interface 603 can be used for signaling or data communication with other node devices.

[0080] Understandable. Figure 6 The structure shown is for illustrative purposes only; the electronic device 60 may also include components that are more advanced than those shown. Figure 6 The more or fewer components shown, or having the same Figure 6 The different configurations shown. Figure 6 The components shown can be implemented using hardware, software, or a combination thereof.

[0081] Based on the above embodiments, the present invention also provides a storage medium storing a computer program. When the computer program is executed by a computer, it causes the computer to execute the scoring and ranking method or ranking model training method based on business object characteristics provided in the above embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.

[0082] Based on the above embodiments, the present invention also provides a program product, which includes a computer program. The processor can execute the computer program to implement the scoring and ranking method or ranking model training method based on business object characteristics provided in the embodiments of the present invention. For specific implementation, please refer to the method embodiments, which will not be repeated here.

[0083] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0084] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the objectives of the embodiments of the present invention, depending on actual needs.

[0085] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0086] It should be noted that if the function is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes: USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, and other media capable of storing program code.

[0087] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A scoring and ranking method based on business object characteristics, characterized in that, The method includes: Obtain a pre-trained ranking model; wherein the ranking model is trained by pre-constructed sample pairs, and the training objective is to make the predicted score of the sample features with high business indicator values ​​in the sample pairs higher than the predicted score of the sample features with low business indicator values. Obtain different versions of the business object to be evaluated, and extract the features corresponding to each version of the business object to be evaluated; The features for each version are input into the ranking model to obtain the predicted score corresponding to the features for each version; The different versions of the business objects to be evaluated are sorted according to the predicted scores.

2. The scoring and ranking method based on business object characteristics according to claim 1, characterized in that, Before obtaining the pre-trained ranking model, the method further includes: Construct a training dataset; wherein the training dataset contains multiple sample pairs; each sample pair consists of features corresponding to two different versions of the same business object, and the business indicator values ​​of the two different versions differ in the online experiment; Input the sample pairs into the machine learning model to obtain the feature prediction score for each feature and calculate the score difference; When the score difference is less than or equal to a preset boundary value, a loss value is calculated based on a preset ranking loss function, and the parameters of the machine learning model are updated according to the loss value before continuing to process the next sample pair. Training stops after all sample pairs have been processed, and the ranking model is obtained.

3. The scoring and ranking method based on business object characteristics according to claim 2, characterized in that, Construct the training dataset, including: Obtain the business metric values ​​corresponding to different versions of each business object; When the difference between the business indicator values ​​of any two target versions of each business object is greater than or equal to a preset significant difference value, the features corresponding to the two target versions are obtained. The sample pairs are formed by the features of the two target versions, and all the sample pairs are used to form the training dataset.

4. The scoring and ranking method based on business object characteristics according to claim 2, characterized in that, Obtaining the sample pair further includes: Distribute different versions of each business object to the target users; Collect behavioral data for each target user; The business metric value for each version is determined based on the behavioral data of the target users corresponding to each version.

5. The scoring and ranking method based on business object characteristics according to claim 2, characterized in that, The method further includes: Obtain new sample pairs corresponding to each business object; The ranking model is updated with the new sample pairs according to the preset update cycle.

6. The scoring and ranking method based on business object characteristics according to claim 1, characterized in that, Obtain different versions of the business object to be evaluated, and extract the features corresponding to each version of the business object to be evaluated, including: Extract the features from different versions of the business object to be evaluated; and / or, The contextual information related to the business object to be evaluated is used as the feature.

7. The scoring and ranking method based on business object characteristics according to any one of claims 1-6, characterized in that, The method further includes: Obtain the ranking results of the different versions of the business objects to be evaluated; Based on the ranking results, the version with the highest predicted score will be pushed to the target user; or, Based on the ranking results, the two versions with the highest predicted scores were selected as candidate test subjects for testing.

8. A method for training a ranking model, characterized in that, The method includes: Construct a training dataset; wherein the training dataset contains multiple sample pairs; each sample pair consists of features corresponding to two different versions of the same business object, and the business indicator values ​​of the two different versions differ in the online experiment; Input the sample pairs into the machine learning model to obtain the feature prediction score for each feature and calculate the score difference; When the score difference is less than or equal to a preset boundary value, a loss value is calculated based on a preset ranking loss function, and the parameters of the machine learning model are updated according to the loss value before continuing to process the next sample pair. Training stops after all sample pairs have been processed, resulting in the ranking model.

9. A scoring and ranking device based on business object characteristics, characterized in that, include: An acquisition module is used to obtain a pre-trained ranking model; wherein the ranking model is trained by pre-constructed sample pairs, and the training objective is to make the predicted score of the sample features with high business indicator values ​​in the sample pairs higher than the predicted score of the sample features with low business indicator values. The acquisition module is also used to acquire different versions of the business object to be evaluated, and the extraction module is used to extract the features corresponding to each version of the business object to be evaluated. The prediction module is used to input the features of each version into the ranking model to obtain the prediction score corresponding to the content features of each version; The sorting module is used to sort the different versions of the business objects to be evaluated according to the predicted scores.

10. A sorting model training device, characterized in that, include: A construction module is used to construct a training dataset; wherein the training dataset contains multiple sample pairs; each sample pair consists of features corresponding to two different versions of the same business object, and the business indicator values ​​of the two different versions differ in online experiments; The training module is used to input sample pairs into the machine learning model, obtain the feature prediction score of each feature and calculate the score difference; when the score difference is less than or equal to a preset boundary value, a loss value is calculated based on a preset ranking loss function, and the parameters of the machine learning model are updated according to the loss value before continuing to process the next sample pair; when all the sample pairs have been processed, training stops and a ranking model is obtained.

11. An electronic device, characterized in that, It includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor to implement the sorting method according to any one of claims 1-8.

12. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the sorting method as described in any one of claims 1-8.