An evaluation method and system for the quality of AI technology-based exposure posts and the possibility of singling
By constructing a standardized sample pool, introducing similar historical posts as training context, and conducting dual-task joint training, the accuracy issues of content quality and order conversion probability assessment in post evaluation were resolved, achieving efficient model iteration updates and recommendation ranking optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU QINGMENG NETWORK TECHNOLOGY CO LTD
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-21
Smart Images

Figure CN122199117B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and internet information processing technology, specifically relating to a method and system for evaluating the quality and likelihood of closing deals in information-sharing posts based on AI technology. Background Technology
[0002] With the rapid development of internet content platforms, e-commerce shopping guide platforms, and community-based consumer decision-making platforms, the number of user-posted product deals, discount information posts, and shopping guides continues to grow. Deal posts typically refer to text content surrounding information such as product prices, promotional activities, purchase channels, inventory availability, target audience, and user reviews. Their core value lies in helping platform users quickly identify product information with purchase reference value and assisting the platform in content distribution and transaction conversion. Because deal posts possess both informational attributes and potential transaction guidance attributes, platforms often aim to assess their likelihood of generating sales while controlling content quality, prioritizing the display of higher-quality deal posts that are more likely to lead to transactions within limited traffic resources.
[0003] However, in practical applications, the content of "deal-breaking" posts is often from diverse sources, inconsistent in expression, and contains significant textual noise. Furthermore, different publishers exhibit marked differences in product naming, price descriptions, discount rules, and the presentation of timely information. On the one hand, while some "deal-breaking" posts are highly attractive, they may suffer from exaggerated titles, missing price information, incomplete discount conditions, insufficient timeliness, or questionable authenticity. On the other hand, some "deal-breaking" posts, while using relatively ordinary text, may possess high sales potential due to factors such as product category, discount strength, or the timing of user demand. Therefore, relying solely on manual review makes it difficult to achieve efficient, low-cost, and consistent filtering across a large volume of content. Conversely, relying solely on single behavioral metrics such as click-through rates and reading time can easily confuse "eye-catching but low-quality" content with "authentic but bland" content, making it difficult to simultaneously assess content quality and sales potential.
[0004] In the prior art, Chinese patent document CN111008278A discloses a content recommendation method and apparatus. Its technical approach includes: acquiring content to be classified and performing content recognition; selecting a corresponding content classification model for classification based on the recognition results; subsequently recalling the content based on a recall strategy; then ranking and secondary ranking according to the recommendation model and ranking algorithm model; and finally, reviewing the recalled content in conjunction with user feedback. This patent also explicitly points out that low-quality content can be filtered at the source to improve ranking accuracy, and that indicators such as user negative feedback, click-through rate, conversion rate, and reading time can be used for further review and processing of online content.
[0005] The aforementioned existing technologies are beneficial for recommending and ranking general information, text and image content, or video content, and can improve issues such as the posting of low-quality content, inaccurate recommendations, and decreased user experience to some extent. However, when these solutions are directly applied to the scenario of breaking news posts, the following shortcomings still exist: First, existing technologies mainly focus on content classification, recall, and ranking, paying more attention to recommendation accuracy and filtering of low-quality content, and do not adequately consider the strong coupling relationship between price, discounts, channels, timeliness, and transaction results unique to breaking news posts; Second, existing technologies are usually based on general content feedback indicators, making it difficult to coordinate the evaluation of the "information quality" and "likelihood of conversion" of breaking news posts, and therefore cannot finely distinguish between the two different objectives of "worth watching" and "worth recommending and likely to result in a sale"; Third, in the scenario of breaking news posts, training samples often suffer from problems such as high noise, inconsistent labeling, and sparse high-quality transaction samples. If general content classification or ranking training methods are still used, it is easy to cause model learning bias, which in turn affects the stability and usability of the evaluation results in actual sales guidance business.
[0006] Therefore, there is an urgent need for a new technical solution that can target the content object of "deals" – which has the characteristics of price discounts and transaction conversion – and can more accurately assess the likelihood of a sale while taking into account the quality of the post content. This solution can also be used to improve recommendation ranking and traffic allocation, thereby improving the distribution efficiency of "deals" and the overall conversion effect of the platform. Summary of the Invention
[0007] To overcome the problems of existing technologies, this paper provides a method and system for evaluating the quality and conversion probability of news tips posts based on AI technology. This method aims to accurately identify the content quality and transaction conversion potential of news tips posts, and improve the effectiveness of recommendation ranking and online conversion results.
[0008] To achieve the above-mentioned technical objectives, the present invention provides the following technical solution.
[0009] In a first aspect, this invention provides a method for evaluating the quality and likelihood of closing a deal based on AI technology-based tipping posts, comprising the following steps: S1. Obtain historical leaked post samples, extract post text, product entity information, price discount fields, as well as exposure, click, order and transaction data, and build the original sample pool; S2. The original sample pool is processed by text standardization, field alignment and behavior feature normalization. Based on expert experience rules and rule models, quality candidate labels and single candidate labels are generated respectively. Sample anomaly labels are generated by combining the preset overfitting discrimination index. Samples with conflict degree exceeding the preset threshold are removed by consistency comparison to obtain the mixed fine-tuning sample set. S3. For the target sample in the mixed fine-tuning sample set, according to the text semantic similarity and conversion behavior similarity, retrieve a preset number of related historical breaking posts from the historical breaking post library, and concatenate them with the target sample to form the training context. S4. Input the training context into the pre-trained model, and use the parameter efficient fine-tuning method to perform dual-task joint training of the quality level prediction task and the order probability prediction task to obtain the tip post evaluation model. S5. Input the post to be evaluated into the post evaluation model, output the quality level score and the probability of order completion score, and generate a comprehensive ranking adjustment value to send to the recommendation system to control the recall priority, ranking position or exposure quota. At the same time, collect the exposure, click, order and transaction feedback data after recommendation, write it back to the historical post library to update the mixed fine-tuning sample set and drive the model to iterative update.
[0010] Specifically, in step S1, the product entity information includes at least a brand field, a product field, and a specification field, and the price discount field includes at least a price field, a discount field, and a time field; The price field is used to represent one or more of the original price, the final price, and the price after coupon; the discount field is used to represent one or more of the coupon, the discount for purchases over a certain amount, the rebate, or the gift; and the time field is used to represent one or more of the posting time, the start time of the activity, and the end time of the activity.
[0011] Specifically, in step S2, the text standardization includes one or more of word segmentation, noise reduction, synonym merging, and entity recognition; the field alignment includes mapping product entity information and price discount fields with different expressions into a unified field; and the behavior feature normalization includes scaling the exposure, click, order, and transaction data.
[0012] Specifically, in step S2, the expert experience rules include at least one or more of the following: completeness rules, authenticity rules, timeliness rules, expression standardization rules, and risk rules; the rule model is used to output rule quality scores and rule success scores, and to generate the quality candidate labels and success candidate labels based on the rule quality scores and rule success scores, respectively.
[0013] Specifically, in step S2, the overfitting discrimination index includes at least one or more of validation loss bias, output confidence dispersion, and class boundary volatility; the conflict degree is used to characterize the degree of inconsistency between quality candidate labels, single candidate labels, and sample anomaly labels, and is determined based on one or more of label differences, confidence differences, and ranking differences.
[0014] Specifically, in step S3, the text semantic similarity is the degree of similarity between the target sample and the historical post in at least two of the fields of post text, product entity information, and price discount; the conversion behavior similarity is the degree of similarity between the target sample and the historical post in at least two of the fields of exposure rate, click rate, order rate, and conversion rate; and the number of associated historical posts is 3 to 20.
[0015] Specifically, in step S3, the training context is formed by splicing the target sample and the associated historical leaked posts in a preset order. The preset order includes sorting in descending order by text semantic similarity, sorting in descending order by conversion behavior similarity, or sorting in descending order by the fusion result of text semantic similarity and conversion behavior similarity.
[0016] Specifically, in step S4, the parameter fine-tuning method is any one of LoRA fine-tuning, QLoRA fine-tuning, or adapter fine-tuning; the dual-task joint training is performed by simultaneously optimizing the quality level prediction loss and the order probability prediction loss, wherein the quality level prediction loss is used to constrain the model's output error on the quality level score, and the order probability prediction loss is used to constrain the model's output error on the order probability score.
[0017] Specifically, in step S5, the comprehensive ranking adjustment value is obtained based on the quality level score, the order completion probability score, and the risk penalty value; when the risk penalty value exceeds the preset risk threshold, the corresponding post to be evaluated is subject to demotion, traffic restriction, or transfer to the manual review queue; the feedback data includes at least exposure data, click data, order data, and transaction data; after writing the feedback data back to the historical post database, posts to be evaluated that deviate from the predicted results to the actual transaction results by more than the preset deviation threshold are selected as difficult samples and added to the mixed fine-tuning sample set for subsequent model iteration updates.
[0018] Secondly, this invention also provides an evaluation system for the quality and likelihood of closing deals on leaked information posts based on AI technology, used to implement the evaluation method described in the first aspect, including: The sample construction module is used to obtain historical leaked post samples, extract post text, product entity information, price discount fields, as well as exposure, click, order and transaction data, and build the original sample pool. The standardization and hybrid annotation module is used to perform text standardization, field alignment and behavioral feature normalization on the original sample pool. It generates quality candidate labels and single candidate labels based on expert experience rules and rule models, respectively, and generates sample anomaly labels by combining preset overfitting discrimination index. Samples with conflict degree exceeding preset threshold are eliminated by consistency comparison to obtain a hybrid fine-tuned sample set. The similar sample retrieval module is used to retrieve a preset number of related historical posts from the historical post library based on text semantic similarity and conversion behavior similarity for the target sample in the mixed fine-tuning sample set, and then concatenate them with the target sample to form the training context. The model training module is used to input the training context into the pre-trained model and perform dual-task joint training of the quality level prediction task and the order probability prediction task using an efficient parameter fine-tuning method to obtain the tip-off post evaluation model. The online evaluation and update module is used to input the post to be evaluated into the post evaluation model, output the quality level score and the order completion probability score, generate a comprehensive ranking adjustment value and send it to the recommendation system, and collect the exposure, click, order and transaction feedback data after recommendation and write it back to the historical post library to update the mixed fine-tuning sample set and drive the model to iterative update.
[0019] This invention constructs an original sample pool containing post text, product entity information, price discount fields, and behavioral feedback data. It then combines expert experience rules, rule models, and overfitting criteria to perform consistency comparisons and conflict pruning on the samples, reducing the interference of noisy samples on the training boundary. Furthermore, during the model training phase, it introduces historical breakout posts similar to the target samples in textual semantics and conversion behavior as training context, enabling the model to learn more fully at local decision boundaries. Through joint training of a quality level prediction task and a conversion probability prediction task, it simultaneously characterizes the information quality and transaction conversion potential of posts. Based on this, the model's output quality level score and conversion probability score are used to generate a comprehensive ranking adjustment value and fed back to the recommendation system. Simultaneously, it combines online exposure, clicks, orders, and transaction data for closed-loop updates, thereby achieving continuous optimization and traffic allocation for high-value breakout posts. Therefore, this invention not only improves the accuracy and stability of breakout post quality assessment and conversion probability assessment but also enhances model convergence efficiency and online conversion effects. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the overall process of the method of the present invention.
[0021] Figure 2 This is a schematic diagram of the module composition of the system of the present invention.
[0022] Figure 3 This is a schematic diagram of the original sample pool construction and preprocessing process.
[0023] Figure 4 This is a schematic diagram of the process for generating a mixed fine-tuning sample set.
[0024] Figure 5 A schematic diagram illustrating the process of retrieving and training context for similar historical leaked posts.
[0025] Figure 6 A schematic diagram of the model training process for evaluating the tip-off post.
[0026] Figure 7 This is a schematic diagram illustrating the online evaluation, sorting control, and updating process for posts that are pending evaluation.
[0027] Figure 8 This is a schematic diagram of the feedback data write-back and hard sample update process.
[0028] Figure 9 A diagram showing the comparison of the prediction effects of different schemes.
[0029] Figure 10 This is a comparative diagram showing the change in prediction accuracy with training rounds under the condition of having and not having similar historical leaked posts.
[0030] Figure 11 This is a diagram illustrating the comparison of transaction results in online A / B testing. Detailed Implementation
[0031] The technical solution of the present invention will be further described in detail below with reference to the embodiments of the present invention. It should be understood that the following specific embodiments are only used to explain the present invention and are not intended to limit the scope of protection of the present invention. Conventional substitutions, equivalent transformations or combinations of the technical features made by those skilled in the art without departing from the technical concept of the present invention should all fall within the scope of protection of the present invention.
[0032] This invention provides a method and system for evaluating the quality and conversion probability of "discount posts" based on AI technology. It is applicable to e-commerce shopping guide platforms, content community platforms, discount aggregation platforms, consumer decision-making platforms, or other platform scenarios with a need for distributing product information posts. The "discount post" typically contains text content including product information, price information, discount information, purchase channels, timeliness information, or purchase suggestions. It possesses both content information attributes and a strong correlation with actual order placement, transaction completion, and other transaction behaviors. Therefore, this invention does not merely perform general text classification of discount posts, but simultaneously evaluates their content quality and conversion probability. The evaluation results are used for recall, ranking, and exposure control in the recommendation system to improve the platform's ability to identify high-value discount posts and the efficiency of traffic allocation.
[0033] refer to Figure 1The overall technical approach of this invention is as follows: First, historical breakout posts are collected and an original sample pool is constructed. Then, the original sample pool is processed by text standardization, field alignment, and behavioral feature normalization, and a hybrid fine-tuning sample set is constructed based on expert experience rules, rule models, and sample anomaly identification mechanisms. Subsequently, for target samples in the hybrid fine-tuning sample set, historical breakout posts that are similar to them in terms of text semantics and conversion behavior are retrieved and used as training context input to the pre-trained model. The quality level prediction task and the order completion probability prediction task are jointly trained using a parameter-efficient fine-tuning method to obtain a breakout post evaluation model. After that, the breakout post to be evaluated is input into the breakout post evaluation model, which outputs a quality level score and an order completion probability score, and generates a comprehensive ranking adjustment value to be sent to the recommendation system. Finally, the exposure, click, order, and transaction feedback data after online recommendation are collected, feedback tags are formed, and written back to the historical breakout post library to update the hybrid fine-tuning sample set and drive the model to iteratively update, thereby forming a continuously self-optimizing closed-loop evaluation system.
[0034] I. System Composition refer to Figure 2 In one embodiment, the system of the present invention includes a sample construction module, a standardization and hybrid annotation module, a similar sample retrieval module, a model training module, and an online evaluation and update module.
[0035] The sample construction module is used to obtain historical leaked post samples, extract post text, product entity information, price discount fields, and exposure, click, order and transaction data to build the original sample pool.
[0036] The standardization and hybrid annotation module is used to perform text standardization, field alignment and behavioral feature normalization on the original sample pool. It generates quality candidate labels and single candidate labels based on expert experience rules and rule models, respectively, and generates sample anomaly labels by combining preset overfitting discrimination index. Samples with conflict degree exceeding preset threshold are eliminated by consistency comparison to obtain a hybrid fine-tuned sample set.
[0037] The similar sample retrieval module is used to retrieve a preset number of related historical posts from the historical post library based on text semantic similarity and conversion behavior similarity for the target sample in the mixed fine-tuning sample set, and then concatenate them with the target sample to form the training context.
[0038] The model training module is used to input the training context into the pre-trained model and perform dual-task joint training of the quality level prediction task and the order probability prediction task using an efficient parameter fine-tuning method to obtain the tip-off post evaluation model.
[0039] The online evaluation and update module is used to input the post to be evaluated into the post evaluation model, output the quality level score and the probability of order completion score, generate a comprehensive ranking adjustment value and send it to the recommendation system, and collect the exposure, click, order and transaction feedback data after recommendation and write it back to the historical post library to update the mixed fine-tuning sample set and drive the model to iterative update.
[0040] In practical deployment, the above modules can be deployed on the same server or deployed separately using a distributed service approach. The sample construction module, standardization and hybrid annotation module, and model training module are preferably deployed in an offline training environment, while the online evaluation and update module is preferably deployed in an online service environment.
[0041] II. Implementation Process 2.1 Step S1: Construct the original sample pool refer to Figure 3 In this invention, historical leaked post samples are first obtained, and information related to content quality assessment and order completion assessment is extracted for each historical leaked post to construct an original sample pool.
[0042] In one implementation, the extracted information includes: post text, product entity information, price discount field, and exposure, click, order and transaction data.
[0043] The post text may include a title, body, summary, tags, or notes; product entity information may include brand, product name, specifications, capacity, applicable scenarios, or category; the price discount field may include original price, final price, coupon information, discount rules, rebate information, gift information, limited-time event information, and channel information; behavioral data is used to characterize user feedback after the post goes live. For consistent representation, the price discount field may also include a time field to characterize one or more of the following: posting time, event start time, and event end time.
[0044] To ensure sample availability, the samples in the original sample pool are preferably derived from real online data within a preset time window. This preset time window can be the last 7 days, 30 days, 90 days, 180 days, or longer. Those skilled in the art can rationally select the time window based on business cycles, activity cycles, and product category characteristics.
[0045] Step S1 mainly realizes the unified collection and organization of historical data. The specific fields extracted can be added or removed according to the business scenario, but at least they should be able to support subsequent text quality judgment and conversion behavior evaluation.
[0046] 2.2 Step S2: Generate a mixed fine-tuning sample set This step's process is as follows: Figure 4As shown, the purpose of this step is to generate a high-quality sample set from the original sample pool that is more suitable for fine-tuning the model, so as to reduce the interference of noisy samples, inconsistent label samples and abnormal samples on the training process.
[0047] 2.2.1 Data Preprocessing In one embodiment, the following preprocessing operation is performed on the original sample pool: (1) Text standardization.
[0048] The system performs word segmentation, noise reduction, synonym merging, and entity recognition on post text. Noise reduction includes removing meaningless characters, duplicate symbols, residual webpage tags, garbled text, or obviously irrelevant text. Synonym merging unifies fields with different expressions but the same meaning, such as merging "final price," "discounted price," and "actual price" into a single price field. Entity recognition identifies information such as brand, product, specifications, price, discount threshold, channel, and timeliness.
[0049] (2) Field alignment.
[0050] Map inconsistent information from different posts to a unified field. For example, map “JD self-operated”, “JD self-operated”, and “JD official self-operated” to the same channel field, and map “500ml”, “500 ml”, and “500 mL” to the same specification field.
[0051] (3) Normalization of behavioral characteristics.
[0052] Exposure, click, order, and transaction data are standardized in terms of scale to reduce the differences in numerical dimensions caused by different traffic locations, time windows, and recommendation positions. Normalization methods can include one or more of the following: min-max normalization, Z-score standardization, logarithmic transformation, or bucketing mapping.
[0053] 2.2.2 Generation of Quality Candidate Tags and Order Candidate Tags After preprocessing, the system generates quality candidate labels and order candidate labels based on expert experience rules and rule models, respectively.
[0054] In one implementation, the expert experience rules include at least one or more of the following rules: Completeness rules are used to determine whether a post contains at least a preset number of information items from the categories of product, price, discount, channel, and timeliness. The authenticity rule is used to determine whether there are obvious conflicts or anomalies between the price, discount conditions, and product description; The timeliness rule is used to determine whether the activity described in the post is still valid, or whether there are expired activities that have not been marked. Expression standards are used to determine whether the title matches the body text and whether there are obvious exaggerations or misleading expressions. Risk rules are used to determine whether there are risks such as false traffic generation, illegal traffic redirection, missing information, or abnormally low prices.
[0055] The rule model can be a traditional supervised model trained on existing historical data, used to output rule quality scores and rule success scores. The rule model can employ logistic regression models, tree models, shallow neural network models, or other models suitable for structured feature inputs.
[0056] Quality candidate tags are preferably used to characterize the quality level of a post, and can be in the form of binary, tri-class, or multi-classification; conversion candidate tags are preferably used to characterize the subsequent conversion potential of a post, and can also be in the form of binary, tiered, or probability interval division.
[0057] 2.2.3. Generation of Sample Anomaly Markers To further improve the quality of the sample set, this invention introduces a sample anomaly labeling mechanism based on an overfitting discrimination index. The sample anomaly labeling is used to identify samples that exhibit significant instability, contradictory labels, or semantic-behavioral inconsistencies during training.
[0058] In one embodiment, the overfitting discrimination metric includes at least one or more of validation loss bias, output confidence dispersion, and class boundary volatility.
[0059] Among them, validation loss bias can represent the degree of difference in the loss of a certain sample in different rounds of training or different validation trade-offs; Output confidence dispersion can represent the degree of fluctuation in the model's output confidence for a given sample during multiple training processes; Class boundary volatility can represent the degree to which the class to which a sample belongs changes frequently during multiple rounds of training.
[0060] The system generates sample anomaly labels based on the overfitting discrimination index. For example, when a historical post shows frequent jumps between high and low quality in multiple training rounds, or when its textual semantics are obviously complete but its transaction behavior is extremely abnormal and cannot be explained by traffic, it can be labeled as an anomaly sample.
[0061] 2.2.4 Consistency Comparison and Pruning After obtaining the quality candidate label, the single candidate label, and the sample anomaly label, the consistency of the three is compared, and sample pruning is performed according to the degree of conflict.
[0062] In one embodiment, the conflict degree is used to characterize the degree of inconsistency between quality candidate labels, single candidate labels, and sample anomaly labels, and the conflict degree may be determined based on one or more of label differences, confidence differences, and ranking differences.
[0063] When the conflict level of a sample exceeds a preset threshold, it is either removed from the main training sample set or added to a pool of samples awaiting review, after which a decision is made on whether to remove it. This method yields a mixed fine-tuning sample set with lower noise, higher label consistency, and suitability for fine-tuning large models.
[0064] Step S2 is not simply data cleaning, but rather an optimization of training sample quality by combining multi-source label generation with outlier sample identification, thereby improving the generalization ability and stability of the subsequent model.
[0065] 2.3 Step S3: Constructing the training context The process of step S3 is as follows: Figure 5 As shown, its core lies in the fact that instead of directly inputting the target sample into the model for fine-tuning, it retrieves several historical posts similar to the target sample and uses them together to form the training context, thereby enhancing the model's ability to learn the decision boundaries of the samples during the training phase.
[0066] In one embodiment, for each target sample in the mixed fine-tuning sample set, the system retrieves a preset number of related historical breaking news posts from the historical breaking news post library according to text semantic similarity and conversion behavior similarity.
[0067] 2.3.1 Text Semantic Similarity Textual semantic similarity is used to characterize the degree of similarity between the target sample and historical samples in at least two of the fields of post text, product entity information, and price discount.
[0068] In one implementation, the system can convert the target sample and historical samples into vector representations respectively, and calculate their textual semantic similarity using cosine similarity, dot product similarity, or other similarity measures.
[0069] 2.3.2 Similarity of conversion behavior Conversion behavior similarity is used to characterize the similarity between target samples and historical samples in at least two of the following: exposure rate, click-through rate, order rate, and conversion rate.
[0070] For target samples that do not yet have real online behavior, prior behavioral features can be constructed based on their product entity information, price discount fields, release time, channel information, and text structure, and then compared with historical samples.
[0071] 2.3.3 Training Context Construction After completing the similar sample retrieval, the associated historical posts and the target sample are concatenated to form the training context.
[0072] Preferably, the number of related historical posts is 3 to 20, more preferably 5 to 10.
[0073] The splicing order can include sorting by text semantic similarity in descending order, sorting by conversion behavior similarity in descending order, or sorting by the fusion result of the two in descending order.
[0074] In practical implementation, a structured template can be used, where the target sample and its associated historical breaking news posts are organized according to the format of "text content - structured fields - tag information" before being input into the model. By introducing a small number of similar historical breaking news posts during the training phase, this invention can provide the model with a local context that is closer to the business decision boundary, which helps to improve the model's training convergence efficiency and task discrimination ability.
[0075] 2.4 Step S4: Training the evaluation model for leaked posts Step S4 is used to train the evaluation model for leaked posts based on the training context, and its process is as follows: Figure 6 As shown.
[0076] 2.4.1 Overall Model Structure In one embodiment, the tip-off evaluation model includes at least an input layer, an encoding layer, a feature fusion layer, a task output layer, and a parameter fine-tuning layer.
[0077] (1) Input layer This layer is used to receive the training context. The training context includes at least the post text of the target sample, product entity information, price discount field, and corresponding information from several related historical posts. The input layer can convert the text content into a sequence of lexical units or subwords, and convert structured fields into preset tags, field embeddings, or discrete identifiers, which, together with the text sequence, form the model input.
[0078] (2) Coding layer This layer is used to semantically encode the text sequence and structured fields of the input layer, outputting a contextual representation. The encoding layer preferably employs a pre-trained model's backbone network structure to learn the contextual dependencies between the target sample and associated historical samples. The pre-trained model can be a language model with a self-attention mechanism, or any other pre-trained network capable of contextual modeling of sequence information.
[0079] (3) Feature fusion layer This is used to fuse the target sample representation, associated historical leak post representation, and structured field representation output by the encoding layer to form a shared feature representation for subsequent task prediction. The fusion method can be concatenation, weighted summation, attention aggregation, gated fusion, or a combination thereof. Preferably, the feature fusion layer outputs a shared vector representation of uniform length, which can be used by both the quality level prediction task and the order probability prediction task.
[0080] (4) Task output layer It includes a quality level prediction sublayer and a success probability prediction sublayer. The quality level prediction sublayer outputs the quality level score; the success probability prediction sublayer outputs the success probability score. The quality level prediction sublayer can be implemented using a classification head or a regression head; the success probability prediction sublayer can use binary classification probability output or probability regression output. The two sublayers share the main parameters of the encoding layer and the feature fusion layer, but each has independent output parameters.
[0081] (5) Parameter High-Efficiency Fine-Tuning Layer This is used to introduce trainable adaptive parameters at certain layers of a pre-trained model, while keeping all or most of the backbone parameters frozen, thereby reducing training resource consumption. The parameter-efficient fine-tuning layer is preferably set in the attention-related layer, projection layer, or intermediate transformation layer within the encoding layer, enabling the model to adapt to the scenario of breaking news posts with limited incremental parameters.
[0082] With the above structure, the model of the present invention is not a simple black-box scoring model, but includes at least an input layer for receiving training context, an encoding layer for semantic modeling, a feature fusion layer for constructing shared representations, and an output layer corresponding to the two tasks, so that those skilled in the art can complete the model implementation based on the contents disclosed in the specification.
[0083] 2.4.2. Basic Model and Efficient Parameter Fine-Tuning Method In this invention, the evaluation model for leaked posts preferably uses a pre-trained model as the base model and is trained through efficient parameter fine-tuning.
[0084] Among them, the efficient parameter fine-tuning method can be at least any one of LoRA fine-tuning, QLoRA fine-tuning, or adapter fine-tuning.
[0085] In implementations employing LoRA fine-tuning, it is preferable to introduce low-rank adaptation parameters in the attention-related layer of the coding layer to achieve domain adaptation while keeping the backbone parameters frozen.
[0086] In implementations employing QLoRA fine-tuning, memory usage can be further reduced through low-bit quantization.
[0087] In an implementation using adapter fine-tuning, lightweight adapter modules can be inserted between several transformation layers of the coding layer to achieve feature transfer related to the leaked post scenario.
[0088] Among the aforementioned efficient parameter fine-tuning methods, LoRA fine-tuning refers to a method that efficiently fine-tunes parameters by introducing low-rank trainable parameters into some network layers of a pre-trained model to adapt the base model to the task, while keeping all or most of the backbone parameters unchanged. This method allows for model training in specific scenarios with fewer new parameters, thereby reducing training costs and storage overhead.
[0089] QLoRA fine-tuning refers to a parameter fine-tuning method that further combines LoRA fine-tuning with low-bit quantization technology. It reduces memory usage and computational resource consumption by quantizing and storing the basic model parameters and using low-rank trainable parameters for task adaptation. This method is suitable for model training and deployment under conditions of limited computing resources.
[0090] Adapter fine-tuning refers to inserting adapter modules between several network layers of a pre-trained model, and training these adapter modules to adapt the model to a specific task, while keeping all or most of the backbone model parameters unchanged. This approach can improve the model's adaptability to specific domain tasks without significantly altering the original model structure.
[0091] This invention does not limit the specific type and parameter size of the base model, as long as the selected model can model the training context and output the quality level prediction result and the order probability prediction result.
[0092] 2.4.3 Dual-Task Joint Training The key point of this step is: First, the model performs joint training on two tasks; Second, the dual-task joint training shall include at least a quality level prediction task and a single-order probability prediction task.
[0093] In one embodiment, the quality grade prediction task is used to output the quality grade score of the tip post, and the order completion probability prediction task is used to output the order completion probability score of the tip post.
[0094] Quality level prediction tasks can be implemented using either classification or regression methods; order probability prediction tasks can be implemented using either binary classification probability output or probability regression methods.
[0095] During model training, both the quality level prediction loss and the order conversion probability prediction loss are optimized simultaneously to obtain a news tip evaluation model that balances content quality judgment and transaction conversion judgment capabilities. For the loss function form, learning rate, batch size, training epochs, low-rank matrix rank, and dropout parameters, those skilled in the art can make reasonable selections based on existing fine-tuning techniques and computing power, as long as the above dual-task joint training can be completed and a usable news tip evaluation model is obtained.
[0096] 2.5 Step S5: Online evaluation, ranking control and updating The process for this step is as follows: Figure 8 As shown; after training is completed, the post to be evaluated is input into the post evaluation model, which outputs the quality level score and the probability of closing a deal score, and generates a comprehensive ranking adjustment value accordingly.
[0097] In one implementation, the overall ranking adjustment value is obtained based on the quality level score, the probability of closing a deal score, and the risk penalty value.
[0098] The risk penalty value is used to characterize the risk of authenticity, timeliness, or missing information in the post being evaluated.
[0099] When the risk penalty value exceeds the preset risk threshold, the corresponding post will be subject to demotion, traffic restriction, or transfer to the manual review queue.
[0100] Subsequently, the overall ranking adjustment value is sent to the recommendation system to control the recall priority, ranking position, or exposure quota of the corresponding tip post.
[0101] For example, for posts with high quality ratings, high likelihood of closing a deal, and low risk penalties, their recall priority and ranking can be increased; for posts with low quality ratings or high risk penalties, their exposure quota can be reduced or manual review can be required.
[0102] In this invention, step S5 includes not only online evaluation and ranking control, but also feedback data collection and model updating. Specifically, the system collects feedback data on the exposure, clicks, orders, and transactions of the tip-off post after it has been recommended, and writes the feedback data back to the historical tip-off post library to update the hybrid fine-tuning sample set and drive iterative model updates.
[0103] Preferably, the feedback data can be further used to generate exposure feedback tags, click feedback tags, order feedback tags, and transaction feedback tags.
[0104] More preferably, the system prioritizes selecting posts that deviate from the predicted results and the actual transaction results by more than a preset deviation threshold as difficult samples to be added to the mixed fine-tuning sample set for subsequent iterative updates.
[0105] For example, if a post predicts a high probability of a sale but the actual transaction is low, the post can be marked as a biased sample and included in the hard sample set; conversely, if a post predicts a moderate probability but the actual transaction is high, it can also be included in the hard sample set.
[0106] By writing back such difficult samples and participating in the next round of training, the model's ability to adapt to boundary samples and new types of samples can be gradually improved.
[0107] Thus, this invention forms a complete closed-loop technical solution from sample construction, sample optimization, similar sample enhancement training, dual-task evaluation, ranking linkage to feedback update.
[0108] III. Preferred Implementation Methods Without affecting the implementation of the present invention, the following can be used as preferred embodiments: Firstly, efficient parameter fine-tuning preferably adopts the LoRA fine-tuning method, and it is preferable to introduce low-rank adaptation parameters in the attention-related layers of the pre-trained model to reduce training resource consumption.
[0109] Secondly, the number of related historical leak posts is preferably 5 to 10, in order to balance the richness of context and training efficiency.
[0110] Third, the fusion of text semantic similarity and conversion behavior similarity can be achieved by weighted summation, and the weights can be selected based on the validation set results.
[0111] Fourth, independent sample pools can be established for different product categories or a category-based training strategy can be adopted to improve the model's ability to identify features in vertical domains.
[0112] Fifth, when providing online services, the evaluation model for whistleblower posts can be linked with the manual review system to trigger a manual review process for high-risk posts.
[0113] Sixth, in the implementation of the model structure, the feature fusion layer preferably adopts a combination of splicing and attention aggregation to take into account both the local features of the target sample and the contextual features of the associated historical samples.
[0114] Seventh, in the task output layer, the quality level prediction sub-layer and the order probability prediction sub-layer preferably share the parameters of the encoding layer and the feature fusion layer, and adopt independent output heads to improve the training efficiency of multi-task. Specific Implementation
[0115] The technical effects of the present invention will be described below with reference to specific embodiments. It should be noted that the following embodiments are used to demonstrate the technical effectiveness of the present invention compared to existing solutions in assessing the quality of promotional posts and the likelihood of order completion, and are not intended to limit the scope of protection of the present invention. Those skilled in the art can make conventional adjustments to the sample size, product category, model base, training rounds, and evaluation indicators based on the concept of the present invention.
[0116] 4.1 Example 1: Verification of the effect of predicting the quality grade and order probability of liquor category information leaks. 4.1.1 Experimental Objective The technical solution described in this invention demonstrates an improvement in the effectiveness of judging the quality level of leaked posts and the probability of orders being placed, compared to the general large-scale model direct prediction solution and the traditional rule model solution.
[0117] 4.1.2 Data Sources and Sample Composition The experimental sample consisted of 180 days' worth of posts about liquor products on a certain e-commerce platform. From the original sample, the post text, product information, price discount fields, posting time, historical exposure, clicks, order volume, and transaction results were extracted. These were then processed according to the method described in this invention, including text standardization, field alignment, and behavioral feature normalization. Finally, consistency comparison and pruning were performed based on expert experience rules, rule models, and sample anomaly markers to form a hybrid fine-tuned sample set.
[0118] In this embodiment, the sample division is shown in Table 1.
[0119] Table 1. Composition of the Sample Set
[0120] 4.1.3 Control Group Setup To verify the technical effect of the present invention, the following three schemes are set up: (1) Control group A: Traditional rule model scheme The quality rating and order completion result are output based solely on price rules, discount rules, timeliness rules, and historical statistics rules.
[0121] (2) Control group B: General large model direct prediction scheme The text of the leaked post and the structured fields are directly input into a general large model for prediction. Mixed sample pruning is not used, the context enhancement is not trained using similar historical leaked posts, and dual-task joint training is not used.
[0122] (3) Experimental Group 1: Scheme of the present invention The technical approach described in this invention is as follows: Original sample pool preprocessing → Mixed fine-tuning sample set generation → Retrieve related historical breaking news posts and construct training context during training phase → Efficient parameter fine-tuning → Joint training of quality level prediction task and order completion probability prediction task → Online evaluation and ranking linkage → Feedback writing and iterative update.
[0123] 4.1.4 Evaluation Indicators The following evaluation metrics are used in this embodiment: Accuracy of quality grade prediction; Accuracy of order completion probability prediction; The actual order generation probability of high-quality, high-success-rate samples; High-value post identification and recall rate; Number of transactions per thousand exposures.
[0124] Among them, the accuracy of quality grade prediction is used to measure the model's ability to judge the quality grade of posts; the accuracy of conversion probability prediction is used to measure the model's ability to judge whether a post will become a sale in the future; and the actual conversion probability of high-quality, high-sales samples is used to characterize the conversion ability of high-potential posts screened by the model in real traffic.
[0125] 4.1.5 Experimental Results The predictive performance indicators for different schemes are shown in Table 2 and Figure 9 As shown.
[0126] Table 2 Comparison of prediction performance of different schemes on the test set
[0127] From Table 2 and Figure 9 As can be seen, compared with the general large-scale model direct prediction scheme, the present invention improves the accuracy of quality level prediction from 15.0% to 65.0%, an improvement of 333.3%; the accuracy of order completion probability prediction from 21.4% to 64.6%, an improvement of 201.9%; and the actual order probability of high-quality, high-order samples from 19.4% to 43.8%. This indicates that the present invention effectively improves the model's ability to identify the quality and transaction potential of breaking news posts through mixed sample pruning and the injection of similar historical breaking news posts during the training phase. In addition, compared with the traditional rule-based model scheme, the present invention also shows better performance in terms of high-value post identification recall and number of transactions per thousand impressions, indicating that the present invention not only improves offline prediction accuracy but also brings about an improvement in actual transaction results in online traffic scenarios.
[0128] 4.2 Example 2: Verification of the impact of mixed sample pruning on model performance 4.2.1 Experimental Objective The study verifies the impact of the consistency comparison and pruning mechanism of "quality candidate labels, single candidate labels, and sample anomaly markers" in this invention on the model training effect.
[0129] 4.2.2 Experimental Setup Under the conditions of the same pre-trained model, the same number of training rounds, and the same test set, the following two groups are set up: Control group C: No conflict sample pruning was performed; Experimental group D: Employs the conflict sample pruning mechanism described in this invention.
[0130] The experimental results are shown in Table 3.
[0131] Table 3. Impact of mixed-sample pruning on model performance
[0132] As shown in Table 3, without changing the basic model, the accuracy of quality grade prediction and order probability prediction are significantly improved after adopting the hybrid sample pruning mechanism of the present invention, and the fluctuation of validation set loss is significantly reduced. This indicates that the present invention reduces the interference of noise labels on the model training boundary by removing abnormal samples with high conflict, thereby improving training stability and generalization performance.
[0133] 4.3 Example 3: Verification of the impact of injecting similar historical leaked posts during the training phase on convergence speed and prediction performance 4.3.1 Experimental Objective This study verifies the effect of the technique of "retrieving similar historical leaked posts and constructing training context during the training phase" in this invention on improving the model's convergence speed and prediction performance.
[0134] 4.3.2 Experimental Setup Under the same base model, the same mixed fine-tuning sample set, and the same dual-task joint training settings, the following two groups are set up: Control group E: Training was performed using only the target samples; Experimental group F: Using the scheme of this invention, 8 related historical breaking news posts were retrieved for each target sample and spliced together to form the training context.
[0135] The experimental results for training context enhancement are shown in Table 4 and Figure 10 As shown.
[0136] Table 4. Impact of Training Context Augmentation on Convergence Speed and Model Performance
[0137] As shown in Table 4, although the present invention slightly increases the training time per round due to the introduction of similar historical leaked posts, the model achieves higher prediction accuracy in fewer rounds, resulting in higher overall training convergence efficiency. This indicates that introducing historical leaked posts similar to the target sample as training context into the training phase can enhance the model's ability to learn local decision boundaries, thereby improving convergence speed and task recognition performance.
[0138] 4.4 Example 4: Online A / B Testing Verification 4.4.1 Experimental Objective Online A / B testing: In a real business operation environment, the test objects or user traffic are divided into at least two groups, and different technical solutions are used to run them. By comparing the differences between the groups in terms of exposure, clicks, orders, transactions, or other business metrics, the actual effectiveness of different technical solutions is verified. This embodiment is to verify the real business effect of the present invention in an online recommendation and ranking scenario.
[0139] 4.4.2 Experimental Setup Newly released informational posts about liquor products were selected as test subjects, and A / B testing was conducted continuously for 14 days: Control group G: The original rule model ranking strategy was used; Experimental group H: The comprehensive ranking adjustment value output by the tip-off evaluation model of this invention is used to control the recall priority and ranking position.
[0140] The results of the online A / B test are shown in Table 5 and Figure 11 As shown.
[0141] Table 5 Online A / B Test Results
[0142] As can be seen, this invention can increase the exposure ratio of high-quality posts and reduce the exposure ratio of low-quality, misleading posts in real online scenarios, and significantly improve click-through rate, order rate, and conversion rate. This demonstrates that this invention can not only improve prediction accuracy in offline experiments, but also effectively transform model evaluation results into online business revenue through ranking collaboration.
[0143] As can be seen from Examples 1 to 4, the present invention has at least the following technical effects: By using a hybrid sample pruning mechanism, the consistency of training samples is improved, and the interference of abnormal samples on the training boundary is reduced, thereby improving the stability and generalization ability of the model. By introducing similar historical leaked posts into the training phase to construct a training context, the model's ability to learn local decision boundaries is improved, and training convergence is accelerated. The system employs a dual-task joint training approach, combining quality level prediction and order conversion probability prediction tasks, while simultaneously considering both post content quality and transaction conversion goals. By using the model output for recommendation ranking control and combining it with online feedback for closed-loop updates, we can achieve continuous optimization and traffic allocation for high-value breaking news posts.
[0144] Therefore, compared with the prior art, the present invention can achieve better prediction accuracy and actual online conversion effect in scenarios such as quality assessment of information post and order conversion probability assessment.
[0145] The above is a description of the embodiments and examples of the present invention. Various modifications to these embodiments and examples will be readily apparent to those skilled in the art. The general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention should not be limited to the embodiments and examples shown herein, but should be accorded the widest scope of protection consistent with the principles and novelty disclosed herein.
Claims
1. A method for evaluating the quality and likelihood of conversion of leaked information posts based on AI technology, characterized in that, Includes the following steps: S1. Obtain historical leaked post samples, extract post text, product entity information, price discount fields, as well as exposure, click, order and transaction data, and build the original sample pool; S2. The original sample pool is processed by text standardization, field alignment and behavior feature normalization. Based on expert experience rules and rule models, quality candidate labels and single candidate labels are generated respectively. Sample anomaly labels are generated by combining the preset overfitting discrimination index. Samples with conflict degree exceeding the preset threshold are removed by consistency comparison to obtain the mixed fine-tuning sample set. S3. For the target sample in the mixed fine-tuning sample set, according to the text semantic similarity and conversion behavior similarity, retrieve a preset number of related historical breaking posts from the historical breaking post library, and concatenate them with the target sample to form the training context. S4. Input the training context into the pre-trained model, and use the parameter efficient fine-tuning method to perform dual-task joint training of the quality level prediction task and the order probability prediction task to obtain the tip post evaluation model. S5. Input the post to be evaluated into the post evaluation model, output the quality level score and the probability of order completion score, and generate a comprehensive ranking adjustment value to send to the recommendation system to control the recall priority, ranking position or exposure quota. At the same time, collect the exposure, click, order and transaction feedback data after recommendation, write it back to the historical post library to update the mixed fine-tuning sample set and drive the model to iterative update.
2. The evaluation method according to claim 1, characterized in that, In step S1, the product entity information includes at least a brand field, a product field, and a specification field, and the price discount field includes at least a price field, a discount field, and a time field; The price field is used to represent one or more of the original price, the final price, and the price after coupon; the discount field is used to represent one or more of the coupon, the discount for purchases over a certain amount, the rebate, or the gift; and the time field is used to represent one or more of the posting time, the start time of the activity, and the end time of the activity.
3. The evaluation method according to claim 1, characterized in that, In step S2, the text standardization includes one or more of word segmentation, noise reduction, synonym merging, and entity recognition; the field alignment includes mapping product entity information and price discount fields with different expressions into a unified field; and the behavior feature normalization includes scaling the exposure, click, order, and transaction data.
4. The evaluation method according to claim 1, characterized in that, In step S2, the expert experience rules include at least one or more of the following: completeness rules, authenticity rules, timeliness rules, expression norms rules, and risk rules. The rule model is used to output rule quality score and rule success score, and to generate the quality candidate label and success candidate label based on the rule quality score and rule success score, respectively.
5. The evaluation method according to claim 1, characterized in that, In step S2, the overfitting discrimination index includes at least one or more of the following: validation loss bias, output confidence dispersion, and class boundary volatility. The conflict degree is used to characterize the degree of inconsistency between quality candidate labels, single candidate labels and sample anomaly labels, and is determined based on one or more of label differences, confidence differences and ranking differences.
6. The evaluation method according to claim 1, characterized in that, In step S3, the text semantic similarity is the degree of similarity between the target sample and the historical post in at least two of the fields of post text, product entity information, and price discount. The conversion behavior similarity is the degree of similarity between the target sample and historical posts in at least two of the following: exposure rate, click-through rate, order rate, and conversion rate. The number of related historical posts is 3 to 20.
7. The evaluation method according to claim 1, characterized in that, In step S3, the training context is formed by splicing the target sample and the associated historical post in a preset order. The preset order includes sorting in descending order by text semantic similarity, sorting in descending order by conversion behavior similarity, or sorting in descending order by the fusion result of text semantic similarity and conversion behavior similarity.
8. The evaluation method according to claim 1, characterized in that, In step S4, the efficient fine-tuning method for the parameters is any one of LoRA fine-tuning, QLoRA fine-tuning, or adapter fine-tuning; The dual-task joint training is performed by simultaneously optimizing the quality level prediction loss and the order probability prediction loss. The quality level prediction loss is used to constrain the model's output error on the quality level score, and the order probability prediction loss is used to constrain the model's output error on the order probability score.
9. The evaluation method according to claim 1, characterized in that, In step S5, the comprehensive ranking adjustment value is obtained based on the quality level score, the order completion probability score, and the risk penalty value; When the risk penalty value exceeds the preset risk threshold, the corresponding post to be evaluated will be subject to demotion, traffic restriction, or transfer to the manual review queue. The feedback data includes at least exposure data, click data, order data, and transaction data; After writing the feedback data back to the historical leak post library, leak posts to be evaluated that deviate from the predicted results and the actual transaction results by more than a preset deviation threshold are selected as difficult samples and added to the mixed fine-tuning sample set for subsequent model iteration updates.
10. A system for evaluating the quality and likelihood of closing deals on leaked information posts based on AI technology, used to implement any one of the evaluation methods described in claims 1-9, characterized in that, include: The sample construction module is used to obtain historical leaked post samples, extract post text, product entity information, price discount fields, as well as exposure, click, order and transaction data, and build the original sample pool. The standardization and hybrid annotation module is used to perform text standardization, field alignment and behavioral feature normalization on the original sample pool. It generates quality candidate labels and single candidate labels based on expert experience rules and rule models, respectively, and generates sample anomaly labels by combining preset overfitting discrimination index. Samples with conflict degree exceeding preset threshold are eliminated by consistency comparison to obtain a hybrid fine-tuned sample set. The similar sample retrieval module is used to retrieve a preset number of related historical posts from the historical post library based on text semantic similarity and conversion behavior similarity for the target sample in the mixed fine-tuning sample set, and then concatenate them with the target sample to form the training context. The model training module is used to input the training context into the pre-trained model and perform dual-task joint training of the quality level prediction task and the order probability prediction task using an efficient parameter fine-tuning method to obtain the tip-off post evaluation model. The online evaluation and update module is used to input the post to be evaluated into the post evaluation model, output the quality level score and the order completion probability score, generate a comprehensive ranking adjustment value and send it to the recommendation system, and collect the exposure, click, order and transaction feedback data after recommendation and write it back to the historical post library to update the mixed fine-tuning sample set and drive the model to iterative update.