A human-computer cooperation implicit sponsorship content detection and intervention method fusing uncertainty perception

By integrating a human-computer collaboration framework that integrates uncertainty perception, and utilizing a large language model and a calibrated support vector machine classifier, the problem of multimodal feature fusion and user intervention in the detection of implicitly sponsored content on social media is solved, achieving more efficient and interpretable detection and intervention results.

CN122634166APending Publication Date: 2026-08-25DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610926929.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing technologies for detecting implicitly sponsored content on social media suffer from problems such as identifying illusions, overgeneralization, uncalibrated confidence levels, and unexplainable decisions. Furthermore, multimodal features are difficult to fuse effectively, resulting in low detection accuracy, poor generalization ability, low efficiency in utilizing human review resources, and difficulty in translating detection results into intervention prompts that users can understand.

Method used

A human-computer collaboration framework integrating uncertainty perception is adopted. Multimodal semantic features are extracted using a pre-trained large-scale language model. Combined with a calibrated support vector machine classifier and a hybrid query strategy, iterative learning is carried out through human review feedback to generate reliable posterior probability outputs. These outputs are then converted into probabilistic disclosure labels to activate informed decision-making by users.

Benefits of technology

It enables more reliable and interpretable detection and intervention of implicit sponsored content under limited annotation resources, improves detection accuracy and user acceptance, optimizes the utilization of manual review resources, and provides a reliable multimodal feature fusion and user-friendly intervention mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122634166A_ABST
    Figure CN122634166A_ABST
Patent Text Reader

Abstract

The application discloses a kind of fusion uncertainty perception man-machine cooperation implicit sponsorship content detection and intervention method, belong to natural language processing field.It includes: using large language model to extract post text depth semantic embedding and image visual features, and with platform metadata, content quality features and sentiment features to build unified structured feature representation;Through the calibration support vector machine classifier with category imbalance penalty weight, learn the optimal classification hyperplane between implicit sponsorship content and natural content, reliable posterior probability is output using Platt scaling;Based on the mixed query strategy of entropy sampling and bayesian inconsistency, the most informative samples are handed over to manual review under limited labeling budget, and manual feedback is continuously converted into model retraining signal;Posterior probability is converted into probabilistic disclosure label and is shown to users.The application realizes more reliable, more interpretable and more suitable for platform governance scene implicit soft advertising detection and probabilistic disclosure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of computer vision, natural language processing, and social media content governance, and in particular to an automatic detection and intervention method for implicitly sponsored content that integrates multimodal feature extraction from a large language model (LLM), calibrated support vector machine classification, uncertainty-driven hybrid query strategy, and probabilistic user intervention under a limited annotation budget constraint. Background Technology

[0002] With the rapid development of social media platforms, user-generated content (UGC) has become a crucial channel for consumers to obtain product information and make purchasing decisions. Benefiting from the inherent sense of authenticity in UGC, commercial collaborations between brands and influencers (KOLs / KOCs) have evolved into one of the most influential marketing forms globally. However, driven by the commercial motivation to maintain the "halo of authenticity," an increasing number of brands and creators are bypassing platform intermediaries and mandatory disclosure obligations, disguising organic content as commercial promotions, thus creating undisclosed sponsored content (USC), which is commercial promotional content that does not, as required by regulations, indicate the sponsorship relationship in the post.

[0003] The widespread existence of USC (Unsolicited Content) not only continuously deceives users, subjecting them to commercial manipulation without their knowledge, but also forces platforms to bear significant reputational risks without receiving corresponding advertising revenue. Research indicates that up to 96% of sponsored posts on major platforms are not disclosed as required, and existing USC detection methods face fundamental technical challenges in the following four areas:

[0004] First, there's the issue of the sparsity and continuous evolution of deceptive strategies. As new product categories emerge and promotional strategies iterate, deceptive signals become increasingly sparse and context-dependent in massive amounts of user-generated content. This limits the generalization ability of purely data-driven AI models, and capturing such signals through manual annotation is both economically and operationally infeasible. Existing human-machine collaborative content moderation frameworks are mostly one-way workflows: generative AI (GenAI) handles large-scale screening and reports uncertain cases to human moderators, with human judgment serving only as the final decision and not being transformed into continuous learning signals. This leads to a severe underestimation of valuable human feedback.

[0005] Second, there is the issue of reliability in end-to-end decision-making by large language models. While LLMs demonstrate significant advantages in semantic representation and contextual reasoning, when used as end-to-end decision-makers for USC's automated review process, their autoregressive generation mechanism carries risks of hallucination and over-generalization, potentially leading to unfair misjudgments of compliant creators or underestimation of feigned commercial intent. As a black-box system, LLMs struggle to provide the statistically reliable confidence levels required in high-risk review scenarios.

[0006] Third, there is the issue of the effectiveness of user intervention regarding USC's detection results. Existing platform disclosure mechanisms generally use binary warning labels (such as "marked as sponsored content"), which are too simplistic, lack transparency, and are easily perceived by users as arbitrary judgments, triggering psychological resistance or generalized suspicion of the platform's overall content, thus achieving the opposite of the intended intervention effect.

[0007] Fourth, there is the issue of effective fusion of multimodal information. Deceptive cues in USC are often hidden in multiple dimensions, including post text, images, comments, and publisher behavior. It is difficult to capture them comprehensively by relying solely on text features. The technical bottleneck of existing methods is how to effectively fuse high-dimensional semantic features with low-dimensional numerical features and avoid feature dominance and information redundancy.

[0008] In summary, there is an urgent need for a USC detection and intervention method that can continuously learn under limited annotation resources, effectively suppress LLM illusion, reliably fuse multimodal features, and transform detection results into an intervention mechanism that can be recognized and accepted by users. Summary of the Invention

[0009] To address the shortcomings of existing end-to-end Large Language Models (LLMs) and traditional supervised learning methods in detecting implicit sponsored content on social media, which suffer from issues such as recognition illusions, overgeneralization, uncalibrated confidence levels, and uninterpretable decisions due to the LLMs acting as the final arbiter, and the continuous dynamic changes in brand data, product categories, text and image expression methods, and implicit promotion strategies on social media platforms, coupled with the dispersion of deceptive cues across multimodal raw observation data including text, images, comments, author information, and platform interaction metrics, existing methods suffer from low accuracy in identifying implicit sponsored content, poor generalization ability, low efficiency in utilizing human review resources, difficulty in effectively integrating multi-source cues under imbalanced feature spaces, and difficulty in converting detection results into user-understandable disclosure prompts, this invention proposes a human-computer collaborative method for detecting and intervening in implicit sponsored content that incorporates uncertainty perception. The core technology of this method is the proposed Uncertainty Awareness-Integrated Human-Machine Collaborative Implicit Sponsored Content Detection and Intervention Framework (U-HGAC). Based on a human-machine collaborative iterative learning framework, this method positions a large language model as a multimodal semantic feature extractor rather than a final detection judge. It utilizes a pre-trained large language model to extract deep semantic embeddings from post text and visual features from post images, and combines platform metadata, content quality features, and sentiment features to construct a unified structured feature representation. On this basis, a calibrated support vector machine classifier with class imbalance penalty weights learns the optimal classification hyperplane between implicitly sponsored content and natural content, and uses Platt scaling to obtain a reliable posterior probability output. Furthermore, a hybrid query strategy combining information entropy sampling and periodic Bayesian inconsistency refinement sampling is employed to submit the most informative sample batches for manual review, and the manual review results are fed back to the reviewed subset to trigger classifier retraining. Finally, the calibrated posterior probability is converted into a probabilistic disclosure label and displayed to users to activate their persuasive knowledge and promote informed decision-making.

[0010] To achieve the above objectives, the specific technical solution of the present invention is as follows:

[0011] A method for detecting and intervening in implicit sponsored content in human-computer collaboration that integrates uncertainty perception includes:

[0012] Step S1: Collect user-generated content posts from social media platforms, obtain multimodal raw observation data, construct a dataset and divide it into an approved subset and an unapproved candidate pool; the multimodal raw observation data includes text modal data, image modal data and platform metadata;

[0013] Step S2: Extract features from the multimodal raw observation data in the dataset, and perform feature space conditionalization on various feature vectors to generate a unified structured feature representation;

[0014] The feature extraction includes: generating deep semantic embedding vectors for post text using a pre-trained large language model, generating visual feature vectors for post images, extracting original platform observation feature vectors from platform metadata, extracting generalization pointing feature vectors through platform metadata and text content, and extracting content quality feature vectors and sentiment feature vectors through LLM semantic inference.

[0015] Step S3: Train a soft-margin support vector machine with class imbalance penalty weights on the approved subset to learn the optimal classification hyperplane between implicit sponsored content and natural content; apply Platt scaling to the support vector machine decision values ​​to obtain the calibrated posterior probability output, which serves as a reliable input to the uncertainty-driven hybrid query strategy.

[0016] Step S4: Based on a hybrid query strategy, select the most informative sample batch from the unreviewed candidate pool and submit it for manual review; add the manual review results to the reviewed subset, trigger the calibration support vector machine classifier to be retrained on the updated reviewed subset, and enter the next round of human-machine collaboration; the hybrid query strategy is to use information entropy-based sampling in non-refining rounds, and to use Bayesian inconsistency-based refined sampling on the high-entropy candidate subset in refined rounds at fixed intervals;

[0017] Step S5: Repeat steps S3 to S4 until the number of reviewed samples reaches the preset labeling budget limit or the number of rounds of human-machine collaboration reaches the preset limit, and obtain the trained calibrated support vector machine classifier.

[0018] Step S6: Input the feature representation of the post to be detected into the trained calibrated support vector machine classifier to obtain the posterior probability that it is latent sponsored content; if the posterior probability exceeds the preset judgment threshold, the post is judged to be latent sponsored content, otherwise the post is judged to be natural content.

[0019] Step S7: Based on the posterior probability obtained in step S6, generate a probabilistic disclosure label and display it to the user. The probabilistic disclosure label dynamically presents the predicted probability value of implicit sponsored content to activate the user's persuasive knowledge and promote informed decision-making.

[0020] Further, in step S1, the dataset is represented as:

[0021]

[0022] in, Indicates the first User-generated content posts , Total number of posts; This indicates that the label is manually reviewed. When this is the case, it indicates that the post is implicitly sponsored content. When the post is displayed, it indicates that the content is natural.

[0023] The dataset is divided into a candidate sample pool and a test sample set. The candidate sample pool is further divided into an approved subset and an unapproved candidate pool based on whether or not it has an approval tag, as shown below:

[0024]

[0025] in, Represents the candidate sample pool. Indicates the first The approved subset of human-machine collaboration. Indicates the first Unreviewed candidate pool for human-machine collaboration.

[0026] Further, in step S1, the text modal data includes the post title, post body, comment text, and topic tags; the image modal data includes the post cover image, images in the body, and product display images; the platform metadata includes the number of likes, favorites, comments, author's followers, author's historical likes, author's historical favorites, author's homepage information, and information related to commercial promotion, including external links or brand mentions.

[0027] Further, in step S2, feature extraction is performed on the multimodal raw observation data in the dataset, specifically including:

[0028] Generate deep semantic embedding vectors from post text using a pre-trained large language model. :

[0029]

[0030] in, This indicates semantic embedding in the post body. This indicates semantic embedding of the post title. This indicates semantic embedding of comments. This indicates semantic embedding of topic tags. Indicates the number of topic tags. This represents the embedding dimension of a large language model. This indicates vector concatenation;

[0031] Generate visual feature vectors from post images using a large model visual encoder. ;

[0032] Extracting raw platform observation feature vectors from platform metadata The original platform observation feature vector includes at least one of the following: number of post comments, number of likes, number of favorites, number of author followers, and author historical interaction metrics. The author historical interaction metrics are the sum of the author's historical likes and historical favorites.

[0033] Extract promotion-targeting feature vectors from platform metadata and text content. The promotion-oriented feature vector is used to indicate whether a post contains external links, explicit mentions of brand or product names, cross-platform traffic generation, purchase entry points, or other commercial promotion leads;

[0034] Extracting content quality feature vectors through LLM semantic inference The content quality feature vector includes topic consistency, lexical diversity, text coherence, text content redundancy, image-text consistency, tag consistency, author's historical topic consistency, tag diversity, and comment sentiment variance.

[0035] Extract sentiment feature vectors using LLM semantic inference or sentiment analysis modules. The sentiment feature vector includes the sentiment score of the post body, the sentiment score of the post title, the sentiment score of the comments, and the sentiment differences between different fields.

[0036] Furthermore, in step S2, feature space conditionalization is performed on various feature vectors to generate a unified structured feature representation, specifically including:

[0037] Principal component analysis is applied to reduce the dimensionality of the multimodal semantic embedding matrix formed by concatenating the deep semantic embedding vector and the visual feature vector to obtain a compact semantic representation. And the cumulative explained variance satisfies:

[0038]

[0039] in, Indicates the first The eigenvalues ​​corresponding to each principal component Indicates the original semantic embedding dimension;

[0040] A linear gain factor is applied to the original platform observation feature vector, generalization orientation feature vector, and content quality feature vector to prevent them from being overwhelmed by high-dimensional semantic features in subsequent similarity calculations; and these are then sequentially concatenated with the compact semantic representation and the sentiment feature vector to form the first... Structured feature representation of user-generated content posts :

[0041]

[0042] in, Indicates the first A compact semantic representation of each user-generated content post; Represents the linear gain factor. =3.

[0043] Furthermore, in step S3, the training process of calibrating the support vector machine classifier specifically includes:

[0044] In the conditional feature space, the optimal classification hyperplane is learned by solving the soft-margin quadratic programming problem, and the samples are implicitly mapped to the infinite-dimensional Hilbert space by the radial basis function kernel to realize the nonlinear decision boundary.

[0045] To address the class imbalance between implicitly sponsored content samples and natural content samples, a class-specific penalty weight is introduced to amplify the misclassification cost of minority class samples.

[0046] The support vector machine decision values ​​are scaled using Platt and then converted into posterior probabilities using a temperature-scaled Sigmoid function.

[0047] Further, in step S3, the optimization objective of the soft-margin support vector machine is:

[0048]

[0049] The constraints are:

[0050]

[0051] in, The normal vector of the classification hyperplane. Indicates the bias term. Represents slack variables. Represents the regularization parameter. This indicates the number of training samples in the reviewed subset;

[0052] The radial basis function kernel is:

[0053]

[0054] in, , The first User-generated content posts and the first A structured feature representation of user-generated content posts;

[0055] Category-specific penalty weights Set to:

[0056]

[0057] in, Indicate category The number of reviewed samples in the database; Indicates the number of categories. ;

[0058] The support vector machine decision function is expressed as follows:

[0059]

[0060] in, Represents the set of support vectors. Represents the Lagrange multipliers. Indicates the optimal bias term;

[0061] The Platt scaling formula is:

[0062]

[0063] in, This is the calibrated posterior probability; Indicates model parameters; and The calibration parameters are obtained by minimizing the negative log-likelihood:

[0064]

[0065] in, This represents the posterior probability prediction value after calibration.

[0066] Furthermore, in step S4, after adding the manual review results to the already reviewed subset, the dataset is updated as follows:

[0067]

[0068]

[0069] in, This is an unapproved sample batch.

[0070] Furthermore, in step S4, in each round of human-machine collaboration, the hybrid query strategy selects query batches according to the following periodic rules:

[0071] In the non-refined round, the information entropy is calculated for all samples in the unreviewed candidate pool, and the batch of samples with the highest information entropy is selected.

[0072] In the refinement round, firstly, a number of samples with the highest information entropy are selected from the unreviewed candidate pool to form a high-entropy candidate subset. Then, the Bayesian inconsistency degree of each sample in the high-entropy candidate subset is approximately calculated using Monte Carlo input perturbation. The batch of samples with the highest Bayesian inconsistency degree is selected. The refinement rounds occur at fixed periodic intervals.

[0073] The hybrid query strategy is expressed by the following formula:

[0074]

[0075] in, Indicates a fixed refining cycle. This represents the high-entropy candidate subset obtained by pre-screening using information entropy. Indicates the degree of Bayesian inconsistency. This represents information entropy.

[0076] Furthermore, in the non-refined round, the information entropy is calculated for all samples in the unreviewed candidate pool, specifically as follows:

[0077]

[0078]

[0079] in, This represents the posterior probability that a sample belongs to implicitly sponsored content. This represents the posterior probability that the sample does not belong to implicitly sponsored content. .

[0080] Furthermore, in the refinement rounds at fixed intervals, the Bayesian inconsistency degree is approximated for each sample in the high-entropy candidate subset using Monte Carlo input perturbation, as follows:

[0081] A high-entropy candidate subset is formed by selecting the samples with the highest information entropy from the unreviewed candidate pool. :

[0082]

[0083] in, This indicates the batch size of samples submitted for manual review in each round. This indicates the first candidate in the unreviewed candidate pool, sorted by information entropy. One sample;

[0084] Perform on each sample in the high-entropy candidate subset Secondary perturbation inference, calculating the average predicted distribution. :

[0085]

[0086] Calculating Bayesian inconsistency based on the average prediction distribution:

[0087]

[0088] in, Indicates the first Secondary input disturbance.

[0089] Furthermore, in step S5, the preset upper limit of the annotation budget is:

[0090]

[0091] in, This indicates the initial number of reviewed samples. Indicates the number of rounds of human-machine collaboration. This indicates the number of samples submitted for manual review in each round.

[0092] Furthermore, in step S7, the probabilistic disclosure label includes "Platform notification: This content exists". "% probability of undisclosed sponsorship content"; among which , This is the posterior probability obtained in step S6.

[0093] Furthermore, the probabilistic disclosure label can be displayed differently according to different probability ranges: when When the risk level is below the first disclosure threshold, no risk warning will be displayed; when... When the value is between the first and second disclosure thresholds, a medium risk probability warning is displayed; when... When the risk exceeds the second disclosure threshold, a high-risk probability warning will be displayed or the platform's governance processes such as manual review, traffic restriction, and labeling will be triggered.

[0094] The beneficial effects of this invention are as follows: This invention combines the multimodal semantic representation capability of a large language model, visual encoding capability, platform metadata structured expression capability, stable discrimination capability of calibrated support vector machine, context judgment capability of manual review, and user cognitive trigger capability of probabilistic disclosure. It can achieve more reliable, interpretable, and probabilistic disclosure of implicit soft advertising in situations where implicit sponsored content samples are sparse, brand data is dynamically changing, deceptive clues are ambiguous, and manual review budgets are limited. Attached Figure Description

[0095] Figure 1 This is a flowchart of a human-computer collaborative implicit sponsorship content detection and intervention method that incorporates uncertainty perception in an embodiment of the present invention.

[0096] Figure 2 This is a schematic diagram of the overall architecture of the U-HGAC human-computer collaboration framework that integrates uncertainty perception in an embodiment of the present invention.

[0097] Figure 3 This is a schematic diagram of the workflow of the multimodal feature extraction module based on LLM in an embodiment of the present invention.

[0098] Figure 4This is a schematic diagram comparing the performance of different confidence level calibration methods for calibrating support vector machine classifiers in an embodiment of the present invention.

[0099] Figure 5 This is a schematic diagram of the hyperparameter sensitivity analysis results of the U-HGAC framework in an embodiment of the present invention.

[0100] Figure 6 This is a schematic diagram illustrating the effect of post type and disclosure type on the activation of persuasive knowledge in the subject experiment of this invention, wherein (a) shows the effect of post type on the activation of persuasive knowledge, and (b) shows the effect of disclosure type on the activation of persuasive knowledge.

[0101] Figure 7 This is a schematic diagram of the interaction effect between disclosure type and uncertainty level in the LLM evaluation simulation experiment in this embodiment of the invention, where (a) is the influence of post type, (b) is the influence of disclosure type, and (c) is the interaction effect of uncertainty level. Detailed Implementation

[0102] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.

[0103] This invention models the USC detection task as a binary classification problem under audit budget constraints, and designs a system within a human-machine collaboration framework to maintain decision accountability with minimal annotation cost. Figure 1 and Figure 2 As shown, the present invention provides a method for detecting and intervening in implicit sponsored content in human-computer collaboration that integrates uncertainty perception, comprising the following steps:

[0104] Step 1: Collect user-generated content posts from social media platforms, obtain multimodal raw observation data, construct a dataset, and divide it into an approved subset and an unapproved candidate pool.

[0105] Given a dataset ,in This represents the total number of posts. From the audited subset Unreviewed candidate pool and test sample set constitute. For the post The raw multimodal observations include text modal data, image modal data, and platform metadata; Labels assigned to human reviewers This indicates that the post is from USC. This indicates that the post contains organic content. The learning objective is to train the model. To maximize the F1 score while satisfying the following audit budget constraints:

[0106]

[0107] in, This represents the initial number of reviewed samples. Let T be the number of samples submitted for human review in each round of human-machine collaboration, and let T be the number of rounds of human-machine collaboration. Under this constraint, the decision coverage of AI is gradually expanded through iterative human-machine collaboration, striving to achieve effective detection of large-scale unreviewed data with minimal annotation.

[0108] In this embodiment, the dataset originates from a leading user-generated content platform, constructed based on the platform's law enforcement disclosure reports, and covers eight major consumer categories, ultimately containing 3,936 text and image posts and 504,445 related comments. Logarithmic transformations were applied to variables exhibiting a significant right-skewed distribution in social interaction metrics (such as the number of followers and likes) to ensure data stability. Labeling was completed by three independent labelers, and the final labels were determined by majority consensus, with a Fleiss-Kappa confidence level of 0.8698 (nearly perfect agreement) among the labelers.

[0109] Step 2: The LLM-based multimodal feature extraction module extracts features from the original multimodal observation data in the dataset and performs feature space conditionalization on various feature vectors to generate a unified structured feature representation.

[0110] like Figure 3 As shown, this module transforms the original multimodal post data into a structured feature representation, integrating three types of complementary signals:

[0111] (a) Pre-trained model representation features. The post content, post title, comments, and topic tags are encoded into dense embedding vectors through the embedding layer of a pre-trained large-scale language model (Qwen3-VL-Embedding-8B in this embodiment), forming text semantic embedding vectors. ,in This is an embedding dimension specific to large language models. Post images are then encoded into visual feature vectors using the Qwen3-VL-Embedding-8B visual encoder. .

[0112] (ii) Original Platform Observation Characteristics. Observable interactive responses and platform-level business leads are obtained directly from post metadata, including social interaction metrics. (Number of comments, likes, favorites, total historical likes, and number of followers of the poster) and binary third-party promotion signals (Reflects whether the post involves the promotion of external brands or products {0,1}).

[0113] (III) Content Quality Features. Content quality vectors are obtained through LLM semantic inference. It captures the inherent characteristics of posts, including topic consistency, lexical diversity, text coherence, text content redundancy, image-text consistency, tag consistency, author's historical topic consistency, tag diversity, and comment sentiment variance. Among these, topic consistency... This is used to measure whether the title, body text, tags, and comments revolve around the same topic; lexical diversity. This is used to measure the richness of vocabulary in the main text, avoiding template-based or repetitive expressions; text coherence. It is used to measure whether the sentences within a text are semantically coherent and logically connected; text content redundancy. This is used to measure whether there are a large number of repetitive sentence structures or semantic repetitions within the main text; consistency between text and graphics. This is used to measure whether the image content matches the text description; tag consistency. This is used to measure whether the topic tags match the content of the text; consistency of the author's historical themes. This is used to measure whether the current post is consistent with the author's previous content themes; tag diversity. This is used to measure whether the tags are rich and cover multiple expressive perspectives; comment sentiment variance. This is used to measure whether the sentiment in the comment section is polarized.

[0114] Simultaneously extracting emotional signals This includes the sentiment score of the post body, the sentiment score of the post title, the sentiment score of the comments, and the sentiment differences between each field. Among these, the sentiment score of the post body... The sentiment score is used to measure the overall sentiment of the main text; the sentiment score of the post title. This is used to measure the sentiment of the title; the sentiment score for comments. This is used to measure the average sentiment response of user comments; the sentiment differences between various fields. It is used to measure whether there is an emotional misalignment between the body text, the title, and the comments.

[0115] All signal groups are extracted within a unified LLM-driven pipeline, ensuring consistency and coherence among heterogeneous feature spaces.

[0116] (iv) Feature Space Conditioning. To improve numerical stability and mitigate the dominant effect of high-dimensional semantic representation: all numerical indicators are standardized using training set statistics; the concatenated multimodal semantic embedding matrix is ​​then subjected to conditional processing. By applying principal component analysis for dimensionality reduction, a compact semantic representation is obtained. Under LLM conditions, more than 92% of the cumulative explained variance ratio is retained:

[0117]

[0118] A linear gain factor is applied to the original platform observation feature vector, the promotion-oriented feature vector, and the content quality feature vector. This is to prevent it from being overwhelmed by high-dimensional semantic features in subsequent similarity calculations. (Post) The final feature representation is constructed through ordered concatenation:

[0119]

[0120] Step 3: Based on the calibrated support vector machine classifier module, train a soft-margin support vector machine with class imbalance penalty weights on the approved subset to learn the optimal classification hyperplane between implicit sponsored content and natural content; apply Platt scaling to the support vector machine decision values ​​to obtain the calibrated posterior probability output, which serves as a reliable input for the uncertainty-driven hybrid query strategy.

[0121] In the conditional feature space In the context of a given training set SVM learns the optimal classification hyperplane by solving the following soft-margin convex quadratic programming problem:

[0122]

[0123]

[0124] To address class imbalance, class-specific penalty weights are introduced. ,in For category The number of reviewed samples is increased, and the risk of missed detections is reduced by increasing the cost of misclassification of minority class samples. A radial basis function (RBF) kernel is used. The sample is implicitly mapped to an infinite-dimensional Hilbert space, realizing the nonlinear decision boundary after dimensionality reduction; Mercer's theorem guarantees that the dual problem under the RBF kernel is strictly convex, ensuring global optimum.

[0125] To obtain the posterior probabilities required for the hybrid query strategy and improve its reliability, the SVM decision values ​​are... Applying Platt scaling, the Sigmoid function is then used to convert it into a posterior probability by temperature scaling:

[0126]

[0127] in, A learnable temperature scaling factor. To calibrate the offset, These are the model parameters. The calibration parameters are estimated by minimizing the negative log-likelihood on the validation set.

[0128]

[0129] Calibrated probability It serves as a reliable input for the uncertainty-driven query strategy in the human-computer collaboration module.

[0130] Step 4: Based on the human-machine collaboration module, a hybrid query strategy is used to select the most informative sample batch from the unreviewed candidate pool and submit it for manual review; the manual review results are added to the reviewed subset, triggering the calibration support vector machine classifier to be retrained on the updated reviewed subset, and entering the next round of human-machine collaboration.

[0131] (I) Human-computer collaboration framework design.

[0132] The human-machine collaboration framework embeds human judgment into a pure AI decision-making process loop, achieving progressive refinement of the confidence decision boundary under limited review budgets. A candidate pool is established. The subset is divided into approved subsets based on collaboration round t. with unapproved subsets ,satisfy Initial annotation set First, it is used to train an SVM classifier, and then calibrated posterior probabilities are obtained through Platt scaling. Based on a predefined query strategy, from... Select the unreviewed sample batch with the highest expected information value. Submit to human reviewers; after annotation is completed, merge the new review sample into the database. The SVM is retrained in the next round. The dataset size evolves across each collaborative round as follows:

[0133]

[0134] Through multiple rounds of human-machine collaboration with a fixed budget, the confidence decision boundary is gradually expanded to cover a large-scale unlabeled data space, continuously improving detection performance with limited manual annotation costs.

[0135] (II) Sample Selection Based on Information Entropy. Information entropy is an efficient first-order approximation of sample confidence. For unapproved posts... The calibration probability vector and prediction entropy are defined as follows:

[0136]

[0137]

[0138] in, Ensure numerical stability. Predict entropy in The maximum value (maximum decision ambiguity) is reached when the value is 0.5. or The confidence level approaches zero (high confidence). Entropy sampling offers a low computational cost advantage in large-scale candidate screening due to its single forward inference requirement. However, entropy sampling conflates uncertainties from different sources, failing to distinguish between cognitive uncertainty (which can be reduced) caused by insufficient model knowledge and accidental uncertainty (which cannot be reduced) caused by inherent data ambiguity. This can lead to repeated prioritization of inherently noisy or highly ambiguous samples, which offer limited value for improving the decision boundary.

[0139] (III) Refined Sampling Based on Bayesian Inconsistency. To more accurately capture uncertainties that are informative for model improvement, Bayesian inconsistency (BD) is introduced as a complementarity criterion to quantify the mutual information between predictions and model parameters:

[0140]

[0141] in, Measuring total forecast uncertainty By approximating accidental uncertainty through posterior expectation, the difference between the two represents cognitive uncertainty. In practice, through... The uncertainty of the approximate parameters for the sub-Monte Carlo input perturbation, the integrated prediction distribution, and the inconsistency score are calculated as follows:

[0142]

[0143]

[0144]

[0145] Compared to entropy sampling, BD more accurately targets reducible uncertainties, producing a higher expected information gain, but with a computational cost approximately one-third that of entropy sampling. times.

[0146] (iv) Hybrid Query Strategy. To balance annotation efficiency and information content, entropy-based screening and periodic BD refinement are integrated into a hybrid strategy, and a candidate pre-filtering mechanism is introduced to reduce computational overhead. The hybrid strategy is applied in each round... Select query batches according to the following periodic rules:

[0147]

[0148] in, To maintain a fixed refining cycle. In most rounds (each In the wheel (round), from the complete candidate pool Efficiently identify samples with high uncertainty; each Each round performs a refinement step, but only applies to high-entropy candidate subsets. ( China's predicted entropy ranking The design is based on the premise that high-inconsistency samples are necessarily in regions of high prediction uncertainty, thus limiting BD calculations to... Internal energy can reduce computational cost while preserving informative candidate samples. Significantly reduced to The overall computational cost ratio is approximately:

[0149]

[0150] (v) Pareto improvement proof of hybrid query strategy.

[0151] set up Indicate that policy S is in the first... The expected classifier performance improvement obtained in the round. Because the BD metric more accurately addresses cognitive uncertainty, it satisfies... Therefore, the hybrid strategy achieves improvements in the cumulative classifier that are no less than those of pure entropy sampling.

[0152]

[0153] This is because the hybrid strategy is in In the remaining rounds, it is exactly the same as entropy sampling, in the remaining The hybrid strategy employs a Bayesian refinement step in each round, which guarantees an equal to or greater classifier improvement. Therefore, the hybrid strategy achieves Pareto improvements over both baseline strategies: its annotation efficiency is no worse than pure entropy sampling, and its computational cost is reduced to approximately 11% compared to the full Bayesian strategy.

[0154] Under the specific parameter settings in this embodiment (| The theoretical calculation upper limit of the cost is 10.92%, which is in complete agreement with the estimate of about 11% in the paper.

[0155] Step 5: Repeat steps 3 to 4 until the number of reviewed samples reaches the preset annotation budget limit, and obtain the trained calibrated support vector machine classifier.

[0156] Step 6: Input the feature representation of the post to be detected into the trained calibration support vector machine classifier to obtain the posterior probability that it is latent sponsored content; if the posterior probability exceeds the preset judgment threshold, the post is judged to be latent sponsored content, otherwise the post is judged to be natural content.

[0157] Step 7: Based on the persuasive knowledge model, a probabilistic disclosure intervention mechanism is used to generate probabilistic disclosure tags and display them to users.

[0158] This invention further transforms the USC detection output into a user-oriented intervention mechanism. Based on the Persuasive Knowledge Model (PKM), when a user becomes aware of a potential persuasive intent, their persuasive knowledge is activated, leading to subsequent cognitive adjustments. This invention hypothesizes that detection results presented as calibrated probabilities ("This post has an X% probability of being implicitly sponsored content") provide tiered and more easily interpreted signals, which are more effective at activating the user's persuasive knowledge compared to binary warnings.

[0159] The display rules for probabilistic disclosure labels are as follows: For posts with a posterior probability higher than the judgment threshold, the predicted probability percentage is dynamically displayed next to the post title, presenting probabilistic disclosure information to the user; for low-confidence posts, no labels are applied to establish a natural baseline. Compared to binary warning conditions (absolute red warning: "Platform announcement: This post has been identified as implicitly sponsored content"), probabilistic disclosure provides continuous and interpretable uncertainty information, helping users form autonomous skepticism rather than passively accepting the platform's judgment.

[0160] Verification Experiment

[0161] I. Framework Evaluation Scenario Construction Based on Real Platform Data

[0162] To verify the effectiveness, robustness, and reliability of the U-HGAC method of this invention, a real multimodal implicit soft advertising dataset was constructed based on the Xiaohongshu platform, and four experimental tasks were carried out on the real-world multimodal dataset. The data came from publicly available clues on platform governance from August 2023 to March 2024, covering eight major consumer product categories, and ultimately included 3936 image and text posts and 504445 comments. Three independent annotators performed binary classification annotation, and the Fleiss' Kappa reached 0.8698, indicating high annotation consistency. The data distribution is shown in Table 1.

[0163] Table 1. Descriptive statistics of the original USC dataset

[0164]

[0165] Experiment Task Description

[0166] The experimental tasks included overall detection performance comparison, feature extractor / classifier / sampling strategy ablation, probability calibration comparison, and hyperparameter sensitivity analysis. The results show that the implementation using LLM feature extraction, SVM classifier, and a hybrid query strategy achieves superior detection performance. All experiments were repeated 10 times, and the average results are reported.

[0167] Task 1 (Overall Detection Performance Comparison)

[0168] U-HGAC was compared with 21 pure LLM baseline methods, including zero-shot and mind-chain cue settings for multimodal LLM and plain text LLM. The experimental results are shown in Table 2. The results show that U-HGAC (hybrid strategy) outperforms all pure LLM baselines (the best baseline: 84.94% F1) with an F1 score of 88.13% and an AUC of 86.04%, and significantly improves precision while maintaining high recall, achieving a more balanced classification performance.

[0169] Table 2 Comparison of overall detection performance of different methods

[0170]

[0171] Task 2 (Component Ablation and Impact Analysis)

[0172] Various combinations of three feature extractors (LLM, CLIP, BERT), four classifiers (SVM, MLP, logistic regression, random forest), and three sampling strategies (entropy sampling, BD sampling, and hybrid strategy) were compared. The experimental results are shown in Table 3. The results indicate that the LLM feature extractor consistently achieves the best performance across all settings (optimal combination SVM + hybrid strategy: F1=88.13%, BERT optimal F1=86.44%, CLIP optimal F1=84.25%), while the hybrid sampling strategy achieves optimal or near-optimal performance in most cases.

[0173] Table 3 Performance Comparison of Different Feature Extractors, Classifiers, and Sampling Strategies

[0174]

[0175] Task 3 (Effect of Confidence Calibration Method)

[0176] Four methods were compared: no calibration, Platt scaling, isotonic regression, and temperature scaling. Experimental results are as follows: Figure 4 As shown, the results indicate that the uncalibrated model exhibits significant bias in the high confidence interval. Platt scaling achieves the lowest expected calibration error (ECE) and Brier score, providing the most stable and accurate calibration results.

[0177] Task 4 (Hyperparameter Sensitivity Analysis)

[0178] Experimental results are as follows Figure 5 As shown, U-HGAC exhibits strong robustness under parameter variations, with the number of collaboration rounds and the number of samples per round having a greater impact on performance than the initial number of labeled samples. SVM achieves optimal overall performance and highest stability under different parameter configurations. The optimal configuration is: initial number of samples = 200, number of collaboration rounds = 20, and number of samples per round = 50.

[0179] Task 5 (Disclosure Mechanism Assessment)

[0180] Through dual-track empirical verification using human subject experiments (n=91, 273 valid observations) and LLM evaluation simulations (n=135, 392 valid observations), this invention demonstrates that:

[0181] (1) Probabilistic disclosure labels significantly improve users' persuasive knowledge activation level (PKA) compared to binary warning labels. Human experimental group: LLM simulation group: also significant, Furthermore, this effect remains stable across posts with different levels of uncertainty.

[0182] (2) Probabilistic disclosure labeling does not trigger a significantly higher tendency for overcorrection (OC) (human experiments: This means that while successfully enhancing the activation of specific persuasive knowledge, it will not produce unintended negative effects such as platform generalization of doubt or overall content trust erosion.

[0183] (3) Significant discrepancies exist between LLM simulation and human experimental results: LLM responses exhibit strong uncertainty-level conditional dependence, and the interaction effect is significant. ), while no significant interaction effect was observed in human experiments ( This divergence reveals the inherent limitation of LLM in lacking human emotional spillover effects, and that pure GenAI simulation cannot replace human judgment in trust-sensitive governance scenarios.

[0184] In summary, compared with existing USC detection technologies, the present invention has the following beneficial technical effects:

[0185] (1) The human-machine collaboration paradigm is upgraded from one-way review to uncertainty-driven bidirectional co-learning, and human cognitive resources are allocated in Pareto optimal way. In the existing content review framework, the collaboration between generative AI and human reviewers is essentially a one-way transfer process. After the AI ​​completes a large-scale initial screening, it reports uncertain cases, and human judgment is used as the final decision. However, this feedback is never systematically transformed into a signal for model improvement. This design flaw leads to a huge waste of human cognitive resources, which means that the model still needs to repeatedly use limited human resources when facing similar cases. It cannot accumulate learning ability from historical interactions, which is especially insufficient in the scenario of continuous evolution of USC deception strategies. This invention constructs an uncertainty-driven bidirectional co-learning mechanism, which continuously transforms human judgment into classifier retraining signals, so that the model decision boundary is dynamically refined with each round of labeling, thereby overcoming the adaptability defects of the existing one-way workflow when deception strategies iterate and evolve. Meanwhile, the hybrid query strategy has achieved rigorous Pareto improvement through theoretical proof: in a total of T rounds of collaboration, the cumulative detection performance gain is no less than that of the pure entropy sampling strategy, while the computational cost is only about 11% of that of the full Bayes strategy. It achieves Pareto optimality in both annotation efficiency and computational efficiency, accurately allocating limited human cognitive resources to the boundary cases where the model is most uncertain and the value of human intervention is highest.

[0186] (2) The decoupled architecture effectively suppresses the illusion risk of large models and enables reliable and accountable decision-making in high-risk review scenarios. Current solutions that directly use LLM as the end-to-end decision-maker for USC detection have serious reliability risks due to the inherent illusion and overgeneralization characteristics of the LLM autoregressive generation mechanism: illusion may lead to false accusations against organic content, thereby unfairly suppressing compliant creators, while overgeneralization may allow real disguised commercial intentions to be allowed, leaving users continuously exposed to implicit commercial persuasion. The black-box decision-making mechanism of LLM also makes it difficult to provide the statistically reliable confidence required for high-risk scenarios. This invention proposes a decoupled architecture of "LLM as a semantic extractor + calibrated statistical classifier", which strictly positions LLM in the semantic feature extraction stage rather than the final decision stage. Its rich semantic understanding capabilities are only used to generate multimodal deep feature representations; the final binary classification decision is completed by a support vector machine calibrated by Platt scaling. This calibrator estimates the calibration parameters on the validation set by minimizing the negative log-likelihood, and can generate statistically reliable posterior probabilities rather than the probabilistic token output of LLM. Experimental results demonstrate that this decoupling design outperforms 21 pure LLM baseline methods in both F1 score and AUC, fundamentally eliminating the interference of LLM illusion on the review of high-risk content, while taking into account the accountability and interpretability of decisions.

[0187] (3) Feature space conditionalization effectively alleviates the imbalance of multimodal heterogeneous feature dimensions and significantly improves the collaborative discrimination ability of high-dimensional and low-dimensional features. USC deception cues are distributed in multimodal heterogeneous feature spaces such as text semantics, visual content, social interaction and platform behavior. However, existing methods generally face the problem of severe imbalance of feature dimensions when fusing different modal features: after the high-dimensional semantic embedding (usually thousands of dimensions) generated by the pre-trained large language model is directly concatenated with low-dimensional numerical behavioral features (such as the number of fans, the number of likes, etc., which are only single-digit dimensions), the high-dimensional semantic features dominate the subsequent similarity calculation, which leads to the key deception cues (such as abnormal interaction patterns, external platform jumps, etc.) carried by the low-dimensional numerical behavioral features being submerged, thus affecting the classifier's comprehensive capture ability of multi-dimensional deception signals. This invention employs a targeted feature space conditionalization processing scheme: principal component analysis is applied to the multimodal semantic embedding matrix for dimensionality reduction, retaining the top 256 principal components that account for over 92% of the cumulative explained variance. This significantly compresses the dimensionality of semantic features while preserving their key discriminative information. A linear gain factor is applied to low-dimensional numerical features to enhance their relative weight in the feature space before similarity calculation, preventing them from being overwhelmed by high-dimensional semantic features. This symmetrical "semantic compression + numerical amplification" design achieves a balanced fusion of the two types of features in a unified feature space, overcoming the feature dominance and information redundancy problems caused by simple splicing in existing technologies. This ensures that the final feature representation balances deep semantic understanding and key behavioral signals, thereby significantly improving the overall discriminative ability and generalization performance of the classifier.

[0188] (4) Platt scaling calibration of SVM provides reliable posterior probability output, providing an accurate information basis for human-machine collaboration in uncertainty perception. The effectiveness of the hybrid query strategy depends on the posterior probability output by the classifier accurately reflecting the model's true prediction uncertainty. However, existing classification models usually directly output uncalibrated decision values ​​or prediction probabilities. Such uncalibrated outputs often have systematic biases, leading to a large number of misjudgments in the high confidence interval (overconfidence problem), thus misleading the uncertainty sampling mechanism to misclassify low-value samples as high-uncertainty samples, wasting valuable manual annotation resources. This invention applies Platt scaling calibration to the SVM decision values ​​and estimates the temperature scaling factor and offset on the validation set by minimizing the negative log-likelihood, reliably converting the decision values ​​into calibrated posterior probabilities. Experimental comparisons show that Platt scaling achieves the best performance in both expected calibration error and Brier score reliability indicators, significantly outperforming schemes such as no calibration, temperature scaling, and isotonic regression. Reliable posterior probabilities enable both entropy sampling and Bayesian inconsistency to accurately quantify the model’s true uncertainty for each unverified sample, ensuring that every investment of human annotation resources is directed to samples that truly have information gain, thus guaranteeing the efficient operation of the entire human-machine collaboration framework from an information foundation level.

[0189] (5) The hybrid query strategy distinguishes between cognitive uncertainty and accidental uncertainty, accurately locates the cognitive blind spots of the model, and achieves strict Pareto improvement. Existing single entropy sampling strategies conflate uncertainties from different sources, failing to distinguish between cognitive uncertainty caused by insufficient model knowledge (which can be reduced by labeling new samples) and accidental uncertainty caused by inherent data ambiguity (which cannot be eliminated no matter how many additional labels are obtained). This defect leads to samples that are inherently noisy or naturally ambiguous being repeatedly prioritized, even though such samples have extremely limited value for improving the decision boundary, resulting in a serious waste of the limited labeling budget. This invention introduces Bayesian inconsistency as a complementary refinement criterion. By quantifying the mutual information between prediction and model parameters, the total uncertainty is decomposed into two components: cognitive uncertainty and accidental uncertainty. Only samples with high cognitive uncertainty are prioritized for manual review. However, the computational cost of Bayesian inconsistency is higher than that of entropy sampling. The direct application of high-inconsistency samples to large-scale candidate pools is impractical in engineering. This invention's hybrid strategy leverages the mathematical property that high-inconsistency samples are necessarily located in high-entropy regions, applying Bayesian refinement only to the pre-filtered high-entropy candidate subset, thus compressing computational costs. Furthermore, theoretical proof demonstrates that the hybrid strategy achieves a strict Pareto improvement over pure entropy sampling (at least one round of refinement is significantly better than entropy sampling), and compared to the full Bayesian strategy, it compresses the upper bound of computational overhead to approximately 11%, achieving Pareto optimality simultaneously in labeling efficiency, information gain, and computational efficiency.

[0190] (6) Define the USC detection output as a cognitive trigger and realize a theoretical closed loop from detection results to user intervention based on the persuasive knowledge model. In the past, USC detection and user intervention were promoted as two independent research streams. The detection system was responsible for judging whether the content was USC, and the intervention mechanism informed the user of the conclusion through a simple binary label ("identified as sponsored content"). There was no theoretical connection between the two. This separation led to the key problem of intervention effect: the binary warning label could not convey the degree of uncertainty of the USC determination. Users often perceived it as the platform's arbitrary judgment, which led to psychological resistance and even generalized suspicion of the platform's overall content. The intervention effect was far lower than expected. This invention redefines the posterior probability of USC detection as a "cognitive trigger" based on the persuasive knowledge model (PKM). It directly transforms statistical uncertainty into a probabilistic disclosure label for users ("This post has an X% probability of being implicitly sponsored content"), providing users with graded and interpretable uncertainty signals. Through dual-track empirical verification using human subject experiments and LLM (Leadership Management Model) simulations, this invention demonstrates that probabilistic disclosure is significantly more effective than binary warnings in activating users' persuasive knowledge, and this effect remains stable across posts with different levels of uncertainty. Simultaneously, it does not trigger a significantly higher tendency for overcorrection; that is, while successfully activating specific persuasive knowledge, it does not induce generalized suspicion or trust erosion regarding the overall content of the platform. This finding establishes, for the first time, a quantitative theoretical link between AI-detected statistical uncertainty and user cognitive response in the field of Information Security (IS), providing a replicable research paradigm for other content governance fields requiring high transparency and decision-making reliability.

[0191] (7) Revealing the cognitive boundaries of the LLM evaluation paradigm and clarifying the theoretical limitations of AI replacing human judgment in trust-sensitive governance scenarios. This invention systematically compares the differences in cognitive responses of two types of evaluators to different disclosure formats by simultaneously conducting experiments with human subjects and LLM evaluation simulations. The experimental results reveal a key cognitive divergence: human subjects activate persuasive knowledge after perceiving the persuasive intent and adopt a relatively stable evaluation strategy, resulting in no significant interaction effect between uncertainty level and disclosure type; while the LLM response model strongly depends on uncertainty level conditions and produces a significant interaction effect with disclosure type. More importantly, the internal consistency of the LLM overcorrection (OC) scale is significantly lower than the acceptable level, indicating that LLM, as a rational intelligent agent, does not possess the emotional spillover effect and platform generalized trust collapse reaction experienced by humans when exposing deception. This finding clarifies the theoretical boundaries of the LLM evaluation paradigm in trust-sensitive governance scenarios: LLM can simulate cognitive adjustment based on probabilistic information, but it cannot simulate the defensive behavior and trust erosion induced by human emotional vulnerability. In such scenarios, LLM simulation systematically underestimates the real user trust destruction effect and cannot be used as a substitute for human judgment. This theoretical contribution of the present invention provides an empirical basis for the boundary conditions of the irreplaceability of AI and humans in the governance of social media ecosystems.

[0192] (8) Significantly reduced annotation costs enable high-performance continuous detection in scenarios with small sample sparse annotations. The core data dilemma faced by the USC detection task is that the spoofing signals themselves are sparse and evolve frequently. Acquiring high-quality labeled data is not only costly (requiring experts to combine external searches to accurately determine sponsorship relationships), but historical data also becomes outdated quickly due to rapid updates in product categories, leading to a continuous decline in the performance of models trained on static datasets when facing dynamically changing UGC environments. This invention adopts a combination of frozen pre-trained large model semantic embedding and human-machine collaborative active learning. Without the need for parameter fine-tuning of the large model (avoiding the high computational cost of fine-tuning), it fully utilizes its general semantic representation capabilities. At the same time, the hybrid query strategy prioritizes the annotation of samples with the greatest cognitive uncertainty, maximizing the information gain of each manual annotation input and significantly reducing the repeated annotation of a large number of low-information samples. Experimental results show that, starting from only 200 initial labeled data, after 20 rounds of iterative human-machine collaboration with 50 samples per round, U-HGAC can achieve an F1 score of 88.13%, surpassing the pure LLM baseline that requires a large amount of labeled data for training. This demonstrates that the present invention can rapidly establish effective detection capabilities in cold-start scenarios where labeled data is extremely scarce, and fundamentally solves the vicious cycle problem of increased costs due to data scarcity in USC detection tasks by adapting to the dynamic evolution of deception strategies through continuous human-machine collaboration.

[0193] (9) The framework design possesses model independence and strong hyperparameter robustness, making it easy to promote and deploy in different platform environments. Addressing the strong dependence of the framework on specific model or parameter configurations in actual platform deployments, this invention balances flexibility and robustness in its framework design. At the feature extraction layer, the LLM embedding extractor can be replaced with pre-trained models of different scales, such as BERT or CLIP. Although LLM features have superior absolute performance, the overall framework architecture remains compatible with the choice of specific extractors. At the classifier layer, SVM can be replaced with MLP, logistic regression, or random forest, and various combinations can achieve reasonable performance. SVM is the best in terms of overall performance and stability. At the hyperparameter setting layer, experiments demonstrate that U-HGAC exhibits strong robustness across a wide range of initial labeled sample counts, number of collaborative rounds, and sample counts per round, with minimal performance differences under different configurations. This feature means that platform operators do not need to perform extensive hyperparameter engineering optimization for local data distribution. They can directly deploy U-HGAC in different social platforms, product categories, and operational scale scenarios, which greatly reduces the actual implementation cost and provides practical protection for large-scale commercial platform governance.

[0194] In summary, this invention addresses the technical problems of existing USC detection technologies, such as insufficient utilization of human feedback, low reliability of model decision-making, imbalance in multimodal feature fusion, high annotation costs, and poor intervention effects, through synergistic effects from multiple levels, including upgrading the human-machine collaboration paradigm, suppressing LLM illusions, balancing the fusion of multimodal features, accurately quantifying uncertainty, closing the theoretical loop of detection intervention, and revealing the cognitive boundaries of LLM. It provides a reliable, efficient, interpretable, and user-cognition-oriented complete technical solution for social media content governance.

[0195] Finally, it should be noted that the above embodiments are intended to illustrate the technical solutions of the present invention and do not constitute any limitation on the present invention. Those skilled in the art should fully understand that modifications to the technical solutions described in the foregoing embodiments or equivalent substitutions for any part or all of the technical features are entirely feasible. Such modifications or substitutions, as long as they do not depart from the scope of protection defined by the claims of the present invention, should be considered reasonable extensions of the present invention.

Claims

1. A method for detecting and intervening in implicit sponsorship content in human-computer collaboration that integrates uncertainty perception, characterized in that, include: Step S1: Collect user-generated content posts from social media platforms, obtain multimodal raw observation data, construct a dataset and divide it into an approved subset and an unapproved candidate pool; the multimodal raw observation data includes text modal data, image modal data and platform metadata; Step S2: Extract features from the multimodal raw observation data in the dataset, and perform feature space conditionalization on various feature vectors to generate a unified structured feature representation; The feature extraction includes: generating deep semantic embedding vectors for post text using a pre-trained large language model, generating visual feature vectors for post images, extracting original platform observation feature vectors from platform metadata, extracting generalization pointing feature vectors through platform metadata and text content, and extracting content quality feature vectors and sentiment feature vectors through LLM semantic inference. Step S3: Train a soft-margin support vector machine with class imbalance penalty weights on the approved subset to learn the optimal classification hyperplane between implicit sponsored content and natural content; apply Platt scaling to the support vector machine decision values ​​to obtain the calibrated posterior probability output, which serves as a reliable input to the uncertainty-driven hybrid query strategy. Step S4: Based on a hybrid query strategy, select the most informative sample batch from the unreviewed candidate pool and submit it for manual review; add the manual review results to the reviewed subset, trigger the calibration support vector machine classifier to be retrained on the updated reviewed subset, and enter the next round of human-machine collaboration; the hybrid query strategy is to use information entropy-based sampling in non-refining rounds, and to use Bayesian inconsistency-based refined sampling on the high-entropy candidate subset in refined rounds at fixed intervals; Step S5: Repeat steps S3 to S4 until the number of reviewed samples reaches the preset labeling budget limit or the number of rounds of human-machine collaboration reaches the preset limit, and obtain the trained calibrated support vector machine classifier. Step S6: Input the feature representation of the post to be detected into the trained calibrated support vector machine classifier to obtain the posterior probability that it is latent sponsored content; if the posterior probability exceeds the preset judgment threshold, the post is judged to be latent sponsored content, otherwise the post is judged to be natural content. Step S7: Based on the posterior probability obtained in step S6, generate a probabilistic disclosure label and display it to the user. The probabilistic disclosure label dynamically presents the predicted probability value of implicit sponsored content to activate the user's persuasive knowledge and promote informed decision-making.

2. The method according to claim 1, characterized in that, In step S2, feature extraction is performed on the multimodal raw observation data in the dataset, specifically including: Generate deep semantic embedding vectors from post text using a pre-trained large language model. : in, This indicates semantic embedding in the post body. This indicates semantic embedding of the post title. This indicates semantic embedding of comments. This indicates semantic embedding of topic tags. Indicates the number of topic tags; Generate visual feature vectors from post images using a large model visual encoder. ; Extracting raw platform observation feature vectors from platform metadata The original platform observation feature vector includes at least one of the following: number of post comments, number of likes, number of favorites, number of author followers, and author historical interaction metrics. The author historical interaction metrics are the sum of the author's historical likes and historical favorites. Extract promotion-targeting feature vectors from platform metadata and text content. The promotion-oriented feature vector is used to indicate whether a post contains external links, explicit mentions of brand or product names, cross-platform traffic generation, purchase entry points, or other commercial promotion leads. Extracting content quality feature vectors through LLM semantic inference The content quality feature vector includes topic consistency, lexical diversity, text coherence, text content redundancy, image-text consistency, tag consistency, author's historical topic consistency, tag diversity, and comment sentiment variance. Extract sentiment feature vectors using LLM semantic inference or sentiment analysis modules. The sentiment feature vector includes the sentiment score of the post body, the sentiment score of the post title, the sentiment score of the comments, and the sentiment differences between different fields.

3. The method according to claim 1 or 2, characterized in that, In step S2, feature space conditionalization is performed on various feature vectors to generate a unified structured feature representation, specifically including: Principal component analysis is applied to the multimodal semantic embedding matrix formed by concatenating the deep semantic embedding vector and the visual feature vector to reduce dimensionality, resulting in a compact semantic representation that retains at least 92% of the cumulative explained variance of the original feature space. A linear gain factor is applied to the original platform observation feature vector, promotion orientation feature vector, and content quality feature vector, and these are then concatenated with the compact semantic representation and the sentiment feature vector in an ordered manner to form a structured feature representation of user-generated content posts.

4. The method according to claim 1, characterized in that, Step S3, the training process of calibrating the support vector machine classifier, specifically includes: In the conditional feature space, the optimal classification hyperplane is learned by solving the soft-margin quadratic programming problem, and the samples are implicitly mapped to the infinite-dimensional Hilbert space by the radial basis function kernel to realize the nonlinear decision boundary. To address the class imbalance between implicitly sponsored content samples and natural content samples, a class-specific penalty weight is introduced to amplify the misclassification cost of minority class samples. The support vector machine decision values ​​are scaled using Platt and then converted into posterior probabilities using a temperature-scaled Sigmoid function.

5. The method according to claim 4, characterized in that, In step S3, the optimization objective of the soft-margin support vector machine is: The constraints are: in, The normal vector of the classification hyperplane. Indicates the bias term. Represents slack variables. Represents the regularization parameter. This indicates the number of training samples in the reviewed subset; The radial basis function kernel is: in, , The first User-generated content posts and the first A structured feature representation of user-generated content posts; Category-specific penalty weights Set to: in, Indicates category The number of reviewed samples in the database; Indicates the number of categories. ; The support vector machine decision function is expressed as follows: in, Represents the set of support vectors. Represents the Lagrange multipliers. Indicates the optimal bias term; The Platt scaling formula is: in, This is the calibrated posterior probability; Indicates model parameters; and The calibration parameter is obtained by minimizing the negative log-likelihood.

6. The method according to claim 5, characterized in that, In step S4, in each round of human-machine collaboration, the hybrid query strategy selects query batches according to the following periodic rules: In the non-refined round, the information entropy of all samples in the unreviewed candidate pool is calculated, and the batch of samples with the highest information entropy is selected as the query batch. In the refinement round, firstly, select a number of samples with the highest information entropy from the unreviewed candidate pool to form a high-entropy candidate subset. Then, calculate the Bayesian inconsistency of each sample in the high-entropy candidate subset using Monte Carlo input perturbation. Select the batch of samples with the highest Bayesian inconsistency as the query batch. The refining cycles occur at fixed periodic intervals.

7. The method according to claim 6, characterized in that, In the non-refined round, the information entropy is calculated for all samples in the unreviewed candidate pool, specifically as follows: in, This represents the posterior probability that a sample belongs to implicitly sponsored content. This represents the posterior probability that the sample does not belong to implicitly sponsored content. Represents information entropy. .

8. The method according to claim 7, characterized in that, In the refinement rounds at fixed intervals, the Bayesian inconsistency degree is approximated for each sample in the high-entropy candidate subset using Monte Carlo input perturbation, as follows: A high-entropy candidate subset is formed by selecting the samples with the highest information entropy from the unreviewed candidate pool. : in, This indicates the batch size of samples submitted for manual review in each round. This indicates the first candidate in the unreviewed candidate pool, sorted by information entropy. One sample; Perform on each sample in the high-entropy candidate subset Secondary perturbation inference, calculating the average predicted distribution. : Calculating Bayesian inconsistency based on the average prediction distribution : in, Indicates the first Secondary input disturbance.

9. The method according to claim 1, characterized in that, In step S7, the probabilistic disclosure label is displayed differently according to different probability ranges: when the posterior probability is lower than the first disclosure threshold, no risk warning is displayed; when the posterior probability is between the first and second disclosure thresholds, a medium risk probability warning is displayed; when the posterior probability is higher than the second disclosure threshold, a high risk probability warning is displayed or the platform's manual review, traffic restriction, and tagging governance process is triggered.

10. The method according to claim 1, characterized in that, In step S1, the text modal data includes the post title, post body, comment text, and topic tags; the image modal data includes the post cover image, images in the body, and product display images; the platform metadata includes the number of likes, favorites, comments, author's followers, author's historical likes, author's historical favorites, author's homepage information, and information related to commercial promotion, including external links or brand mentions.