A Multi-Dimensional Feature Interaction Influencer Recommendation Method Based on Account Topic Model
Through a multi-dimensional feature interaction method based on the account theme model, the shortcomings of the existing influencer recommendation algorithm when considering the diversity of account content, emotional factors and brand interaction relationships are solved, and more accurate and personalized influencer recommendations are achieved.
Patent Information
- Application Number
- CN202510295257.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-13
AI Technical Summary
When selecting the most suitable influencer, existing influencer recommendation algorithms cannot fully consider the diversity of account content, emotional factors, and the complex interaction between the brand and influencer, resulting in insufficient accuracy and personalization of recommendations.
Using a multi-dimensional feature interaction method based on the account theme model, by constructing an account theme model SMATM, multi-dimensional features such as account theme features, image features, industry probability distribution features, emotional features and attribute features are extracted, and deep interaction between features is achieved through deep networks and cross networks, and finally optimize the recommended results through sorting losses and matching losses.
It improves the comprehensiveness and accuracy of account feature representation, effectively captures the complex interaction between multi-dimensional features, improves the accuracy and personalization of recommendations, and optimizes the sorting and filtering of recommendation results.
Smart Images

Figure CN119807539B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a multi-dimensional feature interaction influencer recommendation method based on an account topic model. Background Art
[0002] With the rapid development of social media, social media marketing has gradually become the core means for enterprises to promote brands, enhance user stickiness, and improve market competitiveness. In this process, influencers have become the bridge between brands and target users. With their professionalism in specific fields, profound social media influence, and close interaction with the audience, they can help brands increase exposure and trust in a short period of time. How to accurately select the most suitable influencers to achieve efficient brand communication has become an important issue in influencer marketing.
[0003] Early influencer recommendation algorithms mainly focused on how to mine influencers with great potential from a large number of social media users, usually relying on traditional features such as the number of followers, the influence of historical posts, and the relationship strength in the social network. These methods, based on the statistical analysis of user basic data, can roughly screen out influencers with high communication potential. For example, by analyzing the interaction frequency between influencers and their followers, the number of likes and comments, etc., the communication effectiveness of a certain influencer can be roughly judged. However, with the diversification of social media platform functions, such traditional feature-dependent methods have gradually revealed many deficiencies. Simply relying on the number of followers and historical data ignores the diversity of content, emotional factors, and the fit with the brand, resulting in insufficient accuracy and personalization of recommendations and being unable to meet more complex and refined marketing needs.
[0004] To address these deficiencies, influencer recommendation algorithms have started to incorporate content relevance analysis. Specifically, by analyzing the themes, styles of the content posted by influencers and their degree of match with brand requirements, the accuracy of recommendations can be significantly improved. Such content-relevance-based recommendation methods can conduct in-depth analysis of the text content posted by influencers through natural language processing (NLP) techniques to determine the theme categories of their content. Meanwhile, deep learning methods have been gradually applied to influencer recommendation systems, especially in the modeling of account historical activity data. Through deep neural networks (DNNs), more complex feature extraction can be performed on the historical behavior data of influencers. By jointly analyzing the past interaction records and posted content of influencers, it is possible to automatically discover which behavioral patterns are highly attractive to the target user group of a specific brand. This method can provide more personalized recommendations, helping brands more accurately identify those influencers who can best attract their target audience. Although certain progress has been made in existing algorithms related to influencer marketing, how to accurately select the most suitable influencers for a certain brand or event among numerous influencers remains a challenging issue. Existing algorithms often lack sufficient consideration of the diversity of account content, emotional factors, and the complex interaction relationships between influencers and brands when analyzing the content of influencer accounts. In addition, with the continuous development of social media content, the emergence of new content forms, interaction methods, and changes in social behaviors have made it necessary for influencer recommendation algorithms to be continuously updated and optimized to adapt to these new requirements and challenges.
[0005] One of the core issues in influencer recommendation lies in how to effectively represent and understand the characteristics of an account. In previous work, the recommendation methods constructed for influencer recommendation tasks mostly relied on text information or image features to obtain content-related features of the account to measure the account's expression. However, in the actual complex application environment, the influencer recommendation task is not simply a matter of simple content relevance matching, but a multi-dimensional and multi-level recommendation problem. The recommendation system needs to consider not only the relevance of account content, but also incorporate features from multiple dimensions such as emotional factors, industry attributes, and the attributes of the account itself into the representation in order to comprehensively reflect the characteristics and value of the account.
[0006] In the influencer recommendation task, in addition to solving how to effectively represent and understand account features, another core issue is how to achieve accurate matching and efficient recommendation of brands and influencer accounts based on account features. The core goal of influencer recommendation is to help brands find influencers that highly match the product characteristics, marketing objectives, and target audience. Compared with traditional recommendation tasks, the data characteristics of the influencer recommendation task are abstract, heterogeneous, and multi-dimensional. Therefore, previous inventions or research work mainly focused on solving the first core issue in the influencer recommendation task. Given the complexity and particularity of the influencer recommendation task, further optimizing the influencer recommendation framework can be achieved by solving the second core issue, that is, how to achieve accurate matching and efficient recommendation of brands and influencer accounts. This issue requires more attention to the complex correlation and interaction relationships between multi-dimensional features. Summary of the Invention
[0007] In view of the above deficiencies of the prior art, the present application provides a multi-dimensional feature interaction influencer recommendation method based on an account topic model.
[0008] In a first aspect, the present application proposes a multi-dimensional feature interaction influencer recommendation method based on an account topic model, including the following steps:
[0009] Receive and integrate the multi-dimensional heterogeneous data of each target account through the input layer of the influencer recommendation framework, where the multi-dimensional heterogeneous data includes a brand data set, an influencer data set, and a cooperation relationship data set;
[0010] Construct an account topic model based on the multi-dimensional heterogeneous data, and perform multi-dimensional feature extraction on each target account based on the account topic model to obtain account topic features, account image features, account industry probability distribution features, account sentiment features, and account attribute features;
[0011] For the account topic features, account image features, account industry probability distribution features, account sentiment features, and account attribute features, project the sparse features from the high-dimensional sparse space to the low-dimensional dense space through the embedding matrix of the index lookup feature embedding layer to obtain sparse feature embedding vectors; use the Z-Score normalization method to process the dense features and splice them with the sparse feature embedding vectors, and use the splicing result as the input of the feature interaction layer of the influencer recommendation framework;
[0012] Based on a deep network and a cross network, simulate the non-linear relationship between features in the input data in the feature interaction layer and extract deep-level information to achieve deep interaction between features and obtain a feature interaction result;
[0013] Calculate the matching prediction values between the brand and each influencer in the feature interaction result through a fully connected layer, sort the influencers according to the matching prediction values, generate a recommended sorted list, and optimize the recommended sorted list using two loss functions, namely sorting loss and matching loss, to obtain the final recommendation result.
[0014] In some embodiments, the influencer recommendation framework includes: an input layer, a feature embedding layer, a feature interaction layer, and a sorting output layer;
[0015] The input layer is used to receive and integrate heterogeneous data inputs of different dimensions;
[0016] The feature embedding layer is used to divide the data into sparse and dense feature sets, and apply an embedding mechanism to map the sparse features to a low-dimensional space and splice them with the dense features;
[0017] The feature interaction layer is used to achieve deep interaction between brand features and influencer features through two sub-networks, namely a cross network and a deep network;
[0018] The sorting output layer is used to calculate the matching scores of each potential influencer through a fully connected layer and generate a recommendation list.
[0019] In some embodiments, the constructing an account topic model according to the multi-dimensional heterogeneous data includes:
[0020] Perform text cleaning, word segmentation, and hashtag extraction on the text data in the brand dataset, influencer dataset, and cooperation relationship dataset to generate a text set and a hashtag set;
[0021] Convert the text set and the image data in each dataset to the same feature space to form corresponding text embeddings and image embeddings;
[0022] Calculate the importance weights and correlation weights of the hashtag set, construct a hashtag weighted matrix, assign differential weights to the image data according to the hashtag weighted matrix to obtain a weighted image set, and perform clustering and dimensionality reduction on the weighted image set to obtain a representative clustering set;
[0023] Calculate the similarity between the text embedding and the image embedding through a cosine function, and determine the account topic set according to the representative clustering set and the similarity.
[0024] In some embodiments, the calculating the importance weights and correlation weights of the hashtag set, constructing a hashtag weighted matrix, and assigning differential weights to the image data according to the hashtag weighted matrix to obtain a weighted image set includes:
[0025] Define the text of a certain post published by a certain account as , and the corresponding set of hashtag labels is . Then for each , represents the total number of times that this label appears in all posts of the entire account. We define the importance weight as , and introduce a smoothing factor to enhance the robustness of the weight. Its calculation is shown in formula (1):
[0026]
[0027] Define the relevance weight as , where represents the number of labels used in the post. Use the pre-trained text model RoBERTa to calculate the text representation and of the representation . Its calculation formula is shown in formula (2):
[0028]
[0029] Finally, based on the two normalized weights and , introduce parameters and , where adjusts the importance weight. The larger the value, the greater the weight of the post usage frequency. On the contrary, adjusts the relevance weight. The higher the value, the greater the relevance weight between the post text content and the label. Under the constraint condition , define the hashtag weighted weight The calculation formula is shown in formula (3):
[0030]
[0031] Adjust the weight of the text to the image by adjusting the values of and .
[0032] In some embodiments, clustering and dimensionality reduction are performed on the weighted image set to obtain a representative clustering set, including:
[0033] For each data point , calculate its average distance from other points in the same cluster and its average distance to the nearest other cluster, the calculation formulas are shown in (4) and (5) respectively:
[0034]
[0035]
[0036] where is the cluster containing points, where is the cluster not containing points, is the distance between the data points and ;
[0037] The calculation formula of the silhouette coefficient is shown in formula (6).
[0038]
[0039] Among them, the overall silhouette coefficient is the average value of the silhouette coefficients of all data points. The closer the value is to 1, the better the clustering effect. During the iterative optimization process, the model gradually approaches the optimal parameter combination to optimize the clustering effect, and finally outputs the optimal clustering parameters and their corresponding silhouette coefficients as the representative clustering set.
[0040] In some embodiments, performing multi-dimensional feature extraction on each of the target accounts based on the account theme model to obtain account theme features, account image features, account industry probability distribution features, account emotion features, and account attribute features, including:
[0041] Processing the image and text sets of each account according to the account theme set to obtain account theme words and theme distributions, and at the same time performing account theme feature extraction on the content published by the account audience to obtain audience theme words and theme distributions;
[0042] Preprocessing the image data set of each account, inputting the preprocessed image into the model to obtain 2048-dimensional features, using global average pooling to compress the feature dimension, and calculating the average value of all image feature vectors under each account to obtain account image features;
[0043] Cleaning and preprocessing the text data of each account, extracting emojis, and calculating the VAD value of the account through dictionary-based emotion feature extraction, emotion feature analysis, emotion classification model application, and emotion polarity adjustment to obtain account emotion features;
[0044] Calculating the number of fans of each account, the average number of comments and likes per post to form account attribute features.
[0045] In some embodiments, for the account theme features, account image features, account industry probability distribution features, account sentiment features, and account attribute features, the sparse features are projected from a high-dimensional sparse space to a low-dimensional dense space through an embedding matrix of an index lookup feature embedding layer to obtain sparse feature embedding vectors; the dense features are processed using the Z-Score normalization method and then concatenated with the sparse feature embedding vectors, and the concatenation result is used as the input to the influencer recommendation framework feature interaction layer, including:
[0046] The numerical features among the account theme features, account image features, account industry probability distribution features, account sentiment features, and account attribute features are classified as dense features, and the text features are classified as sparse features;
[0047] The embedding objective of sparse features is to map a sparse high-dimensional topic word index to a low-dimensional dense vector space. Through the embedding matrix , the corresponding sparse features are projected from a high-dimensional sparse space to a low-dimensional dense space to achieve the mapping, as shown in formulas (7)-(9):
[0048]
[0049]
[0050]
[0051] is the embedding vector of the sparse feature, and represent the vocabulary size and the embedding dimension respectively. For multi-valued features, the average value of the embedding vectors of the sparse features is used as the final vector;
[0052] The processing objective of dense features is mainly to combine them with the embedding results of sparse features through preprocessing and transformation. By calculating the difference between each feature value and the mean of the feature, and then dividing by the standard deviation of the feature for normalization, the final output is the concatenation of all the embedding vectors and the normalized dense features shown in formula (10):
[0053]
[0054] In some embodiments, the matching prediction values between the brand and each influencer in the calculation of the feature interaction result through the fully connected layer are calculated, the influencers are sorted according to the matching prediction values, a recommended sorted list is generated, and two loss functions, namely the sorting loss and the matching loss, are used to optimize the recommended sorted list to obtain the final recommendation result, including:
[0055] Calculate the matching prediction values between the brand and each influencer according to the result of feature interaction , and the calculation formula is as shown in formula (11):
[0056] (11)
[0057] Among them, is the weight vector for log odds, The expression of
[0058] (12)
[0059] Learn the difference between the true value and the predicted value of the model through a custom multi-task loss function , The definition of
[0060] (13)
[0061] Among them represents the matching loss, represents the sorting loss, is the set of all weights of the model, represents the square of the L2 norm of and are the weight weighting parameters of the two losses, is the L2 regularization coefficient;
[0062] Reduce the influence of simple samples on model training, and the implementation of the focal loss is as shown in formula (14):
[0063] (14)
[0064] Among them is the probability that the model predicts as the positive class, calculated from the basic binary cross-entropy loss, is the class balance weight, used to adjust the weights between different classes, is the focusing parameter, controlling the weights of easy and difficult samples;
[0065] The goal of the ranking task is to optimize the relative order between positive and negative samples when the model makes predictions, making the scores of positive samples as high as possible compared to negative samples. Positive and negative samples are randomly sampled according to a preset ratio, and the loss function calculation formula for each pair of positive and negative samples is shown in Equation (15):
[0066] (15)
[0067] Define the average pairwise ranking loss in a training batch as:
[0068] (16)
[0069] where represents the loss function for each pair of positive and negative samples, represents the score of the positive sample, represents the score of the negative sample.
[0070] In a second aspect, the present application proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0071] In a third aspect, the present application proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0072] Advantages of the present invention:
[0073] 1. More comprehensive and accurate account feature representation: Existing technologies mainly rely on text or image features, ignoring factors such as sentiment tendency and industry attributes. This technology constructs a more three-dimensional account feature system through the account topic model (SMATM) and multi-dimensional feature extraction, improving the recommendation accuracy.
[0074] 2. Effectively capture the complex interaction relationships between multi-dimensional features: This technology decomposes the interaction relationships of multi-dimensional features into two subtasks of matching and ranking. Through cross networks and deep network structures, it makes up for the deficiencies of traditional methods in feature interaction and capturing non-linear relationships, improving the accuracy and personalization of recommendations.
[0075] 3. Improve the interpretability and adaptability of the recommendation system: By integrating deep learning methods and social interaction data, this technology makes the recommendation process more transparent, enabling brands to better understand the reasons for recommendations. At the same time, it has good adaptability and can cope with the continuous changes in social media content.
[0076] 4. Optimize the sorting and screening of recommended results: Through two loss functions, namely sorting loss and matching loss, this technology optimizes the order and relevance of the recommended list, helps brands find the most suitable cooperative influencers faster, and improves the efficiency and effectiveness of marketing activities. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 It is the overall flowchart of the present invention.
[0078] Figure 2 It is a schematic diagram of the influencer recommendation framework.
[0079] Figure 3 It is a schematic diagram of the calculation of the cross layer and the depth layer. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0080] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein; on the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be completely conveyed to those skilled in the art.
[0081] In a first aspect, the present application proposes a multi-dimensional feature interaction influencer recommendation method based on an account topic model, as Figure 1 shown, including the following steps:
[0082] S100: Receive and integrate multi-dimensional heterogeneous data of each target account through the input layer of the influencer recommendation framework, where the multi-dimensional heterogeneous data includes a brand data set, an influencer data set, and a cooperation relationship data set;
[0083] In some embodiments, the influencer recommendation framework includes: an input layer, a feature embedding layer, a feature interaction layer, and a sorting output layer;
[0084] The input layer is used to receive and integrate heterogeneous data inputs of different dimensions;
[0085] The feature embedding layer is used to divide the data into sparse and dense feature sets, and apply an embedding mechanism to map the sparse features to a low-dimensional space and splice them with the dense features;
[0086] The feature interaction layer is used to achieve deep interaction between brand features and influencer features through two sub-networks, namely a cross network and a depth network;
[0087] The sorting output layer is used to calculate the matching scores of each potential influencer through a fully connected layer and generate a recommended list.
[0088] Among them, as Figure 2As shown in the figure, the input layer receives and integrates heterogeneous data inputs of different dimensions. In order to improve the representation ability of the multi-dimensional features of the account, the present invention first innovatively proposes to use the account theme model SMATM of the account subject to mine the account theme information. In addition, the present invention also uses different feature extraction methods based on multiple dimensions such as the sentiment dimension, industry attribute, and audience attribute to obtain a more accurate representation of the account.
[0089] The feature embedding layer divides the data into sparse and dense feature sets, and applies an embedding mechanism to map the sparse features to a low-dimensional space, which is concatenated with the dense features for use by the feature interaction layer.
[0090] The feature interaction layer is the core part of the influencer recommendation framework (MFI-IR) to achieve feature interaction. It realizes the deep interaction between brand features and influencer features through two sub-networks, namely the cross network and the deep network. The cross network simulates the non-linear relationship between features through multiple layers of interaction units, while the deep network extracts the deep information of features through stacked hidden layers.
[0091] After the feature interaction layer, the model enters the sorting and output stage. Through the fully connected layer, the model calculates the matching scores of each potential influencer. These scores are then used to generate a recommendation list. In order to optimize the recommendation results, the model adopts two loss functions, namely the sorting loss and the matching loss. The sorting loss is used to ensure the consistency of the order of the recommendation list with the user's preferences, while the matching loss is used to measure the matching degree between the recommended influencer and the brand's needs.
[0092] S200: Construct an account theme model SMATM according to the multi-dimensional heterogeneous data, and perform multi-dimensional feature extraction on each target account based on the account theme model to obtain account theme features, account image features, account industry probability distribution features, account sentiment features, and account attribute features;
[0093] In some embodiments, the constructing an account theme model according to the multi-dimensional heterogeneous data includes:
[0094] Perform text cleaning, word segmentation, and hashtag extraction on the text data in the brand dataset, influencer dataset, and cooperation relationship dataset to generate a text set and a hashtag set;
[0095] Convert the text set and the image data in each dataset to the same feature space to form corresponding text embeddings and image embeddings;
[0096] Calculate the importance weight and relevance weight of the hashtag set, construct a hashtag weighted matrix, assign differential weights to the image data according to the hashtag weighted matrix, obtain a weighted image set, and perform clustering and dimensionality reduction on the weighted image set to obtain a representative clustering set;
[0097] Calculate the similarity between the text embedding and the image embedding through a cosine function, and determine the account topic set according to the representative clustering set and the similarity.
[0098] Among them, when constructing the account topic model, two key problems need to be solved: one is how to effectively extract and fuse the multi-modal features of text and images to achieve accurate topic modeling; the other is how to handle data sparsity to ensure that the extracted topics are both representative and interpretable. Therefore, the present invention proposes an account topic model SMATM, the core of which is to use feature extraction and clustering algorithms to extract topics from multi-modal account data. The model input data includes text and images , and data quality is the basis for topic extraction. The data preprocessing stage includes text cleaning, word segmentation, and hashtag extraction to generate a text set and a hashtag set . Based on the BERTopic and CLIP models, text and images are transformed into the same feature space to form corresponding embeddings and . To improve the accuracy of the clustering results and reduce data sparsity, a parameter learning module is introduced to adapt the best dimensionality reduction and clustering parameters for the learned image data of each account . At the same time, for hashtag data, considering its importance and relevance to the post text, a hashtag weighted matrix is constructed to assign differential weights to different posts. Under the guidance of the best parameters, clustering and dimensionality reduction are performed on the weighted image set to extract a representative clustering set to distinguish data categories. In the text embedding stage, the similarity between the text embedding and the image representative embedding is calculated through a cosine function to obtain the best description of the image data by the text in the post, and these descriptions constitute the topic set of each account .
[0099] In some embodiments, the calculating the importance weight and relevance weight of the hashtag set, constructing a hashtag weighted matrix, assigning differential weights to the image data according to the hashtag weighted matrix, and obtaining a weighted image set includes:
[0100] Define the text of a certain post published by an account as , the corresponding set of hashtag labels is , then for each , represents the total number of times the label appears in all posts of the entire account. We define the importance weight as , and introduce a smoothing factor to enhance the robustness of the weight, and its calculation is shown in formula (1):
[0101]
[0102] Define the relevance weight as , where represents the number of tags used in the post. Use the pre-trained text model RoBERTa to calculate the text representation of a certain post and 's representation , and its calculation formula is shown in formula (2):
[0103]
[0104] Finally, based on the two normalized weights and , introduce parameters and , where adjusts the importance weight. The larger the value, the greater the weight of the post usage frequency. On the contrary, adjusts the relevance weight. The higher the value, the greater the relevance weight between the post text content and the tag. Under the constraint condition , define the hashtag weighted weight The calculation formula is shown in formula (3):
[0105]
[0106] By adjusting the and values to adjust the weight of the text to the image.
[0107] Among them, when mining the theme of the account content, the hashtag plays a key role. We believe that: the higher the relevance between the hashtag and the text, the more prominent the theme of the post; the more content a user posts under a specific hashtag, the more the posts under this hashtag are close to the theme of the account. Based on this, a hashtag weighting module is introduced to learn the text weight of the hashtag and provide weighting for image clustering.
[0108] In some embodiments, clustering and dimensionality reduction of the weighted image set to obtain a representative clustering set includes:
[0109] For each data point , calculate its average distance from other points in the same cluster and the average distance to the nearest other cluster , and the calculation formulas are shown in (4) and (5) respectively:
[0110]
[0111]
[0112] Where is the cluster containing points, where is the cluster not containing points, is the data point and the distance between;
[0113] The calculation formula of the silhouette coefficient is shown in formula (6).
[0114] (6)
[0115] Where the overall silhouette coefficient is the average value of the silhouette coefficients of all data points. The closer the value is to 1, the better the clustering effect. During the iterative optimization process, the model gradually approaches the optimal parameter combination to optimize the clustering effect, and finally outputs the optimal clustering parameters and their corresponding silhouette coefficients as the representative clustering set.
[0116] Among them, in order to alleviate the data sparsity problem faced by traditional multi-modal topic models when facing account data, a clustering parameter learning module is used to evaluate different combinations of clustering parameters before clustering, and the parameters are selected and optimized. In this module, UMAP is used for data dimensionality reduction, and HDBSCAN is used for clustering analysis. Combining the empirical values of UMAP and HDBSCAN parameters in related work, and the sampling experiment results for the scale of social media account data, for the key parameters number of neighboring nodes and minimum distance parameter as well as the minimum size grouping of clustering clusters and the clustering density parameter , during the clustering parameter learning process, the model gradually explores the influence of parameters on the clustering result. After each clustering, the silhouette coefficient is calculated through the CEM-SS module.
[0117] S300: Project the account theme features, account image features, account industry probability distribution features, account sentiment features, and account attribute features from the high-dimensional sparse space to the low-dimensional dense space through the embedding matrix of the index search feature embedding layer to obtain sparse feature embedding vectors; use the Z-Score normalization method to process the dense features and splice them with the sparse feature embedding vectors, and use the splicing result as the input of the influencer recommendation framework feature interaction layer;
[0118] In some embodiments, the multi-dimensional feature extraction of each target account based on the account theme model to obtain account theme features, account image features, account industry probability distribution features, account sentiment features, and account attribute features includes:
[0119] Process the image and text sets of each account according to the account theme set to obtain account theme words and theme distributions, and at the same time perform account theme feature extraction on the content published by the account audience to obtain audience theme words and theme distributions;
[0120] Preprocess the image data set of each account, input the preprocessed images into the model to obtain 2048-dimensional features, use global average pooling to compress the feature dimensions, and calculate the average value of all image feature vectors under each account to obtain account image features;
[0121] Clean and preprocess the text data of each account, extract emojis, and calculate the VAD value of the account through dictionary-based sentiment feature extraction, sentiment feature analysis, application of sentiment classification models, and sentiment polarity adjustment to obtain account sentiment features;
[0122] Calculate the number of fans of each account, the average number of comments and likes per post to form account attribute features.
[0123] Among them, the core of the account theme feature extraction process is implemented by applying the SMATM framework proposed in the present invention. The input of SMATM is the image and text sets composed of accounts in the data set. By calculating the representation embeddings of the image texts in these sets in the same space, and then performing image clustering based on the image embeddings under the text weights, representative image clusters are obtained. Subsequently, in the text content, find the text with the highest similarity to the representative image representation to represent the core theme words of the account content. In the account theme feature extraction link, account audience theme extraction is performed on the text and image sets of the audience accounts collected centered on the brand or influencer account, and the theme analysis results of the target account audience are obtained to enhance the dimensions of the account features. As the input of the final influencer recommendation framework, the theme features of the brand or influencer account and its audience account set will be presented in the form of account or audience theme words and theme distributions.
[0124] Image feature extraction is also one of the core tasks of image analysis. In the present invention, a model based on ResNeXt WSL is selected for account image feature extraction and industry feature prediction. As a derivative variant of ResNeXt, the ResNeXt WSL model is trained using a weakly supervised learning strategy, that is, by mining potential features and patterns using a large-scale Instagram image dataset that is not fully annotated, and then fine-tuning on the fully annotated ImageNet dataset to adapt to task requirements and improve classification accuracy, making it flexible and efficient in processing large-scale datasets. At the same time, during the training process of the ResNeXt WSL model, the concept of "cardinality" is introduced, that is, the number of parallel paths in the network. Specifically, first, the images in the dataset are re-integrated into an account dataset according to the account. During the feature extraction process, the social media image dataset in units of accounts is first preprocessed. Each image undergoes normalization processing, specifically including adjusting the image size, cropping, and normalization. The preprocessed images are input into the pre-trained model resnext101_32x8d_wsl to obtain 2048-dimensional features of the social media images, which can comprehensively reflect the main shape, texture, and color features of the images. To obtain the image feature representation of the account, first, global average pooling is used to compress the feature dimension of each image, and then the average value of all image feature vectors under each account is calculated to obtain a single vector representing the account image features;
[0125] In the process of feature extraction in the present invention, the image feature extraction model is also adjusted based on the research on the industry information and image classification method of brand accounts, and used as a classification model. The classification model is trained using a set of brand images with industry attributes, and the image set of influencer accounts is analyzed on the trained model to provide the industry feature distribution of account image information for auxiliary recommendation decisions for candidate influencer accounts, thereby enhancing the interpretability of recommendation decisions. Specifically, the calculation of industry features depends on the output of the image classification model. The images posted by influencer accounts are classified, and a set of industry features is generated for each account. To achieve this goal, the image classification model is adjusted to enable the model to predict the probability distribution of industry-related labels. In this task, based on the image feature extraction model ResNeXt WSL, modifications are made, the output layer of the model is adjusted, the number of neurons in the fully connected layer is increased to 12, and the softmax activation function is used to output the probability value of each industry feature. During the model fine-tuning training process, the brand account data in the dataset is used, with the brand category as the label, and it is divided into a training set and a validation set. Subsequently, based on the fine-tuned model, all images of each account are classified, and the probability distribution of each photo under the account is averaged to obtain the image industry probability distribution of the influencer account.
[0126] In the influencer recommendation and matching task, the emotional dimension of an account is equally crucial. The basic emotion theory in the field of emotion computing holds that humans are born with a set of basic emotions. This study uses the three-dimensional emotion theory in emotion psychology - the VAD model - to quantitatively analyze emotional expression and intensity. The VAD model argues that emotions can be described through three dimensions: emotional polarity (Valence), arousal level (Arousal), and dominance (Dominance). Specifically, emotional polarity reflects the "good or bad" or "positive or negative" tendency of an emotion, usually represented by a continuous spectrum from positive to negative. Through the emotional polarity dimension, the emotional tendency of an emotion can be identified: positive indicates a positive emotion, and negative indicates a negative emotion. The arousal level measures the intensity of an emotion or the strength of an emotional reaction, which describes the physiological arousal level of an emotional experience, from low to high. This dimension helps to understand the intensity or energy of an emotional experience: high arousal represents intense emotions, and low arousal indicates calm emotions. Dominance assesses the sense of control or dominance in an emotional experience. Low-dominance emotions indicate passivity or being controlled, while high-dominance emotions convey a sense of mastery or control. High-dominance emotions are usually associated with a positive self-perception, while low-dominance emotions are associated with feelings of passivity and vulnerability. In terms of algorithm implementation, the calculation of the account's emotional characteristics is not only based on the text content but also incorporates the emotional influence of an emotion dictionary, emojis, and the prediction of an emotion classification model to form a multi-dimensional emotional analysis index - the VAD value. To effectively conduct emotional analysis, the account's text data is first cleaned and preprocessed, and the emojis in the posts are extracted. The core modules of emotional feature calculation include dictionary-based emotional feature extraction, emotional feature analysis, application of an emotion classification model, and emotional polarity adjustment, etc.
[0127] S400: Based on a deep network and a cross network, simulate the non-linear relationship between features in the input data in the feature interaction layer and extract deep information to achieve deep interaction between features and obtain a feature interaction result;
[0128] In some embodiments, the method of projecting the sparse features from a high-dimensional sparse space to a low-dimensional dense space to obtain sparse feature embedding vectors by looking up the embedding matrix of the feature embedding layer through the account theme features, account image features, account industry probability distribution features, account emotional features, and account attribute features; and using the Z-Score normalization method to process the dense features and then concatenate them with the sparse feature embedding vectors, and using the concatenation result as the input of the feature interaction layer of the influencer recommendation framework includes:
[0129] Divide the numerical features in the account theme features, account image features, account industry probability distribution features, account emotional features, and account attribute features into dense features, and divide the text features into sparse features;
[0130] The embedding objective of sparse features is to map the sparse high-dimensional topic word index into a low-dimensional dense vector space. Through the embedding matrix , the corresponding sparse features are projected from the high-dimensional sparse space to the low-dimensional dense space to achieve the mapping, as shown in formulas (7)-(9):
[0131]
[0132]
[0133]
[0134] is the embedding vector of sparse features, and represent the vocabulary size and the embedding dimension respectively. For multi-valued features, the average value of the embedding vectors of sparse features is used as the final vector;
[0135] Among them, the input layer takes the brand data , influencer data and the cooperation relationship data between the two as the original input.
[0136] Taking the social media account as the main body, based on the multi-dimensional characterization method of the account, the topic features of the brand account and the influencer account are calculated , the image features are extracted , and the industry probability distribution and the sentiment dimension value are calculated.
[0137] In addition, we calculate the number of fans of the account, the average number of comments and likes per post, and form the account attribute features .
[0138] Subsequently, using the account-related features as dictionary values, the brand and influencer feature dictionaries are initialized. By generating all possible combinations of brand and influencer IDs, the feature pairs between the brand and the influencer are constructed. Each combination of features will be labeled as a "positive sample" or a "negative sample", and this cooperation relationship is recorded. During the generation of feature pairs, each brand and influencer feature value is assigned a unique ID and stored in the feature dataset. During the storage process, the model associates the feature values with the brand ID and the influencer ID. After completing the calculation, storage and annotation of all features, the algorithm generates the final set of feature pairs , and uses it as the input of the model for subsequent learning and optimization phases.
[0139] Followed by the feature embedding layer. First, the features of the input layer brands and influencers, as shown in Table 1, are divided into sparse features and dense features according to their characteristics.
[0140] Table 1
[0141]
[0142] The main processing objective of dense features is to combine their embedding results with those of sparse features through preprocessing and transformation. By calculating the difference between each feature value and the mean of the feature, and then dividing by the standard deviation of the feature for standardization, the final output is the concatenation of all embedding vectors and the normalized dense features shown in formula (10):
[0143]
[0144] The feature interaction layer is the key to learning effective feature interactions. In the influencer recommendation research framework proposed in the present invention, to achieve meaningful interactions between brand features and influencer features, the feature interaction layer is based on the deep network and cross network modules in DCN_V2 to construct the cross network and deep network modules in the MFI-IR recommendation model. Figure 2 The calculation processes of a cross layer and a deep layer are respectively shown. When constructing the recommendation framework, for different data sets, the number and combination method of the cross layer and the deep layer can be dynamically adjusted according to their data characteristics.
[0145] S500: Calculate the matching prediction values between the brand and each influencer in the feature interaction result through a fully connected layer, sort the influencers according to the matching prediction values to generate a recommended sorted list, and optimize the recommended sorted list using two loss functions, namely the sorting loss and the matching loss, to obtain the final recommendation result.
[0146] In some embodiments, the step of calculating the matching prediction values between the brand and each influencer in the feature interaction result through a fully connected layer, sorting the influencers according to the matching prediction values to generate a recommended sorted list, and optimizing the recommended sorted list using two loss functions, namely the sorting loss and the matching loss, to obtain the final recommendation result, includes:
[0147] Calculate the matching prediction values between the brand and each influencer according to the result of feature interaction , and the calculation formula is as shown in formula (11):
[0148] (11)
[0149] Wherein, is the weight vector for log odds, The expression is as shown in formula (12):
[0150] (12)
[0151] To further optimize the recommendation results, it is crucial to design an appropriate loss function. Combining with the influencer recommendation task, the model loss function should consider the effects of both the ranking task and the matching task, that is, not only the accuracy of the predicted values needs to be concerned, but also the relative order of the prediction results should be optimized. Therefore, during model training, through a custom multi-task loss function learn the difference between the true values and the predicted values of the model, is defined as shown in formula (13):
[0152] (13)
[0153] where represents the matching loss, represents the ranking loss, is the set of all weights of the model, represents the square of the L2 norm of and are the weight weighting parameters for the two losses, is the L2 regularization coefficient;
[0154] The matching task can be regarded as a binary classification prediction problem. Considering that the number of positive samples in the sample dataset is very small and the ratio of positive and negative samples is extremely unbalanced, the focal loss is used in this part of the loss function to optimize the model matching effect. The focal loss reduces the impact of simple samples on model training by adjusting the basic binary cross-entropy loss, thereby improving the model's learning ability for difficult samples. The implementation of the focal loss is as shown in formula (14):
[0155] (14)
[0156] where is the probability that the model predicts as the positive class, calculated from the basic binary cross-entropy loss, is the class balance weight, used to adjust the weights between different classes, is the focusing parameter, controlling the weights of easy and difficult samples;
[0157] The goal of the ranking task is to optimize the relative order between positive and negative samples during model prediction, making the scores of positive samples as high as possible higher than those of negative samples. Positive and negative samples are randomly sampled according to a preset ratio, and the calculation formula for the loss function of each pair of positive and negative samples is as shown in formula (15):
[0158] (15)
[0159] Define the average pairwise ranking loss in a training batch as:
[0160] (16)
[0161] where represents the loss function for each pair of positive and negative samples, represents the positive sample score, represents the negative sample score. In the ranking task, the model needs to optimize the relative order between positive and negative samples, making the positive sample score as high as possible compared to the negative sample. The loss of this pair of samples is measured by calculating a value related to the difference between the positive and negative sample scores, using a form like formula (15), where margin represents the set interval value, and a loss will occur when the difference between the positive sample score and the negative sample score is less than the interval.
[0162] Furthermore, the experimental parameters of the recommendation framework model and the results of the comparative experiments are as follows:
[0163] 1. MFI-IR model parameters and settings:
[0164] The recommendation effect of the MFI-IR model is affected by multiple factors. The most crucial ones are the structure of the five feature interaction layers (see Table 2 in detail) and the key parameters of the model. Table 2 details the definitions, descriptions of each symbol, and the value ranges of the parameters used in the experiment.
[0165] Table 2 Structure of the feature interaction layer of the recommendation algorithm and its related key parameters
[0166]
[0167] Based on the key experimental parameters of the recommendation algorithm included in the table, the present invention uses the grid search method in the experiment to achieve parameter exploration and optimization. During the grid search process, the parameters are combined into multiple different configurations and are successively used for model training. For each parameter configuration, cross-validation is implemented to evaluate the model performance, thereby ensuring the reliability of the optimization results. The model training process is all based on a unified training set and validation set, and the performance of the model is comprehensively evaluated with reference to five core indicators: AUC, CAUC, P@10, MRR, and MAP.
[0168] Taking into account the trade-off between global indicators and local ranking indicators, when the experimental parameters are determined as shown in Table 3, the comprehensive performance of the model is optimal, achieving a good balance between feature expression ability and model generalization.
[0169] Table 3 Optimal parameter configuration of the MFI-IR model
[0170]
[0171] 2. Baseline Experiments and Experimental Data:
[0172] Baseline Models:
[0173] To verify the effectiveness of the influencer recommendation algorithm proposed in the present invention, the present invention selects eight representative baseline models for comparative analysis, covering three categories: traditional recommendation methods, multi-modal learning methods, and conceptual representation methods:
[0174] (1) RAND: A random recommendation benchmark model that establishes a null hypothesis baseline by assigning random ranking scores to all brand-influencer pairs, used to verify the recommendation effectiveness of other methods.
[0175] (2) MIV: An influence evaluation model. The present invention introduces an activity factor module into its original architecture to enhance its adaptability to the brand-influencer matching scenario by dynamically adjusting the influence weight.
[0176] (3) MIR: A multi-modal ranking learning model that constructs a brand-influencer correlation evaluation matrix through heterogeneous data fusion and uses a collaborative attention mechanism to achieve cross-modal feature alignment.
[0177] (4) HPTR: A multi-modal historical pooling triple model that combines the pooling representation learning of historical interaction data with triple loss optimization to capture the dynamic interaction pattern between brands and influencers through time series modeling.
[0178] (5) Wsim: A multi-task similarity ranking framework that uses a trainable inner product space to calculate visual-text bimodal similarity, jointly optimizes the participation prediction and ranking tasks, and realizes the joint embedding of the feature space.
[0179] (6) MOR: A multi-dimensional portrait ranking model that constructs a three-dimensional representation system of individual images, target audiences, and cooperation preferences based on a cross-modal co-attention mechanism, and improves the recommendation accuracy by designing a multi-dimensional fusion ranking function.
[0180] (7) CATR: A concept alignment triple model that uses a semantic mapping method based on a concept space to optimize the triple loss under ontology constraints, strengthening the semantic consistency between brand requirements and influencer features.
[0181] (8) CAMERA: A conceptual influencer ranking framework that realizes deep semantic matching based on the concept hierarchy structure by establishing a brand-influencer concept graph and using a graph attention network for concept node propagation.
[0182] (9) MFI-IR(mmLDA): The influencer recommendation algorithm proposed in the present invention replaces the account topic feature extraction method with mmLDA used by other existing influencer recommendation methods. Experimental Results and Analysis:
[0183] To verify the effectiveness of the MFI-IR model proposed in the present invention, a number of comparative experiments of recommendation system benchmark models were conducted, and the results are shown in Table 4. In the experiment, to ensure the comprehensiveness and fairness of the experiment, each model was trained and tested on the same dataset, and the evaluation index system introduced above was used. In terms of model configuration, the optimal parameter configuration of the MFI-IR model is shown in Table 3 in the above text.
[0184] Table 4 Comparison between MFI-IR Model and Baseline Models
[0185]
[0186] The experiment shows that the MFI-IR model proposed in the present invention has demonstrated significant performance improvement in multiple key evaluation indicators, and shows stronger discrimination ability and recommendation accuracy compared with traditional baseline models.
[0187] Specifically, in terms of the accuracy indicators related to the recommendation matching task, the AUC value of MFI-IR is 0.9371, which is about 6% higher than 0.884 of the MOR model. This improvement indicates that MFI-IR can more effectively distinguish positive samples from negative samples, has stronger discriminant ability, and thus provides more accurate recommendation results for users. The CAUC value of MFI-IR is 0.7841, which is about 7.8% higher than 0.727 of MOR. CAUC further examines the ranking of positive and negative samples within the same category on the basis of AUC. The improvement of MFI-IR reflects that the model has stronger ranking ability within the same category and can more accurately judge the matching degree between the brand and the influencer. The P@10 of MFI-IR reaches 0.7191, which is significantly higher than 0.019 and 0.005 of the MIV and RAND models, and the improvement amplitude is close to 35 times. This shows that MFI-IR can accurately screen out relevant positive samples among the top 10 influencers in the recommendation list, greatly improving the practicality and accuracy of the recommendation, especially suitable for situations where users hope to quickly obtain accurate recommendations.
[0188] In terms of the metrics related to recommendation ranking, the MAP value of MFI-IR is 0.9079, which is significantly higher than 0.238 of MOR. The improvement amplitude is about 3.8 times, indicating that the overall performance of MFI-IR in multiple queries is more accurate and can provide higher-quality recommendations in the brand and influencer matching tasks. The MRR value of MFI-IR is 0.328. Although it is lower than 0.602 of CAMERA and 0.565 of MOR, it is still better than other models. This shows that although there is still room for improvement in MFI-IR in this metric, compared with other baseline models, it can still ensure that relevant positive samples appear at a higher position in the recommendation list, thereby improving the user's click-through rate and satisfaction. In addition, the MedR value of MFI-IR is 3. Although it does not reach 2 of the existing solutions, it is much lower than 54 of RAND and 49 of MIV. This means that the positive samples are in an earlier position in the recommendation list, and relevant influencers can be presented to users more quickly, further improving the efficiency of the recommendation system and the user experience.
[0189] In addition, compare the situation of using the account theme model SMATM proposed in the present invention for account theme extraction and using the basic multi-modal LDA model mentioned in the existing method for account theme extraction. From the experimental results, the MFI-IR model using SMATM is superior to the MFI-IR(mmLDA) model using mmLDA in all metrics. The AUC value is increased from 0.9113 to 0.9371, the CAUC value is increased from 0.7631 to 0.7841, P@10 is increased from 0.6402 to 0.7191, and MAP is increased from 0.884 to 0.9079. This shows that the SMATM model proposed in the present invention can more effectively extract account theme features, provide more accurate input for the MFI-IR model, thereby further improving the overall performance of the model and making it perform more excellently in recommendation matching and ranking.
[0190] Generally speaking, compared with the traditional recommendation system model, MFI-IR demonstrates higher recommendation efficiency through precise optimization of the sorting of positive and negative samples and similar positive and negative samples, can provide more personalized and accurate recommendation services for users, and improves the overall quality of brand and influencer matching.
[0191] In a second aspect, the present application proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0192] In a third aspect, the present application proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0193] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0194] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0195] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.
[0196] In the embodiments provided in this disclosure, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there can be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the device or unit can be in electrical, mechanical or other forms.
[0197] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0198] In addition, in each embodiment of the present disclosure, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0199] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, all or part of the processes in the above-described embodiment methods of the present disclosure may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments may be implemented. The computer program may include computer program code, and the computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0200] The above are only the preferred embodiments of the present invention. It should be noted that for those skilled in the art, without departing from the premise of the technical solution of the present invention, several modified and improved technical solutions should also be regarded as falling within the scope protected by the claims of the present invention.
Claims
1. A multi-dimensional feature interactive influencer recommendation method based on an account topic model, characterized by: The following steps are involved: Receiving and integrating multi-dimensional heterogeneous data of each target account through an input layer of an influencer recommendation framework, wherein the multi-dimensional heterogeneous data includes a brand dataset, an influencer dataset, and a partnership dataset; Constructing an account topic model according to the multi-dimensional heterogeneous data, including performing text cleaning, word segmentation and hashtag extraction on the text data in the brand dataset, influencer dataset and partnership dataset to generate a text set and a hashtag set; Convert the text set and the image data in each data set into the same feature space to form corresponding text embedding and image embedding; Calculating the importance weight and relevance weight of the hashtag set, constructing a hashtag weighting matrix, assigning differentiated weights to the image data according to the hashtag weighting matrix, obtaining a weighted image set, clustering and reducing the dimension of the weighted image set, and obtaining a representative clustering set; Calculating the similarity between the text embedding and the image embedding by using a cosine function, and determining an account topic set according to the representative cluster set and the similarity; Based on the account topic model, multi-dimensional features are extracted for each target account to obtain account topic features, account image features, account industry probability distribution features, account sentiment features, and account attribute features; The account theme features, account image features, account industry probability distribution features, account sentiment features and account attribute features are projected from high-dimensional sparse space to low-dimensional dense space by indexing the embedding matrix of the feature embedding layer to obtain a sparse feature embedding vector; the dense features are processed using the Z-Score normalization method and then spliced with the sparse feature embedding vector, and the splicing result is used as the input of the feature interaction layer of the influencer recommendation framework; Based on the deep network and the cross network, the nonlinear relationship between the features of the input data is simulated in the feature interaction layer and the deep information is extracted to achieve the deep interaction between the features and obtain the feature interaction result; The matching prediction value between the brand and each influencer in the feature interaction result is calculated through a fully connected layer, the influencers are sorted according to the matching prediction value, a recommended sorting list is generated, and the sorting loss and matching loss loss functions are used to optimize the recommended sorting list to obtain the final recommendation result.
2. The method according to claim 1, characterized in that: The influencer recommendation framework includes: an input layer, a feature embedding layer, a feature interaction layer, and a ranking output layer; The input layer is used to receive and integrate heterogeneous data inputs of different dimensions; The feature embedding layer is used to divide the data into sparse and dense feature sets, and apply an embedding mechanism to map the sparse features to a low-dimensional space and concatenate them with the dense features; The feature interaction layer is used to achieve deep interaction between brand features and influencer features through two sub-networks, a cross network and a deep network; The ranking output layer is used to calculate the matching score of each potential influencer through the fully connected layer and generate a recommendation list.
3. The method according to claim 2, characterized in that: The calculating the importance weight and the relevance weight of the hashtag set, constructing a hashtag weighting matrix, assigning differentiated weights to the image data according to the hashtag weighting matrix, and obtaining a weighted image set includes: Define the text of a post published by an account as , the corresponding hashtag tag set is , then for each , represents the total number of times the tag appears in all posts of the entire account. We define the importance weight as , and introduce the smoothing factor To improve the robustness of the weight, its calculation is shown in formula (1): Define the relevance weight as ,in Represents the number of tags used in the post, and uses the pre-trained text model RoBERTa to calculate the text representation of a post and hashtag Characterization , and its calculation formula is shown in formula (2): Finally, based on the two normalized weights and , introduce parameters and ,in Adjust the importance weight. The larger the value, the greater the weight of the post usage frequency. On the contrary, Adjust the relevance weight. The higher the value, the greater the relevance weight between the post text content and the tag. Next, define the hashtag weight The calculation formula is shown in formula (3): By adjusting and The value of adjusts the weight of the text over the image.
4. The method according to claim 3, characterized in that: The step of clustering and reducing the dimension of the weighted image set to obtain a representative cluster set includes: For each data point , calculate the average distance between it and other points in the same cluster and the average distance to the nearest other cluster , the calculation formulas are shown in (4) and (5) respectively: in is included A cluster of points, where Does not include Clusters of points, It is a data point and The distance between The calculation formula of the silhouette coefficient is shown in formula (6); Among them, the overall silhouette coefficient It is the average value of the silhouette coefficient of all data points. The closer the value is to 1, the better the clustering effect is. During the iterative optimization process, the model gradually approaches the optimal parameter combination to optimize the clustering effect, and finally outputs the optimal clustering parameters and their corresponding silhouette coefficients as a representative clustering set.
5. The method according to claim 4, characterized in that: The multi-dimensional feature extraction is performed on each target account based on the account topic model to obtain account topic features, account image features, account industry probability distribution features, account sentiment features and account attribute features, including: Processing the image and text sets of each account according to the account theme set to obtain account theme words and theme distribution, and extracting account theme features from the content published by the account audience to obtain audience theme words and theme distribution; Preprocess the image dataset of each account, input the preprocessed images into the model to obtain 2048-dimensional features, use global average pooling to compress the feature dimensions, calculate the average of all image feature vectors under each account, and obtain the account image features; Clean and preprocess the text data of each account, extract emoticons, and calculate the VAD value of the account through dictionary-based sentiment feature extraction, sentiment feature analysis, sentiment classification model application and sentiment polarity adjustment to obtain the account sentiment features; The number of followers of each account, the average number of comments and likes for each post are calculated to form the account attribute characteristics.
6. The method according to claim 5, characterized in that: The account subject features, account image features, account industry probability distribution features, account sentiment features and account attribute features are projected from a high-dimensional sparse space to a low-dimensional dense space by indexing the embedding matrix of the feature embedding layer to obtain a sparse feature embedding vector; The dense features are processed by using a Z-Score normalization method and then concatenated with the sparse feature embedding vector, and the concatenated result is used as the input of the feature interaction layer of the influencer recommendation framework, including: The numerical features among the account topic features, account image features, account industry probability distribution features, account sentiment features, and account attribute features are classified as dense features, and the text features are classified as sparse features; The goal of sparse feature embedding is to map the sparse high-dimensional topic word index into a low-dimensional dense vector space through the embedding matrix , the corresponding sparse features are projected from the high-dimensional sparse space to the low-dimensional dense space to achieve mapping, as shown in formulas (7)-(9): is the embedding vector of sparse features, and They represent the vocabulary size and embedding dimension respectively. For multi-valued features, the average of the embedding vectors of sparse features is used as the final vector; The main goal of dense feature processing is to combine it with the embedding results of sparse features through preprocessing and conversion. The difference between each feature value and the mean of the feature is calculated and then divided by the standard deviation of the feature for standardization. The final output is the concatenation of all embedded vectors and normalized dense features as shown in formula (10): 。 7. The method according to claim 6, characterized in that: The matching prediction value between the brand and each influencer in the feature interaction result is calculated through the fully connected layer, the influencers are sorted according to the matching prediction value, a recommended sorting list is generated, and the recommended sorting list is optimized by using two loss functions, namely, sorting loss and matching loss, to obtain the final recommendation result, including: Calculate the match prediction value between the brand and each influencer based on the results of feature interactions , the calculation formula is shown in formula (11): (11) in, is the weight vector for the log-odds, The expression of is shown in formula (12): (12) Through a custom multi-task loss function Learn the difference between the model's true value and its predicted value, The definition of is shown in formula (13): (13) in represents the matching loss, represents the sorting loss, is the set of all weights of the model, express The square of the L2 norm of and is the weighting parameter of the two losses, is the L2 regularization coefficient; To reduce the impact of simple samples on model training, the implementation of focal loss is shown in formula (14): (14) in is the probability of the model predicting the positive class, which is calculated by the basic binary cross loss entropy. is the category balance weight, which is used to adjust the weights between different categories. is the focusing parameter, which controls the weight of difficult and easy samples; The goal of the sorting task is to enable the model to optimize the relative order between positive and negative samples during prediction, so that the score of the positive sample is as high as possible than that of the negative sample. Positive and negative samples are randomly sampled according to a preset ratio, and the loss function calculation formula for each pair of positive and negative samples is defined as shown in formula (15): (15) The average pairwise ranking loss in a training batch is defined as: (16) in, Represents the loss function of each pair of positive and negative samples, represents the positive sample score, Represents the negative sample score.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-dimensional collision identification method for multi-platform virtual identity account
CN111160130A
Social recommendation method based on cross-topic contrastive learning
WO2025011132A1