Marketing expert precise matching and intelligent decision system based on multi-modal data mining
By constructing dynamic profiles through heterogeneous feature extraction and multimodal fusion networks, the problem of insufficient multimodal data processing capabilities is solved, the objectivity of marketing expert selection and the transparency of decision-making are improved, and highly credible intelligent decision support is provided.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU YUEZHENG NETWORK INFORMATION TECH CO LTD
- Filing Date
- 2026-02-26
- Publication Date
- 2026-05-29
AI Technical Summary
Existing marketing influencer matching methods cannot effectively handle multimodal data, lack cross-modal association modeling capabilities, and the recommendation results lack interpretability, making it difficult to achieve systematic and quantitative intelligent decision-making.
A heterogeneous feature extraction module is used to extract features from multimodal data. Dynamic profiles are constructed through multimodal fusion networks and temporal machine learning models. Combined with an interpretable decision generation module, interpretable intelligent decision reports are generated.
It enables deep feature extraction and dynamic profile construction from multimodal data, improving the objectivity and transparency of the influencer selection process and providing highly credible marketing decision support.
Smart Images

Figure CN122114989A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent marketing technology, and in particular to a marketing expert precision matching and intelligent decision-making system based on multimodal data mining. Background Technology
[0002] With the rapid development of the digital marketing industry, brands are increasingly relying on multimodal data from social media platforms when selecting marketing influencers. This data includes text and image content, video materials, live stream clips, and user interactions. This data is not only massive in quantity and complex in type, but also highly unstructured, multi-sourced, heterogeneous, and rapidly changing over time. To improve marketing effectiveness, brands often want systems that can accurately identify influencer content styles, fan base characteristics, content creation pace, and potential dissemination risks, and match these with their own product attributes, target audience, and marketing scenarios.
[0003] Existing influencer matching methods generally suffer from several key technical bottlenecks. On the one hand, existing systems have limited capabilities in processing multimodal data, mostly analyzing only text tags, basic statistical features, or simple content tags. They cannot achieve cross-modal correlation modeling, let alone capture the dynamic trends of influencer content style, audience characteristics, and interactive behavior evolving over time. On the other hand, existing recommendation mechanisms typically only provide a single score or simple ranking result, lacking the ability to explain the "reasons for recommendation" and "potential risks," making it difficult for brands to judge the rationality and risk level of recommendations. Furthermore, existing case retrieval and comparison methods are mostly driven by human experience, lacking structured comparison mechanisms and the ability to form systematic, quantitative, and reusable intelligent decision-making basis. Summary of the Invention
[0004] This invention provides a marketing influencer precision matching and intelligent decision-making system based on multimodal data mining. It is a comprehensive system that can realize multimodal deep feature extraction, dynamic profile construction, multidimensional matching calculation, and interpretable recommendation, thereby improving the scientificity, transparency, and reliability of influencer selection and marketing decisions.
[0005] The marketing influencer precision matching and intelligent decision-making system based on multimodal data mining includes a heterogeneous feature extraction module, a dynamic profile construction module, a multidimensional matching degree calculation module, and an interpretable decision generation module, among which; The heterogeneous feature extraction module is used to extract features from multimodal raw data and generate a heterogeneous feature set; The dynamic profile building module is used to process the heterogeneous feature set through a multimodal fusion network and a temporal machine learning model to build a dynamic multimodal fusion profile containing dynamic labels. The multidimensional matching degree calculation module is used to perform multidimensional mapping and alignment between the dynamic multimodal fusion profile and the marketing demand description through machine learning algorithms, and calculate and generate a multidimensional matching degree vector. The interpretable decision generation module is used to perform weighted aggregation and sorting of the multidimensional matching degree vector, and to verify it in conjunction with the interpretable sub-module based on case reasoning and rule engine, and output an interpretable intelligent decision report.
[0006] Optionally, the heterogeneous feature extraction module includes: Collect multimodal raw data of target marketing influencers and their associated fan groups. The multimodal raw data includes text and image content data, video and live stream data, and cross-platform interactive behavior data. The multimodal raw data is cleaned and standardized preprocessed to obtain preprocessed multimodal raw data. The pre-trained machine learning model is used to extract modality-specific features from the preprocessed multimodal raw data, generating image-text semantic feature vectors, visual emotion feature vectors, and behavioral sequence feature vectors, which are then aggregated to form the heterogeneous feature set.
[0007] Optionally, the modality-specific feature extraction of the preprocessed multimodal raw data using a pre-trained machine learning model includes: The pre-trained image and text understanding model is used to extract semantic features from the image and text content data in the preprocessed multimodal raw data, generating image and text semantic feature vectors; A pre-trained video understanding model is used to extract visual and emotional features from the video and live stream data in the pre-processed multimodal raw data, generating a visual and emotional feature vector. A pre-trained behavior sequence model is used to extract sequence patterns from the cross-platform interactive behavior data in the preprocessed multimodal raw data to generate behavior sequence feature vectors. The image and text semantic feature vector, the visual emotion feature vector, and the behavior sequence feature vector are aggregated to form the heterogeneous feature set.
[0008] Optionally, the dynamic profile building module includes: The heterogeneous feature set is input into a multimodal fusion network based on an attention mechanism to perform cross-modal correlation analysis and feature alignment, and output a unified multimodal fusion feature representation. The multimodal fusion feature representation is input into the temporal machine learning model to analyze the evolution of features over time, and to construct a group feature profile that represents the comprehensive influence of the influencer and the stability of the fan base, as well as an individual in-depth profile that represents the influencer's creative cycle and public opinion risk index. The group feature profile and the individual depth profile are fused together, and a dynamic multimodal fusion profile containing multiple dynamic tags is generated and continuously updated based on a preset dynamic update triggering mechanism.
[0009] Optionally, the step of inputting the heterogeneous feature set into an attention-based multimodal fusion network for cross-modal correlation analysis and feature alignment, and outputting a unified multimodal fusion feature representation, includes: The image and text semantic feature vectors, visual emotion feature vectors, and behavioral sequence feature vectors from the heterogeneous feature set are respectively input into a cross-modal attention network based on the Transformer architecture. The correlation weights between different modal features are calculated to generate intermodal aligned feature representations. The intermodal aligned feature representations are weighted, fused, and dimensionality reduced to generate the unified multimodal fusion feature representation.
[0010] Optionally, the multidimensional matching degree calculation module includes: Receive the dynamic multimodal fusion profile from the dynamic profile building module, and receive the marketing requirement description input by the brand. The marketing demand description is parsed and structured using the natural language processing model in the machine learning algorithm to generate a structured demand representation; The multi-dimensional mapping alignment model in the machine learning algorithm is invoked to perform similarity calculation and weight allocation between the dynamic label vector in the dynamic multimodal fusion profile and each vector in the structured demand representation, and outputs the semantic relevance score of the content dimension, the emotional consistency score of the audience dimension, and the scene adaptability score of the scene dimension respectively. The semantic relevance score, the sentiment consistency score, and the scenario suitability score are combined and normalized according to preset dimensions to generate the multi-dimensional matching vector for this marketing need.
[0011] Optionally, the step of using the natural language processing model in the machine learning algorithm to parse and structure the marketing demand description to generate a structured demand representation includes: The pre-trained language model in the machine learning algorithm is used to perform word segmentation, entity recognition, and semantic encoding on the marketing demand description to generate a preliminary semantic vector. The initial semantic vectors are classified and mapped using preset attribute extraction templates, and then filled into a structured framework of four dimensions: product attributes, target audience, marketing scenarios, and core appeals. The filled structured framework is vectorized and encoded to generate the structured requirement representation, which consists of product attribute vectors, target audience vectors, marketing scenario vectors, and core appeal vectors.
[0012] Optionally, the interpretable decision generation module includes: The system receives the multidimensional matching degree vector from the multidimensional matching degree calculation module, and performs weighted aggregation sorting on each dimension component of the multidimensional matching degree vector based on preset machine learning model weights to generate an initial expert recommendation sequence. The initial influencer recommendation sequence and the corresponding multi-dimensional matching degree vector are input into the interpretable module based on case reasoning and rule engine to retrieve similar historical marketing cases, compare and analyze the performance differences of the current recommended influencers and historical case influencers in multiple dimensions, and generate the matching advantage prompts and the potential risk warnings. By combining the initial expert recommendation sequence, the multi-dimensional matching degree vector, the matching advantage prompts, and the potential risk warnings, the data is formatted, assembled, and visualized according to a preset report template, ultimately generating an interpretable intelligent decision report containing complete recommendation criteria and risk analysis.
[0013] Optionally, the weighted aggregation and sorting of the components of the multidimensional matching degree vector based on preset machine learning model weights to generate an initial expert recommendation sequence includes: Based on the preset dimensional importance weights, the components of each dimension in the multidimensional matching degree vector are weighted and summed to generate a comprehensive matching score for each target marketing influencer. All target marketing influencers are sorted in descending order based on the comprehensive matching score to generate the initial influencer recommendation sequence. Each record in the sequence contains an influencer identifier and its corresponding multidimensional matching degree vector and the comprehensive matching score.
[0014] Optionally, the step of inputting the initial influencer recommendation sequence and the corresponding multi-dimensional matching degree vector into an interpretable module based on case reasoning and a rule engine to retrieve historical similar marketing cases, comparing and analyzing the performance differences of the currently recommended influencers and historical case influencers in multiple dimensions, and generating the matching advantage prompt and the potential risk warning includes: The interpretable module uses a similarity algorithm to retrieve the set of successful cases and the set of failed cases with the highest similarity from the historical case library based on the features of the current multidimensional matching degree vector. The multidimensional matching vector of the currently recommended influencer is compared dimension by dimension with the historical matching vectors of the corresponding influencers in the set of successful cases and the set of failed cases, and the difference value is calculated. According to the preset difference threshold rules, if the difference value of a certain dimension exceeds the success case threshold, a matching advantage prompt is generated; if the difference value of a certain dimension exceeds the failure case threshold, a potential risk warning is generated.
[0015] The beneficial effects of this invention are: 1. This invention employs a heterogeneous feature extraction module to perform multi-level structured extraction of semantic feature vectors from images and text, visual emotion feature vectors, and behavioral sequence feature vectors. In the dynamic profile construction module, it utilizes an attention-based multimodal fusion network and a temporal machine learning model to achieve cross-modal alignment, enhanced semantic association, and capture of temporal evolution patterns, thereby generating a continuously updated dynamic multimodal fusion profile. This invention solves the technical challenges of traditional user profiling, such as reliance on a single data source, lack of temporal dynamism, and difficulty in depicting deep behavioral patterns. It significantly improves the stability, precision, and timeliness of the profile results, providing a highly reliable input foundation for subsequent decision-making.
[0016] 2. This invention, through a multi-dimensional matching degree calculation module, maps and aligns the dynamically multimodal fused profile processed by dynamic profiling with the structured marketing demand representation in multiple dimensions. It then constructs a three-dimensional multi-dimensional matching degree vector composed of semantic relevance score, sentiment consistency score, and scenario adaptability score, achieving precise quantitative matching between influencer characteristics and brand needs. This invention not only automatically identifies fine-grained correlations between influencer content themes, fan audience characteristics, and scenario-based delivery capabilities, but also eliminates inconsistencies in human judgment through a standardized calculation process, significantly improving the objectivity, robustness, and repeatability of the influencer selection process.
[0017] 3. This invention, through an interpretable decision generation module, compares and analyzes the initial influencer recommendation sequence obtained from weighted aggregation and ranking with historical successful and unsuccessful cases using a case-based reasoning approach. Combined with a rule engine, it generates matching advantage suggestions and potential risk warnings, ultimately forming an interpretable intelligent decision report that includes visual charts, dimensional score explanations, and historical case comparisons. This mechanism addresses the industry pain point of traditional recommendation models' "black box output that cannot be explained," enabling brands not only to obtain ranking results but also to understand the basis for recommendations, sources of risk, and influencing factors. This improves decision transparency, security, and practicality, providing higher-value decision support for actual marketing campaigns. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the system flow according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the dynamic image construction module in an embodiment of the present invention. Detailed Implementation
[0020] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should also be noted that, to make the embodiments more comprehensive, the following embodiments are the best and preferred embodiments, and those skilled in the art can use other alternative methods to implement some well-known technologies; moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.
[0021] like Figures 1-2 As shown, the marketing influencer precision matching and intelligent decision-making system based on multimodal data mining includes a heterogeneous feature extraction module, a dynamic profile construction module, a multidimensional matching degree calculation module, and an interpretable decision generation module, wherein; The heterogeneous feature extraction module is used to extract features from multimodal raw data and generate heterogeneous feature sets. The specific steps are as follows: Data collection is carried out by calling the open application programming interfaces provided by social media platforms, starting with the unique account identifier of the target marketing influencer.
[0022] The collected multimodal raw data were divided into three categories.
[0023] The first category is text and image content data, which includes the text content, title content, topic tag content, and accompanying image files of historical text and image posts published by target marketing influencers. A unique content identifier is recorded for each text and image content.
[0024] The second category is video and live stream data, which includes metadata of short videos, metadata of long videos, video frame sequences, audio streams, and live stream recording playback data and real-time bullet screen text published by target marketing influencers, and is associated with video identifiers, live stream session identifiers and corresponding timestamps.
[0025] The third category is cross-platform interactive behavior data, which includes the comment and reply texts of target marketing influencers under their published content, the like, forward, and favorite behavior records of fans, as well as the corresponding timestamp sequences. It also includes the mention and interaction relationship chains between target marketing influencers and fans on different platforms, and uses user identifiers, content identifiers, and timestamps to establish a complete sequence of behavioral events.
[0026] During the data collection process, all records are sorted according to timestamps, and each record is associated with and stored with the corresponding influencer and fan identifiers, forming a multimodal raw data warehouse organized by time, providing a data foundation for subsequent preprocessing and feature extraction.
[0027] Perform data cleaning and format standardization operations on the collected multimodal raw data.
[0028] During the data cleaning phase, duplicate records are first removed by comparing content identifiers and timestamps, empty records whose content has been removed are deleted, and noisy data containing garbled characters and illegal encoded characters are filtered out. For image and text content data, as well as comment and reply text, text normalization processing is performed to remove meaningless special characters, standardize punctuation, and standardize character encoding formats.
[0029] For video and live stream data, keyframe extraction technology is used to convert the original video stream into a sequence of image frames sampled at fixed time intervals. Continuous frames are extracted according to the preset sampling time step, and each video sample preserves a set of image frames with a stable frame rate.
[0030] For cross-platform interactive behavior data, user identifiers and timestamps are used as the primary keys to aggregate discrete behavior logs. Individual like behavior records, forwarding behavior records, collection behavior records, and comment reply behavior records are linked together in chronological order to form a continuous behavior sequence structure.
[0031] In the standardization preprocessing stage, all text data is uniformly converted to UTF-8 encoding, and the text is segmented and labeled using a unified word segmentation rule. All image frames are scaled to a fixed resolution of 224×224 pixels, maintaining a consistent aspect ratio, resulting in a standardized image frame sequence. All action timestamps are uniformly converted to Greenwich Mean Time (GMT) format and truncated or padded according to a unified time precision.
[0032] Through the above processing, preprocessed multimodal raw data with standardized structure and uniform quality are obtained, providing a standardized data source for subsequent model input.
[0033] Modality-specific feature extraction is performed on the pre-processed multimodal raw data using a pre-trained machine learning model to generate image-text semantic feature vectors, visual emotion feature vectors, and behavioral sequence feature vectors, respectively. Based on these, a heterogeneous feature set is formed, including the following steps: 1. The pre-trained language model BERT based on the Transformer architecture is used as the image and text understanding model. For each piece of image and text content data, the corresponding text content, title content and topic tag content are concatenated in a fixed order to construct a single input sequence, and sentence start mark and sentence end mark are added. The input sequence is then fed into the BERT model.
[0034] The BERT model is composed of multiple layers of self-attention networks and feedforward networks stacked together. During the encoding process, the self-attention network performs weighted summation of the word vectors at each position in the input sequence and the word vectors at other positions in the sequence to learn the complete contextual semantic relationship. The feedforward network performs non-linear transformation on the intermediate representation after self-attention processing, thereby obtaining a deep semantic representation.
[0035] The output vector corresponding to the start marker position of the sentence is extracted from the last hidden state of the BERT model. This output vector is used as a 768-dimensional dense vector. This dense vector comprehensively represents the overall semantic information of the corresponding image and text content data. This dense vector is defined as the image and text semantic feature vector and is associated with the corresponding image and text content record one by one for storage.
[0036] 2. A deep convolutional neural network ResNet-50, pre-trained on the large visual dataset ImageNet, is used as the backbone network of the video understanding model. The obtained standardized image frame sequence is input into the ResNet-50 network frame by frame.
[0037] For each frame, ResNet-50 extracts multi-scale visual features through multiple convolutional layers, batch normalization layers, and residual connection structures. Finally, a 2048-dimensional visual feature vector is generated through a global average pooling layer. For all frame features of the same video sample, average pooling is performed on the feature dimensions in chronological order. All frame feature vectors are aggregated into a 2048-dimensional aggregated visual feature vector by averaging element-wise, which represents the comprehensive visual features of the entire video in terms of spatial content.
[0038] For the audio stream corresponding to the video and the live stream bullet screen text, speech recognition processing is first performed on the audio stream to obtain the audio transcribed text. The audio transcribed text is then merged with the key bullet screen text to form a sentiment text sequence, which is then input into a sentiment classifier based on a Long Short-Term Memory (LSTM) network. This LSM network encodes the sentiment tendency in the text sequence through memory units and gating mechanisms, outputting a sentiment hidden vector at the end of the sequence. Subsequently, it is mapped through a fully connected output layer to a sentiment probability distribution vector containing positive sentiment probability values, negative sentiment probability values, and neutral sentiment probability values.
[0039] The 2048-dimensional aggregated visual feature vector and the emotion probability distribution vector are concatenated along the feature dimension to obtain a composite feature vector containing 2051 components. This composite feature vector is defined as the visual emotion feature vector and is associated with and stored in relation to the corresponding video and live stream data records.
[0040] 3. Aggregate the associated fans of each target marketing influencer by fan identifier, and sort the interaction records of each fan on each platform by timestamp to construct a fan-level behavior sequence. Each behavior record is mapped to a fixed-dimensional embedding vector, which is obtained by concatenating the behavior type embedding and the target content identifier embedding. The behavior type embedding is used to distinguish different behavior types such as likes, reposts, favorites, and comments, and the target content identifier embedding is used to represent the specific content object of the behavior.
[0041] The aforementioned behavioral sequences are processed using a behavioral sequence model pre-trained on a Long Short-Term Memory (LSTM) network architecture. The embedded vector sequence, arranged chronologically, is sequentially input into the network. The LSTM network controls the flow of information over time through input, forget, and output gates, utilizing hidden states to transmit behavioral pattern information across multiple positions in the sequence, capturing the interaction rhythm and behavioral preferences of fans within a specific timeframe.
[0042] For each behavior sequence, after processing the entire sequence, the hidden state output of the Long Short-Term Memory network at the last time step of the sequence is read. This hidden state output is regarded as a 256-dimensional feature vector, which is used to represent the behavior pattern and activity level of the fan in the current observation window. This feature vector is defined as the behavior sequence feature vector and associated with the corresponding fan behavior sequence for storage.
[0043] 4. Using the target marketing influencer as the primary key, aggregate the aforementioned image and text semantic feature vectors, visual emotional feature vectors, and behavioral sequence feature vectors to form a corresponding heterogeneous feature set.
[0044] For each target marketing influencer, the semantic feature vectors of all text and image content under that influencer's name are first aggregated according to content identifiers. Then, if necessary, influencer-level text and image semantic feature expressions are obtained by element-wise averaging or weighted averaging of multiple semantic feature vectors. Next, the visual sentiment feature vectors corresponding to all videos and live streams published by that influencer are aggregated to form influencer-level visual sentiment feature expressions.
[0045] For the behavioral sequence feature vector, a core fan set is selected from the associated fan group. The core fan behavioral sequence feature vector set that meets the criteria is filtered out by using the fan activity threshold and the interaction frequency threshold. The multiple behavioral sequence feature vectors in the set are averaged element by element to generate the average behavioral sequence feature vector that represents the interaction characteristics of the influencer's fans.
[0046] Finally, the influencer-level text-image semantic feature representation, influencer-level visual emotional feature representation, and fan average behavioral sequence feature vector are combined in key-value pair form, where the key is the feature type identifier and the value is the corresponding feature vector, forming a structured heterogeneous feature set. This heterogeneous feature set simultaneously retains text-image semantic feature vectors, visual emotional feature vectors, and behavioral sequence feature vectors within a single data structure, providing a complete feature input foundation for subsequent multimodal fusion analysis in the dynamic profile construction module.
[0047] The dynamic profile building module processes heterogeneous feature sets through a multimodal fusion network and a temporal machine learning model to construct a dynamic multimodal fusion profile containing dynamic labels. The specific steps are as follows: The heterogeneous feature sets are input into an attention-based multimodal fusion network for cross-modal correlation analysis and feature alignment, outputting a unified multimodal fusion feature representation. The steps are as follows: 1. Construct and run a cross-modal attention network based on the Transformer encoder architecture. The network consists of three parallel self-attention sub-layers, which respectively receive image-text semantic feature vectors, visual sentiment feature vectors, and behavioral sequence feature vectors from heterogeneous feature sets as input.
[0048] Each self-attention sublayer models the input modality features through a multi-head self-attention mechanism and a feedforward sublayer. The self-attention mechanism enhances the structural information and semantic dependencies within the modality through linear mapping of Query, Key, and Value and weighted summation.
[0049] Subsequently, a cross-modal attention mechanism is introduced: using the image-text semantic feature vector as the query vector, and the visual sentiment feature vector and behavioral sequence feature vector as the key vector and value vector, the attention weights of the image-text modality to the visual modality and behavioral modality are calculated to obtain an image-text feature representation with context awareness; when the visual sentiment feature vector is used as the query vector, the same weight calculation is performed with the image-text semantic feature vector and behavioral sequence feature vector to obtain a visual feature representation with cross-modal semantic association.
[0050] Through the aforementioned cross-attention calculation process, the model learns cross-modal association information such as "whether the text topic is consistent with the video image expression" and "the accompanying relationship between specific behavioral patterns and specific topic content", and outputs a set of inter-modal alignment feature representations aligned in a unified semantic space.
[0051] 2. The alignment feature representations between modalities are fused and dimensionality reduced. Specifically, the alignment representations from the image-text semantic feature vectors, visual sentiment feature vectors, and behavioral sequence feature vectors are concatenated to form a high-dimensional fused feature matrix.
[0052] The fused feature matrix is input into a weight generation network containing a fully connected layer and a Softmax activation function, and the weight coefficients corresponding to the three modalities are output.
[0053] The obtained weight coefficients are used to perform a weighted summation on the alignment feature representations of the three modalities to generate a single preliminary fusion vector.
[0054] Subsequently, the initial fusion vector is input into a bottleneck network consisting of one or more fully connected layers. Dimensionality reduction is performed through linear mapping and nonlinear activation, and a dense feature vector with a fixed dimension of 512 is output. This dense feature vector is defined as a unified multimodal fusion feature representation and is used as input to subsequent temporal machine learning models.
[0055] The multimodal fusion feature representation is input into a time-series machine learning model to analyze the evolution of features over time, and group feature profiles and individual deep profiles are constructed respectively. The steps are as follows: 1. The multimodal fusion features are segmented over time and input into a Long Short-Term Memory (LSTM) network. Specifically, the unified multimodal fusion feature representation is divided into natural week time windows, forming a sequence of feature segments arranged chronologically, with each time window corresponding to a fusion feature vector. This feature segment sequence is then input into the LTM network, which serves as the temporal machine learning model. The LTM network controls the update of the control state through input gates, forget gates, and output gates, thus preserving long-term dependent information while suppressing the accumulation of invalid historical information. After processing the entire time series, the network outputs a hidden state vector at the last time step. This hidden state vector contains temporal evolution information throughout the entire observation period and is defined as a dynamic evolution feature for subsequent profile construction.
[0056] 2. Constructing group feature profiles and individual in-depth profiles based on dynamic evolutionary features. A dual-branch classification and regression model is constructed, with the aforementioned dynamic evolutionary features as input.
[0057] The first branch is the group feature extraction branch. It calculates the moving average and variance of the influencer's content interaction rate over the past four weeks through a multi-layer fully connected network. This set of indicators is used as the influencer's comprehensive influence score. At the same time, based on the autocorrelation function value of the fan interaction behavior time series, a fan group stability score is generated. The comprehensive influence score and the group stability score are then assembled into a group feature profile vector.
[0058] The second branch is the individual feature extraction branch. It also uses dynamic evolution features as input and calculates a regularity index of content release intervals over time through another fully connected network. This regularity index forms the creation cycle score. Then, by comparing the similarity between the dynamic evolution features and expert-annotated negative public opinion event feature templates, a public opinion risk index is obtained. The creation cycle score and the public opinion risk index are then assembled into an individual deep profile vector.
[0059] The final output of the group feature profile and the individual in-depth profile are used to describe the overall behavioral characteristics of the fan group and the in-depth characteristics of the individual influencer in terms of content creation and public opinion risk, respectively.
[0060] The process involves fusing group feature profiles and individual deep profiles, and generating and continuously updating a dynamic multimodal fusion profile containing multiple dynamic tags based on a preset dynamic update trigger mechanism. The steps are as follows: 1. Constructing a Basic Fusion Profile: The aforementioned group feature profile vector and individual deep profile vector are directly concatenated along the feature dimension to form an extended feature vector, which is defined as the basic fusion profile. The basic fusion profile includes multiple indicators such as influencer comprehensive influence score, fan group stability score, creation cycle score, and public opinion risk index, providing a basic structure for dynamic tag generation.
[0061] 2. Configure and monitor dynamic update triggering mechanisms: Construct three types of triggering conditions: one is the content publishing triggering condition, which continuously detects whether new text and image content data and video content data are generated by connecting to the content publishing monitoring service of influencer accounts. Once a new content publishing event is captured, an update signal is generated; another is the magnitude jump triggering condition, which continuously monitors the time series of the total number of influencer followers. When the growth rate of the number of followers exceeds a preset threshold of 5% within 24 consecutive hours, an update signal is generated; and the third is the periodic triggering condition, which automatically generates an update signal at the end of each consecutive seven-day cycle through the system clock.
[0062] When any one of the above three triggering conditions is met, the dynamic update triggering mechanism sends a triggering instruction to the portrait update process.
[0063] 3. Incremental learning and updates are performed based on a dynamic update trigger mechanism to generate a dynamic multimodal fusion profile. Upon receiving a trigger instruction from the dynamic update trigger mechanism, the system loads the basic fusion profile calculated within the latest time period, compares it with existing profiles stored in the database, and calculates gradient information based on the differences between the old and new profiles. An online gradient descent algorithm is used to incrementally adjust the weight coefficients and numerical parameters corresponding to each dynamic label in the existing profile, without completely retraining the overall model. In cases where new data indicates a significant increase in the recent interaction rate of influencers and a simultaneous rise in the public opinion risk index, the online gradient descent algorithm will increase the values of "influence" related labels in the profile and simultaneously increase the weights of "risk" related labels to reflect the true state of the current stage. The latest profile obtained after the above incremental update process is defined as a dynamic multimodal fusion profile. This dynamic multimodal fusion profile consists of multiple dynamic labels with time-sensitive attributes. Each dynamic label contains values and weights that are updated over time, used to support the subsequent multidimensional matching degree calculation module and interpretable decision generation module for accurate matching and intelligent decision-making.
[0064] The multi-dimensional matching degree calculation module is used to map and align dynamic multimodal fusion profiles with marketing demand descriptions in multiple dimensions using machine learning algorithms, and to calculate and generate a multi-dimensional matching degree vector. The specific steps are as follows: It receives dynamic multimodal fusion profiles from the dynamic profile building module, as well as marketing requirement descriptions input by the brand, specifically: The system receives a dynamic multimodal fusion profile from the dynamic profile building module via an internal data bus. This dynamic multimodal fusion profile is cached in memory as a structured data object. The data object contains several dynamic label vectors with weights and values. Each dynamic label vector is used to describe the target marketing influencer's characteristics in terms of content theme performance, fan group characteristics, interaction behavior patterns, and public opinion risk status.
[0065] Simultaneously, the system receives marketing requirement descriptions in text format through a graphical user interface for brands. The marketing requirement description is a continuous natural language text containing the target product category, product positioning, desired audience characteristics, campaign scenario description, and the brand message to be conveyed. After input validation, the system caches the dynamic multimodal fusion profile object and the marketing requirement description text in session-level memory areas, providing data input for subsequent natural language parsing and matching degree calculation.
[0066] The marketing needs description is parsed and structured using natural language processing models in machine learning algorithms, generating a structured needs representation that includes product attribute vectors, target audience vectors, marketing scenario vectors, and core appeal vectors. This process includes the following sub-steps: 1. Utilize pre-trained language models from machine learning algorithms to perform word segmentation, entity recognition, and semantic encoding on marketing demand descriptions, generating preliminary semantic vectors, specifically: We selected BERT, a pre-trained language model based on the Transformer architecture, as the natural language processing model. The complete marketing needs description text was input into the BERT model. The model first segmented the text according to its built-in segmentation rules. Based on the segmentation results, it performed entity recognition, labeling entity fragments related to product name, functional characteristics, user group description, region, age group, gender characteristics, campaign time and scenario, and marketing needs.
[0067] Subsequently, the BERT model uses its multi-layer Transformer encoder to perform multi-head self-attention computation and feedforward network transformation on the segmented sequence to obtain the contextual semantic representation of each position. After encoding, the output vector corresponding to the [CLS] marker position is extracted from the last hidden state as a dense vector representing the overall semantics of the entire marketing demand description. This dense vector is set as the initial semantic vector with a vector dimension of 768.
[0068] 2. Using pre-defined attribute extraction templates, the initial semantic vectors are categorized and mapped, and then populated into a structured framework encompassing four dimensions: product attributes, target audience, marketing scenarios, and core appeals. Specifically: The system pre-builds an attribute extraction template, which consists of a four-dimensional structured framework. The four dimensions correspond to the product attribute dimension, target audience dimension, marketing scenario dimension, and core appeal dimension, respectively.
[0069] Each dimension contains several slots. The product attribute dimension includes product category slots, price range slots, and functional focus slots. The target audience dimension includes region slots, age group slots, and gender slots. The marketing scenario dimension includes activity type slots, time period slots, and communication channel slots. The core appeal dimension includes brand image slots, conversion target slots, and communication style slots.
[0070] The system inputs the initial semantic vectors and their corresponding entity annotations into a multi-classification model. This model, based on trained classification weights, determines the correspondence between semantic information and the various slots. Semantic fragments related to product characteristics are mapped to corresponding slots in the product attribute dimension; semantic fragments related to audience descriptions are mapped to corresponding slots in the target audience dimension; semantic fragments reflecting the timing and environment of the campaign are mapped to corresponding slots in the marketing scenario dimension; and semantic fragments related to the message delivery style and brand objectives are mapped to corresponding slots in the core appeal dimension. Through this classification and mapping process, the four dimensions in the attribute extraction template are filled into a framework containing structured text values.
[0071] 3. The filled-in structured framework is vectorized and encoded to generate a structured requirement representation. This representation consists of product attribute vectors, target audience vectors, marketing scenario vectors, and core appeal vectors, specifically: For each dimension that has been filled, the system constructs a corresponding embedding layer, mapping the text values in the slots to fixed-dimensional embedding vectors. For each dimension, the embedding vectors of all slots in that dimension are averaged using a pooling operation to obtain a single-dimensional representation. Specifically, pooling is performed on the product attribute dimension to obtain a 128-dimensional product attribute vector; pooling is performed on the target audience dimension to obtain a 128-dimensional target audience vector; pooling is performed on the marketing scenario dimension to obtain a 128-dimensional marketing scenario vector; and pooling is performed on the core appeal dimension to obtain a 128-dimensional core appeal vector.
[0072] The four vectors mentioned above together constitute a structured requirement representation, which is used for subsequent multi-dimensional mapping alignment calculations.
[0073] The multi-dimensional mapping alignment model configured in the machine learning algorithm is invoked to map the dynamic multimodal fusion profile and structured requirement representation to the same semantic space. Semantic relevance score, sentiment consistency score, and scene fit score are obtained through three types of similarity calculation sub-models. This step specifically includes the following sub-steps: 1. Input each dynamic tag vector in the dynamic multimodal fusion profile and the product attribute vector in the structured demand representation into the first similarity calculation sub-model to calculate and obtain the semantic relevance score, specifically: The first similarity calculation sub-model employs a dual-tower neural network structure. The left tower receives dynamic tag vectors related to the influencer's content theme, professional field, and historical product category performance as input, while the right tower receives product attribute vectors from the structured demand representation as input. Each tower consists of one or more fully connected networks that perform dimensionality transformation and non-linear activation on the input vectors, mapping them to a unified embedding space.
[0074] After completing the encoding transformation, the system calculates the cosine similarity between the two output vectors. The obtained similarity value is compressed to a value range of zero to one using the Sigmoid activation function. This value is used as the semantic relevance score of the content dimension to represent the degree of matching between the influencer's content focus and the brand's product attributes.
[0075] 2. Input the relevant dynamic tag vectors from the dynamic multimodal fusion profile and the target audience vectors from the structured demand representation into the second similarity calculation sub-model to calculate and obtain the sentiment consistency score, specifically: The second similarity calculation sub-model also employs a dual-tower neural network structure. The left tower receives dynamic label vectors selected from the dynamic multimodal fusion profile, which are related to fan sentiment, interaction style, and audience composition. The right tower receives the target audience vector from the structured demand representation. Both towers encode and transform their respective inputs through fully connected layers and nonlinear activation functions, ensuring that the dynamic label features and target audience features are mapped to the same vector space.
[0076] After encoding, the system calculates the dot product similarity of the two vectors, inputs the dot product result into the sigmoid activation function, and outputs a sentiment consistency score between zero and one. This score is used to measure the degree of alignment between the characteristics of the target marketing influencer's fan base and the characteristics of the brand's target audience in terms of value orientation, engagement methods, and content acceptance style.
[0077] 3. The relevant dynamic tag vectors from the dynamic multimodal fusion profile, along with the marketing scenario vectors and core appeal vectors from the structured demand representation, are input into the third similarity calculation sub-model to calculate and obtain the scenario fit score, specifically: The third similarity calculation sub-model employs a three-input neural network structure. The system selects dynamic tag vectors from the dynamic multimodal fusion profile that are related to the influencer's past performance in various advertising scenarios, content release time distribution, peak interaction periods, and content dissemination rhythm. These vectors are then combined with marketing scenario vectors and core appeal vectors. The three sets of vectors are concatenated along their feature dimensions to form a single high-dimensional input vector, which is then fed into a multilayer perceptron network. The multilayer perceptron sequentially performs deep fusion of the input features through multiple fully connected layers and nonlinear activation layers, capturing the mapping relationship between the influencer's historical scenario performance and the requirements of the current marketing scenario.
[0078] The network output layer generates a scalar value, which the system compresses to between zero and one using the Sigmoid function. This scalar value is defined as the scenario fit score, which represents the target marketing influencer's ability to adapt to the current marketing scenario and core appeal.
[0079] The semantic relevance score, sentiment consistency score, and scenario fit score are combined and normalized according to preset dimensions to generate a multi-dimensional matching vector for this marketing need. The specific steps include the following: 1. Multiply the semantic relevance score, sentiment consistency score, and scene fit score by preset dimension weight coefficients to obtain the weighted dimension scores, specifically: Based on the strategic focus of this marketing campaign, the system retrieves pre-defined dimension weight coefficients from the configuration database. The configuration database pre-defines weight combinations for different delivery strategies: the semantic relevance score reflects the importance of content theme matching, the sentiment consistency score reflects the importance of audience matching, and the scenario adaptability score reflects the importance of scenario matching. The system multiplies each raw score by its corresponding weight coefficient to obtain three weighted dimension scores.
[0080] 2. Concatenate the weighted dimension scores to form a preliminary combined vector, specifically: The system arranges the weighted semantic relevance score, weighted sentiment consistency score, and weighted scenario suitability score in a fixed dimensional order to form a preliminary three-dimensional combination vector. This preliminary combination vector numerically reflects the comprehensive performance of the three matching dimensions under the current marketing needs.
[0081] 3. The initial combined vector is normalized to ensure that the scores of each dimension fall within a uniform numerical range, ultimately generating a multi-dimensional matching degree vector, specifically: The system employs a min-max normalization method to linearly transform the three components of the initial combined vector. The system calculates the maximum and minimum scores for each dimension from historical matching records and transforms the current three-dimensional scores to a uniform numerical range of zero to one hundred using a preset formula. After normalization, a three-dimensional multi-dimensional matching degree vector is obtained. In this vector, each component represents the content dimension matching degree, audience dimension matching degree, and scene dimension matching degree, respectively, which can be used by the subsequent interpretable decision generation module for sorting, filtering, and decision report generation.
[0082] The interpretable decision generation module performs weighted aggregation and sorting of multi-dimensional matching degree vectors, and verifies them in conjunction with the interpretable sub-module based on case reasoning and rule engine, outputting an interpretable intelligent decision report. The specific steps are as follows: The system receives a multidimensional matching degree vector from the multidimensional matching degree calculation module, and performs weighted aggregation and sorting of each dimension component of the multidimensional matching degree vector based on preset machine learning model weights to generate an initial expert recommendation sequence, specifically: This step is the aggregation and sorting stage of the interpretable decision generation process: The system first receives the multidimensional matching degree vector calculated for each target marketing influencer from the multidimensional matching degree calculation module. The multidimensional matching degree vector is a three-dimensional vector containing semantic relevance dimension score, sentiment consistency dimension score, and scenario adaptability dimension score, all three scores have been normalized to a uniform numerical range.
[0083] Converting the three scores into a single sorting criterion involves the following sub-steps: 1. Based on the preset dimension importance weights, the components of each dimension are weighted and summed to generate a comprehensive matching score, specifically: The system loads a pre-defined dimension importance weight model. This model configures fixed weight coefficients for the semantic relevance dimension, sentiment consistency dimension, and scene adaptability dimension, with the semantic relevance weight configured as 0.5, the sentiment consistency weight configured as 0.3, and the scene adaptability weight configured as 0.2.
[0084] For each target marketing influencer, the system reads three dimension scores from their multi-dimensional matching vector and calculates a comprehensive matching score using a weighted summation method. The calculation process is as follows: the semantic relevance score is multiplied by 0.5, the sentiment consistency score by 0.3, and the scenario suitability score by 0.2. Then, the products of the three are added together to obtain a single scalar result, which serves as the influencer's comprehensive matching score. After calculation, this comprehensive matching score, along with the corresponding influencer's unique identifier and multi-dimensional matching vector, is stored in a list data structure in memory.
[0085] 2. Based on the overall matching score, all target marketing influencers are ranked to generate an initial influencer recommendation sequence, specifically as follows: The system calls a sorting program on the aforementioned list data structure. The sorting program uses the quicksort algorithm, with the comprehensive matching score as the key, to sort all target marketing influencer records in descending order. After sorting, the list is organized into an ordered sequence, which is defined as the initial influencer recommendation sequence.
[0086] Each record in the initial expert recommendation sequence contains three fields: an expert unique identifier, a multidimensional matching degree vector, and a comprehensive matching score. This initial expert recommendation sequence provides the basic input for subsequent interpretable analysis and report generation.
[0087] The initial influencer recommendation sequence and its corresponding multi-dimensional matching vector are input into an interpretable module based on case-based reasoning and a rule engine. This module retrieves similar historical marketing cases, compares and analyzes the performance differences of the currently recommended influencers with those of historical cases across multiple dimensions, and generates matching advantage suggestions and potential risk warnings. Specifically: This step is the interpretability analysis phase: Historical marketing case studies provide explanations of advantages and risk warnings for the current recommendation results. The system inputs the records prioritized in the initial influencer recommendation sequence and their corresponding multi-dimensional matching vectors into the interpretability module based on case-based reasoning and a rule engine, performing case retrieval, difference analysis, and rule determination. This step includes the following sub-steps: 1. Retrieve successful and failed case sets from the historical case database based on a similarity algorithm, specifically: The interpretable module selects several top-ranked influencers from the initial influencer recommendation sequence, reads the multi-dimensional matching vector of each influencer, and uses the K-nearest neighbor algorithm as the similarity algorithm to search the historical case database. The historical case database stores the historical matching vector of a corresponding influencer and the actual marketing effect label for each historical marketing case. The marketing effect label distinguishes between successful and unsuccessful cases.
[0088] For the current expert's multidimensional matching degree vector, the module calculates the Euclidean distance between it and the matching degree vectors of all cases in the historical case library. Then, it selects several successful cases with the smallest Euclidean distance to form a successful case set, selects several failed cases with the smallest Euclidean distance to form a failed case set, and stores the historical matching degree vector and case identifier of each case in the set into a temporary data structure.
[0089] 2. Compare the current recommended influencer's multi-dimensional matching vector with the historical matching vectors in the success and failure case sets, dimension by dimension, and calculate the difference value. Specifically: The multi-dimensional matching vector of the currently recommended influencer is compared dimension-by-dimensionally with the historical matching vectors of each case in the successful case set. First, for the successful case set, the average historical matching score of all successful cases is calculated for each dimension, forming the successful case average vector. Then, the corresponding successful case average vector is subtracted from the current influencer's multi-dimensional matching vector for each dimension to obtain the successful case difference vector.
[0090] Similarly, for the set of failed cases, the average historical matching score of all failed cases is calculated for each dimension to form a failed case average vector. Then, the current influencer's multidimensional matching vector is subtracted from the corresponding failed case average vector for each dimension to obtain a failed case difference vector. Each component in the above two difference vectors represents the degree of deviation of the current influencer from the average level of successful cases and the average level of failed cases in the corresponding dimension.
[0091] 3. Generate matching advantage suggestions and potential risk warnings based on preset difference threshold rules, specifically: Load the preset difference threshold rules. Each difference threshold rule configures two threshold parameters for each dimension: the difference threshold for successful cases and the difference threshold for failed cases. The difference threshold for successful cases is set to +0.15, and the difference threshold for failed cases is set to -0.1.
[0092] The algorithm iterates through each dimension component of the success case difference vector. When a dimension component is greater than the success case difference threshold, it indicates that the score for that dimension is significantly higher than the average level of historical success cases. A matching advantage alert is generated, recording the influencer's identifier, dimension name, and the specific difference value for that dimension. Then, it iterates through each dimension component of the failure case difference vector. When a dimension component is less than the failure case difference threshold, it indicates that the score for that dimension is significantly lower than the average level of historical failure cases. A potential risk warning is generated, similarly recording the influencer's identifier, dimension name, and the specific difference value for that dimension.
[0093] Ultimately, all matching advantage tips and potential risk warnings are stored in a structured record format for later report generation.
[0094] The system integrates the initial expert recommendation sequence, multi-dimensional matching degree vector, matching advantage hints, and potential risk warnings. Following a pre-set report template, it is formatted, assembled, and visualized to generate an interpretable intelligent decision-making report containing complete recommendation justification and risk analysis. Specifically: This step is the results output stage: It presents the complete recommendation basis and risk analysis to the brand through visualization and text description. The system integrates the initial influencer recommendation sequence, multi-dimensional matching degree vector, matching advantage hints, and potential risk warnings to generate a structured and interpretable intelligent decision-making report. This step includes the following sub-steps: 1. Convert the sorting results and score information into a visual chart in the form of a leaderboard, specifically: The system calls a chart generation library to read several top-ranked influencers from the initial influencer recommendation sequence. Using each influencer's unique identifier as the x-axis and their corresponding overall matching score as the y-axis, a ranked bar chart is created. Simultaneously, for each influencer, the standard deviation of the three dimensions' scores is calculated from the multi-dimensional matching degree vector. This standard deviation serves as a quantitative indicator of matching degree balance and is superimposed on the corresponding bar chart as error bars, allowing the chart to simultaneously present both the overall matching level and the balance between dimensions. After generation, the chart is rendered as an image file and cached.
[0095] 2. Categorize and organize the matching advantage suggestions and potential risk warnings to generate text analysis paragraphs, specifically: The system creates a text buffer and categorizes matching advantage tips and potential risk warnings based on influencer identifiers. For the top-ranked influencers, the system first reads their multi-dimensional matching vector, combining semantic relevance score, sentiment consistency score, and scene suitability score into an explanatory text describing the influencer's quantitative performance across these three dimensions. Subsequently, the system lists all matching advantage tips associated with that influencer in dimensional order, and then lists all potential risk warnings associated with that influencer in dimensional order.
[0096] For each prompt and warning, the system retrieves the case identifiers and average scores on the corresponding dimensions from the historical case database to supplement the description of "the specific difference ratio compared to relevant historical cases". This content is then integrated into a continuous natural language paragraph to form a text analysis paragraph for each expert.
[0097] 3. Assemble the visual charts, text analysis paragraphs, and report header and footer information according to the report template to output an interpretable intelligent decision-making report, specifically: The system loads a pre-set report template file. The report template uses an HTML structure and includes a cover information area, a marketing needs summary area, a core recommendation list area, an influencer in-depth analysis area, and a closing information area. The system inserts generated visualization charts into the core recommendation list area, inserts generated text analysis paragraphs in the influencer in-depth analysis area according to the influencer ranking, fills in a summary of the marketing needs description at the beginning of the report, and adds a generation timestamp and system version number at the end of the report.
[0098] After the above content is filled in, the system calls the rendering engine to convert the HTML report into a PDF document, which is then defined as the final interpretable intelligent decision-making report. This report includes ranking results, multi-dimensional matching metrics, matching advantage tips, and potential risk warnings, which brands can directly use for influencer selection and campaign decisions.
[0099] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0100] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A marketing influencer precision matching and intelligent decision-making system based on multimodal data mining, characterized in that: It includes a heterogeneous feature extraction module, a dynamic profile construction module, a multi-dimensional matching degree calculation module, and an interpretable decision generation module, among which; The heterogeneous feature extraction module is used to extract features from multimodal raw data and generate a heterogeneous feature set; The dynamic profile building module is used to process the heterogeneous feature set through a multimodal fusion network and a temporal machine learning model to build a dynamic multimodal fusion profile containing dynamic labels. The multidimensional matching degree calculation module is used to perform multidimensional mapping and alignment between the dynamic multimodal fusion profile and the marketing demand description through machine learning algorithms, and calculate and generate a multidimensional matching degree vector. The interpretable decision generation module is used to perform weighted aggregation and sorting of the multidimensional matching degree vector, and to verify it in conjunction with the interpretable sub-module based on case reasoning and rule engine, and output an interpretable intelligent decision report.
2. The marketing influencer precision matching and intelligent decision-making system based on multimodal data mining according to claim 1, characterized in that, The heterogeneous feature extraction module includes: Collect multimodal raw data of target marketing influencers and their associated fan groups. The multimodal raw data includes text and image content data, video and live stream data, and cross-platform interactive behavior data. The multimodal raw data is cleaned and standardized preprocessed to obtain preprocessed multimodal raw data. The pre-trained machine learning model is used to extract modality-specific features from the preprocessed multimodal raw data, generating image-text semantic feature vectors, visual emotion feature vectors, and behavioral sequence feature vectors, which are then aggregated to form the heterogeneous feature set.
3. The marketing influencer precision matching and intelligent decision-making system based on multimodal data mining according to claim 2, characterized in that, The modality-specific feature extraction of the preprocessed multimodal raw data using a pre-trained machine learning model includes: The pre-trained image and text understanding model is used to extract semantic features from the image and text content data in the preprocessed multimodal raw data, generating image and text semantic feature vectors; A pre-trained video understanding model is used to extract visual and emotional features from the video and live stream data in the pre-processed multimodal raw data, generating a visual and emotional feature vector. A pre-trained behavior sequence model is used to extract sequence patterns from the cross-platform interactive behavior data in the preprocessed multimodal raw data to generate behavior sequence feature vectors. The image and text semantic feature vector, the visual emotion feature vector, and the behavior sequence feature vector are aggregated to form the heterogeneous feature set.
4. The marketing influencer precision matching and intelligent decision-making system based on multimodal data mining according to claim 2, characterized in that, The dynamic profile construction module includes: The heterogeneous feature set is input into a multimodal fusion network based on an attention mechanism to perform cross-modal correlation analysis and feature alignment, and output a unified multimodal fusion feature representation. The multimodal fusion feature representation is input into the temporal machine learning model to analyze the evolution of features over time, and to construct a group feature profile that represents the comprehensive influence of the influencer and the stability of the fan base, as well as an individual in-depth profile that represents the influencer's creative cycle and public opinion risk index. The group feature profile and the individual depth profile are fused together, and a dynamic multimodal fusion profile containing multiple dynamic tags is generated and continuously updated based on a preset dynamic update triggering mechanism.
5. The marketing influencer precision matching and intelligent decision-making system based on multimodal data mining according to claim 4, characterized in that, The process of inputting the heterogeneous feature set into an attention-based multimodal fusion network for cross-modal correlation analysis and feature alignment, and outputting a unified multimodal fusion feature representation, includes: The image and text semantic feature vectors, visual emotion feature vectors, and behavioral sequence feature vectors from the heterogeneous feature set are respectively input into a cross-modal attention network based on the Transformer architecture. The correlation weights between different modal features are calculated to generate intermodal aligned feature representations. The intermodal aligned feature representations are weighted, fused, and dimensionality reduced to generate the unified multimodal fusion feature representation.
6. The marketing influencer precision matching and intelligent decision-making system based on multimodal data mining according to claim 4, characterized in that, The multidimensional matching degree calculation module includes: Receive the dynamic multimodal fusion profile from the dynamic profile building module, and receive the marketing requirement description input by the brand. The marketing demand description is parsed and structured using the natural language processing model in the machine learning algorithm to generate a structured demand representation; The multi-dimensional mapping alignment model in the machine learning algorithm is invoked to perform similarity calculation and weight allocation between the dynamic label vector in the dynamic multimodal fusion profile and each vector in the structured demand representation, and outputs the semantic relevance score of the content dimension, the emotional consistency score of the audience dimension, and the scene adaptability score of the scene dimension respectively. The semantic relevance score, the sentiment consistency score, and the scenario suitability score are combined and normalized according to preset dimensions to generate the multi-dimensional matching vector for this marketing need.
7. The marketing influencer precision matching and intelligent decision-making system based on multimodal data mining according to claim 6, characterized in that, The step of using the natural language processing model in the machine learning algorithm to parse and structure the marketing demand description to generate a structured demand representation includes: The pre-trained language model in the machine learning algorithm is used to perform word segmentation, entity recognition, and semantic encoding on the marketing demand description to generate a preliminary semantic vector. The initial semantic vectors are classified and mapped using preset attribute extraction templates, and then filled into a structured framework of four dimensions: product attributes, target audience, marketing scenarios, and core appeals. The filled structured framework is vectorized and encoded to generate the structured requirement representation, which consists of product attribute vectors, target audience vectors, marketing scenario vectors, and core appeal vectors.
8. The marketing influencer precision matching and intelligent decision-making system based on multimodal data mining according to claim 6, characterized in that, The interpretable decision generation module includes: The system receives the multidimensional matching degree vector from the multidimensional matching degree calculation module, and performs weighted aggregation sorting on each dimension component of the multidimensional matching degree vector based on preset machine learning model weights to generate an initial expert recommendation sequence. The initial influencer recommendation sequence and the corresponding multi-dimensional matching degree vector are input into the interpretable module based on case reasoning and rule engine to retrieve similar historical marketing cases, compare and analyze the performance differences of the current recommended influencers and historical case influencers in multiple dimensions, and generate the matching advantage prompts and the potential risk warnings. By combining the initial expert recommendation sequence, the multi-dimensional matching degree vector, the matching advantage prompts, and the potential risk warnings, the data is formatted, assembled, and visualized according to a preset report template, ultimately generating an interpretable intelligent decision report containing complete recommendation criteria and risk analysis.
9. The marketing influencer precision matching and intelligent decision-making system based on multimodal data mining according to claim 8, characterized in that, The weighted aggregation and sorting of each dimension component of the multidimensional matching degree vector based on the preset machine learning model weights to generate an initial expert recommendation sequence includes: Based on the preset dimensional importance weights, the components of each dimension in the multidimensional matching degree vector are weighted and summed to generate a comprehensive matching score for each target marketing influencer. All target marketing influencers are sorted in descending order based on the comprehensive matching score to generate the initial influencer recommendation sequence. Each record in the sequence contains an influencer identifier and its corresponding multidimensional matching degree vector and the comprehensive matching score.
10. The marketing influencer precision matching and intelligent decision-making system based on multimodal data mining according to claim 8, characterized in that, The process of inputting the initial influencer recommendation sequence and the corresponding multi-dimensional matching vector into an interpretable module based on case reasoning and a rule engine, retrieving historical similar marketing cases, comparing and analyzing the performance differences of the currently recommended influencers and historical case influencers in multiple dimensions, and generating the matching advantage prompts and potential risk warnings includes: The interpretable module uses a similarity algorithm to retrieve the set of successful cases and the set of failed cases with the highest similarity from the historical case library based on the features of the current multidimensional matching degree vector. The multidimensional matching vector of the currently recommended influencer is compared dimension by dimension with the historical matching vectors of the corresponding influencers in the set of successful cases and the set of failed cases, and the difference value is calculated. According to the preset difference threshold rules, if the difference value of a certain dimension exceeds the success case threshold, a matching advantage prompt is generated; if the difference value of a certain dimension exceeds the failure case threshold, a potential risk warning is generated.