Multi-mode intention perception method and system based on CTR improvement
Through PCA dimensionality reduction and feature extraction network, combined with Bayesian network and CTR data feedback, the accuracy and personalized adjustment of multimodal user intention perception are solved, achieving more accurate user intention recognition and CTR improvement.
Patent Information
- Application Number
- CN202510479173.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to accurately and efficiently perceive and understand multimodal user intentions, especially when processing complex and variable multimodal data, it is susceptible to redundant information interference, and it is difficult to adapt to dynamic changes and personalized adjustments of user intentions.
PCA dimensionality reduction and feature extraction network are used to build an identification model for each mode, and the model parameters are optimized through Adam, combined with Bayesian network to perform multimodal information fusion, and typical intention modes of different CTR user groups are calculated as the benchmark point, and user intentions are adjusted according to the cross value and CTR data feedback.
It improves the accuracy and personalized perception of user intention recognition, improves CTR, increases the possibility of user clicks, and optimizes user satisfaction and loyalty.
Smart Images

Figure CN120336760A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent multimodal interaction and intention recognition, and particularly to a multimodal intention perception method and system based on CTR improvement. Background Art
[0002] In the current digital age, the interaction between users and various intelligent systems is becoming increasingly frequent and diverse, involving multiple modalities such as text, image, voice, video, etc. These multimodal data contain rich user intention information, which is of great significance for improving user experience, optimizing product recommendations, enhancing marketing effects, etc. However, how to accurately and efficiently perceive and understand the multimodal intentions of users remains a major challenge currently.
[0003] Traditional intention perception methods often focus on the processing of single-modal data, such as analyzing user intentions only based on text or recognizing user behaviors only through images. These methods are unable to cope when dealing with complex and variable multimodal data and are difficult to comprehensively and accurately capture the true intentions of users. In addition, when dealing with high-dimensional feature data, traditional methods are easily interfered by redundant information, resulting in low recognition efficiency and even misjudgment.
[0004] With the development of deep learning technology, although some progress has been made in multimodal intention perception, there are still many deficiencies. For example, the information fusion between different modalities is often not sufficient, and the conditional dependence relationship between modalities is not fully considered; there is a lack of an effective adaptation mechanism for the dynamic changes of user intentions; in terms of personalized intention perception, it is difficult to make precise adjustments according to the behavior characteristics of user groups. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a multimodal intention perception method and system based on CTR improvement, which can improve the accuracy of user intention recognition, increase CTR and the possibility of user clicks.
[0006] To solve the above technical problems, the technical solution of the present invention is as follows:
[0007] In the first aspect, a multimodal intention perception method based on CTR improvement, the method includes:
[0008] Obtain the multimodal data of the user;
[0009] Reduce the dimensionality of the high-dimensional features of the multimodal data through PCA, construct a feature extraction network, and extract key feature data;
[0010] According to each modality, construct a corresponding recognition model, and optimize the model parameters through Adam to obtain a single-modal recognition model;
[0011] Based on the unimodal recognition model, identify the key feature data and obtain the recognition results for each modality;
[0012] Combine the recognition results with the Bayesian network for multimodal information fusion to obtain the user's preliminary intention;
[0013] According to the historical data, calculate the typical intention patterns of different CTR user groups as the first reference point and the second reference point;
[0014] Calculate the cross values between the user's preliminary intention and the first reference point and the second reference point;
[0015] Adjust the user's intention according to the cross values and the CTR data feedback to achieve multimodal intention perception.
[0016] Furthermore, perform dimensionality reduction on the high-dimensional features of the multimodal data through PCA, construct a feature extraction network, and extract the key feature data, including:
[0017] Preprocess the collected multimodal data to obtain the preprocessed multimodal data;
[0018] Convert the preprocessed multimodal data into a matrix, where each row represents a sample and each column represents a feature, and calculate the covariance matrix of the data matrix;
[0019] Perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues and the corresponding eigenvectors,
[0020] According to the magnitudes of the eigenvalues, select the eigenvectors corresponding to the top k largest eigenvalues as the principal components;
[0021] Construct a transformation matrix from the principal components and multiply it with the original high-dimensional data matrix to obtain the projected low-dimensional data matrix;
[0022] Construct a feature extraction network based on the low-dimensional data matrix.
[0023] Furthermore, construct a corresponding recognition model for each modality and optimize the model parameters through Adam to obtain the unimodal recognition model, including:
[0024] Construct LSTM, CNN, and RNN recognition models for each modality. Among them, construct an LSTM recognition model for the text modality, a CNN recognition model for the image modality, and an RNN recognition model for the speech modality;
[0025] Initialize the model parameters and set the objective function for optimizing the model parameters during training;
[0026] Configure the Adam optimizer to train the model, optimize the model parameters, and reduce the loss function value to obtain the optimized unimodal recognition model.
[0027] Furthermore, combine the recognition results with the Bayesian network for multi-modal information fusion to obtain the user's preliminary intention, including:
[0028] Construct the structure of the Bayesian network, including nodes and edges. Among them, each node represents the recognition result of a modality or the user's intention, and the directed edges between nodes represent the conditional dependence relationships between modalities or between modalities and intentions;
[0029] Set the prior probability for each node in the Bayesian network and the conditional probability for each edge;
[0030] According to the prior probability and conditional probability, combined with the single-modal recognition results, calculate the posterior probability of the user's intention to obtain information of different modalities;
[0031] Integrate the information of different modalities to form a comprehensive preliminary intention of the user.
[0032] Furthermore, according to historical data, calculate the typical intention patterns of different CTR user groups as the first reference point and the second reference point, including:
[0033] According to historical data, calculate the CTR value of each user and divide the users into different groups;
[0034] According to the CTR value of each user, extract the features of the user's intention, and based on the extracted features, through Calculate the typical intention pattern of each user group, where F′ j is the comprehensive intention of the weighted user group G j , w f is the weight of the interaction frequency F f , w s is the weight of the modality switching frequency F s , w c is the weight of the occurrence times O c of a specific operation, w d is the weight of the interaction duration D i , N j is the number of users in the user group G j , F f,i is the interaction frequency of user i, F s,i is the modality switching frequency of user i, O c,i is the occurrence times of a specific operation of user j, D i,i is the interaction duration of user i, and i is the index;
[0035] According to the typical intention patterns, set the high CTR group as the first reference point and the low CTR group as the second reference point.
[0036] Further, calculating the intersection values of the user's preliminary intention with the first reference point and the second reference point includes:
[0037] By calculating the intersection values of the user's preliminary intention with the first reference point and the second reference point, where Cr is the intersection value of the user's preliminary intention with the first reference point and the second reference point, w A and w B are the weights of the first reference point and the second reference point, U is the feature vector of the user's preliminary intention, B1 and B2 are the feature vectors of the first reference point and the second reference point, and are the cosine similarities between the user's preliminary intention and the first reference point and the second reference point, ɑ A and α B are the weights of the feature interaction, and are the sums of the element-wise products of the features between the user's preliminary intention and the first reference point and the second reference point respectively, β is the weight of the reference point interaction, and k is the index.
[0038] Further, according to the intersection values and the CTR data feedback, adjusting the user's intention to achieve multi-modal intention perception includes:
[0039] Obtaining the CTR data of the user in the actual interaction process, and performing correlation analysis on the intersection values and the obtained CTR data to obtain the analysis result;
[0040] According to the analysis results of the intersection values and the obtained CTR data, formulating a strategy for adjusting the user's intention, and implementing and monitoring the effect according to the adjustment strategy;
[0041] According to the monitoring results, optimizing the user intention recognition model and the interaction process to achieve multi-modal intention perception.
[0042] In a second aspect, a multi-modal intention perception system based on CTR improvement includes:
[0043] An acquisition module for acquiring the multi-modal data of the user;
[0044] An identification module for reducing the dimension of the high-dimensional features of the multi-modal data through PCA, constructing a feature extraction network, and extracting key feature data; constructing a corresponding identification model according to each modality, and optimizing the model parameters through Adam to obtain a single-modal identification model; identifying the key feature data according to the single-modal identification model to obtain the identification results of each modality;
[0045] A processing module, configured to combine the recognition result with a Bayesian network to perform multimodal information fusion, so as to obtain a preliminary user intention; calculate typical intention patterns of different CTR user groups based on historical data as a first reference point and a second reference point; calculate the cross values between the preliminary user intention and the first reference point and the second reference point; and adjust the user intention according to the cross values and the CTR data feedback to achieve multimodal intention perception.
[0046] In a third aspect, a computing device includes:
[0047] One or more processors;
[0048] A storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method described above.
[0049] In a fourth aspect, a computer-readable storage medium stores a program that, when executed by a processor, implements the method described above.
[0050] The above solutions of the present invention at least include the following beneficial effects:
[0051] Through PCA dimensionality reduction and a feature extraction network, key features in multimodal data are effectively extracted, redundant information is reduced, and the recognition efficiency is improved. A dedicated recognition model is constructed for each modality, and the Adam optimizer is used to optimize the model parameters to ensure that the model can accurately capture the intention information of each modality. Multimodal information fusion is performed in combination with a Bayesian network, comprehensively considering the conditional dependence relationships between different modalities, to obtain a more comprehensive preliminary user intention. Calculate the typical intention patterns of different CTR user groups based on historical data as reference points, which helps to understand the behavioral characteristics of different user groups and achieve personalized intention perception. By calculating the cross values between the preliminary user intention and the reference points, the similarity and difference between the user intention and the reference points are quantified. According to the cross values and the CTR data feedback, the user intention is dynamically adjusted to make the intention perception more adaptable to the changes of different users and scenarios. By optimizing the user intention recognition model and the interaction process, the recognition accuracy of the user intention is improved, so as to more accurately meet the user's needs, improve the user satisfaction and loyalty, and accurate intention perception helps to improve the CTR and increase the possibility of user clicks. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a schematic flowchart of a multimodal intention perception method based on CTR improvement provided by an embodiment of the present invention.
[0053] Figure 2 is a schematic diagram of a multimodal intention perception system based on CTR improvement provided by an embodiment of the present invention. Detailed Implementation Manner
[0054] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0055] As Figure 1 shown, an embodiment of the present invention proposes a multi-modal intention perception method based on CTR improvement, and the method includes the following steps:
[0056] Step 11, obtaining multi-modal data of a user;
[0057] Step 12, reducing the dimension of high-dimensional features of the multi-modal data through PCA, constructing a feature extraction network, and extracting key feature data;
[0058] Step 13, constructing a corresponding recognition model according to each modality, and optimizing the model parameters through Adam to obtain a single-modal recognition model;
[0059] Step 14, recognizing the key feature data according to the single-modal recognition model to obtain the recognition results of each modality;
[0060] Step 15, combining the recognition results with a Bayesian network for multi-modal information fusion to obtain a preliminary user intention;
[0061] Step 16, calculating typical intention patterns of different CTR user groups according to historical data as the first reference point and the second reference point;
[0062] Step 17, calculating the cross values of the preliminary user intention with the first reference point and the second reference point;
[0063] Step 18, adjusting the user intention according to the cross values and CTR data feedback to achieve multi-modal intention perception.
[0064] In the embodiment of the present invention, the acquisition and fusion of multi-modal data provide more comprehensive user intention information. PCA dimensionality reduction and the feature extraction network improve the efficiency and accuracy of data processing. The combination of the single-modal recognition model and the Bayesian network takes into account modal specificity and information fusion. The calculation of the reference points and cross values provides a quantitative basis for intention adjustment. Through feedback adjustment, continuous optimization of intention perception is achieved. More accurate intention perception can provide more personalized services and recommendations, and by optimizing the user intention, key business indicators such as click-through rate and conversion rate are improved.
[0065] In a preferred embodiment of the present invention, step 11 may include:
[0066] Obtain the user's multimodal data, including text data, image data, and audio data.
[0067] In a preferred embodiment of the present invention, step 12 may include:
[0068] Step 121: Preprocess the collected multimodal data to obtain preprocessed multimodal data;
[0069] Step 122: Convert the preprocessed multimodal data into a matrix, where each row represents a sample and each column represents a feature, and calculate the covariance matrix of the data matrix;
[0070] Step 123: Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors,
[0071] Step 124: According to the magnitudes of the eigenvalues, select the eigenvectors corresponding to the top k largest eigenvalues as the principal components;
[0072] Step 125: Construct a transformation matrix from the principal components and multiply it with the original high-dimensional data matrix to obtain the projected low-dimensional data matrix;
[0073] Step 126: Construct a feature extraction network based on the low-dimensional data matrix.
[0074] In the embodiments of the present invention, noise, missing values, and outliers are removed to improve data quality, and feature extraction is performed: meaningful features are extracted from the original data. The matrix form facilitates mathematical operations and algorithm processing. The covariance matrix reflects the linear relationship between features and helps to understand the data structure. The variance matrix is the basis for dimensionality reduction algorithms such as PCA. Eigenvalues represent the variance of the data in the direction of the corresponding eigenvectors and reflect the main directions of data variation. Eigenvectors point to the directions where the data changes the most and help to capture the main features of the data. Eigenvalues and eigenvectors provide a basis for selecting the principal components. Eigenvalue decomposition is an orthogonal transformation that preserves the internal structure of the data. Selecting the top k principal components projects the high-dimensional data into a low-dimensional space, reducing the computational complexity. The eigenvectors corresponding to the largest eigenvalues capture the main variations of the data and retain the key information. The low-dimensional data matrix occupies less storage space, and the processing speed of the low-dimensional data is faster, improving the efficiency of the algorithm. The feature extraction network can automatically learn the effective features in the low-dimensional data, and the network can be adjusted according to different tasks and data characteristics to improve adaptability.
[0075] The specific steps in the embodiments of the present invention are as follows:
[0076] Step 121: Preprocess the collected multimodal data to check for missing values. For numerical data, the mean, median, or mode can be used for filling; for categorical data, the mode can be used for filling or special markers can be introduced. If the proportion of missing values is too high, consider deleting the sample or feature. Use the interquartile range method to identify outliers and replace them with reasonable values according to the situation. Identify and delete completely duplicate samples to avoid data redundancy. Convert numerical data to zero mean and unit variance to eliminate the influence of dimensions and scale the data to the [0, 1] interval; convert categorical data to numerical form and convert text data to numerical representation, such as bag-of-words model, TF-IDF, or word embeddings; extract the mean or use the sliding window method to capture time dependence. Ensure that data from different modalities are aligned in time or space, unify the time format, handle the problem of inconsistent time intervals, and perform synonym replacement, random insertion, random deletion, etc. to enhance the diversity of the text.
[0077] Step 122: Convert the preprocessed multimodal data into a matrix. Assume that after preprocessing, there are N samples and each sample has M features. Construct an N×M data matrix X, where the rows represent samples and the columns represent features. Calculate the mean of each feature and subtract the mean from the data to obtain the centered data matrix. The elements of the covariance matrix C represent the covariance between feature i1 and feature j1, and calculate the covariance matrix through , where N1 is the number of samples, is the centered value of the i1-th feature of the k1-th sample, is the centered value of the j1-th feature of the k1-th sample, and obtain an M×M covariance matrix, which reflects the linear relationship between features.
[0078] Step 123: Perform eigenvalue decomposition on the covariance matrix to obtain: C = VΛV T , where VΛV T is an orthogonal matrix, and each column is an eigenvector. Solve the eigenvalues and eigenvectors through the power iteration method. The eigenvalues represent the data variance magnitude in the direction of the eigenvectors, and the eigenvectors represent the main change directions of the data, obtaining the eigenvalue vector and the eigenvector matrix.
[0079] Step 124: Sort the eigenvalues in descending order, determine the eigenvectors corresponding to each eigenvalue, and select the eigenvectors corresponding to the top k largest eigenvalues to form an M×k principal component matrix. The principal components represent the k directions with the largest variance in the data, obtaining the principal component matrix.
[0080] Step 125: Use the selected principal component matrix as the transformation matrix to project the centered data matrix into a low-dimensional space to obtain an N×k low-dimensional data matrix.
[0081] Step 126: Select a fully connected neural network and set the network structure. The input layer has the number of nodes equal to the number of features of the low-dimensional data. The hidden layer designs the number of layers and nodes according to the task complexity, introduces non-linearity using an activation function. The output layer designs the number of output nodes according to the task type, uses an appropriate activation function, uses the mean squared error as the loss function, and uses Adam to update the network parameters to minimize the loss function. Calculate the output through forward propagation, calculate the gradient through backpropagation, update the parameters, and iteratively optimize the network performance to obtain a trained feature extraction network.
[0082] In a preferred embodiment of the present invention, the above step 13 may include:
[0083] Step 131: Construct LSTM, CNN, and RNN recognition models according to each modality. Specifically, construct an LSTM recognition model for the text modality, a CNN recognition model for the image modality, and an RNN recognition model for the voice modality.
[0084] Step 132: Initialize the model parameters and set the objective function for optimizing the model parameters during training.
[0085] Step 133: Configure the Adam optimizer, train the model, optimize the model parameters, reduce the loss function value, and obtain an optimized single-modal recognition model.
[0086] In the embodiments of the present invention, the most suitable model structure is selected for each modality, making full use of the characteristics of each modality. The Adam optimizer accelerates the model training process, improves the training efficiency, and the optimized model performs better in its respective modality, improving the recognition accuracy. The model structure and optimization method can be adjusted according to needs to adapt to different tasks and data. In scenarios where multiple modalities of data need to be processed, more accurate recognition results can be provided, and the more precise recognition ability can enhance user satisfaction and improve the interaction experience.
[0087] The specific steps in the embodiments of the present invention are as follows:
[0088] Step 131: Analyze the characteristics of data modalities, including text modality, image modality, and voice modality.
[0089] For the text modality, an LSTM model is constructed. The LSTM can effectively process long sequence data and capture long-term dependencies in the text. The input layer receives word embeddings or word vectors, the LSTM layer processes the sequence data, and the output layer performs classification or regression. Specifically, the text LSTM model has an input layer with the word embedding dimension, an LSTM layer with the number of hidden units set according to the task complexity, and an output layer designed according to the task type.
[0090] The image modality constructs a CNN model. The CNN can automatically learn the spatial features of images, such as edges, textures, etc. The input layer receives image pixels, the convolutional layer extracts features, the pooling layer reduces the dimension, and the fully connected layer performs classification or regression. Among them, the image CNN model includes that the input layer is the image size, the convolutional layer is the convolutional kernel size, number, and activation function, the pooling layer is the pooling method and pooling window size, and the fully connected layer is the number of nodes set according to the task complexity.
[0091] The voice modality constructs an RNN model. The RNN can process the time series characteristics of voice signals and capture the dependencies between frames. The input layer receives voice features, the RNN layer processes sequence data, and the output layer performs classification or regression. Among them, the voice RNN model includes that the input layer is the feature dimension, the RNN layer is the number of hidden units and activation function, and the output layer is designed according to the task type. Obtain LSTM, CNN, and RNN recognition models for text, image, and voice modalities.
[0092] Step 132, initialize the model parameters, including weight initialization: use Xavier initialization to ensure that the initial weight values are reasonable. Bias initialization: usually initialized to zero or a small constant. Set the objective function: through the cross-entropy loss function to measure the difference between the predicted probability and the true label, and use the mean squared error loss function to measure the squared difference between the predicted value and the true value.
[0093] Step 133, configure the Adam optimizer and perform hyperparameter settings, including learning rate: set to 0.001 or 0.0001; beta1 and beta2: set to 0.9 and 0.999 respectively and constants; send the data into the model, calculate the output value and the loss function value, calculate the gradient of the loss function with respect to the model parameters, use the Adam optimizer to update the model parameters, adjust the parameters to minimize the loss function. The number of samples used each time the parameters are updated affects the training speed and stability. The number of times the entire dataset is traversed. When the performance on the validation set no longer improves, terminate the training early to prevent overfitting. After multiple iterative trainings, the loss function value gradually decreases, and the performance of the model on the validation set reaches the optimal.
[0094] In a preferred embodiment of the present invention, the above step 14 may include:
[0095] Extract the key feature data that has been preprocessed and feature-extracted from the multi-modal dataset, input the preprocessed data into the corresponding single-modal recognition model, the model performs forward propagation, calculates the output value, and obtains the recognition results of text, image, and voice modalities.
[0096] In the embodiments of the present invention, the key feature data after preprocessing and feature extraction removes noise and redundant information, retains the most valuable features, enables the model to focus more on key information, improves the accuracy and effectiveness of the input. For different modalities such as text, image, and voice, single-modal recognition models are respectively constructed and optimized, which can make full use of the characteristics of each modality and improve the professionalism of recognition. Obtaining the recognition results of text, image, and voice modalities can more comprehensively understand the user input or scene information.
[0097] In a preferred embodiment of the present invention, step 15 may include:
[0098] Step 151, constructing the structure of the Bayesian network, including nodes and edges. Among them, each node represents the recognition result of a modality or the user intention, and the directed edges between the nodes represent the conditional dependence relationships between modalities or between modalities and intentions;
[0099] Step 152, setting the prior probability for each node in the Bayesian network and setting the conditional probability for each edge;
[0100] Step 153, according to the prior probability and conditional probability, combined with the single-modal recognition results, calculating the posterior probability of the user intention to obtain information of different modalities;
[0101] Step 154, integrating the information of different modalities to form a comprehensive preliminary user intention.
[0102] In the embodiments of the present invention, the Bayesian network can effectively model and reason about uncertainties, process complex multi-modal information, and through the prior probability and conditional probability, achieve the effective fusion of multi-modal information, improve the accuracy of intention recognition, the model structure is flexible, can be dynamically adjusted according to new evidence, adapt to different users and scenarios, and the quantified posterior probability provides strong support for decision-making, helps to formulate more reasonable strategies. Based on the comprehensive preliminary user intention, more personalized services and recommendations can be provided to enhance the user experience. The application of the Bayesian network provides new ideas and methods for multi-modal data processing and analysis, and promotes technological innovation in related fields.
[0103] The specific steps in the embodiments of the present invention are as follows:
[0104] Step 151, determining nodes: Each node represents the recognition result of a modality. Among them, the text modality node represents the recognition result of text data; the image modality node represents the recognition result of image data; the voice modality node represents the recognition result of voice data; the user intention node: represents the user's final intention or behavioral goal.
[0105] Determine edges: Directed edges represent the conditional dependency relationships between nodes. They point from the parent node to the child node, indicating that the state of the parent node affects the probability distribution of the child node.
[0106] Edge construction principle: For the dependency relationships between modalities, if the recognition result of one modality may affect the recognition of another modality, an edge is added. Dependency relationship between modality and intention: The recognition result of each modality may affect the user's intention. Therefore, an edge is drawn from each modality node to the user intention node. Determine reasonable dependency relationships using domain expert knowledge, including Text modality node → User intention node, Image modality node → User intention node, Voice modality node → User intention node, and Text modality node → Voice modality node. And use a graphical tool to draw the Bayesian network structure.
[0107] Step 152, Set prior probabilities: Prior probabilities represent the probability that a node takes a certain state in the absence of other evidence. Statistically analyze the occurrence frequencies of each node state from historical data. For example, for the text modality node: P(Positive sentiment)=0.6, P(Negative sentiment)=0.4, and for the user intention node: P(Purchase intention)=0.3, P(Consultation intention)=0.7.
[0108] Set conditional probabilities: Conditional probabilities represent the probability that a child node takes a certain state given that the parent node takes a specific state. Statistically analyze the joint probability of the parent node state and the child node state from historical data, and then calculate the conditional probability. For example, statistically analyze the probability that the user intention is to purchase when the text has a positive sentiment. Use a decision tree classification model to learn the conditional probabilities and set the conditional probabilities according to the experience of domain experts. P(Purchase intention|Positive sentiment)=0.8, P(Consultation intention|Negative sentiment)=0.9.
[0109] Step 153, According to the prior probabilities and conditional probabilities, obtain the recognition results of each modality from Step 14, apply Bayes' theorem, and calculate the posterior probability of each user intention. I is the user intention, is the recognition result of the i2-th modality, is the posterior probability, P(I) is the prior probability, is the conditional probability, is the probability of the evidence.
[0110] Step 154, Use the posterior probabilities calculated by the Bayesian network as the contributions of each modality to the user intention. Weightedly fuse the posterior probabilities of different modalities, select the intention with the highest posterior probability as the user's preliminary intention, and set a threshold. Only when the highest posterior probability exceeds the threshold is the intention determined; otherwise, the intention is considered unclear.
[0111] In a preferred embodiment of the present invention, the above Step 16 may include:
[0112] Step 161: Calculate the CTR value of each user based on historical data and divide the users into different groups.
[0113] Step 162: Extract the features of the user intent according to the CTR value of each user, and calculate the typical intent pattern of each user group based on the extracted features. Here, F′ is the comprehensive intent of the weighted user group G j , w j is the weight of the interaction frequency F f , w f is the weight of the modality switching frequency F s , w s is the weight of the occurrence times O c of a specific operation, w c is the weight of the interaction duration D d , N i is the number of users in the user group G j , F j is the interaction frequency of user i, F f,i is the modality switching frequency of user i, O s,i is the occurrence times of a specific operation of user j, D c,i i,i is the interaction duration of user i, and i is the index.
[0114] Step 163: Set the high CTR group as the first reference point and the low CTR group as the second reference point according to the typical intent pattern.
[0115] In the embodiments of the present invention, by quantitatively analyzing user behaviors and intents, data support is provided for decision-making, improving the accuracy and effectiveness of decision-making. For users in different CTR groups, more targeted marketing strategies can be formulated to improve marketing effects. Deeply understanding user intents and behavior patterns helps to provide more personalized services and recommendations, enhancing the user experience. The setting of reference points and the support for business decisions help to optimize business processes and improve business effects. By understanding user intents and behavior patterns, product design can be optimized and user satisfaction can be improved.
[0116] The specific steps in the embodiments of the present invention are as follows:
[0117] Step 161: CTR is the ratio of the number of user clicks to the number of displays. For each user, count the total number of clicks and the total number of displays, and calculate the CTR value of each user.
[0118] Users are divided into different groups according to the CTR value, and the division is made through a custom threshold. For example, high CTR group: CTR > 0.3; medium CTR group: 0.1 ≤ CTR ≤ 0.3; low CTR group: CTR < 0.1.
[0119] Step 162: According to the CTR value of each user, feature extraction is performed. The interaction frequency is the number of times a user interacts with the system, the modality switching frequency is the number of times a user switches between different modalities, the occurrence frequency of a specific operation is the number of times a user performs a specific operation, and the interaction duration is the average duration of each user interaction. For each user, the values of the above features are statistically calculated, and feature aggregation can be considered with a time window. For each user group, the weighted comprehensive intention pattern is calculated through a formula.
[0120] Step 163: Select the user group with a high CTR value as the first reference point. The typical intention pattern of this group represents the characteristics of highly engaged users. Select the user group with a low CTR value as the second reference point. The typical intention pattern of this group represents the characteristics of lowly engaged users.
[0121] In a preferred embodiment of the present invention, the above step 17 may include:
[0122] By Calculate the cross value of the user's preliminary intention with the first reference point and the second reference point. Among them, Cr is the cross value of the user's preliminary intention with the first reference point and the second reference point, w A and w B are the weights of the first reference point and the second reference point, U is the feature vector of the user's preliminary intention, B1 and B2 are the feature vectors of the first reference point and the second reference point, and are the cosine similarities between the user's preliminary intention and the first reference point and the second reference point, α A and α B are the weights of the feature interaction, and are the sums of the element-wise products of the features between the user's preliminary intention and the first reference point and the second reference point respectively, β is the weight of the reference point interaction, and k is the index.
[0123] In the embodiments of the present invention, by calculating the cross - value between the user's preliminary intention and the reference point, the similarity and difference between the user's intention and the reference point can be quantified. By setting weights for the first reference point and the second reference point, the attention degree of the model to different reference points can be flexibly adjusted to adapt to different business scenarios and requirements. By calculating the weights of feature interactions and the sum of element - by - element products of features between the user's preliminary intention and the reference point, the interactions between features can be captured, improving the expressive ability of the model. The calculation of the cross - value provides a quantitative basis for decision - making, helps to more objectively evaluate the relationship between the user's intention and the reference point, formulate more reasonable strategies. By understanding the relationship between the user's intention and the reference point, the business process can be optimized, improving business effectiveness and user satisfaction.
[0124] In a preferred embodiment of the present invention, the above - mentioned step 18 may include:
[0125] Step 181, obtain the CTR data of the user during the actual interaction process, and perform correlation analysis on the cross - value and the obtained CTR data to obtain an analysis result;
[0126] Step 182, formulate a strategy for adjusting the user's intention based on the cross - value and the analysis result of the obtained CTR data, and implement and monitor the effect of the adjustment strategy according to the adjustment strategy;
[0127] Step 183, optimize the user intention recognition model and the interaction process according to the monitoring result to achieve multimodal intention perception.
[0128] In the embodiments of the present invention, by obtaining the actual CTR data and performing correlation analysis on the cross - value and the CTR data, the CTR data reflects the user's behavior in the actual interaction. The correlation analysis with the cross - value helps to understand the relationship between the user's intention and behavior. The correlation analysis reveals the potential rules between the user's intention and CTR, providing strong support for optimizing the strategy. Based on the analysis results of the cross - value and the CTR data, more targeted adjustment strategies can be formulated to improve the accuracy of user intention recognition. Implementing and monitoring the effect of the adjustment strategy, the strategy can be dynamically optimized according to the actual effect to achieve continuous improvement. Optimizing the user intention recognition model according to the monitoring result can improve the accuracy and generalization ability of the model, better adapt to different users and scenarios. By optimizing the model and process to achieve multimodal intention perception, the user's intention can be more comprehensively understood, improving the accuracy of intention recognition.
[0129] The specific steps in the embodiments of the present invention are as follows:
[0130] Step 181: Extract actual interaction CTR data from the actual interaction records between the user and the system, including user ID, number of clicks, number of impressions, and interaction timestamps; align the CTR data and cross-value data according to the user ID and timestamp, calculate the Pearson correlation coefficient between the CTR and the cross-value for each user, analyze the relationship between the CTR and the cross-value using statistical methods, identify the key cross-value factors affecting the CTR, and display the analysis results through a scatter plot.
[0131] Step 182: Based on the analysis results of Step 181, determine the key cross-value factors affecting the CTR. For users with a high number of modality switches, optimize the interaction design to reduce unnecessary modality switches. For users with a low occurrence of specific operations, provide operation guides and prompts, and provide personalized content and recommendations according to the user's historical interaction patterns. Adjust the system interaction process, interface design, or recommendation algorithm according to the strategy, test the effect of the adjusted strategy on a small scale, collect data, and gradually deploy the strategy to the entire system according to the test results. Observe the increase or decrease in the CTR after the adjustment.
[0132] Step 183: Identify the impact of the adjustment strategy on the CTR and user engagement, and discover deficiencies in the user intention recognition model or interaction process; according to the monitoring results, add or adjust cross-value features, and retrain the model using the latest user interaction data. Integrate information in multiple modalities such as text, images, and voices to improve the comprehensiveness of intention recognition and enhance the model's capabilities.
[0133] As Figure 2 shown, an embodiment of the present invention also provides a multi-modal intention perception system 20 based on CTR improvement, including:
[0134] An acquisition module 21 for acquiring multi-modal data of a user.
[0135] An identification module 22 for reducing the dimension of high-dimensional features of the multi-modal data through PCA, constructing a feature extraction network, and extracting key feature data; constructing a corresponding identification model according to each modality, and optimizing the model parameters through Adam to obtain a single-modal identification model; identifying the key feature data according to the single-modal identification model to obtain the identification results of each modality.
[0136] A processing module 23 for combining the identification results with a Bayesian network to perform multi-modal information fusion to obtain a preliminary user intention; calculating typical intention patterns of different CTR user groups according to historical data as the first reference point and the second reference point; calculating the cross-values between the preliminary user intention and the first reference point and the second reference point; adjusting the user intention according to the cross-value and CTR data feedback to achieve multi-modal intention perception.
[0137] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A multi-modal intent perception method based on CTR improvement, characterized in that The method includes: Obtain the multi-modal data of the user; Reduce the high-dimensional features of the multi-modal data through PCA, construct a feature extraction network, and extract key feature data; According to each modality, construct a corresponding recognition model, and optimize the model parameters through Adam to obtain a single-modal recognition model; According to the single-modal recognition model, identify the key feature data to obtain the recognition results of each modality; Combine the recognition results with the Bayesian network for multi-modal information fusion to obtain the user's preliminary intention; According to historical data, calculate the typical intention patterns of different CTR user groups as the first reference point and the second reference point; Calculate the cross values of the user's preliminary intention with the first reference point and the second reference point; Adjust the user's intention according to the cross values and CTR data feedback to achieve multi-modal intention perception.
2. The multi-modal intention perception method based on CTR improvement according to claim 1, wherein, Reduce the high-dimensional features of the multi-modal data through PCA, construct a feature extraction network, and extract key feature data, including: Preprocess the collected multi-modal data to obtain preprocessed multi-modal data; Convert the preprocessed multi-modal data into a matrix, where each row represents a sample and each column represents a feature, and calculate the covariance matrix of the data matrix; Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors; According to the magnitudes of the eigenvalues, select the eigenvectors corresponding to the top k largest eigenvalues as the principal components; Construct a transformation matrix from the principal components and multiply it with the original high-dimensional data matrix to obtain the projected low-dimensional data matrix; Construct a feature extraction network based on the low-dimensional data matrix.
3. The multi-modal intention perception method based on CTR improvement according to claim 1, characterized in that, According to each modality, construct a corresponding recognition model, and optimize the model parameters through Adam to obtain a single-modal recognition model, including: According to each modality, construct LSTM, CNN, and RNN recognition models. Among them, an LSTM recognition model is constructed for the text modality, a CNN recognition model is constructed for the image modality, and an RNN recognition model is constructed for the voice modality; Initialize the model parameters and set the objective function for optimizing the model parameters during training; Configure the Adam optimizer to train the model, optimize the model parameters, and reduce the loss function value to obtain the optimized single-modal recognition model.
4. The multi-modal intent perception method based on CTR improvement according to claim 1, wherein Combine the recognition results with the Bayesian network for multi-modal information fusion to obtain the user's preliminary intention, including: Construct the structure of the Bayesian network, including nodes and edges. Among them, each node represents the recognition result of a modality or the user's intention, and the directed edges between the nodes represent the conditional dependence relationships between modalities or between modalities and intentions; Set the prior probability for each node in the Bayesian network and the conditional probability for each edge; According to the prior probability and conditional probability, combine the single-modal recognition results to calculate the posterior probability of the user's intention to obtain information of different modalities; Integrate the information of different modalities to form a comprehensive user's preliminary intention.
5. The multi-modal intention perception method based on CTR improvement according to claim 1, characterized in that According to historical data, calculate the typical intention patterns of different CTR user groups as the first reference point and the second reference point, including: According to historical data, calculate the CTR value of each user and divide the users into different groups; Extract the features of user intent according to the CTR value of each user, and according to the extracted features, through calculate the typical intent patterns of each user group, where F′ j is the comprehensive intent of the weighted user group G j , w f is the weight of the interaction frequency F f , w s is the weight of the modality switching frequency F s ; w c is the weight of the occurrence times O c of a specific operation, w d is the weight of the interaction duration D i ; N j is the number of users in the user group G j ; F f,i is the interaction frequency of user i, F s,i is the modality switching frequency of user i, O c,i is the occurrence times of a specific operation of user j, D i,i is the interaction duration of user i, where i is the index; According to the typical intention pattern, the high-CTR group is set as the first reference point, and the low-CTR group is set as the second reference point.
6. The multi-modal intention perception method based on CTR improvement according to claim 1, characterized in that, Calculate the cross values of the user's preliminary intention with the first reference point and the second reference point, including: By calculating the cross value of the user's preliminary intention with the first reference point and the second reference point, where Cr is the cross value of the user's preliminary intention with the first reference point and the second reference point, w A and w B are the weights of the first reference point and the second reference point, U is the feature vector of the user's preliminary intention, B1 and B2 are the feature vectors of the first reference point and the second reference point, and are the cosine similarities of the user's preliminary intention with the first reference point and the second reference point, α A and α B are the weights of the feature interaction, and are the sums of the element-wise products of the features between the user's preliminary intention and the first reference point and the second reference point respectively, β is the weight of the reference point interaction, and k is the index.
7. The multi-modal intention perception method based on CTR improvement according to claim 6, wherein, Adjust the user's intention according to the cross values and CTR data feedback to achieve multi-modal intention perception, including: Obtain the CTR data of the user during the actual interaction process, and conduct correlation analysis on the cross values and the obtained CTR data to obtain the analysis result; Formulate a strategy for adjusting the user's intention according to the cross values and the analysis result of the obtained CTR data, and implement and monitor the effect of the adjustment strategy according to the adjustment strategy; Optimize the user intention recognition model and interaction process according to the monitoring result to achieve multi-modal intention perception.
8. A multi-modal intention perception system based on CTR improvement, which implements the method according to any one of claims 1 to 7, characterized in that Including: An acquisition module for acquiring the multi-modal data of the user; An identification module for reducing the dimensionality of high-dimensional features of the multi-modal data through PCA, constructing a feature extraction network, and extracting key feature data; constructing a corresponding identification model according to each modality, and optimizing the model parameters through Adam to obtain a single-modal identification model; identifying the key feature data according to the single-modal identification model to obtain the identification results of each modality; A processing module for combining the identification results with a Bayesian network for multi-modal information fusion to obtain the user's preliminary intention; Calculate the typical intention patterns of different CTR user groups based on historical data as the first reference point and the second reference point; Calculate the cross values of the user's preliminary intention with the first reference point and the second reference point; adjust the user's intention according to the cross values and CTR data feedback to achieve multi-modal intention perception.
9. A computing device, characterized in that, Including: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A program is stored in the computer-readable storage medium, and when the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.