Outbound method, system and device based on multi-modal user portrait fusion and medium
By integrating multimodal data and optimizing dynamic outbound calling strategies, the problem of existing outbound calling systems being unable to respond quickly to market changes has been solved, enabling personalized outbound calling and improving customer experience and conversion rates.
Patent Information
- Application Number
- CN202510967156.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-28
AI Technical Summary
Existing outbound calling systems rely on static rules and pre-set scripts, which cannot quickly respond to market changes and customer needs. They also lack multimodal information fusion, resulting in poor customer experience and low conversion rates.
By acquiring multimodal data, constructing a four-dimensional feature vector, using the K-means++ algorithm for clustering, and optimizing the user profile template based on user feedback data, the outbound calling strategy is dynamically adjusted to achieve personalized outbound calling.
It improves the accuracy and efficiency of outbound calls, reduces ineffective communication, enhances marketing success rates and customer experience, and ensures the long-term effectiveness and adaptability of user profile templates.
Smart Images

Figure CN121030291A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent outbound call, in particular to an outbound call method and system based on multi-modal user portrait fusion, equipment and medium. BACKGROUND
[0002] In the current competitive market environment, enterprises need to improve marketing conversion rate, optimize customer service experience and enhance customer loyalty through efficient customer communication means. Outbound call, as an important way to actively reach customers, is widely used in product promotion, after-sales service, customer follow-up and other scenarios. However, the needs, preferences and behavior characteristics of different customers are significantly different. If a unified outbound call method is used, it may lead to customer dissatisfaction, low connection rate and poor conversion rate. Therefore, developing personalized and precise outbound call strategies can effectively improve outbound call efficiency, enhance customer satisfaction and achieve optimal resource allocation and business goals.
[0003] Currently, most outbound call systems mainly rely on static rules and preset dialogue trees for strategy development, usually based on customer's basic information (such as age, gender, region) or simple historical behavior data for classification and matching of fixed dialogue templates. This method, although to some extent, realizes customer grouping, has the following defects: first, the strategy updating cycle is long, which cannot quickly respond to market changes and customer needs; second, the data source is single, lacking the fusion of multi-modal information such as voice emotion and social behavior, resulting in incomplete portrait; third, the strategy matching rule is simple, difficult to cover complex and variable long-tail scenarios, easily causing poor customer experience and low conversion rate. SUMMARY
[0004] The present application provides an outbound call method, system, device and medium based on multi-modal user portrait fusion, which can improve the accuracy and efficiency of outbound call.
[0005] An embodiment of the present application provides an outbound call method based on multi-modal user portrait fusion, comprising:
[0006] Obtaining multi-modal data of a user, wherein the multi-modal data includes user information, text data, voice data and image data;
[0007] Respectively extracting features and weighting fusion of the multi-modal data to construct a four-dimensional feature vector representing the user's basic attributes, consumption behavior, voice interaction and social media, and performing clustering processing on the four-dimensional feature vector through K-means++ algorithm to obtain user grouping results, wherein the four-dimensional feature vector;
[0008] The user clustering result is matched with a preset user portrait template to determine a user type, the user type is input into a preset outbound call strategy mapping table, a corresponding outbound call strategy is matched, and outbound calls are performed according to the outbound call strategy, wherein threshold determination criteria of each dimension in the user portrait template are dynamically adjusted and optimized based on an outbound call score, and the outbound call score is calculated according to user feedback data in an outbound call process.
[0009] The embodiment of the application can obtain multi-modal data, more comprehensively understand user characteristics, and further improve subsequent user portrait accuracy; the four-dimensional feature vector can be constructed to convert different types of data into a calculable feature vector; clustering can divide users into several groups, facilitating subsequent differentiated strategy formulation; matching the user clustering result with the user portrait template can quickly and accurately map the clustering result to a specific portrait, avoiding re-modeling each time; inputting the user type into the outbound call strategy mapping table can realize personalized outbound calls, ensure the accuracy of outbound calls, reduce ineffective communication, improve marketing success rate, avoid ineffective outbound calls to users who are not interested, and improve customer experience; adjusting the portrait determination criteria according to user feedback data can continuously optimize the user portrait template, ensuring long-term effectiveness and self-adaptability of the user portrait template. Compared with the prior art, the application can improve the accuracy and efficiency of outbound calls.
[0010] Further, the multi-modal data of the user is obtained, specifically:
[0011] The initial multi-modal data of the user is obtained, missing values of the initial multi-modal data are identified, the missing values are filled according to the type of the missing values using a corresponding filling method, and a first processing result is obtained;
[0012] The IQR method or the Z-score method is used to detect abnormal values of the first processing result, and the abnormal values are corrected or deleted to obtain a second processing result;
[0013] The second processing result is standardized to obtain a third processing result, and the category type features of the third processing result are one-hot encoded or label encoded to obtain the multi-modal data.
[0014] The multi-modal data is obtained, the user characteristics can be more comprehensively understood, and the subsequent user portrait accuracy is improved.
[0015] Further, the multi-modal data is respectively subjected to feature extraction and weighted fusion to construct a four-dimensional feature vector representing the basic attributes, consumption behavior, voice interaction and social media of the user, specifically:
[0016] According to the category of the multi-modal data, a corresponding feature extraction method is selected to perform feature extraction, and basic attribute feature data, consumption behavior feature data, voice interaction feature data and social media feature data are obtained respectively.
[0017] The basic attribute feature data, the consumption behavior feature data, the voice interaction feature data and the social media feature data are fused to construct an initial four-dimensional feature vector representing the basic attributes, consumption behavior, voice interaction and social media of the user, and the initial four-dimensional feature vector is normalized to obtain the four-dimensional feature vector.
[0018] In this way, by constructing a four-dimensional feature vector, different types of data can be converted into a computable feature vector.
[0019] Further, the corresponding feature extraction method is selected according to the category of the multi-modal data to perform feature extraction, and basic attribute feature data, consumption behavior feature data, voice interaction feature data and social media feature data are obtained respectively, specifically:
[0020] The user information is numerically encoded and segmented and normalized to obtain the basic attribute feature data;
[0021] The text data is classified by consumption type, the amount frequency is counted, and the promotion sensitivity is analyzed to obtain the consumption behavior feature data;
[0022] Acoustic features and semantic features of the voice data are extracted to obtain the voice interaction feature data, wherein the acoustic features include speech rate, tone fluctuation and interaction interruption rate, and the semantic features include emotional tendency and preference rhetoric keywords;
[0023] The image data is subjected to image recognition to obtain the user's interest circle, KOL attention relationship and content interaction features to obtain the social media feature data.
[0024] In this way, by extracting basic attribute feature data, consumption behavior feature data, voice interaction feature data and social media feature data, it is helpful to construct a thinking feature vector, and different types of data can be converted into a computable feature vector.
[0025] Further, the four-dimensional feature vector is clustered by the K-means++ algorithm to obtain a user clustering result, specifically:
[0026] The elbow rule is used to determine the target cluster number, the cluster centers are initialized through the K-means++ algorithm, and iterative cluster analysis is performed on the four-dimensional feature vector to obtain a cluster result until the silhouette coefficient of the cluster result meets a preset threshold, and a cluster label to which each user belongs is output to determine the user grouping result.
[0027] In this way, the users can be divided into several groups through clustering and grouping, facilitating subsequent differentiated strategy making.
[0028] Further, the user feedback data includes a connection rate, a call duration, a hang-up rate and a customer emotion fluctuation amplitude, and the outbound call score is calculated according to the user feedback data in the outbound call process, specifically:
[0029] Obtaining a call record in the outbound call process and analyzing the call record to obtain initial user feedback data;
[0030] Standardizing the initial user feedback data to obtain the connection rate, the call duration, the hang-up rate and the customer emotion fluctuation amplitude;
[0031] According to the calculation method of each user feedback data, the connection rate, the call duration, the hang-up rate and the customer emotion fluctuation amplitude are calculated to obtain a connection rate score, a call duration score, a hang-up rate score and a customer emotion fluctuation score;
[0032] According to a preset weight, the connection rate score, the call duration score, the hang-up rate score and the customer emotion fluctuation score are weighted and summed to obtain the outbound call score.
[0033] In this way, the portrait judgment standard is continuously adjusted according to the user feedback data, the user portrait template can be continuously optimized, and the long-term effectiveness and adaptive ability of the user portrait template are ensured.
[0034] Another embodiment of the application also provides an outbound call device based on multi-modal user portrait fusion, comprising: an acquisition module, a fusion module and an outbound call module;
[0035] The acquisition module is configured to acquire multi-modal data of a user, wherein the multi-modal data includes user information, text data, voice data and image data;
[0036] The fusion module is configured to perform feature extraction and weighted fusion on the multi-modal data respectively to construct a four-dimensional feature vector representing the basic attributes, consumption behavior, voice interaction and social media of the user, and perform clustering processing on the four-dimensional feature vector through a K-means++ algorithm to obtain a user grouping result, wherein the four-dimensional feature vector;
[0037] The outbound call module is configured to match the user grouping result with a preset user portrait template to determine a user type, input the user type into a preset outbound call strategy mapping table, match a corresponding outbound call strategy, and perform outbound call according to the outbound call strategy.
[0038] The embodiment of the present application can obtain multi-modal data to comprehensively understand user characteristics and improve subsequent user portrait accuracy, convert different types of data into calculable feature vectors by constructing a four-dimensional feature vector, divide users into several groups by clustering and grouping, facilitate subsequent differentiated strategy formulation, quickly and accurately map clustering results to specific portraits by matching user grouping results with a user portrait template, avoid re-modeling each time, realize personalized outbound call by inputting user types into an outbound call strategy mapping table, ensure the accuracy of outbound call, reduce ineffective communication, improve marketing success rate, avoid ineffective outbound call to users who are not interested, and improve customer experience, continuously optimize the user portrait template by adjusting portrait judgment criteria according to user feedback data, and ensure long-term effectiveness and adaptive ability of the user portrait template. Compared with the prior art, the present application can improve the accuracy and efficiency of outbound call.
[0039] Further, the fusion module includes an extraction unit and a fusion unit, specifically:
[0040] The extraction unit is configured to select a corresponding feature extraction method according to the category of the multi-modal data to perform feature extraction, and obtain basic attribute feature data, consumption behavior feature data, voice interaction feature data, and social media feature data.
[0041] The fusion unit is configured to fuse the basic attribute feature data, the consumption behavior feature data, the voice interaction feature data, and the social media feature data, construct an initial four-dimensional feature vector representing the basic attributes, consumption behavior, voice interaction, and social media of the user, and normalize the initial four-dimensional feature vector to obtain the four-dimensional feature vector.
[0042] Another embodiment of the present application also provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the steps of the outbound call method based on multi-modal user portrait fusion as described in the present application are implemented.
[0043] Another embodiment of the present application also provides a computer readable storage medium item, comprising: a stored computer program, when the computer program runs, controls a device where the computer readable storage medium is located to execute steps of the outbound call method based on multi-modal user portrait fusion as described in the present application. BRIEF DESCRIPTION OF DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described in the following are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0045] Figure 1 is a flowchart of an embodiment of the outbound call method based on multi-modal user portrait fusion provided by the present application;
[0046] Figure 2 is a flowchart of steps S201 to S202 provided by the present application;
[0047] Figure 3 is a K-SSE curve diagram provided by the present application;
[0048] Figure 4 is a structural diagram of an embodiment of the outbound call system based on multi-modal user portrait fusion provided by the present application. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solutions and advantages of the present application more clear, the technical solutions in the present application will be described clearly and completely in the following with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the present application; the terms "include" and "have" and any variations thereof in the specification and claims of the present application and the above drawing description are intended to cover not exclusive inclusion.
[0051] In the description of the embodiments of the present application, the technical terms "first", "second" and the like are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "multiple" is more than two, unless otherwise explicitly and specifically limited.
[0052] Reference herein to "embodiments" means that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily all refer to the same embodiment, nor is it necessarily independent or alternative to other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0053] In the description of the embodiments of the present application, the term "and / or" is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " herein generally represents an "or" relationship between the front and rear associated objects.
[0054] In the description of the embodiments of the present application, the term "multiple" refers to more than two (including two), and similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).
[0055] In the description of the embodiments of the present application, unless otherwise explicitly specified and limited, the technical terms "mounting", "connection", "connection", "fixing" and the like should be understood in a broad sense, for example, it can be fixedly connected, or it can be detachably connected, or it can be integrated; it can be mechanical connection, or it can be electrical connection; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the internal communication of two elements or the interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the embodiments of the present application can be understood according to the specific circumstances.
[0056] In the competitive market environment, outbound call is an important means for enterprises to actively reach out to customers, improve marketing conversion rate and service experience. However, there are significant differences between customers, and uniform outbound call methods can easily cause customers to be dissatisfied, have low connection rate and poor conversion rate, etc., so it is imperative to develop personalized and precise outbound call strategies. At present, most outbound call systems rely on static rules and preset dialogue trees, and classify and match fixed dialogue templates based on customer basic information or simple historical behavior data. This way has a long strategy update cycle and cannot quickly respond to market changes and customer needs, and lacks the fusion of multi-modal information, making it difficult to cover complex and variable long-tail scenarios, and the customer experience and conversion rate are poor.
[0057] Referring to Figure 1 To improve the accuracy and efficiency of outbound calls, an embodiment of the present application provides an outbound call method based on multi-modal user portrait fusion, comprising steps S101 to S103.
[0058] Step S101, acquiring multi-modal data of a user, wherein the multi-modal data comprises user information, text data, voice data and image data;
[0059] In some embodiments, the multi-modal data of the user is acquired, specifically: acquiring initial multi-modal data of the user, identifying missing values of the initial multi-modal data, and filling the missing values according to the type of the missing values to obtain a first processing result; detecting outliers of the first processing result using IQR method or Z-score method, and correcting or deleting the outliers to obtain a second processing result; performing standardization processing on the second processing result to obtain a third processing result, and performing one-hot encoding or label encoding on the category type features of the third processing result to obtain the multi-modal data. Specifically, first, the initial multi-modal data of the user is collected from CRM systems, customer service systems, APPs, Web platforms, etc. through API interfaces, database queries, log collection, etc., wherein the initial multi-modal data comprises user information, text data, voice data and image data; then, according to the type of the initial multi-modal data, a corresponding identification method is selected to identify the missing values of the initial multi-modal data, and according to the type of the missing values (numeric features, category features, text / voice / image data), a corresponding filling method is selected to fill the missing values to obtain a first processing result; then, IQR (interquartile range), Z-score, etc. are used to detect outliers of the first processing result, if the outliers are obviously wrong (such as age of 200 years old), they are directly deleted, if the outliers are possibly extreme but reasonable (such as extremely high income), they are corrected (such as winsorize or replaced with upper and lower limit values) to obtain a second processing result; then, Min-Max standardization or Z-score standardization is used to standardize the data of different dimensions (such as age and consumption amount) in the second processing result to the same scale to obtain a third processing result; finally, if the category type features of the third processing result are unordered category type features (such as gender, region), one-hot encoding is performed, for example: gender (male, female) -> [1, 0] or [0, 1], if the category type features of the third processing result are ordered category type features (such as education level: high school, undergraduate, master), label encoding is performed, for example: high school -> 0, undergraduate -> 1, master -> 2, thus completing the encoding operation of the third processing result, and the multi-modal data can be obtained.
[0060] It should be noted that user information includes: basic customer information, transaction records, etc.; text data includes user messages and comments on websites and apps, customer service chat logs, questionnaires, etc.; image data includes user-uploaded profile pictures, ID photos, active social media images such as lifestyle photos, or images pushed to users; and voice data includes outbound call voice messages and voice chat logs between customer service and customers, etc.
[0061] It should be noted that pandas' isnull() method is used to detect missing values for user information (such as age, income, and spending amount); while for text data, voice data, and image data, if a field is empty or cannot be parsed, it is also considered missing.
[0062] It should be noted that for numerical features (such as age and income), if the data distribution is relatively uniform, the mean or median should be used to fill the data; if the data shows obvious skewness, the median should be used. For categorical features (such as gender and occupation), the mode (the category with the highest frequency) should be used to fill the data. For text / voice / image data, if there are missing values, consider using a null label (such as "unknown") or deleting the sample.
[0063] By acquiring multimodal data, we can gain a more comprehensive understanding of user characteristics, thereby improving the accuracy of subsequent user profiling.
[0064] Step S102: Feature extraction and weighted fusion are performed on the multimodal data respectively to construct a four-dimensional feature vector representing the user's basic attributes, consumption behavior, voice interaction and social media. The four-dimensional feature vector is then clustered using the K-means++ algorithm to obtain the user grouping results.
[0065] Please refer to Figure 2 In some embodiments, feature extraction and weighted fusion are performed on the multimodal data respectively to construct a four-dimensional feature vector representing the user's basic attributes, consumption behavior, voice interaction and social media, including steps S201 to S202.
[0066] Step S201: Select the corresponding feature extraction method according to the category of the multimodal data to extract features, and obtain basic attribute feature data, consumer behavior feature data, voice interaction feature data and social media feature data respectively;
[0067] In some embodiments, the step of selecting a corresponding feature extraction method based on the category of the multimodal data to extract features, thereby obtaining basic attribute feature data, consumer behavior feature data, voice interaction feature data, and social media feature data, specifically involves: performing numerical encoding and segmented normalization on the user information to obtain the basic attribute feature data; performing consumption type classification, amount frequency statistics, and promotion sensitivity analysis on the text data to obtain the consumer behavior feature data; extracting acoustic and semantic features from the voice data to obtain the voice interaction feature data, wherein the acoustic features include speech rate, tone fluctuation, and interaction interruption rate, and the semantic features include emotional tendency and preferred keywords; and performing image recognition on the image data to obtain the user's interest circles, KOL follow relationships, and content interaction features, thereby obtaining the social media feature data.
[0068] In some embodiments, the user information is numerically encoded and segmented normalized to obtain the basic attribute feature data. Specifically, firstly, for user information (static tags such as age, gender, region, income, and occupation), for age / income, segmented normalization is used to map the original data to the 0-1 interval. The feature extraction formula is: feature value = (actual value - minimum value) / (maximum value - minimum value), for example, age feature = (actual age - 18) / (65 - 18). Secondly, for gender / region, one-hot encoding is used. Encoding converts categorical data into binary vectors (e.g., male is represented as 1,0, female as 0,1); then, for occupation, weights are assigned based on industry influence (e.g., technology industry = 0.8, education industry = 0.6, etc.); finally, the results are weighted and fused to obtain the basic attribute feature data (static value dimension), which is the first dimension of the four-dimensional feature vector. The relevant calculation formula is V1 = 0.4 * income + 0.3 * occupation + 0.2 * region + 0.1 * age.
[0069] By extracting basic attribute feature data, consumer behavior feature data, voice interaction feature data, and social media feature data, it is possible to construct a thought feature vector, which can transform different types of data into computable feature vectors.
[0070] In some embodiments, the text data is classified into consumption types, the frequency of purchase amounts is statistically analyzed, and the sensitivity to promotions is analyzed to obtain the consumption behavior characteristic data. Specifically, the following steps are taken: First, a category preference matrix is constructed, for example, luxury goods = 0.9 and necessities = 0.3, to obtain the consumption types; second, for the consumption amount / frequency, Z-score standardization is used to statistically analyze the frequency of purchase amounts; then, for the sensitivity to promotions, the product of the participation rate and the discount intensity is calculated to reflect the user's sensitivity to promotional activities; finally, the results are weighted and fused to obtain the consumption behavior characteristic data (economic sensitivity dimension), which is the second dimension of the four-dimensional feature vector. The relevant calculation formula is: V2 = 0.5 * promotion sensitivity + 0.3 * consumption type + 0.2 * amount fluctuation.
[0071] In some embodiments, acoustic and semantic features of the speech data are extracted to obtain the speech interaction feature data. The acoustic features include speech rate, intonation fluctuation, and interaction interruption rate. The semantic features include sentiment tendency and preferred dialogue keywords. Specifically: First, based on the speech data, the number of words per minute is calculated and normalized to obtain the speech rate. Second, the MFCC (Melbourne Frequency Cepstral Coefficient) features of the speech data are extracted and the standard deviation is calculated to obtain the intonation fluctuation. Then, the ratio of the number of abnormal terminations in the speech data to the total number of interactions is calculated to obtain the interaction interruption rate. Subsequently, keyword sentiment analysis is performed on the speech data, with positive sentiment = 0.8 and negative sentiment = 0.2, to obtain the preferred dialogue. Finally, the results are weighted and fused to obtain the speech interaction feature data (communication trait dimension), which is the third dimension of the four-dimensional feature vector. The relevant calculation formula is: V3 = 0.4 * intonation fluctuation + 0.3 * interruption rate + 0.2 * dialogue + 0.1 * speech rate.
[0072] In some embodiments, image recognition is performed on the image data to obtain the user's interest circles, KOL follow relationships, and content interaction features, thereby obtaining the social media feature data. Specifically, the image data is first subjected to image recognition to extract content that the user is interested in, and the user's interest circles, KOL follow relationships, and content interaction features are determined based on the content that the user is interested in. For interest circles, the community participation depth (views: comments: reposts = 1:3:5) is calculated based on the user's social data; for KOL follow, the influence of the followed objects is weighted (top KOLs = 1.0, mid-tier KOLs = 0.6); for content interaction features, the ratio of user dwell time to content length is calculated based on user dwell time data. Then, the obtained interest circles, KOL follow relationships, and content interaction features are weighted and fused to obtain the social media feature data (interest influence dimension), which is the fourth dimension of the four-dimensional feature vector. The relevant calculation formula is: V4 = 0.5 * interest circles + 0.3 * KOL follow + 0.2 * content dwell time.
[0073] By extracting basic attribute feature data, consumer behavior feature data, voice interaction feature data, and social media feature data, it is possible to construct a thought feature vector, which can transform different types of data into computable feature vectors.
[0074] Step S202: The basic attribute feature data, the consumption behavior feature data, the voice interaction feature data, and the social media feature data are fused to construct an initial four-dimensional feature vector representing the user's basic attributes, consumption behavior, voice interaction, and social media. The initial four-dimensional feature vector is then normalized to obtain the four-dimensional feature vector.
[0075] In some embodiments, the basic attribute feature data V1, the consumption behavior feature data V2, the voice interaction feature data V3, and the social media feature data V4 are fused to construct an initial four-dimensional feature vector [V1, V2, V3, V4] representing the user's basic attributes, consumption behavior, voice interaction, and social media, for example [0.82, 0.75, 0.31, 0.68].
[0076] In some embodiments, the initial four-dimensional feature vector is normalized to obtain the four-dimensional feature vector. Specifically, since different features in the four-dimensional feature vector may have different dimensions and value ranges, such as age being between 0 and 100, and income being between several hundred and tens of thousands, normalization is needed to scale the data proportionally to make it fall into a specific interval, which is [0,1] or [-1,1], in order to eliminate such dimension differences and avoid some features having too much influence due to large numerical ranges. Among them, the min-max normalization method can be used to process, and all feature values are scaled to the [0,1] interval through Min-Max to eliminate dimension differences. When the effect is not good, the Z-score normalization method is tried to eliminate dimension differences between different features, avoid some features having too much influence due to large numerical ranges, and ensure the accuracy and stability of the clustering algorithm.
[0077] It should be noted that min-max normalization can scale features to the [0,1] interval, with the formula Xnorm = (X - Xmin) / (Xmax - Xmin), where X is the original feature value, and Xmin and Xmax are the minimum and maximum values of the feature, respectively.
[0078] It should be noted that Z-score normalization is based on the mean and standard deviation of features to normalize them so that the features have zero mean and unit variance. The formula is Xnorm=(X-μ) / σ, where μ is the mean of the feature and σ is the standard deviation of the feature.
[0079] For example, after normalization, the four-dimensional feature vectors of different users can be: User 1: [0.82, 0.75, 0.31, 0.68]; User 2: [0.43, 0.92, 0.57, 0.39]; User 3: [0.61, 0.54, 0.42, 0.87].
[0080] In some embodiments, the step of clustering the four-dimensional feature vector using the K-means++ algorithm to obtain user grouping results specifically involves: determining the target number of clusters using the elbow rule, initializing cluster centers using the K-means++ algorithm, and performing iterative clustering analysis on the four-dimensional feature vector to obtain clustering results until the silhouette coefficient of the clustering results meets a preset threshold. Finally, the cluster label to which each user belongs is output to determine the user grouping results. Specifically, firstly, K is gradually increased from 1 (typically within the range of 2-10), the K-means++ algorithm is run each time and the sum of squared errors (SSE) is calculated. The horizontal axis represents the K value, the vertical axis represents the corresponding SSE value, and the points are connected to form a line graph (e.g., ...). Figure 3 As shown in the figure, the inflection point in the line graph is identified, which is the location where the decrease in SSE slows down significantly, and the K value corresponding to the inflection point is used as the target number of clusters. Next, the K-means++ method is used to initialize the cluster centers. This method selects the initial cluster centers through probability distribution, making the initial cluster centers more evenly distributed and accelerating convergence. Then, for each sample point, the distance to each initial cluster center is calculated, and it is assigned to the cluster where the nearest initial cluster center is located. The cluster center of each cluster is updated to the mean of all points in the cluster. The above process is repeated until the silhouette coefficient of the clustering result is less than a preset threshold (such as 0.5). The silhouette coefficient of the clustering result is used to evaluate the clustering effect. If the silhouette coefficient does not meet the preset threshold (such as 0.5), the K value is adjusted and the K-means algorithm is run again until the silhouette coefficient meets the preset threshold. At this time, the cluster label of each user is output to determine the user grouping result.
[0081] In some embodiments, the relevant formula for calculating the sum of squared errors is: In the formula, C i For the i-th cluster, μ i Let x be the centroid of the cluster and x be a sample point.
[0082] By clustering users, users can be divided into several groups, which facilitates the formulation of subsequent differentiated strategies.
[0083] Step S103: Match the user grouping results with the preset user profile template to determine the user type, and input the user type into the preset outbound call strategy mapping table to obtain the corresponding outbound call strategy, and make outbound calls according to the outbound call strategy. The threshold judgment criteria of each dimension in the user profile template are dynamically adjusted and optimized based on the outbound call score, and the outbound call score is calculated based on user feedback data during the outbound call process.
[0084] In some embodiments, the user grouping results are matched with a preset user profile template to determine the user type. Specifically, the thought feature vector of each cluster (e.g., [0.82, 0.25, 0.56, 0.78]) is compared with the user profile template. Based on the high and low thresholds of each dimension (V1~V4) (e.g., high ≥ 0.67, medium ≥ 0.33, low < 0.33), it is determined which type of profile the cluster belongs to, thereby determining the user profile type (e.g., "Type A") of each user. For example, the user profile template uses high, medium, and low to distinguish the performance of each feature in its respective dimension. Assuming the thresholds are 0.67 and 0.33, the feature distribution of [0.82, 0.25, 0.56, 0.78] is [high V1, low V2, medium V3, high V4], corresponding to the user profile [Type A].
[0085] Table 2 is a user profile template provided in this application:
[0086]
[0087]
[0088] It should be noted that the user profile templates are predetermined based on industry experience and expert advice. Specifically, the feature distribution is defined as eight user profile types: Type A [High V1, Low V2, Medium V3, High V4], Type B [Low V1, High V2, High V3, High V4], and Type C [Medium V1, High V2, Low V3, Low V4].
[0089] In some embodiments, the user type is input into a preset outbound call strategy mapping table to match the corresponding outbound call strategy, and an outbound call is made according to the outbound call strategy. Specifically, after the user type is determined, the user type is input into the preset outbound call strategy mapping table to find the corresponding outbound call strategy. The outbound call strategy mapping table predefines the mapping relationship between user profile types and outbound call strategies. Then, after the outbound call strategy is determined, the outbound call system automatically dials the user's phone number according to the strategy content.
[0090] Table 3 shows the mapping relationship between the user profile types and outbound calling strategies provided in this application:
[0091] User portrait Outbound call strategy 1. Type A Push private customized service, first launch of scarce goods 2. Type B Fission coupons, community flash sale activities 3. Type C Price trend reminder, competitor comparison tool 4. Type D Pre-sale lock for large promotion, KOL co-branded set 5. Type E Live streaming flash shopping, interest-free for a limited time 6. Type F Periodic purchase subscription, family combination set discount 7. Type G Free shipping insurance, exclusive quality inspection report 8. Type H All-network price guarantee: Promise to "buy expensive and make up for the difference"
[0092] In some embodiments, the user feedback data includes connection rate, call duration, hang-up rate, and customer emotional fluctuation amplitude. The outbound call score is calculated based on the user feedback data during the outbound call process. Specifically, this involves: acquiring call records during the outbound call process and analyzing the call records to obtain initial user feedback data; standardizing the initial user feedback data to obtain connection rate, call duration, hang-up rate, and customer emotional fluctuation amplitude; calculating the connection rate, call duration, hang-up rate, and customer emotional fluctuation amplitude according to the calculation methods for each of the user feedback data to obtain connection rate score, call duration score, hang-up rate score, and customer emotional fluctuation score; and weighting and summing the connection rate score, call duration score, hang-up rate score, and customer emotional fluctuation score according to preset weights to obtain the outbound call score. Specifically, firstly, during a call between the outbound calling system and a customer, the system automatically records detailed call information, including the call start time, end time, call status (connected, not connected, hung up, etc.), and call content (recording). Then, the call records are processed using speech recognition, semantic analysis, and sentiment analysis to determine initial user feedback data. Next, this initial user feedback data is cleaned and standardized to obtain processed connection rate, call duration, hang-up rate, and customer sentiment fluctuations. Finally, based on the goals and importance of the outbound calling business, [the system is further refined]. Each indicator is assigned a weight; for example, connection rate is weighted at 30%, call duration at 25%, hang-up rate at 25%, and customer sentiment fluctuation at 20%. Based on the standardized data and the assigned weights, the score for each indicator is calculated: Indicator Score = Standardized Value × Indicator Weight. This yields the connection rate score, call duration score, hang-up rate score, and sentiment fluctuation score. Finally, the scores for each indicator are summed to obtain the outbound call score. The formula for calculating the outbound call score is: Outbound Call Score = Connection Rate Score + Call Duration Score + Hang-up Rate Score + Customer Sentiment Fluctuation Score.
[0093] It should be noted that the process involves several steps: Speech Recognition: Automatic Speech Recognition (ASR) technology is used to convert the recorded call into text data for subsequent analysis. Semantic Analysis: Semantic analysis is performed on the converted text data to identify the customer's speech content, keywords, tone, etc., to determine the customer's level of interest and needs. Sentiment Analysis: Through voice sentiment analysis technology, the characteristics of the customer's voice tone, speaking speed, volume, etc., in the recorded call are analyzed to identify the customer's emotional state, such as happiness, anger, annoyance, etc.
[0094] It's important to note the following metrics: Connection rate: This reflects the effectiveness of outbound calling and is fundamental for customer contact. A higher connection rate means more opportunities to convey information to customers. It's calculated by the ratio of connected calls to total outbound calls. Call duration: This reflects customer interest and engagement to some extent. Longer calls may indicate a customer's greater willingness to learn about the content. It's calculated by measuring the duration of each call, from start to finish. Hang-up rate: This identifies calls that are hung up by the customer or due to system timeouts. It's calculated by the ratio of hung-up calls to total calls. A low hang-up rate usually indicates that the outbound content and approach are more readily accepted by customers, avoiding premature communication interruptions. Customer emotional fluctuation range: Based on voice sentiment analysis, this assesses the degree of change in customer emotions, such as the range from calm to excitement or from happiness to annoyance. Positive and stable customer emotions are beneficial for business progress. Understanding customer emotions helps evaluate outbound call effectiveness and adjust communication strategies.
[0095] It should be noted that the outbound call score is a comprehensive evaluation indicator that reflects the effectiveness of outbound calls and can be used for subsequent outbound call strategy optimization and user profile adjustment.
[0096] It should be noted that, based on the outbound call score, the thresholds for judging high, medium, and low values in the four dimensions V1-V4 for each user profile are adjusted using the following rules: If the outbound call score is higher than the set value (e.g., 80), it indicates that the current user profile is accurately positioned, and the judgment thresholds for the corresponding dimensions can be appropriately modified according to the data distribution (mean, median) of the user clusters. For example, if the outbound call score is higher than 80, the original V1 threshold is [0.33, 0.67], and the V1 value of this cluster is 0.18, then the threshold is modified to [0.31, 0.67]. If the outbound call score is lower than the set low threshold (e.g., 20), it indicates that the current user profile is inaccurately positioned, and the weights need to be significantly adjusted. A threshold close to the vector is selected for adjustment; if the vector value is greater than the threshold, the threshold is increased, and vice versa. For example, if the original V1 threshold is [0.33, 0.67], and the V1 value of this cluster is 0.31, then the threshold is modified to [0.30, 0.67].
[0097] It should be noted that within a preset time period (such as weekly), the user threshold judgment criteria for the four dimensions of V1-V4 are updated according to the weight adjustment rules. Based on the new thresholds, user profiles are classified and corresponding outbound calling strategies are matched to enter the next round of outbound calling operations, continuously optimizing the accuracy of user profile templates.
[0098] By continuously adjusting the user profile judgment criteria based on user feedback data, the user profile template can be continuously optimized, ensuring its long-term effectiveness and adaptability.
[0099] This invention, through acquiring multimodal data, provides a more comprehensive understanding of user characteristics, thereby improving the accuracy of subsequent user profiling. By constructing four-dimensional feature vectors, different types of data can be transformed into computable feature vectors. Clustering can divide users into several groups, facilitating the development of differentiated strategies. Matching user grouping results with user profile templates allows for quick and accurate mapping of clustering results to specific profiles, avoiding remodeling each time. Inputting user types into the outbound call strategy mapping table enables personalized outbound calls, ensuring accuracy, reducing ineffective communication, improving marketing success rates, avoiding invalid calls to uninterested users, and enhancing customer experience. Continuously adjusting profile judgment criteria based on user feedback data allows for continuous optimization of the user profile template, ensuring its long-term effectiveness and adaptability. Compared to existing technologies, this application improves the accuracy and efficiency of outbound calls.
[0100] like Figure 4 As shown, based on the above method embodiments, corresponding apparatus embodiments are provided;
[0101] One embodiment of the present invention provides an outbound calling system based on multimodal user profile fusion, comprising: an acquisition module 100, a fusion module 200, and an outbound calling module 300;
[0102] The acquisition module 100 is used to acquire the user's multimodal data, wherein the multimodal data includes user information, text data, voice data and image data;
[0103] The fusion module 200 is used to extract features and perform weighted fusion on the multimodal data respectively to construct a four-dimensional feature vector representing the user's basic attributes, consumption behavior, voice interaction and social media. The four-dimensional feature vector is then clustered using the K-means++ algorithm to obtain the user grouping results.
[0104] The outbound call module 300 is used to match the user segmentation results with a preset user profile template to determine the user type, input the user type into a preset outbound call strategy mapping table, match the corresponding outbound call strategy, and make outbound calls according to the outbound call strategy. The threshold judgment criteria of each dimension in the user profile template are dynamically adjusted and optimized based on the outbound call score, and the outbound call score is calculated based on user feedback data during the outbound call process.
[0105] In some embodiments, the fusion module 200 includes an extraction unit and a fusion unit, specifically:
[0106] The extraction unit is used to select the corresponding feature extraction method according to the category of the multimodal data to extract features, and obtain basic attribute feature data, consumer behavior feature data, voice interaction feature data and social media feature data respectively;
[0107] The fusion unit is used to fuse the basic attribute feature data, the consumption behavior feature data, the voice interaction feature data, and the social media feature data to construct an initial four-dimensional feature vector representing the user's basic attributes, consumption behavior, voice interaction, and social media, and to normalize the initial four-dimensional feature vector to obtain the four-dimensional feature vector.
[0108] It is understood that the above-described device embodiments correspond to the method embodiments of the present invention, and can implement the outbound calling method based on multimodal user profile fusion provided by any of the above-described method embodiments of the present invention.
[0109] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can specifically be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0110] Based on the above embodiments of the outbound calling method based on multimodal user profile fusion, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the outbound calling method based on multimodal user profile fusion of any embodiment of the present invention.
[0111] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the terminal device.
[0112] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0113] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting various parts of the terminal device via various interfaces and lines.
[0114] Based on the above-described method embodiments, another embodiment of the present invention provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the outbound calling method based on multimodal user profile fusion as described in any of the above-described method embodiments of the present invention.
[0115] The modules / units integrated in the device / terminal equipment, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0116] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for outbound calling based on multimodal user profile fusion, characterized in that, include: Acquire user's multimodal data, wherein the multimodal data includes user information, text data, voice data, and image data; Feature extraction and weighted fusion are performed on the multimodal data respectively to construct a four-dimensional feature vector representing the user's basic attributes, consumption behavior, voice interaction and social media. The four-dimensional feature vector is then clustered using the K-means++ algorithm to obtain the user grouping results. The user segmentation results are matched with a preset user profile template to determine the user type. The user type is then input into a preset outbound call strategy mapping table to obtain the corresponding outbound call strategy. Outbound calls are then made according to the outbound call strategy. The threshold judgment criteria for each dimension in the user profile template are dynamically adjusted and optimized based on the outbound call score. The outbound call score is calculated based on user feedback data during the outbound call process.
2. The outbound calling method based on multimodal user profile fusion according to claim 1, characterized in that, The acquisition of the user's multimodal data specifically includes: The user's initial multimodal data is obtained, and the missing values of the initial multimodal data are identified. The missing values are filled using the corresponding filling method according to the type of the missing values, and a first processing result is obtained. The outliers in the first processing result are detected using the IQR method or the Z-score method, and the outliers are corrected or deleted to obtain the second processing result. The second processing result is standardized to obtain a third processing result, and the categorical features of the third processing result are one-hot encoded or labeled to obtain the multimodal data.
3. The outbound calling method based on multimodal user profile fusion according to claim 1, characterized in that, The process involves feature extraction and weighted fusion of the multimodal data to construct a four-dimensional feature vector representing the user's basic attributes, consumption behavior, voice interaction, and social media activity. Specifically: Based on the category of the multimodal data, the corresponding feature extraction method is selected to extract features, and basic attribute feature data, consumer behavior feature data, voice interaction feature data and social media feature data are obtained respectively; The basic attribute feature data, the consumption behavior feature data, the voice interaction feature data, and the social media feature data are fused to construct an initial four-dimensional feature vector representing the user's basic attributes, consumption behavior, voice interaction, and social media. The initial four-dimensional feature vector is then normalized to obtain the final four-dimensional feature vector.
4. The outbound calling method based on multimodal user profile fusion according to claim 1, characterized in that, The step involves selecting a corresponding feature extraction method based on the category of the multimodal data to extract features, thereby obtaining basic attribute feature data, consumer behavior feature data, voice interaction feature data, and social media feature data, specifically: The user information is numerically encoded and segmented and normalized to obtain the basic attribute feature data. The text data is classified by consumption type, statistically analyzed by amount frequency, and analyzed by promotion sensitivity to obtain the consumption behavior characteristic data. The acoustic and semantic features of the speech data are extracted to obtain the speech interaction feature data, wherein the acoustic features include speech rate, intonation fluctuation and interaction interruption rate, and the semantic features include emotional tendency and preferred speech keywords; Image recognition is performed on the image data to obtain the user's interest circles, KOL follow relationships, and content interaction characteristics, thus obtaining the social media feature data.
5. The outbound calling method based on multimodal user profile fusion according to claim 1, characterized in that, The user grouping results are obtained by clustering the four-dimensional feature vectors using the K-means++ algorithm, specifically as follows: The elbow rule is used to determine the target number of clusters. The cluster centers are initialized using the K-means++ algorithm. Iterative cluster analysis is performed on the four-dimensional feature vectors to obtain the clustering results until the silhouette coefficient of the clustering results meets the preset threshold. The cluster label of each user is then output to determine the user grouping results.
6. The outbound calling method based on multimodal user profile fusion according to claim 1, characterized in that, The user feedback data includes connection rate, call duration, hang-up rate, and customer emotional fluctuation. The outbound call score is calculated based on the user feedback data during the outbound call process, specifically: Obtain call logs during outbound calls and analyze the call logs to obtain initial user feedback data; The initial user feedback data is standardized to obtain connection rate, call duration, hang-up rate, and customer emotional fluctuation range. The connection rate, call duration, hang-up rate, and customer emotional fluctuation range are calculated according to the calculation methods of each user feedback data to obtain connection rate score, call duration score, hang-up rate score, and customer emotional fluctuation score. The outbound call score is obtained by weighting and summing the connection rate score, call duration score, hang-up rate score, and customer emotion fluctuation score according to preset weights.
7. An outbound calling system based on multimodal user profile fusion, characterized in that, include: Acquisition module, fusion module, and outbound call module; The acquisition module is used to acquire the user's multimodal data, wherein the multimodal data includes user information, text data, voice data, and image data; The fusion module is used to extract features and perform weighted fusion on the multimodal data respectively to construct a four-dimensional feature vector representing the user's basic attributes, consumption behavior, voice interaction and social media. The four-dimensional feature vector is then clustered using the K-means++ algorithm to obtain the user grouping results. The outbound calling module is used to match the user segmentation results with a preset user profile template to determine the user type, input the user type into a preset outbound calling strategy mapping table, match the corresponding outbound calling strategy, and make outbound calls according to the outbound calling strategy. The threshold judgment criteria of each dimension in the user profile template are dynamically adjusted and optimized based on the outbound calling score, and the outbound calling score is calculated based on user feedback data during the outbound calling process.
8. The outbound calling system based on multimodal user profile fusion according to claim 7, characterized in that, The fusion module includes an extraction unit and a fusion unit, specifically: The extraction unit is used to select the corresponding feature extraction method according to the category of the multimodal data to extract features, and obtain basic attribute feature data, consumer behavior feature data, voice interaction feature data and social media feature data respectively; The fusion unit is used to fuse the basic attribute feature data, the consumption behavior feature data, the voice interaction feature data, and the social media feature data to construct an initial four-dimensional feature vector representing the user's basic attributes, consumption behavior, voice interaction, and social media, and to normalize the initial four-dimensional feature vector to obtain the four-dimensional feature vector.
9. A terminal device, characterized in that, include: One or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the outbound calling method based on multimodal user profile fusion as claimed in any one of claims 1-6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the outbound calling method based on multimodal user profile fusion as described in any one of claims 1-6.