International propagation effect accurate evaluation method and system based on large language model
Through a large language model, analyzing user behavior logs, dividing life cycle stages and generating interest tag collections, the problem of dynamic changes in user interests in traditional evaluation methods is solved, and accurate evaluation of international communication effects and differentiated strategy formulation is achieved.
Patent Information
- Application Number
- CN202510707858.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-02
AI Technical Summary
The existing international communication evaluation methods cannot effectively capture the timing changes of user interests, and user behavior and interests are highly dynamic. The traditional user portrait method is based on static tags and lacks modeling and identification of changes in user behavior across life cycles and multiple stages, resulting in push strategies lag behind the actual needs of users.
Through a large language model, the user behavior log is analyzed, the user is divided into the preset life cycle stage, the interest label collection is generated and the weight change rate is calculated, and the matching degree is dynamically corrected by the life cycle stage and interest drift trend characteristics is constructed to achieve accurate evaluation of the propagation effect.
Real-time tracking and refined classification of user interests is realized, the timing evolution laws of interest are accurately captured, differentiated strategies are supported, and the coarse-grained defects of traditional methods are avoided, and the timeliness and accuracy of communication effect evaluation is improved.
Smart Images

Figure CN120579554A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of international communication effect evaluation, and specifically to a method and system for accurately evaluating international communication effects based on a large language model. Background Art
[0002] With the acceleration of global information dissemination and the rise of diverse communication channels such as social media, news portals, and short video platforms, the reach and user response of international communication content have become important indicators of communication quality. In particular, in areas such as national image communication, cross-cultural marketing, and international brand building, accurately and dynamically assessing the actual impact of communication content on users in different regions and over time is a key factor in improving the relevance and effectiveness of communication strategies.
[0003] Existing international communication assessment methods generally have the following shortcomings: user modeling is static and rough, and traditional user profiling methods are mostly based on static labels or long-term average behavioral characteristics, which make it difficult to capture the temporal changes in user interests; the content and interest matching dimension is single, and current matching algorithms mostly rely on keyword overlap or basic semantic relevance; it is impossible to effectively utilize trend information in user behavior trajectories. User behavior and interests are highly dynamic, and existing methods rarely consider the drift trend of user interests, and lack modeling and identification of user behavior change processes across the life cycle and multiple stages.
[0004] In view of the above problems, the existing technology is in urgent need of improvement. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a method and system for accurately evaluating the international communication effect based on a large language model.
[0006] In order to achieve the above object, the technical solution of the present invention is as follows:
[0007] In a first aspect, the present invention discloses a method for accurately evaluating the international communication effect based on a large language model, comprising the following steps:
[0008] Obtaining a user's behavior log data set, and dividing the user into a preset life cycle stage based on the behavior log data set;
[0009] Extracting the user's text interaction content within a preset continuous time window and generating a set of interest tags using a large language model; the set of interest tags includes semantic keywords and corresponding weight coefficients;
[0010] Calculate the weight change rate of the same semantic keywords in adjacent time windows to generate user interest drift trend characteristics;
[0011] Input the dissemination content to be evaluated into the large language model to generate a dissemination content feature vector, and calculate the basic matching degree between the dissemination content feature vector and the interest tag set in the current time window;
[0012] Dynamically modifying the basic matching degree based on the user's life cycle stage and the interest drift trend characteristics to obtain a final matching degree;
[0013] Extract the rate of change of the user's final matching degree between adjacent life cycle stages and the duration of each stage as a path pattern feature vector; perform parameter matching on the path pattern feature vector and a preset propagation path pattern library, and determine the propagation effect evaluation level based on the calculation results of the parameter matching;
[0014] Users are divided into user groups according to the communication effect evaluation level, and intervention strategies are implemented for the user groups.
[0015] In a second aspect, the present invention discloses a precise evaluation system for international communication effects based on a large language model, comprising:
[0016] A user behavior extraction module is used to obtain a user's behavior log data set and classify the user into a preset life cycle stage based on the behavior log data set;
[0017] The interest drift extraction module is used to extract the user's text interaction content within a preset continuous time window and generate a set of interest tags using a large language model; the interest tag set includes semantic keywords and corresponding weight coefficients; the weight change rate of the same semantic keyword in adjacent time windows is calculated to generate user interest drift trend characteristics;
[0018] A basic matching module is used to input the dissemination content to be evaluated into a large language model to generate a dissemination content feature vector, and calculate the basic matching degree between the dissemination content feature vector and the interest tag set in the current time window;
[0019] A matching correction module, configured to dynamically correct the basic matching degree based on the user's life cycle stage and the interest drift trend characteristics to obtain a final matching degree;
[0020] The effect evaluation module is used to extract the change rate of the user's final matching degree between adjacent life cycle stages and the duration of each stage as a path pattern feature vector; perform parameter matching on the path pattern feature vector and a preset propagation path pattern library, and determine the propagation effect evaluation level based on the calculation results of the parameter matching;
[0021] The user intervention module is used to divide users into user groups according to the communication effect evaluation level and implement intervention strategies for the user groups.
[0022] Compared with the prior art, the present invention has the following beneficial effects:
[0023] 1. By analyzing the time series data in user behavior logs and combining multiple thresholds to dynamically determine the user lifecycle stage, this solves the lag problem of traditional static user profiling and enables real-time tracking and refined classification of user status;
[0024] 2. Utilize a large language model to extract semantic keywords and weight coefficients from user text interactions, calculate the weight change rate within adjacent time windows, and generate interest drift trend features to accurately capture the temporal evolution of user interests, avoiding the confusion between short-term fluctuations and long-term trends that traditional methods often confuse.
[0025] 3. By extracting the change rate and duration of user matching in adjacent lifecycle stages, we construct a path pattern feature vector and perform parameter matching with the historical pattern library to identify the typical propagation path of user behavior, providing a multi-dimensional quantitative basis for evaluating the propagation effect.
[0026] 4. Use multiple thresholds to divide the communication effect level, combined with a dynamic parameter matching mechanism, to achieve refined hierarchical evaluation of the communication effect, support differentiated strategy formulation, and avoid the coarse-grained defects of traditional binary classification methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 This is an overall block diagram of a method according to a first embodiment of the present invention;
[0029] Figure 2 This is a flow chart of a method according to embodiment 1 of the present invention;
[0030] Figure 3 This is an overall block diagram of the system according to the second embodiment of the present invention. DETAILED DESCRIPTION
[0031] The technical solutions of the present invention will be described clearly and completely below with reference to the embodiments. It is obvious that the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.
[0032] Application Overview: In existing technologies, the evaluation of international communication effectiveness mainly relies on static indicators such as click-through rate and forwarding volume, which are difficult to reflect the dynamic evolution of user interests. Traditional user portrait methods are based on long-term average behavioral data and cannot capture the behavioral differences of users at different stages of their life cycle. Communication content matching algorithms generally use keyword matching or basic semantic similarity calculations, and lack the mining of deep semantic associations in texts. When a multinational news platform evaluated the communication effect of climate issue reports on users in Southeast Asia, it was found that user click behavior showed periodic fluctuations over time, but the existing evaluation model could not explain the dynamic correlation between changes in interest and communication content, resulting in push strategy adjustments lagging behind actual user needs.
[0033] In order to solve the above problems, the inventors noticed that user interests are time-sensitive and have stage-evolution characteristics, and it is necessary to build a dynamic user model. By analyzing the time series characteristics in user behavior logs, it was found that there is a correlation between user activity and registration duration, and it was proposed to divide users into different life cycle stages to reflect their state migration. In order to solve the problem of semantic parsing of text interaction content, the deep semantic understanding ability of the large language model is used to generate fine-grained interest tags. In order to solve the impact of interest drift on matching, the interest evolution trend is quantified by calculating the weight change rate of adjacent time windows. Finally, a dynamic evaluation framework based on the linkage of life cycle stage correction and trend characteristics is formed to achieve accurate adaptation of dissemination content and user status.
[0034] Example 1:
[0035] like Figure 1-2 As shown, a method for accurately evaluating international communication effects based on a large language model includes the following steps: obtaining a user's behavior log dataset, and dividing the user into preset life cycle stages based on the behavior log dataset; extracting the user's text interaction content in a preset continuous time window, and generating a set of interest tags through a large language model; calculating the weight change rate of the same semantic keywords in adjacent time windows to generate user interest drift trend characteristics; inputting the communication content to be evaluated into the large language model to generate a communication content feature vector, and calculating the basic matching degree between the communication content feature vector and the interest tag set in the current time window; dynamically correcting the basic matching degree based on the user's life cycle stage and the interest drift trend characteristics to obtain the final matching degree; extracting the change rate of the user's final matching degree between adjacent life cycle stages and the stage duration as a path pattern feature vector; performing parameter matching on the path pattern feature vector and a preset communication path pattern library, and determining the communication effect evaluation level based on the calculation result of the parameter matching; dividing the user into user groups based on the communication effect evaluation level, and implementing intervention strategies for the user groups.
[0036] The present application further proposes that the behavior log data set includes the first interaction timestamp, the active days sequence and the recent interaction frequency; the life cycle stages include the novice period, the growth period, the loyalty period, the drift period and the latent period; the life cycle stage division specifically includes: determining the user registration base date based on the first interaction timestamp, constructing a time coordinate system with natural weeks as units, and calculating the time span ΔT between the current time and the user registration base date; when ΔT is less than or equal to the preset novice period, it is determined to be the novice period; when ΔT is greater than the preset novice period and the number of active days is increasing, it is determined to be the growth period; calculating the proportion of active days for four consecutive weeks based on the user's active days sequence, and determining it to be the loyalty period when it exceeds the first threshold; calculating the average interaction change rate between weeks based on the recent interaction frequency, and determining it to be the drift period when the average interaction change rate exceeds the second threshold and shows a continuous downward trend; determining it to be the latent period when the user is in the loyalty period and the number of consecutive days without interaction exceeds the third threshold; determining it to be the latent period when the number of consecutive days without interaction of the user in the drift period exceeds the fourth threshold.
[0037] The first interaction timestamp refers to the date and time when a user first interacts with the content. This can be achieved by recording the time of the user's first click, view, or comment. It is used to determine the user's registration base date and calculate the starting point of their lifecycle on the platform. The active days sequence refers to the set of active days recorded by a user within a preset statistical period. Specifically, it can be calculated by whether the number of daily logins or interactions exceeds a threshold. This is used to analyze the persistence and volatility of user activity trends. The recent interaction frequency refers to the number of user interactions within a recently set time range. This can be achieved by calculating the number of clicks or comments per unit time. This is used to detect real-time changes in user activity. The novice period, growth period, loyalty period, drift period, and latent period are different stages in the user lifecycle. They can be divided into different stages through a dynamic combination of time span and behavioral indicators to reflect the entire process from initial contact to waning interest. The time span ΔT refers to the difference between the current time and the time of the user's first interaction. It can be calculated by aligning the time in calendar weeks and used to quantify the length of the user's lifecycle on the platform. The preset novice period refers to the maximum duration of the novice stage set according to business needs. For example, it can be set to the first two weeks after registration to distinguish the behavioral characteristics of new users from mature users. The increasing trend of active days means that the number of active days of users shows a gradual increase in consecutive statistical periods. Specifically, it can be judged by calculating whether the difference between the active days of adjacent periods is positive, which is used to identify users in the interest growth stage. The number of consecutive days without interaction refers to the number of consecutive natural days in which the user has not generated any interactive behavior. Specifically, it can be counted by monitoring the difference between the last interaction time and the current time in the user behavior log to determine whether the user has entered the interest fading stage.
[0038] Specifically, the lifecycle stage classification process achieves refined classification through a dynamic combination of time span and behavioral indicators. For example, the user registration base date is based on the first interaction timestamp. By calculating the time span ΔT between the current time and the base date, combined with the preset onboarding period, the user is determined to be in the novice phase. For users beyond the novice phase, the increasing trend of their active days sequence is further analyzed to determine whether they have entered the growth phase. In the post-growth phase, if the proportion of active days for four consecutive weeks exceeds the first threshold, the user has established stable usage habits and is considered to be in the loyalty phase. When the average rate of change in recent interaction frequency exceeds the second threshold and shows a continuous downward trend, it indicates that user interest is waning and the user enters the drift phase. If the number of consecutive days of no interaction for users in the loyalty or drift phase exceeds the third or fourth threshold, respectively, the user is considered to be in the latent phase, indicating that the user is on the verge of churn. The entire process dynamically tracks and categorizes user status through time series analysis of multi-dimensional indicators.
[0039] Traditional user segmentation methods are typically based on static labels or behavioral data from a single point in time. For example, segmentation based solely on registration duration or cumulative active days fails to reflect the dynamic evolution of user interests. However, this solution uses a comprehensive calculation of multi-dimensional time series data, such as the first interaction timestamp, active days sequence, and recent interaction frequency, combined with dynamic comparison of preset thresholds, to more accurately capture the state transition patterns of users throughout their life cycle, from the initiation of interest to the fading of interest. For example, by analyzing the increasing trend of the active days sequence rather than the total amount, it is possible to effectively distinguish between users in the growth stage and the mature stage; by correlating the number of consecutive days without interaction with thresholds at different stages, it is possible to avoid misjudging the short period of silence of loyal users as churn.
[0040] Through the above technical solution, this application can adjust the life cycle stage division results in real time according to the dynamic changes of user behavior data, solving the problem of misjudgment of user status caused by traditional methods relying on static indicators. For example, when a user transitions from the growth stage to the loyalty stage, the system can timely update the stage label by comparing the threshold of the proportion of active days for four consecutive weeks, providing an accurate basis for user status for subsequent content matching. At the same time, by designing the judgment conditions of the drift period and the latent period, it can trigger an early warning mechanism when the user's interest begins to wane but has not yet completely disappeared, thereby gaining a time window for implementing recovery strategies.
[0041] This application further proposes a method for generating interest drift trend features: defining a time window length T, extracting text interaction content within the time window [tT, t]; performing semantic analysis on the text interaction content through a large language model, and outputting a set of semantic keywords and corresponding weight coefficients; obtaining a weight change rate by calculating the ratio of the difference in weight coefficients of the same semantic keywords in adjacent time windows to the time window length T; normalizing the weight change rate, and retaining feature items whose absolute values exceed a preset sensitivity threshold to form an interest drift trend feature vector.
[0042] The time window length T refers to the time interval used to analyze user behavior. This can be implemented using a fixed period or dynamically adjusted, for example, 7 or 30 days to meet the data collection requirements of different communication scenarios. This parameter balances the ability to capture short-term fluctuations with the ability to capture long-term trends. The semantic keyword set and corresponding weight coefficients refer to the results of topic extraction and importance quantification of user-generated text content using a large language model. Specifically, an attention mechanism can be used to calculate keyword weights, reflecting the user's core focus during a specific time period. The weight change rate is calculated by comparing the weight difference of the same keyword in adjacent time windows with the time span. This can be achieved using differential calculation combined with a time decay function to quantify the rate of evolution of user interests. Normalization is the process of mapping the weight change rate to a uniform numerical range. Minimum and maximum normalization or Z-score normalization can be used to eliminate the impact of dimensional differences on feature analysis. The preset sensitivity threshold is the critical value used to screen for significant changes in interest. It can be dynamically adjusted based on historical data distribution or business needs, for example, setting it to twice the standard deviation to filter out noise.
[0043] Specifically, in the time window division stage, the current time point t is used as the benchmark, and a time interval of length T is intercepted forward as the analysis range, such as selecting text interaction data for the past four weeks. The text in the window is semantically parsed through a large language model, and a set of weighted keywords is output. For example, keywords such as "cultural differences" and "brand value" in user comments are given a weight coefficient of 0.2-0.8. When comparing data from adjacent windows, the rate of change of the weight difference relative to the time span is calculated for repeated keywords. For example, the weight of a keyword increases from 0.4 to 0.6 within two weeks, and the rate of change is (0.6-0.4) / 14=0.0143 / day. After normalization, feature items whose absolute values exceed the threshold are retained, such as keywords with a change rate exceeding 0.01, to form a feature vector reflecting the direction and intensity of user interest migration.
[0044] Traditional methods typically only count keyword occurrences or calculate simple growth rates, failing to effectively distinguish natural fluctuations from genuine interest shifts. This solution, by introducing a dynamic time window partitioning mechanism and combining it with the deep semantic parsing capabilities of a large language model, accurately captures the temporal evolution of user interests. Furthermore, by correlating the weight change rate with the time span, and combining normalization with threshold filtering, the sensitivity and reliability of interest drift detection are effectively improved.
[0045] Through the above technical solution, this application can accurately identify the dynamic migration trajectory of user interests, providing a quantitative basis for the subsequent correction of the matching degree of dissemination content. By eliminating the impact of time span differences on weight changes, misjudgment caused by inconsistent analysis periods is avoided. At the same time, a sensitive threshold screening mechanism is used to reduce data noise interference, ensuring that the extracted drift trend characteristics have clear business interpretation and operability.
[0046] This application further proposes that the calculation process of the basic matching degree includes: the feature vector of the disseminated content includes an implicit topic distribution vector and an attribute tag vector; the disseminated content text is input into a large language model to extract the implicit topic distribution vector; the attribute tag vector is generated through a predefined content attribute classifier; the basic matching degree is calculated: Match_base = αcos(V_theme, V_keyword) + βcos(V_attr, V_keyword), where α+β=1, V_keyword is a vector composed of weight coefficients of semantic keywords in the interest tag set, V_theme is the implicit topic distribution vector, and V_attr is the attribute tag vector.
[0047] The implicit topic distribution vector refers to a multidimensional vector representation formed after semantic understanding of the dissemination content through a large language model. Specifically, it can be implemented using the output vector of the semantic encoding layer of the large language model, which can reflect the implicit topic distribution characteristics in the text. The attribute label vector refers to a vector that encodes structured attributes such as the type, field, and emotional tendency of the dissemination content. Specifically, it can be generated by a pre-trained multi-label classification model. For example, after the dissemination content is input into the classifier, a vector containing labels such as politics, economy, and culture is output. α and β in the basic matching degree calculation are adjustable weighting coefficients. For example, α can be 0.6 and β can be 0.4. They are used to balance the contribution of topic relevance and attribute matching in the overall matching degree. Cosine similarity calculation is used to measure the similarity between the dissemination content and the user's interest tags in the vector space. Specifically, it can be implemented using the standardized vector dot product method.
[0048] Specifically, the content is first fed into a large language model for semantic parsing. The model's deep semantic encoding module extracts a latent topic distribution vector (TDV). This TDV captures the underlying semantic themes within the text. For example, when analyzing international news, it might identify thematic dimensions such as "geopolitics" and "economic cooperation." Simultaneously, the content is fed into a pre-trained content attribute classifier, generating a structured attribute tag vector encompassing domain classification, sentiment polarity, and content format. For example, short videos might be tagged with attributes such as "entertainment," "positive sentiment," and "dynamic imagery." Subsequently, the semantic keywords and their weight coefficients within the user's current time window are converted into vectors and subjected to cosine similarity calculations with the content's topic vector and attribute vector. Finally, a weighted summation is performed to generate a basic matching metric. The weight coefficients can be configured based on the application scenario. For example, in the context of international brand communication, the attribute matching weight coefficient β can be increased to enhance the adaptability of the content format to user preferences.
[0049] Existing technologies typically rely solely on keyword matching or single-dimensional semantic similarity calculations, such as relying solely on TF-IDF keyword overlap or basic topic model matching. This solution, by integrating a dual matching mechanism that uses implicit topic distribution vectors and attribute tag vectors, can simultaneously capture the deep semantic features and structured attribute characteristics of the disseminated content. Furthermore, the dynamic adjustment of weighting coefficients allows for flexible optimization of matching calculations based on dissemination goals. For example, in scenarios where content timeliness needs to be enhanced, the weight coefficient α of the topic vector can be increased to quickly respond to changes in user interests.
[0050] Through the above technical solution, this application can solve the problem of misjudgment caused by the single dimension of traditional matching methods. For example, when the dissemination content differs from the user's interests in the topic dimension, but is highly consistent in the attribute dimension, the compensation effect of attribute matching can prevent high-quality content from being incorrectly filtered. At the same time, the dual vector matching mechanism can reduce the error caused by single feature noise. For example, when there are interfering keywords in the user's interest tags, the semantic aggregation characteristics of the topic vector can effectively suppress the influence of noise.
[0051] The present application further proposes that the process of dynamically correcting the basic matching value to obtain the final matching degree includes: establishing a mapping table of life cycle stages and correction coefficients; determining the trend factor δ based on the interest drift trend characteristics: when the weight coefficient of the semantic keywords related to the dissemination content in the user interest tag set shows an increasing trend, it is judged as drifting in the same direction; when it shows a decaying trend, it is judged as drifting in the opposite direction; calculating the final matching degree M = M_base×γ×δ, where γ is the correction coefficient corresponding to the current life cycle stage.
[0052] The lifecycle stage refers to the different stages of user behavior data. Specifically, registration time, active days, and interaction frequency can be used as classification indicators to categorize users into different stages, such as the novice, growth, and loyalty stages. The correction coefficient refers to an adjustment parameter associated with the user's lifecycle stage. This can be implemented using a preset numerical mapping table. For example, the correction coefficient for growth-stage users is higher than that for latent-stage users. This coefficient reflects the differences in user responses to disseminated content at different stages. The interest drift trend feature is a quantitative indicator that reflects the direction of user interest changes. Specifically, it can be generated by calculating the normalized value of the rate of change of semantic keyword weights within adjacent time windows to capture the migration direction of user interests. The trend factor δ is a dynamically adjusted parameter generated based on the direction of interest drift. Specifically, it can be implemented using a piecewise function or threshold judgment mechanism. When an increase in interest related to the disseminated content is detected, δ is used to amplify the basic matching degree; otherwise, the matching degree is suppressed.
[0053] Specifically, when the user is in the growth stage and the weight of keywords related to the target dissemination content in the interest tag continues to rise, the system obtains the correction coefficient γ corresponding to this stage from the mapping table and calculates δ = 1.2 based on the interest drift trend characteristics. Multiplying the basic matching degree M_base = 0.8 by these two coefficients, we get the final matching degree M = 0.8 × 1.1 × 1.2 = 1.056. This value exceeds the baseline threshold, indicating that the current dissemination content is highly consistent with the user's dynamic interests. When the user enters the latent period and the weight of the relevant interest tag continues to decline, the system automatically adjusts δ to 0.8. At this time, even if the basic matching degree remains the same, the final matching degree will be corrected to 0.8 × 0.9 × 0.8 = 0.576, which is lower than the response threshold. This dynamic adjustment mechanism enables the evaluation results to reflect the dual influence of the user's life cycle status and interest evolution trend in real time.
[0054] Traditional matching calculation methods rely solely on static user profiles and content features, failing to distinguish between the responses of novice and loyal users and failing to consider the temporal dimension of interest changes. For example, when dealing with users in the waning interest phase, existing methods may still maintain a high matching score based on historical data, rendering push strategies ineffective. This solution, however, addresses the lack of adaptability of a single matching model in dynamic scenarios by introducing a dual adjustment mechanism: a lifecycle correction coefficient and a trend factor. This ensures that the evaluation results are more accurately aligned with the user's actual state.
[0055] Through the above technical solution, this application achieves dynamic optimization of the matching degree of dissemination content, which is specifically manifested in three aspects: first, the difference in response sensitivity of users in different states to the same content is distinguished through the correction coefficient of the life cycle stage; second, the trend factor is used to capture the real-time change direction of user interests to avoid evaluation bias caused by interest drift; third, static semantic matching is combined with dynamic behavioral characteristics to improve the timeliness and accuracy of dissemination effect prediction. For example, when user interests drift in the same direction, it can identify potential high-response groups in advance, providing a decision basis for precise push.
[0056] The present application further proposes a method for constructing a path pattern feature vector, including extracting a migration sequence of adjacent life cycle stages, wherein the migration sequence is composed of multiple continuous life cycle stages; calculating a stage duration vector, which is composed of the number of days of each stage; extracting a final matching degree change rate vector, which is composed of the matching degree difference between adjacent stages divided by the stage duration; and normalizing the stage duration vector and the change rate vector and then splicing them to generate a path pattern feature vector.
[0057] Among them, the life cycle stage transition sequence refers to the behavioral state transition path of users in multiple consecutive time periods. Specifically, it can be achieved by recording the user's activity level, interaction frequency and other behavioral indicators in different time units to divide the stage types, which is used to characterize the evolution of user behavior. The stage duration vector refers to the quantified value of the span of each stage in the time dimension. Specifically, it can be achieved by using natural weeks or natural days as the time unit to count the duration, which is used to reflect the stability of users in different behavioral states. The final matching degree change rate vector refers to the rate of change of content matching between adjacent stages. Specifically, it can be achieved by using differential calculation combined with time window length for normalization processing, which is used to capture the dynamic correlation between user interest migration and content reach effect. Normalized splicing refers to scaling vectors of different dimensions to a unified interval through linear transformation and then merging the dimensions. Specifically, it can be achieved by using the minimum-maximum normalization method to process the data and then connecting the vector elements in sequence, which is used to eliminate the impact of data scale differences on pattern matching.
[0058] Specifically, the user's migration process between different lifecycle stages is modeled as an ordered sequence of stages, and the duration of each stage is quantified as a numerical vector of time spans. By calculating the ratio of the relative change in content matching between adjacent stages to the time span, an indicator vector reflecting the rate of change in matching efficiency is obtained. The two indicators, time span and change rate, are standardized and combined into a multidimensional feature vector, so that the spatiotemporal characteristics of the user behavior path and the content response effect can be uniformly represented in a machine-processable numerical form. This feature vector, as an input parameter, can be used to calculate similarity with a preset propagation path pattern library, thereby identifying the typical behavior pattern category to which the user belongs.
[0059] Compared to existing technologies, traditional methods typically only count the absolute value of matching within a single phase or analyze the duration of an isolated phase, failing to consider the dynamic nature of matching efficiency across inter-phase migration paths. This solution, by constructing a composite feature vector that integrates time span and matching change rate, can simultaneously capture the temporal patterns of user state transitions and the evolving trends of content reach. This addresses the shortcomings of existing technologies, which model user behavior trajectories in a single dimension and fail to reflect inter-phase correlations.
[0060] Through the above technical solution, this application achieves quantitative modeling of user behavior paths across their lifecycle, effectively improving the accuracy of identifying long-term behavioral patterns in communication effectiveness evaluation. By jointly encoding the time dimension and the effect change dimension, it is possible to accurately distinguish the behavioral evolution characteristics of different user groups, providing reliable data support for the development of differentiated communication intervention strategies.
[0061] This application further proposes a parameter matching process. The propagation path pattern library is dynamically constructed based on the migration patterns of historical user groups between life cycle stages and the final matching degree. Each pattern in the pattern library contains quantitative parameters: the shortest threshold value of the duration between stages, the slope range of the matching index change, and the associated weights of adjacent stages. Calculate the duration matching degree, which is measured by the difference between the duration of the current stage and the shortest threshold. Calculate the slope matching degree, and count the proportion of the final matching degree change rate that falls within the preset slope range. Calculate the comprehensive matching degree, combine the duration matching degree and the slope matching degree, and perform weighted summation according to the associated weights. The pattern corresponding to the comprehensive matching degree with the highest value is used as the classification result of the user's current behavior.
[0062] Among them, the propagation path pattern library refers to a data set that stores the migration patterns of historical user groups between life cycle stages. It can be implemented using a database table structure to store typical patterns of duration and matching changes between different stages. The minimum threshold for duration between stages refers to the minimum time span required for migration to a specific stage. It can be determined by statistically analyzing the distribution of durations of corresponding stages in historical data, and is used to screen migration behaviors that conform to time patterns. The slope range of the matching index change refers to the upper and lower limits of the allowable matching change rate. It can be determined by grouping the historical user matching change rates through a clustering algorithm, and is used to identify matching fluctuations that conform to expected trends. The adjacent stage association weight refers to the relative importance parameter of the time factor and the slope factor in the migration patterns of different stages. It can be obtained through expert experience or machine learning model training, and is used to adjust the contribution of different dimensional features to the classification results.
[0063] Specifically, the parameter matching process first extracts the quantitative parameters of each mode from the propagation path pattern library. For the current user's stage migration sequence, the difference between its stage duration vector and the shortest threshold of each mode in the pattern library is calculated and converted into a duration matching degree. At the same time, the proportion of entries in the user's final matching rate change vector that meet the preset slope range of the pattern library is counted as the slope matching degree. The two matching degrees are weighted and summed according to the preset association weights in the pattern library to obtain a comprehensive matching degree. By comparing the comprehensive matching degree values of all modes, the mode corresponding to the highest value is selected as the basis for classifying the user behavior, thereby determining the type of propagation path to which it belongs.
[0064] Compared to existing technologies, traditional methods typically classify paths based on a single metric or fixed rules, such as considering only whether a stage duration meets a target or whether the change in matching exceeds an absolute threshold. Such methods struggle to adapt to the dynamic correlation between time and trend characteristics in migration patterns across different stages, and can easily lead to classification bias. This solution, by constructing a propagation path pattern library containing multi-dimensional quantitative parameters and introducing a weighted comprehensive matching calculation mechanism, can more accurately capture the similarity between user behavior trajectories and typical patterns, effectively improving the accuracy and interpretability of classification results.
[0065] Through the above-mentioned technical solution, this application solves the technical problems of the existing technology in the single-dimension and rigid rules for user propagation path classification. By dynamically constructing a pattern library that includes both time and trend dimensions and adopting a weighted comprehensive matching mechanism, multi-feature collaborative matching of user behavior migration paths is achieved. This matching method can more accurately identify the type of propagation stage a user is in, providing a reliable basis for subsequent evaluation and grading, and thus optimizing the formulation of intervention strategies for different user groups.
[0066] This application further proposes that the process of determining the communication effect evaluation level includes: setting several evaluation level thresholds θ_1, θ_2,..., θ_n, comparing the comprehensive matching degree S with the evaluation level threshold; when S∈[θk,θ{k+1}), it is determined to be the kth level communication effect.
[0067] The evaluation level threshold refers to the pre-set numerical interval boundary used to distinguish different levels of communication effects. Specifically, it can be determined by the percentiles of the comprehensive matching degree distribution corresponding to different communication effects in historical user groups. For example, the first 20% is set to θ_1, the middle 50% is set to θ_2, and the last 30% is set to θ_3. A threshold division standard is established through statistical methods to solve the problem of coarse granularity of single threshold division in traditional methods. Among them, the comprehensive matching degree refers to a quantitative indicator of the degree of match between the user behavior path and the preset communication path pattern. Specifically, it can be obtained through a weighted calculation of time matching degree and slope matching degree, for example, using the formula S = W × S_time + (1-W) × S_slope. The reliability of the evaluation results is improved through the fusion of multi-dimensional parameters.
[0068] Specifically, the determination of the communication effect evaluation level can be achieved based on the following process: first, a judgment framework containing multiple evaluation level thresholds is constructed, for example, setting θ_1 = 0.7, θ_2 = 0.5, and θ_3 = 0.3, and storing the thresholds in the system configuration module; second, the user's current comprehensive matching degree S is calculated. For example, when S = 0.65, it is determined that the user belongs to the second-level communication effect corresponding to the interval θ_2 to θ_1; further, the system can trigger differentiated logging strategies based on different evaluation levels, such as automatically generating communication effect analysis reports for high-level users. During the threshold update phase, a sliding window mechanism can be used to dynamically adjust the thresholds, for example, recalculating the percentile thresholds every quarter based on the comprehensive matching degree distribution of the latest user group.
[0069] Compared to existing technologies, traditional methods typically use a single threshold for binary evaluation, distinguishing only between effective and ineffective communication, failing to reflect the gradient of communication effects. This solution, however, establishes a multi-level evaluation system using multiple thresholds, categorizing communication effects into high, medium, and low levels. Combined with a dynamic threshold adjustment mechanism, this more accurately reflects the degree of match between user behavior paths and target communication patterns.
[0070] Through the above technical solution, this application realizes the refined division of communication effect evaluation levels, so that user behavior paths with different matching degrees obtain corresponding quantitative ratings, providing accurate classification basis for subsequent differentiated intervention strategies, such as prioritizing the allocation of high-quality communication resources for high-level users and launching a compensation push mechanism for low-level users, effectively improving the pertinence of communication strategies and resource allocation efficiency.
[0071] This application further proposes a method for implementing an intervention strategy for user groups, including dividing users into a strong response group, a potential conversion group, and an inefficient reach group according to the level of communication effect evaluation; increasing the push frequency for the strong response group; and pushing compensatory content to the potential conversion group and the inefficient reach group. The specific process is: extracting a set of semantic keywords with a positively increasing weight change rate in the current interest drift trend characteristics, screening candidate push content from the content library based on the semantic keyword set, calculating the similarity between the candidate push content and the user's interest tag set, and selecting the top N with the largest similarity for targeted push.
[0072] Among them, the communication effect evaluation level refers to the degree of fit between the user behavior calculated by parameter matching and the preset communication path pattern. Specifically, the threshold interval division method can be used to map the numerical results into discrete level labels, which are used to quantitatively evaluate the actual effect of the communication strategy. The interest drift trend feature refers to a dynamic indicator that reflects the change of user interests over time. Specifically, it can be generated through statistical analysis of the change rate of semantic keyword weights in adjacent time windows, and is used to identify the direction of user interest migration. The semantic keyword set refers to the core semantic units and their weight coefficients extracted from the text interaction content through a large language model. Specifically, it can be implemented by word embedding clustering and attention weight allocation algorithm to characterize the user's current focus. Similarity calculation refers to the measurement of the degree of match between candidate content and user interest tags. Specifically, it can be implemented by using the cosine similarity algorithm combined with multi-dimensional feature vector weighted calculation to screen the push content that is most relevant to the user's interest trend.
[0073] Specifically, user group divisions are dynamically adjusted based on the level of communication effect evaluation. The strong response group represents a user group with a high content match and continuous activity; the potential conversion group represents a user group with fluctuating match but positive interest trends; and the inefficient reach group represents a user group with a persistently low match and declining activity. Differentiated intervention strategies are adopted for different groups: the push frequency is increased for the strong response group to enhance communication effects; a compensatory push mechanism is activated for the potential conversion group and the inefficient reach group. Dynamic screening conditions are constructed by extracting keywords with increasing weights in interest drift trends, retrieving candidate content from the content library, calculating the vector similarity between the candidate content and the user's interest tags, and selecting the top N similarity-ranked content for targeted push. This process combines real-time interest trend analysis with historical content library retrieval to achieve a precise content compensation strategy.
[0074] In some specific implementations, the content library can be constructed using semantic indexing technology based on a large language model. The similarity calculation for candidate content can incorporate weighted matching results of topic distribution vectors and attribute tag vectors. Push frequency adjustment can set an upper threshold based on the user's current lifecycle stage to avoid a degradation of user experience caused by excessive pushes.
[0075] Compared to existing technologies, existing user segmentation methods typically segment users based on static behavioral indicators and lack dynamic awareness of interest migration trends, causing push strategies to lag behind users' actual needs. This method, by combining communication effect evaluation levels with interest drift trend characteristics, can identify user status changes in real time and filter compensatory content based on the weight changes of semantic keyword sets, ensuring that push strategies align with user interest evolution.
[0076] Through the above technical solution, this application can dynamically adjust user grouping criteria based on communication effects, implement differentiated intervention strategies for different groups, effectively improve the efficiency of content reach for potential conversion groups, and reduce resource waste in inefficient reach groups. At the same time, the semantic keyword screening mechanism based on interest drift trends can accurately capture the direction of user interest migration and ensure that the compensatory push content is highly matched with the user's current focus.
[0077] Example 2:
[0078] like Figure 3 As shown in the figure, a precise evaluation system for international communication effects based on a large language model includes:
[0079] A user behavior extraction module is used to obtain a user's behavior log data set and classify the user into a preset life cycle stage based on the behavior log data set;
[0080] The interest drift extraction module is used to extract the user's text interaction content within a preset continuous time window and generate a set of interest tags using a large language model; the interest tag set includes semantic keywords and corresponding weight coefficients; the weight change rate of the same semantic keyword in adjacent time windows is calculated to generate user interest drift trend characteristics;
[0081] A basic matching module is used to input the dissemination content to be evaluated into a large language model to generate a dissemination content feature vector, and calculate the basic matching degree between the dissemination content feature vector and the interest tag set in the current time window;
[0082] A matching correction module, configured to dynamically correct the basic matching degree based on the user's life cycle stage and the interest drift trend characteristics to obtain a final matching degree;
[0083] The effect evaluation module is used to extract the change rate of the user's final matching degree between adjacent life cycle stages and the duration of each stage as a path pattern feature vector; perform parameter matching on the path pattern feature vector and a preset propagation path pattern library, and determine the propagation effect evaluation level based on the calculation results of the parameter matching;
[0084] The user intervention module is used to divide users into user groups according to the communication effect evaluation level and implement intervention strategies for the user groups.
[0085] The above content is merely an example and explanation of the structure of the present invention. Those skilled in the art may make various modifications or additions to the described specific embodiments or replace them in a similar manner. As long as they do not deviate from the structure of the invention or exceed the scope defined by the claims, they should all fall within the scope of protection of the present invention.
[0086] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0087] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to specific embodiments. Obviously, many modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for accurately evaluating the effect of international communication based on a large language model, characterized by: The following steps are involved: Obtaining a user's behavior log data set, and dividing the user into a preset life cycle stage based on the behavior log data set; Extracting the user's text interaction content within a preset continuous time window and generating a set of interest tags using a large language model; the set of interest tags includes semantic keywords and corresponding weight coefficients; Calculate the weight change rate of the same semantic keywords in adjacent time windows, and then generate user interest drift trend characteristics; Input the dissemination content to be evaluated into the large language model to generate a dissemination content feature vector, and calculate the basic matching degree between the dissemination content feature vector and the interest tag set in the current time window; Dynamically modifying the basic matching degree based on the user's life cycle stage and the interest drift trend characteristics to obtain a final matching degree; Extract the change rate of the user's final matching degree between adjacent life cycle stages and the duration of each stage as the path pattern feature vector; Performing parameter matching on the path pattern feature vector and a preset propagation path pattern library, and determining a propagation effect evaluation level based on a calculation result of the parameter matching; Users are divided into user groups according to the communication effect evaluation level, and intervention strategies are implemented for the user groups.
2. The method for accurately evaluating international communication effects based on a large language model according to claim 1, characterized in that: The behavior log dataset includes the first interaction timestamp, active days sequence, and recent interaction frequency; The life cycle stages include the novice stage, the growth stage, the loyalty stage, the drift stage and the latent stage; The life cycle stage division specifically includes: determining the user registration base date based on the first interaction timestamp, constructing a time coordinate system with natural weeks as units, and calculating the time span ΔT between the current time and the user registration base date; When ΔT is less than or equal to the preset novice period, it is determined to be the novice period; When ΔT is greater than the preset novice period and the number of active days is increasing, it is determined to be in the growth stage; Calculate the percentage of active days for four consecutive weeks based on the user's active days sequence. If the percentage exceeds the first threshold, it is considered a loyalty period. Calculate the average interaction change rate between weeks based on the recent interaction frequency. When the average interaction change rate exceeds the second threshold and shows a continuous downward trend, it is determined to be a drift period. When a user is in the loyalty period and the number of consecutive days without interaction exceeds the third threshold, it is determined to be in the latent period; when a user in the drift period has no consecutive days without interaction exceeding the fourth threshold, it is determined to be in the latent period.
3. The method for accurately evaluating international communication effects based on a large language model according to claim 2, characterized in that: The method for generating the interest drift trend feature is as follows: Define the time window length T and extract the text interaction content within the time window [tT,t]; Perform semantic analysis on text interaction content through a large language model, and output a set of semantic keywords and corresponding weight coefficients; The weight change rate is obtained by calculating the ratio of the weight coefficient difference of the same semantic keyword in adjacent time windows to the time window length T; The weight change rate is normalized, and the feature items whose absolute values exceed the preset sensitivity threshold are retained to form the interest drift trend feature vector.
4. The method for accurately evaluating international communication effects based on a large language model according to claim 3 is characterized by: The calculation process of the basic matching degree includes: The propagation content feature vector includes an implicit topic distribution vector and an attribute tag vector; Input the dissemination content text into the large language model to extract the implicit topic distribution vector; Generate an attribute tag vector through a predefined content attribute classifier; Calculate the basic matching degree: Match_base = α*cos(V_theme, V_keyword) + β*cos(V_attr, V_keyword), where α+β=1, V_keyword is the vector composed of the weight coefficients of the semantic keywords in the interest tag set, V_theme is the implicit topic distribution vector, and V_attr is the attribute tag vector.
5. The method for accurately evaluating international communication effects based on a large language model according to claim 4 is characterized by: The process of dynamically correcting the basic matching value to obtain the final matching value includes: Establish a mapping table between life cycle stages and correction factors; Determine the trend factor δ based on the interest drift trend characteristics: When the weight coefficient of the semantic keywords related to the dissemination content in the user interest tag set shows an increasing trend, it is determined to be drifting in the same direction, δ>1; when it shows a decreasing trend, it is determined to be drifting in the opposite direction, δ<1; Calculate the final matching degree M = M_base × γ × δ, where γ is the correction coefficient corresponding to the current life cycle stage.
6. The method for accurately evaluating international communication effects based on a large language model according to claim 5, characterized in that: The process of constructing the path pattern feature vector includes: Extract the adjacent life cycle stage migration sequence S = {s_1→s_2→...→s_n}, where s_n represents the life cycle stage; Calculate the phase duration vector D = (d_1, d_2, ..., d_{n-1}), where d_{n-1} represents the duration of the life cycle phase s_{n-1}; Extract the final matching rate change vector R = (r_1, r_2, ..., r_{n-1}), where r_{n-1} = (M_{n-1} - M_{j-2}) / d_{n-1}; Normalize and concatenate vector D and vector R to generate a path pattern feature vector.
7. The method for accurately evaluating international communication effects based on a large language model according to claim 6, characterized in that: The parameter matching process includes: The propagation path pattern library is dynamically constructed based on the migration patterns and final matching degree of historical user groups between life cycle stages. Each pattern in the pattern library contains the following quantitative parameters: the minimum duration threshold T_min between stages, the slope range of the matching index change S_range, and the association weight W of adjacent stages; Calculate the duration matching degree S_time = 1 / (1 + |D-T_min|); Calculate the slope matching degree S_slope = ΣI(R_i∈S_range) / n, where I is the indicator function; Calculate the comprehensive matching degree S = W × S_time + (1-W) × S_slope; The pattern corresponding to the comprehensive matching degree S with the highest numerical value is taken as the classification result of the user's current behavior.
8. The method for accurately evaluating international communication effects based on a large language model according to claim 7 is characterized by: The process of determining the communication effect evaluation level includes: Set several evaluation level thresholds θ_1, θ_2, ..., θ_n, and compare the comprehensive matching degree S with the evaluation level thresholds; When S∈[θ_k,θ_{k+1}), it is determined to be the kth level propagation effect.
9. The method for accurately evaluating international communication effects based on a large language model according to claim 8, characterized in that: The process of implementing intervention strategies on a user group includes: Divide users into strong response groups, potential conversion groups, and low-efficiency reach groups based on evaluation levels; Increase push frequency for high-response groups; Push compensatory content to potential conversion groups and inefficiently reached groups: extract a set of semantic keywords with a positively increasing weight change rate from the current interest drift trend characteristics, filter candidate push content from the content library based on the semantic keyword set, calculate the similarity between the candidate push content and the user's interest tag set, and select the top N with the greatest similarity for targeted push.
10. A precise evaluation system for international communication effectiveness based on a large language model, characterized by: A method for accurately evaluating the effect of international communication based on a large language model as described in any one of claims 1 to 9 is used, comprising: A user behavior extraction module is used to obtain a user's behavior log data set and classify the user into a preset life cycle stage based on the behavior log data set; The interest drift extraction module is used to extract the user's text interaction content within a preset continuous time window and generate a set of interest tags using a large language model; the interest tag set includes semantic keywords and corresponding weight coefficients; the weight change rate of the same semantic keyword in adjacent time windows is calculated to generate user interest drift trend characteristics; A basic matching module is used to input the dissemination content to be evaluated into a large language model to generate a dissemination content feature vector, and calculate the basic matching degree between the dissemination content feature vector and the interest tag set in the current time window; A matching correction module, configured to dynamically correct the basic matching degree based on the user's life cycle stage and the interest drift trend characteristics to obtain a final matching degree; The effect evaluation module is used to extract the change rate of the user's final matching degree between adjacent life cycle stages and the duration of each stage as a path pattern feature vector; perform parameter matching on the path pattern feature vector and a preset propagation path pattern library, and determine the propagation effect evaluation level based on the calculation results of the parameter matching; The user intervention module is used to divide users into user groups according to the communication effect evaluation level and implement intervention strategies for the user groups.
Citation Information
Cited By
Big data advertisement label classification system based on AI analysis
CN120822133A
Big data advertisement tag classification system based on ai analysis
CN120822133B
Home intelligent control method and system based on computer vision
CN120949870A
Customer portrait matching and positioning method and system
CN121280084A
Real-time information evolution prediction method and system based on user portrait
CN121434259A