A social media public participation prediction method, medium and computer device

By integrating text and multimodal policy frameworks through the ELM dual-path information processing model, and using LDA and zero-inflated negative binomial regression models to analyze public engagement on social media, this approach solves the problem of unclear policy influence mechanisms in existing technologies, and enables accurate prediction and optimization of policies.

CN122113027APending Publication Date: 2026-05-29HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI UNIV OF TECH
Filing Date
2026-02-10
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing research has failed to systematically reveal how social media strategies affect public participation, and the effectiveness of communication strategies is mostly limited to relevance descriptions, lacking actionable optimization solutions.

Method used

We employ the ELM dual-path information processing model, integrating textual and multimodal strategy frameworks. We use the Latent Dirichlet Allocation (LDA) model for topic identification and prediction, and combine it with a zero-inflated negative binomial regression model to analyze public engagement on social media, thus constructing a multi-dimensional behavioral indicator model.

Benefits of technology

It enables accurate prediction of public engagement on social media, provides theoretical basis and practical optimization strategies, enhances communication effectiveness, and supports customized communication strategies for governments, businesses, and NGOs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113027A_ABST
    Figure CN122113027A_ABST
Patent Text Reader

Abstract

The application discloses a social media public participation degree prediction method, medium and computer equipment, wherein the method comprises the following steps: selecting a social media platform, obtaining a green travel related original data set on the social media platform, and obtaining high-quality samples after cleaning and screening; text preprocessing is performed on the high-quality samples to obtain effective samples, and a latent Dirichlet allocation model is selected for training and prediction, and the clustering effect of the latent Dirichlet allocation model is highly dependent on the selection of the number of themes; the number of themes is determined by comprehensively considering the perplexity and consistency evaluation indexes; text themes are extracted according to the theme recognition result of the latent Dirichlet allocation model, a prediction variable set is constructed by fusing multi-modal features, a social media public participation degree prediction model is established, and a participation degree prediction result and key driving factors are output. The application discloses the mechanism of the social media in spreading green travel, and provides a scientific and direct basis for customizing and optimizing a spreading strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary fields of information dissemination and machine learning, and more specifically to a method, medium, and computer device for predicting public engagement on social media based on the ELM dual-path information processing model. Background Technology

[0002] Social media, with its explosive information dissemination, deep interactivity, and widespread penetration (Sun et al., 2020a; Dong et al., 2017), has become a core platform for government departments, businesses, and non-governmental organizations (NGOs) to promote low-carbon policies, disseminate environmental knowledge, and shape green concepts. Research shows that effective communication by organizations on social media can significantly enhance public environmental awareness and willingness to act, such as promoting waste sorting (Sujata et al., 2019), increasing plastic recycling (de Fano et al., 2022), and guiding low-carbon consumption (Castro-Santa et al., 2023). Therefore, designing and implementing effective social media strategies has become a key lever for various organizations to motivate public participation in green travel.

[0003] Existing research on social media strategy identification tends to be fragmented: strategic elements (such as content themes, visual design, and interaction formats) are often discussed in isolation, failing to construct a unified framework integrating "content features" (what is said) and "non-content features" (how it is said) (Kim et al., 2021; Hallac et al., 2021). This results in an inability to systematically reveal the intrinsic mechanisms by which strategies influence public participation. Furthermore, from a methodological perspective, textual analyses of communication strategies are mostly limited to manual coding or simple word frequency statistics, making it difficult to quantify thematic heterogeneity and uncover rules governing relationships between themes. In terms of strategy effectiveness, evaluations of strategy effectiveness also largely remain at the level of relevance description, failing to propose empirically based and actionable optimization schemes to improve communication effectiveness. Summary of the Invention

[0004] This invention provides a method, medium, and computer equipment for predicting public engagement on social media based on an ELM dual-path information processing model. This method integrates the Fine Processing Assumption Model (ELM) with a multimodal strategy framework to reveal the mechanism by which social media promotes green travel and proposes a collaborative governance optimization strategy. This invention can at least solve one of the aforementioned technical problems.

[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A method for predicting public engagement on social media includes the following steps: S1. Select a social media platform that is credible, representative, highly interactive, and widely penetrated, obtain the original dataset related to green travel from the social media platform, and obtain high-quality samples after cleaning and screening. S2. Text preprocessing is performed on high-quality samples to obtain effective samples, and the Latent Dirichlet Allocation (LDA) model is selected for training and prediction. The clustering effect of the LDA model is highly dependent on the selection of the number of topics K. The number of topics is determined by combining two evaluation indicators: perplexity and consistency. S3. Extract text topics based on the Latent Dirichlet Allocation (LDA) topic identification results, integrate multimodal features to construct a predictive variable set, establish a social media public engagement prediction model based on the predictive variable set, and output engagement prediction results and key driving factors.

[0006] Furthermore, in S1, the acquisition of the original dataset employs a strategy combining keyword-targeted search and a two-layer crawler. This process specifically includes: S1.a1. Conduct a preliminary search using core concept terms and specific behavioral terms as keywords; S1.a2, Layered Data Collection: The first layer uses the Houyi Data Collector to batch crawl metadata of posts published on the social media platform, including ID account, posting time, posting content, interaction metrics, media resource URL, and publisher homepage URL. The second layer crawls account attribute information based on the publisher homepage URL, including VIP badge, number of followers, number of posts, official certification information, and industry category.

[0007] Furthermore, in S1, the original dataset contains a large amount of noise. To ensure that the samples represent the organization's propagation behavior, high-quality samples are constructed through cleaning and screening. This process specifically includes: S1.b1. Filter official media accounts based on authentication information and exclude content from personal user accounts; S1.b2. Remove duplicate entries based on the unique URL of the social media platform; S1.b3. Manually remove invalid content, including posts with prohibited comments, commercial advertisements, and irrelevant topics; S1.b4 Filter short text topics with fewer than 20 characters.

[0008] Furthermore, the text preprocessing process for the high-quality samples specifically includes: S2.1 Perform symbol cleaning, deleting special characters, abbreviations, and emojis; S2.2 Filtering stop words: The Harbin Institute of Technology general stop word list is used as the filtering basis, and the characteristic words of the social media platform are added to build a customized stop word library. S2.3 Deduplicating and summarizing the word segmentation results, while manually annotating and adding green travel-specific terms to improve the domain semantic capture capability; S2.4. Create a word-document matrix (DTM) as input to the Latent Dirichlet Distribution (LDA) model.

[0009] Furthermore, in S2, the perplexity index is used to evaluate the model's ability to predict new documents in the effective samples. The lower the value, the better the generalization of the new documents, and vice versa. The consistency index is used to calculate the semantic similarity between high-frequency words in the topic in the effective samples. The higher the value, the stronger the topic interpretability, and vice versa.

[0010] Furthermore, in S3, based on the detailed approximation model (ELM) dual-path framework, the variables studied include dependent variables, core independent variables, and moderating variables; The dependent variable is social media engagement, which is measured using a multi-dimensional set of behavioral indicators, including the number of reposts, comments, and likes for a single post. The core independent variables include text-based strategies on the core path and multimodal strategies on the edge path. Text-based strategies include text topics and text length, while multimodal strategies include account characteristics, conversational communication skills, headline characteristics, media richness, and publication time. The moderating variable is the industry category, which is divided into five categories based on certification information: government departments, news media, enterprises, public institutions, and non-profit organizations.

[0011] Furthermore, the Extensibility-Based Model (ELM) provides a core framework for understanding how individuals process persuasive information. This model indicates that information recipients primarily process information through two paths: First, the core path: when individuals have high motivation and ability, they will deeply process complex cues related to the core quality of information and think carefully. At this time, high-quality and profound information will guide individuals to generate positive feedback and lasting cognition, and increase their level of involvement. The second is the marginal path. When an individual lacks motivation or ability, they rely on simple, superficial heuristic cues to make quick judgments, forming shallow cognition and short-term attitude tendencies, and becoming less involved in things.

[0012] Furthermore, in S3, based on the study of variables, a quantitative model is constructed. Specifically, the numerical distribution characteristics of the variables are analyzed, and the zero-inflated negative binomial model ZINB is selected to conduct regression analysis on the relationship between information strategy and social media participation in green travel. The Zero-Inflated Negative Binomial (ZINB) model is a two-step counting model that automatically divides the dependent variable data into two latent groups: a zero-inflated group ZI consisting of all zero values ​​and a negative binomial group NB consisting of zero and non-zero values. The ZINB model uses Logit regression on the zero-inflated group ZI to study the probability that the dependent variable is zero, and performs negative binomial regression on the negative binomial group NB to study the direct driving effect of influencing factors on the dependent variable. The distribution function formula for the zero-inflated negative binomial model ZINB is:

[0013] Where P(·) represents the probability of the social media public engagement count result occurring given the model parameters, and Y i This represents the public engagement count for the i-th social media message. It is the probability that the dependent variable is zero. It is the average of the dependent variable. It is a divergence parameter that is independent of the covariates. Represents the gamma distribution function; The established zero-inflation negative binomial model ZINB formula is as follows:

[0014] Where Logit(·) represents the logarithmic probability function, and Both represent intercepts. and These are a series of text-based policy variables and multimodal policy variables, respectively. and Let represent the corresponding estimated coefficients, and let γ and δ represent the estimated coefficients of the corresponding variables in the negative binomial count part, respectively.

[0015] A computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the above-described social media public engagement prediction method.

[0016] A computer device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the above-described social media public engagement prediction method.

[0017] The beneficial effects of this invention are reflected in: 1. At the theoretical level, this paper is the first to systematically integrate the dual-dimensional characteristics of the ELM model and social media strategies, constructing a "textual-multimodal" dual-path framework: the former (core path) focuses on the intrinsic quality of information (text topic, length), while the latter (peripheral path) integrates surface cues (account characteristics, conversational communication skills, headline design, media richness, and publication time). This framework not only bridges the gap between content and non-content strategies but also provides a unified theoretical foundation for analyzing how organizations influence public participation through differentiated persuasion paths by embedding Uses and Gratifications Theory (U&G), Media Richness Theory (MRT), and Dialogue Communication Theory (DCT).

[0018] 2. At the methodological level, it innovatively integrates computational social science with traditional econometric analysis: On the one hand, it adopts unsupervised topic modeling to quantify the distribution of green travel text topics, revealing their heterogeneity in different industry organizations, and uses association rules to identify high-frequency topic combinations; on the other hand, in response to the over-dispersion and zero-inflation characteristics of public participation data (likes, comments, and reposts), it adopts zero-inflation negative binomial regression (ZINB) to accurately estimate the main effect of the dual-path strategy and the moderating effect of industry categories.

[0019] 3. In practice, empirical testing can provide direct evidence for governments, enterprises and NGOs to customize and optimize communication strategies. For example, based on the new mechanism of market incentives to improve policy implementation, an optimization strategy for the collaborative governance of "government credibility, enterprise innovation and media communication power" can be proposed. Attached Figure Description

[0020] The accompanying drawings, which are provided to further illustrate this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application.

[0021] Figure 1 This is a hypothetical model diagram of the overall prediction method in this invention.

[0022] Figure 2 This is a data cleaning flowchart according to an embodiment of the present invention.

[0023] Figure 3 This is a schematic diagram illustrating the topic distribution of multiple entities publishing green travel tweets according to an embodiment of the present invention.

[0024] Figure 4 This is a diagram illustrating the impact of government regulation in an embodiment of the present invention.

[0025] Figure 5 This is a diagram illustrating the influence of the regulatory role of enterprises in an embodiment of the present invention.

[0026] Figure 6 This is a diagram illustrating the effect of media regulation in an embodiment of the present invention.

[0027] Figure 7 This is a structural block diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] It should be noted that the meaning of "and / or" throughout the text includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, solution B, or a solution that simultaneously satisfies A and B. Furthermore, "multiple" refers to two or more. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0030] This invention continues the approach of existing literature, comprehensively using the number of reposts, comments, and likes to measure the public's participation in green travel on social media. In this implementation case, it provides a method for predicting public participation in social media based on the ELM dual-path information processing model. The overall prediction method research assumptions the model as follows: Figure 1 As shown.

[0031] The Estimated Availability Model (ELM) provides a core framework for understanding how individuals process persuasive information. This model posits that information recipients primarily process information through two paths: the Central Route and the Peripheral Route. When individuals are highly motivated and capable, they deeply process complex cues related to the core quality of the information (such as argument strength, relevance, and logic), engaging in careful consideration. In this case, high-quality and insightful information guides individuals to generate positive feedback and lasting cognition, increasing their involvement. The Peripheral Route, on the other hand, is where individuals, lacking sufficient motivation or ability, rely on simple, superficial heuristics (such as the attractiveness of the information source, information length, visual elements, and emotional arousal) to make quick judgments, forming shallow cognition and short-lived attitudinal tendencies, resulting in lower levels of involvement.

[0032] ELM provides a framework for individuals facing persuasive information and is widely used in research on various social media messages in different contexts. Core path cues are mostly complex analytical cues that reflect the quality of information, such as comment content, text topic, etc.; peripheral path cues are mostly heuristic cues other than text content, such as source credibility, media richness, etc.

[0033] This implementation case presents a method for predicting public engagement on social media based on the ELM dual-path information processing model. Sina Weibo is chosen as the social media platform for this study. The specific implementation steps include: Step 1: Data Collection and Screening The data in this invention comes from green travel-related content published by organizations on the Sina Weibo platform. According to Sina Weibo's 2023 annual financial report (Corporation, 2024), as of the end of the fourth quarter, Weibo's monthly active users reached 598 million, accounting for approximately 54.76% of the total number of Chinese internet users during the same period (Center, 2024). On the Weibo platform, government departments, enterprises, mainstream media, and other organizations promote low-carbon living to the public through text, images, and videos, and the rapid spread of information fosters a cultural atmosphere of nationwide participation in green travel. The data collection period is set from September 2020 to December 2023. It should be noted that Sina Weibo adjusted its industry classification identification method starting in 2024; to ensure the comparability of industry category variables across periods, the data ends on December 31, 2023.

[0034] Data acquisition employs a strategy combining keyword-targeted search and a two-layer crawler. First, the advanced search function on Weibo is used to conduct an initial search using core conceptual terms (such as "green travel," "low-carbon travel," and "environmentally friendly travel") and specific behavioral terms (such as "cycling," "walking," "subway," and "bus"). Then, layered data collection is implemented: the first layer uses a crawler to batch-crawl Weibo post metadata, covering Weibo ID, posting time, posting content (including title, @mentions, and external links), interaction metrics (number of reposts / comments / likes), media resource URLs, and the poster's homepage URL; the second layer crawls account attribute information based on the homepage URL, including VIP badges, number of followers, number of posts, official certification information, and industry category affiliation.

[0035] The original dataset contained a large amount of noise. To ensure that the samples represent the organization's dissemination behavior, this invention constructs high-quality samples through a four-step cleaning process: First, "official Weibo" accounts are filtered based on authentication information to exclude content from personal users; second, duplicate entries are removed based on the unique URL of the Weibo posts; subsequently, invalid content, including posts with prohibited comments, commercial advertisements, and irrelevant topics, is manually removed; finally, short texts with fewer than 20 characters are filtered out because such data has low information content and is not conducive to subsequent topic modeling (Xu et al., 2021, Le and Mikolov, 2014). After the above four-step cleaning process, 36,270 valid organizational Weibo samples were finally obtained. The cleaning process is as follows: Figure 2 As shown.

[0036] Step 2: Topic Modeling for Text-Based Strategies Existing research on identifying topics in social media texts mainly relies on manual coding (Bonsón et al., 2015; Chen et al., 2020) or supervised machine learning (Lee et al., 2018; Wang and Yang, 2020), which suffers from limitations such as limited sample size, subjective bias, and incomplete topic coverage. To accurately analyze the content design strategies of various organizations in promoting green travel, this invention adopts the Latent Dirichlet Allocation (LDA) topic model (Blei et al., 2003). Its advantage lies in that it does not require a pre-set topic framework, but automatically infers the potential topic structure from the corpus through a probability generation process, combining large sample processing capability with the objectivity of topic discovery (Xu et al., 2021, Tafesse et al., 2015, Cheng et al., 2021, Sun et al., 2020b). The clustering effect of LDA is highly dependent on the selection of the number of topics (K). This invention determines the number of topics by combining two evaluation indicators: perplexity and consistency. The former is used to evaluate the model's predictive ability for new documents (Blei et al., 2003), and the lower the value, the better the generalization. The latter is used to calculate the semantic similarity between high-frequency words within a topic (Röder et al., 2015), and the higher the value, the stronger the topic interpretability.

[0037] Furthermore, text preprocessing is a key prerequisite for topic modeling, and includes four steps: First, symbol cleaning is performed to remove special characters, abbreviations, and emojis; second, stop words are filtered, using the Harbin Institute of Technology's general stop word list as the filtering basis and supplementing it with platform feature words to build a customized stop word library; then, the word segmentation results are deduplicated and summarized, while green travel-specific terms are manually annotated and added to improve the domain semantic capture capability; finally, a word-document matrix (DTM) (Chun et al., 2021) is created as the input to the LDA model.

[0038] Step 3: Examining the driving mechanisms of social media engagement 3-1 Variable Operations Based on the ELM dual-path framework, the variables studied in this invention include dependent variables, core independent variables, and moderating variables, among which: The dependent variable was social media engagement, which was comprehensively measured using multi-dimensional behavioral indicators, including the number of reposts, comments, and likes for a single Weibo post. To comprehensively reflect the intensity of public participation, the three indicators were aggregated into a composite indicator, following the method of Sun et al. (2020b). Subsequent analysis also confirmed that the three indicators were highly correlated.

[0039] The core independent variables include text-based strategies on the core path and multimodal strategies on the peripheral path. Text-based strategies include text topic and text length, while multimodal strategies include account characteristics, conversational communication skills, headline characteristics, media richness, and publication time.

[0040] The moderating variable is the industry category, which is divided into five categories based on certification information: government departments, media, enterprises, public institutions, and public welfare organizations (NGOs).

[0041] For a detailed list of all variable categories, names, and measurement metrics, please refer to Table 1.

[0042] Table 1. Measurement Table of Each Variable

[0043] 3-2. Model Construction The process for selecting the measurement model in this invention is as follows: (1) Considering that the dependent variables are all count data, Poisson or negative binomial regression is commonly used.

[0044] (2) The Poisson model strictly requires that the mean and variance of the dependent variable be equal. The dependent variable data in this invention exhibits a significant tendency to be overdispersed, and violating the above assumptions will lead to an estimation bias of "underestimating the standard error and overestimating the significance" (Hilbe, 2017). Choosing the negative binomial regression model is superior to the traditional Poisson model.

[0045] (3) There are a large number of zero values ​​in the number of reposts, comments, and likes, accounting for 66.93%, 67.41%, and 49.93% respectively. The count distribution is seriously unbalanced. The zero-inflation model can effectively solve this problem. In addition, the Vuong statistic used to test whether there are too many zero frequencies in the model is significantly greater than 1.96, indicating that the zero-inflation model is better than the non-inflation model.

[0046] In summary, this invention determines that a zero-inflated negative binomial model (ZINB) (Long and Freese, 2005) should be used to perform regression analysis on the relationship between information strategy and social media engagement in green travel.

[0047] The ZINB model is a two-step counting model that automatically divides the dependent variable data into two latent groups: a zero-inflated group (ZI) consisting entirely of zero values, and a negative binomial group (NB) consisting of zero and non-zero values ​​(Moreno-Mondejar et al., 2021). The model uses Logit regression on the former to study the probability of the dependent variable being zero, and negative binomial regression on the latter to study the direct driving effect of influencing factors on the dependent variable (Long and Freese, 2005). The distribution function of the ZINB model is shown in the following formula:

[0048] Where P(·) represents the probability of the social media public engagement count result occurring given the model parameters, and Y i This represents the public engagement count for the i-th social media message. It is the probability that the dependent variable is zero. It is the average of the dependent variable. It is a divergence parameter that is independent of the covariates. Represents the gamma distribution function; Combining the above variables, the ZINB model is established as shown in the following formula:

[0049] Where Logit(·) represents the logarithmic probability function, and Both represent intercepts. and These are a series of text-based policy variables and a multimodal policy variable, respectively. and Let represent the corresponding estimated coefficients, and let γ and δ represent the estimated coefficients of the corresponding variables in the negative binomial count part, respectively.

[0050] Based on the social media public engagement prediction method based on the ELM dual-path information processing model provided in the above embodiments of the present invention, the experimental results and analysis findings are as follows: I. Thematic Analysis of Text-Based Strategies 1-1. Text Topic Recognition Results Based on LDA topic modeling of 36,270 organizational microblogs, the optimal number of topics was determined to be 8 by balancing the model's generalization ability and semantic consistency.

[0051] 1-2. Thematic Distribution of Green Travel Tweets by Multiple Entities like Figure 3 The chart shows the distribution of social media text topics used by organizations across different industry sectors to promote green travel. Among them, Figure 3 (a) The left column represents the publishing entity, the right column represents the text topic, and the stream width represents the proportion of each industry account disclosing the corresponding topic; Figure 3 (b) reflects the information disclosure situation of different departments within the three major categories of organizations: government, media and business.

[0052] II. Empirical Analysis 2-1. Analysis of the direct effects of social media strategies The regression results of the zero-inflated negative binomial model (ZINB) are shown in Table 2. The ln(α) of all three models is significantly greater than 0, indicating that the negative binomial model is superior to Poisson regression. Furthermore, the Vuong values ​​of all three models are significantly greater than 1.96, demonstrating that the zero-inflated negative binomial regression model is suitable for the research data of this invention.

[0053] The regression results for the Zero Inflation (ZI) group show that: 1. None of the eight text topic variables significantly affect the probability of zero reposts or likes on Weibo, but they significantly negatively affect the probability of zero comments, and their influence coefficients are generally larger than those of the multimodal strategy variables. This indicates that text topics reflecting content quality are the key factors determining whether a green travel Weibo post can attract public comments. 2. VIP membership, number of followers, use of the Mention feature, and media richness negatively affect the probability of a zero value for social media green travel engagement (GTSME). Among these, the number of followers has the largest absolute value of its influence coefficient, indicating that the larger the follower base, the lower the probability that a Weibo post will go unnoticed.

[0054] Regression results for the negative binomial group (NB) show that the text strategy significantly affects GTSME. All eight text topics significantly increased the number of shares, comments, and likes. >0, p-value <0.001), and the impact was generally greater than other multimodal policy variables. Among them, Topic 6 and Topic 7 had the greatest positive effect. In addition, text length significantly suppressed social media engagement in green travel ( <0, p-value <0.001). Therefore, H1a and H1b are supported.

[0055] Multimodal strategies have a limited impact on GTSME. (1) VIP access and accounts with a large number of followers have a significant positive effect. >0, p-value <0.001). Interestingly, although, according to common sense, organizational accounts with high and stable posting frequency are often more trusted by the public, the research results show that the total number of posts by an organization has a significant negative effect on the number of reposts and likes, with a very low impact strength, and does not affect the number of comments. Hypotheses H2a and H2b are valid, while hypothesis H2c is not valid. (2) Mention and Weblink, two forms of dialogue communication, play drastically different roles in social media participation in green travel. The former plays a positive role ( >0, p-value <0.001), the latter significantly suppresses the dependent variable ( <0, p-value <0.001). These two results further validate the "dialogue cycle principle" and "user retention principle" proposed by the dialogue communication theory. Hypotheses H3a and H3b are valid. (3) The title format has no significant effect on GTSME, but the title length has a significant inhibitory effect ( <0, p-value <0.001). Zhuoming's concise and powerful short headlines are often more effective in directly conveying the value of social media information. H4a is not true, H4b is true. (4) Media richness can significantly increase the number of likes ( =0.059, p-value <0.001, but does not affect the number of reposts and comments, so the H5 part is valid. (5) Posting information during peak user activity periods can significantly increase the number of reposts ( =-0.078, p-value <0.05), but had no impact on the number of comments and likes. Organizations should appropriately consider the timing of information release when promoting green travel on social media platforms. The H7 part is valid.

[0056] Table 2. Results of ZINB Main Effects Regression

[0057] 2-2 Analysis of the Moderating Effect of Industry Categories This invention introduces an interaction term between account category and text topic variables into the aforementioned empirical model to investigate whether account category can moderate the impact of text topic on social media engagement with green travel. The results are illustrated as follows: Figures 4-6 The moderating effect diagram shown here only considers the three types of accounts with the largest number of posts: government, media, and enterprises, and only shows a few cases where the moderating effect is significant.

[0058] Depend on Figure 4 It can be seen that posts by government departments on Weibo promoting low-carbon living knowledge and public transportation incentives are more likely to be forwarded and commented on than those by non-governmental departments, but this boost is not significant.

[0059] Combination Figure 4 and Figure 5 It is evident that, in the process of increasing public participation in green travel on social media under the theme of influencing public transportation incentives (Topic 7), corporate accounts play a more significant and substantial positive moderating role, particularly in increasing likes on Topic 7. Furthermore, as the proportion of content posted by corporate Weibo accounts on Topic 6 (new energy vehicle infrastructure) continues to increase, the slope of the curve also rises, indicating a continuously strengthening positive moderating effect of the account type.

[0060] Figure 6 This demonstrates that the type of media account suppressed the effectiveness of Topic 1 (traffic conditions and safety tips) and Topic 5 (public transportation system construction) in increasing social media engagement for green travel. In other words, posts about traffic conditions, safety tips, and public transportation system construction from non-media Weibo accounts were more likely to be shared, commented on, and liked by the public than posts from official media accounts. Therefore, assuming H7 is true...

[0061] III. Key Research Findings and Recommendations 3-1. Textual strategies along the central route significantly impact social media engagement with green travel. Analysis revealed a significant impact of text strategies related to core path cues and reflecting content quality on public social media engagement. All eight text themes showed significant positive effects, with the themes of new energy vehicle infrastructure construction, low-carbon policy planning and technological innovation, and public transportation incentives exhibiting the most pronounced positive effects. For any social media engagement behavior—forwarding, commenting, or liking—the theme "new energy vehicle infrastructure construction" had the strongest influence, reflecting public attention to some extent. On one hand, weak charging infrastructure (Topic 6) is a significant obstacle to individual purchases of new energy vehicles; organizations can alleviate public concerns by frequently disclosing information such as "community power grid construction" and "charging pile installation and charging base station construction." On the other hand, public transportation incentives play a positive role. In the social media context, the public actively participates in promotional information released by organizations to encourage green travel, such as discounts on public transport and subway card top-ups, the "Green Travel - Carbon Inclusiveness" platform launched by Gaode Maps, and Alipay's Ant Forest carbon credit program. This aligns with the findings of (Liu et al., 2021), who argued that corporate accounts posting content related to "rewards," "lotteries," and "gifts" are directly relevant to netizens' interests, and that monetary and relational incentives can achieve higher levels of customer engagement.

[0062] Furthermore, the analysis also found that lengthy microblog posts were associated with lower social media engagement. This may be because longer texts, containing more information and deeper meaning, require readers to have sufficient comprehension and knowledge, and netizens often spend more time and effort interacting with such posts. Microblogs, on the other hand, are information platforms for fragmented reading; in order to obtain more useful information in a shorter time, the public prefers to read concise and focused posts.

[0063] 3-2. Further Discussion Based on Frequent Itemset Mining The regression results show that the significant positive effect of text topics is generally greater than that of other non-textual information strategy variables, indicating that topic arrangement and content design are crucial aspects of organizing and promoting green travel information strategies. To further explore the combined use and matching of text topics, this invention further mines frequent itemsets of text topics in posts with high engagement. Frequent itemset mining is a commonly used data mining method with wide applications in social media text topic or content analysis (Tanantong and Ramjan, 2021, Zuo et al., 2019). Using this method, patterns that are not accidental can be discovered from large amounts of text data, thus providing a basis for social media promotion. The implementation steps are as follows: (1) Process the percentage of each post under each topic. When the product of the text length and the topic percentage is less than 10 characters, the topic is considered to be inconspicuous and assigned a value of 0; otherwise, it is assigned a value of 1.

[0064] (2) Use Shannon information entropy to reduce the dimensionality of the number of reposts, comments and likes (Shannon, 1948) to form the GTSME value of each Weibo post.

[0065] (3) Using the K-median method, the classification samples were clustered into four groups: High, Medium-high, Medium-low, and Low based on the GTSME value. The data from the High and Medium-high groups were used to mine frequent itemsets of the text topics.

[0066] (4) The Apriori algorithm was developed to find the frequent itemset with a support greater than or equal to the minimum support. Support(A,B) represents the frequency of posts that include both topics A and B in the total number of posts (Li et al., 2022). Due to the large difference in the number of microblog samples from various organizations, setting a uniform minimum support would lead to an imbalance in the output of frequent itemsets. The threshold needs to be adjusted repeatedly according to the number of samples. Finally, the value of Min_sup was set to be in the range of [0.15, 0.2]. The mining results are shown in Table 3.

[0067] Table 3. Results of Association Rule Mining

[0068] The results of frequent itemset mining indicate that the popularity of some Weibo posts is not accidental; certain fixed combinations of text themes can attract a high level of public participation in green travel on social media. For example, the text theme combination of Topic 3 (civilized cycling initiative) and Topic 7 (public transportation incentives) frequently appears in posts published by government departments that demonstrate high levels of green travel participation on social media, accounting for approximately 19.6%, with about 44 Weibo samples exhibiting this text theme combination.

[0069] 3-3. Multi-level suggestions The findings of this invention provide several insightful observations on how different organizations can effectively utilize social media to promote green travel.

[0070] First, optimize text strategies and strengthen the combination of themes and demand response. Empirical evidence shows that "new energy vehicle infrastructure" and "public transportation incentives" have the most significant impact on increasing participation. Therefore, the government can collaborate with enterprises to release combined posts such as "charging pile layout planning + carbon credit redemption"; enterprises can link technology disclosure (such as BYD's battery range) with immediate incentives (such as "riding bikes to redeem food delivery coupons"). Simultaneously, a mechanism for identifying trending comments should be established to enhance the tracking and analysis of posts with high engagement on social media, deepen the text mining of online comments, capture public concerns, and drive content iteration.

[0071] Secondly, the focus of implementing a multimodal strategy should be restructured. Empirical evidence shows that VIP certification, number of followers, headline characteristics, and dialogue communication skills can significantly improve topic engagement. Therefore, when releasing information, organizations should ideally have authoritative endorsements, such as government / enterprise collaboration with professional institutions (e.g., the Environmental Impact Assessment Center of the Chinese Academy of Sciences) to release data verification content (e.g., "Calculation of Emission Reduction by Cycling"). Emphasis should be placed on converting fan economy into revenue; for example, enterprises can collaborate with transportation KOLs to conduct "carbon credit challenges" (e.g., promoting shared bicycles with cycling bloggers); governments can invite public figures to serve as "green travel ambassadors" to expand the reach and reach of their message. Furthermore, headlines are key to attracting public clicks; organizations should focus on their guiding and attractive qualities when crafting headlines. At the same time, media format selectivity should be strengthened, such as increasing the proportion of text / images / videos, but avoiding complex layouts (e.g., long videos with embedded external links); external links should be used cautiously, and key information should be embedded when necessary (e.g., converting discount details into infographics).

[0072] Finally, an industry-wide collaborative communication network should be built. Empirical results show that industry category has a moderating effect on the influence of certain text topics on social media engagement. Therefore, the government can take the lead in building credibility and leverage its policy decoding advantages to promptly interpret relevant policies, such as the Ministry of Finance's new energy subsidy details, but should avoid lengthy texts; enterprises should focus on incentivizing innovation and developing scenario-based incentive tools, such as Meituan's "ride-hailing for food delivery coupons"; and media should streamline communication channels and establish a cross-platform content distribution matrix, such as in-depth interpretations on WeChat official accounts combined with Weibo topic traffic generation.

[0073] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the social media public engagement prediction method described above.

[0074] See Figure 7 The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the social media public engagement prediction method described above.

[0075] This invention also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps of the social media public engagement prediction method described above.

[0076] It is understood that the systems, devices, and storage media provided in the embodiments of the present invention correspond to the methods provided in the embodiments of the present invention, and the explanations, examples, and beneficial effects of the relevant content can be referred to the corresponding parts of the above-mentioned social media public engagement prediction method.

[0077] It should be noted that those skilled in the art will understand that all or part of the steps implemented in the embodiments of the present invention can be implemented entirely or partially by software, hardware, firmware, or any combination thereof. When implemented in hardware, it can be implemented entirely or partially by purchasing standard parts or modifications. When implemented in software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid state disks (SSDs)).

[0078] In summary, this invention addresses the problems of fragmented identification of social media strategies in existing research, textual analysis of communication strategies often limited to manual coding or simple word frequency statistics, and strategy effectiveness testing mostly remaining at the level of relevance description. It provides a method for predicting public engagement on social media based on the ELM dual-path information processing model. Based on data from 36,270 organizational microblog posts, this method integrates the Fine Processing Likelihood Model (ELM) with a multimodal strategy framework to reveal the mechanism by which social media promotes green travel. The main contributions are as follows: First, the applicability of the ELM dual-path approach to environmental communication was verified: on the core path, the text theme, especially incentives for new energy vehicle facilities and public transportation, is the core element driving public participation, while on the peripheral path, account authority and dialogue communication skills play a supporting role. Secondly, the study discovered the moderating effect of industry categories: the incentive measures pushed by enterprise accounts were significantly more effective than those of the government, revealing to some extent a new mechanism for market incentives to compensate for policy implementation. Based on this, an optimization strategy of collaborative governance of "government credibility, enterprise innovation, and media communication" was proposed.

[0079] It should be understood that the examples and embodiments described herein are for illustrative purposes only and are not intended to limit the invention. Those skilled in the art can make various modifications or changes based on them. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the protection scope of the invention.

Claims

1. A method for predicting public engagement on social media, characterized in that, Includes the following steps: S1. Select a social media platform that is credible, representative, highly interactive, and widely penetrated, obtain the original dataset related to green travel from the social media platform, and obtain high-quality samples after cleaning and screening. S2. Text preprocessing is performed on high-quality samples to obtain effective samples, and the Latent Dirichlet Allocation (LDA) model is selected for training and prediction. The clustering effect of the LDA model is highly dependent on the selection of the number of topics K. The number of topics is determined by combining two evaluation indicators: perplexity and consistency. S3. Extract text topics based on the Latent Dirichlet Allocation (LDA) topic identification results, integrate multimodal features to construct a predictive variable set, establish a social media public engagement prediction model based on the predictive variable set, and output engagement prediction results and key driving factors.

2. The social media public engagement prediction method as described in claim 1, characterized in that, In S1, the acquisition of the original dataset employs a strategy combining keyword-targeted search and a two-layer crawler. This process specifically includes: S1.a1. Conduct a preliminary search using core concept terms and specific behavioral terms as keywords; S1.a2, Layered Data Collection: The first layer uses the Houyi Data Collector to batch crawl metadata of posts published on the social media platform, including ID account, posting time, posting content, interaction metrics, media resource URL, and publisher homepage URL. The second layer crawls account attribute information based on the publisher homepage URL, including VIP badge, number of followers, number of posts, official certification information, and industry category.

3. The social media public engagement prediction method as described in claim 1, characterized in that, In S1, the original dataset contains a lot of noise. To ensure that the samples represent the tissue propagation behavior, high-quality samples are constructed through cleaning and screening. This process specifically includes: S1.b1. Filter official media accounts based on authentication information and exclude content from personal user accounts; S1.b2. Remove duplicate entries based on the unique URL of the social media platform; S1.b3. Manually remove invalid content, including posts with prohibited comments, commercial advertisements, and irrelevant topics; S1.b4 Filter short text topics with fewer than 20 characters.

4. The social media public engagement prediction method as described in claim 1, characterized in that, The text preprocessing process for the high-quality samples specifically includes: S2.1 Perform symbol cleaning, deleting special characters, abbreviations, and emojis; S2.2 Filtering stop words: The Harbin Institute of Technology general stop word list is used as the filtering basis, and the characteristic words of the social media platform are added to build a customized stop word library. S2.3 Deduplicating and summarizing the word segmentation results, while manually annotating and adding green travel-specific terms to improve the domain semantic capture capability; S2.

4. Create a word-document matrix (DTM) as input to the Latent Dirichlet Distribution (LDA) model.

5. The social media public engagement prediction method as described in claim 1, characterized in that, In S2, the perplexity index is used to evaluate the model's ability to predict new documents in the effective samples. The lower the value, the better the generalization of the new documents, and vice versa. The consistency index is used to calculate the semantic similarity between high-frequency words in the topic in the effective samples. The higher the value, the stronger the topic interpretability, and vice versa.

6. The social media public engagement prediction method as described in claim 1, characterized in that, In S3, based on the detailed approximation model (ELM) dual-path framework, the variables studied include dependent variables, core independent variables, and moderating variables. The dependent variable is social media engagement, which is measured using a multi-dimensional set of behavioral indicators, including the number of reposts, comments, and likes for a single post. The core independent variables include text-based strategies on the core path and multimodal strategies on the edge path. Text-based strategies include text topics and text length, while multimodal strategies include account characteristics, conversational communication skills, headline characteristics, media richness, and publication time. The moderating variable is the industry category, which is divided into five categories based on certification information: government departments, news media, enterprises, public institutions, and non-profit organizations.

7. The social media public engagement prediction method as described in claim 6, characterized in that, The Extensibility-Based Model (ELM) provides a core framework for understanding how individuals process persuasive information. The model indicates that information recipients primarily process information through two paths: First, the core path: when individuals have high motivation and ability, they will deeply process complex cues related to the core quality of information and think carefully. At this time, high-quality and profound information will guide individuals to generate positive feedback and lasting cognition, and increase their level of involvement. The second is the marginal path. When an individual lacks motivation or ability, they rely on simple, superficial heuristic cues to make quick judgments, forming shallow cognition and short-term attitude tendencies, and becoming less involved in things.

8. The social media public engagement prediction method as described in claim 6, characterized in that, In S3, based on the study of variables, a quantitative model is constructed. Specifically, the numerical distribution characteristics of the variables are analyzed, and the zero-inflated negative binomial model ZINB is selected to conduct regression analysis on the relationship between information strategy and social media participation in green travel. The Zero-Inflated Negative Binomial (ZINB) model is a two-step counting model that automatically divides the dependent variable data into two latent groups: a zero-inflated group ZI consisting of all zero values ​​and a negative binomial group NB consisting of zero and non-zero values. The ZINB model uses Logit regression on the zero-inflated group ZI to study the probability that the dependent variable is zero, and performs negative binomial regression on the negative binomial group NB to study the direct driving effect of influencing factors on the dependent variable. The distribution function formula for the zero-inflated negative binomial model ZINB is: Where P(·) represents the probability of the social media public engagement count result occurring given the model parameters, and Y i This represents the public engagement count for the i-th social media message. It is the probability that the dependent variable is zero. It is the average of the dependent variable. It is a divergence parameter that is independent of the covariates. Represents the gamma distribution function; The established zero-inflation negative binomial model ZINB formula is as follows: Where Logit(·) represents the logarithmic probability function, and Both represent intercepts. and These are a series of text-based policy variables and a multimodal policy variable, respectively. and Let represent the corresponding estimated coefficients, and let γ and δ represent the estimated coefficients of the corresponding variables in the negative binomial count part, respectively.

9. A computer-readable storage medium, characterized in that, The system contains a computer program that, when executed by a processor, causes the processor to perform the steps of the social media public engagement prediction method as described in any one of claims 1-8.

10. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the social media public engagement prediction method as described in any one of claims 1-8.