A variable-granularity region-wise stance aggregation method and system
By constructing a multi-model integration framework and a dynamic regional division mechanism, the systematic bias and data representativeness issues in regional position prediction were resolved, achieving a more scientific and reliable aggregated measurement of regional positions and improving the accuracy and robustness of predictions.
Patent Information
- Application Number
- CN202511457971.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-10-13
AI Technical Summary
Existing technologies for regional position prediction suffer from problems such as systematic bias caused by a single model, difficulty in fusion of heterogeneous data, neglect of differences in data quality and representativeness, and lack of dynamic adaptive capability in static prediction, resulting in insufficient robustness and accuracy of prediction results.
A multi-model integration framework is constructed to run sub-models of social media, regional historical features, and multi-source trend data in parallel. A dynamic partitioning mechanism of core and variable regions is introduced, using the stable core region as a reliable anchor point. The fusion weights are adjusted through optimization algorithms, and a data representativeness correction mechanism is established.
It improves the accuracy, robustness, and dynamic adaptability of regional position prediction, effectively overcomes the systematic bias of single models, and achieves deep fusion of heterogeneous data and reduction of sample bias.
Smart Images

Figure CN120930886B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of position prediction technology, and specifically to a method and system for position aggregation measurement in variable granularity regions. Background Technology
[0002] In fields such as social sciences and market analysis, accurately quantifying and predicting the attitudes (such as support / opposition, like / dislike, optimism / pessimism) of people in a specific region towards a particular issue, product, or event has significant application value.
[0003] With the widespread adoption of the internet and social media, data sources have become unprecedentedly abundant, including massive amounts of user-generated content (UGC), traditional survey data, and regional macroeconomic and social characteristic data. These data are characterized by heterogeneous sources, high timeliness, and inconsistent quality.
[0004] Currently, most research on position prediction relies on a single data source or method. For example:
[0005] Traditional research methods involve collecting sample data through questionnaires and other means, then statistically analyzing the data to infer the overall regional stance. This method is costly, time-consuming, and susceptible to sample selection bias that can affect the generalizability of the results.
[0006] Social media analytics-based methods analyze online text using natural language processing to uncover user opinions. While this approach is highly timely, it faces challenges such as biases in user representation and inaccurate contextual understanding.
[0007] Regional characteristic-based statistical models utilize historical data on a region's population, economy, and culture to build regression or classification models for prediction. This method is relatively macro-level and struggles to capture rapid, dynamic changes in stance caused by unforeseen events.
[0008] Therefore, how to effectively integrate multi-source heterogeneous data, combine analytical models of different granularities, and consider the differences in the representativeness of data in different regions, so as to achieve a more accurate, robust, and dynamic aggregation measurement of regional positions, is a core problem that urgently needs to be solved in the current technical field.
[0009] Current technical solutions in the field of regional position measurement mainly suffer from the following problems:
[0010] First, there is the systematic bias introduced by a single model. Relying on a single data source (such as social media data only or survey data only) or a single model (such as time series model only or feature model only) for prediction cannot avoid the inherent systematic bias of that data source or model, resulting in insufficient robustness and accuracy of the prediction results.
[0011] Second, heterogeneous data fusion is difficult. It's challenging to effectively integrate structured regional characteristic data (such as economic indicators), semi-structured survey data, and unstructured social media text data. Simple methods like weighted averages cannot fully utilize the deeper information contained within various data types.
[0012] Third, it ignores differences in data quality and representativeness. When using online data such as social media, existing methods usually assume that the sample is uniform and representative, but fail to consider that the representativeness of the data may vary greatly across different geographical regions, leading to model distortion in some areas.
[0013] Fourth, static prediction lacks dynamic adaptability. Most models are static; once trained, their parameters and weights are fixed. This makes it difficult for the model to adapt to rapid changes in the external information environment and to dynamically adjust the contribution of various data sources and the model in the final prediction. Summary of the Invention
[0014] To address the aforementioned technical problems, this invention provides a variable-granularity regional stance aggregation measurement method and system. Its core lies in constructing a multi-model integration framework that runs multiple sub-models in parallel based on social media, regional historical characteristics, and multi-source trend data. It innovatively introduces a dynamic partitioning mechanism for core and variable regions, utilizing the stable core region as a reliable anchor point to constrain and optimize the fusion weights of the multiple models. Simultaneously, a data representativeness correction mechanism is established to overcome sample bias. Through these design features, this invention aims to systematically solve the above challenges and provide a more scientific and reliable variable-granularity regional stance aggregation solution.
[0015] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0016] In a first aspect, the present invention provides a method for measuring the aggregation of positions in a variable-grained region, comprising:
[0017] Acquire and preprocess multi-source heterogeneous data, including social media content, regional economic characteristics, and multi-source position trend data;
[0018] Multiple independent sub-models are run in parallel, including a position analysis model based on social media content, a decision-making model based on regional historical economic characteristics, and a time-series prediction model based on network volume fitting and baseline trend correction, to generate preliminary position prediction results for each region.
[0019] Based on the stability of historical positions and the consistency of position prediction results of each sub-model, the regional dynamics are divided into core areas and variable areas.
[0020] Using the position prediction results of the core region as a hard constraint, the global optimal weights that make the weight distribution of each sub-model, the core region and the variable region most balanced are solved by an optimization algorithm.
[0021] For each region, the representativeness of a specific data source is assessed, and the globally optimal weights are locally adjusted based on the difference between the sample proportion and the standard reference proportion.
[0022] Using locally modified weights, the preliminary position prediction results of each sub-model, as well as the position prediction results of the core region and the variable region, are weighted and fused to output the final regional position aggregation measure value.
[0023] In one embodiment, the acquisition and preprocessing of multi-source heterogeneous data, including social media content, regional economic characteristics, and multi-source position trend data, specifically includes:
[0024] Social media content ,in, Representing the first user, represent The published text content, represent Geographic location information;
[0025] Regional economic characteristics ,in, Represents the first area. represent Economic characteristic data, represent Historical regional stance labels;
[0026] Multi-source position trend data ,in, Representing the first event, represent The corresponding position trend time series.
[0027] In one embodiment, the stance analysis model based on social media content specifically includes:
[0028] By utilizing users' geolocation information within social media content, users can be assigned to corresponding regions;
[0029] Using a large language model, for each user Published text content Conduct position analysis by designing prompts that contain keywords related to specific issues. This guides large language models to determine and quantify the stance expressed in text content. Resulting in a quantitative value of the position :
[0030] ;
[0031] For each region The quantitative values of all user positions within the region are statistically summed to obtain the preliminary position prediction results of the position analysis model for region r. :
[0032] ;
[0033] For the region The collection of all users within.
[0034] In one embodiment, the decision-making model based on regional historical economic characteristics specifically includes:
[0035] Use Area Historical economic characteristics data Corresponding to historical regional stance labels Matching is performed to form a training dataset, which is then used to train the decision model based on gradient boosting decision trees;
[0036] The latest regional feature data is input into the trained decision model. In the process, the preliminary position prediction results of the decision model for region r were obtained. .
[0037] In one embodiment, the time-series prediction model based on network volume fitting and baseline trend correction specifically includes:
[0038] Extract target words from social media content, combine the text content posted by users with relevant target words and pre-set prompt templates, and feed them into a large model to obtain a prediction of the user's stance.
[0039] This study integrates various text content attributes related to online buzz and assigns weights to these attributes based on actual circumstances. It then employs the Pettitt test to identify salient points based on the published text content and corresponding comments. Finally, it allocates online buzz for different users based on attribute weights and salient points.
[0040] By integrating multi-source position trend data and employing distance-weighted or error-minimizing fitting methods, a baseline position trend time series curve is generated. ;
[0041] It is necessary to construct a time series curve of online popularity based on each user's stance and their online influence. ;
[0042] The time series curve of online popularity is fused with the time series curve of baseline position trend, and then input into a time series prediction model based on the ARIMA model. In the process, the time-series prediction model for the region is obtained. Preliminary position forecast results : , This represents historical time series data related to region r.
[0043] In one embodiment, the step of dynamically dividing the region into a core region and a variable region based on historical position stability and the consistency of position prediction results from each sub-model specifically includes:
[0044] Condition a): The current regional stance has remained consistent throughout the past set of historical cycles;
[0045] Condition b): The position analysis model and the decision model predict the same regional position as in condition a);
[0046] If conditions a) and b) are met, and the region's stance is supportive, then the current region belongs to the core support region set. If conditions a) and b) are met, and the region's stance is "oppose," then the current region belongs to the core set of opposing regions. Other regions belong to the variable region set. The core support area and the core opposition area are collectively referred to as the core area.
[0047] In one embodiment, the step of using the position prediction results of the core region as a hard constraint and solving for the globally optimal weights that make the weight distribution of each sub-model, the core region, and the variable region most balanced through an optimization algorithm specifically includes:
[0048] The optimization objective is set to minimize the standard deviation of the weights of each sub-model, the core region, and the variable region:
[0049] ;
[0050] The standard deviation of the weight vector is represented by... Represents the weight vector. This represents the arithmetic mean of the weights. ; These represent the weights of the position analysis model, the decision model, the time series prediction model, the core region, and the variable region, respectively.
[0051] Constraints:
[0052] The final fusion score for all core support regions must be greater than a set positive threshold: ; Indicates the first Core support areas index, Indicates the first The final integration score for each core support area;
[0053] The final fusion score for all core opposition regions must be less than a predetermined negative threshold: ; Indicates the first One core area of opposition index, Indicates the first The final integration score for each core opposing region;
[0054] All weights must be non-negative and sum to 1: , ;
[0055] Using a sequential least squares programming algorithm, the optimal weight combination that minimizes the objective function is iteratively solved while satisfying all constraints. ; These are the weights of the optimized position analysis model, the decision model, the time series prediction model, the core region, and the variable region, respectively.
[0056] In one embodiment, the step of evaluating the representativeness of a specific data source for each region and locally adjusting the globally optimal weight based on the difference between the sample proportion and the standard reference proportion specifically includes:
[0057] For each region r, calculate the corresponding percentage of users using social media. and population ratio :
[0058] ;
[0059] ;
[0060] like If the representativeness of social media content is deemed insufficient, the weights of the position analysis model will be partially adjusted.
[0061] In one embodiment, the step of using locally modified weights to perform weighted fusion of the preliminary position prediction results of each sub-model, and outputting the final regional position aggregation measure value, specifically includes:
[0062] Using locally corrected weights, the preliminary predictions of each sub-model, along with the regional positions of the core region and variable regions, are weighted and fused to output the regional... Final regional position aggregation measure :
[0063] ;
[0064] These represent the weights of the locally modified position analysis model, the decision model, the time series prediction model, the core region, and the variable region, respectively. The figures show the preliminary position prediction results for region r using the position analysis model, decision-making model, and time-series prediction model, respectively. As for the regional stance of the core region, when region r is the core support region, then When region r is the core opposing region, then ; .
[0065] In a second aspect, the present invention provides a computer system including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method of any embodiment of the first aspect.
[0066] The core idea of this invention is to construct an integrated system consisting of multiple independent prediction sub-models, and to learn the optimal fusion weights of each sub-model through a constrained optimization algorithm based on the division of core region and variable region. At the same time, a data representativeness correction mechanism is introduced to finally output a high-precision regional position aggregation measurement result.
[0067] Specifically:
[0068] This invention constructs a comprehensive stance measurement framework that can integrate multi-source heterogeneous data and multiple prediction models to overcome the limitations of a single model and improve the accuracy and robustness of predictions.
[0069] This invention dynamically classifies regions based on the historical stability of regional positions and the consistency of model predictions, and uses the classification results as reliable anchors to constrain and optimize the weights of multi-model fusion, thereby improving the reliability of the fusion results.
[0070] This invention establishes a data representativeness correction mechanism to assess the sample bias of data (especially social media data) in different regions, and adaptively adjusts the weights of the corresponding models in the fusion process accordingly, thereby reducing prediction distortion caused by sample bias.
[0071] Compared with existing technologies, this invention, through its systematic multi-model fusion and dynamic correction framework, achieves significant beneficial effects in terms of the accuracy, robustness, adaptability, and data fusion depth of position measurement, specifically reflected in:
[0072] First, accuracy and robustness are significantly improved: by fusing multiple prediction models based on different principles, the systematic bias of a single model is effectively offset. Strong constraint optimization based on the core region ensures the reliability of the fusion results in key areas, making the overall prediction results more robust.
[0073] Second, it achieves deep integration of heterogeneous data: This invention does not simply overlay data, but transforms different data sources into independent prediction models and performs intelligent integration at the model layer, which can better uncover the deep logic behind various types of data.
[0074] Third, it enhances the model's dynamic adaptability: the framework of this method can dynamically re-divide the region and optimize the weights based on the input of new data, enabling it to adapt to the dynamic characteristics of position changes over time.
[0075] Fourth, it effectively overcomes the problem of data sample bias: the original data representativeness correction mechanism can identify and quantify the representativeness of social media data in different regions, and adjust the model weights accordingly, which significantly improves the accuracy of prediction results in sparse or biased regions. Attached Figure Description
[0076] Figure 1 This is a flowchart of the method in an embodiment of the present invention.
[0077] Figure 2 This is a schematic diagram of the overall framework structure in an embodiment of the present invention.
[0078] Figure 3 This is a schematic diagram illustrating the process of a stance analysis model based on social media content in an embodiment of the present invention.
[0079] Figure 4 This is a schematic diagram of the decision-making model based on regional historical characteristics in an embodiment of the present invention. Detailed Implementation
[0080] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.
[0081] like Figure 1 As shown, a variable-grained regional position aggregation measurement method of the present invention includes the following steps:
[0082] S1, acquire and preprocess multi-source heterogeneous data, including social media content, regional economic characteristics, and multi-source position trend data;
[0083] S2 runs multiple independent sub-models in parallel, including a position analysis model based on social media content, a decision-making model based on regional historical economic characteristics, and a time-series prediction model based on network volume fitting and baseline trend correction, to generate preliminary position prediction results for each region.
[0084] S3. Based on the historical stability of the position and the consistency of the position prediction results of each sub-model, the regional dynamics are divided into a core area and a variable area.
[0085] S4 uses the position prediction results of the core area as a hard constraint, and solves the global optimal weights that make the weight distribution of each sub-model, the core area and the variable area most balanced through the optimization algorithm.
[0086] S5 assesses the representativeness of a specific data source for each region and makes local adjustments to the globally optimal weights based on the difference between the sample proportion and the standard reference proportion.
[0087] S6 uses locally modified weights to perform weighted fusion of the preliminary position prediction results of each sub-model as well as the position prediction results of the core area and the variable area, and outputs the final regional position aggregation measure value.
[0088] like Figure 2 As shown, the detailed technical solution of the present invention will be described in the following sections.
[0089] 1. Acquisition and preprocessing of multi-source data.
[0090] The following content was obtained from social media platforms, official statistical databases, and third-party research institutions, and then cleaned, aligned, and standardized to lay the foundation for subsequent modeling:
[0091] Social media content ,in, Representing the first user, represent The published text content, represent Geographic location information;
[0092] Regional economic characteristics ,in, Represents the first area. represent Economic characteristic data, represent Historical regional stance labels;
[0093] Multi-source position trend data ,in, Representing the first event, represent The corresponding position trend time series.
[0094] 2. Construct multidimensional position prediction sub-models in parallel.
[0095] like Figure 3 As shown, three independent prediction sub-models were constructed and run in parallel for three different types of data: a stance analysis model based on social media content, a decision-making model based on regional historical economic characteristics, and a time-series prediction model based on network volume fitting and baseline trend correction. Preliminary stance prediction results were generated from micro, macro, and time-series dimensions. Detailed model descriptions are as follows:
[0096] (1) Stance analysis model based on social media content:
[0097] User-region matching: Utilizes geolocation information from users' social media content to precisely match them to the corresponding region.
[0098] Position classification: Using a Large Language Model (LLM), for each user The published text content undergoes a stance analysis. By designing prompts containing key issue keywords, the LLM (Leadership Management Analyzer) is guided to determine and quantify the stance expressed in the published text content, obtaining a quantified stance value. -1, 0, and 1 represent opposing, neutral, and supporting, respectively.
[0099] ;
[0100] Preliminary aggregation of regional stances: The quantitative values of all user stances within each region are statistically summed to obtain the regional stance prediction results based on social media content, i.e., the preliminary stance prediction results of the stance analysis model for region r. :
[0101] .
[0102] (2) Decision-making model based on regional historical characteristics:
[0103] See Figure 4 Model training: using historical regions Economic characteristics data As input features, they are compared with corresponding historical regional stance labels. Matching is performed to form a training dataset. XGBoost gradient boosting is used to improve the decision tree model. The training process aims to build a powerful ensemble model, which is then iteratively optimized. Specifically, the model starts with a simple initial prediction (usually zero) and trains a new weak learner (such as a decision tree) in each iteration to fit the residuals (i.e., errors) between the predictions of all current weak learners and the actual labels. This process can be represented as:
[0104] .
[0105] in, It is the ensemble model in the kth round. This is the model for round k-1. It is a newly trained weak learner. It is the learning rate.
[0106] Model prediction: The latest regional characteristic data is input into the trained model to obtain regional position prediction results based on macro-historical economic characteristics. .
[0107] (3) Time series prediction model based on network volume fitting and baseline trend correction:
[0108] Stance Acquisition: The core of this stage is to extract target keywords from social media content to capture user preferences and opinions. Social media content and relevant target keywords, combined with pre-defined prompts, are fed into a large model to predict user stances.
[0109] Voice Volume Allocation: In this stage, various tweet attributes related to voice volume are first integrated, such as the number of tweets and user interaction. Then, weights are assigned to these attributes based on actual circumstances to ensure the voice volume allocation more closely reflects reality. Next, based on social media tweets and comments, the Pettitt test is used to identify salient points, which are key nodes in voice volume changes. Based on this, the online voice volume of different users is allocated, with particular attention paid to those allocated a weight greater than 0.5%, to more accurately reflect public opinion dynamics.
[0110] Benchmark trend curve generation: By integrating multiple third-party surveys and market data sources (multi-source position trend data), and employing methods such as distance weighting or minimizing error fitting, a relatively neutral and reliable benchmark position trend time series curve is generated. .
[0111] Network popularity curve correction: Based on each user's stance and their online influence, a time series curve of network popularity is constructed. .
[0112] Time series forecasting: The network popularity time series curve is fused with the baseline position trend time series curve and input into the ARIMA time series forecasting model to obtain the time series forecasting model for the region. Preliminary position forecast results .
[0113] 3. Division of core and variable regions based on position stability.
[0114] Based on the stability of the position in historical data and the consistency of the preliminary predictions of each sub-model, and considering: a) the position has remained consistent and has shown a clear advantage over multiple historical periods; b) the predictions of sub-model one and sub-model two are highly consistent and point to the same position; all regions are dynamically divided into core regions with stable positions that can serve as anchor points, and variable regions with changeable positions that require focused prediction. Core Supporting Region Set Core opposition area set and variable region set The following relationship is satisfied:
[0115] ;
[0116] For the entire region, and These are any two distinct regions; the region partitioning results will be weighted during subsequent model fusion and will serve as a hard constraint.
[0117] 4. Multi-model fusion weight learning based on constraint optimization.
[0118] The position prediction results of the core area are used as inviolable hard constraints. Optimization algorithms such as Sequential Least Squares Programming (SLSQP) are used to iteratively calculate the optimal global fusion weights that maximize the balance of contributions from each sub-model (i.e., minimize the standard deviation of the weights). Detailed steps are as follows:
[0119] Objective function: To ensure a balanced contribution from each sub-model and avoid extreme weight distribution, the optimization objective is set to minimize the standard deviation of the weight values of each sub-model, the core region, and the variable region.
[0120] .
[0121] Apply hard constraints based on the core region partitioning results:
[0122] (1) The final fusion score of all core support regions must be greater than a set positive threshold: ; Indicates the first Core support areas index, Indicates the first The final integration score for each core support area;
[0123] (2) The final fusion score of all core opposition regions must be less than a set negative threshold: ; Indicates the first One core area of opposition index, Indicates the first The final integration score for each core opposing region;
[0124] (3) All weights must be non-negative and sum to 1:
[0125] ;
[0126] .
[0127] The core support zone has a supportive regional stance, the core opposition zone has an opposition regional stance, and the variable zone has a neutral regional stance. When calculating the final fusion score, the regional stance results are numerically represented: 1 represents support, -1 represents opposition, and 0 represents neutrality.
[0128] The final integration score of the region
[0129] ;
[0130] If the region is a core support area, then the region's regional stance... If this area is a core area of opposition, . . The figures represent the position prediction results for the region using the position analysis model, the decision model, and the time series prediction model, respectively. These represent the weights of the position prediction results output by the position analysis model, the position prediction results output by the decision model, the position prediction results output by the time series prediction model, the weights of the regional positions in the core area, and the weights of the regional positions in the variable area, respectively.
[0131] Iterative optimization: Employing mature optimization algorithms such as Sequential Least Squares Programming (SLSQP), the optimal weight combination that minimizes the objective function is iteratively solved while satisfying all constraints. .
[0132] 5. Introduce a data representativeness correction mechanism.
[0133] For each region r, calculate the corresponding percentage of users using social media. and population ratio :
[0134] ;
[0135] ;
[0136] like If the representativeness of social media content is deemed insufficient, the weights of the position analysis model will be partially adjusted.
[0137] 6. Generate the final aggregation measure results.
[0138] The final weights, after global optimization and local correction, are applied to weighted and fused the preliminary predictions of each sub-model, as well as the regional positions of the core region and the variable region, to output the regional position of each region. Final, high-precision position aggregation metric score :
[0139] ;
[0140] These represent the weights of the locally modified position analysis model, the decision model, the time series prediction model, the core region, and the variable region, respectively. The figures show the preliminary position prediction results for region r using the position analysis model, decision-making model, and time-series prediction model, respectively. As for the regional stance of the core region, when region r is the core support region, then When region r is the core opposing region, then ; .
[0141] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0142] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple steps or stages, which are not necessarily completed at the same time, but may be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but may be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0143] In one embodiment, a computer system is provided, which may be a server. The computer system includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data used in the methods described above. The network interface communicates with external terminals via a network connection. The computer program is executed by the processor to implement the methods described above.
[0144] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0145] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.
[0146] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A variable-granularity region-wise aggregate measurement method, comprising: The method comprises the following steps: acquiring and preprocessing multi-source heterogeneous data, including social media content, regional economic characteristics, and multi-source stance trend data; running multiple independent sub-models in parallel, including a stance analysis model based on social media content, a decision model based on regional historical economic characteristics, and a time series prediction model based on network voice fitting and benchmark trend correction, to generate preliminary stance prediction results for each region; dividing the regional dynamics into core regions and variable regions according to historical stance stability and the consistency of the stance prediction results of each sub-model; using the stance prediction results of the core regions as hard constraints, solving the global optimal weights of each sub-model, core regions, and variable regions by an optimization algorithm to make the weight distribution most balanced; for each region, evaluating the representativeness of specific data sources and locally correcting the global optimal weights according to the difference between the sample proportion and the standard reference proportion; using the locally corrected weights to weight and fuse the preliminary stance prediction results of each sub-model and the stance prediction results of the core regions and variable regions, and outputting the final regional stance aggregation measure value.
2. The method of claim 1, wherein, The method of acquiring and preprocessing multi-source heterogeneous data, including social media content, regional economic characteristics, and multi-source stance trend data, specifically comprises: Social media content wherein, representing a first user, representing published text content, representing geographical location information; regional economic characteristics wherein, representing a first region, representing economic characteristic data, representing historical regional stance labels; Multi-source stance trend data wherein, representing a first event, representing a corresponding stance trend time series.
3. The method of claim 1, wherein the method is a variable-granularity regional standing-pole aggregation measurement method. The stance analysis model based on social media content specifically comprises: allocating users in the social media content to corresponding regions using the geographic location information of the users; Using a large language model, the text content published by each user is analyzed to determine the stance of the user . A prompt word containing a specific issue keyword is designed to guide the large language model to determine the stance expressed by the text content and to quantify it . A stance quantification value is obtained . : ; Statistical aggregation is performed on all user stance quantification values in each region to obtain a preliminary stance prediction result of the stance analysis model for the region r : ; is a set of all users within a region is a set of all users within a region 4. The method of claim 1, wherein, The decision model based on regional historical economic characteristics specifically comprises: Use region Historical economic feature data corresponding historical region stance label matching, forming a training data set, training a decision model based on gradient boosting decision tree; inputting the latest regional feature data into the trained decision model obtaining a preliminary stance prediction result of the decision model for the region r .
5. The method of claim 2, wherein, The time series prediction model based on network voice fitting and benchmark trend correction specifically comprises: extracting target words from social media content, combining the text content published by users and related target words with a pre-set prompt template, and inputting them into a large model to obtain a prediction of the user's stance; integrating various published text content attributes related to network voice, assigning weights to these attributes according to actual conditions; based on the published text content and the corresponding comments, using the Pettitt test technique to identify significant points; and assigning network voice to different users according to the attribute weight distribution and significant points; Fusing multi-source position trend data, using distance weighting or minimum error fitting method, a baseline position trend time series curve is generated ; According to the standpoints of each user and the network voice quantity of the user, a network heat time series curve is constructed ; The network heat time series curve is fused with the benchmark stand trend time series curve, and input into an ARIMA model-based time series prediction model In the time series prediction model, the region The preliminary stand prediction result : , The historical time series data related to the region r is represented.
6. The method of claim 1, wherein, The method of dividing the regional dynamics into core regions and variable regions according to historical stance stability and the consistency of the stance prediction results of each sub-model specifically comprises: condition a): the current regional stance has remained consistent in the past set number of historical periods; condition b): the prediction results of the stance analysis model and the decision model for the current regional stance are consistent and point to the same regional stance as condition a); The current region is assigned to the set of core supporting regions if condition a) and condition b) are met and the regional stance is supporting The current region is assigned to the set of core opposing regions if condition a) and condition b) are met and the regional stance is opposing Other regions are assigned to the set of variable regions Core supporting regions and core opposing regions are collectively referred to as core regions.
7. The method of claim 6, wherein the method is a variable-granularity regional standing-pole aggregation measurement method. The method of using the stance prediction results of the core regions as hard constraints, solving the global optimal weights of each sub-model, core regions, and variable regions by an optimization algorithm to make the weight distribution most balanced, specifically comprises: The optimization objective is set to minimize the standard deviation of the weights of each sub-model, core region, and variable region: ; a standard deviation representing the weight vector, a weight vector, an arithmetic mean value of each weight; ; respectively represent a weight of a stance analysis model, a weight of a decision model, a weight of a timing prediction model, a weight of a core region, and a weight of a variable region. constraint condition: The final fusion score of all core support regions must be greater than a set positive threshold: ; denotes the index of the th core support region , denotes the final fusion score of the th core support region; All the core opposing region final fusion scores must be less than a set negative threshold: ; denotes the index of the th core opposing region , denotes the final fusion score of the th core opposing region; All weights must be non-negative and sum to 1: ; Adopting the sequence least square programming algorithm, the optimal weight combination which can minimize the optimization objective function is solved iteratively under the premise of satisfying all constraints ; are the weight of the optimized stand analysis model, the weight of the decision model, the weight of the timing prediction model, the weight of the core area, and the weight of the variable area, respectively.
8. The method of claim 1, wherein, The method of evaluating the representativeness of specific data sources for each region and locally correcting the global optimal weights according to the difference between the sample proportion and the standard reference proportion specifically comprises: For each region r, the corresponding proportion of users using social media is calculated and the population : ; ; If the social media content is considered to be under-representative, the weight of the stance analysis model is locally corrected.
9. The method of claim 1, wherein, The local corrected weight is used to weight and fuse the preliminary position prediction results of each sub-model, and output a final regional position aggregation measure value, and the method specifically comprises the following steps: The preliminary prediction results of each sub-model, the regional stand of the core region and the regional stand of the variable region are fused by using the locally corrected weights, and the regional stand of the core region The final regional stand aggregation measure value : ; respectively represent the weight of the stance analysis model, the weight of the decision model, the weight of the timing prediction model, the weight of the core region, and the weight of the variable region after local correction, respectively represent the preliminary stance prediction results of the stance analysis model, the decision model, and the timing prediction model on the region r, is the region stance of the core region, when the region r is a core support region, then ; when the region r is a core opposition region, then ; .
10. A computer system comprising a memory and a processor, said memory storing a computer program, characterized in that, The processor executes the computer program to realize the steps of the method in any one of claims 1 to 9.
Citation Information
Patent Citations
Standard detection method, system and equipment for providing judgment basis, medium and product
CN118657156A
Automobile quality risk assessment system and method based on novel data fusion technology
CN120181592A