Regional policy dynamic comparison system and method based on large language model

By using a regional policy dynamic comparison system based on a large language model, changes in policy texts are monitored in real time. The system extracts structured features using a large language model and combines Bayesian models and Wilcoxon signed-rank tests to generate visual comparison reports. This solves the problems of low efficiency, strong subjectivity, and difficulty in quantification in traditional policy comparison analysis, and achieves efficient and accurate cross-regional policy analysis.

CN121502382APending Publication Date: 2026-02-10JIANGXI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511691720.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Traditional policy comparative analysis is inefficient, highly subjective, difficult to quantify policy effects, and difficult to conduct precise analysis across regions.

Method used

A regional policy dynamic comparison system based on a large language model is adopted, including a policy radar module, a policy data processing engine, a statistical significance unit, and a visualization comparison module. It generates a visualization comparison report by monitoring policy text changes in real time, extracting structured features using a large language model, quantifying the effect using a Bayesian model, using Wilcoxon signed-rank test for significance labeling, and optimizing prompt words using deep learning.

Benefits of technology

It enables efficient and accurate cross-regional policy analysis, significantly improves the extraction accuracy of key policy clauses and implicit conditions, quantifies policy effects and generates intuitive decision-making basis, and helps users identify policy implementation gaps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502382A_ABST
    Figure CN121502382A_ABST
Patent Text Reader

Abstract

The invention provides a regional policy dynamic comparison system and method based on a large language model, and the method comprises the steps: monitoring a government official website and a white list site in real time, and automatically triggering collection when detecting that the policy text change frequency exceeds a preset threshold value; using a large language model to extract policy structured features according to the dynamic cue words; accessing a government open database, and quantifying the actual effect of the policy through a Bayesian structure time sequence model; calculating the statistical significance of the policy term difference through a Wilcoxon symbol rank test, and carrying out statistical significance labeling; the cue word is dynamically optimized through a deep deterministic strategy gradient reinforcement learning agent; and generating a visual comparison report containing statistical saliency labels and a deviation degree comparison matrix. According to the invention, automatic analysis and dynamic comparison of policies in each region are realized, and the problem that cross-region policy analysis is difficult to quantify, verify and make decisions is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of regional policy comparison, and in particular to a regional policy dynamic comparison system and method based on a large language model. BACKGROUND

[0002] Under the background of regional economic coordinated development, differentiated policies are introduced to attract industrial resources. Current policy information is scattered on government websites, local government platforms, and third-party databases, and presents a multi-source heterogeneous characteristic (PDF / Word / web text). With the acceleration of policy update frequency (the average policy revision rate in the Yangtze River Delta in 2023 reached 37%), the traditional manual comparison mode has been difficult to meet the efficient and accurate cross-regional policy analysis needs.

[0003] Traditional policy comparison analysis usually relies on manual work, and often has the following defects:

[0004] Low efficiency: due to the scattered storage of policy texts, manual collection needs to repeatedly deal with different platforms' data interfaces and formats; key clauses (such as subsidy amounts, applicable objects) rely on manual word-by-word analysis, and when facing more than 50 pages of documents, the information omission rate is more than 28%; lacking an automatic policy change capturing mechanism, unable to respond to high-frequency revisions;

[0005] Strong subjectivity: the same policy clause analyzed by different personnel has significant differences; some ambiguous expressions (such as "give appropriate subsidies") lead to the lack of quantitative standards; multi-clause nested logic (such as "meet A and B or C") leads to an increased misjudgment rate;

[0006] Shallow comparison: only surface clause differences (such as subsidy values) can be extracted, and some implicit factors such as regional policy motivation, implementation cost, and policy effect are easily overlooked. SUMMARY

[0007] Therefore, the purpose of the present application is to provide a regional policy dynamic comparison system and method based on a large language model to solve the above-mentioned problems.

[0008] The regional policy dynamic comparison system based on a large language model according to the present application comprises:

[0009] A policy radar module for real-time monitoring of government websites and white-listed sites, and automatically triggering collection when detecting that the policy text change frequency exceeds a preset threshold;

[0010] A policy data processing engine comprising:

[0011] A text analysis unit for extracting policy structured features using a large language model according to dynamically optimized prompt words;

[0012] An effect verification unit is configured to access a government open database and quantify actual effects of a policy by using a Bayesian structural time series model;

[0013] A statistical significance unit is configured to calculate statistical significance of differences in policy provisions by using a Wilcoxon signed-rank test and mark the statistical significance;

[0014] A prompt word optimization module is configured to dynamically optimize prompt words by using a deep deterministic policy gradient reinforcement learning agent;

[0015] and a visual comparison module is configured to generate a visual comparison report and a deviation comparison matrix containing statistical significance marks.

[0016] The application further provides a regional policy dynamic comparison method based on a large language model, which is used to implement the above-mentioned regional policy dynamic comparison system based on a large language model.

[0017] Real-time monitoring of government official websites and whitelist sites is performed, and collection is automatically triggered when it is detected that the frequency of policy text changes exceeds a preset threshold;

[0018] Structured features of the policy are extracted according to dynamic prompt words by using a large language model;

[0019] Access is made to a government open database, and actual effects of the policy are quantified by using a Bayesian structural time series model;

[0020] Statistical significance of differences in policy provisions is calculated by using a Wilcoxon signed-rank test, and statistical significance marks are marked;

[0021] Prompt words are dynamically optimized by using a deep deterministic policy gradient reinforcement learning agent;

[0022] A visual comparison report and a deviation comparison matrix containing statistical significance marks are generated.

[0023] Furthermore, the real-time monitoring of government official websites and whitelist sites, when it is detected that the frequency of policy text changes exceeds a preset threshold, automatically triggers collection, includes:

[0024] Verification is made on whether a policy publishing agency belongs to a predefined administrative agency coding list;

[0025] A Hamming distance of the policy text is calculated by using a local sensitive hashing algorithm, and a version change alarm is generated when the distance exceeds 5%, and collection of policy information after the change is automatically triggered.

[0026] Furthermore, the structured features of the policy are extracted according to dynamic prompt words by using a large language model, which includes:

[0027] Based on the policy type selected by the user, generate a prompt word template Φ(T) containing dynamic placeholders;

[0028] Compare the core policy paragraphs with similar reference policy summaries Inject the prompt template Φ(T) to obtain the prompt:

[0029] Prompt word = Φ(T) + "Policy text: " + Substr(Policy text, L) + "Reference case: " + Where L is the truncation length of the core policy paragraph;

[0030] Input the prompt words into the large language model and obtain the raw output. ;

[0031] Validate the original output using a JSON-Schema validator. Does it conform to the predefined structure?

[0032] ={type: "object",required: ["Applicable Object", "Key Indicators", "Validity Period"],…};

[0033] If the verification is successful, the original output will be accepted. As extracted policy structural features;

[0034] If the verification fails, the error handling process is triggered: discard the original output. And regenerate the prompt words.

[0035] Furthermore, the access to the government's open database and the quantification of the actual policy effects using a Bayesian structured time series model include:

[0036] Retrieve datasets related to policy indicators from government open databases;

[0037] Calculate semantic matching degree: ,in, The matching degree coefficient, For policy indicators Keyword set, In line with policy indicators Related datasets A collection of tags;

[0038] Constructing a Bayesian structured time series model: ,in, Let be the observed value in year t. For local linear trend terms, For the Fourier seasonal term, The policy intervention coefficient. This is a policy intervention variable; it is 1 if implemented, and 0 otherwise. For random error term, The standard deviation of noise;

[0039] Sampling iterations N times, estimating the posterior distribution ;

[0040] Calculate the policy contribution rate: , ,in, for The posterior mean estimate, This is the average of observations before the policy was implemented. This refers to the number of years prior to the policy's implementation.

[0041] when If the 95% confidence interval does not include 0, then the policy is considered valid.

[0042] Furthermore, the calculation of the statistical significance of policy clause differences using the Wilcoxon signed-rank test, and the subsequent annotation of statistical significance, includes:

[0043] For each policy indicator Calculate the region pairing difference: Where k = 1, 2, ..., t, and t is the number of samples in each year. For policy indicators The actual value of k in year a in region a. For policy indicators The actual value of k in year b in region b;

[0044] The absolute value of the pairing difference by region | Sort in ascending order and assign ranks;

[0045] Calculate the rank sum of the positive differences respectively Rank sum of negative differences ;

[0046] Take W=min( , ) as a statistic;

[0047] Determine the p-value by looking up the table based on the sample size n;

[0048] Significance is marked based on p-value.

[0049] Furthermore, the method of dynamically optimizing prompt words through deep deterministic gradient reinforcement learning agents includes:

[0050] Initial prompts are generated using an Actor network with a deep deterministic policy gradient proxy. And input it into the large language model;

[0051] Calculate the quality score of the prompt words based on the output of the large language model: =α·accuracy + β·coverage - γ·redundancy, where accuracy is the accuracy of the key points output by the large language model, coverage is the coverage of policy clauses, and redundancy is the output redundancy;

[0052] Generate optimization prompts using policy networks: ,in, The basic prompt words generated for the Actor network, For noise;

[0053] Based on the quality score of the prompt words Update the Critic network weights.

[0054] Furthermore, the generation of a visual comparison report and a deviation comparison matrix containing statistical significance annotations includes:

[0055] Based on the enterprise attribute parameters input by the user, the policy adaptability of each region is calculated. The formula is: Adaptability = Σ(Policy clause weight × Enterprise matching degree) × Regional competitiveness coefficient.

[0056] Generate a policy evolution decision tree diagram that includes the policy adaptability of each region, where the color depth of the nodes represents the strength of policy effectiveness.

[0057] Furthermore, the generation of a visual comparison report and a deviation comparison matrix containing statistical significance annotations also includes:

[0058] Receive structured policy clause data extracted from a large language model, including subsidy amount, applicable targets, and validity period;

[0059] Receives Bayesian model output based on open government data, including actual policy realization value and contribution rate indicators;

[0060] In the comparison matrix, each policy indicator is assigned a separate row, and each comparison area is assigned a separate column. The matrix also contains three types of fields: clause value, actual value, and deviation.

[0061] The policy text values ​​extracted by the large language model are filled into the clause value field, the average annual actual realization value calculated by the Bayesian model is filled into the actual value field, and the absolute difference between the clause value and the actual value is calculated and filled into the deviation field.

[0062] In summary, the regional policy dynamic comparison method based on a large language model of this invention accurately captures and automatically extracts dynamic updates of policies across multiple regions through real-time monitoring and intelligent data collection mechanisms. It utilizes deep parsing technology based on the large language model to structurally extract core clauses of policies in each region, and employs dynamic prompts to inject policy type adaptation templates, context fragments, and JSON-Schema validation, significantly improving the extraction accuracy of key policy clauses and implicit conditions. It quantifies the actual effects of policies through a Bayesian structured time series model and uses the Wilcoxon signed-rank test to compare differences in the implementation of policy clauses between regions. Through reinforcement learning-driven prompt optimization, the prompt strategy is dynamically adjusted based on the accuracy of policy parsing, enabling the model to improve parsing accuracy through continuous interaction. Finally, it generates a heatmap and deviation matrix with integrated statistical annotations, intuitively presenting the differences in policies across regions, assisting users in quickly identifying policy implementation gaps, and providing users with comprehensive and accurate decision-making basis. This invention significantly improves the automation and scientific rigor of regional policy analysis and dynamic comparison, solving the difficulties of quantifying, verifying, and deciding on cross-regional policies in traditional analysis.

[0063] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by means of embodiments of the invention. Attached Figure Description

[0064] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0065] Figure 1 This is a flowchart of the regional policy dynamic comparison method based on a large language model according to Embodiment 1 of the present invention;

[0066] Figure 2 This is a system block diagram of the regional policy dynamic comparison system based on a large language model according to Embodiment 2 of the present invention;

[0067] Figure 3 This is a structural diagram of the policy data processing engine in Embodiment 2 of the present invention. Detailed Implementation

[0068] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.

[0069] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.

[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0071] Example 1

[0072] Please see Figure 1 This invention proposes a method for dynamic comparison of regional policies based on a large language model, which includes steps S101 to S106:

[0073] S101 monitors government websites and whitelisted sites in real time, and automatically triggers data collection when the frequency of policy text changes exceeds a preset threshold.

[0074] Further optional, the real-time monitoring of government websites and whitelisted sites, automatically triggering data collection when the frequency of policy text changes exceeds a preset threshold, includes:

[0075] Verify whether the policy-issuing agency belongs to the predefined list of administrative agency codes;

[0076] The Hamming distance of the policy text is calculated using the Locality Sensitive Hash algorithm. When the distance exceeds 5%, a version change alarm is generated, which automatically triggers the collection of policy information for the changed version.

[0077] Understandably, during the data collection phase, authoritative institutional verification and change detection algorithms ensure the legality and timeliness of the data source, ultimately guaranteeing that the policy source is reliable and the version is up-to-date. The 5% Hamming distance threshold is a balanced value derived from extensive testing, avoiding oversensitivity while effectively capturing changes.

[0078] Specifically, when verifying the authority of an institution, the process begins by constructing a predefined administrative agency code database. This database, a structured database, stores national administrative division codes (e.g., 110000 represents Beijing) and their corresponding lists of legitimate policy-issuing agencies, such as the Beijing Municipal Human Resources and Social Security Bureau and the Shanghai Municipal Commission of Economy and Information Technology. The database employs a two-level index architecture: the first level contains the administrative division codes, and the second level lists the names of authoritative agencies within that region. Then, the administrative division codes are automatically parsed from the policy information source URL (e.g., code 310000 is extracted from the domain "sh.gov.cn"). The name of the policy-issuing agency is then matched against the corresponding regional agency list in the code database. If the agency name exists in the pre-stored list for that region, the verification passes; otherwise, the data collection process terminates and an illegal source alert is recorded. This ultimately solves the verification problem of documents issued by different agencies under the same domain, preventing third-party platforms from forging policies.

[0079] When detecting version changes, the policy text is segmented to generate a 256-bit MinHash fingerprint. A sliding window is used to process long texts, ensuring that local features are fully captured. Then, the Hamming distance between the old and new policy texts, i.e., the difference ratio, is calculated using a Locality Sensitive Hashing algorithm: \text{Distance} = 1 - \frac{\text{Number of identical hash bits}}{256}. When the distance exceeds a 5% threshold (5% being the optimal value verified through large-scale testing), a structured alarm log is automatically generated, including a change location marker. A differential algorithm is then used to locate the specific modified chapter, such as "Chapter 3 Identification Conditions," re-collecting only the modified chapter rather than the entire text, reducing bandwidth consumption. Ultimately, this approach overcomes the limitations of traditional keyword matching, accurately identifying implicit changes such as synonym replacements and clause order adjustments.

[0080] S102 uses a large language model to extract policy structure features based on dynamic prompts.

[0081] Further, optionally, the extraction of policy structured features based on dynamic prompt words using a large language model includes:

[0082] Based on the policy type selected by the user, such as talent policy or tax policy, generate a prompt word template Φ(T) containing dynamic placeholders;

[0083] The prompt template for talent policies is: "Extract talent policy characteristics: [Applicable targets], [Subsidy amount], [Application conditions], [Validity period]";

[0084] The prompt template for tax policies is: "Extract tax policy characteristics: [Applicable enterprises], [tax reduction / exemption rate], [application materials], [validity period]";

[0085] Compare the core policy paragraphs (excluding the first 300 characters, L=300) with similar reference policy summaries. Inject the prompt template Φ(T) to obtain the prompt:

[0086] Prompt word = Φ(T) + "Policy text: " + Substr(Policy text, L) + "Reference case: " + Where L is the cut length, which is a fixed value;

[0087] Input the prompt words into the large language model and obtain the raw output. ;

[0088] Validate the original output using a JSON-Schema validator. Does it conform to the predefined structure?

[0089] ={type: "object",required: ["Applicable Object", "Key Indicators", "Validity Period"],…};

[0090] If the verification is successful, the original output will be accepted. As extracted policy structural features;

[0091] If the verification fails, the error handling process is triggered: discard the original output. And regenerate the prompt words.

[0092] Understandably, combining dynamic templates to adapt to policy types, context injection to enhance semantic understanding, and JSON validation to ensure data structuring can effectively address the accuracy issues in automatic policy text parsing, thereby significantly improving the precision and reliability of policy feature extraction. Specifically, by using dynamic placeholders to adapt to multiple policy types (such as talent / taxation), injecting core paragraphs and reference cases to resist ambiguous expressions, employing JSON-Schema mandatory validation to eliminate unstructured output, and using an error regeneration mechanism to ensure result usability, the extraction fluctuations caused by the complexity of policy texts are ultimately resolved.

[0093] S103 connects to the government's open database and uses a Bayesian structured time series model to quantify the actual effects of policies.

[0094] Further optional, the access to the government's open database and the quantification of the actual policy effects using a Bayesian structured time series model include:

[0095] Retrieve datasets related to policy indicators from government open databases;

[0096] Calculate semantic matching degree: ,in, The matching coefficient is used only when η > 0.8, based on the dataset. , For policy indicators Keyword set, In line with policy indicators Related datasets A collection of tags;

[0097] Constructing a Bayesian structured time series model: , , , ,

[0098] in, Let be the observed value in year t. For local linear trend terms, For the Fourier seasonal term, The policy intervention coefficient. This is a policy intervention variable; it is 1 if implemented, and 0 otherwise. For random error term, The standard deviation of noise;

[0099] Estimate the posterior distribution by sampling N times. ;

[0100] Calculate the policy contribution rate: , ,in, for The posterior mean estimate, This is the average of observations before the policy was implemented. This refers to the number of years prior to the policy's implementation.

[0101] when If the 95% confidence interval does not include 0, then the policy is considered valid.

[0102] Understandably, the first step is to retrieve datasets related to policy indicators (such as enterprise subsidy application records and tax reduction ledgers) from government open databases. Semantic matching degree calculations (used when η>0.8) ensure a strong correlation between the data and policy objectives. The matching degree coefficient is based on the intersection ratio of keyword sets (such as R&D expense deduction and high-tech enterprise certification) and data tags, filtering out unstructured interference data such as meeting notices and budget preparation instructions to ensure the accuracy of the analysis data.

[0103] Then, a Bayesian structured time series model is constructed, in which the local linear trend term is used to capture the long-term natural growth trend of the number of enterprises and economic indicators, the Fourier seasonal term is used to fit the cyclical fluctuations such as the quarterly reporting peak and the year-end fiscal settlement, and the policy intervention coefficient is used to quantify the policy impact effect through a dummy variable (after implementation = 1) to separate the policy-specific contribution value.

[0104] Then, through Markov chain Monte Carlo sampling iterations N times (usually N>10000), the model generates the posterior distribution of the policy contribution rate. Instead of a single-point estimate, this probabilistic representation can show the confidence interval of the policy effect (e.g., a 95% probability that it falls between 2.8% and 7.2%), providing a quantitative basis for decision-making.

[0105] Then through Calculate the policy contribution rate, where, This is the average of observations prior to implementation.

[0106] when When the 95% confidence interval does not include 0 (e.g., [2.8%, 7.2%] are all positive), the system automatically marks the policy as valid. For example, the confidence interval for the interest subsidy policy for first-time borrowers in a certain region is [3.1%, 5.9%], which significantly promotes financing for micro and small enterprises; while the confidence interval for a similar policy in another region includes 0 values ​​(e.g., [-1.2%, 4.3%]), further auditing of the implementation process is required.

[0107] This embodiment solves the problem of policy effect attribution by constructing a Bayesian model. It separates the policy impact by intervening variables, and the matching degree threshold (such as 0.8) can effectively filter irrelevant data. Finally, the contribution rate is used to quantify the policy effect, realize the scientific evaluation of the policy effect, and overcome the limitation of traditional statistical models in handling small sample government data.

[0108] S104. The statistical significance of the differences in policy clauses is calculated using the Wilcoxon signed-rank test, and the statistical significance is marked.

[0109] Further optionally, the calculation of the statistical significance of the policy clause differences using the Wilcoxon signed-rank test, and the subsequent annotation of statistical significance, includes:

[0110] For each policy indicator Calculate the region pairing difference: Where k = 1, 2, ..., t, and t is the sample size for each year, i.e., the total number of years included in the comparison. For example, if comparing data from 2015 to 2024, then t = 10. For policy indicators The actual value of k in year a in region a. For policy indicators The actual value of k in year b in region b;

[0111] The absolute value of the pairing difference by region | Sort in ascending order and assign ranks (ignore zero differences);

[0112] Calculate the rank sum of the positive differences respectively Rank sum of negative differences , , ,in, For | |Rank of ascending order;

[0113] Take W=min( , ) as a statistic;

[0114] Determine the p-value by looking up the table based on the sample size n;

[0115] If n≤20, then look up the p value in the Wilcoxon critical value table;

[0116] If n>20, then the value of p can be obtained by approximation using a normal distribution: ,in, , , As the expected value, Standard deviation;

[0117] The significance level is marked based on the p-value, and the marking rules are as follows:

[0118] ,

[0119] in, This indicates a highly significant difference (confidence level greater than 99%). Indicates a significant difference (confidence level greater than 95%). This indicates a weakly significant difference (confidence level greater than 90%). An empty label indicates no significant difference.

[0120] Understandably, for each policy indicator, a time-aligned paired sample is formed by constructing a series of historical actual values ​​for the corresponding policy indicator (such as the subsidy amount for high-tech enterprises) for different regions. For example, when comparing policies in the Yangtze River Delta and the Pearl River Delta, the annual difference of this indicator between the two regions from 2015 to 2024 is calculated, and t sets of observation data are established.

[0121] When sorting by difference, the system implicitly encodes region attributes, allowing positive and negative ranks to directly reflect the superiority or inferiority relationship between regions. For example, when the difference is +5 (region a is superior to region b), this rank will be included in the positive rank sum. The final statistic W = min( , Directly quantify the strength of the advantage of region a relative to region b.

[0122] Based on the statistic W and p-value, a quantitative criterion for determining the significance of differences is formed. When the sample size n≤20, a critical value table is used for accurate inference; when n>20, a normal approximation is used to ensure computational efficiency, ensuring that confidence labels (e.g., p<0.01 label) can be output in comparisons of regions of different sizes. (Indicating highly significant differences). The significance labeling rule is essentially a visual encoding of regional policy competitiveness. For example, in a policy evaluation of a certain district, if the R&D expense deduction ratio indicator is labeled... This indicates that the policy effect in this region is significantly better than that in its comparison regions, while the talent housing subsidy is marked as... If the value is not significant, it indicates that there is no significant difference between regions, and further analysis of the execution details is required. The higher the significance (e.g., ... The higher the statistical significance of the difference between the actual value realized by the policy provisions and the value promised in the policy document, the higher the consistency between the policy implementation effect and the commitment. For enterprise users, regions with high significance should be selected, as this means that the policy implementation is more reliable and the policy benefits are more easily realized.

[0123] This scheme supports not only dual-region comparisons but also N-region network comparisons. A complete map of policy competitiveness can be generated by constructing inter-regional difference matrices (such as 435 paired comparisons of 31 provincial-level administrative regions).

[0124] This embodiment, through the deep integration of statistical testing and regional coding, not only enables cross-regional quantitative comparison of policy effects, but also constructs an analytical link from micro-level differences in clauses to macro-level regional competitiveness, effectively solving the core problems of isolated regional analysis and lack of horizontal reference in traditional policy evaluation.

[0125] S105 uses deep deterministic gradient reinforcement learning to dynamically optimize prompt words for the agent.

[0126] Further optionally, the step of dynamically optimizing prompt words through a deep deterministic strategy gradient reinforcement learning agent includes:

[0127] Initial prompts are generated using an Actor network with a deep deterministic policy gradient proxy. And input it into the large language model;

[0128] Calculate the quality score of the prompt words based on the output of the large language model: =α·accuracy + β·coverage - γ·redundancy, where accuracy is the accuracy of the key points output by the large language model (0~1), coverage is the coverage of policy clauses (0~1), and redundancy is the output redundancy (0~1).

[0129] Generate optimization prompts using policy networks: ,in, The basic prompt words generated for the Actor network, For noise, It is a state vector that includes a task description and historical results. For policy network weights;

[0130] Based on the quality score of the prompt words The weights of the Critic network are updated using gradient descent.

[0131] Understandably, the Actor network generates initial prompts based on the current state vector (which includes the task description and historical optimization results), while the Critic network evaluates the quality scores of the prompts. =α·Accuracy + β·Coverage - γ·Redundancy provides value feedback, forming the driving signal for policy gradient updates. Accuracy weights ensure the precision of policy clause parsing, coverage weights guarantee the depth of implicit condition discovery, and redundancy is used as a penalty to suppress overgeneralization. A weight configuration of α=0.7, β=0.3, and γ=0.1 can significantly improve the extraction rate of key clauses. Finally, noise is injected into the policy exploration process to introduce relevant randomness and avoid getting trapped in local optima through temporary perturbations. For example, when optimizing tax policy prompts, noise can help the model discover the implicit correlation between the R&D expense deduction ratio and the recognition of high-tech enterprises.

[0132] This embodiment uses a trial-and-error feedback loop of reinforcement learning to dynamically optimize prompt words, enabling the prompt word generation strategy to have environmental adaptability. It can automatically adapt to different policy types (such as the semantic differences between talent policies and environmental protection policies) and parsing objectives (such as comparison of key clauses or prediction of impact), significantly reducing the cost of manual parameter tuning while improving the policy parsing accuracy of the large language model.

[0133] S106 generates a visual comparison report and a deviation comparison matrix that include statistical significance annotations.

[0134] Understandably, in visualization reports of significance annotation results based on the Wilcoxon test, users should prioritize selecting annotations. (This refers to the region where p < 0.01). This is because... The marker indicates a highly significant difference between the actual realized value of a policy provision and the value promised in the policy document, with a confidence level greater than 99%. This means that in these regions, the consistency between policy implementation and commitment is extremely high, and enterprises can more reliably obtain the expected policy benefits. For example, if a region's talent subsidy policy is marked as... Therefore, the actual amount of subsidies issued in the region is likely to be very close to the amount promised in the policy documents, making it more secure for companies to choose to set up operations in the region.

[0135] In contrast, labeling (p<0.05) and The region (p<0.1), while indicating a significant difference between the policy implementation effect and the promised value, shows a gradually decreasing confidence level. Therefore, users can prioritize regions with this label when making their selections. In order to avoid the risk of inadequate policy implementation in certain areas and ensure that the benefits brought by the policies can be fully enjoyed.

[0136] Further optionally, the generation of a visual comparison report and a deviation comparison matrix containing statistical significance annotations includes:

[0137] Based on the enterprise attribute parameters input by the user, the policy adaptability of each region is calculated. The formula is: Adaptability = Σ(Policy clause weight × Enterprise matching degree) × Regional competitiveness coefficient.

[0138] Generate a policy evolution decision tree diagram, where the color depth of the nodes represents the strength of policy effectiveness.

[0139] Specifically, after receiving enterprise attribute parameters, the policy suitability of region a is calculated using the following formula:

[0140] ,

[0141] Where E is the enterprise attribute vector. (e.g., size, industry, qualifications), C represents the set of policy clauses. (such as subsidy amount, tax rate). For the terms The weighting is as follows: key clauses have a weight of 0.4, and ordinary clauses have a weight of 0.3. This is the clause matching degree function (values ​​from 0 to 1; the output is 1 when the enterprise fully meets the clause requirements, and it is calculated according to the feature coverage ratio when there is a partial match, i.e., the ratio of the number of matched features to the total number of features). Let be the competitiveness coefficient of region a. .

[0142] Generate a policy evolution decision tree diagram, where the node types include root nodes, branch nodes, and leaf nodes. The root node is the enterprise attribute selection, such as "size = large", the branch node is the priority dimension selection, such as "focus on subsidy intensity?", and the leaf node represents the recommended region and its suitability value.

[0143] The color depth of leaf nodes is encoded using the following formula: Where s is the policy effectiveness score, which is derived from the analysis of policy influence by the large language model, ranging from 0 to 10 points; λ=0.8 is the Sigmoid scaling factor, which is a fixed parameter.

[0144] Output RGB color values: (0, 0, \lfloor 255 \times \text{color depth} \rfloor) (blue gradient).

[0145] Understandably, the system receives characteristic parameters such as enterprise size, industry, and qualifications (e.g., the "Specialized, Refined, and Innovative Small and Medium-sized Enterprises" identifier) ​​and generates policy matching scores for each region through a policy fit formula.

[0146] The generated tree diagram contains three types of nodes: the root node is a company attribute filter (e.g., annual revenue > 500 million), the branch nodes are policy priority selectors (e.g., priority subsidy intensity), and the leaf nodes display recommended regions and their suitability. The node color depth maps the policy effectiveness score (0-10 points) to an RGB blue gradient using the Sigmoid function (e.g., effectiveness score of 8 points corresponds to RGB(0, 0, 204)), enabling companies to quickly identify high-value policy combinations.

[0147] The report integrates significance markers from the Wilcoxon test and uses heatmaps to compare the deviations between the fulfilled and committed values ​​of policy provisions. For example, if a region's R&D subsidy provisions are marked with significance markers... Furthermore, the decision tree tree nodes are displayed in dark blue, allowing users to prioritize that area.

[0148] Further, optionally, the generation of the visual comparison report and deviation comparison matrix containing statistical significance annotations further includes:

[0149] Receive structured policy clause data extracted from a large language model, including subsidy amount, applicable targets, and validity period;

[0150] Receives Bayesian model output based on open government data, including actual policy realization value and contribution rate indicators;

[0151] In the comparison matrix, each policy indicator is assigned a separate row, and each comparison area is assigned a separate column. The matrix also contains three types of fields: clause value, actual value, and deviation.

[0152] The policy text values ​​extracted by the large language model are filled into the clause value field, the average annual actual realization value calculated by the Bayesian model is filled into the actual value field, and the absolute difference between the clause value and the actual value is calculated and filled into the deviation field.

[0153] Understandably, the system first receives structured clause data extracted from policy texts by a large language model (e.g., "Postdoctoral research funding subsidy of 300,000 yuan / person" is parsed into JSON format {"Subsidy type":"Postdoctoral","Amount":30,"Unit":"Ten thousand yuan"}); simultaneously integrating the actual realized value and policy contribution rate indicators calculated by the Bayesian model based on the government's open database, where the actual realized value = ×(1+ ), This is the average of observations before the policy was implemented. Policy contribution rate, policy contribution rate =5%, ,in, for The posterior mean estimate, This is the average of observations before the policy was implemented. The number of years prior to the policy's implementation. Let t be the observed value in year t;

[0154] The design employs a matrix-style visualization framework of regions and indicators. Each policy provision (such as subsidy amount) is displayed in a separate row, and each region to be evaluated is displayed in a separate column, forming a three-dimensional data structure that includes the provision value, the actual value, and the degree of deviation. This allows for the rapid identification of differences in regional policy competitiveness (e.g., a region may offer substantial subsidies but have a low fulfillment rate, while another region may make conservative promises but implement them effectively).

[0155] A deviation field is generated by calculating the absolute difference between the clause value and the actual value, and statistical significance is marked using a Wilcoxon test, providing users with comprehensive and accurate decision-making support. Users can use this to avoid areas where promises are made but not kept (e.g., deviation > 20% and marked with a negative value). (Through these provisions), policymakers can also identify implementation weaknesses (such as policy provisions with a contribution rate of less than 10%) and optimize policies accordingly.

[0156] In summary, the regional policy dynamic comparison method based on a large language model of this invention accurately captures and automatically extracts dynamic updates of policies across multiple regions through real-time monitoring and intelligent data collection mechanisms. It utilizes deep parsing technology based on the large language model to structurally extract core clauses of policies in each region, and employs dynamic prompts to inject policy type adaptation templates, context fragments, and JSON-Schema validation, significantly improving the extraction accuracy of key policy clauses and implicit conditions. It quantifies the actual effects of policies through a Bayesian structured time series model and uses the Wilcoxon signed-rank test to compare differences in the implementation of policy clauses between regions. Through reinforcement learning-driven prompt optimization, the prompt strategy is dynamically adjusted based on the accuracy of policy parsing, enabling the model to improve parsing accuracy through continuous interaction. Finally, it generates a heatmap and deviation matrix with integrated statistical annotations, intuitively presenting the differences in policies across regions, assisting users in quickly identifying policy implementation gaps, and providing users with comprehensive and accurate decision-making basis. This invention significantly improves the automation and scientific rigor of regional policy analysis and dynamic comparison, solving the difficulties of quantifying, verifying, and deciding on cross-regional policies in traditional analysis.

[0157] Example 2

[0158] Please see Figure 2 andFigure 3 The present invention proposes a regional policy dynamic comparison system based on a large language model, the system comprising:

[0159] The policy radar module is used to monitor government websites and whitelisted sites in real time. It automatically triggers data collection when the frequency of policy text changes exceeds a preset threshold.

[0160] The policy data processing engine includes:

[0161] The text parsing unit is used to extract policy structure features based on dynamically optimized prompt words using a large language model;

[0162] The effect verification unit is used to access the government's open database and quantify the actual effect of policies through a Bayesian structured time series model.

[0163] The statistical significance unit is used to calculate the statistical significance of policy clause differences using the Wilcoxon signed-rank test and to label the statistical significance.

[0164] The prompt word optimization module is used to dynamically optimize prompt words through a deep deterministic policy gradient reinforcement learning agent;

[0165] It also includes a visualization comparison module for generating visualization comparison reports and deviation comparison matrices that include statistical significance annotations.

[0166] Further optionally, the policy radar module is also used for:

[0167] Verify whether the policy-issuing agency belongs to the predefined list of administrative agency codes;

[0168] The Hamming distance of the policy text is calculated using the Locality Sensitive Hash algorithm. When the distance exceeds 5%, a version change alarm is generated, which automatically triggers the collection of policy information for the changed version.

[0169] Further optionally, the text parsing unit is also used for:

[0170] Based on the policy type selected by the user, generate a prompt word template Φ(T) containing dynamic placeholders;

[0171] Compare the core policy paragraphs with similar reference policy summaries Inject the prompt template Φ(T) to obtain the prompt:

[0172] Prompt word = Φ(T) + "Policy text: " + Substr(Policy text, L) + "Reference case: " + Where L is the truncation length of the core policy paragraph;

[0173] Input the prompt words into the large language model and obtain the raw output. ;

[0174] Validate the original output using a JSON-Schema validator. Does it conform to the predefined structure?

[0175] ={type: "object",required: ["Applicable Object", "Key Indicators", "Validity Period"],…};

[0176] If the verification is successful, the original output will be accepted. As extracted policy structural features;

[0177] If the verification fails, the error handling process is triggered: discard the original output. And regenerate the prompt words.

[0178] Further optionally, the effect verification unit is also used for:

[0179] Retrieve datasets related to policy indicators from government open databases;

[0180] Calculate semantic matching degree: ,in, The matching degree coefficient, For policy indicators Keyword set, In line with policy indicators Related datasets A collection of tags;

[0181] Constructing a Bayesian structured time series model: ,in, Let be the observed value in year t. For local linear trend terms, For the Fourier seasonal term, The policy intervention coefficient. This is a policy intervention variable; it is 1 if implemented, and 0 otherwise. For random error term, The standard deviation of noise;

[0182] Sampling iterations N times, estimating the posterior distribution ;

[0183] Calculate the policy contribution rate: , ,in, for The posterior mean estimate, This is the average of observations before the policy was implemented. This refers to the number of years prior to the policy's implementation.

[0184] when If the 95% confidence interval does not include 0, then the policy is considered valid.

[0185] Further optionally, the statistical significance unit is also used for:

[0186] For each policy indicator Calculate the region pairing difference: Where k = 1, 2, ..., t, and t is the number of samples in each year. For policy indicators The actual value of k in year a in region a. For policy indicators The actual value of k in year b in region b;

[0187] The absolute value of the pairing difference by region | Sort in ascending order and assign ranks;

[0188] Calculate the rank sum of the positive differences respectively Rank sum of negative differences ;

[0189] Take W=min( , ) as a statistic;

[0190] Determine the p-value by looking up the table based on the sample size n;

[0191] Significance is marked based on p-value.

[0192] Further optionally, the prompt word optimization module is also used for:

[0193] Initial prompts are generated using an Actor network with a deep deterministic policy gradient proxy. And input it into the large language model;

[0194] Calculate the quality score of the prompt words based on the output of the large language model: =α·Accuracy + β·Coverage - γ·Redundancy, where accuracy is the accuracy of key points output by the large language model, coverage is the coverage of policy clauses, and redundancy is the output redundancy; optimized prompt words are generated using a policy network. ,in, The basic prompt words generated for the Actor network, For noise;

[0195] Based on the quality score of the prompt words Update the Critic network weights.

[0196] Further optionally, the visualization comparison module is also used for:

[0197] Based on the enterprise attribute parameters input by the user, the policy adaptability of each region is calculated. The formula is: Adaptability = Σ(Policy clause weight × Enterprise matching degree) × Regional competitiveness coefficient.

[0198] Generate a policy evolution decision tree diagram that includes the policy adaptability of each region, where the color depth of the nodes represents the strength of policy effectiveness.

[0199] Further optionally, the visualization comparison module is also used for:

[0200] Receive structured policy clause data extracted from a large language model, including subsidy amount, applicable targets, and validity period;

[0201] Receives Bayesian model output based on open government data, including actual policy realization value and contribution rate indicators;

[0202] In the comparison matrix, each policy indicator is assigned a separate row, and each comparison area is assigned a separate column. The matrix also contains three types of fields: clause value, actual value, and deviation.

[0203] The policy text values ​​extracted by the large language model are filled into the clause value field, the average annual actual realization value calculated by the Bayesian model is filled into the actual value field, and the absolute difference between the clause value and the actual value is calculated and filled into the deviation field.

[0204] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A regional policy dynamic comparison system based on a large language model, characterized in that, The system includes: The policy radar module is used to monitor government websites and whitelisted sites in real time. It automatically triggers data collection when the frequency of policy text changes exceeds a preset threshold. The policy data processing engine includes: The text parsing unit is used to extract policy structure features based on dynamically optimized prompt words using a large language model; The effect verification unit is used to access the government's open database and quantify the actual effect of policies through a Bayesian structured time series model. The statistical significance unit is used to calculate the statistical significance of policy clause differences using the Wilcoxon signed-rank test and to label the statistical significance. The prompt word optimization module is used to dynamically optimize prompt words through a deep deterministic policy gradient reinforcement learning agent; It also includes a visualization comparison module for generating visualization comparison reports and deviation comparison matrices that include statistical significance annotations.

2. A method for dynamic comparison of regional policies based on a large language model, used to implement the system for dynamic comparison of regional policies based on a large language model as described in claim 1, characterized in that, The method includes: Real-time monitoring of government websites and whitelisted sites; automatic data collection when policy text changes are detected to exceed a preset threshold. Use a large language model to extract policy structure features based on dynamic prompts; Access the government's open database and quantify the actual effects of policies using a Bayesian structured time series model; The statistical significance of the differences in policy provisions was calculated using the Wilcoxon signed-rank test, and the statistical significance was then labeled. Dynamically optimize prompt words for agents through deep deterministic gradient reinforcement learning; Generates a visual comparison report and a deviation comparison matrix that include statistical significance annotations.

3. The method for dynamic comparison of regional policies based on a large language model according to claim 2, characterized in that, The real-time monitoring of government websites and whitelisted sites automatically triggers data collection when the frequency of policy text changes exceeds a preset threshold, including: Verify whether the policy-issuing agency belongs to the predefined list of administrative agency codes; The Hamming distance of the policy text is calculated using the Locality Sensitive Hash algorithm. When the distance exceeds 5%, a version change alarm is generated, which automatically triggers the collection of policy information for the changed version.

4. The method for dynamic comparison of regional policies based on a large language model according to claim 2, characterized in that, The method of using a large language model to extract policy structured features based on dynamic prompts includes: Based on the policy type selected by the user, generate a prompt word template Φ(T) containing dynamic placeholders; Compare the core policy paragraphs with similar reference policy summaries Inject the prompt template Φ(T) to obtain the prompt: Prompt word = Φ(T) + "Policy text: " + Substr(Policy text, L) + "Reference case: " + Where L is the truncation length of the core policy paragraph; Input the prompt words into the large language model and obtain the raw output. ; Validate the original output using a JSON-Schema validator. Does it conform to the predefined structure? ={type: "object",required: ["Applicable Object", "Key Indicators", "Validity Period"],…}; If the verification is successful, the original output will be accepted. As extracted policy structural features; If the verification fails, the error handling process is triggered: discard the original output. And regenerate the prompt words.

5. The method for dynamic comparison of regional policies based on a large language model according to claim 2, characterized in that, The access to the government's open database, and the quantification of the actual policy effects using a Bayesian structured time series model, includes: Retrieve datasets related to policy indicators from government open databases; Calculate semantic matching degree: ,in, The matching degree coefficient, For policy indicators Keyword set, In line with policy indicators Related datasets A collection of tags; Constructing a Bayesian structured time series model: ,in, Let be the observed value in year t. For local linear trend terms, For the Fourier seasonal term, The policy intervention coefficient, This is a policy intervention variable; it is 1 if implemented, and 0 otherwise. For random error term, The standard deviation of noise; Sampling iterations N times, estimating the posterior distribution ; Calculate the policy contribution rate: , ,in, for The posterior mean estimate, This is the average of observations before the policy was implemented. This refers to the number of years prior to the policy's implementation. when If the 95% confidence interval does not include 0, then the policy is considered valid.

6. The method for dynamic comparison of regional policies based on a large language model according to claim 2, characterized in that, The calculation of the statistical significance of policy clause differences using the Wilcoxon signed-rank test, and the subsequent annotation of statistical significance, includes: For each policy indicator Calculate the region pairing difference: Where k = 1, 2, ..., t, and t is the number of samples in each year. For policy indicators The actual value of k in year a in region a. For policy indicators The actual value of k in year b in region b; The absolute value of the pairing difference by region | Sort in ascending order and assign ranks; Calculate the rank sum of the positive differences respectively Rank sum of negative differences ; Take W=min( , ) as a statistic; Determine the p-value by looking up the table based on the sample size n; Significance is marked based on p-value.

7. The method for dynamic comparison of regional policies based on a large language model according to claim 2, characterized in that, The method of dynamically optimizing prompt words through deep deterministic gradient reinforcement learning agents includes: Initial prompts are generated using an Actor network with a deep deterministic policy gradient proxy. And input it into the large language model; Calculate the quality score of the prompt words based on the output of the large language model: =α·accuracy + β·coverage - γ·redundancy, where accuracy is the accuracy of the key points output by the large language model, coverage is the coverage of policy clauses, and redundancy is the output redundancy; Generate optimization prompts using policy networks: ,in, The basic prompt words generated for the Actor network, For noise; Based on the quality score of the prompt words Update the Critic network weights.

8. The method for dynamic comparison of regional policies based on a large language model according to claim 2, characterized in that, The generation of a visual comparison report and a deviation comparison matrix containing statistical significance annotations includes: Based on the enterprise attribute parameters input by the user, the policy adaptability of each region is calculated. The formula is: Adaptability = Σ(Policy clause weight × Enterprise matching degree) × Regional competitiveness coefficient; Generate a policy evolution decision tree diagram that includes the policy adaptability of each region, where the color depth of the nodes represents the strength of policy effectiveness.

9. The method for dynamic comparison of regional policies based on a large language model according to claim 2, characterized in that, The generation of the visual comparison report and deviation comparison matrix, which include statistical significance annotations, also includes: Receive structured policy clause data extracted from a large language model, including subsidy amount, applicable targets, and validity period; Receives Bayesian model output based on open government data, including actual policy realization value and contribution rate indicators; In the comparison matrix, each policy indicator is assigned a separate row, and each comparison area is assigned a separate column. The matrix also contains three types of fields: clause value, actual value, and deviation. The policy text values ​​extracted by the large language model are filled into the clause value field, the average annual actual realization value calculated by the Bayesian model is filled into the actual value field, and the absolute difference between the clause value and the actual value is calculated and filled into the deviation field.