Policy declaration auditing method based on large language model

By automating the policy application review process using a large language model based on the Transformer architecture, the problems of low efficiency and poor accuracy in traditional manual review have been solved. This has enabled an efficient and fair policy application review process, improving review efficiency and credibility.

CN120931231APending Publication Date: 2025-11-11天元大数据信用管理有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511001168.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional policy application and review processes rely on manual methods, which are inefficient, susceptible to subjective factors, and difficult to guarantee in terms of accuracy and impartiality. Furthermore, the lack of unified review standards leads to inconsistent review results and insufficient credibility.

Method used

Employing a large language model based on the Transformer architecture, the system achieves automated policy application review through data preprocessing, training, review, and feedback optimization. Data preprocessing includes cleaning and word segmentation; the training phase uses labeled datasets to fine-tune the model; and the review phase generates visualizations and collects feedback for model optimization.

Benefits of technology

The review process has significantly shortened the review time, improved review efficiency and accuracy, ensured the consistency and fairness of the results, reduced labor costs, and enhanced the credibility of policy application review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931231A_ABST
    Figure CN120931231A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large language models, in particular to a policy declaration auditing method based on a large language model. The policy declaration auditing method based on the large language model comprises the steps of collecting policy declaration materials, converting the policy declaration materials into a text format, and performing cleaning and word segmentation processing; a model based on a Transform architecture is used as a basic model, and the marked policy declaration auditing data set is used for training and fine tuning of the model; inputting the pre-processed policy declaration material into the trained large language model, generating an audit result, and performing visual display; the auditing result is fed back to the declaration enterprise, feedback suggestions are collected, and the auditing result is evaluated and analyzed; and optimizing and adjusting the large language model according to the feedback opinion and the evaluation result. According to the policy declaration auditing method based on the large language model, the auditing time is shortened, the time cost and the labor cost are reduced, the influence of subjective factors in manual auditing is avoided, and the accuracy and the fairness of an auditing result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large language model technology, and in particular to a policy application review method based on a large language model. Background Technology

[0002] Traditional policy application review processes primarily rely on manual review. Reviewers need to spend a significant amount of time and effort reading and understanding the application materials, assessing their compliance, authenticity, and feasibility. This method is not only inefficient but also susceptible to the subjective influence of reviewers, making it difficult to guarantee the accuracy and impartiality of the review results.

[0003] For example, in some large policy application projects, reviewers may need to process hundreds or thousands of application materials. Long hours of work can easily lead to fatigue and lack of concentration, thereby increasing the risk of review errors.

[0004] (ii) Insufficient information processing capabilities

[0005] With the increasing number of policy applications and the growing complexity of the application materials, manual review methods are proving inadequate in terms of information processing capabilities. Reviewers find it difficult to conduct a comprehensive and detailed analysis of a large number of application materials in a short period of time, easily overlooking important information and thus affecting the quality of the review.

[0006] Taking a company's application for a certain enterprise certification policy as an example, the application materials may include information on various aspects such as the company's R&D investment, intellectual property status, and the proportion of scientific and technological personnel. The reviewers need to verify and analyze this information one by one, which is a huge workload.

[0007] (iii) Lack of unified review standards

[0008] Different reviewers may have differing interpretations of policy application standards, leading to inconsistencies in review results. Furthermore, the lack of unified review standards and procedures makes the review process prone to arbitrariness and irregularities, impacting the credibility of policy application reviews.

[0009] For example, for the same policy application, different reviewers may have different understandings and grasps of certain key indicators, resulting in different review results. This will not only confuse the applicant companies, but also undermine the fairness and authority of the policy application review process.

[0010] To improve the efficiency and accuracy of policy application review and to standardize and regulate the review process, this invention proposes a policy application review method based on a large language model. Summary of the Invention

[0011] To overcome the shortcomings of existing technologies, this invention provides a simple and efficient policy application review method based on a large language model.

[0012] This invention is achieved through the following technical solution:

[0013] A policy application review method based on a large language model, characterized by the following steps:

[0014] Step S1: Data Preprocessing

[0015] Collect policy application materials, including basic enterprise information, description of the application project and relevant supporting documents, and convert them into text format. Clean and segment the text data.

[0016] Step S2: Large Language Model Training

[0017] The model is based on the Transformer architecture and trained and fine-tuned using a labeled policy declaration review dataset.

[0018] Step S3: Policy Application and Review

[0019] The pre-processed policy application materials are input into a trained large language model. The model is then called to analyze and understand the application materials based on the learned policy application review rules and standards, and to evaluate and generate review results, including whether the application is compliant and meets policy requirements.

[0020] The review results are visualized to provide reviewers with an intuitive reference; reviewers then use the visualized results to further examine and judge the application materials.

[0021] Step S4: Feedback and Optimization of Audit Results

[0022] The review results will be fed back to the applicant companies, informing them of the review status and any existing problems.

[0023] Collect feedback from reviewers and applicant companies, and evaluate and analyze the review results;

[0024] Based on feedback and evaluation results, the large language model is optimized and adjusted to continuously improve its accuracy and adaptability.

[0025] In step S1, cleaning the text data includes removing special characters, HTML tags, and extra spaces from the text data.

[0026] The cleaned data is segmented into words, breaking the text down into individual words or phrases for processing by a large language model.

[0027] In step S1, a data cleaning script is written using Python, and regular expressions are used to remove special characters from the text.

[0028] The Chinese word segmentation tool jieba was used to segment the cleaned text data.

[0029] In step S2, the BERT model, the Tongyi Qianwen series model, or the GPT series model are used as the base model. The base model is trained using a labeled policy application and review dataset so that it learns the rules and standards of policy application and review.

[0030] In step S2, the BERT model is fine-tuned using the Hugging Face Transformers library.

[0031] In step S3, the audit results are displayed in the form of charts or reports using visualization tools such as Matplotlib or Plotly.

[0032] Use Matplotlib to create a bar chart to show the approval rate of different application projects.

[0033] A policy application review system based on a large language model is provided to implement the above method, including a data collection and preprocessing module, a large language model training module, an application review module, and a feedback and optimization module.

[0034] The data collection and preprocessing module is responsible for collecting policy application materials, including basic information of enterprises, descriptions of application projects and relevant supporting documents, and converting them into text format, as well as cleaning and segmenting the text data.

[0035] The large language model training module is responsible for training and fine-tuning the model using a Transformer-based model as the base model and a labeled policy declaration and review dataset.

[0036] The application review module is responsible for inputting the pre-processed policy application materials into the trained large language model, calling the model to analyze and understand the application materials according to the learned policy application review rules and standards, and evaluating and generating review results, including whether the application is compliant and meets policy requirements;

[0037] The review results are visualized to provide reviewers with an intuitive reference; reviewers then use the visualized results to further examine and judge the application materials.

[0038] The feedback and optimization module is responsible for providing the review results to the applicant companies, informing them of the review status and any issues encountered; collecting feedback from reviewers and applicant companies, evaluating and analyzing the review results; and optimizing and adjusting the large language model based on the feedback and evaluation results to continuously improve the model's accuracy and adaptability.

[0039] A policy application review device based on a large language model is characterized by comprising a memory and a processor; the memory is used to store computer programs, and the processor is used to execute the computer programs to implement the above-mentioned method steps.

[0040] A readable storage medium, characterized in that: a computer program is stored on the readable storage medium, and the computer program, when executed by a processor, implements the above-described method steps.

[0041] The beneficial effects of this invention are: the policy application review method based on a large language model greatly shortens the review time, reduces time and labor costs, avoids the influence of subjective factors in manual review, and improves the accuracy and fairness of the review results. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Appendix Figure 1 This is a schematic diagram of the policy application review method based on a large language model according to the present invention. Detailed Implementation

[0044] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions in the embodiments of this invention will be clearly and completely described below in conjunction with the embodiments of this invention. Obviously, the described embodiments are merely...

[0045] These are some, but not all, embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort should fall within the scope of protection of the present invention.

[0046] This policy application review method based on a large language model includes the following steps:

[0047] Step S1: Data Preprocessing

[0048] Collect policy application materials, including basic enterprise information, description of the application project and relevant supporting documents, and convert them into text format. Clean and segment the text data.

[0049] Step S2: Large Language Model Training

[0050] The model is based on the Transformer architecture and trained and fine-tuned using a labeled policy declaration review dataset.

[0051] Step S3: Policy Application and Review

[0052] The pre-processed policy application materials are input into a trained large language model. The model is then called to analyze and understand the application materials based on the learned policy application review rules and standards, and to evaluate and generate review results, including whether the application is compliant and meets policy requirements.

[0053] The review results are visualized to provide reviewers with an intuitive reference; reviewers then use the visualized results to further examine and judge the application materials.

[0054] During the review process, large language models can quickly and accurately analyze and evaluate application materials, greatly improving review efficiency. Visual presentations allow reviewers to understand the review results more intuitively, facilitating decision-making.

[0055] Step S4: Feedback and Optimization of Audit Results

[0056] The review results will be fed back to the applicant companies, informing them of the review status and any existing problems.

[0057] Collect feedback from reviewers and applicant companies, and evaluate and analyze the review results;

[0058] Based on feedback and evaluation results, the large language model is optimized and adjusted to continuously improve its accuracy and adaptability. Timely feedback and optimization enable the large language model to continuously learn and improve, enhancing its performance in policy application review.

[0059] In step S1, cleaning the text data includes removing special characters, HTML tags, and extra spaces from the text data.

[0060] The cleaned data is segmented into words, breaking the text down into individual words or phrases for processing by a large language model.

[0061] In step S1, a data cleaning script is written using Python, and regular expressions are used to remove special characters from the text.

[0062] The Chinese word segmentation tool jieba was used to segment the cleaned text data.

[0063] During data collection, it is crucial to ensure the completeness and accuracy of the application materials to avoid impacting subsequent review results due to missing or incorrect data. Data cleaning and word segmentation are performed to improve the large language model's ability to understand and process text data.

[0064] In step S2, the BERT model, the Tongyi Qianwen series model, or the GPT series model are used as the base model. The base model is trained using a labeled policy application and review dataset so that it learns the rules and standards of policy application and review.

[0065] During training, appropriate optimization algorithms and hyperparameter tuning methods are employed to improve the model's performance and generalization ability. Choosing a suitable large language model is crucial, as different models may differ in their language understanding and processing capabilities. Fine-tuning the model and training it using labeled datasets allows it to better adapt to policy application review tasks.

[0066] Fine-tuning is a crucial step in training Large Language Models (LLMs). It refers to further training the pre-trained model using task-specific data to adapt it to a particular application scenario. Fine-tuning methods include full-parameter fine-tuning, low-rank adaptation (LoRA), and adapters. Fine-tuning can improve the model's accuracy in specific tasks (such as healthcare and finance) and enhance its contextual understanding capabilities.

[0067] In step S2, the BERT model is fine-tuned using the Hugging Face Transformers library.

[0068] (1) Define a custom PyTorch dataset class named PolicyDataset to process text data and convert it into a format suitable for sequence classification tasks using the BERT model. The code is as follows:

[0069]

[0070]

[0071] (2) A pre-trained BERT model and its corresponding word segmenter were loaded using Hugging Face's transformers library, and the model was configured for sequence classification tasks. The code is as follows:

[0072] # Initialize the tokenizer and model

[0073] tokenizer=BertTokenizer.from_pretrained('bert-base-chinese')

[0074] model=BertForSequenceClassification.from_pretrained('bert-base-chinese',num_labels=2)

[0075] (3) Define some hyperparameters for training the BERT model, including batch size, maximum text length, optimizer learning rate, and training iterations. The code is as follows:

[0076] # Define training parameters

[0077] batch_size = 16

[0078] max_length = 512

[0079] learning_rate = 2e-5

[0080] num_epochs=3

[0081] (4) Encapsulate the declaration text data and labeling data into a custom PolicyDataset class, and convert it into a batch-loadable data iterator using DataLoader. The code is as follows:

[0082] # Create dataset and data loader

[0083] texts = [...] # Declaration text data

[0084] labels = [...] # Labeling

[0085] policy_dataset=PolicyDataset(texts,labels,tokenizer,max_length)

[0086] policy_dataloader=DataLoader(policy_dataset,batch_size=batch_size,shuffle=True)

[0087] (5) Define the optimizer and loss function, as shown in the code below:

[0088] # Define the optimizer and loss function

[0089] optimizer=AdamW(model.parameters(),lr=learning_rate)

[0090] loss_fn=torch.nn.CrossEntropyLoss()

[0091] (6) Move the model and data to the specified device (GPU or CPU) and train it. The code is as follows:

[0092]

[0093]

[0094] In step S3, the audit results are displayed in the form of charts or reports using visualization tools such as Matplotlib or Plotly.

[0095] Use Matplotlib to create a bar chart to show the approval rate of different application projects.

[0096] The policy application and review system based on a large language model is used to implement the above methods, including a data collection and preprocessing module, a large language model training module, an application and review module, and a feedback and optimization module.

[0097] The data collection and preprocessing module is responsible for collecting policy application materials, including basic information of enterprises, descriptions of application projects and relevant supporting documents, and converting them into text format, as well as cleaning and segmenting the text data.

[0098] The large language model training module is responsible for training and fine-tuning the model using a Transformer-based model as the base model and a labeled policy declaration and review dataset.

[0099] The application review module is responsible for inputting the pre-processed policy application materials into the trained large language model, calling the model to analyze and understand the application materials according to the learned policy application review rules and standards, and evaluating and generating review results, including whether the application is compliant and meets policy requirements;

[0100] The review results are visualized to provide reviewers with an intuitive reference; reviewers then use the visualized results to further examine and judge the application materials.

[0101] The feedback and optimization module is responsible for providing the review results to the applicant companies, informing them of the review status and any issues encountered; collecting feedback from reviewers and applicant companies, evaluating and analyzing the review results; and optimizing and adjusting the large language model based on the feedback and evaluation results to continuously improve the model's accuracy and adaptability.

[0102] The policy application review device based on a large language model includes a memory and a processor; the memory is used to store computer programs, and the processor is used to execute the computer programs to implement the above-described method steps.

[0103] The readable storage medium stores a computer program that, when executed by a processor, implements the above-described method steps.

[0104] Compared with existing technologies, this policy application review method based on a large language model has the following characteristics:

[0105] 1) Improved review efficiency

[0106] By using large language models to automatically analyze and review policy application materials, review time can be significantly shortened and review efficiency improved. Large language models can process large volumes of application materials in a short time and quickly generate review results, saving policy management departments substantial manpower and time costs.

[0107] For example, in the past, manual review of a single application could take several days or even weeks, but with the use of large language models, the review time can be shortened to a few hours or even less.

[0108] 2) Improved the accuracy of the review process.

[0109] Large language models possess powerful language understanding and analysis capabilities, enabling them to accurately comprehend the content of application materials and objectively and impartially assess the compliance, authenticity, and feasibility of the application. Compared to manual review, large language models can avoid the influence of subjective factors, improving the accuracy and impartiality of review results.

[0110] Large language models can accurately extract and analyze key information in application materials, reducing review errors caused by human negligence or misjudgment.

[0111] 3) Standardized review process implemented.

[0112] It provides a unified set of policy application review standards and procedures, ensuring consistency and standardization through the learning and application of large language models. Different reviewers can conduct reviews based on the same standards and procedures, reducing discrepancies in review results for identical application materials, guaranteeing fairness and impartiality, and enhancing the credibility of policy application review.

[0113] 4) Promoted policy optimization

[0114] By analyzing and providing feedback on the review results, problems and shortcomings in the policy application process can be identified in a timely manner, providing a basis for policy optimization and adjustment. Policy management departments can then improve and refine policies based on the review results, enhancing their relevance and effectiveness.

[0115] For example, if it is found that a large number of applicant companies have difficulty understanding or do not meet the requirements of a certain policy clause, the policy management department can revise and interpret the clause to make the policy more reasonable and easier to implement.

[0116] The embodiments described above are merely one specific implementation of the present invention. Ordinary changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included within the protection scope of the present invention.

Claims

1. A policy application review method based on a large language model, characterized in that: Includes the following steps: Step S1: Data Preprocessing Collect policy application materials, including basic enterprise information, description of the application project and relevant supporting documents, and convert them into text format. Clean and segment the text data. Step S2: Large Language Model Training The model is based on the Transformer architecture and trained and fine-tuned using a labeled policy declaration review dataset. Step S3: Policy Application and Review The pre-processed policy application materials are input into a trained large language model. The model is then called to analyze and understand the application materials based on the learned policy application review rules and standards, and to evaluate and generate review results, including whether the application is compliant and meets policy requirements. The review results are visualized to provide reviewers with an intuitive reference; reviewers then use the visualized results to further examine and judge the application materials. Step S4: Feedback and Optimization of Audit Results The review results will be fed back to the applicant companies, informing them of the review status and any existing problems. Collect feedback from reviewers and applicant companies, and evaluate and analyze the review results; Based on feedback and evaluation results, the large language model is optimized and adjusted to continuously improve its accuracy and adaptability.

2. The policy application review method based on a large language model according to claim 1, characterized in that: In step S1, cleaning the text data includes removing special characters, HTML tags, and extra spaces from the text data. The cleaned data is segmented into words, breaking the text down into individual words or phrases for processing by a large language model.

3. The policy application review method based on a large language model according to claim 2, characterized in that: In step S1, a data cleaning script is written using Python, and regular expressions are used to remove special characters from the text. The Chinese word segmentation tool jieba was used to segment the cleaned text data.

4. The policy application review method based on a large language model according to claim 1, characterized in that: In step S2, the BERT model, the Tongyi Qianwen series model, or the GPT series model are used as the base model. The base model is trained using a labeled policy application and review dataset so that it learns the rules and standards of policy application and review.

5. The policy application review method based on a large language model according to claim 4, characterized in that: In step S2, the BERT model is fine-tuned using the Hugging Face Transformers library.

6. The policy application review method based on a large language model according to claim 1, characterized in that: In step S3, the audit results are displayed in the form of charts or reports using visualization tools such as Matplotlib or Plotly. Use Matplotlib to create a bar chart to show the approval rate of each application project.

7. A policy application review system based on a large language model, characterized in that: The method for implementing any one of claims 1 to 6 includes a data acquisition and preprocessing module, a large language model training module, an application review module, and a feedback and optimization module. The data collection and preprocessing module is responsible for collecting policy application materials, including basic information of enterprises, descriptions of application projects and relevant supporting documents, and converting them into text format, as well as cleaning and segmenting the text data. The large language model training module is responsible for training and fine-tuning the model using a Transformer-based model as the base model and a labeled policy declaration and review dataset. The application review module is responsible for inputting the pre-processed policy application materials into the trained large language model, calling the model to analyze and understand the application materials according to the learned policy application review rules and standards, and evaluating and generating review results, including whether the application is compliant and meets policy requirements; The review results are visualized to provide reviewers with an intuitive reference; reviewers then use the visualized results to further examine and judge the application materials. The feedback and optimization module is responsible for providing the review results to the applicant companies, informing them of the review status and any issues encountered; collecting feedback from reviewers and applicant companies, evaluating and analyzing the review results; and optimizing and adjusting the large language model based on the feedback and evaluation results to continuously improve the model's accuracy and adaptability.

8. A policy application review device based on a large language model, characterized in that: It includes a memory and a processor; the memory is used to store a computer program, and the processor is used to execute the computer program to implement the steps of the method as described in any one of claims 1 to 6.

9. A readable storage medium, characterized in that: The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 6.