Public website page analysis processing method, storage medium and system

Through content data analysis and large-scale model verification of public website pages, the deficiencies in the existing technology in the quality assessment of policy interpretation on public websites have been addressed, rapid assessment and ranking of policy importance have been achieved, and the quality of public services and governance efficiency have been improved.

CN120632239APending Publication Date: 2025-09-12MINJIANG UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510972147.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing public websites and policy interpretation pages lack scientific, comprehensive and effective analysis methods, which makes it difficult to accurately measure the quality and importance of policy interpretations, hindering the improvement of public service quality and the optimization of governance efficiency.

Method used

By obtaining content data from policy notification pages, detecting links, analyzing dimension labels, calculating the importance scores of responsible units, and using the large-model RAG knowledge base to verify the validity of the subject content and generate importance prompt words, we can achieve rapid assessment and ranking of policy importance.

Benefits of technology

It improves the efficiency of page information extraction without increasing additional computing power, can better evaluate the importance and quality of policy interpretation, and provide intuitive analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632239A_ABST
    Figure CN120632239A_ABST
Patent Text Reader

Abstract

The invention discloses a public website page analysis processing method, a storage medium and a system, and the method comprises the following steps: S1, obtaining a public website which comprises a policy notification page, and obtaining page content data of the policy notification page; s2, analyzing the acquired page content data, and acquiring a notification file title corresponding to a page in the page content data; s3, detecting links in the policy notification page, judging whether notification file title characters and policy interpretation characters exist in a first link or not, and if yes, executing the step S4, obtaining a first page pointed by the first link, and analyzing the first page to obtain a plurality of dimension tags; by assigning the importance score of the policy, the technical effect of quickly and intuitively analyzing the importance of the policy in the page can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for analyzing public website pages, and in particular to a method capable of performing analysis and preprocessing on policy interpretation. Background Art

[0002] Public websites are open to the general public and bear the heavy responsibility of publicizing government policies. There is also a great demand for studying these pages. In the digital age, public websites serve as important platforms for public information release, policy interpretation, and public services. They are key channels for safeguarding the public's right to know, right to participate, and right to oversight. Policy interpretation pages are the core link in helping the public accurately understand the connotations of public policies and promoting their effective implementation. However, the quality of various public websites and policy interpretation pages currently varies significantly, and there is a lack of a scientific, comprehensive, and effective analysis method. This has created significant obstacles to analyzing page quality and targeted optimization and improvement.

[0003] Existing evaluation methods have significant limitations: some focus solely on a website's technical performance, such as page response speed and link validity, while overlooking core elements such as the accuracy, completeness, and accessibility of policy interpretation content. Others focus on superficial indicators like the number of information released and the frequency of updates, failing to fully measure the actual value of policy interpretation in helping the public understand and apply policies to solve practical problems. Furthermore, the lack of unified and comparable evaluation standards for public websites operated by different sectors and entities makes cross-sector and cross-entity website quality comparisons difficult, hindering the analysis and dissemination of best practices.

[0004] In practice, the lack of an accurate importance scoring method makes it difficult for readers to quickly interpret important policies, hindering the improvement of public service quality and the optimization of public governance effectiveness. Therefore, a method that can comprehensively and accurately process public websites and policy interpretation pages by integrating multiple factors is urgently needed to meet the urgent needs of optimizing public information services and responding to public needs. Summary of the Invention

[0005] To this end, it is necessary to provide a public website page analysis and processing method, which includes the following steps: S1, obtaining a public website, the public website including a policy notification page, and obtaining page content data of the policy notification page.

[0006] S2. Analyze the acquired page content data to obtain a notification file title corresponding to the page in the page content data;

[0007] S3. Detect the links in the policy notification page and determine whether the first link contains the title of the notification document and the policy interpretation. If so, proceed to the following steps:

[0008] S4. Obtain a first page pointed to by the first link, analyze the first page, and obtain a plurality of dimension tags;

[0009] S5. Marking the hierarchical correspondence between the dimension tags according to the title and paragraph information in the first page;

[0010] S6. Obtain the corresponding field of the dimension label named responsible unit to obtain a list of responsible units for each policy;

[0011] S7. Calculate the importance score Dp of the p-th policy: Set the initial score S for the q-th responsible unit corresponding to the p-th policy.

[0012]

[0013] in,

[0014] Nq is the total number of times the qth responsible unit appears in the responsible units of other policies;

[0015] S8. Sort each policy in descending order according to its importance score and then display it.

[0016] In one embodiment of the present application, step S9 is further performed: setting the initial score to S×tq, checking whether the responsible unit q belongs to the first verification list, if not, setting tq=0.5; if the verification result is yes, otherwise setting tq=1.0.

[0017] In one embodiment of the present application, a verification step is further performed to obtain the corresponding website of the responsible unit q and verify whether the policy notification page exists on the homepage of the website. If so, t=1.5 is set.

[0018] In one embodiment of the present application, the following steps are further included to verify the validity of each subject matter:

[0019] The specific steps are to verify the information under the dimension label named main content, extract the measure entry, and check whether the content of the measure entry is consistent with the content data of the policy notification page.

[0020] In one embodiment of the present application, the steps are further included: segmenting the text into units of measures, generating importance prompt words based on the policy importance scores corresponding to the measures, merging the segmented text and the importance prompt words into a whole material, storing the whole material into a large model RAG knowledge base, and calculating a text vector.

[0021] A public website page analysis and processing storage medium stores a computer program, which performs the following steps when run: S1, obtaining a public website, the public website including a policy notification page, and obtaining page content data of the policy notification page.

[0022] S2. Analyze the acquired page content data to obtain a notification file title corresponding to the page in the page content data;

[0023] S3. Detect the links in the policy notification page and determine whether the first link contains the title of the notification document and the policy interpretation. If so, proceed to the following steps:

[0024] S4. Obtain a first page pointed to by the first link, analyze the first page, and obtain a plurality of dimension tags;

[0025] S5. Marking the hierarchical correspondence between the dimension tags according to the title and paragraph information in the first page;

[0026] S6. Obtain the corresponding field of the dimension label named responsible unit to obtain a list of responsible units for each policy;

[0027] S7. Calculate the importance score Dp of the p-th policy: Set the initial score S for the q-th responsible unit corresponding to the p-th policy.

[0028]

[0029] in,

[0030] Nq is the total number of times the qth responsible unit appears in the responsible units of other policies;

[0031] S8. Sort each policy in descending order according to its importance score and then display it.

[0032] In one embodiment of the present application, the computer program further executes step S9 when being run: setting the initial score to S×tq, checking whether the responsible unit q belongs to the first verification list, if not, setting tq=0.5; if the verification result is yes, otherwise setting tq=1.0.

[0033] In one embodiment of the present application, the computer program further executes the steps of obtaining the corresponding website of the responsible unit q, checking whether the policy notification page exists on the homepage of the website, and setting t=1.5 if it exists.

[0034] In one embodiment of the present application, steps are also included for verifying the validity of each subject content: specifically, verifying the information under the dimension label named main content, extracting measure entries, and checking whether the content of the measure entries is consistent with the content data of the policy notification page; steps are also performed to segment the text based on measures, and generate importance prompt words based on the policy importance scores corresponding to the measures, merge the segmented text and the importance prompt words into a whole material, store the whole material into the large model RAG knowledge base, and calculate the text vector.

[0035] A public website page analysis and processing system is provided with the public website page analysis and processing storage medium as described in the preceding item.

[0036] Unlike existing technologies, this solution can improve the efficiency of user page parsing by assigning detection values, using existing models without requiring additional computing power. Ultimately, this solution can achieve better results in extracting page information. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a flow chart of the public website page analysis and processing method described in the specific implementation method;

[0038] Figure 2 Schematic diagram of a storage medium for analyzing and processing public website pages according to a specific embodiment;

[0039] Figure 3 This is a schematic diagram of a public website page analysis and processing system according to a specific implementation method. DETAILED DESCRIPTION

[0040] In order to explain the technical content, structural features, achieved objectives and effects of the technical solution in detail, the following is a detailed description in conjunction with specific embodiments and accompanying drawings.

[0041] The following detailed description is an exemplary description and is intended to provide further detailed description of the present invention. Unless otherwise indicated, all technical terms used in the present invention have the same meaning as those generally understood by those skilled in the art to which the present invention belongs. The terms used in the present invention are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0042] In the description of this application, the term "and / or" is used to describe a logical relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A exists, B exists, and both A and B exist. In addition, the character " / " in this document generally indicates that the objects before and after are in a logical "or" relationship.

[0043] In this application, terms such as "first" and "second" are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantity, priority or sequence relationship between these entities or operations.

[0044] Without further limitations, in this application, the words "include", "comprise", "have" or other similar expressions used in the sentences are intended to cover non-exclusive inclusion. These expressions do not exclude the presence of additional elements in the process, method or product including the elements, so that the process, method or product including a series of elements may include not only those defined elements, but also other elements not explicitly listed, or elements inherent to such process, method or product.

[0045] In this application, expressions such as "greater than," "less than," and "exceed" are understood to exclude the number itself; expressions such as "above," "below," and "within" are understood to include the number itself. In addition, in the description of the embodiments of this application, "multiple" means two or more (including two), and similar expressions related to "multiple" are also understood in this way, such as "multiple groups" and "multiple times," unless otherwise specifically limited.

[0046] In the description of the embodiments of the present application, the space-related expressions used, such as "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "vertical", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc., indicate the orientation or position relationship based on the orientation or position relationship shown in the specific embodiments or drawings, and are only for the convenience of describing the specific embodiments of the present application or facilitating the reader's understanding, and do not indicate or imply that the device or component referred to must have a specific position, a specific orientation, or be constructed or operated in a specific orientation. Therefore, it should not be understood as a limitation on the embodiments of the present application.

[0047] Unless otherwise expressly specified or limited, in the description of the embodiments of the present application, the terms "installed", "connected", "connected", "fixed", "set", etc. used should be understood in a broad sense. For example, the "connection" can be a fixed connection, a detachable connection, or an integrated setting; it can be a mechanical connection, an electrical connection, or a communication connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be the internal connection of two elements or the interaction relationship between two elements. For those skilled in the art of the present application, the specific meanings of the above terms in the embodiments of the present application can be understood according to the specific circumstances.

[0048] A public website refers to a website established for the purpose of providing public services, and is generally a non-profit website. In some cases of this application, a public page website can be a website that specializes in providing policy notifications, or it can be a website of a department or unit with public service functions, where the URL under the website's root domain name provides a policy notification page or some links provide a policy notification page.

[0049] See also Figure 1 , is a public website page analysis and processing method, comprising the following steps: S1, obtaining a public website, the public website including a policy notification page, and obtaining page content data of the policy notification page.

[0050] S2. Analyze the acquired page content data to obtain a notification file title corresponding to the page in the page content data;

[0051] S3. Detect the links in the policy notification page and determine whether the first link contains the title of the notification document and the policy interpretation. If so, proceed to the following steps:

[0052] S4. Obtain a first page pointed to by the first link, analyze the first page, and obtain a plurality of dimension tags;

[0053] S5. Marking the hierarchical correspondence between the dimension tags according to the title and paragraph information in the first page;

[0054] S6. Obtain the corresponding field of the dimension label named responsible unit to obtain a list of responsible units for each policy;

[0055] S7. Calculate the importance score Dp of the p-th policy: Set the initial score S for the q-th responsible unit corresponding to the p-th policy.

[0056]

[0057] Nq is the total number of times the qth responsible unit appears in the responsible units of other policies;

[0058] S8. Sort each policy in descending order according to its importance score and then display it.

[0059] Generally speaking, in important policy notification pages, the subject is usually the document title of this notification. At the end of the policy notification page, there is usually an interpretation link with the document title + policy interpretation. In the interpretation link, there can be a specific analysis of the above policy, usually a text analysis. Automatic acquisition of the above text analysis can save labor costs. After obtaining the first link, text extraction is performed on the first page pointed to by the first link, and a big data model analysis is performed to obtain dimension labels. The dimension labels can be dimension labels of multiple levels, such as titles, subtitles, paragraphs, text, boldface, responsible units, policies, measures, deadlines, etc. The dimension labels can be set according to user needs, and the relative value of the dimension label is the specific text extraction result. There can be a hierarchical relationship between dimension labels. In one embodiment of the application, the hierarchical relationship between multiple dimension labels is marked according to the title and paragraph information, which helps to sort out the logical relationship of the text. Then analyze and sort out the list of responsible unit relationships corresponding to each policy.

[0060] In a specific embodiment, it is necessary to evaluate the importance of the policy, and the evaluation is based on the following assumptions. When a policy has multiple responsible units, and the fewer policies each responsible unit is responsible for at the same time, it means that the policy requires a more specific response and the policy has a higher importance. The policy dimension label and the responsible unit dimension label can be a one-to-one, one-to-many or vacant relationship. S can be set to any constant as needed. In step S7, the importance score of each responsible unit is S / Nq, and the importance score of the policy is the sum of the importance scores of all Q responsible units corresponding to it.

[0061] In this way, by assigning a score to the importance of the policy, a technical effect of quickly and intuitively analyzing the importance of the policy on the page can be achieved. Then, step S8 is performed to sort the policies in descending order according to their importance scores and then display them, which can provide users with better analysis results.

[0062] In order to better assign policy importance scores, in some further embodiments, the nature of the responsible units is adjusted. For example, the scores in government-related units are higher than those in general enterprises. Based on the above, step S9 is also performed: S9, the initial score is set to S×tq, and it is verified whether the responsible unit q belongs to the first verification list. If not, tq is set to a value less than 1, for example, it can be set to tq=0.5; if the verification result is yes, otherwise tq=1.0 is set. Among them, a list of responsible units with high priority is set in the first verification list, and the responsible units can be preset according to user needs. The above settings can better optimize the importance score.

[0063] In one embodiment of the present application, a verification step is also performed to obtain the corresponding website of the responsible unit q and verify whether the policy notification page exists on the homepage of the website. If so, t is set to a value greater than 1, such as t = 1.5. The existence of a corresponding website for the responsible unit q and the presence of the policy notification page on the homepage further demonstrate the importance of the policy.

[0064] Some other embodiments further include a step of verifying the validity of each subject content: this step can be performed using another analysis model independent of steps S1-S8, such as the Tongyi Qianwen, CHATGPT model for semantic verification. Specifically, the solution also performs steps to verify the information under the dimension label named main content, extract measure items, and check whether the content of the measure items is consistent with the content data of the policy notification page. Through the above verification steps, the validity of the subject content can be verified, and the misidentification caused by overfitting of large models can be avoided. It can also verify whether the main content is consistent with the measure items to avoid misjudgment.

[0065] In one embodiment of the present application, the steps are also included, segmenting the text into units of measures, generating importance prompt words according to the policy importance scores corresponding to the measures, merging the segmented text and the importance prompt words into a whole material, storing the whole material into the large model RAG knowledge base, and calculating the text vector. Retrieval-Augmented Generation (RAG) is a model that combines retrieval and generation technologies. It generates answers or content by referencing information from an external knowledge base, has strong interpretability and customization capabilities, and is suitable for multiple natural language processing tasks such as question-answering systems, document generation, and intelligent assistants. By re-storing the measure text and the importance prompt words into the RAG knowledge base, the large model can have a deep understanding of the importance of the measures.

[0066] Others such as Figure 2 In the embodiment shown, a public website page analysis and processing storage medium 200 is provided, which stores a computer program. When the computer program is run, the following steps are executed: S1, obtaining a public website, wherein the public website includes a policy notification page, and obtaining page content data of the policy notification page.

[0067] S2. Analyze the acquired page content data to obtain a notification file title corresponding to the page in the page content data;

[0068] S3. Detect the links in the policy notification page and determine whether the first link contains the title of the notification document and the policy interpretation. If so, proceed to the following steps:

[0069] S4. Obtain a first page pointed to by the first link, analyze the first page, and obtain a plurality of dimension tags;

[0070] S5. Marking the hierarchical correspondence between the dimension tags according to the title and paragraph information in the first page;

[0071] S6. Obtain the corresponding field of the dimension label named responsible unit to obtain a list of responsible units for each policy;

[0072] S7. Calculate the importance score Dp of the p-th policy: Set the initial score S for the q-th responsible unit corresponding to the p-th policy.

[0073]

[0074] in,

[0075] Nq is the total number of times the qth responsible unit appears in the responsible units of other policies;

[0076] S8. Sort each policy in descending order according to its importance score and then display it.

[0077] In a specific embodiment, it is necessary to evaluate the importance of the policy, and the evaluation is based on the following assumptions. When a policy has multiple responsible units, and the fewer policies each responsible unit is responsible for at the same time, it means that the policy requires a more specific response, and the policy has a higher importance. The policy dimension label and the responsible unit dimension label can be a one-to-one, one-to-many or vacant relationship. S can be set to any constant as needed. In step S7, the importance score of each responsible unit is S / Nq, and the importance score of the policy is the sum of the importance scores of all Q responsible units corresponding to it. In this way, by assigning points to the importance scores of policies, the technical effect of quickly and intuitively analyzing the importance of policies on the page can be achieved. Then proceed to step S8, sort each policy in descending order according to its importance score and display it, which can provide users with better analysis results.

[0078] In some embodiments of the present application, when the computer program is executed, the following steps are further performed: S9, setting the initial score to S×tq, verifying whether the responsible unit q belongs to the first verification list, and if not, setting tq = 0.5; if the verification result is yes, otherwise setting tq = 1.0. Adjusting the nature of the responsible unit can better assign policy importance. The existence of a corresponding website for the responsible unit q and the presence of a policy notification page on the homepage further demonstrate the importance of the policy.

[0079] In one embodiment of the present application, the computer program further executes the steps of obtaining the corresponding website of the responsible unit q, checking whether the policy notification page exists on the homepage of the website, and setting t=1.5 if it exists.

[0080] In one embodiment of the present application, the following steps are further included to verify the validity of each subject matter:

[0081] Specifically, verify the information under the dimension label named "main content", extract the measure entry, and check whether the content of the measure entry is consistent with the content data of the policy notification page;

[0082] The text is then segmented by measure, and importance cues are generated based on the corresponding policy importance scores. The segmented text and importance cues are then combined into a single piece of material, which is then stored in the large-scale RAG knowledge base. The text vector is then calculated. This verification step verifies the validity of the subject matter, preventing misidentifications caused by overfitting in the large-scale model. It also verifies the consistency between the main content and the measure items, thus avoiding misjudgments. Re-storing the measure text and importance cues in the RAG knowledge base allows the large-scale model to gain a deeper understanding of the importance of the measure.

[0083] In such Figure 3 In the illustrated embodiment, a public website page analysis and processing system is provided. The processing system includes the public website page analysis and processing storage medium 200 described above. The storage medium 200 stores a computer program that, when executed, performs the following steps: S1. Obtain a public website, wherein the public website includes a policy notification page, and obtain page content data of the policy notification page.

[0084] S2. Analyze the acquired page content data to obtain a notification file title corresponding to the page in the page content data;

[0085] S3. Detect the links in the policy notification page and determine whether the first link contains the title of the notification document and the policy interpretation. If so, proceed to the following steps:

[0086] S4. Obtain a first page pointed to by the first link, analyze the first page, and obtain a plurality of dimension tags;

[0087] S5. Marking the hierarchical correspondence between the dimension tags according to the title and paragraph information in the first page;

[0088] S6. Obtain the corresponding field of the dimension label named responsible unit to obtain a list of responsible units for each policy;

[0089] S7. Calculate the importance score Dp of the p-th policy: Set the initial score S for the q-th responsible unit corresponding to the p-th policy.

[0090]

[0091] in,

[0092] Nq is the total number of times the qth responsible unit appears in the responsible units of other policies;

[0093] S8. Sort each policy in descending order according to its importance score and then display it.

[0094] In a specific embodiment, it is necessary to evaluate the importance of the policy, and the evaluation is based on the following assumptions. When a policy has multiple responsible units, and the fewer policies each responsible unit is responsible for at the same time, it means that the policy requires a more specific response, and the policy has a higher importance. The policy dimension label and the responsible unit dimension label can be a one-to-one, one-to-many or vacant relationship. S can be set to any constant as needed. In step S7, the importance score of each responsible unit is S / Nq, and the importance score of the policy is the sum of the importance scores of all Q responsible units corresponding to it. In this way, by assigning points to the importance scores of policies, the technical effect of quickly and intuitively analyzing the importance of policies on the page can be achieved. Then proceed to step S8, sort each policy in descending order according to its importance score and display it, which can provide users with better analysis results.

[0095] In some embodiments of the present application, when the computer program is executed, the following steps are further performed: S9, setting the initial score to S×tq, verifying whether the responsible unit q belongs to the first verification list, and if not, setting tq = 0.5; if the verification result is yes, otherwise setting tq = 1.0. Adjusting the nature of the responsible unit can better assign policy importance. The existence of a corresponding website for the responsible unit q and the presence of a policy notification page on the homepage further demonstrate the importance of the policy.

[0096] It should be noted that although the above embodiments have been described herein, this does not limit the scope of patent protection of the present invention. Therefore, based on the innovative concept of the present invention, changes and modifications to the embodiments described herein, or equivalent structural or equivalent process transformations made using the contents of the present invention's description and drawings, and direct or indirect application of the above technical solutions to other related technical fields, are all included in the scope of protection of the present invention's patent.

Claims

1. A method for analyzing and processing public website pages, characterized in that: The method comprises the following steps: S1, obtaining a public website, wherein the public website includes a policy notification page, and obtaining page content data of the policy notification page; S2. Analyze the acquired page content data to obtain a notification file title corresponding to the page in the page content data; S3. Detect the links in the policy notification page and determine whether the first link contains the title of the notification document and the policy interpretation. If so, proceed to the following steps: S4. Obtain a first page pointed to by the first link, analyze the first page, and obtain a plurality of dimension tags; S5. Marking the hierarchical correspondence between the dimension tags according to the title and paragraph information in the first page; S6. Obtain the corresponding field of the dimension label named responsible unit to obtain a list of responsible units for each policy; S7. Calculate the importance score Dp of the p-th policy: Set the initial score S for the q-th responsible unit corresponding to the p-th policy. Among them, Q is the number of responsible units corresponding to the pth policy; Nq is the total number of times the qth responsible unit appears among the responsible units of other policies; S8. Sort each policy in descending order according to its importance score and then display it.

2. The method for analyzing and processing public website pages according to claim 1, characterized in that: Also perform step: S9, set the initial score to S×tq, check whether the responsible unit q belongs to the first verification list, if not, set tq=0.5; if the verification result is yes, otherwise set tq=1.

0.

3. The method for analyzing and processing public website pages according to claim 2, characterized in that: A verification step is also performed to obtain the corresponding website of the responsible unit q and to verify whether the policy notification page exists on the homepage of the website. If so, t=1.5 is set.

4. The method for analyzing and processing public website pages according to claim 1, wherein: It also includes steps to verify the validity of each topic content: The specific steps are to verify the information under the dimension label named main content, extract the measure entry, and check whether the content of the measure entry is consistent with the content data of the policy notification page.

5. The method for analyzing and processing public website pages according to claim 4, characterized in that: The method also includes the steps of segmenting the text into units of measures, generating importance prompt words according to the policy importance scores corresponding to the measures, merging the segmented text and the importance prompt words into a whole material, storing the whole material into a large model RAG knowledge base, and calculating a text vector.

6. A public website page analysis and processing storage medium, characterized in that: A computer program is stored, and when the computer program is run, the following steps are executed: S1, obtaining a public website, the public website including a policy notification page, and obtaining page content data of the policy notification page. S2. Analyze the acquired page content data to obtain a notification file title corresponding to the page in the page content data; S3. Detect the links in the policy notification page and determine whether the first link contains the title of the notification document and the policy interpretation. If so, proceed to the following steps: S4. Obtain a first page pointed to by the first link, analyze the first page, and obtain a plurality of dimension tags; S5. Marking the hierarchical correspondence between the dimension tags according to the title and paragraph information in the first page; S6. Obtain the corresponding field of the dimension label named responsible unit to obtain a list of responsible units for each policy; S7. Calculate the importance score Dp of the p-th policy: Set the initial score S for the q-th responsible unit corresponding to the p-th policy. in, Nq is the total number of times the qth responsible unit appears in the responsible units of other policies; S8. Sort each policy in descending order according to its importance score and then display it.

7. The public website page analysis and processing storage medium according to claim 6, characterized in that: When the computer program is run, it further executes the following steps: S9, setting the initial score to S×tq, checking whether the responsible unit q belongs to the first verification list, if not, setting tq=0.5; if the verification result is yes, otherwise setting tq=1.

0.

8. The public website page analysis and processing storage medium according to claim 7, characterized in that: When the computer program is run, it further performs the steps of obtaining the corresponding website of the responsible unit q, checking whether the policy notification page exists on the homepage of the website, and setting t=1.5 if it exists.

9. The public website page analysis and processing storage medium according to claim 7, characterized in that: It also includes steps to verify the validity of each topic content: Specifically, verify the information under the dimension label named "main content", extract the measure entry, and check whether the content of the measure entry is consistent with the content data of the policy notification page; The steps are also performed to segment the text into units of measures, generate importance prompt words according to the policy importance scores corresponding to the measures, merge the segmented text and the importance prompt words into a whole material, store the whole material into the large model RAG knowledge base, and calculate the text vector.

10. A public website page analysis and processing system, characterized in that: The processing system is provided with a public website page analysis and processing storage medium according to any one of claims 4 to 9.