Policy mining and intelligent interaction platform based on AI large model

By adopting policy mining and intelligent interaction platforms based on AI models in policy text analysis, a policy analysis model and multi-dimensional feature extraction module of Transformer architecture are built, which solves the problems of low structural recognition accuracy, insufficient feature extraction, and lack of scientificity in the existing technology of policy text analysis, and achieves a more efficient and scientific policy text analysis.

CN119963001APending Publication Date: 2025-05-09STATE GRID SICHUAN ELECTRIC POWER CO TIANFU NEW DISTRICT POWER SUPPLY CO +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510083723.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the analysis of policy texts, the existing technology has problems such as low accuracy in structural identification, insufficient comprehensive feature extraction, and lack of scientificity and interpretability of the impact assessment results.

Method used

Using a policy mining and intelligent interaction platform based on AI models, intelligent analysis and scientific management of policy texts are realized by building a policy analysis model, multi-dimensional feature extraction module and an impact assessment model based on timing features of Transformer architecture.

Benefits of technology

The accuracy of policy text structure recognition is improved, comprehensive text feature representation is obtained, the scientificity and accuracy of influence assessment is enhanced, and the shortcomings existing in the existing technology are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963001A_ABST
    Figure CN119963001A_ABST
Patent Text Reader

Abstract

The invention provides a policy mining and intelligent interaction platform based on an AI large model, and belongs to the technical field of electrical digital data processing. The mining steps in the platform comprise: firstly, preprocessing and standardizing a policy text; then, a policy analysis model of a Transform architecture is used for carrying out structure identification; obtaining text features through a word vector model and a multi-dimensional feature extraction module; and finally, establishing an influence evaluation model based on the time sequence features, the network features and the hierarchical features, and realizing scientific analysis and management of the policy text. According to the platform, through fusion of deep learning and natural language processing technologies, especially in the aspect of structure identification, limitation of a traditional rule method needs to be overcome; in the aspect of feature extraction, effective extraction of deep semantic features needs to be realized; in the aspect of influence assessment, a scientific assessment model comprehensively considering multi-dimensional factors needs to be established. According to the method, the problem of inaccurate policy text analysis influence evaluation in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electronic digital data processing, and in particular, relates to a policy mining and intelligent interaction platform based on an AI big model. Background Art

[0002] With the increasing number and complexity of policies, intelligent analysis and management of policy texts has become an important topic of current research. Traditional policy text analysis methods mainly use rule-based and statistical methods to extract features and identify structures of policy texts through predefined text processing rules and statistical models. These methods usually include steps such as text preprocessing, word segmentation and tagging, feature extraction, and structured storage. In the text preprocessing stage, encoding conversion and format normalization are used; in the word segmentation and tagging stage, dictionary-based word segmentation algorithms and statistical models are used for part-of-speech tagging; in the feature extraction stage, word frequency features and position features are mainly considered; in the structured storage stage, relational databases are used to store and manage policy texts.

[0003] However, traditional technologies have many limitations in the policy text analysis process. First, rule-based text processing methods lack flexibility and cannot adapt to the diverse policy text formats and expressions, resulting in low accuracy in structural recognition. Second, traditional feature extraction methods rely too much on surface features and fail to fully tap into the deep semantic information and temporal features of policy texts, affecting the comprehensiveness and accuracy of feature representation. Third, existing policy impact assessment methods are mainly based on simple statistical indicators and fail to comprehensively consider the timeliness, relevance, and hierarchy of policy texts, resulting in a lack of scientificity and interpretability in the assessment results. These problems have seriously restricted the effectiveness and application value of policy text analysis.

[0004] In other words, there are technical problems in the existing technology regarding the inaccurate feature extraction and impact assessment of policy texts. Summary of the invention

[0005] In view of this, the present invention provides a policy mining and intelligent interaction platform based on AI big model, which can solve the technical problems of inaccurate policy text feature extraction and influence assessment in the prior art. The present invention is implemented as follows: the present invention provides a policy mining and intelligent interaction platform based on AI big model, including a policy mining module, a policy information base module and an indicator database module; the policy mining module is used to obtain policy text data, perform feature extraction, structure recognition and influence analysis on the policy text data, and generate a policy analysis report; the policy information base module is used to classify, store and manage the policy text data, build a policy data warehouse, and realize the archival storage of the policy text data; the indicator database module is used to store policy time series feature data, policy influence scores and policy storage level identifiers; the policy mining module is used to perform the following steps: ... Row feature extraction and structure recognition, wherein feature extraction includes word segmentation and part-of-speech tagging, extracting policy keyword sequences from the policy text data, establishing a policy text vector matrix, converting the policy keyword sequences into policy word vectors, and generating the policy feature vectors; constructing a Transformer architecture policy parsing model, performing multi-level clause structure recognition on the policy text data, and generating the policy document structured data; extracting the policy time series feature data, calculating the policy time decay weight value according to the policy release time, forming the data update sequence, integrating the policy node association strength value and the policy time decay weight value, calculating the policy influence score, and generating the policy storage level identifier.

[0006] Among them, the steps of obtaining policy text data are to process the policy text data for encoding consistency, convert all texts into UTF8 encoding format, filter text, remove special characters, redundant spaces and line breaks in the text, normalize the formats of numbers, dates, currencies, etc. in the text, segment the text, identify paragraph marks in the text, segment the text according to the natural paragraph structure of the policy text, standardize the text length, truncate overly long paragraphs, and set the maximum number of characters in a single paragraph to 2000.

[0007] Among them, the steps for generating the policy feature vector are specifically to establish a word vector training corpus, use the bag-of-words model to vectorize the text, use the word embedding model to train the word vector, select the continuous bag-of-words model for word vector learning, set the word vector dimension to 300 dimensions, the context window size to 5, the number of negative sampling to 5, perform dimensionality reduction processing on the trained word vector, and use the principal component analysis method to compress the word vector to 100 dimensions.

[0008] Among them, the construction steps of the Transformer architecture policy parsing model are specifically to design a multi-layer encoder based on the attention mechanism. The number of encoder layers is set to 6, each layer contains a multi-head self-attention layer and a feedforward neural network layer, and a clause structure recognition layer is designed. A hierarchical structure recognition strategy is adopted to sequentially identify the text structure of the chapter level, clause level, and item level, post-process the recognition results, merge the fragmented structural units, and adjust the abnormal hierarchical relationships.

[0009] Among them, it also includes the step of calculating the contribution value of policy keywords, and the policy keyword contribution value is obtained through a contribution calculation equation group, and the contribution calculation equation group includes a timing weight calculation equation, a position weight calculation equation and a semantic association calculation equation. The timing weight calculation equation is used to calculate the keyword timeliness impact factor value, the position weight calculation equation is used to calculate the keyword structure importance value, and the semantic association calculation equation is used to calculate the keyword topic relevance value.

[0010] It also includes the steps of building a policy association network, specifically establishing the association relationship between policy nodes based on the text similarity value, connecting policy documents with a similarity greater than 0.7, calculating the association strength between nodes, considering factors such as text similarity, release time differences, policy levels, etc., comprehensively calculating the association strength value, optimizing the association network, and removing weakly associated connections.

[0011] It also includes the steps of establishing policy archiving rules, specifically setting the archiving threshold based on the time decay weight, triggering the archiving operation when the weight value is lower than 0.3, compressing the archived data, establishing the archiving index, implementing the regular cleanup mechanism of the archived data, and safely deleting expired data.

[0012] It also includes the step of building a data warehouse, specifically designing a hybrid storage system that includes structured data and semi-structured data, using a relational database to store standardized policy data in the first database, and using a document database to store policy data in a flexible format in the second database, establishing a multi-dimensional index, and realizing a data synchronization mechanism.

[0013] It also includes the steps of building a multidimensional indicator data management system, specifically designing an indicator data structure that includes policy time series characteristics, influence indicators, and storage levels, establishing an indicator calculation engine, supporting real-time calculation and updating of indicator data, implementing version management of indicator data, recording the change history of indicator values, and establishing a monitoring mechanism for indicator data.

[0014] It also includes the step of generating a policy analysis report, specifically, sorting the policy texts based on the influence scores, generating an analysis report containing the basic information, structural characteristics, correlation relationships, and impact assessment of the policy texts, storing the influence scores and storage level identifiers in the indicator database, and storing the policy texts in a hierarchical manner according to the storage level identifiers.

[0015] Compared with the prior art, the present invention provides a policy mining and intelligent interaction platform based on AI big model, and the artificial intelligence-based policy text feature extraction and influence assessment method proposed in the present invention realizes intelligent analysis and scientific management of policy texts by constructing a policy parsing model of Transformer architecture, a multidimensional feature extraction module, and an influence assessment model based on time series features. The method first uses a deep learning model to perform structured analysis on the policy text, then uses multidimensional feature extraction technology to obtain text features, and finally performs influence assessment based on time series features and association networks.

[0016] Compared with traditional technologies, the solution of the present invention achieves accurate recognition of the structure of policy texts through a multi-level encoder based on the Transformer architecture in terms of structural recognition, overcoming the limitations of rule-based methods. In terms of feature extraction, the word vector model and multi-dimensional feature extraction technology are used to fully acquire the semantic, statistical and format features of policy texts. In terms of impact assessment, a scientific assessment model is established by integrating temporal features, network features and hierarchical features, thereby improving the accuracy and interpretability of the assessment results.

[0017] The present invention solves the core technical problem of policy text feature extraction and influence assessment based on artificial intelligence. The principle is as follows: automatic recognition of text structure is realized through a deep learning model, avoiding the limitations of manual rules; deep representation of text features is realized through multi-dimensional feature extraction technology, overcoming the one-sidedness of traditional feature extraction methods; a scientific evaluation mechanism is realized through a multi-factor fusion influence assessment model, solving the technical problem of inaccurate policy text feature extraction and influence assessment in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A flowchart of the steps executed by the policy mining module of the present invention; Figure 2 This is a page diagram of the energy policy mining and intelligent interaction platform based on the AI ​​big model in Example 2; Figure 3 This is a policy highlights diagram of the platform in Example 2; Figure 4 This is the policy data analysis diagram in Example 2. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0020] The present invention provides a policy mining and intelligent interaction platform based on AI big model, including a policy mining module, a policy information base module and an indicator database module; the policy mining module is used to obtain policy text data, perform feature extraction, structure recognition and influence analysis on the policy text data, and generate a policy analysis report; the policy information base module is used to classify, store and manage the policy text data, build a policy data warehouse, and realize archival storage of the policy text data; the indicator database module is used to store policy time series feature data, policy influence scores and policy storage level identifiers; Figure 1 As shown, the policy mining module is used to perform the following steps: S01, obtaining the policy text data, and performing data cleaning and data standardization processing on the policy text data based on the preset text processing protocol; S02, performing text error detection, identifying abnormal text formats based on the preset text rule library, using the preset text dictionary to perform text correction, and generating standardized text data; S03, performing word segmentation and part-of-speech tagging on the standardized text data, and extracting a policy keyword sequence from the policy text data; S04, establishing a policy text vector matrix, converting the policy keyword sequence into a policy word vector, and generating the policy feature vector; S05, extracting the language feature data, statistical feature data and format feature data of the policy feature vector to form the policy multidimensional feature vector; S06, calculating a policy text similarity matrix, and calculating text similarity values ​​between the policy text data based on the policy multidimensional feature vector; S07, constructing a Transformer architecture policy parsing model, performing multi-level clause structure recognition on the policy text data, and generating the policy document structured data; S08, calling the policy information base module to build a policy data warehouse, and storing the policy document structured data into the policy information base module, wherein the structured data in the policy document structured data is stored in a first database, and the semi-structured data in the policy document structured data is stored in a second database; S09, calculating the policy keyword weight value, and calculating the policy keyword contribution value of the policy keyword sequence based on the word frequency weight coefficient, the position weight coefficient and the semantic association coefficient; S10, generating a policy association network, constructing a policy node association relationship based on the text similarity value, and calculating a policy node association strength value; S11, constructing a data allocation matrix, hierarchically classifying the structured data of the policy document, and distributing and storing the data according to the data update frequency value and the storage priority value; S12, extracting the policy time series feature data, calculating the policy time decay weight value according to the policy release time, and forming the data update sequence; S13, calling the policy information base module to establish a policy archiving rule, and archiving and storing historical policy data through the policy information base module based on the policy time decay weight value; S14, integrating the policy node association strength value and the policy time decay weight value, calculating the policy influence score, and generating the policy storage level identifier; S15, sorting the policy text data by importance based on the policy influence score, generating the policy analysis report, storing the policy influence score and the policy storage level identifier in the indicator database module, and calling the policy information base module to perform hierarchical storage of the policy text data according to the policy storage level identifier; The policy keyword contribution value is obtained by the contribution calculation equation group, which includes a time series weight calculation equation, a position weight calculation equation and a semantic association calculation equation; The temporal weight calculation equation is used to calculate the value of the keyword timeliness impact factor, the input parameters of the temporal weight calculation equation include the first appearance time value of the keyword, the current time value, and the temporal reference coefficient, and the output parameter of the temporal weight calculation equation is the temporal weight coefficient; The position weight calculation equation is used to calculate the keyword structure importance value, the input parameters of the position weight calculation equation include the keyword level position value, the keyword occurrence frequency value, and the document position weight value, and the output parameter of the position weight calculation equation is the position importance coefficient; Among them, the semantic association calculation equation is used to calculate the keyword topic relevance value, the input parameters of the semantic association calculation equation include keyword vector value, topic vector value, word co-occurrence frequency value, and the output parameter of the semantic association calculation equation is the semantic association coefficient.

[0021] The specific implementation methods of the above steps are described in detail below. The specific implementation method of step S01 is to pre-process the obtained policy text data. First, the policy text data is processed for encoding consistency, and all texts are uniformly converted into UTF8 encoding format to ensure the uniformity of text encoding. Then text filtering is performed to remove special characters, redundant spaces and line breaks in the text, and the formats of numbers, dates, currencies, etc. in the text are standardized, and different forms of expression are unified into a standard format. Then the text is segmented to identify paragraph marks in the text and segment according to the natural paragraph structure of the policy text. Finally, the text length is standardized, and the overlong paragraphs are truncated. The maximum number of characters in a single paragraph is set to 2000 characters. The paragraphs that exceed the length are appropriately segmented to ensure the rationality of the text length. Through data cleaning and standardization, a standardized data basis is provided for subsequent text analysis.

[0022] The specific implementation method of step S02 is to perform text error correction. First, a text format check is performed based on a preset text rule library, which contains normative requirements such as punctuation usage rules, paragraph format rules, and special character usage rules. Texts that do not meet the rules are marked and the exception type is recorded. Then, the text is corrected using a preset text dictionary, which contains error correction resources such as a common error word list, a homophone word list, and a similar word list. The minimum edit distance algorithm is used to calculate the similarity between the text to be corrected and the standard entry in the dictionary. When the similarity exceeds the threshold of 0.85, the text to be corrected is replaced with the corresponding standard entry. For punctuation errors, intelligent correction is performed based on the context to ensure the standardization of punctuation use. Finally, the corrected text is manually reviewed to ensure the accuracy of the text correction.

[0023] The specific implementation method of step S03 is to perform text segmentation and part-of-speech tagging. First, a hybrid segmentation method based on dictionaries and statistics is adopted, and a two-way maximum matching algorithm is used for preliminary segmentation. The segmentation dictionary includes a domain-specific dictionary and a general dictionary. The segmentation results are disambiguated, and a sequence tagging model based on conditional random fields is used to identify and process ambiguous words. Then part-of-speech tagging is performed, and a hidden Markov model is used to tag the words after segmentation. The part-of-speech tagging set includes nouns, verbs, adjectives and other part-of-speech categories. The tagging results are post-processed, adjacent similar part-of-speech tags are merged, and the part-of-speech tags of special words are adjusted. Finally, a policy keyword sequence is extracted, and noun and verb keywords are screened based on part-of-speech tags and word frequency statistics, and a keyword sequence is formed according to the order of appearance in the text.

[0024] The specific implementation method of step S04 is to construct a policy word vector. First, a word vector training corpus is established, and a large amount of policy text is collected as training data. The bag-of-words model is used to vectorize the text and calculate the frequency of occurrence of words in the document. Then, the word embedding model is used for word vector training, and the continuous bag-of-words model is selected for word vector learning. The word vector dimension is set to 300 dimensions, the context window size is 5, and the number of negative sampling is 5. The trained word vector is subjected to dimensionality reduction processing, and the principal component analysis method is used to compress the word vector to 100 dimensions, retaining the main semantic information. Finally, each word in the policy keyword sequence is mapped to a corresponding word vector to form a policy feature vector matrix.

[0025] The specific implementation method of step S05 is to extract multidimensional features. First, language features are extracted, including linguistic features such as part-of-speech distribution features, syntactic structure features, and semantic relationship features. The distribution ratio of different parts of speech in the text is calculated, syntactic dependencies are identified, and a semantic network is constructed. Then statistical features are extracted, including statistical features such as word frequency features, word length features, and character distribution features. The frequency distribution of words is calculated, the distribution law of word length is analyzed, and the usage of characters is counted. Then format features are extracted, including typesetting format features such as paragraph structure features, punctuation usage features, and space distribution features. The distribution of paragraph length is analyzed, the frequency of punctuation usage is counted, and the format pattern of the text is identified. Finally, the various extracted features are integrated to form a policy multidimensional feature vector.

[0026] The specific implementation method of step S06 is to calculate the text similarity. First, the policy multidimensional feature vector is standardized, and the feature value is normalized to the interval of 0 to 1 using the range normalization method. Then the cosine similarity between the feature vectors is calculated to construct a text similarity matrix. The similarity matrix is ​​threshold filtered, and the similarity threshold is set to 0.7, and the similarity values ​​below the threshold are set to 0. Finally, the similarity matrix is ​​normalized to ensure the comparability of the similarity values.

[0027] The specific implementation method of step S07 is to construct a policy parsing model. First, a multi-layer encoder is designed based on the Transformer architecture of the attention mechanism. The number of encoder layers is set to 6, and each layer contains a multi-head self-attention layer and a feedforward neural network layer. Then, a clause structure recognition layer is designed, and a hierarchical structure recognition strategy is adopted to sequentially identify the text structure of the chapter level, clause level, and item level. The recognition results are post-processed, the fragmented structural units are merged, and the abnormal hierarchical relationships are adjusted. Finally, hierarchical policy document structure data is generated, including structured hierarchical relationships and text content.

[0028] The specific implementation method of step S08 is to build a data warehouse. First, the storage structure of the data warehouse is designed to establish a hybrid storage system containing structured data and semi-structured data. Structured data is stored in the first database, and a relational database is used to store standardized policy data. Semi-structured data is stored in the second database, and a document database is used to store policy data in a flexible format. Then, a data index is established, and a multi-dimensional index is established for the stored policy data to support fast retrieval and access. Finally, a data synchronization mechanism is implemented to ensure data consistency between the two databases.

[0029] The specific implementation method of step S09 is to calculate the keyword weight. First, the word frequency weight is calculated, and the importance of the keyword is calculated using the word frequency inverse document frequency algorithm. Then the position weight is calculated, and the position importance is calculated based on the position distribution of the keyword in the document. Then the semantic association weight is calculated, and the semantic importance is calculated based on the semantic relevance between the keyword and the document topic. Finally, the three weight coefficients are combined to calculate the comprehensive contribution value of the keyword.

[0030] The specific implementation method of step S10 is to construct an association network. First, the association relationship between policy nodes is established based on the text similarity value, and the policy documents with similarity greater than 0.7 are connected. Then the association strength between nodes is calculated, and the association strength value is comprehensively calculated by considering factors such as text similarity, release time difference, and policy level. Finally, the association network is optimized, weakly associated connections are removed, and the main association relationships are retained.

[0031] The specific implementation method of step S11 is to perform data allocation. First, a data allocation matrix is ​​constructed to classify policy data based on its update frequency and storage priority. Data with high update frequency and high priority are allocated to high-performance storage areas. Data with low update frequency and low priority are allocated to ordinary storage areas. Then a data scheduling mechanism is implemented to dynamically adjust data distribution according to data access patterns. Finally, a data backup strategy is established to ensure data security and reliability.

[0032] The specific implementation method of step S12 is to extract time series features. First, a time series is constructed based on the policy release time, and the time distribution characteristics of the policy text are analyzed. Then the time decay weight is calculated, and the timeliness weight of the policy is calculated using an exponential decay function, with the decay coefficient set to 0.1. Then a data update sequence is generated, and the policy data is prioritized according to the time decay weight. Finally, an update strategy is established to determine the time interval and update method for data updates.

[0033] The specific implementation method of step S13 is to realize archival storage. First, establish archiving rules, set the archiving threshold based on the time decay weight, and trigger the archiving operation when the weight value is lower than 0.3. Then perform data compression to compress the archived data to reduce the storage space occupied. Then establish an archiving index to ensure the retrievability of the archived data. Finally, implement a regular cleanup mechanism for archived data to safely delete expired data.

[0034] The specific implementation method of step S14 is to calculate the influence score. First, obtain the association strength value and time decay weight value of the policy node, and standardize the two indicators. Then design an influence calculation model, comprehensively consider factors such as association strength, timeliness, and policy level, and calculate the comprehensive influence score of the policy. Then set the storage level identifier according to the influence score, and divide the policy into three storage levels: high, medium, and low. Finally, verify the calculation results to ensure the rationality of the impact assessment.

[0035] The specific implementation method of step S15 is to generate an analysis report. First, the policy texts are sorted based on the influence scores to determine the importance ranking of the policies. Then a policy analysis report is generated, which includes the basic information, structural characteristics, correlation relationships, influence assessment, etc. of the policy texts. Then, the influence scores and storage level identifiers are stored in the indicator database to establish an index structure for the indicator data. Finally, the policy texts are stored hierarchically according to the storage level identifiers to achieve hierarchical management of data.

[0036] The specific implementation method of the policy information library module is to establish a distributed policy data storage system. First, a multi-level storage architecture is constructed, including an online storage layer, a near-line storage layer, and an offline storage layer. The online storage layer uses a high-performance database to store hot policy data. The near-line storage layer uses a distributed file system to store regular policy data. The offline storage layer uses an archiving storage system to store historical policy data. Then, a data hierarchical storage strategy is implemented to manage data hierarchically according to the access frequency and importance of policy data. Then a data synchronization mechanism is established to ensure data consistency between different storage layers. Finally, a data backup and recovery mechanism is implemented to ensure the security of policy data.

[0037] The specific implementation method of the indicator database module is to build a multi-dimensional indicator data management system. First, the indicator data structure is designed, including policy time series characteristics, influence indicators, storage levels and other dimensions. Then, an indicator calculation engine is established to support real-time calculation and update of indicator data. Then, version management of indicator data is implemented to record the change history of indicator values. Finally, a monitoring mechanism for indicator data is established to issue alarms for abnormal indicator values. Through multi-dimensional indicator management, scientific evaluation and effective management of policy data can be achieved.

[0038] As a further limitation of the above implementation scheme, the specific implementation mode of the present invention also includes the following features: the present invention can be implemented using electronic devices, specifically by building a distributed server cluster architecture to deploy and run the intelligent interactive platform. The server cluster adopts a master-slave architecture, the master server is responsible for core algorithm operations and data processing, the slave server is responsible for data storage and load balancing, and a high availability mechanism is configured to ensure the stable operation of the platform. The platform supports users to access the system through the Web or mobile terminal to conduct policy impact forecasting and analysis. Specifically, users can input specific policies of their concern into the platform for interference analysis, and predict the potential impact of policy adjustments on the market environment and the company itself by inputting business parameters including but not limited to corporate revenue data, cost structure, market share, etc., combined with the platform's existing policy data model. Based on accumulated historical data and real-time market data, the platform will use a multi-dimensional comparative analysis method to provide users with data-supported policy impact assessment reports to assist companies in formulating corresponding business strategy adjustment plans.

[0039] As a preferred embodiment of the present invention, the intelligent interactive platform is particularly suitable for policy analysis in the field of new energy and electric energy technology. Specifically, the policy information library module of the platform specifically constructs a professional knowledge map in the field of new energy, covering policy text data in multiple sub-fields such as photovoltaic power generation, wind power generation, energy storage technology, smart grid, and hydrogen energy utilization. In the policy mining module, the system pre-sets a terminology dictionary and technical standard library unique to the new energy industry, including but not limited to professional field knowledge such as power generation efficiency indicators, environmental protection compliance requirements, grid connection technical specifications, and subsidy policy terms, thereby improving the accuracy of parsing new energy policy texts. In the indicator database module, a unique evaluation indicator system for the new energy industry is set, such as technical parameters such as photovoltaic module conversion efficiency, wind turbine available hours, energy storage system response time, and grid frequency regulation capability, as well as economic and environmental indicators including cost per kilowatt-hour, subsidy intensity, and carbon emission reduction. Users can enter specific parameters such as installed capacity, investment scale, and technical route. The platform will analyze the impact of various policies on project feasibility, economy, and technical route selection based on existing policy models and combined with the characteristics of the new energy market. For example, for distributed photovoltaic projects on rooftops throughout the county, the platform can evaluate the superposition effects of multi-dimensional policies such as land policies, electricity price subsidy policies, and grid connection policies to provide decision-making support for project development; for the construction of large-scale wind power bases, it can analyze policy elements such as land approval, environmental impact assessment requirements, and external transmission channels to predict key constraints on project advancement. At the same time, the platform can also model and analyze innovative policy mechanisms such as carbon trading policies and green electricity trading policies related to new energy to help companies grasp policy dividends. Through this professional policy analysis tool, the present invention can effectively support new energy companies in making scientific decisions under a rapidly changing policy environment and promote the healthy development of the new energy industry.

[0040] The calculation involved in the present invention is described in detail below.

[0041] 1. The word vector calculation process is as follows: ; In the formula, is the word vector; is the weight coefficient; For word features; is the bias term; is the feature dimension.

[0042] 2. The policy text vector matrix is ​​represented as follows: ; In the formula, is the text vector matrix; For the Document No. The eigenvalue of the dimension; is the number of documents; is the number of feature dimensions.

[0043] 3. The timing weight calculation equation is expressed as follows: ; In the formula, is the time series weight coefficient; is the current time; is the first appearance time; is the time attenuation coefficient, ranging from 0.1 to 0.5; is the frequency of occurrence; is the frequency reference value; is the weight coefficient, and ; is the adjustment coefficient; is the timing error term, ranging from -0.1 to 0.1.

[0044] 4. The position weight calculation equation is expressed as follows: ; In the formula, is the position weight coefficient; The maximum number of document levels; is the current level position; is the frequency of position occurrence; is the total frequency; is the level depth; is the adjustment coefficient, and ; is the position error term, ranging from -0.05 to 0.05.

[0045] 5. The semantic association calculation equation is expressed as follows: ; In the formula, is the semantic association coefficient; is the keyword vector; is the subject vector; is the co-occurrence frequency; is the co-occurrence weight coefficient; is the semantic error term, ranging from -0.1 to 0.1.

[0046] 6. The text similarity calculation equation is expressed as follows: ; In the formula, is the text similarity; are the feature vectors of the two documents; is the feature weight; is the similarity error term, ranging from -0.05 to 0.05.

[0047] 7. The policy influence score calculation formula is as follows: ; In the formula, Score points for influence; is the time decay weight; is the network association weight; is the policy level weight; is the influence weight coefficient, and ; is the score error term, ranging from -0.1 to 0.1.

[0048] Parameter acquisition method description: Obtained through training on a large corpus, using stochastic gradient descent optimization; Determine the optimal value through statistical analysis of historical data; The weights are determined through expert scoring; Obtained through co-occurrence matrix analysis; Determined through multiple rounds of cross validation.

[0049] Explanation of the equation principle: Word vector calculation adopts a combination of linear combination and nonlinear activation, taking into account the combined effect of features; temporal weight calculation adopts an exponential decay function, which reflects the nonlinear decay characteristics of timeliness over time; position weight calculation comprehensively considers hierarchical position, frequency of occurrence and depth information, and uses normalization to ensure weight comparability; semantic association calculation is based on the cosine similarity principle, and co-occurrence information is introduced to enhance semantic relevance; similarity calculation adopts weighted cosine similarity method, taking into account the differences in the importance of features; influence score calculation adopts a multi-dimensional weighted combination method to achieve a comprehensive influence assessment.

[0050] As a further preferred embodiment of the present invention, the platform has built a dynamic policy impact assessment mechanism specifically for the timeliness characteristics of renewable energy power generation. Specifically, in the process of extracting policy time series characteristic data, the system innovatively introduces the time-varying characteristic parameters of renewable energy power generation, and analyzes the correlation between the timeliness of the policy and the physical characteristics of renewable energy power generation. This mechanism is mainly reflected in the following aspects: First, in the time series weight calculation equation, the system introduces the periodic fluctuation function of renewable energy power generation. For photovoltaic power generation, the seasonal changes and diurnal fluctuation characteristics of sunshine intensity are considered, and the light resource conditions are associated with the timeliness of the policy. For example, the timeliness evaluation of the peak and valley electricity price policy will be dynamically adjusted in combination with the sunrise and sunset time of photovoltaic power generation; for wind power generation, the seasonal distribution characteristics of wind resources are introduced, and the annual utilization hours of wind farms are matched and analyzed with the policy implementation cycle. The time series weight calculation equation is improved to: ; in, is the comprehensive timing weight, is the base weight coefficient, is the time attenuation coefficient, is the seasonal adjustment function, is a resource availability function.

[0051] Secondly, the dynamic conversion relationship of new energy power generation efficiency is incorporated into the policy impact assessment model. The system has established a database containing technical parameters such as photovoltaic module conversion efficiency, wind turbine power curve, energy storage charging and discharging efficiency, etc. These parameters will be updated with technological progress, thus affecting the implementation effect of the policy. For example, when evaluating photovoltaic subsidy policies, the system will dynamically adjust the policy impact score based on the latest module efficiency and cost data. The model expression is: ; in, For policy In time The influence score of is the technical efficiency function, is the cost change function, is the importance weight of the policy itself.

[0052] In practical applications, the system also considers the relationship between photovoltaic power generation efficiency and temperature, and introduces a temperature correction coefficient: ; in, is the actual power generation efficiency, is the efficiency under standard test conditions, is the temperature coefficient, is the battery operating temperature, It is the standard test temperature (25℃).

[0053] For the power characteristics of wind power, the system uses a piecewise function to describe the power curve of the wind turbine: ; in, is the wind speed, For the cut-in wind speed, is the rated wind speed, To cut out the wind speed, is the air density, is the wind wheel swept area, is the wind energy utilization coefficient, is the rated power.

[0054] Again, in the process of generating the data update sequence, the grid absorption constraint coefficient is introduced: ; in, is the absorption constraint coefficient, is the grid load demand, For renewable energy power generation, is the grid regulation capacity coefficient.

[0055] Finally, the system constructs comprehensive evaluation indicators based on the above parameters: ; in, is the cumulative impact of the policy mix, is the policy set, is the evaluation time step.

[0056] Through this mathematical model system, the present invention realizes accurate quantitative evaluation of the timeliness of new energy policies and provides decision makers with a scientific analysis tool. The model not only takes into account the time attribute of policy implementation, but also incorporates the physical characteristics and technical constraints of new energy power generation, making the evaluation results more objective and accurate, and having strong practical guiding significance.

[0057] Specifically, the principle of the present invention is: the technical solution of the present invention is based on deep learning and natural language processing technology, and realizes the feature extraction and influence assessment of policy texts by constructing a multi-level intelligent analysis framework. In terms of structural recognition, a multi-layer encoder with a Transformer architecture is adopted, and the self-attention mechanism is used to capture the long-range dependencies of the text. The multi-head attention mechanism is used to realize the parallel extraction of structural features of different levels, so as to accurately identify the hierarchical structure of the policy text. In terms of feature extraction, the training of word vectors is realized based on the word embedding model, and the distributed representation of words is learned through the continuous bag of words model. The feature dimension is reduced by combining the principal component analysis method to realize the effective extraction of semantic features; at the same time, the language features, statistical features and format features are obtained through the multi-dimensional feature extraction module to form a comprehensive feature representation.

[0058] The core innovation of this invention is to establish a policy influence assessment model based on time series characteristics. The model first characterizes the timeliness characteristics of the policy through the time decay function, taking into account the difference between the policy release time and the current time; then constructs a policy association network based on text similarity, and calculates the association strength between policy nodes; finally, integrates time series characteristics, network characteristics and hierarchical characteristics to establish a comprehensive assessment model. This multi-factor fusion assessment method conforms to the actual characteristics of policy influence and can accurately reflect the importance and scope of policy influence.

[0059] A specific embodiment 1 of the policy mining module in the present invention is provided below. The specific implementation method of each step in this embodiment 1 is described in detail as follows.

[0060] The specific implementation method of step S01 is to perform data preprocessing on the acquired policy text data. The encoding conversion module is used to perform encoding consistency processing on the policy text data. The module first detects the original encoding format of the text, including encoding types such as ASCII, GB2312, and UTF8, and then converts the detected encoding to the UTF8 encoding format. During the conversion process, the buffer is read byte by byte to ensure the integrity of the characters. Then, the text filtering module is used to perform normalization processing. The module removes special control characters, continuous spaces and redundant line breaks in the text based on regular expression matching rules. At the same time, numbers are represented by unified Arabic numerals, the date format is unified to the year-month-day format, and the currency value is uniformly added with currency symbols and decimal places. The text segmentation module divides the text into natural paragraphs by identifying paragraph mark symbols and semantic boundaries, specifically including segmentation rules based on punctuation marks and segmentation rules based on semantic integrity, wherein the paragraph semantic integrity is guaranteed by calculating the semantic similarity between sentences, and the semantic similarity threshold is set to 0.6. Finally, the length of the segmented text is constrained by the length standardization module, the maximum number of characters in a single paragraph is set to 2,000, and the greedy matching algorithm is used to segment the overlong paragraphs at the semantic boundary to ensure that the segmented paragraphs maintain semantic coherence. Through the above processing, a standardized data foundation is provided for subsequent feature extraction and text analysis.

[0061] The specific implementation method of step S02 is to detect and correct text errors based on a preset text rule library. This step first loads the text rule library, which contains a set of rules such as punctuation usage specifications, paragraph format rules, and special character usage specifications. The text detection module uses a rule-based pattern matching algorithm to mark text that does not meet the specifications, and records the exception types through an error type dictionary, including punctuation errors, format errors, special character errors, and other types. Then the text correction module is called, which performs corrections based on a preset text dictionary. The text dictionary contains error correction resources such as standard word lists, homophone word lists, and similar word lists. The correction process uses an edit distance algorithm to calculate the similarity between the text to be corrected and the standard entries in the dictionary. The edit distance is calculated using the following formula: , where Represents a string Before Characters to String Before The edit distance of characters, is the replacement cost, when When the similarity exceeds the threshold of 0.85, the text to be corrected is replaced with the corresponding standard entry. For punctuation errors, a context-based rule matching algorithm is used for intelligent correction to ensure the standardization of punctuation. Finally, manual review is performed to ensure the accuracy of text correction.

[0062] The specific implementation of step S03 is to perform word segmentation and part-of-speech tagging on the standardized text. The word segmentation module adopts a hybrid word segmentation method based on a dictionary and statistics. First, a bidirectional maximum matching algorithm is used for preliminary word segmentation. The algorithm selects the optimal segmentation path through forward and reverse scanning. The segmentation probability calculation formula is: , where Indicates the first words, is the conditional probability. The word segmentation dictionary includes a domain-specific dictionary and a general dictionary, and the segmentation probability is calculated by weighted combination. The word segmentation results are disambiguated, and a sequence labeling model based on conditional random fields is used to identify and process ambiguous words. The characteristic function of the conditional random field is defined as: , where is the part-of-speech tag, is the observation sequence, is the position index, is the feature function index. The part-of-speech tagging module uses a hidden Markov model. The state transition probability matrix is ​​obtained through statistical training, and the emission probability matrix is ​​calculated based on the part-of-speech tagging corpus. The tagging results are post-processed to merge adjacent part-of-speech tags of the same type and adjust the part-of-speech tags of special words. Finally, the policy keyword sequence is extracted. Keyword extraction is based on part-of-speech tag screening and tf-idf algorithm. The tf-idf value calculation formula is: , where is the word frequency, is the document frequency, The total number of documents.

[0063] The specific implementation of step S04 is to construct a policy word vector. First, a word vector training corpus is established, which contains a large amount of policy text data. The bag-of-words model is used to vectorize the text, and the word vector calculation adopts the following formula: , where is the word vector, is the weight coefficient, For word features, is the bias term, is the feature dimension. Then the word embedding model is used to train the word vector, and the continuous bag of words model is used to learn the word vector. The objective function of the model is: , where is the sequence length, is the context window size, The training parameters include word vector dimension of 300, context window size of 5, and negative sampling number of 5. The word vector obtained by training is reduced in dimension, and the principal component analysis method is used to compress the word vector to 100 dimensions, retaining the main semantic information. Finally, each word in the policy keyword sequence is mapped to the corresponding word vector to form a policy feature vector matrix.

[0064] The specific implementation of step S05 is to extract multi-dimensional features. The language feature extraction module calculates the part-of-speech distribution features, and the distribution feature vector calculation formula is: , where For the The frequency ratio of the class part of speech. The syntactic structure features are obtained through dependency syntactic analysis, and the graph structure is used to represent the syntactic relationship. The edge weight calculation formula is: , where is the word distance, is the relationship strength. The semantic relationship features are obtained by constructing a semantic network. Nodes represent concepts, edges represent semantic relationships, and the relationship strength is calculated by co-occurrence frequency. The statistical feature extraction module calculates the word frequency features. The word frequency distribution is modeled using the Zipf distribution, and the distribution function is: , where is the frequency ranking of words, is the distribution parameter, is a normalization constant. The word length feature is obtained by counting the word length distribution, and the character distribution feature is obtained by calculating the frequency of character n-tuples. The format feature extraction module analyzes the paragraph structure features, including paragraph length distribution, paragraph hierarchy, etc.; the punctuation usage feature is obtained by counting the punctuation usage frequency and position distribution; the space distribution feature is obtained by analyzing the space position and length. Finally, all kinds of features are integrated to form a policy multidimensional feature vector.

[0065] The specific implementation of step S06 is to calculate the text similarity. First, the policy multidimensional feature vector is standardized by using the range standardization method. The standardization formula is: Then the weighted cosine similarity between the feature vectors is calculated. The similarity calculation formula is: , where are the feature vectors of the two documents, is the feature weight, is the similarity error term. The similarity matrix is ​​threshold filtered, and the similarity threshold is set to 0.7. Similarity values ​​below the threshold are set to 0. Finally, the similarity matrix is ​​normalized to ensure the comparability of the similarity values.

[0066] The specific implementation of step S07 is to build a policy parsing model. First, a multi-layer encoder is designed based on the Transformer architecture of the attention mechanism. The number of encoder layers is set to 6, and each layer contains a multi-head self-attention layer and a feedforward neural network layer. The attention calculation formula is: , where is the query matrix, is the key matrix, is the value matrix, is the scaling factor. Multi-head attention is achieved by computing multiple attention results in parallel, and the calculation formula is: , where is the output of a single attention head. Then, the clause structure recognition layer is designed, and a hierarchical structure recognition strategy is adopted to sequentially recognize the text structure of the chapter level, clause level, and item level. The hierarchical recognition adopts the conditional random field model, and the feature functions include text format features, position features, and context features. The recognition results are post-processed to merge the fragmented structural units and adjust the abnormal hierarchical relationships. Finally, the hierarchical policy document structure data is generated, which includes structured hierarchical relationships and text content.

[0067] The specific implementation method of step S08 is to build a data warehouse. First, a hybrid storage architecture is designed. A relational database is used in the first database to store structured data. The data table design follows the third normal form, mainly including policy basic information table, policy structure table, policy relationship table, etc. In the second database, a document database is used to store semi-structured data. The document model design supports flexible data formats and hierarchical structures. Then a multidimensional index system is established, including a primary key index based on a B+ tree, a full-text retrieval index based on an inverted index, and a location retrieval index based on a spatial index. Finally, a data synchronization mechanism is implemented, and a dual-write consistency protocol is used to ensure data consistency between the two databases. The atomicity of the synchronization process is guaranteed by distributed transactions.

[0068] The specific implementation of step S09 is to calculate the contribution of policy keywords. The word frequency weight is calculated using the tf-idf algorithm, taking into account the importance of the word in the document. The position weight is calculated based on the position distribution of the keyword in the document, and the position weight calculation formula is: , where is the maximum number of document levels, is the current level position, is the frequency of position occurrence, is the total frequency, is the hierarchical depth. The semantic association weight is obtained by calculating the semantic similarity between the keyword and the document topic. The semantic association calculation formula is: , where is the keyword vector, is the theme vector, is the co-occurrence frequency.

[0069] The specific implementation of step S10 is to construct a policy association network. First, the association relationship between policy nodes is established based on the text similarity value, and the policy documents with similarity greater than 0.7 are connected. Then the association strength between nodes is calculated, and the association strength calculation formula is: , where is the text similarity, For the release time, is the hierarchical similarity, is the weight coefficient. Finally, the associated network is optimized to remove weakly associated connections, and the optimized network meets the small-world network characteristics.

[0070] The specific implementation method of step S11 is to perform data allocation. First, a data allocation matrix is ​​constructed, and the matrix element calculation formula is: , where is the update frequency value, is the priority value, is the weight coefficient. For data with high update frequency and high priority, it is allocated to the high-performance storage area and stored in solid-state hard disks; for data with low update frequency and low priority, it is allocated to the ordinary storage area and stored in mechanical hard disks. Then the data scheduling mechanism is implemented. The scheduling strategy is based on the load balancing algorithm. The load calculation formula is: , where is the processor occupancy, is the memory usage, is the input and output occupancy rate. Finally, a data backup strategy is established, including scheduled full backup and incremental backup, and the backup cycle is determined by priority.

[0071] The specific implementation of step S12 is to extract time series features. First, a time series is constructed based on the policy release time, and the time series is expressed as: , where is the timestamp. Then calculate the time decay weight, and the decay weight calculation formula is: , where is the current time, is the first appearance time, is the time attenuation coefficient, is the frequency of occurrence, is the frequency reference value. Prioritize the policy data according to the time decay weight and build the data update sequence. Finally, establish the update strategy, and the update frequency is inversely proportional to the time decay weight.

[0072] The specific implementation of step S13 is to realize archival storage. First, establish archiving rules, set archiving thresholds based on time decay weights, and trigger archiving operations when the weight value is lower than 0.3. The archiving process adopts a tiered storage strategy, and the storage tier division formula is: , where is the time decay weight, is the weight range, is the number of storage levels. Then data compression is performed, and the compression ratio calculation formula is: , using lossless compression algorithm to ensure data integrity. Then, an archive index is established. The index structure is implemented using B-tree to support efficient range queries. Finally, a regular cleanup mechanism for archived data is implemented, and the cleanup cycle is proportional to the storage level.

[0073] The specific implementation of step S14 is to calculate the influence score. First, the association strength value and time decay weight value of the policy node are obtained, and the two indicators are standardized. Then, an influence calculation model is designed, and the influence score calculation formula is: , where is the time decay weight, is the network association weight, is the policy level weight, is the influence weight coefficient. According to the influence score, the storage level identifier is set. The level division uses a clustering algorithm to divide the policy into three storage levels: high, medium, and low. Finally, the calculation results are verified and the prediction accuracy and recall rate of the model are calculated.

[0074] The specific implementation of step S15 is to generate an analysis report. First, the policy texts are sorted based on the influence score, and the ranking score calculation formula is: , where Score for influence, For timeliness score, The policy analysis report is then generated, which includes basic policy information, structural characteristics, correlations, impact assessment, etc. The analysis text is automatically generated using natural language generation technology. The impact score and storage level identifier are then stored in the indicator database, and an index structure for the indicator data is established. Finally, the policy text is stored hierarchically according to the storage level identifier, and the storage strategy optimizes resource allocation based on the hierarchical storage model. Through the above steps, intelligent analysis and scientific management of policy texts are achieved, providing data support for policy decisions.

[0075] The specific implementation method of step S16 is to establish an indicator monitoring system. First, a multidimensional indicator system is designed, including dimensions such as time series indicators, influence indicators, and storage indicators. The indicator value calculation adopts a sliding window method, and the window size is determined according to the indicator characteristics. Then, an indicator calculation engine is implemented to support both real-time calculation and batch calculation modes. The consistency of the calculation results is guaranteed by version control. Then, indicator monitoring rules are established, indicator thresholds and alarm levels are set, and alarms are triggered when the indicator value exceeds the threshold range. Finally, a visual display is implemented, using multi-dimensional charts to display indicator change trends, supporting interactive data analysis and drilling.

[0076] In order to better understand and implement the present invention, the following provides an example 2 of a specific application scenario of the present invention: A research institute intends to mine electricity-related policy data and establishes an energy policy mining and intelligent interaction platform based on an AI big model. Figure 2-4 As shown in the figure, during the data preprocessing stage, the project team carried out unified encoding conversion and format standardization on the collected policy texts. After processing by the text filtering module, 28 special characters, redundant spaces and line breaks totaling 1.56 million were identified, and 150,000 digital formats, 80,000 date formats, and 50,000 currency formats were standardized. The text segmentation module divided the policy text into 1.56 million natural paragraphs, of which 98% of the paragraph lengths were controlled within 2,000 characters.

[0077] In the text analysis phase, the project team used a pre-trained word vector model for feature extraction. The training corpus of the word vector model contains 500 million words of policy texts. The word vector dimension is set to 300, the context window size is 5, and the number of negative samples is 5. After 100 rounds of training, the model has an accuracy of 85% on the word similarity test set. After compressing the word vector to 100 dimensions through the principal component analysis method, 92% of the semantic information is retained.

[0078] The policy parsing model with Transformer architecture is used in the structural recognition stage. The encoder is set to 6 layers, and each layer contains 8 attention heads. The model is trained on 5,000 annotated data, using the cross entropy loss function, the learning rate is set to 0.0001, the batch size is 32, and the structural recognition accuracy on the test set reaches 89% after 500 rounds of training. The main recognition results are shown in Table 1 below: Table 1 Recognition results

[0079] In the feature extraction stage, multi-dimensional feature data was obtained. In terms of language features, the distribution ratios of 15 types of parts of speech were counted, and a semantic network containing 500,000 nodes was constructed. In terms of statistical features, the word frequency distribution conforms to the Zipf distribution, with a distribution parameter s of 1.2 and a correlation coefficient of 0.95. In terms of format features, the paragraph length distribution was analyzed, with an average paragraph length of 526 characters and a standard deviation of 168 characters.

[0080] In the temporal feature extraction stage, the time decay weight was calculated based on the policy release time, the decay coefficient λ was set to 0.1, and the frequency reference value f0 was set to 10 times per month. The policy association network contains 12,000 nodes, 78,000 edges, and the average node degree is 13. In the calculation of association strength, the text similarity weight α is 0.4, the time difference weight β is 0.3, and the level similarity weight γ is 0.3.

[0081] The data storage adopts a hybrid storage architecture. The first database uses MySQL to store structured data and establishes 15 data tables. The second database uses MongoDB to store semi-structured data, with 8 collections. The multidimensional index includes primary key index, full-text index and time index, with a total index size of 15GB. The data synchronization adopts a dual-write consistency protocol, and the synchronization delay is controlled within 100ms.

[0082] The impact assessment results show that among all policies, 15% are high-impact policies, 45% are medium-impact policies, and 40% are low-impact policies. The weight coefficients of the impact assessment model are set as follows: time weight θ is 0.3, network weight φ is 0.4, and layer weight ψ is 0.3. The accuracy of the assessment results is 87%, the recall rate is 85%, and the F1 value is 0.86.

[0083] Based on the evaluation results, the system automatically generates a policy analysis report. The report contains basic statistical information of the policy text, structural feature analysis, correlation network diagram, impact evaluation results, etc. The analysis report is generated by combining templates and natural language generation, supporting multi-dimensional data visualization.

[0084] The project established a real-time monitoring system and set 12 core indicators, including 4 time series indicators, 5 impact indicators, and 3 storage indicators. The indicator calculation uses a 5-minute sliding window, and the delay of real-time calculation results is controlled within 1 second. The alarm accuracy of the monitoring system reached 92%, and the false alarm rate was controlled within 5%.

[0085] The traditional technologies used mainly have the following problems: the rule-based structure recognition method has a low accuracy rate of about 65%; feature extraction only considers word frequency and position features, and lacks the extraction of semantic features; influence assessment is based only on simple statistical indicators, and the assessment accuracy rate is about 60%. The technical solution of the present invention has achieved significant progress in the following aspects: high-accuracy structure recognition is achieved through the Transformer architecture, with the accuracy rate increased to 89%; comprehensive text feature representation is obtained through multi-dimensional feature extraction, and the feature dimension is expanded from the original 10 dimensions to 100 dimensions; the scientific nature of the assessment is improved through the multi-factor fusion influence assessment model, and the accuracy rate is increased to 87%. These improvements have significantly improved the effect of policy text analysis and provided more reliable data support for policy decisions.

[0086] Through this implementation, the system solves the core technical problems of policy text feature extraction and impact assessment, and realizes the intelligent and scientific analysis of policy text. The subsequent project team will further optimize the model parameters, expand the application scenarios, and improve the overall performance of the system. According to user feedback, the system has received high evaluations in terms of policy analysis efficiency, accuracy, and usability.

[0087] It should be noted that the variables involved in the present invention are explained in detail as shown in Table 2 and Table 3 below.

[0088] Table 2 Variable explanation table (Part I)

[0089] Table 3 Variable explanation table (Part II)

[0090] The platform in the embodiment can be implemented using electronic devices, such as building a server architecture and deploying the platform of the present invention. Users can interact intelligently with the platform through the input and output of the interface, such as the buttons for policy mining, policy information library, and indicator database in the accompanying drawings, to obtain the impact of the policy, thereby reasonably arranging the company's own production and operation, such as the distribution of current and past electricity consumption, and reducing company costs.

[0091] The policies in the above embodiments may be national policies, corporate / non-profit organization policies, or intervention policies for other resources, such as mines, water conservancy, etc. The above are only specific implementations of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered within the protection scope of the present invention.

Claims

1. A policy mining and intelligent interaction platform based on AI big model, characterized by: It includes a policy mining module, a policy information base module and an indicator database module; the policy mining module is used to obtain policy text data, perform feature extraction, structure recognition and influence analysis on the policy text data, and generate a policy analysis report; the policy information base module is used to classify, store and manage the policy text data, build a policy data warehouse, and realize archival storage of the policy text data; the indicator database module is used to store policy time series feature data, policy influence scores and policy storage level identifiers; the policy mining module is used to perform the following steps: obtain policy text data, perform feature extraction and structure recognition on the policy text data, wherein feature extraction includes Perform word segmentation and part-of-speech tagging, extract the policy keyword sequence in the policy text data, establish a policy text vector matrix, convert the policy keyword sequence into a policy word vector, and generate the policy feature vector; construct a Transformer architecture policy parsing model, perform multi-level clause structure recognition on the policy text data, and generate the policy document structured data; extract the policy time series feature data, calculate the policy time decay weight value according to the policy release time, form the data update sequence, fuse the policy node association strength value and the policy time decay weight value, calculate the policy influence score, and generate the policy storage level identifier.

2. According to the AI ​​big model-based policy mining and intelligent interaction platform of claim 1, it is characterized by: The steps to obtain policy text data are to process the policy text data for encoding consistency, convert all texts into UTF8 encoding format, filter the text, remove special characters, redundant spaces and line breaks in the text, normalize the numbers, dates and currency formats in the text, segment the text, identify paragraph marks in the text, segment the text according to the natural paragraph structure of the policy text, standardize the text length, truncate overly long paragraphs, and set the maximum number of characters for a single paragraph to 2000.

3. The policy mining and intelligent interaction platform based on AI big model according to claim 1 is characterized in that: The steps for generating the policy feature vector are specifically to establish a word vector training corpus, use the bag-of-words model to vectorize the text, use the word embedding model to train the word vector, select the continuous bag-of-words model for word vector learning, set the word vector dimension to 300 dimensions, the context window size to 5, the number of negative sampling to 5, perform dimensionality reduction processing on the trained word vector, and use the principal component analysis method to compress the word vector to 100 dimensions.

4. According to claim 1, a policy mining and intelligent interaction platform based on AI big model is characterized in that: The steps for constructing the Transformer architecture policy parsing model are to design a multi-layer encoder based on the attention mechanism. The number of encoder layers is set to 6, and each layer includes a multi-head self-attention layer and a feedforward neural network layer. A clause structure recognition layer is designed, and a hierarchical structure recognition strategy is adopted to sequentially identify the text structures of the chapter level, clause level, and item level. The recognition results are post-processed, the fragmented structural units are merged, and the abnormal hierarchical relationships are adjusted.

5. According to claim 1, a policy mining and intelligent interaction platform based on AI big model is characterized in that: The step of calculating the contribution value of the policy keyword is also included, and the policy keyword contribution value is obtained through a contribution calculation equation group, and the contribution calculation equation group includes a time series weight calculation equation, a position weight calculation equation and a semantic association calculation equation. The time series weight calculation equation is used to calculate the value of the keyword timeliness impact factor, the position weight calculation equation is used to calculate the keyword structure importance value, and the semantic association calculation equation is used to calculate the keyword topic relevance value; The user stores the level identifier according to the policy and adjusts the energy consumption or power consumption of the relevant industries.

6. According to claim 1, a policy mining and intelligent interaction platform based on AI big model is characterized in that: It also includes the steps of building a policy association network, specifically establishing associations between policy nodes based on text similarity values, connecting policy documents with a similarity greater than 0.7, calculating the association strength between nodes, considering text similarity, release time differences, and policy level factors, comprehensively calculating the association strength value, optimizing the association network, and removing weakly associated connections.

7. According to claim 1, a policy mining and intelligent interaction platform based on AI big model is characterized in that: It also includes the steps of establishing policy archiving rules, specifically setting an archiving threshold based on the time decay weight, triggering the archiving operation when the weight value is lower than 0.3, compressing the archived data, establishing an archiving index, implementing a regular cleanup mechanism for archived data, and securely deleting expired data.

8. According to claim 1, a policy mining and intelligent interaction platform based on AI big model is characterized in that: It also includes the step of building a data warehouse, specifically designing a hybrid storage system that includes structured data and semi-structured data, using a relational database to store standardized policy data in the first database, using a document database to store policy data in a flexible format in the second database, establishing a multi-dimensional index, and implementing a data synchronization mechanism.

9. According to claim 1, a policy mining and intelligent interaction platform based on AI big model is characterized in that: It also includes the steps of building a multidimensional indicator data management system, specifically designing an indicator data structure that includes policy time series characteristics, influence indicators, and storage levels, establishing an indicator calculation engine, supporting real-time calculation and updating of indicator data, implementing version management of indicator data, recording the change history of indicator values, and establishing a monitoring mechanism for indicator data.

10. The policy mining and intelligent interaction platform based on AI big model according to claim 1 is characterized in that: It also includes the steps of generating a policy analysis report, specifically, sorting the policy texts based on the influence scores, generating an analysis report containing the basic information, structural characteristics, correlation relationships, and impact assessment of the policy texts, storing the influence scores and storage level identifiers in the indicator database, and storing the policy texts in a hierarchical manner according to the storage level identifiers.

Citation Information

Cited By

  • Policy document intelligent rule extraction and change comparison method for electricity charge verification

    CN120449861A

  • Green watershed policy collaboration analysis system and method

    CN120994709A

  • Large model question answering system and method for business hall

    CN121413776A