Industrial policy information service system and method based on intelligent analysis
The industrial policy information service system with intelligent analysis solves the problems of data dispersion and superficial analysis in the traditional model, realizes the structuring and precise matching of policy data, improves the pertinence and timeliness of policy services, and increases the success rate of enterprise applications.
Patent Information
- Application Number
- CN202510979225.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional industrial policy information service model lacks pertinence and timeliness in policy services due to the scattered data acquisition and superficial correlation analysis, making it difficult to accurately target corporate needs.
We adopt an industrial policy information service system based on intelligent analysis, obtain multi-source policy data through crawler technology, perform pre-processing, character distribution analysis, four-layer encoding and mapping analysis, generate customized service solutions, and dynamically update them to adapt to changes in corporate needs.
It has achieved the structuring and precise matching of policy data, improved the pertinence and timeliness of policy services, increased the success rate of enterprise applications, and formed a virtuous cycle of two-way empowerment between policies and enterprises.
Smart Images

Figure CN120806871A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent analysis, and particularly relates to an industrial policy information service system and method based on intelligent analysis. BACKGROUND
[0002] In the era of global industrial chain restructuring and policy-driven innovation, industrial policy, as a core tool for guiding resource allocation and solving enterprise development pain points, plays a decisive role in helping enterprises break through technical bottlenecks (such as new energy battery research and development) and seize market opportunities (such as artificial intelligence computing power layout). Due to structural defects in traditional industrial policy information service modes, such as scattered data acquisition defects and surface-level correlation analysis defects, enterprises are trapped in the double dilemma of difficulty in finding and using policies, and policy providers and demanders are out of touch and optimization is lagging. That is, due to the lack of data centralized analysis and intelligent processing of policy information, it is difficult to accurately lock the needs of enterprises, resulting in insufficient pertinence and timeliness of policy services.
[0003] Therefore, the present application provides an industrial policy information service system and method based on intelligent analysis. SUMMARY
[0004] The present application provides an industrial policy information service system and method based on intelligent analysis to solve the above technical problems.
[0005] The present application provides an industrial policy information service system based on intelligent analysis, comprising: A preprocessing module is configured to obtain industrial policy information from a designated platform through a crawler technology, obtain multi-source policy data of different types based on subscription information related to the designated platform, and preprocess the multi-source policy data to obtain a standardized policy data set. A character distribution analysis module is configured to perform character screening on the standardized policy data set based on character screening rules to obtain service characters, determine overall character distribution and same-type character distribution, and obtain a distribution array of same-type characters, wherein the distribution array includes distribution proportion and official importance of same-type characters. A first encoding module is configured to mine the same-type characters from three dimensions of semantics, logic and time to obtain correlation analysis results, and perform first encoding on the same-type characters using a four-layer encoding mechanism. A second encoding module is configured to crawl service fields of the industrial policy information, extract historical iteration orders of same-type characters in the service fields, and generate second encoding. The mapping analysis module is configured to collect multi-dimensional demand data of the target enterprise, convert the multi-dimensional demand data into a keyword vector, establish a mapping relationship based on the distribution array, the first encoding, and the second encoding, and generate a customized service scheme; The updating module is configured to obtain policy docking effect data and demand change data of the target enterprise based on the customized service scheme, and update the character screening rule and the first encoding and the second encoding.
[0006] Preferably, the preprocessing module comprises: The homogenization conversion unit is configured to configure a format parsing adapter for a corresponding specified platform according to a heterogeneous data format of each specified platform and in combination with a latest credibility weight of each specified platform, so as to realize homogenization conversion of the heterogeneous data, wherein the credibility weight of the subscription information set by each specified platform is dynamically adjusted in combination with a conflict rate of historical data, and when multi-source subscription information pushes the same policy, a weighted fusion algorithm is used to eliminate the conflict. The mapping unit is configured to map synonyms to standard terminology based on the homogenization conversion result and in combination with an industrial policy knowledge graph. The triple construction unit is configured to identify a condition-result logical pair related to a policy clause in the homogenization conversion result through syntactic analysis, and construct a structured triple, wherein the structured triple comprises a condition relationship and a result. The supplement unit is configured to automatically supplement ambiguous expressions in the homogenization conversion result according to a domain classification standard to obtain specific categories. The data set acquisition unit is configured to obtain a standardized policy data set based on the mapping result, the structured triple, and the specific categories.
[0007] Preferably, the character distribution analysis module comprises: The first candidate unit is configured to calculate semantic similarity of characters and core domain labels, and determine first candidate characters in the standardized policy data set in combination with TF-IDF values. The second candidate unit is configured to identify characters of a strong constraint modifier word appearing in a policy chapter in the standardized policy data set as second candidate characters based on text structure analysis. The third candidate unit is configured to compare comparison characters in the standardized policy data set with a sample enterprise historical demand keyword library for similarity to obtain third candidate characters. The character unit is configured to obtain service characters based on the first candidate characters, the second candidate characters, and the third candidate characters. The feature determination unit is configured to perform global analysis on the standardized data set according to service characters based on dynamic word cloud time series evolution and theme clustering to obtain overall character distribution features. A hierarchical clustering unit is configured to hierarchically cluster the service characters based on policy publishing departments and industry fields to obtain characters of the same type.
[0008] Preferably, the character distribution analysis module further comprises: A proportion determination unit is configured to calculate a frequency proportion of the characters of the same type in the policy cluster to which the characters belong, eliminate short-term fluctuations by an exponential smoothing method, and calculate a distribution proportion according to a policy publishing region; An importance determination unit is configured to construct an evaluation matrix of publishing subjects, clause constraint strength, and implementation range, and obtain official importance of the characters of the same type. The distribution array is obtained based on the distribution proportion and the official importance.
[0009] Preferably, the first encoding module comprises: A semantic analysis unit is configured to define policy entities, attributes, and association relationships, construct an initial graph by human-computer collaborative labeling, cross-modally fuse embedded representations of the characters of the same type with triples of the initial graph, and calculate semantic correlation degrees between the characters. A logical analysis unit is configured to split a policy into atomic clauses based on a text segmentation algorithm, extract condition-action-result structures of the atomic clauses based on a dependency syntax tree to generate logical tuples, and perform compliance verification on the logical tuples. A time analysis unit is configured to perform time series analysis on historical publishing time, revision time, and abolishment time of the characters of the same type, and determine an evolution track table. A mechanism encoding unit is configured to perform first encoding on the characters of the same type based on association analysis results and a four-layer encoding mechanism based on a field dimension, a function dimension, a feature dimension, and an instance dimension. The association analysis results include semantic correlation degrees between the characters based on a semantic dimension, compliance verification results of the logical tuples based on a logical dimension, and the evolution track table based on a time dimension.
[0010] Preferably, the mapping analysis module comprises: A demand collection unit is configured to collect explicit demands, implicit demands, and potential demands of a target enterprise. The explicit demands are obtained from an enterprise-side interactive platform based on a structured form and a free text enhanced mode. The implicit demands are obtained based on enterprise annual reports, social responsibility reports, and interview recordings. The potential demands are predicted based on mining results of the time dimension of the first encoding module. A model analysis unit is configured to input the explicit demand, the implicit demand and the potential demand into an industrial policy field enhanced large model to output a keyword vector, wherein the industrial policy field enhanced large model is based on a general BERT model, incorporates policy-demand and training corpus, and is trained by using a life cycle label mask of a need scene mask language model and a keyword vector; An extraction and determination unit is configured to extract statistical patterns of the distribution array based on a multi-layer perception, determine inter-code semantic connections of a first code based on a graph convolution network, and capture policy evolution rules of a second code based on a time convolution network; A relationship determination unit is configured to project the keyword vector, the statistical pattern, the inter-code semantic connection and the policy evolution rule into a three-dimensional coordinate system respectively, determine a matching relationship based on a matching policy, a matching relationship based on a policy condition dependency chain, and a matching relationship based on a risk point of a regional difference between a historical eliminated character and the distribution array of the second code; A scheme formulation unit is configured to obtain a strategy for coping with each matching relationship, and generate a customized service scheme.
[0011] Preferably, the second coding module comprises: A difference analysis unit is configured to compare version differences of crawled industrial policy information in a service field by a text difference algorithm, mark addition, modification and deletion operations of the same type of characters, and generate a character iteration track table, wherein the character iteration track table contains operation time and an influence range of each operation; A code generation unit is configured to generate a second code of the same type of characters based on the character iteration track table.
[0012] The present application provides an industrial policy information service method based on intelligent analysis, comprising: Step 1: Obtain industrial policy information from a designated platform through a crawler technology, obtain multi-source policy data of different types based on subscription information related to the designated platform, and obtain a standardized policy data set by preprocessing the multi-source policy data; Step 2: Perform character screening on the standardized policy data set based on character screening rules to obtain service characters, determine overall character distribution and same type character distribution, and obtain a distribution array of same type characters, wherein the distribution array comprises a distribution proportion and an official importance of the same type characters; Step 3: Mine the same type characters from three dimensions of semantics, logic and time to obtain correlation analysis results, and perform first coding on the same type characters by using a four-layer coding mechanism; Step 4: Crawl the service field of the industrial policy information, and extract the historical release of the same type of characters in the service field to generate a second code; Step 5: Collect multi-dimensional demand data of the target enterprise, and convert the multi-dimensional demand data into a keyword vector, establish a mapping relationship based on the distribution array, the first code and the second code, and generate a customized service scheme; Step 6: Obtain policy docking effect data and demand change data of the target enterprise based on the customized service scheme, and update the character screening rule and the first code and the second code.
[0013] Compared with the prior art, the beneficial effects of the present application are as follows: Through the full-link design of multi-source collection-intelligent screening-three-dimensional correlation-iterative coding-precise matching-closed loop optimization, the upgrading of policy data from fragmentation to structuring, character analysis from surface statistics to deep correlation, and demand matching from experience recommendation to intelligent mapping is realized, which improves the accuracy of policy service and the success rate of enterprise declaration, and through the dynamic updating mechanism, the policy iteration and enterprise demand change are continuously adapted, forming a virtuous cycle of policy-enterprise bidirectional empowerment, and ensuring the pertinence and timeliness of policy service.
[0014] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and achieved by the structures specifically pointed out in the written description, claims, and drawings.
[0015] The technical solutions of the present application will be further described in detail below with the help of the drawings and examples. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings are used to provide a further understanding of the present application, and constitute a part of the specification, together with the embodiments of the present application, to explain the present application, and do not constitute a limitation on the present application. In the drawings: Figure 1 It is a structure diagram of an industrial policy information service system based on intelligent analysis in an embodiment of the present application; Figure 2 It is a flowchart of an industrial policy information service method based on intelligent analysis in an embodiment of the present application. DETAILED DESCRIPTION
[0017] The preferred embodiments of the present application will be described below in conjunction with the drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and do not limit the present application.
[0018] The present invention provides an industrial policy information service system based on intelligent analysis, such as Figure 1 As shown, including: A preprocessing module is used to obtain industrial policy information from a designated platform using crawler technology, obtain different types of multi-source policy data based on subscription information related to the designated platform, and preprocess the multi-source policy data to obtain a standardized policy data set; a character distribution analysis module, configured to perform character screening on the standardized policy dataset based on character screening rules to obtain service characters, and determine the overall character distribution and the distribution of characters of the same type to obtain a distribution array of characters of the same type, wherein the distribution array includes: a distribution percentage and an official importance of characters of the same type; A first encoding module is used to mine the same type of characters from the three dimensions of semantics, logic and time, obtain association analysis results, and perform a first encoding on the same type of characters using a four-layer encoding mechanism; A second encoding module is used to crawl the service field of the industrial policy information, extract the iterative order of the same type of characters under the historical release of the service field, and generate a second code; a mapping analysis module for collecting multidimensional demand data of a target enterprise, converting the multidimensional demand data into a keyword vector, establishing a mapping relationship based on the distribution array, the first code, and the second code, and generating a customized service solution; The updating module is used to obtain the policy docking effect data and demand change data of the target enterprise based on the customized service solution, and to update the character screening rules and the first code and the second code.
[0019] In this embodiment, crawler technology refers to intelligent data acquisition technology based on a targeted, focused crawler framework (e.g., Scrapy + Selenium). It crawls policy text from a specific platform using a pre-set URL seed list (e.g., links to the "Policies and Regulations" section of an official website) combined with anti-crawling strategies (e.g., dynamic User-Agent switching and IP proxy pool rotation). Implementation requires configuring character matching rules (e.g., the regular expression "subsidy|tax exemption|qualification certification") to only crawl pages containing policy keywords, with a crawling accuracy of ≥90%. For example, when crawling the "2025 Strategic Emerging Industries Support Policy," the crawler automatically avoids pop-up ads and irrelevant information, extracting only the policy clauses within the main text and appendices. The designated platform refers to the official channel with policy release authority, including national platforms (such as national policy document library), provincial platforms (such as "Guangdong Provincial Affairs" government platform), industry platforms (such as China Automobile Industry Association website), and park platforms (such as Suzhou Industrial Park "Enterprise Policy" column). For example, new energy enterprises need to focus on the designated platform, including the "New Energy Vehicle Industry Development Plan" column and the Ministry of Finance's "Energy Conservation and Emission Reduction Subsidy" page. Subscription information refers to the real-time data push protocol bound to the designated platform. Incremental data is obtained through the API interface of the platform (such as the "Policy Update Webhook" of the government platform), including policy title, release time, abstract, and other metadata. When implementing, the platform's digital signature (such as SM2 national encryption algorithm verification) needs to be verified to ensure the authenticity of the information. For example, after subscribing to the "Science and Technology Department API", when the "New R&D Fee Addition Deduction Policy" is released, the system will receive the metadata in real time and trigger the full-text crawling. Multi-source policy data refers to a collection of policy texts from different platforms and different types, which can be divided into categories such as finance and taxation (such as "Enterprise Income Tax Reduction Policy"), qualification (such as "High-tech Enterprise Identification Method"), and innovation (such as "Key R&D Plan Application Guide"). For example, the data collected by a certain system covers "Tax Incentive Policy", "Science and Technology Department's "Science and Technology Small and Medium-sized Enterprise Evaluation Policy", and local government's "Industrial Park Entry Subsidy Policy". In this example, the standardized policy data set refers to the final data set after merging the mapping results, structured triples, and specific categories, and the preprocessing is the operation process corresponding to the merging of mapping results, structured triples, and specific categories.
[0020] In this example, the character screening rule refers to a multi-dimensional screening logic that integrates semantic similarity, policy weight, and enterprise demand, specifically: Semantic layer: Calculate the cosine similarity (≥0.7) between characters and "core field labels" (such as "intelligent manufacturing") through pre-trained BERT model; Policy layer: Keep characters modified by strong constraint words such as "must" and "should" (such as "environmental protection standards" in "must meet environmental protection standards"); Demand layer: Similarity ≥0.65 with sample enterprise historical demand keyword library (such as "financing difficulties" and "technology upgrading"). When implementing, integrate the three elements through the gradient boosting tree model (XGBoost) to output service characters. Service characters refer to policy core words that have direct value for enterprise decision-making, usually involving rights, constraints, or resource allocation, such as "R&D subsidy amount", "high-tech enterprise application conditions", and "tax reduction ratio". For example, from the "Software Enterprise Value-added Tax Refund Policy", the service characters "value-added tax refund rate 13%" and "annual income ≥1000 million" are selected. Overall character distribution refers to the global characteristics of service characters in the full amount of policy data, including frequency of occurrence (such as "subsidy" appearing in 30% of policies), text location distribution (such as 60% of "qualification recognition" located in the policy "Chapter 2 Reporting Conditions"), and visualization through dynamic word clouds (word size mapping frequency) and heat maps. Same type character distribution refers to the character characteristics after clustering by industry field (such as "new energy" "biomedicine") or policy function (such as "financial support" "talent introduction"). When implemented, hierarchical clustering algorithm is used to classify "photovoltaic subsidy" "wind power subsidy" "hydrogen energy subsidy" as "new energy subsidy" character cluster. Distribution array refers to structured data describing the characteristics of the same type of characters, containing two core dimensions: "distribution proportion" and "official importance", in the format [character ID, distribution proportion, official importance].
[0021] Distribution proportion refers to the probability of a character appearing in its type cluster (such as "photovoltaic subsidy" accounting for 25% of "new energy subsidy" characters), calculated through a sliding time window (quarterly granularity) to eliminate short-term fluctuations. Official importance refers to the policy effectiveness weight of a character, calculated by weighting "release department level (national level 10 points / provincial level 7 points) + clause constraint strength (mandatory clause x 1.5) + implementation range (national x 1.2)", for example, the official importance of "R&D expense addition deduction ratio 175%" in a national policy = 10 (level) x 1.5 (mandatory) x 1.2 (national) = 18 points. In this example, the three dimensions of semantics, logic, and time are: Semantic dimension: refers to the synonymous, near-synonymous or hierarchical relationship between characters, such as "R&D subsidy" and "technology development fund" being synonymous, with semantic similarity (≥0.85) calculated by BERT model and knowledge graph constructed. Logical dimension: refers to the cause-and-effect, conditional relationship in policy clauses, such as "annual R&D investment ≥ 5 million → 15% subsidy can be claimed", with "condition-result" logical pairs extracted through dependency syntax analysis. Time dimension: refers to the historical evolution trajectory of a character, such as "new energy subsidy ratio" 10% in 2023 → 8% in 2024 → 5% in 2025, with trend fitted by time series model (ARIMA). Correlation analysis results: refer to the comprehensive conclusions of three-dimensional mining, including semantic correlation path (such as "subsidy → qualification → tax"), logic rule base (such as "subsidy policy must contain reporting deadline"), and time evolution law (such as "subsidy ratio decreases by 2% annually"). Four-layer coding mechanism: refers to the unique character identification system layered by "field-function-feature-instance".
[0022] The first encoding refers to the character encoding generated by a four-layer mechanism, such as "New Energy Field + Financial Support + R&D + 2025 No. 1 Policy" encoding as "01F00100001", which associates semantics, logic, and time characteristics. In this embodiment, the service field refers to the sub-industry or business scenario covered by the policy, such as "Artificial Intelligence Computing Infrastructure", "Biopharmaceutical CDMO Platform", and "Cross-border E-commerce Pilot Zone". Historical releases refer to the collection of policy texts in the past 3-5 years in the service field, such as "2020-2025 Artificial Intelligence Industry Policy Compilation". The same type of character iteration order refers to the evolution trajectory of characters with policy updates, including new additions (such as "Generative AI Subsidy" first appeared in 2023), revisions (such as "Subsidy ratio from 10% to 15%"), and eliminations (such as "Fuel vehicle subsidies" abolished in 2025). Implement text difference algorithm (such as Myers algorithm) to compare historical policy versions. The second encoding refers to the identification of character iteration characteristics, with a format of [timestamp hash, iteration type, evolution weight], where: Timestamp hash: SHA-256 hash of the first 8 bits of the policy release time (such as "20250701" → "a3b7c9d1"); Iteration type: A (new) / M (modified) / D (eliminated); Evolution weight: 0-10 points (such as "Generative AI subsidy" assigned 10 points due to significant impact).
[0023] Multi-dimensional demand data refers to the full-dimensional demands of enterprises in policy adaptation, which is implemented by combining LDA topic model and BERT semantic extraction. Keyword vector refers to the conversion of multi-dimensional demand into a 768-dimensional numerical vector.
[0024] Mapping relationship refers to the association between demand vector and policy characteristics (distribution array, first / second encoding), which is constructed through "graph convolution network (GCN) + attention mechanism" model, highlighting high importance policies (such as official importance ≥ 8 points, character weight increased by 2 times). Customized service plan refers to the enterprise-specific policy guidelines generated based on the mapping relationship. In this embodiment, policy docking effect data refers to quantitative indicators after the implementation of the plan, such as "subsidy application success rate 85%", "qualification recognition cycle 28 days", and "policy fund arrival rate 100%", which are collected through the API of the enterprise reporting system and government public data. The demand change data refers to the demand update caused by the business adjustment of an enterprise, such as "expanding overseas market → adding 'cross-border e-commerce comprehensive test area policy' demand" and "merger and reorganization → adding 'antitrust review exemption' demand". The update character screening rule refers to the screening logic optimized based on the effect data, such as "the success rate of matching a certain character is less than 30% for a long time, and the weight of the character is reduced" and "a high-frequency demand character (such as 'green bond issuance') is added". Reinforcement learning (DQN) is used in implementation to improve the matching accuracy as a reward signal. Updating the first code and the second code refers to synchronizing policy iteration and demand change, such as updating the feature layer of the first code when the subsidy ratio is adjusted to 15%, and supplementing the iteration record of the second code when the generative AI subsidy is added, to ensure the timeliness of the code. The beneficial effects of the above technical solutions are: through the full-link design of multi-source collection-intelligent screening-three-dimensional correlation-iterative coding-precise matching-closed-loop optimization, the policy data is upgraded from fragmentation to structuring, the character analysis is upgraded from surface statistics to deep correlation, and the demand matching is upgraded from experience recommendation to intelligent mapping, so that the accuracy of policy service is improved, the success rate of enterprise declaration is improved, and the dynamic updating mechanism is used to continuously adapt to policy iteration and enterprise demand change, forming a virtuous cycle of policy-enterprise bidirectional empowerment, and ensuring the pertinence and timeliness of policy service.
[0025] The present application provides an industrial policy information service system based on intelligent analysis, the preprocessing module comprises: The homologous conversion unit is used for configuring a format analysis adapter to the corresponding specified platform according to the heterogeneous data format of each specified platform and combining the latest credibility weight of each specified platform, so as to realize the homologous conversion of heterogeneous data, wherein the credibility weight of the subscription information set by each specified platform is dynamically adjusted in combination with the conflict rate of the historical data, and when multi-source subscription information pushes the same policy, a weighted fusion algorithm is used to eliminate the conflict. The mapping unit is used for mapping synonyms to standard terminology based on the homologous conversion result and combining the industrial policy knowledge graph; The triple construction unit is used for identifying the condition-result logic pair related to the policy clause in the homologous conversion result through syntax analysis, and constructing a structured triple, wherein the structured triple includes a condition relationship and a result; The supplement unit is used for automatically supplementing the vague expression in the homologous conversion result according to the field classification standard to obtain a specific category; The data set acquisition unit is used for obtaining a standardized policy data set based on the mapping result, the structured triple and the specific category.
[0026] In this embodiment, the credibility weight refers to a quantitative indicator for measuring the authority of the subscription information of the specified platform. The initial weight is set based on the administrative level of the platform (0.8 for a national-level platform, 0.6 for a provincial-level platform, and 0.4 for a municipal-level platform), with a value range of 0-1. The higher the weight, the stronger the information credibility. For example, the initial credibility weight of a national-level policy document library is 0.8, and the initial weight of a certain municipal-level development zone platform is 0.4. The conflict rate of historical data refers to the proportion of differences in the description of the same policy on different platforms. The calculation formula is "conflict field number / total field number x 100%" (for example, if the "subsidy amount" is "500,000" on platform A and "1,000,000" on platform B, then the field is in conflict, and the conflict rate is 1 / 5=20%). In implementation, the weight is dynamically adjusted by monthly statistics of the conflict rate. The weight decreases by 0.1 for every 10% increase in the conflict rate. The format analysis adapter refers to a special data analysis module for different heterogeneous formats, which adapts to various data formats through modular design: For HTML format: use BeautifulSoup to extract the text within the tag (such as the policy provisions in ); For PDF format: use PyPDF2+OCR (Tesseract) to parse the scanned copy and extract the text content (accuracy ≥95%); For Excel format: use pandas to read and map to standard fields (such as mapping the "subsidy standard" column to "policy provisions"). Homologous conversion refers to the process of converting heterogeneous data into a unified structure (such as JSON format {"title":"...","publish time":"...","clause":"..."}), eliminating format differences. In implementation, it is achieved through a field mapping table (such as mapping "enactment date" and "release date" to "publish time"), with a conversion success rate of ≥98%. The industrial policy knowledge graph refers to a semantic association network constructed around policy terms, containing 1200+ core entities (such as "R&D subsidies" and "high-tech enterprises") and 3000+ semantic relationships (such as "contains", "synonymous", and "causal"). Entity attributes include "standard terminology", "common synonyms", and "domain classification". The graph is stored in Neo4j and supports semantic queries (such as querying the synonyms of "technology development funds"). For example, the synonym nodes of "R&D subsidies" in the graph include "technology development funds" and "research funding subsidies", with an association relationship of "synonymous" (similarity ≥0.85).
[0027] Synonyms refer to different expressions in policy texts but have the same meaning, such as "value-added tax, collect and return" and "value-added tax, collect first and return later", "qualification recognition" and "qualification audit". For example, "enterprise R&D expense subsidy" in one policy and "science and technology R&D fund support" in another policy are synonyms, which are mapped to the standard term "R&D subsidy" through knowledge graph. Standard terminology refers to unified vocabulary defined based on national policy norms (such as "National Administrative Organs Document Handling Method") and industry standards, such as "high-tech enterprise identification" and "enterprise income tax reduction and exemption", which are the basis for mapping synonyms. Syntactic analysis refers to parsing the grammatical structure of policy provisions through natural language processing technology (such as StanfordCoreNLP dependency syntactic analyzer), identifying subject-predicate-object, adverbial-complement, and other components, and extracting "condition-result" logical relationships. When implementing, focus on marking sentences containing "if… then…" "only… can…" and other related words, such as for the provision "enterprises with annual revenue ≥ 100 million yuan can enjoy a 15% tax reduction", syntactic analysis identifies "annual revenue ≥ 100 million yuan" as the condition and "enjoy a 15% tax reduction" as the result. The condition-result logical pair of policy provisions refers to the association between "preliminary conditions" and "corresponding benefits / constraints" in policies. Conditions usually contain quantitative indicators (such as "revenue ≥ 100 million" and "number of employees ≥ 50"), and results are usually policy benefits (such as "subsidy 2 million") or constraints (such as "not allowed to operate across regions"). Structured triples refer to logical units stored in the form of <condition, relationship, result>, where "relationship" includes "implication" (such as "condition implies result") and "equivalence" (such as "condition is equivalent to result"). Triples are written into knowledge graphs through SPARQL language, supporting logical reasoning (such as "satisfying condition A → triggering result B"), such as extracting the triple <annual revenue ≥ 100 million, implies, tax reduction 15%> from the provision, which clearly defines the logical association between conditions and results. Domain classification standard refers to the industry classification specifications published by the state or industry, such as "National Economic Industry Classification" (GB / T4754-2017) and "Strategic Emerging Industry Classification (2018)", which include clear definitions of subfields such as "high-end equipment manufacturing" and "new energy vehicles". Fuzzy expressions refer to abstract words in policy texts that do not clearly specify the specific scope, such as "key support areas", "strategic industries", and "high-tech products", which need to be supplemented with specific categories in combination with domain standards. For example, the fuzzy expression "key support for strategic emerging industries" in a policy needs to be supplemented with specific fields such as "new generation information technology", "high-end equipment manufacturing", and "new energy vehicles" in the "Strategic Emerging Industry Classification". The specific category refers to the accurate classification result corresponding to the fuzzy expression, which is determined through the matching of the field classification standard and the policy context. In implementation, a text classification model (such as a BERT fine-tuning model) is used, and the category with a similarity of ≥0.7 between the fuzzy expression and the standard classification label is taken as the supplementary result. The mapping result refers to the unified expression set after the synonym is mapped to the standard term, such as "technical development fund" and "scientific research fund subsidy" being unified as "R&D subsidy", to ensure the consistency of the term. The standardized policy data set refers to the final data set after the mapping result, the structured triple, and the specific category are fused, which has the characteristics of term unification, logical clarity, and category explicitness, and the data format is Parquet (supporting efficient query), including fields such as policy ID, standard term, condition, result, specific field, and release time. In implementation, data verification (such as logical consistency check: triple with empty condition needs to be marked as abnormal) is used to ensure the quality. The beneficial effects of the above technical solutions are: through the progressive processing of heterogeneous data adaptation conversion → term standardization mapping → logical structuring extraction → fuzzy expression precision, the pain points of multi-platform policy data format disorder, term inconsistency, and logical ambiguity are solved, the data consistency is improved, the subsequent character analysis efficiency is improved, and a high-quality data foundation is laid for accurate matching of enterprise needs.
[0028] The present application provides an industrial policy information service system based on intelligent analysis, and the character distribution analysis module comprises: A first candidate unit is used to calculate the semantic similarity between characters and core field labels, and determine the first candidate character in the standardized policy data set by combining the TF-IDF value; A second candidate unit is used to identify the characters of strong constraint modifiers in the policy chapter in the standardized policy data set based on text structure analysis, as the second candidate character; A third candidate unit is used to compare the similarity between the comparison characters in the standardized policy data set and the sample enterprise historical demand keyword library, to obtain the third candidate character; A character unit is used to obtain service characters based on the first candidate character, the second candidate character, and the third candidate character; A feature determination unit is used to perform global analysis on the standardized data set according to the service characters based on dynamic word cloud time evolution and theme clustering, to obtain the overall character distribution feature; A hierarchical clustering unit is used to perform hierarchical clustering on the service characters based on the policy publishing department and the industrial field to obtain the same type of characters.
[0029] In this embodiment, the core field label refers to a standardized classification label covering the main industries of the national economy, which is determined in accordance with the “Classification of Strategic Emerging Industries (2018)” and “Classification of National Economic Industries (GB / T 4754-2017)”, including 30+ labels such as “high-end equipment manufacturing”, “new energy vehicles”, “biological medicine”, etc., each of which is associated with 50+ field keywords (such as “power battery” and “charging infrastructure” associated with “new energy vehicles”). Semantic similarity refers to the cosine similarity (range 0-1) between the character vector and the core field label vector, which is calculated by the 768-dimensional vector output by the BERT model. When implemented, a threshold value of ≥0.7 is set to filter characters that are strongly related to the field (such as the similarity of 0.82 between “solid-state battery” and the “new energy vehicle” label, which is retained). TF-IDF value refers to the product of term frequency (TF) and inverse document frequency (IDF). Term frequency is the number of times a character appears in a single policy, and inverse document frequency is log(total number of documents / number of documents containing the character). It highlights characters that appear frequently in a specific policy but are not commonly seen globally (such as “first set” in “high-end equipment policy”, which has a high frequency but is rare in other policies, with a high TF-IDF value). The first candidate character refers to a character that meets both “semantic similarity ≥0.7” and “TF-IDF value in the top 30%”. For example, “power battery energy density” is listed as the first candidate character because it has a similarity of 0.85 with the “new energy vehicle” label and a TF-IDF value in the top 20%. Text structure analysis refers to the structured identification of policy text chapters and clause levels. Through regular expression matching of title identifiers such as “Chapter 1 General Provisions” and “1.1 Declaration Conditions”, combined with layout features such as paragraph spacing and font size, the policy chapters “General Provisions”, “Declaration Conditions”, “Subsidy Standards”, and “Penalties” are divided, and the accuracy of the analysis is ≥90%. Policy chapter: refers to a structural unit in a policy text with a clear function, such as “declaration conditions” (specifying enterprise qualifications), “subsidy standards” (clearing subsidy amounts), “implementation period” (limiting time range), etc., which is the carrier of the core content of the policy. Strong constraint modifier: refers to adverbs or conjunctions that express compulsion or obligation, such as “must”, “should”, “may not”, “strictly prohibited”, etc. These words usually modify characters related to policy red lines or necessary conditions (such as “30 days” in “must submit materials within 30 days”). The second candidate character refers to a character directly modified by a strong constraint modifier in a policy chapter. For example, in the “declaration conditions” chapter, “1 item or more core invention patents” in “should have 1 item or more core invention patents” is modified by the strong constraint word “should”, and is listed as the second candidate character. Sample enterprise historical demand keyword library refers to the collection of 1000+ enterprise (covering large, medium and small micro enterprises and various industries) historical policy demand keyword set, including "financing difficult", "R&D investment large", "qualification certification complex" and other demands, each demand is associated with 3-5 keywords (such as "R&D investment large" is associated with "R&D subsidy" and "additional deduction").
[0030] The contrast character refers to the character to be screened in the standardized policy data set (such as "number of intellectual property rights" and "environmental protection approval process"). The similarity comparison refers to the semantic similarity (calculated by Sentence-BERT) between the contrast character and the demand keywords in the keyword library, and the threshold is set to be greater than or equal to 0.6, and the characters strongly related to the enterprise demand are screened (such as the similarity between "number of intellectual property rights" and "qualification certification complex" demand is 0.68, which is retained). The third candidate character refers to the character with a similarity greater than or equal to 0.6 with the sample enterprise historical demand keyword library, such as "software copyright registration" with a similarity of 0.72 with the "high-tech enterprise declaration" demand keyword, which is listed as the third candidate character. The service character refers to the final character determined by the "voting mechanism" by combining the first, second and third candidate characters: if the character appears in two or more candidate sets, it will be selected; if the character only appears in one candidate set, it needs to be manually reviewed (if two of the three field experts agree, it will be retained), such as "R&D expense ratio ≥ 5%", which appears in both the first candidate (related to the "biomedicine" label) and the second candidate (modified by "must"), which is determined as the service character. In this embodiment, the dynamic word cloud time sequence evolution refers to the word cloud sequence generated with time as the axis (monthly / quarterly), and the size of the word cloud changes dynamically with the frequency of its appearance in the policy in that period, which intuitively shows the trend of the character's popularity (such as "carbon footprint" in the word cloud size from "small" to "large" in 2023Q1-Q4, reflecting the rising attention). When implementing, use the WordCloud library of Python to generate, and update once a quarter.
[0031] Theme clustering refers to clustering service characters using unsupervised learning algorithms (such as LDA) to mine potential themes (such as "funding support", "qualification recognition", "talent introduction"). When implementing, the optimal number of themes (usually 8-12) is determined by Perplexity, and the clustering results are visualized by t-SNE dimension reduction, and a contour coefficient greater than or equal to 0.7 indicates effective clustering. Overall character distribution characteristics refer to the global regularity of service characters, including high-frequency themes (such as "funding support characters account for 35%"), cross-domain common characters (such as "tax reduction and exemption" appears in multiple field policies), and time-sensitive hotspots (such as "metaverse" first became a high-frequency character in 2024 policies).
[0032] Policy publishing department refers to the administrative subject publishing the policy, which is divided into national, provincial and municipal levels according to the hierarchy, and the policy coverage and effectiveness of different levels of departments are different (national policy is applicable nationwide, and municipal policy is limited to the local area). Industry field refers to the subdivided industry focused by the policy, such as "integrated circuit", "industrial internet" and "green building", which corresponds to the core field label and is the key dimension of character clustering. Hierarchical clustering refers to a two-step clustering strategy of clustering according to the industry field (such as "integrated circuit") first and then clustering according to the policy publishing department (national / provincial) in each field second, using hierarchical clustering algorithm (Ward method) and Euclidean distance measurement of character frequency to ensure that the characters in the same cluster are consistent in the field and the level. The same type of character refers to the character set formed after hierarchical clustering, and the characters in the same cluster have the same industry field and similar policy publishing level, for example, "14nm chip R&D subsidy" and "EDA tool procurement subsidy" are clustered as "national integrated circuit R&D subsidy" because they belong to the "integrated circuit" field and are published by the national department. The beneficial effects of the above technical solution are: through the process of three candidate character fusion screening + hierarchical clustering, the strong relevance of service characters and industry fields, the necessity of policy constraints and the matching of enterprise needs are ensured, and the precision of service characters is improved through hierarchical clustering to realize fine classification of characters, which lays a high-quality data foundation for subsequent distribution array construction and coding analysis. The present application provides an industrial policy information service system based on intelligent analysis, and the character distribution analysis module further comprises: The proportion determination unit is used for calculating the frequency proportion of the same type of character in the policy cluster, eliminating short-term fluctuations by exponential smoothing method, and calculating the distribution proportion according to the policy publishing area; The importance determination unit is used for constructing the evaluation matrix of publishing subject level, clause constraint strength and implementation range, and obtaining the official importance of the same type of character; Wherein, the distribution array is obtained based on the distribution proportion and the official importance.
[0033] In this embodiment, the window refers to a statistical period divided by time, usually using a quarterly sliding window (window length = 3 months, step = 1 month), which is used to capture the short-term trend of character frequency.
[0034] Policy cluster refers to the grouping of policy texts by theme (e.g., "financial support" "qualification recognition"), achieved through the K-Means algorithm (K value = 8, based on silhouette coefficient determination). Policies in the same cluster have similar core content, such as "new energy vehicle purchase subsidy policy" "photovoltaic industry financial support policy" are clustered as "financial support type policy" because they both involve financial support. Frequency ratio refers to the ratio of the frequency of a certain character in the policy cluster it belongs to. For example, "photovoltaic subsidy" appears 20 times in "financial support type policy", and the total number of characters in the cluster is 80, so the frequency ratio = 20 ÷ 80 = 25%. Exponential smoothing is a time series processing method used to eliminate short-term random fluctuations. It calculates the smoothed value by giving higher weight to recent data (smoothing coefficient α = 0.6), formula: S_t = α × x_t + (1 - α) × S_{t-1} (S_t is the smoothed value, x_t is the current window value, S_{t-1} is the value before smoothing). When implemented using pandas.ewm, the distribution ratio curve is smoother (fluctuation amplitude is reduced by 40%).
[0035] Policy release area is the policy release range divided by administrative region, such as eastern (Beijing-Tianjin-Hebei, Yangtze River Delta), central (Hunan-Hubei-Jiangxi), western (Chengdu-Chongqing), northeast (Heilongjiang-Jilin-Liaoning). When implemented, the "implementation scope" field in the policy (such as "This policy applies to Guangdong Province") is mapped to the regional classification. Distribution ratio is the final ratio after regional division and exponential smoothing, forming a "region-time-ratio" three-dimensional data, such as "eastern region Q1 new energy subsidy character ratio 25%". Release subject level is the administrative level of the policy release department, divided into three levels according to authority: national level (weight 0.45), provincial level (weight 0.35), municipal level (weight 0.2). When implemented, the department name is matched automatically, and the highest level is taken in special cases (such as joint release by multiple departments). Clause constraint strength is the degree of compulsion of policy clauses, divided into "mandatory type" (such as "must" "should" "shall not", assigned 8 points), "guidance type" (such as "encourage" "support", assigned 4 points), "explanation type" (such as "defined as" "includes", assigned 2 points). When implemented, regular expressions are used to match modifiers (such as r"must|should|shall not"), and the modified object is located by combining dependency syntax analysis, with an accuracy of ≥90%. Implementation scope is the spatial or industry breadth covered by the policy, divided into national (10 points), provincial (6 points), municipal (3 points), and specific park (1 point). When implemented, the "applicable scope" field (such as "high-tech enterprises nationwide") is matched, and the highest value is taken for cross-regional policies (such as "East China" is calculated as 6 points at the provincial level). The evaluation matrix is a [character number x 3] matrix taking the publishing subject level, the clause constraint strength and the implementation range as row indexes, and the same type characters as columns, and the weight is calculated by improving the analytic hierarchy process (AHP-entropy weight method): The initial weight is determined by expert scoring (avoiding subjective bias, and 3 or more policy experts participate); The entropy weight method is modified (based on index volatility, and the weight of the index with large coefficient of variation is increased); The maximum weight is normalized (the sum is 1), and the matrix element is the weighted sum of "index score x weight".
[0036] The distribution array is a two-dimensional array of structured storage "distribution proportion" and "official importance", and the format is [character ID, regional distribution proportion (eastern / middle / western), time series proportion (last 4 quarters), official importance score]. For example, the distribution array of the character "photovoltaic subsidy 001" is ["PV001", [25%, 15%, 10%], [20%, 22%, 25%, 23%], 14.5 points]. The beneficial effects of the above technical scheme are: through the time window + regional distribution proportion calculation (combined with exponential smoothing denoising), the character distribution characteristics are more stable, the three-dimensional evaluation matrix + improved AHP official importance quantification, the objectivity (entropy correction) and policy attributes (level / strength / range) are considered, and the subsequent matching precision is improved.
[0037] The application provides an industrial policy information service system based on intelligent analysis, and the first coding module comprises: A semantic analysis unit is configured to define policy entities, attributes and associated relationships, construct an initial graph through human-computer collaborative labeling, perform cross-modal fusion of embedded representations of characters of the same type and triples of the initial graph, and calculate semantic correlation between characters. A logical analysis unit is configured to split a policy into atomic clauses based on a text segmentation algorithm, extract a condition-action-result structure of the atomic clauses using a dependency syntax tree to generate logical tuples, and perform compliance verification on the logical tuples. A time analysis unit is configured to perform time series analysis on historical release time, revision time and abolition time of characters of the same type, and determine an evolution track table. A mechanism coding unit is configured to code the characters of the same type based on the associated analysis results and using a four-layer coding mechanism based on a field dimension, a function dimension, a feature dimension and an instance dimension. The associated analysis results include semantic correlation between characters based on a semantic dimension, compliance verification results of logical tuples based on a logical dimension, and an evolution track table based on a time dimension.
[0038] In this embodiment, policy entities refer to core concepts with independent meanings in industrial policy, covering policy elements (such as "R&D subsidies" and "high-tech enterprise identification"), constraints (such as "environmental protection standards" and "revenue thresholds"), and implementing subjects (such as "provincial science and technology departments" and "industrial park management committees"). Attributes are parameters that describe the characteristics of policy entities, such as the attributes of "R&D subsidies" including "subsidy ratio (15%)" "reporting period (every March)" and "beneficiary enterprise type (small and medium-sized enterprises)". Association relationships are semantic connections between policy entities, mainly including: Synonyms (such as "R&D subsidies" and "technology development funds"); Inclusion (such as "strategic emerging industry subsidies" containing "artificial intelligence subsidies"); Causal relationship (such as "through qualification" → "obtain subsidies"). Relationships are identified through BERT model semantic similarity calculation (≥0.85 is determined as synonymous) and rule engine (such as "if the clause contains 'because…' then it is a causal relationship"). Human-computer collaborative annotation is the annotation method for constructing the initial graph, the process is: Machine pre-annotation: automatically identify policy entities, attributes, and relationships using entity recognition models (such as BERT-NER, F1 value ≥0.9); Manual verification and correction: domain experts (policy researchers, enterprise lawyers) sample review (20% sampling rate) of machine annotation results, correct wrong associations (such as changing the misjudged "inclusion relationship" to "parallel relationship"); Iterative optimization: every 1000 data annotation, use the correction results to fine-tune the model to improve the accuracy of subsequent annotation. Initial graph: the basic semantic network generated by human-computer collaborative annotation, stored in Neo4j graph database, nodes are policy entities, edges are association relationships and annotated with weights (such as synonym relationship weight 0.9, causal relationship weight 0.8). The embedding representation of characters of the same type refers to the conversion of characters of the same type (such as "R&D subsidies" and "equipment subsidies" belonging to the same "subsidy category") into low-dimensional dense vectors, generated by the industrial policy BERT model (768-dimensional vector). The smaller the vector distance, the closer the semantic distance between characters. Specifically: load the pre-trained BERT model using PyTorch, input the character text, and take the [CLS] vector as the embedding representation. Characters with a cosine similarity ≥0.75 are considered semantically close. Triple refers to the basic structure of "entity-relation-entity" in the knowledge graph, such as <R&D subsidies, contains, subsidy ratio>, <enterprise, meets, reporting conditions>, which is the basic unit of semantic association. Cross-modal fusion refers to the fusion of character embedding representation (text modality) and graph triple (structured modality) into unified features, and the enhancement of semantic representation through attention mechanism (such as using triple relationship weight as attention weight), for example, the embedding vector of "R&D subsidy" and the triple <R&D subsidy, contains, subsidy ratio> are fused, and the "contains" relationship weight 0.8 makes the fused vector more prominent "subsidy ratio". Semantic correlation degree refers to the quantitative indicator of semantic similarity between characters, represented by the cosine similarity of the fused vector, with a value range of 0-1, ≥0.8 indicating strong correlation. Specifically, the cosine similarity is calculated using scikit-learn, and the correlation between "R&D subsidy" and "technology development fund" is 0.88. Text segmentation algorithm refers to the algorithm for splitting policy text into independent clauses based on punctuation marks (such as ". " and "; ") and layout features (such as paragraph spacing), combined with regular expression matching of clause identifiers such as "first" and "(one)". Atomic clause refers to the smallest independent clause in a policy that contains complete logical relationships, such as "enterprises with annual R&D investment ratio ≥5% can enjoy 15% tax reduction" and "reporting materials must be submitted by March each year". Dependency syntax tree refers to a tree diagram representing sentence structure using syntactic relationships (such as subject-predicate, verb-object, and adjective-noun), with the root node being the core verb and the leaf nodes being nouns, adjectives, etc., used to extract the logical structure of the sentence. Condition-action-result structure refers to the logical chain implied in the atomic clause: Condition: prerequisite constraint (such as "annual revenue ≥100 million yuan"); Action: core behavior (such as "apply for subsidy"); Result: consequence caused by the action (such as "obtain 2 million yuan subsidy"). Logical tuple refers to the conversion of condition-action-result structure into structured data, with the format (condition, action, result), for example, the atomic clause "enterprises with annual revenue ≥100 million yuan can apply for 2 million yuan subsidy" is extracted as the logical tuple ("annual revenue ≥100 million yuan", "apply", "2 million yuan subsidy"). Compliance verification refers to checking whether the logical tuple meets the general rules of the policy, and the rule library includes: Integrity rule: the tuple must contain condition, action, and result (such as missing "condition" will result in verification failure); Reasonableness rule: the result must match the condition (such as "annual revenue 100 million yuan" corresponding to "subsidy 20 million yuan" is unreasonable); Consistency rule: the same type of tuple in the same policy must be logically consistent (such as the condition of "R&D subsidy" cannot be both "≥500 million" and "≥1000 million"). Implementation method: Use Prolog logical programming language to build a rule base, and perform pattern matching verification on tuples. Tuples with a rate of ≥85% are considered to be in compliance. Historical release time refers to the timestamp of the first publication of the policy (accurate to the day), such as "2023-05-10", which is the basis for the character time attribute. Revision time refers to the timestamp of subsequent modifications to the policy, such as "2024-03-15". Each revision may result in changes to the character content (e.g., "subsidy ratio changed from 10% to 12%"). Abolition time refers to the timestamp of the policy's expiration, such as "2025-12-31". The character becomes invalid upon the policy's abolition (e.g., "old equipment subsidy" is abolished after 2025). Time series analysis refers to statistical modeling of the frequency and content changes of characters at different time points. The ARIMA model (Autoregressive Integrated Moving Average Model) is used to fit the trend and predict the evolution probability in the next 6 months (e.g., the probability of "subsidy ratio" decreasing). Evolution trajectory table refers to a structured table that records the time evolution of characters of the same type, containing fields: character ID, historical release time, revision time, abolition time, content change description (e.g., "2024-03 revision: subsidy ratio from 10% to 12%"), and predicted evolution trend. For example, in the evolution trajectory table of "new energy vehicle purchase subsidy", the record for 2023 is "subsidy ratio 15%", the record for 2024 is "12%", and the predicted record for 2025 is "10%". In this example, the correlation analysis result refers to the comprehensive analysis conclusion of the semantic correlation degree (semantic dimension), logical tuple compliance (logical dimension), and evolution trajectory (time dimension), such as "R&D subsidy" and "high-tech enterprise" with a semantic correlation degree of 0.85, logical tuple compliance, and a 2% decrease in subsidy ratio per year from 2023 to 2025. Four-layer coding mechanism: refers to the unique coding system of characters based on the "field-function-feature-instance" hierarchy, with the following coding rules for each layer: Field dimension: based on the "National Economic Industry Classification" coding, such as "C35-high-end equipment manufacturing" and "D44-power supply"; Function dimension: classified by policy function, such as "B-subsidy", "Z-qualification recognition", and "J-supervision"; Feature dimension: marks semantic / logical features, such as "YF-R&D" and "HJ-environmental protection" (based on semantic correlation), "QT-conditional" and "JG-result" (based on logical tuples); Instance dimension: unique identification of characters, such as 6-digit numbers (000001-999999), associated with the timestamp of the evolution trajectory table. The first code refers to a unique character identifier generated by a four-layer coding mechanism, and the format is "field dimension-function dimension-feature dimension-instance dimension", for example, the first code of "high-end equipment manufacturing field research and development subsidy (conditional, revised in 2024)" is "C35-B-YF-QT-001234", wherein "001234" is associated with the 2024 revision record in the evolution track table. The beneficial effects of the above technical solutions are: through the progressive processing of semantic association graph construction, logical tuple extraction and verification, time evolution analysis and four-layer code solidification, the conversion of characters from text description to structured knowledge is realized, which provides precise feature support for subsequent enterprise demand matching, which is semantically associated, logically traceable and time-evolving.
[0039] The present application provides an industrial policy information service system based on intelligent analysis, and the mapping analysis module comprises: A demand acquisition unit is configured to acquire explicit demand, implicit demand and potential demand of a target enterprise, wherein the explicit demand is obtained from an enterprise end interactive platform based on a structured form and a free text enhanced mode, the implicit demand is obtained by mining enterprise annual reports, social responsibility reports and interview recordings, and the potential demand is predicted based on the mining results of the time dimension of the first coding module; A model analysis unit is configured to input the explicit demand, implicit demand and potential demand into an industrial policy field enhanced large model, and output a keyword vector, wherein the industrial policy field enhanced large model is based on a general BERT model, and is trained by a language model with a life cycle label mask and a keyword vector, and by integrating policy-demand and training corpus; An extraction and determination unit is configured to extract statistical patterns of the distribution array based on a multi-layer perception, determine inter-code semantic connections of the first code based on a graph convolution network, and capture policy evolution rules of the second code based on a time convolution network; A relationship determination unit is configured to respectively project the keyword vector, statistical pattern, inter-code semantic connection and policy evolution rule into a three-dimensional coordinate system, determine a matching relationship based on a matching policy, a matching relationship based on a policy condition dependency chain, and a matching relationship based on a risk point of a historical eliminated character and a regional difference of the distribution array based on the second code; A scheme development unit is configured to obtain a strategy for coping with each matching relationship, and generate a customized service scheme.
[0040] In this embodiment, the explicit demand is the policy appeal expressed by the enterprise. Structured form is a standardized form with preset fields such as "demand type", "field of application", and "urgency level" (drop-down selection of "financial support" and "qualification recognition"), collected through an enterprise-side interactive platform (such as a web portal or app), with data format unified as JSON ({"demand type": "subsidy application", "field": "high-end equipment"}). Free-text enhancement is an additional text input box in the form, allowing enterprises to supplement personalized descriptions (such as "hope to learn about the list of materials for subsidy application"), with the supplemented information automatically classified by a TextCNN model (accuracy ≥ 90%) and associated with explicit demands. Enterprise-side interactive platform is a digital interface for enterprises to submit demands, supporting both PC and mobile devices, with integrated identity authentication (such as enterprise CA certificate login) to ensure information authenticity, with an average of ≥ 1000 demand submissions processed per day. Implicit demand is a policy appeal that is not directly expressed by the enterprise but is implied, which needs to be extracted through text mining: Annual report of the enterprise: extract potential demand keywords such as "R&D investment ratio of 15%" and "overseas market expansion" through an LDA topic model (number of topics K=8); Social responsibility report: use sentiment analysis (VADER model) to identify related demands such as "carbon neutralization target" and "green production"; Interview recording: after ASR transcription (accuracy ≥ 95%), use BERT to extract "supply chain financing difficulties" and other colloquial demands. Potential demand: future demand based on policy time evolution trend prediction, such as "carbon footprint certification demand" derived from "2026 carbon tariff policy implementation". Mining basis: time dimension analysis results of the first coding module (such as policy iteration cycle and character elimination rules), with an LSTM model (time step = 12) to predict policy hotspots in the next 1-2 years, with a prediction error ≤ 15%. For example: based on the historical evolution trajectory of "new energy subsidy reduction" (20% reduction per year from 2023 to 2025), it is predicted that "new energy enterprises need to shift to technology cost reduction" in 2026, leading to the potential demand "support for intelligent manufacturing policies". Industry policy field enhanced large model: a natural language processing model based on general BERT, optimized for the field, with policy-demand semantic understanding capabilities. Training corpus: fusion of 50,000+ policy texts and 30,000+ enterprise demand cases (explicit / implicit / potential demand), building a "policy-demand" parallel corpus. Life cycle label of scene mask: Mask the life cycle stage of enterprises (start-up / growth / maturity / decline) (such as [Stage=Growth]), and train the model to identify the differences in demand at different stages (such as focusing on "venue subsidies" in the start-up stage and "financing support" in the growth stage). Training process: Pre-training with Masked Language Model (MLM) (10 million iterations), fine-tuning with enterprise demand data (learning rate 5e-5), and making the model's accuracy on the policy-demand matching task ≥92%. Keyword vector: Convert enterprise demand into a 768-dimensional dense vector, and vector cosine similarity ≥0.8 indicates similar demand. Generation method: Input demand text (such as "new energy vehicle enterprises + annual R&D investment 100 million + carbon footprint management"), and output vector from the last layer of the model, with vector values reflecting the semantic correlation strength between demand and policy terms. For example, the keyword vector of "R&D subsidy application" has a vector cosine similarity of 0.85 with "technology development fund support", indicating a high degree of correlation. Multi-layer Perceptron (MLP) is used to extract statistical patterns of distribution arrays (such as "new energy subsidy characters account for 25% + official importance 18 points"). The statistical pattern of the distribution array refers to the quantification of the distribution proportion and official importance, such as "high importance characters (≥15 points) account for more than 30% in the eastern region" and "characters with quarterly growth ≥5% are concentrated in the 'digital economy' field". Graph Convolutional Network (GCN) is used to learn the semantic connection between the first encoding, with encoding as nodes and semantic correlation as edges to construct an adjacency matrix (correlation ≥0.7 as an edge), and through 2-layer GCN (hidden layer dimension 128) to extract the topological relationship features between encodings (such as the association path of "R&D subsidies" → "high-tech enterprises" → "tax exemptions"). The inter-encoding semantic connection of the first encoding refers to the association between different encodings in terms of semantics / logic (such as "01F00100001-new energy subsidies" and "01Z00200003-high-tech enterprise identification" connected due to policy clause association). Temporal Convolutional Network (TCN) is used to capture the policy evolution rules of the second encoding (such as the time series characteristics of "subsidy proportion gradually reduced over the years"), with a convolution kernel size of 3 and an expansion coefficient of 2, to extract multi-scale time features (short-term fluctuations / long-term trends), and output a time feature vector (64 dimensions). The policy evolution rule of the second encoding refers to the iterative pattern of characters over time, such as "new characters are concentrated in Q1 release" and "the average survival period of eliminated characters is 3 years". The three-dimensional coordinate system is a space constructed with "statistical feature dimension", "semantic correlation dimension" and "time evolution dimension" as axes, which is used to visualize the matching degree (the closer the distance, the higher the matching degree) of demand vector and policy features. The matching relationship based on matching policies refers to the similarity matching of demand keyword vector and policy feature vector. The Euclidean distance in three-dimensional space (threshold ≤ 0.3) is calculated to determine the matching degree. For example, the distance between "R&D subsidy demand" and "Provincial R&D expense addition deduction policy" is 0.25, which is determined as high matching. The matching relationship based on policy condition dependency chain refers to the "premise-result" dependency matching between policy clauses (such as "complete environmental impact assessment → obtain pollution discharge permit → apply for environmental protection subsidy"). By analyzing the logic tuples of the first encoding, the condition dependency chain is generated, and it is checked whether the enterprise meets the precondition (such as "environmental impact assessment has been completed"). For example, an enterprise needs "environmental protection subsidy", and the system matches the dependency chain "environmental impact assessment passed → subsidy application". If the enterprise has completed the environmental impact assessment, the matching is successful. The matching relationship based on risk points refers to the identification of potential risks in policy adaptation, including: Historical elimination character risk (such as matching to "fuel car subsidy", but the second encoding shows that the character has been eliminated in 2025); Regional difference risk (such as an enterprise in the western region, matching to "eastern region special subsidy", and the distribution array shows that the western region accounts for <5%). Risk points are identified through knowledge graph reasoning (rule base contains 200+ risk patterns), and risk levels are divided into "high / medium / low" (high risk needs to be prompted first). The strategy corresponding to each matching relationship is a response plan for different matching results: High matching policy provides "list of application materials + time node table" (such as "prepare patent certificate in March, submit system in April"); Condition dependency chain matching is supplemented by "precondition handling guide" (such as "environmental impact assessment handling process and required materials"); Risk point matching gives "alternative policy recommendation + risk avoidance suggestion" (such as "fuel car subsidy has been eliminated, recommend new energy vehicle replacement subsidy"). Customized service plan is an enterprise-specific report integrating the above strategies, including: Policy matching list (sorted by "matching degree x official importance", Top3 policies with details link); Implementation path diagram (Gantt chart form, with key nodes and responsible persons marked); Risk warning and response (such as "Policy expires in December 2025, it is recommended to report one month in advance"), for example, the scheme of a new energy enterprise includes "New Energy Vehicle Purchase Subsidy Rules" (matching degree 92%), reporting steps (including patent certificate template), risk warning ("Need to complete the reporting in Q3, overdue invalid").
[0041] The beneficial effects of the above technical solutions are: through the technical chain of full-dimensional demand capture-field enhancement model transformation-multi-modal feature fusion-three-dimensional matching reasoning, the precise adaptation of enterprise demand and policy is realized, and the intelligence and effectiveness of policy service are significantly improved. The present application provides an industrial policy information service system based on intelligent analysis, the second encoding module comprises: The difference analysis unit is used for comparing the version difference of the crawled industrial policy information in the service field by a text difference algorithm, marking the addition, modification and deletion operation of the same type of characters, and generating a character iteration track table, wherein the character iteration track table contains operation time and influence range of each operation. The encoding generation unit is used for generating the second encoding of the same type of characters based on the character iteration track table.
[0042] In this embodiment, the text difference algorithm refers to an algorithm for identifying the difference between two text versions, which adopts Myers difference algorithm (time complexity O(N+M), N and M are the number of characters in two versions), and locates the difference position by calculating the shortest editing distance (the minimum number of insertion, deletion and replacement operations). When implemented, the difflib library of Python is used to output the difference type (addition / modification / deletion) and specific characters, and the accuracy is greater than or equal to 95%, for example, comparing "New Energy Vehicle Subsidy Policy (2024 version)" with "New Energy Vehicle Subsidy Policy (2025 version)", the algorithm identifies the modification difference of "subsidy ratio from 15% to 12%". The service field refers to the subdivided industrial field focused on by the industrial policy, such as "photovoltaic industry", "industrial robot" and "biological medicine", which corresponds to the core field label, and is the range limitation of version comparison (only comparing the policy versions of the same field). The version difference of the crawled industrial policy information refers to the content difference of the policy text published at different times in the same service field, including the increase and decrease of clauses, the adjustment of numerical values, the modification of expressions, etc., for example, the "photovoltaic industry" policy 2023 version contains "single crystal component subsidy", the 2024 version adds "perovskite component subsidy", and the 2025 version deletes "polycrystalline component subsidy", which is the version difference. The same type of characters refers to characters belonging to the same policy function category, such as "photovoltaic power generation subsidy" and "wind power installation subsidy" belonging to "subsidy category", or "high-tech enterprise identification" and "specialized and new product certification" belonging to "qualification category". New operation refers to the same type of character that appears for the first time in the new version of the policy, marked with "+", for example, "green hydrogen production subsidy" appears for the first time in the 2025 version of "hydrogen energy industry policy", marked as a new operation.
[0043] Modification operation refers to the change of the same type of character in the new version (such as numerical adjustment, condition change), marked with "Δ", for example, "R&D subsidy ratio" is adjusted from "10%" in 2024 to "12%" in 2025, marked as a modification operation. Delete operation: refers to the same type of character that no longer appears in the new version, marked with "-", for example, "coal-fired unit modification subsidy" is deleted in the 2025 version of "traditional manufacturing industry upgrading policy", marked as a deletion operation. Character iteration track table is a structured table 1 that records the version changes of the same type of character, for example: Table 1 Structured table
[0044] Second encoding refers to the unique identification of character time evolution characteristics generated based on the character iteration track table, the encoding rule is [timestamp hash]-[iteration type]-[evolution weight]-[impact range code], the specific composition is: Timestamp hash: SHA-256 hash of operation time (e.g. "20250310"→"a7b3c9d2"), ensure time uniqueness; Iteration type: use letters to represent operation type (A=New, M=Modify, D=Delete); Evolution weight: a quantitative index of 0-10 points, calculated based on "character impact range (nationwide x 1.2 / region x 0.8) + policy level (national x 1.5 / provincial x 1.0)" (e.g. national new character gets 10 points); Impact range code: 2-letter code (e.g. "GN=nationwide, BJ=Beijing, JS=Jiangsu"), for example, the second encoding of the modification of the national photovoltaic subsidy ratio on March 10, 2025 (evolution weight 8 points) is "a7b3c9d2-M-8-GN". Use the hashlib library of Python to generate timestamp hash; Iteration type is mapped as "+→A", "Δ→M", "-→D"; Evolution weight is calculated by a preset formula (e.g. weight=impact range coefficient x policy level coefficient x base score 5), impact range code is mapped based on administrative division code table (e.g. "310000→SH" represents Shanghai).
[0045] The beneficial effects of the above technical solution are: through the text difference algorithm, the iterative trajectory of the policy character is accurately captured, the second code containing time, type and weight is generated combined with the structured trajectory table, the quantitative record of the whole life cycle evolution of the policy character is realized, the key features in the time dimension (such as preferentially matching the newly added / high weight character) are provided for subsequent enterprise demand matching, the timeliness of policy service is improved.
[0046] The application provides an industrial policy information service method based on intelligent analysis, as shown in the figure, comprising: Figure 2 Step 1: obtaining industrial policy information from a specified platform through a crawler technology, and obtaining multi-source policy data of different types based on subscription information related to the specified platform, and pre-processing the multi-source policy data to obtain a standardized policy data set; Step 2: performing character screening on the standardized policy data set based on a character screening rule to obtain service characters, and determining overall character distribution and same-type character distribution to obtain a distribution array of same-type characters, wherein the distribution array comprises: distribution proportion and official importance of same-type characters; Step 3: mining the same-type characters from three dimensions of semantics-logic-time to obtain correlation analysis results, and performing first coding on the same-type characters by using a four-layer coding mechanism; Step 4: crawling service fields of the industrial policy information, and extracting an iteration order of same-type characters in the historical publication of the service fields to generate second coding; Step 5: collecting multi-dimensional demand data of a target enterprise, and converting the multi-dimensional demand data into a keyword vector, establishing a mapping relationship based on the distribution array, the first coding and the second coding, and generating a customized service scheme; Step 6: obtaining policy docking effect data and demand change data of the target enterprise based on the customized service scheme, and updating the character screening rule and the first coding and the second coding. The beneficial effects of the above technical solution are: through the whole-link design of multi-source collection-intelligent screening-three-dimensional correlation-iterative coding-accurate matching-closed loop optimization, the upgrading of policy data from fragmentation to structuring, character analysis from surface statistics to deep correlation, and demand matching from experience recommendation to intelligent mapping is realized, the accuracy of policy service is improved, the success rate of enterprise declaration is improved, and through the dynamic updating mechanism, the policy iteration and enterprise demand change are continuously adapted, forming a virtuous cycle of policy-enterprise bidirectional empowerment, and ensuring the pertinence and timeliness of policy service.
[0047]
[0048] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. An industrial policy information service system based on intelligent analysis, characterized in that: include: A preprocessing module is used to obtain industrial policy information from a designated platform using crawler technology, obtain different types of multi-source policy data based on subscription information related to the designated platform, and preprocess the multi-source policy data to obtain a standardized policy data set; a character distribution analysis module, configured to perform character screening on the standardized policy dataset based on character screening rules to obtain service characters, and determine the overall character distribution and the distribution of characters of the same type to obtain a distribution array of characters of the same type, wherein the distribution array includes: a distribution percentage and an official importance of characters of the same type; A first encoding module is used to mine the same type of characters from the three dimensions of semantics, logic and time, obtain association analysis results, and perform a first encoding on the same type of characters using a four-layer encoding mechanism; A second encoding module is used to crawl the service field of the industrial policy information, extract the iterative order of the same type of characters under the historical release of the service field, and generate a second code; a mapping analysis module for collecting multidimensional demand data of a target enterprise, converting the multidimensional demand data into a keyword vector, establishing a mapping relationship based on the distribution array, the first code, and the second code, and generating a customized service solution; The updating module is used to obtain the policy docking effect data and demand change data of the target enterprise based on the customized service solution, and to update the character screening rules and the first code and the second code.
2. The industrial policy information service system based on intelligent analysis according to claim 1 is characterized in that: The preprocessing module includes: A homologous conversion unit is used to configure a format parsing adapter for each designated platform based on the heterogeneous data format of each designated platform and the latest credibility weight of each designated platform, thereby achieving homologous conversion of heterogeneous data. The credibility weight of the subscription information set by each designated platform is dynamically adjusted based on the conflict rate of historical data. When multiple sources of subscription information push the same policy, a weighted fusion algorithm is used to eliminate conflicts. A mapping unit, which is used to map synonyms to standard terms based on the homology conversion results and in combination with the industrial policy knowledge graph; A triple construction unit is used to identify condition-result logical pairs related to policy clauses in homologation conversion results through syntactic analysis, and construct structured triples, wherein the structured triples include conditional relations and results; A supplementing unit, configured to automatically supplement the fuzzy expressions in the homology conversion results to obtain specific categories according to a field classification standard; The dataset acquisition unit is used to obtain a standardized policy dataset based on the mapping results, structured triples and specific categories.
3. The industrial policy information service system based on intelligent analysis according to claim 1 is characterized in that: The character distribution analysis module includes: A first candidate unit, configured to calculate the semantic similarity between a character and a core domain label, and determine a first candidate character in the standardized policy dataset in combination with a TF-IDF value; A second candidate unit is used to identify characters of strong constraint modifiers appearing in policy sections in the standardized policy dataset based on text structure analysis as second candidate characters; A third candidate unit is used to compare the comparison characters in the standardized policy data set with the sample enterprise historical demand keyword library for similarity to obtain a third candidate character; a character unit, configured to obtain a service character based on the first candidate character, the second candidate character, and the third candidate character; A feature determination unit, configured to perform a global analysis of the standardized data set according to service characters based on the temporal evolution of the dynamic word cloud and topic clustering to obtain overall character distribution features; The hierarchical clustering unit is used to hierarchically cluster the service characters based on the policy issuing department and the industry field to obtain characters of the same type.
4. The industrial policy information service system based on intelligent analysis according to claim 3 is characterized in that: The character distribution analysis module further includes: The proportion determination unit is used to calculate the frequency ratio of the same type of characters in each policy cluster within each window, eliminate short-term fluctuations through exponential smoothing, and statistically distribute the proportion according to the policy release area; Importance determination unit, used to construct an evaluation matrix of the issuing entity level, clause constraint strength, and implementation scope, and to obtain the official importance of characters of the same type; A distribution array is obtained based on the distribution ratio and official importance.
5. The industrial policy information service system based on intelligent analysis according to claim 1 is characterized in that: The first encoding module includes: A semantic analysis unit is used to define policy entities, attributes, and relationships, construct an initial graph through human-machine collaborative annotation, and perform cross-modal fusion of the embedded representations of characters of the same type with the triples of the initial graph to calculate the semantic relevance between characters; a logic analysis unit, configured to split the policy into atomic clauses based on a text segmentation algorithm, extract the condition-action-result structure of the atomic clauses using a dependency syntax tree to generate logic tuples, and perform compliance verification on the logic tuples; A time analysis unit is used to perform time series analysis on the historical release time, revision time, and abolition time of characters of the same type to determine an evolution trajectory table; a mechanism encoding unit, configured to perform a first encoding on the characters of the same type based on the association analysis result and using a four-layer encoding mechanism consisting of a domain dimension, a function dimension, a feature dimension, and an instance dimension; Among them, the association analysis results include: the semantic association between characters based on the semantic dimension, the compliance verification results of logical tuples based on the logical dimension, and the evolution trajectory table based on the time dimension.
6. The industrial policy information service system based on intelligent analysis according to claim 1 is characterized in that: The mapping analysis module includes: A demand collection unit is configured to collect the explicit, implicit, and potential demands of the target enterprise, wherein the explicit demands are obtained from the enterprise-side interactive platform based on a structured form and free text enhancement model, the implicit demands are mined based on corporate annual reports, social responsibility reports, and interview recordings, and the potential demands are predicted based on the mining results of the time dimension in the first coding module; A model analysis unit is configured to input the explicit, implicit, and potential demands into an enhanced large-scale model in the field of industrial policy, and output a keyword vector, wherein the enhanced large-scale model in the field of industrial policy is based on a general BERT model, incorporates policy-demand and training corpus, and is trained using a language model masked with lifecycle labels requiring scenario masking, and keyword vectors; an extraction and determination unit, configured to extract statistical patterns of the distribution array based on a multi-layer perceptron, determine semantic connections between codes of the first code based on a graph convolutional network, and capture policy evolution laws of the second code based on a temporal convolutional network; a relationship determination unit, configured to project the keyword vectors, statistical patterns, semantic connections between codes, and policy evolution laws into a three-dimensional coordinate system, respectively, and determine matching relationships based on matching policies, matching relationships based on policy condition dependency chains, and matching relationships based on risk points of historically eliminated characters in the second code and regional differences in distribution arrays; The solution formulation unit is used to obtain strategies for dealing with each matching relationship and generate customized service solutions.
7. The industrial policy information service system based on intelligent analysis according to claim 1 is characterized in that: The second encoding module includes: A difference analysis unit is used to compare the version differences of the crawled industrial policy information in the service field using a text difference algorithm, mark the addition, modification, and deletion operations of the same type of characters, and generate a character iteration trajectory table, wherein the character iteration trajectory table includes the operation time and the impact range of each operation; A code generation unit is used to generate a second code for the same type of character based on the character iteration trajectory table.
8. An industrial policy information service method based on intelligent analysis, characterized in that: include: Step 1: Obtain industrial policy information from a designated platform using crawler technology, and based on subscription information related to the designated platform, obtain different types of multi-source policy data, and pre-process the multi-source policy data to obtain a standardized policy dataset; Step 2: Filtering the standardized policy dataset based on character filtering rules to obtain service characters, and determining the overall character distribution and the distribution of characters of the same type to obtain a distribution array of characters of the same type, wherein the distribution array includes: distribution percentage and official importance of characters of the same type; Step 3: mining the same type of characters from the three dimensions of semantics, logic, and time to obtain association analysis results, and performing a first encoding on the same type of characters using a four-layer encoding mechanism; Step 4: crawl the service field of the industrial policy information, extract the iterative order of the same type of characters in the historical release of the service field, and generate a second code; Step 5: Collect multi-dimensional demand data of the target enterprise, convert the multi-dimensional demand data into a keyword vector, establish a mapping relationship based on the distribution array, the first code, and the second code, and generate a customized service plan; Step 6: Obtain the policy docking effect data and demand change data of the target enterprise based on the customized service solution, and update the character screening rules and the first code and the second code.
Citation Information
Cited By
Policy text multi-modal acquisition and evolution graph analysis method and policy text multi-modal acquisition and evolution graph analysis system
CN122174844A