Policy matching and pushing method and system based on large model technology

By constructing a multi-dimensional enterprise profile database and combining it with large model technology, the problems of inconsistent label dimensions and limited semantic understanding in traditional policy recommendation systems have been solved. This has enabled accurate matching and dynamic adjustment of policies and enterprise data, improving the accuracy and adaptability of recommendations.

CN120849709APending Publication Date: 2025-10-28天元大数据信用管理有限公司

Patent Information

Application Number
CN202510978922.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Traditional policy recommendation systems suffer from inconsistent label dimensions and limited semantic understanding, making it difficult to effectively match multi-dimensional data of policies and enterprises. In particular, they lack systematic solutions for the differentiated handling of first-release and non-first-release policies.

Method used

We employ large-scale model technology to construct a multi-dimensional enterprise profile library. By combining embedding models and large-scale models with machine learning algorithms, we achieve cross-modal semantic alignment, dynamically update matching rules, and adapt to real-time changes in policies and enterprise data.

Benefits of technology

It achieves cross-modal semantic alignment, improves the accuracy and adaptability of policy recommendations, dynamically adjusts the credibility of conditions, solves the problems of inconsistent label dimensions and limited semantic understanding in traditional systems, and meets government compliance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849709A_ABST
    Figure CN120849709A_ABST
Patent Text Reader

Abstract

The invention discloses a policy matching and pushing method and system based on a large model technology, and relates to the technical field of big data processing, and the implementation of the method comprises the following steps: constructing a multi-dimensional enterprise portrait library comprising enterprise basic information, honor information, qualification information, judicial risks, administrative punishment and project information; for the first policy, vectorizing a policy text and an enterprise portrait through an Embedding model, calculating semantic similarity and generating a recommendation list; for a non-first policy, policy conditions are deduced based on a historical reward list, condition credibility is adjusted through a large model and a machine learning model, and recommendation priorities are calculated in a weighted manner; and a data real-time updating and model feedback optimization mechanism is established, and the matching precision is dynamically improved. According to the method, cross-modal semantic alignment can be realized, recommendation accuracy is improved, matching rules are dynamically updated, and the method can adapt to real-time changes of policies and enterprise data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data processing technology, specifically to a policy matching and push method and system based on big model technology. Background Technology

[0002] Traditional policy recommendation systems primarily rely on manual analysis of policy application criteria into enterprise tags (such as industry type, revenue size, number of patents, etc.), and then use tag matching to filter target enterprises. However, this method has the following fundamental flaws:

[0003] (1) Inconsistent label dimensions: Policy application conditions may involve qualitative descriptions (such as “outstanding innovation capabilities”), industry standards (such as “in line with the national strategic emerging industries”) or dynamic indicators (such as “R&D investment growth rate in the past three years”), while enterprise labels are mostly static structured data (such as business registration information). The two have significant differences in semantic expression, data form and evaluation dimensions, which leads to the failure of direct matching.

[0004] (2) Limitations of semantic understanding: Traditional methods rely on rule engines or simple keyword matching, which cannot capture the deep semantics in policy texts (such as implicit conditions and contextual relationships). For example, "high-tech enterprise" may correspond to different recognition standards (national level / provincial level), which traditional systems find difficult to distinguish.

[0005] In recent years, large-scale models have achieved qualitative breakthroughs in natural language processing (NLP) technology, such as keyword extraction through named entity recognition (NER). However, they have not solved the problems of multi-dimensional data fusion and dynamic credibility calculation. The development of large-scale model technologies (such as GPT and BERT) has provided new paths for semantic understanding, but their application in policy matching still lacks systematic solutions, especially for differentiated processing mechanisms for first-time and non-first-time policies. Summary of the Invention

[0006] The technical objective of this invention is to address the above-mentioned shortcomings by providing a policy matching and recommendation method and system based on large model technology, which can achieve cross-modal semantic alignment, improve recommendation accuracy, dynamically update matching rules, and adapt to real-time changes in policy and enterprise data.

[0007] The technical solution adopted by this invention to solve its technical problem is:

[0008] A policy matching and push method based on large model technology, the implementation of which includes the following steps:

[0009] (1) Construct a multi-dimensional enterprise profile database that includes basic enterprise information, honor information, qualification information, legal risks, administrative penalties, and project information;

[0010] (2) For the first policy, the policy text and enterprise profile are vectorized by the Embedding model, the semantic similarity is calculated and a recommendation list is generated;

[0011] (3) For policies that are not first-time policies, the policy conditions are inferred based on the historical list of subsidies and rewards, and the credibility of the conditions is adjusted by using large models and machine learning models to calculate the recommendation priority in a weighted manner.

[0012] (4) Establish a real-time data update and model feedback optimization mechanism to dynamically improve matching accuracy.

[0013] This method constructs a multi-dimensional enterprise profile database, combines vectorized semantic matching with dynamic credibility calculation, breaks through the dimensional barriers between policy tags and enterprise tags, and achieves cross-modal semantic alignment; it distinguishes the matching logic between first-release policies and non-first-release policies, thereby improving recommendation accuracy; and it dynamically updates the matching rules to adapt to real-time changes in policy and enterprise data.

[0014] Furthermore, the Embedding model adopts a BERT-like large model and incorporates an attention mechanism to weight key policy conditions.

[0015] Furthermore, step (3) is achieved by integrating large model semantic analysis and XGBoost machine learning algorithm based on the credibility adjustment model.

[0016] Furthermore, step (1), which involves constructing a multi-dimensional enterprise profile database and performing data governance, specifically includes:

[0017] Data was extracted from heterogeneous systems from multiple sources, such as the State Administration for Industry and Commerce database, Credit China, and the State Intellectual Property Office, using ETL (Extract-Transform-Load) tools.

[0018] Entity linking technology (such as DBpedia Spotlight) can be used to eliminate data ambiguity (such as the problem of different companies having the same name);

[0019] Construct an enterprise knowledge graph with the enterprise as the central node, linking various attributes and relationships (such as "owning patents" and "receiving honors"), and supporting graph query and reasoning.

[0020] Furthermore, step (2) involves vectorizing the policy application eligibility criteria and user profiles using an embedding model, comparing the similarity between policy tags and user profiles, and thus providing a recommended application list for the first policy. Specific implementation includes:

[0021] Policy text processing: Encode policy application conditions using a large model (such as BERT-base) to generate policy semantic vectors. Capture entities (such as "high-tech enterprises"), attributes (such as "annual revenue ≥ 50 million yuan"), and logical relationships (such as "and / or" conditions) in the text; Example: The policy clause "supports high-tech enterprises in the field of new-generation information technology with R&D investment accounting for no less than 5% in the past three years" is encoded and the vector contains semantic features such as "new-generation information technology", "R&D investment ratio", and "high-tech enterprises";

[0022] Enterprise profiling vectorization: Structured data (such as R&D investment ratio) in the enterprise profiling database is normalized and converted into numerical vectors; unstructured data (such as project descriptions) is generated into text vectors using the TextCNN model; and multimodal vectors (numerical + text) are fused into a comprehensive enterprise vector through a feature fusion layer (such as a fully connected neural network).

[0023] Similarity calculation: The cosine similarity formula is used to calculate the matching degree between the policy vector and the enterprise vector.

[0024]

[0025] Set a threshold θ (e.g., 0.6) to filter out companies with S≥θ as candidates. In this process, an attention mechanism needs to be introduced to dynamically weight key conditions in the policy terms (e.g., "R&D investment ratio" has a higher weight than "establishment time") when calculating similarity, thereby improving matching accuracy.

[0026] Furthermore, in step (3), based on the enterprise tags in the historical subsidy list, the eligibility criteria for policy application (enterprise-specific) are inferred, and credibility is adjusted through a large model and machine learning to obtain a weighted priority for list recommendations; specifically, this includes:

[0027] Policy condition reverse inference: Analyze the common tags of enterprises in the historical subsidy list, and use the big model to generate natural language descriptions of policy application conditions (such as "it is speculated that the policy tends to support enterprises with tags A, B, and C"); Example: If 80% of the enterprises in the historical list have "≥3 invention patents" and "50-200 employees", the big model infers that these two items are the core conditions, and the initial confidence levels are set to 80% and 100% respectively;

[0028] Credibility Adjustment Model: Large Model-Driven: Utilizes a GPT-like model to analyze the semantic relationships of historical policy conditions and outputs suggested credibility adjustments (e.g., "Environmental compliance" requirement, credibility +20%); Machine Learning Assistance: Uses an XGBoost model to train on historical data, with inputs including enterprise labels and policy text features, and outputs a corrected credibility value (e.g., a policy in a certain region has a higher weight for "local tax payment", corrected credibility +15%); Fusion Strategy: Generates the final credibility through weighted averaging.

[0029] Recommendation priority calculation: A weighted sum is calculated based on the conditions met by the enterprise, using the formula: ∑(C i *f i ), where f i To determine whether a company meets the conditions, if it meets f i =1, does not satisfy f i =0; C i This indicates the credibility of the application conditions for the i-th policy. The recommendation results are sorted in descending order of priority to generate a recommendation list.

[0030] This invention also claims a policy matching and push system based on large model technology, comprising:

[0031] The multi-dimensional enterprise profile database construction module is used to build a multi-dimensional enterprise profile database that includes basic enterprise information, honors information, qualification information, legal risks, administrative penalties, and project information.

[0032] The first policy implementation module uses the Embedding model to vectorize policy texts and enterprise profiles, calculates semantic similarity, and generates a recommendation list.

[0033] The non-first-release policy implementation module infers policy conditions based on historical reward and subsidy lists, adjusts the credibility of conditions through large models and machine learning models, and calculates the recommendation priority with weighted averages.

[0034] The feedback and optimization module is used to establish a real-time data update and model feedback optimization mechanism to dynamically improve matching accuracy.

[0035] This system can implement the policy matching and push method based on the large model technology mentioned above.

[0036] Furthermore, the specific architecture of the system includes:

[0037] The data layer includes data source interface units and data processing units;

[0038] Data sources include: government information disclosure platforms (policy texts), enterprise credit information disclosure systems, and the State Intellectual Property Office API;

[0039] Data processing includes: batch data cleaning using Apache Spark, and real-time data streaming using Flink;

[0040] The model layer includes the base model and custom models;

[0041] The base models include: the Hugging Face open-source BERT model and a finely tuned version of GPT-3.5;

[0042] Custom models include: a feature fusion model developed based on PyTorch and an XGBoost confidence adjustment model;

[0043] The application layer includes a policy management platform and an enterprise portal;

[0044] Policy management platform: Allows government users to upload policy texts and view recommended lists;

[0045] Enterprise Portal: Push matching policies to enterprises and display matching details.

[0046] The present invention also claims a policy matching push device based on large model technology, comprising: at least one memory and at least one processor;

[0047] The at least one memory is used to store a machine-readable program;

[0048] The at least one processor is used to call the machine-readable program to implement the above method.

[0049] The present invention also claims a computer-readable medium storing computer instructions that, when executed by a processor, implement the above-described method.

[0050] Compared with existing technologies, the policy matching and push method and system based on large model technology of the present invention have the following advantages:

[0051] 1. Precise cross-dimensional semantic matching. By using a large model, the unstructured semantics of policy texts (such as "outstanding innovation capabilities") and the structured / unstructured data of enterprise profiles (such as patent numbers and textual descriptions of R&D investment) are converted into a unified semantic vector, solving the core problem of inconsistent dimensions between policy labels and enterprise labels in traditional methods.

[0052] 2. Differentiated scenario adaptation to enhance recommendation generalization ability. For first-time policy scenarios: without relying on historical reward and subsidy data, the recommendation list is generated directly through semantic parsing of a large model, solving the dilemma of traditional systems that "cannot recommend without historical data"; for non-first-time policy scenarios: combining historical experience with real-time semantic analysis, the credibility of conditions is dynamically adjusted through "large model inference + machine learning optimization" (such as increasing the weight of "local tax amount" based on regional policy characteristics), avoiding the overfitting problem caused by simply relying on historical data and improving recommendation accuracy.

[0053] 3. Enhanced interpretability, meeting government compliance requirements. During the vectorized matching process, key matching features between policy vectors and enterprise vectors are visualized, facilitating policy implementers' verification of the rationality of the recommendation logic. The credibility adjustment model outputs detailed conditional weights, meeting the government system's compliance requirements for traceable recommendation results and enhancing government user trust.

[0054] 4. Multimodal data fusion for a more comprehensive enterprise profile. A three-dimensional enterprise profile library with six dimensions is constructed, which not only includes static enterprise data, but also uses knowledge graph technology to link multi-dimensional enterprise attributes (such as the relationship between "enterprise-patent-technology field"), supports complex logical reasoning (such as determining whether an enterprise belongs to the policy target group of "patent-intensive + high R&D investment"), and enhances the richness of matching dimensions. Attached Figure Description

[0055] Figure 1 This is a flowchart illustrating the policy matching and push method based on large model technology provided in an embodiment of the present invention. Detailed Implementation

[0056] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0057] This invention provides a policy matching and push method based on large model technology. The implementation of this method includes the following steps:

[0058] (1) Construct a multi-dimensional enterprise profile database that includes basic enterprise information, honor information, qualification information, legal risks, administrative penalties, and project information;

[0059] (2) For the first policy, the policy text and enterprise profile are vectorized by the Embedding model, the semantic similarity is calculated and a recommendation list is generated;

[0060] (3) For policies that are not first-time policies, the policy conditions are inferred based on the historical list of subsidies and rewards, and the credibility of the conditions is adjusted by using large models and machine learning models to calculate the recommendation priority in a weighted manner.

[0061] (4) Establish a real-time data update and model feedback optimization mechanism to dynamically improve matching accuracy.

[0062] In step (2), the Embedding model adopts a BERT-like large model and combines an attention mechanism to weight the key policy conditions.

[0063] Step (3) is achieved by integrating large model semantic analysis and XGBoost machine learning algorithm based on the credibility adjustment model.

[0064] This method, by constructing a multi-dimensional enterprise profile database and combining vectorized semantic matching with dynamic credibility calculation, can achieve the following:

[0065] Breaking down the dimensional barriers between policy labels and enterprise labels to achieve cross-modal semantic alignment;

[0066] Differentiate the matching logic between initial and non-initial policies to improve recommendation accuracy;

[0067] The matching rules are dynamically updated to adapt to real-time changes in policies and enterprise data.

[0068] The specific implementation process of this method is as follows:

[0069] 1. Construct an enterprise profile database encompassing six dimensions (including basic enterprise information, honors information, qualification information, legal risks, administrative penalties, and project information), and conduct data governance.

[0070] (1) Data was extracted from multiple heterogeneous systems such as the industrial and commercial database, Credit China, and the State Intellectual Property Office using the ETL (Extract-Transform-Load) tool.

[0071] (2) Use entity linking technology (such as DBpedia Spotlight) to eliminate data ambiguity (such as the problem of different companies having the same name).

[0072] (3) Construct an enterprise knowledge graph with the enterprise as the central node, linking various attributes and relationships (such as "owning patents" and "receiving honors"), and supporting graph query and reasoning.

[0073] 2. Initial Policy Matching Process (No Historical Subsidy List): The policy application eligibility criteria and user profiles are vectorized using an embedding model. The similarity between policy tags and user profiles is compared to generate a recommended application list for the initial policy.

[0074] (1) Policy text processing: Use a large model (such as BERT-base) to encode the policy application conditions and generate policy semantic vectors. Capture entities (such as "high-tech enterprises"), attributes (such as "annual revenue ≥ 50 million yuan"), and logical relationships (such as "and / or" conditions) in the text. Example: The policy clause "Support high-tech enterprises in the field of next-generation information technology with R&D investment accounting for no less than 5% in the past three years" is encoded so that the vector contains semantic features such as "next-generation information technology", "R&D investment ratio", and "high-tech enterprises".

[0075] (2) Enterprise Profile Vectorization: The structured data (such as the proportion of R&D investment) in the enterprise profile database is normalized and converted into numerical vectors; the unstructured data (such as project descriptions) is generated into text vectors using the TextCNN model; and the multimodal vectors (numerical + text) are fused into a comprehensive enterprise vector through a feature fusion layer (such as a fully connected neural network).

[0076] (3) Similarity Calculation: The cosine similarity formula is used to calculate the matching degree between the policy vector and the enterprise vector.

[0077]

[0078] Set a threshold θ (e.g., 0.6) to filter out companies with S≥θ as candidates. In this process, an attention mechanism needs to be introduced to dynamically weight key conditions in the policy terms (e.g., "R&D investment ratio" has a higher weight than "establishment time") when calculating similarity, thereby improving matching accuracy.

[0079] 3. Non-first-release policy matching process (with historical subsidy lists): Based on the enterprise tags in the historical subsidy lists, the eligibility conditions for policy application (enterprise-specific) are inferred. The credibility is adjusted by a large model and machine learning, and the weighted list recommendation priority is obtained.

[0080] (1) Policy Conditions Backward Inference: Analyze the common labels of enterprises in the historical subsidy list and use a large model to generate natural language descriptions of policy application conditions (e.g., "It is speculated that this policy tends to support enterprises with labels A, B, and C"). Example: If 80% of the enterprises in the historical list have "≥3 invention patents" and "50-200 employees", the large model infers that these two items are the core conditions, with initial confidence levels set at 80% and 100%, respectively.

[0081] (2) Credibility Adjustment Model: A large-scale model is used to analyze the semantic relationships of historical policy conditions using a GPT-like model, outputting suggested credibility adjustments (e.g., for "environmental compliance" requirements, credibility +20%). Machine Learning Assistance: Historical data is trained using an XGBoost model, with inputs including company labels and policy text features, outputting a corrected credibility value (e.g., for a certain region, policies have a higher weighting for "local tax payments," correcting credibility +15%). Fusion Strategy: A weighted average is used to generate the final credibility score.

[0082] (3) Recommendation Priority Calculation: The conditions met by the enterprise are weighted and summed, and the formula is: ∑(C i *f i ), where f i To determine whether a company meets the conditions, if it meets f i =1, does not satisfy f i =0; C i This indicates the credibility of the application conditions for the i-th policy. The recommendation results are sorted in descending order of priority to generate a recommendation list.

[0083] Suppose a policy has three core conditions, with the following credibility weights: Condition 1: High-tech enterprise qualification (C1 = 0.3), Condition 2: Established for more than 3 years with ≥5% [company name missing] (C2 = 0.4), Condition 3: Employee size of 50-200 people (C3 = 0.2).

[0084] Company A: Meets conditions 1 and 2, but does not meet condition 3;

[0085] Priority score = 0.3*1 + 0.4*1 + 0.2*0 = 0.7.

[0086] Company B: Meets conditions 2 and 3, but does not meet condition 1;

[0087] Priority score = 0.3*0 + 0.4*1 + 0.2*1 = 0.6 Result: Company A's score (0.7) is higher than Company B's (0.6), therefore Company A has a higher priority in the recommended list.

[0088] This method constructs a corporate profile database encompassing six dimensions, including basic enterprise information and honors information. It utilizes ETL tools to extract multi-source data and builds a knowledge graph for data governance. It differentiates between first-release and non-first-release policies. For first-release policies, it uses large-scale models such as BERT to vectorize the policy text and corporate profiles, and combines an attention mechanism to calculate semantic similarity to generate a recommendation list. For non-first-release policies, it infers policy conditions based on historical reward lists using large-scale models, and adjusts the credibility of the conditions using a fusion of GPT-like models and XGBoost algorithms, weighting the calculation of recommendation priorities. Simultaneously, it establishes a real-time data update and model feedback optimization mechanism to dynamically improve matching accuracy. This method overcomes the bottleneck of inconsistent label dimensions in traditional policy recommendation systems, achieving accurate matching between policies and enterprises, and is suitable for intelligent push scenarios for different types of policies.

[0089] This invention also provides a policy matching and push system based on large model technology, comprising:

[0090] The multi-dimensional enterprise profile database construction module is used to build a multi-dimensional enterprise profile database that includes basic enterprise information, honors information, qualification information, legal risks, administrative penalties, and project information.

[0091] The first policy implementation module uses an embedding model to vectorize policy texts and enterprise profiles, calculates semantic similarity, and generates a recommendation list.

[0092] The non-first-release policy implementation module infers policy conditions based on historical reward and subsidy lists, adjusts the credibility of conditions through large-scale models and machine learning models, and calculates the recommendation priority with weighted averages.

[0093] The feedback and optimization module is used to establish a real-time data update and model feedback optimization mechanism to dynamically improve matching accuracy.

[0094] This system can implement the policy matching and push method based on large model technology described in the above embodiments.

[0095] The specific implementation process is as follows:

[0096] 1. The multi-dimensional enterprise profile database construction module constructs an enterprise profile database containing 6 dimensions (including basic enterprise information, honors information, qualification information, legal risks, administrative penalties, and project information) and performs data governance.

[0097] (1) Data was extracted from multiple heterogeneous systems such as the industrial and commercial database, Credit China, and the State Intellectual Property Office using the ETL (Extract-Transform-Load) tool.

[0098] (2) Use entity linking technology (such as DBpedia Spotlight) to eliminate data ambiguity (such as the problem of different companies having the same name).

[0099] (3) Construct an enterprise knowledge graph with the enterprise as the central node, linking various attributes and relationships (such as "owning patents" and "receiving honors"), and supporting graph query and reasoning.

[0100] 2. The initial policy implementation module executes the initial policy matching process (without a historical list of subsidies): It uses an embedding model to vectorize the eligibility criteria for policy application and user profiles, compares the similarity between policy tags and user profiles, and thus provides a recommended application list for the initial policy.

[0101] (1) Policy text processing: Use a large model (such as BERT-base) to encode the policy application conditions and generate policy semantic vectors. Capture entities (such as "high-tech enterprises"), attributes (such as "annual revenue ≥ 50 million yuan"), and logical relationships (such as "and / or" conditions) in the text. Example: The policy clause "Support high-tech enterprises in the field of next-generation information technology with R&D investment accounting for no less than 5% in the past three years" is encoded so that the vector contains semantic features such as "next-generation information technology", "R&D investment ratio", and "high-tech enterprises".

[0102] (2) Enterprise Profile Vectorization: The structured data (such as the proportion of R&D investment) in the enterprise profile database is normalized and converted into numerical vectors; the unstructured data (such as project descriptions) is generated into text vectors using the TextCNN model; and the multimodal vectors (numerical + text) are fused into a comprehensive enterprise vector through a feature fusion layer (such as a fully connected neural network).

[0103] (3) Similarity Calculation: The cosine similarity formula is used to calculate the matching degree between the policy vector and the enterprise vector.

[0104]

[0105] Set a threshold θ (e.g., 0.6) to filter out companies with S≥θ as candidates. In this process, an attention mechanism needs to be introduced to dynamically weight key conditions in the policy terms (e.g., "R&D investment ratio" has a higher weight than "establishment time") when calculating similarity, thereby improving matching accuracy.

[0106] 3. Non-first-release policy implementation module executes non-first-release policy matching process (with historical subsidy list): Based on the enterprise tags in the historical subsidy list, infer the policy application eligibility conditions (enterprise scope), and use a large model as the main approach + machine learning to adjust the credibility and weight the list recommendation priority.

[0107] (1) Policy Conditions Backward Inference: Analyze the common labels of enterprises in the historical subsidy list and use a large model to generate natural language descriptions of policy application conditions (e.g., "It is speculated that this policy tends to support enterprises with labels A, B, and C"). Example: If 80% of the enterprises in the historical list have "≥3 invention patents" and "50-200 employees", the large model infers that these two items are the core conditions, with initial confidence levels set at 80% and 100%, respectively.

[0108] (2) Credibility Adjustment Model: A large-scale model is used to analyze the semantic relationships of historical policy conditions using a GPT-like model, outputting suggested credibility adjustments (e.g., for "environmental compliance" requirements, credibility +20%). Machine Learning Assistance: Historical data is trained using an XGBoost model, with inputs including company labels and policy text features, outputting a corrected credibility value (e.g., for a certain region, policies have a higher weighting for "local tax payments," correcting credibility +15%). Fusion Strategy: A weighted average is used to generate the final credibility score.

[0109] (3) Recommendation Priority Calculation: The conditions met by the enterprise are weighted and summed, and the formula is: ∑(C i *f i ), where f i To determine whether a company meets the conditions, if it meets f i =1, does not satisfy f i =0; C i This indicates the credibility of the application conditions for the i-th policy. The recommendation results are sorted in descending order of priority to generate a recommendation list.

[0110] 4. The feedback and optimization module establishes a real-time data update and model feedback optimization mechanism to dynamically improve matching accuracy.

[0111] The specific architecture of the system includes:

[0112] The data layer includes data source interface units and data processing units.

[0113] Data sources include: government information disclosure platforms (policy texts), enterprise credit information disclosure systems, and the State Intellectual Property Office API;

[0114] Data processing includes: batch data cleaning using Apache Spark and real-time data streaming using Flink.

[0115] The model layer contains the base model and custom models.

[0116] The base models include: the Hugging Face open-source BERT model and a finely tuned version of GPT-3.5;

[0117] Custom models include: a feature fusion model developed based on PyTorch and an XGBoost credibility adjustment model.

[0118] The application layer includes a policy management platform and an enterprise portal.

[0119] Policy management platform: Allows government users to upload policy texts and view recommended lists;

[0120] Enterprise Portal: Push matching policies to enterprises and display matching details.

[0121] This invention also provides a policy matching and push device based on large model technology, comprising: at least one memory and at least one processor;

[0122] The at least one memory is used to store a machine-readable program;

[0123] The at least one processor is used to call the machine-readable program to implement the policy matching and push method based on large model technology described in the above embodiments.

[0124] This invention also provides a computer-readable medium storing computer instructions. When executed by a processor, the computer instructions cause the processor to perform the policy matching and push method based on large model technology described in the above embodiments. Specifically, a system or apparatus equipped with a storage medium storing software program code that implements the functions of any of the embodiments described above can be provided, and the computer (or CPU or MPU) of the system or apparatus can read and execute the program code stored in the storage medium.

[0125] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.

[0126] Examples of storage media used to provide program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.

[0127] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0128] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion unit execute some and all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0129] The present invention has been shown and described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above embodiments, those skilled in the art will know that more embodiments of the present invention can be obtained by combining the code review methods in the different embodiments. These embodiments are also within the protection scope of the present invention.

Claims

1. A policy matching and push method based on large model technology, characterized in that, The implementation of this method includes the following steps: (1) Construct a multi-dimensional enterprise profile database that includes basic enterprise information, honor information, qualification information, legal risks, administrative penalties, and project information; (2) For the first policy, the policy text and enterprise profile are vectorized by the Embedding model, the semantic similarity is calculated and a recommendation list is generated; (3) For policies that are not first-time policies, the policy conditions are inferred based on the historical list of subsidies and rewards, and the credibility of the conditions is adjusted by using large models and machine learning models to calculate the recommendation priority in a weighted manner. (4) Establish a real-time data update and model feedback optimization mechanism to dynamically improve matching accuracy.

2. The policy matching and push method based on large model technology according to claim 1, characterized in that, The Embedding model adopts a BERT-like large model and combines an attention mechanism to weight key policy conditions.

3. The policy matching and push method based on large model technology according to claim 1, characterized in that, Step (3) is achieved by integrating large model semantic analysis and XGBoost machine learning algorithm based on the credibility adjustment model.

4. The policy matching and push method based on large model technology according to claim 1, characterized in that, Step (1), which involves constructing a multi-dimensional enterprise profile database and performing data governance, specifically includes: Extracting data from multi-source heterogeneous systems using ETL tools; Entity linking technology is used to eliminate data ambiguity; Construct an enterprise knowledge graph with the enterprise as the central node, linking attributes and relationships across various dimensions, and supporting graph query and reasoning.

5. A policy matching and push method based on large model technology according to claim 1 or 2, characterized in that, Step (2) involves vectorizing the policy application eligibility criteria and user profiles using an embedding model, comparing the similarity between policy tags and user profiles, and thus providing a recommended application list for the first policy. Specific implementation includes: Policy text processing: Encode policy application conditions using a large model to generate policy semantic vectors. Capture entities, attributes, and logical relationships in text; Enterprise profiling vectorization: Structured data in the enterprise profiling database is normalized and converted into numerical vectors; unstructured data is used to generate text vectors using the TextCNN model; and multimodal vectors are fused into a comprehensive enterprise vector through a feature fusion layer. Similarity calculation: The cosine similarity formula is used to calculate the matching degree between the policy vector and the enterprise vector. A threshold θ is set to filter out companies with S≥θ as candidates; an attention mechanism is introduced to dynamically weight key conditions in policy clauses when calculating similarity, thereby improving matching accuracy.

6. The policy matching and push method based on large model technology according to claim 5, characterized in that, Step (3) involves inferring the eligibility criteria for policy application based on the enterprise tags in the historical subsidy list, adjusting the credibility through a large model and machine learning, and weighting the list recommendation priority; the specific implementation includes: Policy conditions reverse inference: Analyze the common tags of enterprises in the historical subsidy list and use a large model to generate natural language descriptions of policy application conditions; Credibility Adjustment Model: Large Model-Driven: Utilizes a GPT-like model to analyze the semantic relationships of historical policy conditions and outputs suggestions for adjusting condition credibility; Machine Learning Assistance: Uses an XGBoost model to train historical data, with inputs including enterprise labels and policy text features, and outputs corrected condition credibility values; Fusion Strategy: Generates the final credibility through weighted averaging. Recommendation priority calculation: A weighted sum is calculated based on the conditions met by the enterprise, using the formula: ∑(C i *f i ), where f i To determine whether a company meets the conditions, if it meets f i =1, does not satisfy f i =0; C i This indicates the credibility of the application conditions for the i-th policy. The recommendation results are sorted in descending order of priority to generate a recommendation list.

7. A policy matching and recommendation system based on large-scale model technology, characterized in that, include: The multi-dimensional enterprise profile database construction module is used to build a multi-dimensional enterprise profile database that includes basic enterprise information, honors information, qualification information, legal risks, administrative penalties, and project information. The first policy implementation module uses the Embedding model to vectorize policy texts and enterprise profiles, calculates semantic similarity, and generates a recommendation list. The non-first-release policy implementation module infers policy conditions based on historical reward and subsidy lists, adjusts the credibility of conditions through large models and machine learning models, and calculates the recommendation priority with weighted averages. The feedback and optimization module is used to establish a real-time data update and model feedback optimization mechanism to dynamically improve matching accuracy. The system is capable of implementing the policy matching and push method based on large model technology as described in any one of claims 1 to 6.

8. A policy matching and push system based on large model technology according to claim 7, characterized in that, The specific architecture of the system includes: The data layer includes data source interface units and data processing units; Data sources include: government information disclosure platforms, enterprise credit information disclosure systems, and the State Intellectual Property Office API; Data processing includes: batch data cleaning using Apache Spark, and real-time data streaming using Flink; The model layer includes the base model and custom models; The base models include: the Hugging Face open-source BERT model and a finely tuned version of GPT-3.5; Custom models include: a feature fusion model developed based on PyTorch and an XGBoost confidence adjustment model; The application layer includes a policy management platform and an enterprise portal; Policy management platform: Allows government users to upload policy texts and view recommended lists; Enterprise Portal: Push matching policies to enterprises and display matching details.

9. A policy matching and push device based on large model technology, characterized in that, include: At least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to invoke the machine-readable program to implement the method according to any one of claims 1 to 6.

10. A computer-readable medium, characterized in that, The computer-readable medium stores computer instructions that, when executed by a processor, implement the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Policy data intelligent recommendation method and system

    CN113468418A

  • Project matching method and device based on similar enterprises, equipment and medium

    CN116523473A

  • Policy automation analysis method and device, electronic equipment and storage medium

    CN116681560A

  • A large-model neural network for policy recommendation

    CN119740661A

  • Electric vehicle industry policy document interpretation method and device based on large language model

    CN120106020A

Cited By

  • Method and device for policy-enterprise intelligent matching

    CN121597735A

  • Policy adaptation and navigation method

    CN121685231A

  • Science and technology enterprise portrait information pushing system based on big data

    CN121887857A