Government-enterprise matching method and system based on large model, and storage medium

By using a large-scale model-based government-enterprise matching method, policy documents and enterprise information are processed automatically, solving the problem of low efficiency in government-enterprise matching and achieving efficient government-enterprise information matching.

CN119988745BActive Publication Date: 2026-02-13SI CHUAN YUN ZHI SHENG ZHI NENG KE JI YOU XIAN GONG SI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510205524.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2026-02-13
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The current government-enterprise matching process is inefficient, mainly because the fragmented and complex policy information makes it difficult for enterprises to accurately match suitable government incentive policies.

Method used

A government-enterprise matching method based on a large model is adopted. By performing text recognition and segmentation on policy documents, a pre-trained large model is used to extract policy content and match enterprise information. The application conditions and incentive content in the policy are automatically extracted, and government-enterprise matching is performed based on matching prompts.

Benefits of technology

It eliminates the need for manual policy analysis and government-enterprise matching, improving the efficiency of government-enterprise matching and achieving automated government-enterprise information matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988745B_ABST
    Figure CN119988745B_ABST
Patent Text Reader

Abstract

The application provides a government-enterprise matching method and system based on a large model and a storage medium, and the method comprises the following steps: performing character recognition on a policy document to obtain a policy text, inputting the policy text into a pre-trained large model to perform policy splitting and obtain policy content; performing conditional extraction on the policy content according to an extraction prompt word to obtain policy extraction information, obtaining enterprise information to be matched and government-enterprise matching requirements; determining a matching prompt word according to the government-enterprise matching requirements, and controlling the pre-trained large model to perform government-enterprise matching on the enterprise information to be matched and the policy extraction information according to the matching prompt word to obtain a government-enterprise matching result. According to the embodiment of the application, the pre-trained large model can automatically perform government-enterprise matching on the enterprise information to be matched and the policy extraction information, and it is not necessary to perform policy analysis and government-enterprise matching in an artificial manner, so that the government-enterprise matching efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a government-enterprise matching method and system based on a large model and a storage medium. BACKGROUND

[0002] In the current economic environment, various types of subsidy policies introduced by the government should become a strong support for enterprises in the process of seeking development, but due to the dispersion and complexity of policy information, enterprises often have difficulty in accurately matching the policies suitable for themselves, resulting in a significant reduction in the landing effect of the policies, and therefore, the problem of government-enterprise matching is increasingly valued by people.

[0003] In the existing government-enterprise matching process, the policy information and enterprise information are generally sorted out manually, and the government-enterprise matching is performed based on the sorting results, which leads to low efficiency of government-enterprise matching. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide a government-enterprise matching method and system based on a large model and a storage medium to solve the problem of low efficiency of government-enterprise matching in the prior art.

[0005] The embodiments of the present application are implemented in the following manner: a government-enterprise matching method based on a large model, the method comprising:

[0006] performing character recognition on a policy document to obtain a policy text, and inputting the policy text into a pre-trained large model to perform policy splitting to obtain policy content;

[0007] performing conditional extraction on the policy content according to an extraction prompt word to obtain policy extraction information, and obtaining enterprise information to be matched and government-enterprise matching requirements, the policy extraction information including reporting conditions and subsidy content;

[0008] determining a matching prompt word according to the government-enterprise matching requirements, and controlling the pre-trained large model to perform government-enterprise matching on the enterprise information to be matched and the policy extraction information according to the matching prompt word to obtain a government-enterprise matching result.

[0009] Preferably, before inputting the policy text into the pre-trained large model to perform policy splitting, the method further comprises:

[0010] obtaining a policy sample, a sample prompt word, and a sample extraction word, and inputting the policy sample and the sample prompt word into the large model to perform sample disassembly to obtain a content disassembly sample;

[0011] performing sample extraction on the content disassembly sample according to the sample extraction word to obtain sample extraction information, and obtaining sample enterprise information and a sample matching word;

[0012] According to the sample matching word control the big model to the sample enterprise information and the sample extraction information carries out the government enterprise matching, obtains the sample matching result, and according to the sample matching result, the sample extraction information and the content disassembly sample determine model loss;

[0013] According to the model loss parameter update is carried out to the big model until the big model converges, obtains the pre-training of the big model.

[0014] Preferably, the policy sample and the sample prompt word are input into the big model to carry out sample disassembly, obtain content disassembly sample, including:

[0015] According to the big model, the policy sample is segmented, and sample segmentation is obtained, and the sample segmentation and the sample prompt word are processed by word embedding, to obtain sample word embedding features and prompt word embedding features;

[0016] The sample word embedding features and the prompt word embedding features are converted into vectors, to obtain sample word vectors and prompt word vectors, and the similarity between the sample word vectors and the prompt word vectors is calculated, to obtain a first vector similarity;

[0017] If the first vector similarity is greater than the similarity threshold, the sample sentence where the first vector similarity corresponds to the sample segmentation is determined as a disassembly sentence, and the disassembly sentence is combined to obtain the content disassembly sample.

[0018] Preferably, according to the sample extraction word, the content disassembly sample is extracted to obtain sample extraction information, including:

[0019] The sample word vectors corresponding to the disassembly sentence in the content disassembly sample are combined to obtain a sample combined vector, and the sample combined vector is used to determine a sentence sample semantic;

[0020] The sample extraction word is processed by word embedding to obtain extraction word embedding features, and the extraction word embedding features are converted into extraction word vectors;

[0021] The similarity between the extraction word vectors and the sample combined vectors is calculated to obtain a second vector similarity, and the target matching vector in the sample combined vector is determined according to the second vector similarity;

[0022] The semantic object vocabulary in the corresponding sentence sample semantic of the target matching vector is obtained, and the semantic object vocabulary is extracted to obtain the sample extraction information.

[0023] Preferably, the policy file is subjected to character recognition to obtain a policy text, including:

[0024] The policy file is subjected to gray processing to obtain a gray image, and the gray image is subjected to normalization processing to obtain a normalized image;

[0025] According to different convolution scales, the normalized image is subjected to convolution processing to obtain convolution features, and the convolution features are subjected to feature fusion according to the convolution scales to obtain a feature pyramid;

[0026] Text prediction is performed according to the feature pyramid to obtain a text prediction result, a target text box is determined according to the text prediction result, and the policy text is generated according to the target text box.

[0027] Preferably, the target text box is determined according to the text prediction result, comprising:

[0028] The text existence probability in the text prediction result is obtained, and the text existence probability is compared with a probability threshold;

[0029] If the text existence probability is greater than the probability threshold, the text box corresponding to the text existence probability is determined as a candidate text box, and the overlapping degree between different candidate text boxes is calculated;

[0030] The text box score of the candidate text box is determined according to the overlapping degree and the text existence probability, and the target text box is determined according to the text box score.

[0031] Preferably, the enterprise information to be matched and the government-enterprise matching demand are obtained, comprising:

[0032] The government-enterprise matching instruction is obtained, and the demand information in the government-enterprise matching instruction is extracted to obtain the government-enterprise matching demand;

[0033] The enterprise name and the private domain enterprise information in the government-enterprise matching instruction are obtained, and the public domain enterprise information is queried according to a preset API interface;

[0034] The public domain enterprise information and the private domain enterprise information are combined to obtain the enterprise information to be matched.

[0035] Another object of the embodiment of the application is to provide a government-enterprise matching system based on a large model, comprising:

[0036] A policy splitting module is configured to perform character recognition on a policy file to obtain a policy text, and input the policy text into a large model after pre-training to perform policy splitting and obtain policy content;

[0037] The conditional extraction module is configured to perform conditional extraction on the policy content according to an extraction prompt word, to obtain policy extraction information, and to obtain enterprise information to be matched and a government-enterprise matching demand, wherein the policy extraction information includes a declaration condition and an award content.

[0038] The government-enterprise matching module is configured to determine a matching prompt word according to the government-enterprise matching demand, and to control the pre-trained large model to perform government-enterprise matching on the enterprise information to be matched and the policy extraction information according to the matching prompt word, to obtain a government-enterprise matching result.

[0039] Preferably, the policy splitting module is further configured to:

[0040] obtain a policy sample, a sample prompt word and a sample extraction word, and input the policy sample and the sample prompt word into the large model to perform sample splitting, to obtain a content splitting sample;

[0041] perform sample extraction on the content splitting sample according to the sample extraction word, to obtain sample extraction information, and obtain sample enterprise information and a sample matching word;

[0042] control the large model to perform government-enterprise matching on the sample enterprise information and the sample extraction information according to the sample matching word, to obtain a sample matching result, and determine a model loss according to the sample matching result, the sample extraction information and the content splitting sample;

[0043] perform parameter updating on the large model according to the model loss, until the large model converges, to obtain the pre-trained large model.

[0044] According to the embodiments of the present application, by inputting a policy text into a pre-trained large model to perform policy splitting, the policy content in the policy text can be automatically extracted, by performing conditional extraction on the policy content according to an extraction prompt word, the declaration condition and the award content in the policy content can be automatically extracted, and by performing government-enterprise matching on enterprise information to be matched and policy extraction information based on the pre-trained large model, the policy analysis and government-enterprise matching can be performed automatically without manual intervention, and the government-enterprise matching efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 is a flowchart of the government-enterprise matching method based on a large model provided by the first embodiment of the present application;

[0046] Figure 2 is a structural schematic diagram of the government-enterprise matching system based on a large model provided by the second embodiment of the present application;

[0047] Figure 3 is a schematic diagram of the policy information extraction by the government-enterprise matching system based on a large model provided by the second embodiment of the present application;

[0048] Figure 4 is a schematic diagram of government-enterprise information matching based on a large model provided by the second embodiment of the present application;

[0049] Figure 5 is a structural schematic diagram of a terminal device provided by the third embodiment of the present application. DETAILED DESCRIPTION

[0050] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0051] In order to illustrate the technical solutions of the present application, the following will be described through specific embodiments.

[0052] Embodiment One

[0053] Please refer to Figure 1 is a flowchart of a government-enterprise matching method based on a large model provided by the first embodiment of the present application. The government-enterprise matching method based on a large model can be applied to any device or system. The government-enterprise matching method based on a large model includes the following steps:

[0054] Step S10, performing character recognition on a policy document to obtain a policy text, and inputting the policy text into a pre-trained large model to perform policy splitting to obtain policy content;

[0055] The policy document is a policy information about subsidies or preferential treatment for enterprises issued by a local government. By performing character recognition on the policy document, the image format policy document can be effectively converted into corresponding text information. In this step, by inputting the policy text into the pre-trained large model to perform policy splitting, the policy content in the policy text can be automatically extracted without manual extraction of the text information by the user.

[0056] Optionally, before inputting the policy text into the pre-trained large model to perform policy splitting, the method further includes:

[0057] Obtaining a policy sample, a sample prompt word and a sample extraction word, and inputting the policy sample and the sample prompt word into the large model for sample disassembly to obtain a content disassembly sample; wherein the policy sample, the sample prompt word and the sample extraction word can be set according to requirements, in this step, the large model performs word segmentation on the policy sample, the sample prompt word and the sample extraction word to obtain word units, and performs vector conversion on the word units to obtain word unit vectors, each word unit vector contains semantic information, and the vector relationship between the word units reflects semantic association. The large model learns semantic association relationships through pre-training, and analyzes the vector sequence corresponding to the policy sample based on the language knowledge and pattern recognition capability obtained through pre-training, captures the semantic connection between the sample prompt word and the policy sample through a multi-layer neural network, deeply analyzes the syntax structure and semantic role of the policy sample, determines the logical relationship between the parts, searches for matching key information in the semantic representation of the policy sample according to the understanding of the sample prompt word, and obtains the content disassembly sample;

[0058] According to the sample extraction word, sample extraction is performed on the content disassembly sample to obtain sample extraction information, and sample enterprise information and sample matching words are obtained; wherein the large model analyzes the semantic association between the sample extraction word and the content disassembly sample according to the language knowledge and semantic understanding capability learned through pre-training, finds the part with high matching degree with the sample extraction word vector in the semantic vector representation of the content disassembly sample, and obtains the sample extraction information;

[0059] According to the sample matching word, the large model controls the sample enterprise information and the sample extraction information to perform government-enterprise matching to obtain a sample matching result, determines a model loss according to the sample matching result, the sample extraction information and the content disassembly sample, and updates parameters of the large model according to the model loss until the large model converges, and obtains the large model after pre-training; wherein the large model determines matching parameters based on the sample matching word, obtains matching information in the sample enterprise information and the sample extraction information based on the matching parameters, obtains enterprise matching information and extraction matching information, and can obtain the sample matching result by comparing the enterprise matching information and the extraction matching information.

[0060] Further, the policy sample and the sample prompt word are input into the large model for sample disassembly to obtain a content disassembly sample, including:

[0061] According to the large model, the policy sample is segmented to obtain a sample segmentation, and the sample segmentation and the sample prompt word are subjected to word embedding processing to obtain sample word embedding features and prompt word embedding features; wherein the word embedding processing on the sample segmentation and the sample prompt word can effectively extract the semantic representation of the sample segmentation and the sample prompt word;

[0062] vector conversion is performed on the sample word embedding features and the prompt word embedding features to obtain sample word vectors and prompt word vectors, and a similarity between the sample word vectors and the prompt word vectors is calculated to obtain a first vector similarity; wherein, a Euclidean distance formula can be used to calculate a distance between the sample word vectors and the prompt word vectors to obtain the first vector similarity;

[0063] If the first vector similarity is greater than a similarity threshold, the sample sentence in which the first vector similarity corresponds to the sample word segmentation is determined as a disassembled sentence, and the disassembled sentence is combined to obtain the content disassembled sample; wherein, the similarity threshold can be set as needed, and if the first vector similarity is greater than the similarity threshold, it is determined that the sample sentence in which the first vector similarity corresponds to the sample word segmentation contains policy information, and therefore, the sample sentence in which the first vector similarity corresponds to the sample word segmentation is determined as the disassembled sentence.

[0064] Further, sample extraction is performed on the content disassembled sample according to the sample extraction word to obtain sample extraction information, including:

[0065] The sample word vectors corresponding to the disassembled sentences in the content disassembled sample are combined to obtain a sample combined vector, and a sentence sample semantic is determined according to the sample combined vector; wherein, a vector similarity between the sample combined vector and a preset semantic vector is calculated to obtain a third vector similarity, and a semantic information corresponding to the preset semantic vector with the maximum third vector similarity is determined as the sentence sample semantic.

[0066] Word embedding processing is performed on the sample extraction word to obtain extraction word embedding features, and vector conversion is performed on the extraction word embedding features to obtain extraction word vectors; wherein, through the word embedding processing and the vector conversion on the sample extraction word, semantic information of the sample extraction word can be effectively obtained.

[0067] A similarity between the extraction word vectors and the sample combined vector is calculated to obtain a second vector similarity, and a target matching vector in the sample combined vector is determined according to the second vector similarity; wherein, the sample combined vector corresponding to the maximum second vector similarity is determined as the target matching vector.

[0068] A semantic object vocabulary in the sentence sample semantic corresponding to the target matching vector is obtained, and the semantic object vocabulary is extracted to obtain the sample extraction information; wherein, the sentence sample semantic includes semantic object vocabulary and object description description vocabulary and other information.

[0069] Preferably, character recognition is performed on a policy document to obtain a policy text, including:

[0070] The policy file is subjected to grayscale processing to obtain a grayscale image, and the grayscale image is subjected to normalization processing to obtain a normalized image;

[0071] The normalized image is subjected to convolution processing according to different convolution scales to obtain convolution features, and the convolution features are subjected to feature fusion according to the convolution scales to obtain a feature pyramid; wherein the convolution scales can be set according to requirements, the number of the convolution scales is greater than or equal to two, the normalized image is subjected to convolution processing through different convolution scales to obtain convolution features of different feature scales, so that the feature pyramid after feature fusion can effectively have context feature information.

[0072] Text prediction is performed according to the feature pyramid to obtain a text prediction result, a target text box is determined according to the text prediction result, and the policy text is generated according to the target text box; wherein the feature pyramid features are decoded through an encoder in a large model, text prediction is performed based on the feature decoding result to obtain a text prediction result.

[0073] In this embodiment, determining the target text box according to the text prediction result comprises:

[0074] The text existence probability in the text prediction result is obtained, and the text existence probability is compared with a probability threshold; wherein the text existence probability is used to represent the probability of text existing at a corresponding position, and the probability threshold can be set according to requirements.

[0075] If the text existence probability is greater than the probability threshold, the text box corresponding to the text existence probability is determined as a candidate text box, and the overlapping degree between different candidate text boxes is calculated; wherein if the text existence probability is greater than the probability threshold, it is determined that text exists at the corresponding position, the text box corresponding to the text existence probability is determined as a candidate text box, and the overlapping area between different candidate text boxes is calculated to determine the overlapping degree based on the overlapping area.

[0076] The text box score of the candidate text box is determined according to the overlapping degree and the text existence probability, and the target text box is determined according to the text box score; wherein the text box score is obtained by weighting operation of the overlapping degree and the text existence probability, and the weighting coefficients of the overlapping degree and the text existence probability in the weighting operation process can be set according to requirements. In this step, the candidate text boxes are sorted according to the text box score, the box with a higher probability and a lower overlapping degree with other boxes is retained, the box with a higher overlapping degree is suppressed, and the target text box is obtained. Each text region only retains one optimal target text box.

[0077] Step S20, condition extraction is performed on the policy content according to the extraction prompt word, policy extraction information is obtained, and enterprise information to be matched and government-enterprise matching demand are acquired;

[0078] The policy extraction information includes declaration conditions and subsidy content, the government-enterprise matching demand can be set according to demand, the declaration conditions and the subsidy content in the policy content are automatically extracted by performing condition extraction on the policy content through the extraction prompt word.

[0079] Optionally, the enterprise information to be matched and the government-enterprise matching demand are acquired, and the method comprises the following steps.

[0080] The government-enterprise matching demand is obtained by acquiring a government-enterprise matching instruction and extracting demand information in the government-enterprise matching instruction, wherein the demand information is used to represent the demand of a user for a policy type, and the demand information is matched with a demand query table. The government-enterprise matching demand is obtained, and the demand query table stores a corresponding relationship between different demand information and corresponding government-enterprise matching demand.

[0081] The enterprise name and private domain enterprise information in the government-enterprise matching instruction are acquired, and public domain enterprise information is queried according to a preset API interface. The preset API interface can be set according to demand to acquire enterprise information corresponding to the enterprise name in network public data, and the public domain enterprise information is obtained.

[0082] The public domain enterprise information and the private domain enterprise information are combined to obtain the enterprise information to be matched. By combining the public domain enterprise information and the private domain enterprise information, the accuracy of the enterprise information to be matched is effectively improved.

[0083] Step S30, a matching prompt word is determined according to the government-enterprise matching demand, and a pre-trained large model is controlled to perform government-enterprise matching on the enterprise information to be matched and the policy extraction information according to the matching prompt word, and a government-enterprise matching result is obtained.

[0084] The matching prompt word is obtained by matching the government-enterprise matching demand with a prompt word query table, the prompt word query table stores a corresponding relationship between different government-enterprise matching demand and corresponding matching prompt word, and by determining the matching prompt word, the pre-trained large model can be effectively prompted to match the information of the demand type specified by the user between the enterprise information to be matched and the policy extraction information.

[0085] In this embodiment, the policy text is input into the pre-trained large model for policy splitting, and the pre-trained large model is used for government-enterprise matching of the enterprise information to be matched and the policy extraction information, without the need for manual government-enterprise matching. The conditional extraction of the policy content based on the prompt word can automatically extract the reporting conditions and the subsidy content in the policy content. The pre-trained large model can automatically perform government-enterprise matching on the enterprise information to be matched and the policy extraction information, without the need for manual policy analysis and government-enterprise matching, thereby improving the government-enterprise matching efficiency.

[0086] Embodiment Two

[0087] Please refer to Figure 2 is a structural schematic diagram of a government-enterprise matching system 100 based on a large model provided by the second embodiment of the present application, which comprises:

[0088] The policy splitting module 10 is configured to perform character recognition on a policy file to obtain a policy text, and input the policy text into a pre-trained large model for policy splitting to obtain policy content.

[0089] Optionally, the policy splitting module 10 is further configured to: obtain a policy sample, a sample prompt word and a sample extraction word, and input the policy sample and the sample prompt word into the large model for sample splitting to obtain a content splitting sample;

[0090] According to the sample extraction word, the content splitting sample is extracted to obtain sample extraction information, and sample enterprise information and sample matching words are obtained;

[0091] According to the sample matching word, the large model is controlled to perform government-enterprise matching on the sample enterprise information and the sample extraction information to obtain a sample matching result, and a model loss is determined according to the sample matching result, the sample extraction information and the content splitting sample;

[0092] According to the model loss, the parameters of the large model are updated until the large model converges, and the pre-trained large model is obtained.

[0093] Further, the policy splitting module 10 is further configured to: perform word segmentation on the policy sample according to the large model to obtain a sample word segmentation, and perform word embedding processing on the sample word segmentation and the sample prompt word to obtain sample word embedding features and prompt word embedding features;

[0094] The sample word embedding features and the prompt word embedding features are converted into vectors to obtain sample word vectors and prompt word vectors, and the similarity between the sample word vectors and the prompt word vectors is calculated to obtain a first vector similarity;

[0095] If the first vector similarity is greater than a similarity threshold, the first vector similarity corresponding to a sample sentence in which the sample word segmentation is located is determined as a disassembled sentence, and the disassembled sentences are combined to obtain the content disassembly sample.

[0096] Further, the policy splitting module 10 is further configured to: combine the sample word vectors corresponding to the disassembled sentences in the content disassembly sample to obtain a sample combined vector, and determine a sentence sample semantic according to the sample combined vector;

[0097] The sample extraction word is subjected to word embedding processing to obtain an extraction word embedding feature, and the extraction word embedding feature is subjected to vector conversion to obtain an extraction word vector;

[0098] The similarity between the extraction word vector and the sample combined vector is calculated to obtain a second vector similarity, and a target matching vector in the sample combined vector is determined according to the second vector similarity;

[0099] The semantic object vocabulary in the target matching vector corresponding to the sentence sample semantic is obtained, and the semantic object vocabulary is extracted to obtain the sample extraction information.

[0100] Preferably, the policy splitting module 10 is further configured to: perform grayscale processing on the policy file to obtain a grayscale image, and perform normalization processing on the grayscale image to obtain a normalized image;

[0101] According to different convolution scales, the normalized image is subjected to convolution processing to obtain convolution features, and the convolution features are subjected to feature fusion according to the convolution scales to obtain a feature pyramid;

[0102] Text prediction is performed according to the feature pyramid to obtain a text prediction result, a target text box is determined according to the text prediction result, and the policy text is generated according to the target text box.

[0103] In this embodiment, the policy splitting module 10 is further configured to: obtain a text existence probability in the text prediction result, and compare the text existence probability with a probability threshold;

[0104] If the text existence probability is greater than the probability threshold, the text box corresponding to the text existence probability is determined as a candidate text box, and the overlap degree between different candidate text boxes is calculated;

[0105] The text box score of the candidate text box is determined according to the overlap degree and the text existence probability, and the target text box is determined according to the text box score.

[0106] The conditional extraction module 11 is configured to perform conditional extraction on the policy content according to an extraction prompt word, to obtain policy extraction information, and to obtain enterprise information to be matched and a government-enterprise matching requirement, wherein the policy extraction information includes a declaration condition and an award content.

[0107] Optionally, the conditional extraction module 11 is further configured to obtain a government-enterprise matching instruction, and extract requirement information in the government-enterprise matching instruction to obtain the government-enterprise matching requirement.

[0108] The enterprise name and private domain enterprise information in the government-enterprise matching instruction are obtained, and public domain enterprise information is queried according to a preset API interface.

[0109] The public domain enterprise information and the private domain enterprise information are combined to obtain the enterprise information to be matched.

[0110] The government-enterprise matching module 12 is configured to determine a matching prompt word according to the government-enterprise matching requirement, and control the pre-trained large model to perform government-enterprise matching on the enterprise information to be matched and the policy extraction information according to the matching prompt word, to obtain a government-enterprise matching result.

[0111] Please refer to Figure 3 , the policy information extraction process includes:

[0112] 1) The system operation personnel input a policy file (in PDF format, one file often includes multiple award policies), and an extraction requirement (in text format, including a single policy identifier, a declaration condition, and other extraction conditions)

[0113] 2) The policy file is input into a character recognition module, and policy text content is output.

[0114] 3) The policy text content, the single policy identifier, and the disassembly prompt word framework are input into an LLM, and multiple award policy contents are output.

[0115] 4) The multiple award policy contents, the extraction conditions of the declaration condition and the award content, and the extraction prompt word framework are input into an LLM, and the award content and the declaration condition (structured information) of each award policy are output.

[0116] In this embodiment, the system further includes an extraction configuration module, a prompt word assembly module, a disassembly prompt word framework, and an extraction prompt word framework. The extraction configuration module is configured to output a single policy identifier and a field extraction requirement according to an input extraction requirement. The prompt word assembly module is configured to output a corresponding disassembly prompt word and an extraction prompt word according to an input single policy identifier and a field extraction requirement. The disassembly prompt word framework is configured to output a preset disassembly template. The extraction prompt word framework is configured to output a preset extraction template.

[0117] Please refer to Figure 4 , the government-enterprise information matching process includes:

[0118] 1) Enterprise personnel input private domain enterprise information (non-public enterprise information on the network, including enterprise qualifications, etc.);

[0119] 2) Pull public enterprise information through enterprise information query API, and combine with private enterprise information to form the final enterprise information;

[0120] 3) Input customized government-enterprise matching requirements, and combine with general government-enterprise matching requirements to form the final policy matching requirements;

[0121] 4) Enterprise information, award subsidy content and reporting conditions of each award subsidy policy, policy matching requirements, and matching prompt word framework are input into the large model, and the output is each policy matched by the enterprise and the award subsidy content corresponding to the policy.

[0122] In this embodiment, the large model extracts the award subsidy content and the reporting condition to replace manual extraction, which only needs to manually configure the extraction features once at the initial use (the extraction features are general and applicable to all policy files), so even if the policy file is updated, no more manual work is needed; Therefore, it can reduce the operation and maintenance cost of the government. Some public enterprise information is pulled by the external interface, and the rest of the private enterprise information is input by the enterprise, and the enterprise can input by means of specification documents and natural language, and the large model replaces manual sorting of each piece of enterprise information; Therefore, it can reduce the use cost of the enterprise. The large model fully understands the reporting conditions and enterprise information, and then gives the matching conclusion. Compared with the regular matching method, it improves the generalization ability and can cope with more matching scenarios such as semantic ambiguity and missing enterprise information; Therefore, it can solve the problem of low matching accuracy.

[0123] In this embodiment, by inputting the policy text into the pre-trained large model for policy splitting, the policy content in the policy text can be automatically extracted, the reporting conditions and the award subsidy content in the policy content can be automatically extracted by extracting the prompt words, and the pre-trained large model can automatically match the enterprise information and the policy extraction information, without using manual policy sorting and government-enterprise matching, which improves the government-enterprise matching efficiency.

[0124] Embodiment Three

[0125] Figure 5 is a structural block diagram of a terminal device 2 provided by the third embodiment of the present application. As shown in Figure 5As shown, the terminal device 2 of the embodiment includes a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, for example, a program of the government-enterprise matching method based on a large model. The processor 20 implements the steps in each embodiment of the above-described government-enterprise matching method based on a large model when executing the computer program 22.

[0126] For example, the computer program 22 can be divided into one or more modules stored in the memory 21 and executed by the processor 20 to complete the present application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device can include, but is not limited to, the processor 20 and the memory 21.

[0127] The processor 20 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0128] The memory 21 can be an internal storage unit of the terminal device 2, such as a hard disk or a memory of the terminal device 2. The memory 21 can also be an external storage device of the terminal device 2, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 21 can include both the internal storage unit and the external storage device of the terminal device 2. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 can also be used to temporarily store data that has been output or will be output.

[0129] In addition, each functional module in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0130] When the integrated module is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. The computer readable storage medium can be non-volatile or volatile. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms, etc. The computer readable storage medium can include any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content included in the computer readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable storage medium does not include electrical carrier signals and telecommunication signals.

[0131] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A government-enterprise matching method based on a large model, characterized in that, The method includes: The policy document is subjected to text recognition to obtain the policy text, and the policy text is then input into a pre-trained large model for policy decomposition to obtain the policy content. Based on the extraction prompts, the policy content is extracted to obtain policy extraction information, and the information of enterprises to be matched and the government-enterprise matching needs are obtained. The policy extraction information includes application conditions and subsidy content. Based on the government-enterprise matching requirements, matching prompt words are determined, and the pre-trained large model is controlled to perform government-enterprise matching on the enterprise information to be matched and the policy extraction information based on the matching prompt words, so as to obtain the government-enterprise matching result; Before inputting the policy text into the pre-trained large model for policy decomposition, the following steps are also included: Obtain policy samples, sample prompts, and sample extraction words, and input the policy samples and sample prompts into the large model to decompose the samples and obtain content decomposition samples; Based on the extracted words, the content is decomposed into samples for extraction, and sample extraction information is obtained, along with sample enterprise information and sample matching words. The model is controlled to perform government-enterprise matching on the sample enterprise information and the sample extraction information based on the sample matching words to obtain the sample matching results. The model loss is determined based on the sample matching results, the sample extraction information and the content decomposition of the samples. The parameters of the large model are updated based on the model loss until the large model converges, thus obtaining the pre-trained large model.

2. The government-enterprise matching method based on a large model as described in claim 1, characterized in that, The policy sample and the sample prompts are input into the large model for sample decomposition to obtain content decomposition samples, including: The policy sample is segmented into words according to the large model to obtain sample words, and word embedding processing is performed on the sample words and the sample prompt words to obtain sample word embedding features and prompt word embedding features; The sample word embedding features and the prompt word embedding features are vectorized to obtain sample word vectors and prompt word vectors, and the similarity between the sample word vectors and the prompt word vectors is calculated to obtain the first vector similarity. If the similarity of the first vector is greater than the similarity threshold, then the sample sentence in which the sample word segmentation is located corresponding to the first vector similarity is determined as the decomposition sentence, and the decomposition sentences are combined to obtain the content decomposition sample.

3. The government-enterprise matching method based on a large model as described in claim 2, characterized in that, Based on the extracted words, the content is decomposed into samples for extraction, resulting in sample extraction information, including: The sample word vectors corresponding to the decomposed statements in the content decomposition sample are combined to obtain a sample combination vector, and the semantics of the statement sample are determined based on the sample combination vector. The extracted words from the sample are processed by word embedding to obtain extracted word embedding features, and the extracted word embedding features are then transformed into vectors to obtain extracted word vectors. Calculate the similarity between the extracted word vector and the sample combination vector to obtain a second vector similarity, and determine the target matching vector in the sample combination vector based on the second vector similarity; Obtain the semantic object vocabulary of the target matching vector in the semantics of the corresponding statement sample, and extract the semantic object vocabulary to obtain the sample extraction information.

4. The government-enterprise matching method based on a large model as described in claim 1, characterized in that, The policy document is subjected to text recognition to obtain the policy text, including: The policy document is processed in grayscale to obtain a grayscale image, and the grayscale image is then normalized to obtain a normalized image. The normalized image is convolved according to different convolution scales to obtain convolution features, and the convolution features are fused according to the convolution scale to obtain a feature pyramid. Text prediction is performed based on the feature pyramid to obtain text prediction results. Target text boxes are determined based on the text prediction results, and the policy text is generated based on the target text boxes.

5. The government-enterprise matching method based on a large model as described in claim 4, characterized in that, Determining the target text box based on the text prediction results includes: Obtain the text existence probability in the text prediction result, and compare the text existence probability with a probability threshold; If the probability of the text being present is greater than the probability threshold, then the text box corresponding to the probability of the text being present is determined as a candidate text box, and the overlap between different candidate text boxes is calculated. The text box score of the candidate text box is determined based on the overlap and the probability of text presence, and the target text box is determined based on the text box score.

6. The government-enterprise matching method based on a large model as described in claim 1, characterized in that, Obtain information on companies to be matched and government-enterprise matching needs, including: Obtain the government-enterprise matching instruction and extract the demand information from the government-enterprise matching instruction to obtain the government-enterprise matching demand; Obtain the enterprise name and private domain enterprise information from the government-enterprise matching instruction, and query the public domain enterprise information according to the preset API interface; The public domain enterprise information and the private domain enterprise information are combined to obtain the enterprise information to be matched.

7. A government-enterprise matching system based on a large model, characterized in that, The system includes: The policy segmentation module is used to perform text recognition on policy documents to obtain policy text, and then input the policy text into a pre-trained large model to perform policy segmentation to obtain policy content. The condition extraction module is used to extract the policy content based on extraction prompts to obtain policy extraction information, and to obtain information on enterprises to be matched and government-enterprise matching needs. The policy extraction information includes application conditions and subsidy content. The government-enterprise matching module is used to determine matching prompt words according to the government-enterprise matching requirements, and control the pre-trained big model to perform government-enterprise matching on the enterprise information to be matched and the policy extraction information according to the matching prompt words, so as to obtain the government-enterprise matching result; The policy splitting module is also used for: Obtain policy samples, sample prompts, and sample extraction words, and input the policy samples and sample prompts into the large model to decompose the samples and obtain content decomposition samples; Based on the extracted words, the content is decomposed into samples for extraction, and sample extraction information is obtained, along with sample enterprise information and sample matching words. The model is controlled to perform government-enterprise matching on the sample enterprise information and the sample extraction information based on the sample matching words to obtain the sample matching results. The model loss is determined based on the sample matching results, the sample extraction information and the content decomposition of the samples. The parameters of the large model are updated based on the model loss until the large model converges, thus obtaining the pre-trained large model.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Policy document intelligent analysis and structuring method and system

    CN114021574A

  • Policy information service recommendation method and system based on deep learning, and storage medium

    CN118051607A