Large model-based government and enterprise matching method and system and storage medium
Through the government-enterprise matching method based on the big model, policy content and enterprise information are automatically extracted and matched, and the problem of inefficient government-enterprise matching in the existing technology is solved, and an efficient and automated government-enterprise matching process is achieved.
Patent Information
- Application Number
- CN202510205524.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The inefficiency of government-enterprise matching in the existing technology is mainly due to the dispersion and complexity of policy information, which makes it difficult for enterprises to accurately match policies that suit them.
The government-enterprise matching method based on the big model is adopted, and the policy content is automatically extracted by text recognition and big model splitting of policy documents; then the conditions are drawn based on the extraction prompt words to obtain the declaration conditions and reward and subsidy content; finally, based on the pre-trained big model, the government-enterprise matching information and policy extraction information are treated as the government-enterprise matching.
There is no need to manually sort out policies and enterprise information, which significantly improves the efficiency of government-enterprise matching, and can automatically extract policy content and match, reducing the problem of operation and maintenance costs and low matching accuracy.
Smart Images

Figure CN119988745A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a government-enterprise matching method, system and storage medium based on a large model. Background Art
[0002] In the current economic environment, various reward and subsidy policies issued by the government should be a strong support for enterprises in their pursuit of development. However, due to the dispersion and complexity of policy information, enterprises often find it difficult to accurately match policies that suit them, resulting in a significant reduction in the effectiveness of policy implementation. Therefore, the issue of government-enterprise matching has received more and more attention.
[0003] In the existing government-enterprise matching process, policy information and enterprise information are generally sorted out manually, and government-enterprise matching is performed based on the information sorting results, which leads to low efficiency in government-enterprise matching. Summary of the invention
[0004] The purpose of the embodiments of the present invention is to provide a government-enterprise matching method, system and storage medium based on a large model to solve the problem of low efficiency of government-enterprise matching in the prior art.
[0005] The embodiment of the present invention is implemented as follows: a government-enterprise matching method based on a large model, the method comprising:
[0006] Perform text recognition on the policy document to obtain the policy text, and input the policy text into the pre-trained large model to perform policy splitting to obtain the policy content;
[0007] Conditionally extract the policy content according to the extraction prompt words to obtain policy extraction information, and obtain the enterprise information to be matched and the government-enterprise matching requirements. The policy extraction information includes application conditions and reward and subsidy content;
[0008] According to the government-enterprise matching requirements, matching prompt words are determined, and according to the matching prompt words, the pre-trained large model is controlled to perform government-enterprise matching on the enterprise information to be matched and the policy extraction information to obtain a government-enterprise matching result.
[0009] Preferably, before inputting the policy text into the pre-trained large model for policy splitting, the method further includes:
[0010] Obtaining a policy sample, a sample prompt word, and a sample extraction word, and inputting the policy sample and the sample prompt word into the large model for sample disassembly to obtain a content disassembly sample;
[0011] Extracting samples of the content disassembly sample according to the sample extraction words to obtain sample extraction information, and acquiring sample enterprise information and sample matching words;
[0012] Controlling the large model to match the sample enterprise information and the sample extraction information with each other according to the sample matching words, obtaining a sample matching result, and determining the model loss according to the sample matching result, the sample extraction information and the content disassembly sample;
[0013] The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
[0014] Preferably, the policy sample and the sample prompt word are input into the large model to perform sample decomposition to obtain a content decomposition sample, including:
[0015] Segmenting the policy sample according to the large model to obtain sample segmentations, and performing word embedding processing on the sample segmentations and the sample prompt words to obtain sample word embedding features and prompt word embedding features;
[0016] Performing vector conversion on the sample word embedding feature and the prompt word embedding feature to obtain a sample word vector and a prompt word vector, and calculating the similarity between the sample word vector and the prompt word vector to obtain a first vector similarity;
[0017] If the first vector similarity is greater than a similarity threshold, the sample sentence where the first vector similarity corresponds to the sample word segment is determined as a disassembled sentence, and the disassembled sentences are combined to obtain the content disassembled sample.
[0018] Preferably, the content decomposition sample is sampled according to the sample extraction word to obtain sample extraction information, including:
[0019] Combining the sample word vectors corresponding to the disassembled sentences in the content disassembled sample to obtain a sample combination vector, and determining the sentence sample semantics according to the sample combination vector;
[0020] Performing word embedding processing on the sample extracted words to obtain extracted word embedding features, and performing vector conversion on the extracted word embedding features to obtain extracted word vectors;
[0021] Calculating the similarity between the extracted word vector and the sample combination vector to obtain a second vector similarity, and determining a target matching vector in the sample combination vector according to the second vector similarity;
[0022] The semantic object vocabulary of the target matching vector in the semantics of the sentence sample is obtained, and the semantic object vocabulary is extracted to obtain the sample extraction information.
[0023] Preferably, the policy document is subjected to text recognition to obtain the policy text, including:
[0024] Performing grayscale processing on the policy document to obtain a grayscale image, and performing normalization processing on the grayscale image to obtain a normalized image;
[0025] According to different convolution scales, the normalized image is convolved to obtain convolution features, and the convolution features are fused according to the convolution scale to obtain a feature pyramid;
[0026] Perform text prediction according to the feature pyramid to obtain a text prediction result, determine a target text box according to the text prediction result, and generate the policy text according to the target text box.
[0027] Preferably, determining a target text box according to the text prediction result includes:
[0028] Obtaining a text existence probability in the text prediction result, and comparing the text existence probability with a probability threshold;
[0029] If the text existence probability is greater than the probability threshold, determining the text box corresponding to the text existence probability as a candidate text box, and calculating the overlap between different candidate text boxes;
[0030] A text box score of the candidate text box is determined according to the overlap degree and the text existence probability, and the target text box is determined according to the text box score.
[0031] Preferably, obtain the information of the enterprise to be matched and the government-enterprise matching requirements, including:
[0032] Obtaining a government-enterprise matching instruction, and extracting demand information in the government-enterprise matching instruction to obtain a government-enterprise matching demand;
[0033] Obtain the enterprise name and private domain enterprise information in the government-enterprise matching instruction, and query the public domain enterprise information according to the preset API interface;
[0034] The public domain enterprise information and the private domain enterprise information are combined to obtain the enterprise information to be matched.
[0035] Another object of an embodiment of the present invention is to provide a government-enterprise matching system based on a large model, the system comprising:
[0036] A policy splitting module is used to perform text recognition on policy documents to obtain policy texts, and input the policy texts into a pre-trained large model to perform policy splitting to obtain policy contents;
[0037] A conditional extraction module is used to conditionally extract the policy content according to the extraction prompt words, obtain policy extraction information, and obtain the enterprise information to be matched and the government-enterprise matching requirements. The policy extraction information includes application conditions and reward and subsidy content;
[0038] The government-enterprise matching module is used to determine the matching prompt words according to the government-enterprise matching requirements, and to control the pre-trained large model to perform government-enterprise matching on the enterprise information to be matched and the policy extraction information according to the matching prompt words to obtain the government-enterprise matching results.
[0039] Preferably, the policy splitting module is also used for:
[0040] Obtaining a policy sample, a sample prompt word, and a sample extraction word, and inputting the policy sample and the sample prompt word into the large model for sample disassembly to obtain a content disassembly sample;
[0041] Extracting samples of the content disassembly sample according to the sample extraction words to obtain sample extraction information, and acquiring sample enterprise information and sample matching words;
[0042] Controlling the large model to match the sample enterprise information and the sample extraction information with each other according to the sample matching words, obtaining a sample matching result, and determining the model loss according to the sample matching result, the sample extraction information and the content disassembly sample;
[0043] The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
[0044] The embodiments of the present invention can automatically extract policy contents from policy texts by inputting policy texts into a pre-trained large model for policy splitting, and can automatically extract application conditions and reward and subsidy contents from policy contents by extracting prompt words. Based on the pre-trained large model, the government-enterprise matching can be automatically performed on the enterprise information to be matched and the policy extraction information, without the need for manual policy sorting and government-enterprise matching, thereby improving the efficiency of government-enterprise matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flow chart of a government-enterprise matching method based on a large model provided in the first embodiment of the present invention;
[0046] Figure 2 It is a structural diagram of a government-enterprise matching system based on a large model provided in the second embodiment of the present invention;
[0047] Figure 3 It is a schematic diagram of extracting policy information by a government-enterprise matching system based on a large model provided by the second embodiment of the present invention;
[0048] Figure 4 It is a schematic diagram of government-enterprise information matching by a government-enterprise matching system based on a large model provided in a second embodiment of the present invention;
[0049] Figure 5 It is a schematic diagram of the structure of a terminal device provided in the third embodiment of the present invention. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0051] In order to illustrate the technical solution of the present invention, a specific embodiment is provided below for illustration.
[0052] Embodiment 1
[0053] See also Figure 1 , is a flow chart of a government-enterprise matching method based on a large model provided in the first embodiment of the present invention. The government-enterprise matching method based on a large model can be applied to any device or system. The government-enterprise matching method based on a large model includes the following steps:
[0054] Step S10, performing text recognition on the policy document to obtain the policy text, and inputting the policy text into the pre-trained large model to perform policy splitting to obtain the policy content;
[0055] Among them, policy documents are policy information such as subsidies or preferential treatment for enterprises issued by local governments. By performing text recognition on policy documents, policy documents in image format can be effectively converted into corresponding text information. In this step, by inputting the policy text into the pre-trained large model for policy splitting, the policy content in the policy text can be automatically extracted without the need for users to manually extract text information.
[0056] Optionally, before inputting the policy text into the pre-trained large model for policy splitting, the method further includes:
[0057] Obtain policy samples, sample prompt words and sample extracted words, and input the policy samples and the sample prompt words into the big model for sample disassembly to obtain content disassembly samples; wherein, the policy samples, sample prompt words and sample extracted words can be set according to needs. In this step, the big model performs word segmentation on the policy samples, sample prompt words and sample extracted words to obtain word units, and performs vector conversion on the word units to obtain word unit vectors. Each word unit vector contains semantic information, and the vector relationship between word units reflects the semantic association. The big model learns the semantic association relationship through pre-training. Based on the language knowledge and pattern recognition ability obtained through pre-training, the big model analyzes the vector sequence corresponding to the policy sample, and captures the semantic connection between the sample prompt words and the policy sample through a multi-layer neural network. The big model deeply analyzes the grammatical structure and semantic role of the policy sample, determines the logical relationship between each part, and based on the understanding of the sample prompt words, the big model searches for matching key information in the semantic representation of the policy sample to obtain the content disassembly sample;
[0058] The content disassembly sample is sampled according to the sample extraction word to obtain sample extraction information, and sample enterprise information and sample matching words are obtained; wherein the large model analyzes the semantic association between the sample extraction word and the content disassembly sample based on the language knowledge and semantic understanding ability learned in pre-training, and in the semantic vector representation of the content disassembly sample, finds the part with a high matching degree with the sample extraction word vector to obtain the sample extraction information;
[0059] According to the sample matching words, the big model is controlled to perform government-enterprise matching on the sample enterprise information and the sample extraction information to obtain a sample matching result; the model loss is determined according to the sample matching result, the sample extraction information and the content disassembly sample, and the parameters of the big model are updated according to the model loss until the big model converges to obtain the pre-trained big model; wherein the big model determines matching parameters based on the sample matching words, obtains matching information in the sample enterprise information and the sample extraction information based on the matching parameters, obtains enterprise matching information and extraction matching information, and can obtain the sample matching result by comparing the enterprise matching information and the extraction matching information.
[0060] Furthermore, the policy sample and the sample prompt word are input into the large model to perform sample decomposition to obtain a content decomposition sample, including:
[0061] Segmenting the policy sample according to the large model to obtain sample segmentations, and embedding the sample segmentations and the sample prompt words to obtain sample word embedding features and prompt word embedding features; wherein, by embedding the sample segmentations and the sample prompt words, the semantic representations of the sample segmentations and the sample prompt words can be effectively extracted;
[0062] Performing vector conversion on the sample word embedding feature and the prompt word embedding feature to obtain a sample word vector and a prompt word vector, and calculating the similarity between the sample word vector and the prompt word vector to obtain a first vector similarity; wherein the distance between the sample word vector and the prompt word vector can be calculated using the Euclidean distance formula to obtain the first vector similarity;
[0063] If the first vector similarity is greater than the similarity threshold, the sample sentence where the sample participle corresponding to the first vector similarity is located is determined to be a disassembled sentence, and the disassembled sentences are combined to obtain the content disassembly sample; wherein, the similarity threshold can be set as needed, and if the first vector similarity is greater than the similarity threshold, it is determined that the sample sentence where the sample participle corresponding to the first vector similarity is located contains policy information, and therefore, the sample sentence where the sample participle corresponding to the first vector similarity is located is determined to be a disassembled sentence.
[0064] Furthermore, the content decomposition sample is sampled according to the sample extraction word to obtain sample extraction information, including:
[0065] The sample word vectors corresponding to the disassembled sentences in the content disassembled sample are combined to obtain a sample combination vector, and the sentence sample semantics is determined according to the sample combination vector; wherein the vector similarity between the sample combination vector and the preset semantic vector is calculated to obtain a third vector similarity, and the semantic information of the preset semantic vector corresponding to the maximum third vector similarity is determined as the sentence sample semantics;
[0066] Performing word embedding processing on the sample extracted words to obtain extracted word embedding features, and performing vector conversion on the extracted word embedding features to obtain extracted word vectors; wherein, by performing word embedding processing and vector conversion on the sample extracted words, semantic information of the sample extracted words can be effectively obtained;
[0067] Calculating the similarity between the extracted word vector and the sample combination vector to obtain a second vector similarity, and determining a target matching vector in the sample combination vector according to the second vector similarity; wherein the sample combination vector corresponding to the maximum second vector similarity is determined as the target matching vector;
[0068] The semantic object vocabulary of the target matching vector in the corresponding sentence sample semantics is obtained, and the semantic object vocabulary is extracted to obtain the sample extraction information; wherein the sentence sample semantics includes information such as semantic object vocabulary and object description vocabulary.
[0069] Preferably, the policy document is subjected to text recognition to obtain the policy text, including:
[0070] Performing grayscale processing on the policy document to obtain a grayscale image, and performing normalization processing on the grayscale image to obtain a normalized image;
[0071] According to different convolution scales, the normalized image is convolved respectively to obtain convolution features, and the convolution features are feature fused according to the convolution scale to obtain a feature pyramid; wherein the convolution scale can be set according to demand, and the number of convolution scales is greater than or equal to two, and the normalized image is convolved respectively by different convolution scales to obtain convolution features of different feature scales, so that the feature pyramid after feature fusion can effectively have context feature information;
[0072] Text prediction is performed according to the feature pyramid to obtain a text prediction result, a target text box is determined according to the text prediction result, and the policy text is generated according to the target text box; wherein, the feature pyramid features are decoded by an encoder in the large model, and text prediction is performed based on the feature decoding result to obtain a text prediction result.
[0073] In this embodiment, determining a target text box according to the text prediction result includes:
[0074] Obtaining the text existence probability in the text prediction result, and comparing the text existence probability with a probability threshold; wherein the text existence probability is used to indicate the probability of text existing at the corresponding position, and the probability threshold can be set according to demand;
[0075] If the text existence probability is greater than the probability threshold, the text box corresponding to the text existence probability is determined as a candidate text box, and the overlap between different candidate text boxes is calculated; wherein, if the text existence probability is greater than the probability threshold, it is determined that the corresponding position exists text, the text box corresponding to the text existence probability is determined as a candidate text box, and the overlap area between different candidate text boxes is calculated, and the overlap degree is determined based on the overlap area;
[0076] The text box score of the candidate text box is determined according to the overlap and the text existence probability, and the target text box is determined according to the text box score; wherein, the text box score is obtained by performing a weighted operation on the overlap and the text existence probability, and the weighted coefficients of the overlap and the text existence probability in the weighted operation process can be set according to demand. In this step, the candidate text boxes are sorted according to the text box scores, and boxes with higher probabilities and lower overlap with other boxes are retained, and boxes with higher overlap are suppressed to obtain target text boxes, and only one optimal target text box is retained for each text area.
[0077] Step S20, conditionally extracting the policy content according to the extraction prompt word, obtaining policy extraction information, and obtaining the enterprise information to be matched and the government-enterprise matching requirements;
[0078] Among them, policy extraction information includes application conditions and reward and subsidy content. The government-enterprise matching needs can be set according to needs. By extracting prompt words to perform conditional extraction of policy content, the application conditions and reward and subsidy content in the policy content can be automatically extracted.
[0079] Optionally, obtain the information of the enterprise to be matched and the government-enterprise matching requirements, including:
[0080] Obtain the government-enterprise matching instruction, and extract the demand information in the government-enterprise matching instruction to obtain the government-enterprise matching demand; wherein the demand information is used to represent the user's demand for the policy type, and the demand information is matched with the demand query table. Obtain the government-enterprise matching demand, and the demand query table stores the corresponding relationship between different demand information and the corresponding government-enterprise matching demand;
[0081] Obtain the enterprise name and private domain enterprise information in the government-enterprise matching instruction, and query the public domain enterprise information according to the preset API interface; wherein the preset API interface can be set according to demand to obtain the enterprise information corresponding to the enterprise name in the data publicly available on the network, and obtain the public domain enterprise information;
[0082] The public domain enterprise information and the private domain enterprise information are combined to obtain the enterprise information to be matched; wherein, by combining the public domain enterprise information and the private domain enterprise information, the accuracy of the enterprise information to be matched is effectively improved.
[0083] Step S30, determining a matching prompt word according to the government-enterprise matching requirement, and controlling the pre-trained large model to perform government-enterprise matching on the to-be-matched enterprise information and the policy extraction information according to the matching prompt word, to obtain a government-enterprise matching result;
[0084] Among them, the government-enterprise matching needs are matched with the prompt word query table to obtain matching prompt words. The prompt word query table stores the correspondence between different government-enterprise matching needs and corresponding matching prompt words. By determining the matching prompt words, the pre-trained large model can be effectively prompted to match the user-specified demand type information between the enterprise information to be matched and the policy extraction information.
[0085] In this embodiment, the policy text is input into the pre-trained big model for policy splitting, and the matching enterprise information and policy extraction information are matched based on the pre-trained big model, without the need for manual matching. By extracting prompt words to conditionally extract the policy content, the application conditions and reward and subsidy content in the policy content can be automatically extracted. Based on the pre-trained big model, the matching enterprise information and policy extraction information can be automatically matched based on the pre-trained big model, without the need for manual policy sorting and government-enterprise matching, which improves the efficiency of government-enterprise matching.
[0086] Embodiment 2
[0087] See also Figure 2 , is a schematic diagram of the structure of a government-enterprise matching system 100 based on a large model provided in a second embodiment of the present invention, including:
[0088] The policy splitting module 10 is used to perform text recognition on the policy document to obtain the policy text, and input the policy text into the pre-trained large model to perform policy splitting to obtain the policy content.
[0089] Optionally, the policy splitting module 10 is further used to: obtain policy samples, sample prompt words and sample extraction words, and input the policy samples and the sample prompt words into the large model for sample splitting to obtain content splitting samples;
[0090] Extracting samples of the content disassembly sample according to the sample extraction words to obtain sample extraction information, and acquiring sample enterprise information and sample matching words;
[0091] Controlling the large model to match the sample enterprise information and the sample extraction information with each other according to the sample matching words, obtaining a sample matching result, and determining the model loss according to the sample matching result, the sample extraction information and the content disassembly sample;
[0092] The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
[0093] Furthermore, the policy splitting module 10 is further used to: segment the policy sample according to the large model to obtain sample segmentation, and perform word embedding processing on the sample segmentation and the sample prompt word to obtain sample word embedding features and prompt word embedding features;
[0094] Performing vector conversion on the sample word embedding feature and the prompt word embedding feature to obtain a sample word vector and a prompt word vector, and calculating the similarity between the sample word vector and the prompt word vector to obtain a first vector similarity;
[0095] If the first vector similarity is greater than a similarity threshold, the sample sentence where the first vector similarity corresponds to the sample word segment is determined as a disassembled sentence, and the disassembled sentences are combined to obtain the content disassembled sample.
[0096] Furthermore, the policy splitting module 10 is further used to: combine the sample word vectors corresponding to the disassembled sentences in the content disassembled sample to obtain a sample combination vector, and determine the sentence sample semantics according to the sample combination vector;
[0097] Performing word embedding processing on the sample extracted words to obtain extracted word embedding features, and performing vector conversion on the extracted word embedding features to obtain extracted word vectors;
[0098] Calculating the similarity between the extracted word vector and the sample combination vector to obtain a second vector similarity, and determining a target matching vector in the sample combination vector according to the second vector similarity;
[0099] The semantic object vocabulary of the target matching vector in the semantics of the sentence sample is obtained, and the semantic object vocabulary is extracted to obtain the sample extraction information.
[0100] Preferably, the policy splitting module 10 is further used to: perform grayscale processing on the policy document to obtain a grayscale image, and perform normalization processing on the grayscale image to obtain a normalized image;
[0101] According to different convolution scales, the normalized image is convolved to obtain convolution features, and the convolution features are fused according to the convolution scale to obtain a feature pyramid;
[0102] Perform text prediction according to the feature pyramid to obtain a text prediction result, determine a target text box according to the text prediction result, and generate the policy text according to the target text box.
[0103] In this embodiment, the policy splitting module 10 is further used to: obtain the text existence probability in the text prediction result, and compare the text existence probability with the probability threshold;
[0104] If the text existence probability is greater than the probability threshold, determining the text box corresponding to the text existence probability as a candidate text box, and calculating the overlap between different candidate text boxes;
[0105] A text box score of the candidate text box is determined according to the overlap degree and the text existence probability, and the target text box is determined according to the text box score.
[0106] The conditional extraction module 11 is used to perform conditional extraction on the policy content according to the extraction prompt words, obtain policy extraction information, and obtain the enterprise information to be matched and the government-enterprise matching requirements. The policy extraction information includes application conditions and reward and subsidy content.
[0107] Optionally, the condition extraction module 11 is further used to: obtain a government-enterprise matching instruction, and extract demand information in the government-enterprise matching instruction to obtain a government-enterprise matching demand;
[0108] Obtain the enterprise name and private domain enterprise information in the government-enterprise matching instruction, and query the public domain enterprise information according to the preset API interface;
[0109] The public domain enterprise information and the private domain enterprise information are combined to obtain the enterprise information to be matched.
[0110] The government-enterprise matching module 12 is used to determine matching prompt words according to the government-enterprise matching requirements, and to control the pre-trained large model to perform government-enterprise matching on the enterprise information to be matched and the policy extraction information according to the matching prompt words to obtain government-enterprise matching results.
[0111] See also Figure 3 ,The policy information extraction process includes:
[0112] 1) System operation and maintenance personnel input policy documents (PDF format, one document often includes multiple reward and subsidy policies) and extraction requirements (text format, including single policy identification, application conditions and other extraction conditions)
[0113] 2) Policy documents are input into the text recognition module, which outputs the policy text content;
[0114] 3) Input the policy text content, single policy identifier, and disassembled prompt word framework into LLM, and output it into multiple reward and subsidy policy contents;
[0115] 4) Input multiple reward and subsidy policy contents, application conditions, extraction conditions of reward and subsidy contents, and extraction prompt word framework into LLM, and output the reward and subsidy contents and application conditions (structured information) of each reward and subsidy policy.
[0116] In this embodiment, the system is also provided with an extraction configuration module, a prompt word assembly module, a split prompt word framework and an extraction prompt word framework. The extraction configuration module is used to output a single policy identifier and a field extraction requirement according to the input extraction requirements. The prompt group assembly module is used to output corresponding disassembly prompt words and extraction prompt words according to the input single policy identifier and field extraction requirements. The split prompt word framework is used to output a preset split template. The extraction prompt word framework is used to output a preset extraction template.
[0117] See also Figure 4 , the government-enterprise information matching process includes:
[0118] 1) Enterprise personnel input private enterprise information (non-public enterprise information on the Internet, including enterprise qualifications, etc.);
[0119] 2) Pull public domain enterprise information through the enterprise information query API and combine it with private domain enterprise information to form the final enterprise information;
[0120] 3) Input customized government-enterprise matching requirements and combine them with general government-enterprise matching requirements to form the final policy matching requirements;
[0121] 4) Enterprise information, reward and subsidy content and application conditions of various reward and subsidy policies, policy matching requirements, and matching prompt word framework are input into the large model, and the various policies matched by the enterprise and the reward and subsidy content corresponding to this policy are output.
[0122] In this embodiment, the large model extracts the reward and subsidy content and application conditions instead of manual extraction. It is only necessary to manually configure the extraction features once when using it for the first time (the extraction features are universal and applicable to all policy documents). Even if the policy documents are updated, there is no need to invest manpower; therefore, the government's operation and maintenance costs can be reduced. Part of the public domain information of the enterprise is pulled by the external interface, and the remaining private domain enterprise information is input by the enterprise, and the enterprise is supported to input through description documents, natural language, etc. The large model replaces manual sorting of each piece of enterprise information; therefore, the cost of use of the enterprise can be reduced. The large model fully understands the application conditions and enterprise information, and then gives a matching conclusion. Compared with the regular matching method, its generalization ability is improved, and it can cope with more matching scenarios, such as semantic ambiguity, missing enterprise information, etc.; therefore, the problem of low matching accuracy can be solved.
[0123] In this embodiment, by inputting the policy text into the pre-trained big model for policy splitting, the policy content in the policy text can be automatically extracted, and the policy content can be conditionally extracted by extracting prompt words, and the application conditions and reward and subsidy content in the policy content can be automatically extracted. Based on the pre-trained big model, the government-enterprise matching can be automatically performed between the matching enterprise information and the policy extraction information, without the need for manual policy sorting and government-enterprise matching, thereby improving the efficiency of government-enterprise matching.
[0124] Embodiment 3
[0125] Figure 5 2 is a block diagram of a terminal device 2 provided in the third embodiment of the present application. Figure 5As shown, the terminal device 2 of this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program of a government-enterprise matching method based on a large model. When the processor 20 executes the computer program 22, the steps in each embodiment of the government-enterprise matching method based on a large model are implemented.
[0126] Exemplarily, the computer program 22 may be divided into one or more modules, which are stored in the memory 21 and executed by the processor 20 to complete the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, a processor 20 and a memory 21.
[0127] The processor 20 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0128] The memory 21 may be an internal storage unit of the terminal device 2, such as a hard disk or memory of the terminal device 2. The memory 21 may also be an external storage device of the terminal device 2, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 2. Further, the memory 21 may also include both an internal storage unit and an external storage device of the terminal device 2. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 may also be used to temporarily store data that has been output or is to be output.
[0129] In addition, each functional module in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional unit.
[0130] If the integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in computer-readable storage media can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable storage media do not include electrical carrier signals and telecommunication signals.
[0131] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A government-enterprise matching method based on a large model, characterized in that: The method comprises: Perform text recognition on the policy document to obtain the policy text, and input the policy text into the pre-trained large model to perform policy splitting to obtain the policy content; Conditionally extract the policy content according to the extraction prompt words to obtain policy extraction information, and obtain the enterprise information to be matched and the government-enterprise matching requirements. The policy extraction information includes application conditions and reward and subsidy content; According to the government-enterprise matching requirements, matching prompt words are determined, and according to the matching prompt words, the pre-trained large model is controlled to perform government-enterprise matching on the enterprise information to be matched and the policy extraction information to obtain a government-enterprise matching result.
2. The government-enterprise matching method based on a large model as claimed in claim 1, characterized in that: Before inputting the policy text into the pre-trained large model for policy splitting, it also includes: Obtaining a policy sample, a sample prompt word, and a sample extraction word, and inputting the policy sample and the sample prompt word into the large model for sample disassembly to obtain a content disassembly sample; Extracting samples of the content disassembly sample according to the sample extraction words to obtain sample extraction information, and acquiring sample enterprise information and sample matching words; Controlling the large model to match the sample enterprise information and the sample extraction information with each other according to the sample matching words, obtaining a sample matching result, and determining the model loss according to the sample matching result, the sample extraction information and the content disassembly sample; The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
3. The government-enterprise matching method based on a large model as claimed in claim 2, characterized in that: Input the policy sample and the sample prompt word into the large model to perform sample disassembly to obtain a content disassembly sample, including: Segmenting the policy sample according to the large model to obtain sample segmentations, and performing word embedding processing on the sample segmentations and the sample prompt words to obtain sample word embedding features and prompt word embedding features; Performing vector conversion on the sample word embedding feature and the prompt word embedding feature to obtain a sample word vector and a prompt word vector, and calculating the similarity between the sample word vector and the prompt word vector to obtain a first vector similarity; If the first vector similarity is greater than a similarity threshold, the sample sentence where the first vector similarity corresponds to the sample word segment is determined as a disassembled sentence, and the disassembled sentences are combined to obtain the content disassembled sample.
4. The government-enterprise matching method based on a large model as claimed in claim 3, characterized in that: The content disassembly sample is sampled according to the sample extraction word to obtain sample extraction information, including: Combining the sample word vectors corresponding to the disassembled sentences in the content disassembled sample to obtain a sample combination vector, and determining the sentence sample semantics according to the sample combination vector; Performing word embedding processing on the sample extracted words to obtain extracted word embedding features, and performing vector conversion on the extracted word embedding features to obtain extracted word vectors; Calculating the similarity between the extracted word vector and the sample combination vector to obtain a second vector similarity, and determining a target matching vector in the sample combination vector according to the second vector similarity; The semantic object vocabulary of the target matching vector in the semantics of the sentence sample is obtained, and the semantic object vocabulary is extracted to obtain the sample extraction information.
5. The government-enterprise matching method based on a large model as claimed in claim 1, characterized in that: Perform text recognition on policy documents to obtain policy texts, including: Performing grayscale processing on the policy document to obtain a grayscale image, and performing normalization processing on the grayscale image to obtain a normalized image; According to different convolution scales, the normalized image is convolved to obtain convolution features, and the convolution features are fused according to the convolution scale to obtain a feature pyramid; Perform text prediction according to the feature pyramid to obtain a text prediction result, determine a target text box according to the text prediction result, and generate the policy text according to the target text box.
6. The government-enterprise matching method based on a large model as claimed in claim 5, characterized in that: Determining a target text box according to the text prediction result includes: Obtaining a text existence probability in the text prediction result, and comparing the text existence probability with a probability threshold; If the text existence probability is greater than the probability threshold, determining the text box corresponding to the text existence probability as a candidate text box, and calculating the overlap between different candidate text boxes; A text box score of the candidate text box is determined according to the overlap degree and the text existence probability, and the target text box is determined according to the text box score.
7. The government-enterprise matching method based on a large model as claimed in claim 1, characterized in that: Obtain information about the companies to be matched and government-enterprise matching requirements, including: Obtaining a government-enterprise matching instruction, and extracting demand information in the government-enterprise matching instruction to obtain a government-enterprise matching demand; Obtain the enterprise name and private domain enterprise information in the government-enterprise matching instruction, and query the public domain enterprise information according to the preset API interface; The public domain enterprise information and the private domain enterprise information are combined to obtain the enterprise information to be matched.
8. A government-enterprise matching system based on a large model, characterized in that: The system comprises: A policy splitting module is used to perform text recognition on policy documents to obtain policy texts, and input the policy texts into a pre-trained large model to perform policy splitting to obtain policy contents; A conditional extraction module is used to conditionally extract the policy content according to the extraction prompt words, obtain policy extraction information, and obtain the enterprise information to be matched and the government-enterprise matching requirements. The policy extraction information includes application conditions and reward and subsidy content; The government-enterprise matching module is used to determine the matching prompt words according to the government-enterprise matching requirements, and to control the pre-trained large model to perform government-enterprise matching on the enterprise information to be matched and the policy extraction information according to the matching prompt words to obtain the government-enterprise matching results.
9. The government-enterprise matching system based on a large model as claimed in claim 8, characterized in that: The policy splitting module is also used to: Obtaining a policy sample, a sample prompt word, and a sample extraction word, and inputting the policy sample and the sample prompt word into the large model for sample disassembly to obtain a content disassembly sample; Extracting samples of the content disassembly sample according to the sample extraction words to obtain sample extraction information, and acquiring sample enterprise information and sample matching words; Controlling the large model to match the sample enterprise information and the sample extraction information with each other according to the sample matching words, obtaining a sample matching result, and determining the model loss according to the sample matching result, the sample extraction information and the content disassembly sample; The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Policy key information extraction method and device, storage medium and electronic equipment
CN112035653A
Policy analysis system and method based on intelligent semantic recognition
CN112036841A
Policy matching method, device and apparatus based on enterprise portrait, and medium
CN113723737A
Policy matching method, device and system, electronic equipment and readable storage medium
CN113870083A
Policy document intelligent analysis and structuring method and system
CN114021574A