Loan project identification method and related device facing multiple financial classification standard catalogs
Patent Information
- Application Number
- CN202610861279.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-11
AI Technical Summary
现有技术难以提高其识别精度
[0017] This invention employs the above technical solution, pre-establishing a multi-directory comprehensive knowledge base, which contains knowledge sub-bases corresponding to multiple financial classification standard directories. It also pre-establishes prompt instruction sets corresponding to these multiple financial classification standard directories. Subsequently, upon receiving the text information of the target loan project and the target directory identifier, the corresponding target knowledge sub-base and target prompt instruction set are located based on the target directory identifier. A candidate knowledge set is obtained by retrieving the target knowledge sub-base. The text information, candidate knowledge set, and target prompt instruction set are then input into a pre-set large language model to obtain the recognition result output by the large language model. Based on this, by pre-establishing knowledge sub-bases and prompt instruction sets corresponding to multiple financial classification standard directories, and subsequently locating the corresponding target knowledge sub-base and selecting the corresponding target prompt instruction set based on the target directory identifier, this invention enables unified support for multiple financial classification standard directories through a single large language model. This significantly reduces the cost for financial institutions to repeatedly develop and maintain multiple independent systems to meet different regulatory requirements. Furthermore, by using prompt instruction sets to guide the large language model in generating recognition results, this invention improves recognition accuracy. Finally, when a new financial classification standard catalog is needed, a new knowledge sub-base and prompt instruction set are constructed according to the established structure, without the need to repeatedly develop a new system, thus giving the invention good scalability.
Smart Images

Figure CN122736756A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of financial technology and artificial intelligence, specifically to a method and related apparatus for identifying loan items in multiple sets of financial classification standard catalogs. Background Technology
[0002] With the deepening of policies such as green finance, financial institutions often need to identify the attributes of loan projects based on multiple sets of financial classification standards, such as the "Green Finance Supported Project Catalog", the "Special Statistical System for Agricultural Loans", and the "Green Finance Supported Biodiversity Conservation Project Catalog".
[0003] Existing technologies are typically designed for single financial classification standards. When it is necessary to identify loan attributes based on multiple sets of financial classification standards, an identification system needs to be independently developed, deployed, and maintained for each set, resulting in high costs. Secondly, standards like the "Special Statistical System for Agricultural Loans" contain a large number of non-literal, strict logical rules and judgment conditions, such as the composite definition of "farmer" status and the geographical exclusion method for "rural areas." Existing technologies struggle to improve their identification accuracy. Finally, existing technologies lack flexible scalability and cannot quickly adapt to new financial classification standards.
[0004] Therefore, in loan project identification scenarios facing multiple sets of financial classification standard catalogs, how to reduce identification costs, improve identification accuracy, and have good scalability has become an urgent technical problem to be solved. Summary of the Invention
[0005] In view of this, in order to solve the above-mentioned technical problems, the present invention provides a method and related apparatus for identifying loan items for multiple sets of financial classification standard catalogs.
[0006] The present invention adopts the following technical solution:
[0007] In a first aspect, the present invention provides a method for identifying loan items oriented towards multiple sets of financial classification standard catalogs, including: Receive text information and a target directory identifier for the target loan project; the text information is used to describe the project content of the target loan project. Based on the target directory identifier, the corresponding target knowledge sub-base is located in the preset multi-directory comprehensive knowledge base; the multi-directory comprehensive knowledge base contains knowledge sub-bases corresponding to multiple financial classification standard directories, and each knowledge sub-base contains knowledge units corresponding to multiple entries of the financial classification standard directories; Based on the text information, each knowledge unit in the target knowledge sub-base is retrieved to generate a candidate knowledge set; Based on the target directory identifier, select the corresponding target prompt instruction set from a set of preset prompt instructions; The text information, the candidate knowledge set, and the target prompt instruction set are input into a preset large language model to obtain the recognition result output by the large language model. The target prompt instruction set is used to guide the large language model to generate the recognition result corresponding to the text information based on the candidate knowledge set. The recognition result includes a directory attribution result, which is used to indicate whether the target loan project belongs to the financial classification standard directory corresponding to the target directory identifier. When the directory attribution result indicates that the target loan project belongs to the financial classification standard directory corresponding to the target directory identifier, the recognition result also includes an entry attribution result.
[0008] Optionally, the knowledge unit includes entry content corresponding to the entry, directory source tags, strategy vocabulary, and a first vector representation; The entry content includes an entry identifier, entry name, content description, judgment elements, and reference document name; the content description includes a textual description of the project activity or category definition corresponding to the entry; the judgment elements include criteria or judgment rules used to determine whether it belongs to the corresponding entry; The first vector is generated based on the content description and judgment elements corresponding to the entry content; The directory source tag is used to identify the financial classification standard directory to which the corresponding knowledge unit belongs, so as to distinguish knowledge units in different financial classification standard directories through the directory source tag; The strategy terminology includes positive-specific keywords and negative-specific keywords.
[0009] Optionally, based on the text information, each knowledge unit in the target knowledge sub-base is retrieved to generate a candidate knowledge set, specifically including: Based on the text information, a semantic similarity retrieval strategy is used to retrieve each knowledge unit of the target knowledge sub-base to obtain a first candidate knowledge set; the first candidate knowledge set includes at least one first target knowledge unit, the semantic similarity between the first vector representation and the second vector representation of the first target knowledge unit satisfies a first preset condition, and the second vector representation is obtained by vectorizing the text information; Based on the text information, a keyword retrieval strategy is used to retrieve each knowledge unit of the target knowledge sub-base to obtain a second candidate knowledge set; the second candidate knowledge set includes at least one second target knowledge unit, the text information hits the positive dedicated keyword in the second target knowledge unit, and the text information does not hit the negative dedicated keyword in the second target knowledge unit; Based on the text information, a text similarity retrieval strategy is used to retrieve each knowledge unit of the target knowledge sub-base to obtain a third candidate knowledge set; the third candidate knowledge set includes at least one third target knowledge unit, and the text similarity between the text composed of the content description and judgment elements of the third target knowledge unit and the text information satisfies the second preset condition. The first candidate knowledge set, the second candidate knowledge set, and the third candidate knowledge set are merged and deduplicated to obtain the candidate knowledge set.
[0010] Optionally, the first preset condition is: the number of the first target knowledge units is a first number, and the semantic similarity between the first vector representation and the second vector representation of the first target knowledge unit is higher than the semantic similarity between the first vector representation and the second vector representation of any other knowledge unit in the target knowledge sub-base other than all the first target knowledge units; The second preset condition is that the number of the third target knowledge units is the second number, and the text similarity between the text composed of the content description and judgment elements of the third target knowledge units and the text information is higher than the text similarity between the text composed of the content description and judgment elements of any other knowledge unit in the target knowledge sub-base other than all the third target knowledge units and the text information.
[0011] Optionally, based on the text information, a text similarity retrieval strategy is used to retrieve each of the knowledge units in the target knowledge sub-base to obtain a third candidate knowledge set, specifically including: Based on the target directory identifier, select the corresponding target non-dedicated keyword thesaurus from a set of preset non-dedicated keyword thesauruses; For each knowledge unit in the target knowledge sub-base, the BM25 text similarity algorithm is used to calculate the text similarity between the text information and the text composed of the content description and judgment elements of the knowledge unit, based on the target non-dedicated keyword vocabulary. The second number of knowledge units with the highest text similarity are selected as the third target knowledge units, and all the third target knowledge units constitute the third candidate knowledge set.
[0012] Optionally, the prompt instruction set includes: role setting instructions, knowledge anchoring instructions, guide word instructions, and output format instructions; The role setting instruction is used to set the large language model as a judgment expert in the field to which the corresponding financial classification standard catalog belongs; The knowledge anchoring instruction is used to constrain the large language model to make judgments based solely on the candidate knowledge set; The guiding words contain explanatory notes or judgment points related to the corresponding financial classification standard catalog. The guiding words are used to assist the large language model in semantic understanding and logical reasoning. The output format instruction is used to specify the format in which the large language model outputs the recognition result.
[0013] Optionally, the item attribution result includes: item identifier and determination criteria.
[0014] Secondly, the present invention provides a loan item identification device for multiple sets of financial classification standard catalogs, comprising: The receiving module is used to receive text information and a target directory identifier for the target loan project; the text information is used to describe the project content of the target loan project. The positioning module is used to locate the corresponding target knowledge sub-base in a preset multi-directory comprehensive knowledge base based on the target directory identifier; the multi-directory comprehensive knowledge base contains knowledge sub-bases corresponding to multiple financial classification standard directories, and each knowledge sub-base contains knowledge units corresponding to multiple entries of the financial classification standard directories. The retrieval module is used to retrieve each knowledge unit of the target knowledge sub-base based on the text information, so as to generate a candidate knowledge set; The selected module is used to select the corresponding target prompt instruction set from a set of preset prompt instruction sets based on the target directory identifier; The module is used to input the text information, the candidate knowledge set, and the target prompt instruction set into a preset large language model to obtain the recognition result output by the large language model; the target prompt instruction set is used to guide the large language model to generate the recognition result corresponding to the text information based on the candidate knowledge set; the recognition result includes a directory attribution result, which is used to indicate whether the target loan project belongs to the financial classification standard directory corresponding to the target directory identifier; when the directory attribution result indicates that the target loan project belongs to the financial classification standard directory corresponding to the target directory identifier, the recognition result also includes an entry attribution result.
[0015] Thirdly, the present invention provides a computer device, comprising: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to implement the loan item identification method for multiple financial classification standard catalogs as described above.
[0016] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the loan item identification method for multiple sets of financial classification standard catalogs as described above.
[0017] This invention employs the above technical solution, pre-establishing a multi-directory comprehensive knowledge base, which contains knowledge sub-bases corresponding to multiple financial classification standard directories. It also pre-establishes prompt instruction sets corresponding to these multiple financial classification standard directories. Subsequently, upon receiving the text information of the target loan project and the target directory identifier, the corresponding target knowledge sub-base and target prompt instruction set are located based on the target directory identifier. A candidate knowledge set is obtained by retrieving the target knowledge sub-base. The text information, candidate knowledge set, and target prompt instruction set are then input into a pre-set large language model to obtain the recognition result output by the large language model. Based on this, by pre-establishing knowledge sub-bases and prompt instruction sets corresponding to multiple financial classification standard directories, and subsequently locating the corresponding target knowledge sub-base and selecting the corresponding target prompt instruction set based on the target directory identifier, this invention enables unified support for multiple financial classification standard directories through a single large language model. This significantly reduces the cost for financial institutions to repeatedly develop and maintain multiple independent systems to meet different regulatory requirements. Furthermore, by using prompt instruction sets to guide the large language model in generating recognition results, this invention improves recognition accuracy. Finally, when a new financial classification standard catalog is needed, a new knowledge sub-base and prompt instruction set are constructed according to the established structure, without the need to repeatedly develop a new system, thus giving the invention good scalability. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating a method for identifying loan items based on multiple sets of financial classification standards provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a loan item identification device for multiple sets of financial classification standard catalogs provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be described in detail below. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other implementation methods obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0021] Figure 1 This is a flowchart illustrating a method for identifying loan items based on multiple sets of financial classification standards, provided by an embodiment of the present invention. Figure 1 As shown, this process includes: Step 101: Receive text information and target directory identifier of the target loan project; the text information is used to describe the project content of the target loan project.
[0022] Specifically, the text information may include details such as the specific purpose, content, and involved parties of the target loan project. For example, the text information could be: "This loan application is to support a new energy technology company in purchasing and installing 5 MW of distributed photovoltaic power generation modules and supporting grid-connected equipment on the roof of its factory in an industrial park. After the project is completed, the generated electricity will mainly be used by the company itself, with surplus electricity sold to the grid." The target directory identifier is used to indicate the financial classification standard directory on which this identification is based. For example, a target directory identifier of "green" indicates that the financial classification standard directory on which this identification is based is the "Green Finance Supported Project Directory." A target directory identifier of "agriculture" indicates that the financial classification standard directory on which this identification is based is the "Special Statistical System for Agricultural Loans." A target directory identifier of "biodiversity" indicates that the financial classification standard directory on which this identification is based is the "Green Finance Supported Biodiversity Conservation Project Directory."
[0023] It should be noted that the target directory identifier can be specified by the user. The content of the text information corresponding to different directory identifiers can be different. In a specific example, when the directory identifier is "Special Statistical System for Agricultural Loans", the corresponding text information will at least include the loan purpose, borrower type, and borrower's registered region.
[0024] Step 102: Based on the target directory identifier, locate the corresponding target knowledge sub-base in the preset multi-directory comprehensive knowledge base; the multi-directory comprehensive knowledge base contains knowledge sub-bases corresponding to multiple financial classification standard directories, and each knowledge sub-base contains knowledge units corresponding to multiple entries of the corresponding financial classification standard directories.
[0025] Specifically, the multi-directory integrated knowledge base includes the association between knowledge sub-bases and their corresponding financial classification standard directory directory identifiers. When locating the corresponding target knowledge sub-base in the preset multi-directory integrated knowledge base based on the target directory identifier, the knowledge sub-base associated with the target directory identifier is located as the target knowledge sub-base based on this association.
[0026] The financial classification standard catalogs involved in this invention may include: the "Catalogue of Green Finance Supported Projects", the "Special Statistical System for Agricultural Loans", and the "Catalogue of Green Finance Supported Biodiversity Conservation Projects", etc.
[0027] In this embodiment of the invention, the knowledge unit includes the entry content of the corresponding entry, the directory source tag, the strategy vocabulary and the first vector representation.
[0028] The entry content includes the entry identifier, entry name, content description, judgment elements, and reference document name; the content description includes a textual description of the project activities or classification definition of the corresponding entry; the judgment elements include the criteria or judgment rules used to determine whether it belongs to the corresponding entry.
[0029] The first vector is generated based on the content description and judgment elements of the corresponding entry.
[0030] The directory source tag is used to identify the financial classification standard directory to which the corresponding knowledge unit belongs, so as to distinguish knowledge units in different financial classification standard directories through the directory source tag. In a specific example, the directory source tags are such as source=green, source=agriculture, and source=biodiversity.
[0031] The strategy terminology includes positive and negative dedicated keywords.
[0032] The construction of a multi-category comprehensive knowledge base aims to transform unstructured or semi-structured financial classification standard catalogs into machine-readable and searchable structured knowledge, thereby laying the data foundation for the entire methodology. In a specific example, the construction process of the multi-category comprehensive knowledge base includes: (1) For each of the above multiple financial classification standard directories, receive the entry content, directory source tags and strategy terminology of each knowledge unit corresponding to the financial classification standard directory input by the user.
[0033] In a specific example, for the "Green Finance Supported Project Catalogue," each third-level classification code corresponds to one entry, such as "4.2.2 Construction and Operation of Solar Energy Utilization Facilities." The entry identifier for this entry can be 4.2.2, and the entry name can be "Construction and Operation of Solar Energy Utilization Facilities."
[0034] In another specific example, regarding the "Special Statistical System for Agricultural Loans," one classification definition can correspond to one entry, and one judgment rule and verification logic can also correspond to one entry. Classification definitions can include, for example, the definition of "farmers," the geographical boundary definition of "rural areas" and "urban areas," and the scope of activities for "agricultural, forestry, animal husbandry, and fishery loans." The definition of "farmers" can be: households that have resided for more than one year within the administrative area of a township (excluding the town center), or households that have resided for more than one year within the administrative villages under the jurisdiction of the town center, with local household registration, or households with household registration in another location but residing locally for more than one year and having established stable economic and social relationships locally. Judgment rules and verification logic refer to the inclusion, overlap, and mutual exclusion relationships between statistical indicators, such as "farmer loans = farmer consumption loans + farmer business loans," and "agricultural loans = non-farmer individual agricultural, forestry, animal husbandry, and fishery loans + farmer loans + rural non-financial enterprise and government agency loans + urban non-financial enterprise and government agency agricultural loans."
[0035] (2) A text embedding model is adopted to vectorize the content description and judgment elements in each item into the corresponding first vector representation. The first vector representation can be a fixed-dimensional vector, such as 1024-dimensional. The text embedding model can be the BGE-M3 model (Bidirectional General Embedding Model-M3, M3 version of bidirectional general embedding model), which can effectively capture the deep semantics of Chinese text.
[0036] (3) Store the entry content, directory source tags, strategy thesaurus, and first vector representation of each knowledge unit in a dedicated vector database, such as the open-source Milvus database. A key implementation method is to utilize the partitioning or aggregating functions of the vector database to create an independent index for each financial classification standard directory. This physically or logically isolates the knowledge of different directories, forming multiple knowledge sub-bases. For example, all knowledge units corresponding to the "Green Finance Supported Project Directory" are organized in the "Green Finance Knowledge Sub-base" and indexed, while all knowledge units corresponding to the "Special Statistical System for Agricultural Loans" are organized in the "Agricultural Loans Knowledge Sub-base" and indexed separately. Understandably, this logically isolated knowledge sub-base structure lays the foundation for dynamic directory routing and avoidance of knowledge crosstalk.
[0037] Step 103: Based on text information, retrieve each knowledge unit of the target knowledge sub-base to generate a candidate knowledge set.
[0038] In this embodiment of the invention, based on text information, each knowledge unit of the target knowledge sub-base is retrieved to generate a candidate knowledge set, specifically including: (1) Based on text information, a semantic similarity retrieval strategy is adopted to retrieve each knowledge unit of the target knowledge sub-base and obtain the first candidate knowledge set; the first candidate knowledge set includes at least one first target knowledge unit, the semantic similarity between the first vector representation and the second vector representation of the first target knowledge unit satisfies the first preset condition, and the second vector representation is obtained by vectorizing the text information.
[0039] In a specific example, the first preset condition can be: the number of the first target knowledge units is a first quantity, and the semantic similarity between the first vector representation and the second vector representation of the first target knowledge unit is higher than the semantic similarity between the first vector representation and the second vector representation of any other knowledge unit in the target knowledge sub-base other than all the first target knowledge units.
[0040] Based on this, and using textual information, a semantic similarity retrieval strategy is employed to retrieve each knowledge unit of the target knowledge sub-base, resulting in the first candidate knowledge set, which specifically includes: First, the text information is vectorized using the same text embedding model as when building the multi-directory integrated knowledge base, to obtain a second vector representation.
[0041] Then, for each knowledge unit in the target knowledge sub-base, the semantic similarity between the first vector representation and the second vector representation of the knowledge unit is calculated. The semantic similarity can be cosine similarity.
[0042] Finally, the first number of knowledge units with the highest semantic similarity are selected as the first target knowledge units, and all the first target knowledge units constitute the first candidate knowledge set. The first number can be 5.
[0043] (2) Based on the text information, a keyword retrieval strategy is adopted to retrieve each knowledge unit of the target knowledge sub-base to obtain the second candidate knowledge set; the second candidate knowledge set includes at least one second target knowledge unit, the text information hits the positive special keywords in the second target knowledge unit, and the text information does not hit the negative special keywords in the second target knowledge unit.
[0044] In a specific example, based on text information, a keyword retrieval strategy is used to retrieve each knowledge unit of the target knowledge sub-base, resulting in a second candidate knowledge set. Specifically, this may include: for each knowledge unit in the target knowledge sub-base, performing regular expression matching between the strategy vocabulary of the knowledge unit and the text information. Knowledge units that match positive specific keywords but not negative specific keywords are identified as second target knowledge units, and all second target knowledge units constitute the second candidate knowledge set.
[0045] (3) Based on text information, a text similarity retrieval strategy is adopted to retrieve each knowledge unit of the target knowledge sub-base and obtain a third candidate knowledge set; the third candidate knowledge set includes at least one third target knowledge unit, and the text similarity between the text composed of the content description and judgment elements of the third target knowledge unit and the text information satisfies the second preset condition.
[0046] In a specific example, the second preset condition can be: the number of third target knowledge units is the second number, and the text similarity between the text composed of the content description and judgment elements of the third target knowledge units and the text information is higher than the text similarity between the text composed of the content description and judgment elements of any other knowledge unit in the target knowledge sub-base other than all third target knowledge units and the text information.
[0047] Based on this, and using textual information, a text similarity retrieval strategy is employed to retrieve each knowledge unit of the target knowledge sub-base, resulting in a third candidate knowledge set, which may specifically include: First, based on the target directory identifier, the corresponding target non-dedicated keyword thesaurus is selected from multiple preset non-dedicated keyword thesauruses. It should be noted that the non-dedicated keyword thesaurus corresponding to each directory identifier is pre-defined in this invention.
[0048] Secondly, for each knowledge unit in the target knowledge sub-base, the BM25 text similarity algorithm is used to calculate the text similarity between the text information and the text composed of the content description and judgment elements of the knowledge unit, based on the target non-specific keyword vocabulary.
[0049] Finally, the second number of knowledge units with the highest text similarity are selected as the third target knowledge units, and all the third target knowledge units constitute the third candidate knowledge set.
[0050] (4) Merge the first candidate knowledge set, the second candidate knowledge set and the third candidate knowledge set to remove duplicates and obtain the candidate knowledge set.
[0051] Step 104: Based on the target directory identifier, select the corresponding target prompt instruction set from the preset multiple prompt instruction sets.
[0052] Specifically, the prompt instruction set may include: character setting instructions, knowledge anchoring instructions, guide word instructions, and output format instructions.
[0053] Role setting instructions are used to assign the large language model the role of a judgment expert in the corresponding financial classification standard directory. In a specific example, when the target directory is identified as the "Green Finance Supported Projects Directory," the corresponding role setting instruction could be: "You are a professional green loan judgment expert; please strictly adhere to the 'Green Finance Supported Projects Directory' in your judgments." In another specific example, when the target directory is identified as the "Special Statistical System for Agricultural Loans," the corresponding role setting instruction could be: "You are a professional agricultural loan statistics expert; please strictly adhere to the 'Special Statistical System for Agricultural Loans' in your judgments."
[0054] Knowledge anchoring directives are used to constrain large language models to make decisions based solely on the candidate knowledge set. In a specific example, a knowledge anchoring directive could be: all your decisions must be based solely on the provided 'candidate knowledge set'.
[0055] The guiding instructions contain explanatory notes or judgment points related to the corresponding financial classification standard catalog. These instructions assist the large language model in semantic understanding and logical reasoning. In a specific example, when the target catalog is identified as the "Green Finance Supported Project Catalog," the corresponding guiding instructions could be: For "solar energy utilization," it is necessary to focus on whether the project involves the construction or operation of solar power generation facilities, rather than the manufacturing of solar equipment. In another specific example, when the target catalog is identified as the "Special Statistical System for Agricultural Loans," the corresponding guiding instructions could include: accurately classifying loans into one of the following five categories: "Agricultural Loans," "Agriculture, Forestry, Animal Husbandry and Fishery Loans," "Farmer Loans," "Rural Non-Financial Enterprises and Government Agencies Loans," and "Urban Non-Financial Enterprises and Government Agencies Agricultural Loans." Furthermore, the corresponding guiding instructions could also include: logical guidance for the detailed statistical classification of agricultural loans to assist the model in accurate classification. For example, the determination of "farmer" requires simultaneously meeting the criteria of "long-term residence for more than one year," "mainly engaged in agricultural production," and "unrelated to household registration." "Rural areas" refers to areas other than prefecture-level and above urban administrative districts, municipal towns, county towns, and development zones. Agricultural loans = Non-agricultural individual loans for agriculture, forestry, animal husbandry and fishery + Farmer loans + Rural non-financial enterprises and government agencies loans + Urban non-financial enterprises and government agencies agricultural loans.
[0056] The output format instruction specifies the format of the recognition results output by the large language model, so that the recognition results include information such as whether it belongs to the current directory, the entry identifier, and the judgment criteria. In a specific example, when the target directory identifier is the directory identifier of "Green Finance Supported Projects Directory", the corresponding output format instruction is: output in JSON format, including whether it belongs to "Green Finance Supported Projects Directory", the identifier of the most matching entry, and the judgment criteria.
[0057] Step 105: Input the text information, candidate knowledge set, and target prompt instruction set into the preset large language model to obtain the recognition result output by the large language model; the target prompt instruction set is used to guide the large language model to generate the recognition result corresponding to the text information based on the candidate knowledge set; the recognition result includes the directory attribution result, which is used to indicate whether the target loan project belongs to the financial classification standard directory corresponding to the target directory identifier; when the directory attribution result indicates that the target loan project belongs to the financial classification standard directory corresponding to the target directory identifier, the recognition result also includes the entry attribution result.
[0058] In this embodiment of the invention, the item attribution result may include: item identifier and determination basis. Secondly, when the directory attribution result indicates that the target loan project does not belong to the financial classification standard directory corresponding to the target directory identifier, the identification result may also include the reason for non-attribution.
[0059] In this embodiment of the invention, after inputting text information, a candidate knowledge set, and a target prompt instruction set into a preset large language model, the large language model performs the following reasoning process: First, the large language model reads the role setting instructions in the target prompt instruction set to clarify its current expert role. Simultaneously, the large language model understands the output format instructions to know the structure the final answer should present. Second, the large language model strictly follows the knowledge anchoring instructions, shielding itself from potentially irrelevant or outdated knowledge in its pre-training parameters and focusing entirely on the input candidate knowledge set. Third, the large language model compares the text information with each knowledge unit in the candidate knowledge set one by one, while combining the explanatory notes or judgment points given in the guiding words to obtain intermediate reasoning conclusions. Finally, the large language model generates and outputs the recognition result based on the intermediate reasoning conclusions, and the recognition result conforms to the output format instructions.
[0060] In this embodiment of the invention, after obtaining the recognition result output by the large language model, since the recognition result is in JSON format, the JSON string of the recognition result can be parsed into structured data for display and storage.
[0061] This invention can receive identification requests from external systems (such as a bank's credit approval system) through a standardized application programming interface (API). The identification request can be an HTTP POST request with a JSON object body containing at least two key fields: `project_text` and `target_catalog`. The `project_text` field contains text information about the target loan project. The `target_catalog` field contains a string representing the target catalog identifier. After obtaining the parsed identification result, the API is used to return the result to the external system, achieving seamless integration with existing external systems.
[0062] This invention employs the above technical solution, pre-establishing a multi-directory comprehensive knowledge base, which includes knowledge sub-bases corresponding to multiple financial classification standard directories. It also pre-establishes prompt instruction sets corresponding to these multiple financial classification standard directories. Subsequently, after receiving the text information of the target loan project and the target directory identifier, the corresponding target knowledge sub-base and target prompt instruction set are located based on the target directory identifier. A candidate knowledge set is obtained by retrieving the target knowledge sub-base. The text information, candidate knowledge set, and target prompt instruction set are then input into a pre-set large language model to obtain the recognition result output by the large language model. Based on this, by pre-establishing knowledge sub-bases and prompt instruction sets corresponding to multiple financial classification standard directories, and subsequently locating the corresponding target knowledge sub-base and selecting the corresponding target prompt instruction set based on the target directory identifier, this invention can provide unified support for multiple financial classification standard directories through a single large language model. This significantly reduces the cost for financial institutions to repeatedly develop and maintain multiple independent systems to meet different regulatory requirements. Furthermore, by using prompt instruction sets to guide the large language model in generating recognition results, this invention can improve recognition accuracy. Finally, when a new financial classification standard catalog is needed, a new knowledge sub-base and prompt instruction set are constructed according to the established structure, without the need to repeatedly develop a new system, thus giving the invention good scalability.
[0063] Based on a general inventive concept, the present invention also provides a loan item identification device for multiple sets of financial classification standard catalogs. Figure 2 This is a schematic diagram of the structure of a loan item identification device for multiple sets of financial classification standard catalogs provided in an embodiment of the present invention. Figure 2 As shown, this device includes: The receiving module 21 is used to receive text information and target directory identifier of the target loan project; the text information is used to describe the project content of the target loan project.
[0064] The positioning module 22 is used to locate the corresponding target knowledge sub-base in the preset multi-directory comprehensive knowledge base based on the target directory identifier; the multi-directory comprehensive knowledge base contains knowledge sub-bases corresponding to multiple financial classification standard directories, and each knowledge sub-base contains knowledge units corresponding to multiple entries of the corresponding financial classification standard directories.
[0065] The retrieval module 23 is used to retrieve each knowledge unit of the target knowledge sub-base based on text information in order to generate a candidate knowledge set.
[0066] Select module 24 is used to select the corresponding target prompt instruction set from a set of preset prompt instructions based on the target directory identifier.
[0067] Module 25 is used to input text information, candidate knowledge set, and target prompt instruction set into a preset large language model to obtain the recognition result output by the large language model. The target prompt instruction set is used to guide the large language model to generate the recognition result corresponding to the text information based on the candidate knowledge set. The recognition result includes the directory attribution result, which indicates whether the target loan project belongs to the financial classification standard directory corresponding to the target directory identifier. When the directory attribution result indicates that the target loan project belongs to the financial classification standard directory corresponding to the target directory identifier, the recognition result also includes the entry attribution result.
[0068] Optionally, the knowledge unit includes the entry content of the corresponding entry, the directory source tag, the strategy vocabulary, and the first vector representation.
[0069] The entry content includes the entry identifier, entry name, content description, judgment elements, and reference document name; the content description includes a textual description of the project activities or classification definition of the corresponding entry; the judgment elements include the criteria or judgment rules used to determine whether it belongs to the corresponding entry.
[0070] The first vector is generated based on the content description and judgment elements of the corresponding entry.
[0071] The directory source tag is used to identify the financial classification standard directory to which the corresponding knowledge unit belongs, so as to distinguish knowledge units in different financial classification standard directories through the directory source tag.
[0072] The strategy terminology includes positive and negative dedicated keywords.
[0073] Optionally, the retrieval module 23 may specifically include: The semantic similarity retrieval unit is used to retrieve each knowledge unit of the target knowledge sub-base based on text information and employ a semantic similarity retrieval strategy to obtain a first candidate knowledge set. The first candidate knowledge set includes at least one first target knowledge unit, and the semantic similarity between the first vector representation and the second vector representation of the first target knowledge unit satisfies a first preset condition. The second vector representation is obtained by vectorizing the text information.
[0074] The keyword retrieval unit is used to retrieve each knowledge unit of the target knowledge sub-base based on text information and using a keyword retrieval strategy to obtain a second candidate knowledge set. The second candidate knowledge set includes at least one second target knowledge unit, where the text information hits the positive dedicated keyword in the second target knowledge unit, and the text information does not hit the negative dedicated keyword in the second target knowledge unit.
[0075] The text similarity retrieval unit is used to retrieve each knowledge unit of the target knowledge sub-base based on text information and using a text similarity retrieval strategy to obtain a third candidate knowledge set. The third candidate knowledge set includes at least one third target knowledge unit, and the text similarity between the text composed of the content description and judgment elements of the third target knowledge unit and the text information satisfies the second preset condition.
[0076] The merge and deduplication unit is used to merge and deduplicate the first candidate knowledge set, the second candidate knowledge set, and the third candidate knowledge set to obtain the candidate knowledge set.
[0077] Optionally, the first preset condition is: the number of the first target knowledge units is a first quantity, and the semantic similarity between the first vector representation and the second vector representation of the first target knowledge unit is higher than the semantic similarity between the first vector representation and the second vector representation of any other knowledge unit in the target knowledge sub-base other than all the first target knowledge units.
[0078] The second preset condition is: the number of third target knowledge units is the second number, and the text similarity between the text and text information composed of the content description and judgment elements of the third target knowledge units is higher than the text similarity between the text and text information composed of the content description and judgment elements of any other knowledge unit in the target knowledge sub-base other than all third target knowledge units.
[0079] Optional, text similarity retrieval unit, specifically can be used for: Based on the target directory identifier, select the corresponding target non-dedicated keyword thesaurus from multiple preset non-dedicated keyword thesauruses.
[0080] For each knowledge unit in the target knowledge sub-base, the BM25 text similarity algorithm is used to calculate the text similarity between the text information and the text composed of the content description and judgment elements of the knowledge unit, based on the target non-specific keyword vocabulary.
[0081] The second number of knowledge units with the highest text similarity are selected as the third target knowledge units, and all the third target knowledge units constitute the third candidate knowledge set.
[0082] Optionally, the prompt instruction set includes: character setting instructions, knowledge anchoring instructions, guide word instructions, and output format instructions.
[0083] The role setting command is used to set the large language model as a judgment expert in the field to which the corresponding financial classification standard directory belongs.
[0084] Knowledge anchoring instructions are used to constrain large language models to make decisions based solely on candidate knowledge sets.
[0085] The guiding words contain explanatory notes or judgment points related to the corresponding financial classification standard catalog. The guiding words are used to assist the large language model in semantic understanding and logical reasoning.
[0086] Output format instructions are used to specify the format of the recognition results output by the large language model.
[0087] Optionally, the item attribution results include: item identifier and determination criteria.
[0088] Based on a general inventive concept, the present invention also provides a computer device. Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 3 As shown, computer device 300 includes: At least one processor 310; and, Memory 330 communicatively connected to at least one processor 310; wherein, The memory 330 stores instructions 320 that can be executed by at least one processor 310, which enables the at least one processor 310 to implement the loan item identification method for multiple financial classification standard catalogs as described above.
[0089] Based on a general inventive concept, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the loan item identification method for multiple sets of financial classification standard catalogs as described above.
[0090] It is understood that the same or similar parts in the above embodiments can be referred to each other, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.
[0091] It should be noted that in the description of this invention, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this invention, unless otherwise stated, "a plurality of" means at least two.
[0092] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
[0093] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0094] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0095] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0096] The storage media mentioned above can be read-only memory, disk, or optical disk, etc.
[0097] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0098] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for identifying loan items in multiple sets of financial classification standards, characterized in that, include: Receive text information and target directory identifiers for the target loan project; The text information is used to describe the project content of the target loan project; Based on the target directory identifier, locate the corresponding target knowledge sub-base in the preset multi-directory comprehensive knowledge base; The multi-directory comprehensive knowledge base contains knowledge sub-bases corresponding to multiple financial classification standard directories, and each knowledge sub-base contains knowledge units corresponding to multiple entries of the financial classification standard directories. Based on the text information, each knowledge unit in the target knowledge sub-base is retrieved to generate a candidate knowledge set; Based on the target directory identifier, select the corresponding target prompt instruction set from a set of preset prompt instructions; The text information, the candidate knowledge set, and the target prompt instruction set are input together into a preset large language model to obtain the recognition result output by the large language model; The target prompt instruction set is used to guide the large language model to generate the recognition result corresponding to the text information based on the candidate knowledge set; the recognition result includes the directory attribution result, which is used to indicate whether the target loan project belongs to the financial classification standard directory corresponding to the target directory identifier; When the directory attribution result indicates that the target loan project belongs to the financial classification standard directory corresponding to the target directory identifier, the identification result also includes the entry attribution result.
2. The loan project identification method for multiple sets of financial classification standard catalogs according to claim 1, characterized in that, The knowledge unit includes the entry content corresponding to the entry, directory source tags, strategy vocabulary, and a first vector representation; The entry content includes an entry identifier, entry name, content description, judgment elements, and reference document name; the content description includes a textual description of the project activity or category definition corresponding to the entry; the judgment elements include criteria or judgment rules used to determine whether it belongs to the corresponding entry; The first vector is generated based on the content description and judgment elements corresponding to the entry content; The directory source tag is used to identify the financial classification standard directory to which the corresponding knowledge unit belongs, so as to distinguish knowledge units in different financial classification standard directories through the directory source tag; The strategy terminology includes positive-specific keywords and negative-specific keywords.
3. The loan project identification method for multiple sets of financial classification standard catalogs according to claim 2, characterized in that, Based on the text information, each knowledge unit of the target knowledge sub-base is retrieved to generate a candidate knowledge set, specifically including: Based on the text information, a semantic similarity retrieval strategy is used to retrieve each knowledge unit of the target knowledge sub-base to obtain a first candidate knowledge set; the first candidate knowledge set includes at least one first target knowledge unit, the semantic similarity between the first vector representation and the second vector representation of the first target knowledge unit satisfies a first preset condition, and the second vector representation is obtained by vectorizing the text information; Based on the text information, a keyword retrieval strategy is used to retrieve each knowledge unit of the target knowledge sub-base to obtain a second candidate knowledge set; the second candidate knowledge set includes at least one second target knowledge unit, the text information hits the positive dedicated keyword in the second target knowledge unit, and the text information does not hit the negative dedicated keyword in the second target knowledge unit; Based on the text information, a text similarity retrieval strategy is used to retrieve each knowledge unit of the target knowledge sub-base to obtain a third candidate knowledge set; the third candidate knowledge set includes at least one third target knowledge unit, and the text similarity between the text composed of the content description and judgment elements of the third target knowledge unit and the text information satisfies the second preset condition. The first candidate knowledge set, the second candidate knowledge set, and the third candidate knowledge set are merged and deduplicated to obtain the candidate knowledge set.
4. The loan project identification method for multiple sets of financial classification standard catalogs according to claim 3, characterized in that, The first preset condition is: the number of the first target knowledge units is a first number, and the semantic similarity between the first vector representation and the second vector representation of the first target knowledge unit is higher than the semantic similarity between the first vector representation and the second vector representation of any other knowledge unit in the target knowledge sub-base other than all the first target knowledge units; The second preset condition is that the number of the third target knowledge units is the second number, and the text similarity between the text composed of the content description and judgment elements of the third target knowledge units and the text information is higher than the text similarity between the text composed of the content description and judgment elements of any other knowledge unit in the target knowledge sub-base other than all the third target knowledge units and the text information.
5. The loan project identification method for multiple sets of financial classification standard catalogs according to claim 4, characterized in that, Based on the aforementioned text information, a text similarity retrieval strategy is employed to retrieve each knowledge unit of the target knowledge sub-base, thereby obtaining a third candidate knowledge set, which specifically includes: Based on the target directory identifier, select the corresponding target non-dedicated keyword thesaurus from a set of preset non-dedicated keyword thesauruses; For each knowledge unit in the target knowledge sub-base, the BM25 text similarity algorithm is used to calculate the text similarity between the text information and the text composed of the content description and judgment elements of the knowledge unit, based on the target non-dedicated keyword vocabulary. The second number of knowledge units with the highest text similarity are selected as the third target knowledge units, and all the third target knowledge units constitute the third candidate knowledge set.
6. The loan project identification method for multiple sets of financial classification standard catalogs according to claim 1, characterized in that, The prompt instruction set includes: role setting instructions, knowledge anchoring instructions, guide word instructions, and output format instructions; The role setting instruction is used to set the large language model as a judgment expert in the field to which the corresponding financial classification standard catalog belongs; The knowledge anchoring instruction is used to constrain the large language model to make judgments based solely on the candidate knowledge set; The guiding words contain explanatory notes or judgment points related to the corresponding financial classification standard catalog. The guiding words are used to assist the large language model in semantic understanding and logical reasoning. The output format instruction is used to specify the format in which the large language model outputs the recognition result.
7. The loan project identification method for multiple sets of financial classification standard catalogs according to claim 2, characterized in that, The item attribution results include: item identifier and determination criteria.
8. A loan item identification device for multiple sets of financial classification standard catalogs, characterized in that, include: The receiving module is used to receive text information and target directory identifiers of the target loan project; The text information is used to describe the project content of the target loan project; The positioning module is used to locate the corresponding target knowledge sub-base in a preset multi-directory comprehensive knowledge base based on the target directory identifier; The multi-directory comprehensive knowledge base contains knowledge sub-bases corresponding to multiple financial classification standard directories, and each knowledge sub-base contains knowledge units corresponding to multiple entries of the financial classification standard directories. The retrieval module is used to retrieve each knowledge unit of the target knowledge sub-base based on the text information, so as to generate a candidate knowledge set; The selected module is used to select the corresponding target prompt instruction set from a set of preset prompt instruction sets based on the target directory identifier; The module is used to input the text information, the candidate knowledge set, and the target prompt instruction set into a preset large language model to obtain the recognition result output by the large language model; The target prompt instruction set is used to guide the large language model to generate the recognition result corresponding to the text information based on the candidate knowledge set; the recognition result includes the directory attribution result, which is used to indicate whether the target loan project belongs to the financial classification standard directory corresponding to the target directory identifier; When the directory attribution result indicates that the target loan project belongs to the financial classification standard directory corresponding to the target directory identifier, the identification result also includes the entry attribution result.
9. A computer device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which enable the at least one processor to implement the loan item identification method for multiple financial classification standard catalogs as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the loan item identification method for multiple sets of financial classification standard catalogs as described in any one of claims 1 to 7.