A method and device for generating a credit rating report
By optimizing the prompt word technology using a large language model and a cross-mutation algorithm, the problem of untimely data updates caused by manual information transmission in commercial bank credit ratings was solved, and the efficiency and quality of credit evaluation report generation were improved.
Patent Information
- Application Number
- CN202411466512.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-10-21
AI Technical Summary
In the existing technology, commercial banks' credit rating work for large and medium-sized enterprises mainly relies on manual input and transmission of information, which leads to untimely data updates and affects the efficiency of generating credit evaluation reports.
A large language model is combined with a cross-mutation algorithm and optimized prompt word technology. By obtaining basic information about the company and industry evaluation information, the cross-mutation algorithm is used to optimize the initial prompt words to generate a credit rating report.
It improves the efficiency of generating credit evaluation reports, avoids the problem of untimely data updates caused by manual information transmission, and ensures the quality of the reports.
Smart Images

Figure CN119417592B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of financial technology, and in particular to a method and device for generating a credit rating report. Background Art
[0002] Commercial banks' corporate lending business primarily involves pre-loan, loan-in-progress, and post-loan processes, involving multiple business scenarios such as credit rating and credit approval. Credit rating for large and medium-sized corporate clients is fundamental to how commercial banks manage and control customer credit risk.
[0003] Currently, loan approval is typically assessed based on the credit application materials submitted by each enterprise. These materials typically include a detailed description of the company's financial status through financial statements, due diligence reports, and other means to demonstrate repayment capabilities within a specified timeframe or under specific requirements. However, these materials are still primarily processed through manual, repetitive input and submission offline. This manual submission can lead to delayed updates of key data, hindering decision-making speed and ultimately inefficient credit evaluation report generation.
[0004] Therefore, how to improve the efficiency of generating credit evaluation reports has become an urgent problem that needs to be solved in this field. Summary of the Invention
[0005] This application provides a credit rating report generation method and device, the purpose of which is to improve the efficiency of generating credit evaluation reports.
[0006] In order to achieve the above objectives, this application provides the following technical solutions:
[0007] A method for generating a credit rating report, comprising:
[0008] Obtain basic information of the subject to be rated and industry evaluation information of the subject to be rated;
[0009] Inputting the basic information and the industry evaluation information into a large language model to obtain an evaluation grade of the object to be evaluated;
[0010] Obtaining an initial prompt word set corresponding to the object to be evaluated; the initial prompt word set includes multiple initial prompt words;
[0011] Optimizing the plurality of initial prompt words using a crossover mutation algorithm to obtain a plurality of crossover mutation prompt words;
[0012] Inputting each of the cross-variation prompt words, the scoring prompt words, and the evaluation grade of the object to be evaluated into a large language model to obtain a scoring result for each of the cross-variation prompt words; the scoring prompt words are obtained by correcting the initial scoring prompt words;
[0013] A credit rating report for the object to be rated is generated based on the rating grade of the object to be rated and the scoring result of each cross-variation prompt word.
[0014] Optionally, the step of inputting the basic information and the industry evaluation information into a large language model to obtain an evaluation grade of the object to be evaluated includes:
[0015] Splitting the industry evaluation information to obtain multiple preset evaluation levels;
[0016] For each of the preset evaluation levels, split the preset evaluation level to obtain multiple evaluation criteria;
[0017] Filter out the highest preset evaluation level from all preset evaluation levels and determine it as the target preset evaluation level;
[0018] Inputting a plurality of evaluation criteria corresponding to the target preset evaluation level and the basic information into a large language model to obtain an evaluation result; the evaluation result indicates the amount of the basic information that meets the evaluation criteria;
[0019] Calculate the ratio between the number of basic information that meets the evaluation criteria and the total number of evaluation criteria, and determine it as the standard ratio;
[0020] Determining whether the standard ratio is greater than a preset ratio;
[0021] If the standard ratio is greater than the preset ratio, the target preset evaluation level is determined as the evaluation level of the object to be evaluated;
[0022] If the standard ratio is not greater than the preset ratio, the highest preset evaluation level other than the target preset evaluation level is screened out from all the preset evaluation levels, and after being determined as the target preset evaluation level, the process returns to the step of inputting the multiple evaluation criteria corresponding to the target preset evaluation level and the basic information into the large language model to obtain the evaluation result.
[0023] Optionally, the crossover mutation algorithm is used to perform crossover mutation on the multiple initial prompt words to obtain multiple crossover mutation prompt words, including:
[0024] Converting multiple initial prompt words to obtain multiple initial prompt word vectors; the multiple initial prompt word vectors include the current initial prompt word vector;
[0025] Selecting any two initial prompt word vectors from the plurality of initial prompt word vectors and marking them as a first vector and a second vector;
[0026] Processing the difference between the first vector and the second vector using a transformation function to obtain a new vector;
[0027] Adding the current initial prompt word vector to the new vector to obtain a new prompt word vector;
[0028] For each of the initial prompt word vectors, cross-mutate the new prompt word vector and the initial prompt word vector to obtain a cross-mutated prompt word vector;
[0029] The cross-variation prompt word vector is converted to obtain multiple cross-variation prompt words.
[0030] Optionally, the process of correcting the initial scoring prompt word to obtain the scoring prompt word includes:
[0031] Acquire a prompt word set; the prompt word set includes multiple prompt words;
[0032] For each of the prompt words, input the prompt word into a large language model to obtain an initial credit rating report;
[0033] Inputting the initial credit rating report and initial rating prompt words into a large language model to obtain a rating result;
[0034] Determining whether the scoring result is consistent with a preset scoring result;
[0035] If the scoring result is consistent with the preset scoring result, the initial scoring prompt word is determined as the scoring prompt word;
[0036] If the scoring result is inconsistent with the preset scoring result, the initial scoring prompt word is corrected based on the initial credit rating report and the scoring result, and after the corrected scoring prompt word is determined as the scoring prompt word, the process returns to the step of inputting the initial credit rating report and the initial scoring prompt word into the large language model to obtain the scoring result.
[0037] Optionally, obtaining a prompt word set includes:
[0038] Acquire real data and expanded prompt words; wherein the real data indicates the basic information of the real object to be evaluated, and the expanded prompt words indicate prompt words that expand the basic information according to the format of the basic information;
[0039] Inputting the real data and the expanded prompt word into the large language model to obtain synthesized data, wherein the format of the synthesized data is consistent with that of the real data;
[0040] A preset evaluation function is used to generate a prompt word set corresponding to the synthesized data.
[0041] Optionally, generating a credit rating report for the object to be rated based on the rating of the object to be rated and the scoring result of each cross-variation prompt word includes:
[0042] Screening out the scoring result with the highest score from all the scoring results, and determining the cross-variation prompt word corresponding to the scoring result with the highest score as the best prompt word;
[0043] generating evaluation reasons based on the evaluation level of the object to be evaluated and the optimal prompt word;
[0044] A credit rating report for the object to be rated is generated based on the rating level of the object to be rated and the rating reasons.
[0045] A credit rating report generating device, comprising:
[0046] The first acquisition unit is used to acquire basic information of the object to be rated and industry evaluation information of the object to be rated;
[0047] A first input unit is used to input the basic information and the industry evaluation information into a large language model to obtain an evaluation level of the object to be evaluated;
[0048] A second acquisition unit is configured to acquire an initial prompt word set corresponding to the object to be evaluated; the initial prompt word set includes a plurality of initial prompt words;
[0049] an optimization unit, configured to optimize the plurality of initial prompt words using a crossover mutation algorithm to obtain a plurality of crossover mutation prompt words;
[0050] A second input unit is configured to input each of the cross-variation prompt words, the scoring prompt words, and the evaluation level of the object to be evaluated into a large language model to obtain a scoring result for each of the cross-variation prompt words; the scoring prompt words are obtained by correcting the initial scoring prompt words;
[0051] A generating unit is configured to generate a credit rating report for the object to be rated based on the rating grade of the object to be rated and the scoring result of each cross-variation prompt word.
[0052] Optionally, the first input unit is specifically configured to:
[0053] Splitting the industry evaluation information to obtain multiple preset evaluation levels;
[0054] For each of the preset evaluation levels, split the preset evaluation level to obtain multiple evaluation criteria;
[0055] Filter out the highest preset evaluation level from all preset evaluation levels and determine it as the target preset evaluation level;
[0056] Inputting a plurality of evaluation criteria corresponding to the target preset evaluation level and the basic information into a large language model to obtain an evaluation result; the evaluation result indicates the amount of the basic information that meets the evaluation criteria;
[0057] Calculate the ratio between the number of basic information that meets the evaluation criteria and the total number of evaluation criteria, and determine it as the standard ratio;
[0058] Determining whether the standard ratio is greater than a preset ratio;
[0059] If the standard ratio is greater than the preset ratio, the target preset evaluation level is determined as the evaluation level of the object to be evaluated;
[0060] If the standard ratio is not greater than the preset ratio, the highest preset evaluation level other than the target preset evaluation level is screened out from all the preset evaluation levels, and after being determined as the target preset evaluation level, the process returns to the step of inputting the multiple evaluation criteria corresponding to the target preset evaluation level and the basic information into the large language model to obtain the evaluation result.
[0061] Optionally, the optimization unit is specifically configured to:
[0062] Converting multiple initial prompt words to obtain multiple initial prompt word vectors; the multiple initial prompt word vectors include the current initial prompt word vector;
[0063] Selecting any two initial prompt word vectors from the plurality of initial prompt word vectors and marking them as a first vector and a second vector;
[0064] Processing the difference between the first vector and the second vector using a transformation function to obtain a new vector;
[0065] Adding the current initial prompt word vector to the new vector to obtain a new prompt word vector;
[0066] For each of the initial prompt word vectors, cross-mutate the new prompt word vector and the initial prompt word vector to obtain a cross-mutated prompt word vector;
[0067] The cross-variation prompt word vector is converted to obtain multiple cross-variation prompt words.
[0068] Optionally, the second input unit is specifically configured to:
[0069] Acquire a prompt word set; the prompt word set includes multiple prompt words;
[0070] For each of the prompt words, input the prompt word into a large language model to obtain an initial credit rating report;
[0071] Inputting the initial credit rating report and initial rating prompt words into a large language model to obtain a rating result;
[0072] Determining whether the scoring result is consistent with a preset scoring result;
[0073] If the scoring result is consistent with the preset scoring result, the initial scoring prompt word is determined as the scoring prompt word;
[0074] If the scoring result is inconsistent with the preset scoring result, the initial scoring prompt word is corrected based on the initial credit rating report and the scoring result, and after the corrected scoring prompt word is determined as the scoring prompt word, the process returns to the step of inputting the initial credit rating report and the initial scoring prompt word into the large language model to obtain the scoring result.
[0075] The technical solution provided in this application obtains basic information and industry evaluation information of the object to be rated; inputs the basic information and industry evaluation information into a large language model to obtain the rating of the object to be rated; obtains an initial set of prompt words corresponding to the object to be rated; optimizes multiple initial prompt words using a cross-mutation algorithm to obtain multiple cross-mutation prompt words; inputs each cross-mutation prompt word, a scoring prompt word, and the rating of the object to be rated into the large language model to obtain a scoring result for each cross-mutation prompt word; and generates a credit rating report for the object to be rated based on the rating of the object to be rated and the scoring results of each cross-mutation prompt word. In this application, the best prompt word is selected based on the scoring results of the cross-mutation prompt words, and a credit rating report is generated based on the best prompt word and the rating. This method avoids the problem of untimely data updates caused by manual information transmission. At the same time, the cross-mutation prompt words are optimized, thereby improving the generation efficiency while ensuring the quality of the credit rating report. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0077] Figure 1 A flowchart of a method for generating a credit rating report provided in an embodiment of the present application;
[0078] Figure 2 A flowchart of a method for determining an evaluation level provided in an embodiment of the present application;
[0079] Figure 3 A flowchart of a method for obtaining cross-variation prompt words provided in an embodiment of the present application;
[0080] Figure 4 A flowchart of a method for correcting initial scoring prompt words provided in an embodiment of the present application;
[0081] Figure 5 A flowchart of obtaining a prompt word set provided in an embodiment of the present application;
[0082] Figure 6 A schematic diagram of synthetic data provided in an embodiment of the present application;
[0083] Figure 7 A schematic diagram of correcting initial scoring prompt words provided in an embodiment of the present application;
[0084] Figure 8 A flowchart of generating a credit rating report provided in an embodiment of the present application;
[0085] Figure 9 A schematic diagram of an optimized prompt word provided in an embodiment of the present application;
[0086] Figure 10 A flowchart of a method for generating a qualitative analysis of a credit rating report provided in an embodiment of the present application;
[0087] Figure 11 A schematic diagram of the architecture of a credit rating report generating device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0088] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0089] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0090] like Figure 1 FIG. 1 is a flowchart of a method for generating a credit rating report provided in an embodiment of the present application, comprising the following steps:
[0091] S101: Obtain basic information of the object to be rated and industry evaluation information of the object to be rated.
[0092] The objects to be evaluated include but are not limited to: enterprises, and the basic information of the objects to be rated includes but is not limited to: corporate charters, audit reports, and information retrieved from specific software.
[0093] Optionally, the industry evaluation information of the subject to be evaluated indicates information on how the subject to be evaluated is evaluated based on the evaluation criteria for different industry indicators. For example, each enterprise corresponds to a specific industry, such as service enterprises, newly established service enterprises, and new energy enterprises. The subject to be evaluated is evaluated based on the industry to which the enterprise belongs.
[0094] It is understandable that since the basic information of the object to be rated is stored in the vector database, the basic information of the object to be rated is obtained from the vector database through retrieval and recall, and semantic tags are added to the basic information of the object to be rated and converted into a string format.
[0095] In addition, the information on the evaluation of the rated object according to the evaluation standards of different industry indicators (i.e., the industry evaluation information of the rated object) will be stored in a json file. That is to say, when rating a certain enterprise, the json file of the industry to which the company belongs will be read, and the industry evaluation information of the rated object will be organized into a string format through semantic tags.
[0096] S102: Input the basic information and industry evaluation information into the large language model to obtain the evaluation level of the object to be evaluated.
[0097] Among them, the evaluation level of the rated object indicates the evaluation level of one of the qualitative analysis items in the credit rating report.
[0098] Optionally, the LangChain architecture is used to convert the basic information of the object to be rated and the industry evaluation information of the object to be rated to obtain input text, that is, the basic information of the object to be rated and the industry evaluation information of the object to be rated in text format, so that they can be subsequently input into the large language model in text format.
[0099] It should be noted that the large language model is a pre-trained model, which is based on the pre-acquired sample basic information and sample industry evaluation information of the object to be rated as input, and the evaluation levels manually annotated on the sample basic information and sample industry evaluation information as training targets.
[0100] Optionally, in another embodiment of the present application, the specific implementation of step S102 is as follows: Figure 2 As shown, the following steps are included:
[0101] S201: Split the industry evaluation information to obtain multiple preset evaluation levels.
[0102] It is understandable that the industry evaluation information is split. Specifically, a scoring standard is set first; the industry evaluation information is split according to the scoring standard to obtain multiple preset evaluation levels.
[0103] S202: For each preset evaluation level, split the preset evaluation level to obtain multiple evaluation criteria.
[0104] It should be noted that since the large language model has a weak ability to comprehensively analyze the preset evaluation levels, it is necessary to divide the preset evaluation levels into multiple scoring criteria one by one, and then analyze them using multiple evaluation criteria.
[0105] S203: Filter out the highest preset evaluation level from all preset evaluation levels and determine it as the target preset evaluation level.
[0106] It is understandable that whether the basic information meets the preset evaluation level is judged in descending order of evaluation levels. Therefore, the highest preset evaluation level is first screened out from all preset evaluation levels and determined as the target preset evaluation level.
[0107] S204: Input multiple evaluation criteria and basic information corresponding to the target preset evaluation level into the large language model to obtain an evaluation result.
[0108] The evaluation result indicates the amount of basic information that meets the evaluation criteria.
[0109] Optionally, the LangChain architecture can be used to convert multiple evaluation criteria and basic information corresponding to the target preset evaluation level, obtain the input text, and then input the input text into the large language model.
[0110] S205: Calculate the ratio between the number of basic information items that meet the evaluation criteria and the total number of items that meet the evaluation criteria, and determine the ratio as the standard ratio.
[0111] The specific form of calculating the ratio between the number of basic information that meets the evaluation criteria and the total number of evaluation criteria is: standard ratio = number of basic information that meets the evaluation criteria / total number of evaluation criteria.
[0112] For example, the number of basic information that meets the evaluation criteria is 5, and the total number of evaluation criteria is 8. The ratio between the number of basic information that meets the evaluation criteria and the total number of evaluation criteria is: 0.625.
[0113] S206: Determine whether the standard ratio is greater than the preset ratio.
[0114] If the standard ratio is greater than the preset ratio, step S207 is executed; if the standard ratio is not greater than the preset ratio, step S208 is executed.
[0115] Specifically, assuming that the standard ratio is 0.625 and the preset ratio is 0.5, it is determined whether the standard ratio is greater than the preset ratio. Obviously, the standard ratio is greater than the preset ratio, so step S207 is continued.
[0116] Specifically, assuming that the standard ratio is 0.4 and the preset ratio is 0.5, it is determined whether the standard ratio is greater than the preset ratio. Obviously, the standard ratio is greater than the preset ratio, so step S208 is continued.
[0117] S207: Determine the target preset evaluation level as the evaluation level of the object to be evaluated.
[0118] It can be understood that when the standard ratio is greater than the preset ratio, it means that the object to be rated has met the target preset evaluation standard, and the target preset evaluation level is determined as the evaluation level of the object to be evaluated.
[0119] S208: After screening out the highest preset evaluation level except the target preset evaluation level from all preset evaluation levels and determining it as the target preset evaluation level, the process returns to step S204.
[0120] It can be understood that the highest preset evaluation level except the target preset evaluation level is screened out from all preset evaluation levels, that is, the second highest preset evaluation level among all preset evaluation levels, that is, the preset evaluation level is reduced to one level lower than the target preset evaluation level.
[0121] For example, there are three preset evaluation levels, namely the first preset evaluation level, the second preset evaluation level and the third preset evaluation level, the first preset evaluation level > the second preset evaluation level > the third preset evaluation level, and the first preset evaluation level is the target preset evaluation level. The highest preset evaluation level except the target preset evaluation level is screened out from all the preset evaluation levels as the second preset evaluation level.
[0122] S103: Obtain an initial prompt word set corresponding to the object to be evaluated.
[0123] The initial prompt word set includes multiple initial prompt words (including manually designed prompt words and machine-generated prompt words).
[0124] For example, if the technology-based enterprise to be rated is "XYZ Technology Co., Ltd.", the initial prompt word set obtained for XYZ Technology Co., Ltd. includes but is not limited to: Industry risk assessment: According to the methodology of S&P Global Ratings (China), the risk of the industry in which the enterprise is located must be assessed first. For example, if XYZ Company belongs to the technology hardware and semiconductor industry, its industry risk assessment may be marked as "high". Basic information of the enterprise: including the company's registered capital, establishment time, major shareholders and equity structure, etc. Operating status: involving the company's main business, market positioning, competitive position, management team experience, R&D capabilities, etc. Financial status: including key financial indicators such as debt-to-asset ratio, interest repayment ratio, maturity credit repayment ratio, cash flow, etc. Credit record: the company's historical credit performance, including whether there is a record of overdue repayments, legal proceedings, etc.
[0125] S104: Optimizing the multiple initial prompt words using a crossover mutation algorithm to obtain multiple crossover mutation prompt words.
[0126] It is understandable that the crossover-mutation algorithm can introduce randomness, making it possible to explore multiple potential solutions simultaneously in the entire solution space, which helps to find the global optimal solution rather than being limited to the local optimal solution. In other words, the crossover-mutation algorithm can extract useful information from existing, well-performing prompt words and combine it with other prompt words to generate more effective new prompt words, namely crossover-mutation prompt words.
[0127] Optionally, in another embodiment of the present application, the specific implementation of step S104 is as follows: Figure 3 As shown, the following steps are included:
[0128] S301: Convert multiple initial prompt words to obtain multiple initial prompt word vectors.
[0129] The multiple initial prompt word vectors include the current initial prompt word vector.
[0130] It is understandable that by converting multiple initial prompt words into vectors, a distance metric (such as cosine similarity) can be used to evaluate the similarity between different prompt words, which helps to select and generate more relevant combinations.
[0131] S302: Select any two initial prompt word vectors from the multiple initial prompt word vectors and mark them as a first vector and a second vector.
[0132] It is understandable that in the crossover mutation algorithm, selecting two vectors is the basis for generating new prompt words. These two vectors generate a new vector through some crossover method (such as linear combination, feature exchange, etc.), thereby generating a new prompt word.
[0133] S303: Process the difference between the first vector and the second vector using a transformation function to obtain a new vector.
[0134] The specific expression of using the transformation function to process the difference between the first vector and the second vector is: F(bc), where F is the transformation function, b is the first vector, and c is the second vector.
[0135] S304: Add the current initial prompt word vector to the new vector to obtain a new prompt word vector.
[0136] The specific form of adding the current initial prompt word vector and the new vector is shown in formula (1).
[0137] y=a+F(bc)(1)
[0138] In formula (1), a is the current initial prompt word vector, and y is the new prompt word vector.
[0139] S305: For each initial prompt word vector, cross-mutate the new prompt word vector and the initial prompt word vector to obtain a cross-mutated prompt word vector.
[0140] Specifically, the new prompt word vector and the initial prompt word vector are cross-mutated. Specifically, the intersection point between the new prompt word vector and the initial prompt word vector is determined, and one intersection point or multiple intersection points are selected. The new prompt word vector and the initial prompt word vector are cross-operated (for example, the first half of the new prompt word vector and the second half of the initial prompt word vector are combined) to obtain a cross-mutated prompt word.
[0141] S306: Convert the cross-variation prompt word vector to obtain multiple cross-variation prompt words.
[0142] It can be understood that converting the cross-variation cue word vectors into cross-variation cue words, by converting the vectors into specific cue words, can enable the model to better understand and apply these cues, thereby improving the quality of generated or processed text.
[0143] S105: Input each cross-variation prompt word, the scoring prompt word, and the evaluation level of the object to be evaluated into the large language model to obtain the scoring result of each cross-variation prompt word.
[0144] The scoring prompt words are obtained by correcting the initial scoring prompt words.
[0145] It should be noted that the large language model is a pre-trained model, which is based on pre-acquired sample cross-variant words, sample scoring prompt words and sample evaluation levels as input, and uses the manually labeled scoring results of the sample cross-variant words, sample scoring prompt words and sample evaluation levels as training targets.
[0146] Optionally, in another embodiment of the present application, the initial scoring prompt word is corrected to obtain a specific implementation method of the scoring prompt word, such as Figure 4 As shown, the following steps are included:
[0147] S401: Obtain a prompt word set.
[0148] The prompt word set includes multiple prompt words.
[0149] Optionally, in another embodiment of the present application, the specific implementation of step S401 is as follows: Figure 5 As shown, the following steps are included:
[0150] S501: Acquire real data and expand prompt words.
[0151] The real data indicates the basic information of the real object to be evaluated, and the expanded prompt word indicates the prompt word for expanding the basic information according to the format of the basic information.
[0152] For example, the basic information of the rating object is as follows: Basic information of listed company A: (1) Company annual report (2) Company audit report (3) Company shareholder personal information (4) Other relevant information...
[0153] For example, the expanded prompt is: You are a large language model assistant. You need to generate challenging samples for each task based on the provided task_description and instructions. Generate {num_samples) challenging samples for the following tasks. The generated samples should be challenging and diverse.
[0154] [task_description}
[0155] {instruction)
[0156] Answer in the following format:
[0157] Sample 1:
[0158] <text>
[0159] Sample 2:
[0160] <text>
[0161] …
[0162] S502: Input the real data and the expanded prompt words into the large language model to obtain synthetic data.
[0163] The format of the synthetic data is consistent with that of the real data.
[0164] Understandably, since customer credit rating reports primarily target larger companies, the amount of data available in datasets is typically limited. Similarly, there is a lack of large-scale, standardized, and publicly available datasets on the internet. A key function of the Agent is data augmentation, paving the way for subsequent prompt word optimization. This data augmentation is achieved by leveraging a large language model specifically designed for data synthesis. Based on a well-designed augmented prompt word set and a small amount of real data, the LangChain architecture, large language model, and data augmentation capabilities dynamically generate a batch of machine-generated synthetic data that is similar to real data.
[0165] For example, see Figure 6 By inputting the basic information and expanded prompt words of the object to be rated into the large language model LLM, the synthetic data obtained includes: basic information of listed company B, basic information of listed company C, basic information of listed company D, basic information of listed company E, basic information of listed company F, basic information of listed company G, basic information of listed company II, basic information of listed company I, etc.
[0166] S503: Generate a prompt word set corresponding to the synthesized data using a preset evaluation function.
[0167] The preset evaluation function includes but is not limited to: generating an evaluation function for the prompt word activated by the Agent.
[0168] S402: For each prompt word, input the prompt word into the large language model to obtain an initial credit rating report.
[0169] Among them, the initial credit rating report includes the reasons for the evaluation.
[0170] It should be noted that the large language model is a pre-trained model, which is based on pre-acquired sample prompt words as input and uses the initial credit rating report manually annotated with the sample prompt words as the training target.
[0171] S403: Input the initial credit rating report and the initial rating prompt words into the large language model to obtain a rating result.
[0172] Among them, the initial scoring prompt words are instructions used to guide the model to evaluate and score specific content.
[0173] It is understandable that the scoring results are divided into five levels based on the quality of the initial credit rating report, from 1 to 5, with higher scores representing higher quality of the generated results.
[0174] It should be noted that the large language model is a pre-trained model, which is based on pre-acquired sample credit rating reports and initial scoring prompt words as input, and the scoring results manually annotated on the sample credit rating reports and initial scoring prompt words as training targets.
[0175] In addition, all the large language models mentioned above can be the same or different large language models, and different results can be obtained according to different inputs.
[0176] S404: Determine whether the scoring result is consistent with the preset scoring result.
[0177] If the scoring result is consistent with the preset scoring result, step S405 is executed; if the scoring result is inconsistent with the preset scoring result, step S406 is executed.
[0178] Optionally, the preset scoring result can be a result of manual scoring of the initial credit rating report. Manual scoring can be performed on the front-end page using Argilla (an open source data management and annotation tool), which is the preset scoring result.
[0179] S405: Determine the initial scoring prompt word as the scoring prompt word.
[0180] It is understandable that if the scoring result is consistent with the preset scoring result, it means that the initial scoring prompt word at this time is a scoring word that meets human needs, and then the initial scoring prompt word is determined as the scoring prompt word.
[0181] S406: Correct the initial scoring prompt words based on the initial credit rating report and the scoring result, and determine the corrected scoring prompt words as the scoring prompt words, then return to step S403.
[0182] It is understandable that if the scoring result is inconsistent with the preset scoring result, it means that the initial scoring prompt word at this time is a scoring word that meets human needs. It is necessary to correct the initial scoring prompt word based on the initial credit rating report and the scoring result, and after the corrected scoring prompt word is determined as the scoring prompt word, return to execute step S403.
[0183] For example, if the score result of the initial credit rating report is 3 points and the preset score result is 5 points, then the initial score prompt word needs to be corrected based on the initial credit rating report and the score result so that the score result is adjusted to 5 points.
[0184] To better explain Figure 4 The following content is explained in conjunction with specific application scenarios. Figure 7 The prompt word set indicates the initial scoring prompt words, and the data sample set indicates the initial credit rating report. By inputting the prompt word set and the data sample set into the large language model LLM, the scoring result can be obtained, and the initial scoring prompt words are corrected according to the scoring result.
[0185] S106: Generate a credit rating report for the object to be rated based on the rating level of the object to be rated and the scoring result of each cross-variation prompt word.
[0186] Among them, based on the evaluation grade of the object to be evaluated and the scoring results of each cross-variation prompt word, a credit rating report of the object to be rated is generated. Specifically, the rating reasons are determined based on the evaluation grade of the object to be evaluated and the scoring results of each cross-variation prompt word, and the credit rating report of the object to be rated is generated based on the evaluation reasons.
[0187] Optionally, in another embodiment of the present application, the specific implementation of step S106 is as follows: Figure 8 As shown, the following steps are included:
[0188] S801: Filter out the scoring result with the highest score from all scoring results, and determine the cross-variation prompt word corresponding to the scoring result with the highest score as the best prompt word.
[0189] The best prompt word indicates the best prompt word for each sub-item in the qualitative analysis.
[0190] For example, there are three cross-mutation prompt words, namely the first cross-mutation prompt word, the second cross-mutation prompt word and the third cross-mutation prompt word. The scoring result of the first cross-mutation prompt word is 2 points, the scoring result of the second cross-mutation prompt word is 4 points, and the scoring result of the third cross-mutation prompt word is 3 points. The scoring result with the highest score is screened out from all the scoring results as the scoring result of the second cross-mutation prompt word, and the third cross-mutation prompt word is determined as the best prompt word.
[0191] It should be noted that all companies in the industry can use the best prompt word as the standard, and there is no need to optimize the prompt word again. If it is a different industry, it is necessary to re-optimize the prompt word (i.e. Figure 4 Figure 5 content shown).
[0192] S802: Generate evaluation reasons based on the evaluation level of the object to be evaluated and the best prompt word.
[0193] Optionally, the Langchain framework and the large language model can be used to process the evaluation level and the best prompt words of the evaluation object to obtain the rating reasons for each sub-item.
[0194] S803: Generate a credit rating report for the object to be rated based on the rating level and evaluation reasons of the object to be rated.
[0195] It can be understood that a tree-like json structure is first generated based on the evaluation level and evaluation reasons of the object to be evaluated, and then the json2doc algorithm is used to generate a doc document (i.e., the credit rating report of the object to be rated).
[0196] For a better explanation of all the above, see Figure 9 The schematic diagram of the optimized prompt words is shown. First, manually created expanded prompt words (i.e., prompts) are obtained and added to real data. Synthetic data is generated based on the expanded prompt words and real data. A prompt word set is determined based on the synthetic data. Scoring prompt words are determined based on the prompt word set and the initial scoring prompt words (i.e., a large language model that conforms to human preferences). An initial prompt word set corresponding to the object to be evaluated is obtained. The multiple initial prompt words are optimized using a crossover mutation algorithm to obtain multiple crossover mutation prompt words. Each crossover mutation prompt word, scoring prompt word, and the evaluation grade of the object to be evaluated are input into the large language model to obtain a scoring result for each crossover mutation prompt word. The scoring result with the highest score is screened from all scoring results, and the crossover mutation prompt word corresponding to the scoring result with the highest score is determined as the optimal prompt word.
[0197] In addition, in order to Figure 1 For a brief description of the contents shown, see Figure 10 The flowchart of the method for generating a qualitative analysis of a credit rating report shown in the figure mainly includes two stages. In the first stage, the basic information of the enterprise and industry evaluation information are input into the large language model to obtain the evaluation grade. In the second stage, the prompt words are first optimized to obtain the best prompt words, and then the credit rating report (i.e., a complete qualitative analysis) is generated based on the best prompt words and the evaluation grade.
[0198] In summary, the optimal prompt word is selected based on the cross-variation prompt word scoring results, and a credit evaluation report is generated based on the optimal prompt word and evaluation grade. This method avoids the problem of delayed data updates caused by manual information transmission. Furthermore, the cross-variation prompt word is optimized, thereby improving the efficiency of credit evaluation report generation while ensuring the quality of the report.
[0199] like Figure 11 As shown, it is a schematic diagram of the architecture of a credit rating report generation device provided in an embodiment of the present application. The generation device includes: a first acquisition unit 100, a first input unit 200, a second acquisition unit 300, an optimization unit 400, a second input unit 500 and a generation unit 600.
[0200] The first acquiring unit 100 is configured to acquire basic information of the object to be rated and industry evaluation information of the object to be rated.
[0201] The first input unit 200 is used to input basic information and industry evaluation information into the large language model to obtain the evaluation level of the object to be evaluated.
[0202] The first input unit 200 is specifically used to: split the industry evaluation information to obtain multiple preset evaluation levels; for each preset evaluation level, split the preset evaluation level to obtain multiple evaluation criteria; screen out the highest preset evaluation level from all preset evaluation levels, and determine it as the target preset evaluation level; input multiple evaluation criteria and basic information corresponding to the target preset evaluation level into the large language model to obtain an evaluation result; the evaluation result indicates the number of basic information that meets the evaluation criteria; calculate the ratio between the number of basic information that meets the evaluation criteria and the total number of evaluation criteria, and determine it as the standard ratio; determine whether the standard ratio is greater than the preset ratio; if the standard ratio is greater than the preset ratio, determine the target preset evaluation level as the evaluation level of the object to be evaluated; if the standard ratio is not greater than the preset ratio, screen out the highest preset evaluation level except the target preset evaluation level from all preset evaluation levels, and after determining it as the target preset evaluation level, return to the step of inputting multiple evaluation criteria and basic information corresponding to the target preset evaluation level into the large language model to obtain the evaluation result.
[0203] The second acquisition unit 300 is used to acquire an initial prompt word set corresponding to the object to be evaluated; the initial prompt word set includes multiple initial prompt words.
[0204] The optimization unit 400 is configured to optimize the multiple initial prompt words using a crossover mutation algorithm to obtain multiple crossover mutation prompt words.
[0205] The optimization unit 400 is specifically used to: convert multiple initial prompt words to obtain multiple initial prompt word vectors; the multiple initial prompt word vectors include the current initial prompt word vector; select any two initial prompt word vectors from the multiple initial prompt word vectors and mark them as the first vector and the second vector; use the transformation function to process the difference between the first vector and the second vector to obtain a new vector; add the current initial prompt word vector and the new vector to obtain a new prompt word vector; for each initial prompt word vector, cross-mutate the new prompt word vector and the initial prompt word vector to obtain a cross-mutated prompt word vector; convert the cross-mutated prompt word vector to obtain multiple cross-mutated prompt words.
[0206] The second input unit 500 is used to input each cross-variation prompt word, the scoring prompt word and the evaluation level of the object to be evaluated into the large language model to obtain the scoring result of each cross-variation prompt word; the scoring prompt word is obtained by correcting the initial scoring prompt word.
[0207] The second input unit 500 is specifically used to: obtain a prompt word set; the prompt word set includes multiple prompt words; for each prompt word, input the prompt word into the large language model to obtain an initial credit rating report; input the initial credit rating report and the initial scoring prompt word into the large language model to obtain a scoring result; determine whether the scoring result is consistent with the preset scoring result; if the scoring result is consistent with the preset scoring result, determine the initial scoring prompt word as the scoring prompt word; if the scoring result is inconsistent with the preset scoring result, correct the initial scoring prompt word based on the initial credit rating report and the scoring result, and after determining the corrected scoring prompt word as the scoring prompt word, return to execute the step of inputting the initial credit rating report and the initial scoring prompt word into the large language model to obtain the scoring result.
[0208] The second input unit 500 is specifically used to: obtain real data and expanded prompt words; wherein the real data indicates the basic information of the real object to be evaluated, and the expanded prompt words indicate the prompt words that expand the basic information according to the format of the basic information; input the real data and the expanded prompt words into the large language model to obtain synthetic data, wherein the format of the synthetic data is consistent with that of the real data; and use the preset evaluation function to generate a prompt word set corresponding to the synthetic data.
[0209] The generating unit 600 is configured to generate a credit rating report for the object to be rated based on the rating level of the object to be rated and the scoring result of each cross-variation prompt word.
[0210] The generation unit 600 is specifically used to: filter out the scoring result with the highest score from all scoring results, and determine the cross-variation prompt word corresponding to the scoring result with the highest score as the best prompt word; generate evaluation reasons based on the evaluation level of the object to be evaluated and the best prompt word; and generate a credit rating report for the object to be rated based on the evaluation level and evaluation reasons of the object to be evaluated.
[0211] In summary, the optimal prompt word is selected based on the cross-variation prompt word scoring results, and a credit evaluation report is generated based on the optimal prompt word and evaluation grade. This method avoids the problem of delayed data updates caused by manual information transmission. Furthermore, the cross-variation prompt word is optimized, thereby improving the efficiency of credit evaluation report generation while ensuring the quality of the report.
[0212] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Ordinary technicians in this field can understand and implement it without expending creative work.
[0213] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0214] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating a credit rating report, characterized in that: include: Obtain basic information of the subject to be rated and industry evaluation information of the subject to be rated; Inputting the basic information and the industry evaluation information into a large language model to obtain an evaluation grade of the object to be evaluated; The step of inputting the basic information and the industry evaluation information into a large language model to obtain the evaluation grade of the object to be evaluated includes: Splitting the industry evaluation information to obtain multiple preset evaluation levels; For each of the preset evaluation levels, split the preset evaluation level to obtain multiple evaluation criteria; Filter out the highest preset evaluation level from all preset evaluation levels and determine it as the target preset evaluation level; Inputting a plurality of evaluation criteria corresponding to the target preset evaluation level and the basic information into a large language model to obtain an evaluation result; the evaluation result indicates the amount of the basic information that meets the evaluation criteria; Calculate the ratio between the number of basic information that meets the evaluation criteria and the total number of evaluation criteria, and determine it as the standard ratio; Determining whether the standard ratio is greater than a preset ratio; If the standard ratio is greater than the preset ratio, the target preset evaluation level is determined as the evaluation level of the object to be evaluated; If the standard ratio is not greater than the preset ratio, the highest preset evaluation level other than the target preset evaluation level is screened out from all the preset evaluation levels, and after determining it as the target preset evaluation level, the process returns to the step of inputting the multiple evaluation criteria corresponding to the target preset evaluation level and the basic information into the large language model to obtain an evaluation result; Obtaining an initial prompt word set corresponding to the object to be evaluated; the initial prompt word set includes multiple initial prompt words; Optimizing the plurality of initial prompt words using a crossover mutation algorithm to obtain a plurality of crossover mutation prompt words; The cross-mutation algorithm is used to perform cross-mutation on the multiple initial prompt words to obtain multiple cross-mutation prompt words, including: Converting multiple initial prompt words to obtain multiple initial prompt word vectors; the multiple initial prompt word vectors include the current initial prompt word vector; Selecting any two initial prompt word vectors from the plurality of initial prompt word vectors and marking them as a first vector and a second vector; Processing the difference between the first vector and the second vector using a transformation function to obtain a new vector; Adding the current initial prompt word vector to the new vector to obtain a new prompt word vector; For each of the initial prompt word vectors, cross-mutate the new prompt word vector and the initial prompt word vector to obtain a cross-mutated prompt word vector; Converting the cross-variation prompt word vector to obtain multiple cross-variation prompt words; Inputting each of the cross-variation prompt words, the scoring prompt words, and the evaluation grade of the object to be evaluated into a large language model to obtain a scoring result for each of the cross-variation prompt words; the scoring prompt words are obtained by correcting the initial scoring prompt words; A credit rating report for the object to be rated is generated based on the rating grade of the object to be rated and the scoring result of each cross-variation prompt word.
2. The method according to claim 1, characterized in that The process of correcting the initial scoring prompt word to obtain the scoring prompt word includes: Acquire a prompt word set; the prompt word set includes multiple prompt words; For each of the prompt words, input the prompt word into a large language model to obtain an initial credit rating report; Inputting the initial credit rating report and initial rating prompt words into a large language model to obtain a rating result; Determining whether the scoring result is consistent with a preset scoring result; If the scoring result is consistent with the preset scoring result, the initial scoring prompt word is determined as the scoring prompt word; If the scoring result is inconsistent with the preset scoring result, the initial scoring prompt word is corrected based on the initial credit rating report and the scoring result, and after the corrected scoring prompt word is determined as the scoring prompt word, the process returns to the step of inputting the initial credit rating report and the initial scoring prompt word into the large language model to obtain the scoring result.
3. The method according to claim 2, characterized in that The step of obtaining a prompt word set includes: Acquire real data and expanded prompt words; wherein the real data indicates the basic information of the real object to be evaluated, and the expanded prompt words indicate prompt words that expand the basic information according to the format of the basic information; Inputting the real data and the expanded prompt word into the large language model to obtain synthesized data, wherein the format of the synthesized data is consistent with that of the real data; A preset evaluation function is used to generate a prompt word set corresponding to the synthesized data.
4. The method according to claim 1, wherein Generating a credit rating report for the object to be rated based on the evaluation grade of the object to be rated and the scoring result of each cross-variation prompt word includes: Screening out the scoring result with the highest score from all the scoring results, and determining the cross-variation prompt word corresponding to the scoring result with the highest score as the best prompt word; generating evaluation reasons based on the evaluation level of the object to be evaluated and the optimal prompt word; A credit rating report for the object to be rated is generated based on the rating level of the object to be rated and the rating reasons.
5. A credit rating report generating device, characterized in that: include: The first acquisition unit is used to acquire basic information of the object to be rated and industry evaluation information of the object to be rated; A first input unit is used to input the basic information and the industry evaluation information into a large language model to obtain an evaluation level of the object to be evaluated; Wherein, the first input unit is specifically used for: Splitting the industry evaluation information to obtain multiple preset evaluation levels; For each of the preset evaluation levels, split the preset evaluation level to obtain multiple evaluation criteria; Filter out the highest preset evaluation level from all preset evaluation levels and determine it as the target preset evaluation level; Inputting a plurality of evaluation criteria corresponding to the target preset evaluation level and the basic information into a large language model to obtain an evaluation result; the evaluation result indicates the amount of the basic information that meets the evaluation criteria; Calculate the ratio between the number of basic information that meets the evaluation criteria and the total number of evaluation criteria, and determine it as the standard ratio; Determining whether the standard ratio is greater than a preset ratio; If the standard ratio is greater than the preset ratio, the target preset evaluation level is determined as the evaluation level of the object to be evaluated; If the standard ratio is not greater than the preset ratio, the highest preset evaluation level other than the target preset evaluation level is screened out from all the preset evaluation levels, and after determining it as the target preset evaluation level, the process returns to the step of inputting the multiple evaluation criteria corresponding to the target preset evaluation level and the basic information into the large language model to obtain an evaluation result; A second acquisition unit is configured to acquire an initial prompt word set corresponding to the object to be evaluated; the initial prompt word set includes a plurality of initial prompt words; an optimization unit, configured to optimize the plurality of initial prompt words using a crossover mutation algorithm to obtain a plurality of crossover mutation prompt words; The optimization unit is specifically used for: Converting multiple initial prompt words to obtain multiple initial prompt word vectors; the multiple initial prompt word vectors include the current initial prompt word vector; Selecting any two initial prompt word vectors from the plurality of initial prompt word vectors and marking them as a first vector and a second vector; Processing the difference between the first vector and the second vector using a transformation function to obtain a new vector; Adding the current initial prompt word vector to the new vector to obtain a new prompt word vector; For each of the initial prompt word vectors, cross-mutate the new prompt word vector and the initial prompt word vector to obtain a cross-mutated prompt word vector; Converting the cross-variation prompt word vector to obtain multiple cross-variation prompt words; A second input unit is configured to input each of the cross-variation prompt words, the scoring prompt words, and the evaluation level of the object to be evaluated into a large language model to obtain a scoring result for each of the cross-variation prompt words; the scoring prompt words are obtained by correcting the initial scoring prompt words; A generating unit is configured to generate a credit rating report for the object to be rated based on the rating grade of the object to be rated and the scoring result of each cross-variation prompt word.
6. The device according to claim 5, characterized in that The second input unit is specifically used for: Acquire a prompt word set; the prompt word set includes multiple prompt words; For each of the prompt words, input the prompt word into a large language model to obtain an initial credit rating report; Inputting the initial credit rating report and initial rating prompt words into a large language model to obtain a rating result; Determining whether the scoring result is consistent with a preset scoring result; If the scoring result is consistent with the preset scoring result, the initial scoring prompt word is determined as the scoring prompt word; If the scoring result is inconsistent with the preset scoring result, the initial scoring prompt word is corrected based on the initial credit rating report and the scoring result, and after the corrected scoring prompt word is determined as the scoring prompt word, the process returns to the step of inputting the initial credit rating report and the initial scoring prompt word into the large language model to obtain the scoring result.
Citation Information
Patent Citations
Multi-level visual evaluation report generation method and device based on large language model
CN117194637A
Data generation method and device
CN117851585A