A method and system for generating a credit investigation report based on a large language model
Through a large language model, the fraud risks in corporate voice data are identified, combined with voice sentiment and financial data, differentiated credit survey reports are generated, which solves the problem of cumbersome and inefficient generation of credit survey reports and improves the generation efficiency and accuracy.
Patent Information
- Application Number
- CN202510541283.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-28
AI Technical Summary
In the field of micro-credit, the generation process of credit investigation reports is cumbersome, the workload is large, and it is difficult to effectively identify the risk of fraud in voice data, resulting in inefficient generation.
A large language model is adopted to identify matching voice periods during the enterprise business survey process, combine voice sentiment and financial data, and filter out risk voice periods to generate differentiated credit survey reports.
It improves the efficiency of generating credit investigation reports, reduces the load on the server when the fraud risk is high, and ensures the accuracy and generation efficiency of the report.
Smart Images

Figure CN120067304B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large models, and particularly relates to a method and system for generating a credit investigation report based on a large language model. Background Art
[0002] In the field of small and micro credit, in order to conduct the credit investigation process of small and micro credit reports, it is cumbersome and variable, with high requirements for the personal capabilities of credit officers. The editing and tabulation processes of credit reports are cumbersome and involve a large amount of work. Therefore, how to improve the accuracy and compilation efficiency of credit investigation reports has become an urgent technical problem to be solved.
[0003] To solve the above technical problems, in the invention patent application CN202410158795.2, "Method and System for Generating an Interactive Generative Financial Investigation Report Based on a Large Model", a large model is used to achieve the interactive and automated generation of due diligence investigation reports in the financial field, greatly reducing the working hours of financial practitioners and improving the efficiency and quality of due diligence investigations. However, there are the following technical defects:
[0004] During the process of generating a credit report, it is necessary to comprehensively generate and process data from multiple dimensions of an enterprise, specifically including tax data, enterprise turnover, electricity consumption data, and voice data during the business investigation process, etc., which need to be comprehensively considered. The difference in fraud risk may be manifested in the voice data during the business investigation process. Therefore, how to form a differential processing method for generating a credit investigation report based on the recognition and processing results of voice data during the business investigation process, and then improve the processing efficiency of credit investigation reports has become an urgent technical problem to be solved.
[0005] The present application provides a method and system for generating a credit investigation report based on a large language model. Summary of the Invention
[0006] To achieve the purpose of the present invention, the present invention adopts the following technical solutions:
[0007] Specifically, in the first aspect, the present application provides a method for generating a credit investigation report based on a large language model, specifically including:
[0008] S1 Based on the type of the enterprise, determine the credit-related keywords of the enterprise, and based on the credit-related keywords, determine the matching voice segments during the business investigation process of the enterprise;
[0009] S2 Based on the matching data between different matching voice segments and the credit-related keywords and the distribution data in different matching voice segments, when the credibility of the matching voice segments can meet the requirements, proceed to the next step;
[0010] S3 Obtain the emotion recognition results of the voices in different matching voice time periods, and combine the matching situation between the voice data in different matching voice time periods and the financial data of the enterprise to determine the risky voice time periods in the matching voice time periods;
[0011] S4 Determine the generation processing method of the credit investigation report of the enterprise using the large language model according to the distribution data of different risky voice time periods and the matching situation with the voice data of other matching voice time periods.
[0012] The beneficial effects of the present invention are as follows:
[0013] By using the emotion recognition results of the voices in different matching voice time periods and the matching situation between the voice data in different matching voice time periods and the financial data of the enterprise to determine the risky voice time periods in the matching voice time periods, not only the differences in the manifestations of fear or anxiety caused by fraud in the matching voice time periods are considered singly, but also the differences in the probabilities of fraud risks caused by the deviations from the financial data are considered, realizing the screening of the risky voice time periods with relatively high probabilities of fraud risks.
[0014] Determine the generation processing method of the credit investigation report of the enterprise using the large language model according to the distribution data of different risky voice time periods and the matching situation with the voice data of other matching voice time periods, thereby realizing the determination of the generation processing method of the credit investigation report from two perspectives: the distribution dispersion of the risky voice time periods with fraud risks and the differences in fraud risks caused by the deviations of other matching voice time periods in data such as financial indicators. This not only ensures the efficiency of the generation processing of the credit investigation report in the case of relatively low fraud risks, but also reduces the impact on the server load caused by the immediate generation processing of the credit investigation report in the case of relatively high fraud risks.
[0015] A further technical solution is that the credit-related keywords of the enterprise include electricity consumption data, business indicators, and insurance participation indicators.
[0016] A further technical solution is that the method for determining the matching voice time periods in the business investigation process of the enterprise is as follows:
[0017] Based on the credit-related keywords, determine the matching quantities of different time periods in the voice investigation data in the business investigation process with different credit-related keywords;
[0018] Based on the matching quantities with different credit-related keywords, determine the matching keywords in the time periods;
[0019] Determine whether the time period is a matching voice time period in the business investigation process of the enterprise according to the quantity of the matching keywords.
[0020] A further technical solution is that the matching keyword is a credit-related keyword with a matching quantity greater than a preset matching quantity threshold.
[0021] A further technical solution is that when the quantity of the matching keywords is greater than a preset matching keyword quantity threshold, it is determined that the time period is a matching voice time period during the business investigation process of the enterprise. A further technical solution is that the method for the enterprise to determine the generation processing method of the credit investigation report by using a large language model is as follows:
[0022] Based on the matching situation of the voice data between different risk voice time periods and other matching voice time periods, determine the quantity of the matching voice time periods inconsistent with the financial indicators of the risk voice time period. Determine the voice deviation time period in the risk voice time period according to the quantity of the matching voice time periods inconsistent with the financial indicators of the risk voice time period. Determine the voice deviation coefficient of different risk voice time periods according to the proportion of the voice deviation time periods in different risk voice time periods in the quantity of the matching voice time periods.
[0023] Divide the voice investigation data into different voice time periods according to the unit duration. Determine the proportion of the quantity of the risk voice time periods in different voice time periods according to the distribution data of the risk voice time periods.
[0024] Based on the proportion of the quantity of the risk voice time periods in different voice time periods and the voice deviation coefficient of different risk voice time periods, determine the generation processing method of the credit investigation report by the enterprise using a large language model.
[0025] A further technical solution is that based on the proportion of the quantity of the risk voice time periods in different voice time periods and the voice deviation coefficient of different risk voice time periods, determine the generation processing method of the credit investigation report by the enterprise using a large language model, which specifically includes:
[0026] Take the voice time period with the proportion of the quantity of the risk voice time period greater than the preset risk time period quantity proportion threshold as the fraud risk time period. When the quantity of the fraud risk time periods is not within the preset voice time period quantity range, there is no need to generate a credit investigation report, and it is determined that the enterprise has a fraud risk.
[0027] When the quantity of the fraud risk time periods is within the preset voice time period quantity range, it is also necessary to determine whether the average value of the voice deviation coefficients of different risk voice time periods is within the preset deviation coefficient range. If so, after parsing all the voice investigation data, determine whether it is necessary to perform the generation processing of the credit investigation report according to the parsing result. If not, during the analysis and processing of the voice investigation data, use a large language model to perform the synchronous generation processing of the credit investigation report.
[0028] A further technical solution lies in determining whether it is necessary to generate a credit investigation report according to the parsing result, specifically including:
[0029] When the proportion of statements with fear or anxiety emotions in the voice survey data is greater than the preset proportion of the number of abnormal emotion statements, it is determined that the parsing result indicates a fraud risk, and it is not necessary to generate a credit investigation report;
[0030] When the proportion of statements with fear or anxiety emotions in the voice survey data is not greater than the preset proportion of the number of abnormal emotion statements, it is determined that the parsing result indicates no fraud risk, and a credit investigation report is generated using a large language model.
[0031] In a second aspect, the present invention provides a computer system, including: a memory and a processor connected by communication, and a computer program stored on the memory and capable of running on the processor. When the processor runs the computer program, it executes the above-mentioned method for generating a credit investigation report based on a large language model.
[0032] Other features and advantages will be described in the subsequent specification. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification and the drawings.
[0033] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides detailed descriptions as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] By referring to the accompanying drawings and describing its exemplary embodiments in detail, the above and other features and advantages of the present invention will become more obvious.
[0035] Figure 1 is a flowchart of a method for generating a credit investigation report based on a large language model;
[0036] Figure 2 is a flowchart of a method for determining a matching voice period in the business investigation process of an enterprise;
[0037] Figure 3 is a flowchart for determining that the credibility of the matching voice period can meet the requirements;
[0038] Figure 4 is a flowchart of a method for determining a risk voice period in the matching voice period;
[0039] Figure 5 is a framework diagram of a computer system. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0041] In this application, by using the fraud risk of enterprises in the surveyed voice data, a differentiated generation processing method for credit investigation reports is determined. Specifically, when the fraud risk is relatively high, after the surveyed voice data is analyzed and processed, if it is determined that there is no fraud risk, the generation of the credit investigation report is carried out. When the fraud risk is relatively low, the generation of the credit investigation report is carried out during the parsing and processing of the surveyed voice data, thereby improving the efficiency of the generation processing of the credit investigation report.
[0042] Match voice time periods with the number of credit-related keywords being more than 10.
[0043] Take the credit-related keywords with the number of associated matching voice time periods being more than 3 as voice-associated keywords, and determine the credibility coefficient based on the average value of the total proportion of the duration of different matching voice time periods in the voice survey data and the proportion of the number of voice-associated keywords in the number of credit-related keywords. When the credibility coefficient is greater than 0.4, it is determined that the credibility degree of the matching voice time period can meet the requirements.
[0044] Determine the proportion of the number of sentences with fear or anxiety emotions in different matching voice time periods based on the emotion recognition results of the voices in different matching voice time periods, and use it as the proportion of abnormal sentence numbers. Determine the proportion of the number of financial indicators inconsistent with the financial data based on the matching situation between the voice data in the matching voice time periods and the financial data of the enterprise, and use it as the proportion of deviation indicator numbers;
[0045] Based on the average value of the proportion of abnormal sentence numbers and the proportion of deviation indicator numbers, determine the voice risk coefficient of the matching voice time period. When the voice risk coefficient is greater than 0.4, it is determined that the matching voice time period is a risk voice time period.
[0046] When the number of matching voice time periods inconsistent with the financial indicators of the risk voice time period is more than 2, take it as a voice deviation time period. According to the proportion of the voice deviation time periods in different risk voice time periods in the number of matching voice time periods, determine the voice deviation coefficient of different risk voice time periods. Divide the voice survey data into different voice time periods according to the unit duration, and determine the proportion of the number of risk voice time periods in different voice time periods based on the distribution data of the risk voice time periods;
[0047] Based on the proportion of the number of risky speech segments in different speech time periods and the speech deviation coefficients of different risky speech segments, determine the generation processing method of the credit investigation report by the enterprise using the large language model.
[0048] Using the large model to generate the credit investigation report mainly includes:
[0049] 1. Record the business investigation process;
[0050] 2. Convert the language into text;
[0051] 3. Perform large language model prompt engineering processing on the text, and mark the non-compliant content;
[0052] 4. Perform large language model prompt engineering processing on the investigation content, and refine the business attribute information and attribute data;
[0053] 5. Follow up the attribute information and attribute data, and generate the report in combination with the user's financial indicator data.
[0054] Embodiment 1
[0055] As Figure 1 shown, the present application provides a method for generating a credit investigation report based on a large language model, specifically including:
[0056] S1 Based on the type of the enterprise, determine the credit-related keywords of the enterprise, and based on the credit-related keywords, determine the matching speech segments in the business investigation process of the enterprise;
[0057] Further, the credit-related keywords of the enterprise include electricity consumption data, business indicators, and insurance participation indicators.
[0058] Specifically, as Figure 2 shown, the method for determining the matching speech segments in the business investigation process of the enterprise is:
[0059] Based on the credit-related keywords, determine the matching quantity of different time periods in the speech investigation data in the business investigation process with different credit-related keywords;
[0060] Based on the matching quantity with different credit-related keywords, determine the matching keywords in the time period;
[0061] According to the quantity of the matching keywords, determine whether the time period is the matching speech segment in the business investigation process of the enterprise.
[0062] Further, the matching keyword is a credit-related keyword with a matching quantity greater than a preset matching quantity threshold.
[0063] It should be noted that when the quantity of the matching keywords is greater than the preset matching keyword quantity threshold, it is determined that the time period is a matching voice time period during the business investigation process of the enterprise.
[0064] In another possible embodiment, the method for determining the matching voice time period during the business investigation process of the enterprise is as follows:
[0065] Based on the credit-related keyword, determine the matching quantities of different time periods in the voice investigation data during the business investigation process with different credit-related keywords;
[0066] Based on the matching quantities with different credit-related keywords, determine the sentences in the time period that have credit-related keywords as associated sentences;
[0067] Determine whether the time period is a matching voice time period during the business investigation process of the enterprise according to the quantity of the associated sentences.
[0068] Specifically, when the quantity of the associated sentences is greater than the preset associated sentence quantity threshold, it is determined that the time period is a matching voice time period during the business investigation process of the enterprise.
[0069] In another possible embodiment, the method for determining the matching voice time period during the business investigation process of the enterprise is as follows:
[0070] Based on the credit-related keyword, determine the matching quantities of different time periods in the voice investigation data during the business investigation process with different credit-related keywords. When the total matching quantity of the time period with different credit-related keywords is greater than the preset matching quantity threshold, it is determined that the time period is a matching voice time period;
[0071] When the total matching quantity of the time period with different credit-related keywords is not greater than the preset matching quantity threshold:
[0072] Based on the matching quantities with different credit-related keywords, determine the matching keywords in the time period. When there are no matching keywords in the time period, it is determined that the time period does not belong to the matching voice time period;
[0073] When there are matching keywords in the time period:
[0074] Obtain the quantity of the matching keywords in the time period. When the quantity of the matching keywords in the time period meets the requirements, it is determined that the time period belongs to the matching voice time period;
[0075] When the number of matching keywords in the time period does not meet the requirements: Based on the number of matches with different credit-related keywords, determine the sentences of the surveyed object with credit-related keywords in the time period, and use them as associated sentences. When the number of associated sentences in the time period meets the requirements, it is determined that the time period belongs to the matching voice time period;
[0076] When the number of associated sentences in the time period meets the requirements:
[0077] Based on the number of associated sentences and the number of matches of different credit-related feature words, determine the correlation coefficients of different credit-related feature words. When the number of credit-related feature words with correlation coefficients greater than the preset correlation coefficient threshold meets the requirements, it is determined that the time period belongs to the matching voice time period;
[0078] When the number of credit-related feature words with correlation coefficients greater than the preset correlation coefficient threshold does not meet the requirements:
[0079] Based on the correlation coefficients of credit-related feature words of different credit-related keywords, determine the correlation coefficient of feature words in the time period, and determine whether the time period is the matching voice time period in the business survey process of the enterprise according to the correlation coefficient of feature words.
[0080] It should be noted that when the correlation coefficient of feature words in the time period is greater than the preset correlation coefficient threshold of feature words, it is determined that the time period is the matching voice time period in the business survey process of the enterprise.
[0081] S2 Based on the matching data of different matching voice time periods and the credit-related keywords, and based on the distribution data in different matching voice time periods, when it is determined that the credibility of the matching voice time period can meet the requirements, proceed to the next step;
[0082] Specifically, as Figure 3 shown, determining that the credibility of the matching voice time period can meet the requirements specifically includes:
[0083] Based on the matching data of different matching voice time periods and the credit-related keywords, determine the number of matching voice time periods of different credit-related keywords, and determine the voice-related keywords in the credit-related keywords based on the number of the matching voice time periods;
[0084] Based on the distribution data of different matching voice time periods, determine the total duration proportion of different matching voice time periods in the voice survey data;
[0085] Based on the total duration proportion and the proportion of the number of voice-related keywords in the number of credit-related keywords, determine the credibility coefficient, and determine whether the credibility of the matching voice time period can meet the requirements based on the credibility coefficient.
[0086] Further, the voice-associated keyword is a credit-associated keyword for which the number of voice time periods that match is greater than a preset threshold for the number of voice time periods to be matched.
[0087] It should be noted that the credibility coefficient is determined based on the product of the total duration ratio and the ratio of the number of voice-associated keywords among the credit-associated keywords.
[0088] It can be understood that the value range of the credibility coefficient is between 0 and 1. When the credibility coefficient is greater than a preset credibility coefficient threshold, it is determined that the credibility of the matched voice time period can meet the requirements.
[0089] Further, when the credibility of the matched voice time period cannot meet the requirements, after parsing and processing all the voice survey data, it is determined whether to generate a credit investigation report according to the parsing result.
[0090] Specifically, the parsing result includes the existence of fraud risk and the non-existence of fraud risk.
[0091] In another possible embodiment, determining that the credibility of the matched voice time period can meet the requirements specifically includes:
[0092] Based on the matching data between different matched voice time periods and the credit-associated keywords, determine the number of matched voice time periods for different credit-associated keywords;
[0093] Based on the duration ratio of the matched voice time periods corresponding to different credit-associated keywords in the survey voice data, determine the keyword credibility coefficients for different credit-associated keywords;
[0094] Based on the average value of the keyword credibility coefficients for different credit-associated keywords, determine the credibility coefficient, and based on the credibility coefficient, determine whether the credibility of the matched voice time period can meet the requirements.
[0095] In another possible embodiment, determining that the credibility of the matched voice time period can meet the requirements specifically includes:
[0096] Based on the matching data between different matched voice time periods and the credit-associated keywords, determine the number of matched voice time periods for different credit-associated keywords. When the number of matched voice time periods for different credit-associated keywords all meet the requirements, it is determined that the credibility of the matched voice time period can meet the requirements;
[0097] When there is a credit-associated keyword for which the number of matched voice time periods does not meet the requirements:
[0098] Obtain the number of credit-related keywords for which the number of matching voice segments does not meet the requirements. When the proportion of the number of credit-related keywords for which the number of matching voice segments does not meet the requirements is greater than the preset keyword number proportion threshold, it is determined that the credibility of the matching voice segments cannot meet the requirements;
[0099] When the proportion of the number of credit-related keywords for which the number of matching voice segments does not meet the requirements is not greater than the preset keyword number proportion threshold:
[0100] Obtain the number of matching voice segments in the survey voice data. When the number of the matching voice segments meets the requirements, it is determined that the credibility of the matching voice segments meets the requirements;
[0101] When the number of the matching voice segments does not meet the requirements:
[0102] Based on the durations of different matching voice segments, determine the sum of the proportion of the durations of the matching voice segments in the survey voice data. When the sum of the proportion of the durations of the matching voice segments in the survey voice data meets the requirements, it is determined that the credibility of the matching voice segments meets the requirements;
[0103] When the sum of the proportion of the durations of the matching voice segments in the survey voice data does not meet the requirements: Based on the number of matching voice segments of different credit-related keywords, the durations of different matching voice segments, and the number of matches of credit-related keywords, determine the keyword matching coefficients of different credit-related keywords. When there is no credit-related keyword with a keyword matching coefficient less than the preset keyword matching coefficient threshold, it is determined that the credibility of the matching voice segments meets the requirements;
[0104] When there is a credit-related keyword with a keyword matching coefficient not less than the preset keyword matching coefficient threshold:
[0105] Determine the credibility coefficient based on the keyword matching coefficients of different credit-related keywords, and based on the credibility coefficient, determine whether the credibility of the matching voice segments can meet the requirements.
[0106] S3 Obtain the emotion recognition results of the voices in different matching voice segments, and combine the matching situation between the voice data in different matching voice segments and the financial data of the enterprise to determine the risk voice segments in the matching voice segments;
[0107] Further, the emotion recognition results include fear emotion and anxiety emotion.
[0108] Specifically, as Figure 4 shown, the method for determining the risk voice segments in the matching voice segments is:
[0109] Determine the proportion of the number of sentences with fear or anxiety emotions in different matching voice periods based on the emotion recognition results of the voices in the different matching voice periods, and use it as the proportion of abnormal sentences;
[0110] Determine the proportion of the number of financial indicators inconsistent with the financial data based on the matching situation between the voice data in the matching voice period and the financial data of the enterprise, and use it as the proportion of deviation indicators;
[0111] Based on the average value of the proportion of abnormal sentences and the proportion of deviation indicators, determine the voice risk coefficient of the matching voice period, and use the voice risk coefficient to determine whether the matching voice period is a risk voice period.
[0112] Further, when the voice risk coefficient is greater than the preset voice risk coefficient threshold, it is determined that the matching voice period is a risk voice period.
[0113] It can be understood that when there is no risk voice period in the matching voice period, during the analysis and processing of voice survey data, the large language model is used to perform synchronous generation processing of the credit investigation report.
[0114] In another possible embodiment, the method for determining the risk voice period in the matching voice period is as follows:
[0115] S31 Determine the proportion of the number of financial indicators inconsistent with the financial data based on the matching situation between the voice data in the matching voice period and the financial data of the enterprise, and use it as the proportion of deviation indicators, and combine the number of sentences containing financial indicators to determine the data deviation coefficient of the matching voice period;
[0116] S32 Determine the proportion of the number of sentences with fear or anxiety emotions in different matching voice periods based on the emotion recognition results of the voices in the different matching voice periods, and use it as the proportion of abnormal sentences, and combine the number of matches of credit-related keywords in the sentences with fear or anxiety emotions to determine the emotion abnormality coefficient of the matching voice period;
[0117] S33 Based on the average value of the emotion abnormality coefficient and the data deviation coefficient, determine the voice risk coefficient of the matching voice period, and use the voice risk coefficient to determine whether the matching voice period is a risk voice period.
[0118] S4 Determine the generation processing method of the credit investigation report of the enterprise by using the large language model according to the distribution data of different risk voice periods and the matching situation of the voice data with other matching voice periods.
[0119] Specifically, the method for determining the generation and processing method of the credit investigation report by the enterprise using the large language model is as follows:
[0120] Based on the matching situation of the voice data of different risk voice time periods and other matching voice time periods, determine the number of matching voice time periods inconsistent with the financial indicators of the risk voice time period. Based on the number of matching voice time periods inconsistent with the financial indicators of the risk voice time period, determine the voice deviation time period in the risk voice time period. Based on the proportion of the voice deviation time period in different risk voice time periods in the number of matching voice time periods, determine the voice deviation coefficient of different risk voice time periods;
[0121] Divide the voice survey data into different voice time periods according to the unit duration. Based on the distribution data of the risk voice time period, determine the proportion of the number of risk voice time periods in different voice time periods;
[0122] Based on the proportion of the number of risk voice time periods in different voice time periods and the voice deviation coefficient of different risk voice time periods, determine the generation and processing method of the credit investigation report by the enterprise using the large language model.
[0123] Furthermore, based on the proportion of the number of risk voice time periods in different voice time periods and the voice deviation coefficient of different risk voice time periods, determining the generation and processing method of the credit investigation report by the enterprise using the large language model specifically includes:
[0124] Take the voice time period with the proportion of the number of risk voice time periods greater than the preset risk time period number proportion threshold as the fraud risk time period. When the number of the fraud risk time periods is not within the preset voice time period number interval, there is no need to generate a credit investigation report, and it is determined that the enterprise has a fraud risk;
[0125] When the number of the fraud risk time periods is within the preset voice time period number interval, it is also necessary to determine whether the average value of the voice deviation coefficients of different risk voice time periods is within the preset deviation coefficient interval. If so, after parsing all the voice survey data, determine whether it is necessary to generate a credit investigation report according to the parsing result. If not, during the analysis and processing of the voice survey data, use the large language model to perform synchronous generation processing of the credit investigation report.
[0126] It should be noted that determining whether it is necessary to generate a credit investigation report according to the parsing result specifically includes:
[0127] When the proportion of the sentences with fear or anxiety emotions in the voice survey data is greater than the preset abnormal emotion sentence number proportion, it is determined that the parsing result is that there is a fraud risk, and there is no need to generate a credit investigation report;
[0128] When the proportion of statements with fear or anxiety emotions in the voice survey data is not greater than the preset proportion of abnormal emotion statement quantities, it is determined that there is no fraud risk in the parsing result, and the large language model is used to perform the generation process of the credit investigation report.
[0129] Embodiment 2
[0130] In a second aspect, as Figure 5 shown, the present invention provides a computer system, including: a memory and a processor connected by communication, and a computer program stored on the memory and capable of running on the processor. When the processor runs the computer program, it executes the above-mentioned method for generating a credit investigation report based on a large language model.
[0131] Optionally, the above step S31 includes the following content:
[0132] S311 Based on the matching situation between the voice data in the matching voice period and the financial data of the enterprise, when it is determined that there are no financial indicators inconsistent with the financial data, the data deviation coefficient of the matching voice period is determined to be 0, and step S32 is entered. When there are financial indicators inconsistent with the financial data, step S312 is entered;
[0133] S312 When the quantity of financial indicators inconsistent with the financial data does not meet the requirements, the matching voice period is determined as a risk voice period. When the quantity of financial indicators inconsistent with the financial data meets the requirements, step S313 is entered;
[0134] S313 Take the financial indicators inconsistent with the financial data as deviation financial indicators, and obtain the quantity of statements associated with different deviation financial indicators. When there are no deviation financial indicators with the quantity of associated statements greater than the preset data quantity threshold, step S315 is entered. When there are deviation financial indicators with the quantity of associated statements greater than the preset data quantity threshold, step S314 is entered;
[0135] S314 When the quantity of deviation financial indicators with the quantity of associated statements greater than the preset data quantity threshold does not meet the requirements, the matching voice period is determined as a risk voice period. When the quantity of deviation financial indicators with the quantity of associated statements greater than the preset data quantity threshold meets the requirements, step S315 is entered;
[0136] S315 determines the data deviation coefficient of the matched voice period according to the proportion of the number of deviation indicators and the number of statements containing financial indicators. When the data deviation coefficient of the matched voice period is greater than the preset data deviation coefficient threshold, it is determined that the matched voice period is a risk voice period. When the data deviation coefficient of the matched voice period is not greater than the preset data deviation coefficient threshold, step S32 is entered;
[0137] Optionally, the above step S32 includes the following content:
[0138] Determine the proportion of the number of statements with fear or anxiety emotions in different matched voice periods based on the emotion recognition results of the voices in different matched voice periods, and use it as the proportion of abnormal statements. Combine the number of matches of credit-related keywords in the statements with fear or anxiety emotions to determine the emotion abnormality coefficient of the matched voice period
[0139] S321 determines the proportion of the number of statements with fear or anxiety emotions in different matched voice periods based on the emotion recognition results of the voices in different matched voice periods, and uses it as the proportion of abnormal statements. When the proportion of abnormal statements does not meet the requirements, it is determined that the matched voice period is a risk voice period. When the proportion of abnormal statements meets the requirements, step S322 is entered;
[0140] S322 regards the statements with credit-related keywords among the statements with fear or anxiety emotions as risk statements. When the number of risk statements does not meet the requirements, it is determined that the matched voice period is a risk voice period. When the number of risk statements meets the requirements, step S323 is entered;
[0141] S323 determines the emotion abnormality coefficient of the matched voice period based on the proportion of abnormal statements and the number of matches of credit-related keywords in the statements with fear or anxiety emotions. When the emotion abnormality coefficient of the matched voice period does not meet the requirements, it is determined that the matched voice period is a risk voice period. When the emotion abnormality coefficient of the matched voice period meets the requirements, step S33 is entered.
[0142] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0143] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0144] The foregoing is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and variations can be made to one or more embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.
Claims
1. A method for generating a credit investigation report based on a large language model, characterized in that: Specifically include: Based on the type of the enterprise, determine the credit-related keywords of the enterprise, and based on the credit-related keywords, determine the matching voice time period in the business investigation process of the enterprise; When it is determined that the credibility of the matching voice period can meet the requirement based on the matching data of different matching voice periods and the credit-related keywords and the distribution data in different matching voice periods, the next step is entered; Acquire emotion recognition results of speech in different matching speech periods, and determine risky speech periods in the matching speech periods in combination with matching conditions of speech data in different matching speech periods and financial data of the enterprise; Determine the generation and processing method of the credit investigation report by the enterprise using the large language model according to the distribution data of different risk voice time periods and the matching status of the voice data of other matching voice time periods; The credit-related keywords with more than 3 associated matching voice periods are regarded as voice-related keywords. The credibility coefficient is determined based on the total proportion of the duration of different matching voice periods in the voice survey data and the average proportion of the voice-related keywords in the number of credit-related keywords. When the credibility coefficient is greater than 0.4, it is determined that the credibility of the matching voice period meets the requirements; Determine the percentage of sentences with fear or anxiety in different matching speech periods based on the emotion recognition results of the speech in different matching speech periods, and use it as the percentage of abnormal sentences; determine the percentage of financial indicators that are inconsistent with the financial data based on the matching of the speech data in the matching speech period with the financial data of the enterprise, and use it as the percentage of deviation indicators; Based on the average of the proportion of abnormal sentences and the proportion of deviation indicators, the voice risk coefficient of the matching voice period is determined. When the voice risk coefficient is greater than 0.4, the matching voice period is determined to be a risky voice period. The method for determining the generation and processing method of the credit investigation report by the enterprise using the large language model is as follows: Determine the number of matching voice periods that are inconsistent with the financial indicators of the risk voice period based on the matching conditions of voice data of different risk voice periods and other matching voice periods; determine the voice deviation periods in the risk voice period based on the number of matching voice periods that are inconsistent with the financial indicators of the risk voice period; and determine the voice deviation coefficients of different risk voice periods based on the proportion of voice deviation periods in different risk voice periods in the number of matching voice periods; Divide the voice survey data into different voice time periods according to the unit duration, and determine the proportion of risky voice time periods in different voice time periods according to the distribution data of the risky voice time periods; The speech periods whose proportion of risk speech periods is greater than the preset risk period proportion threshold are regarded as fraud risk periods. When the number of the fraud risk periods is not within the preset speech period number range, there is no need to generate a credit investigation report, and it is determined that the enterprise has a fraud risk. When the number of fraud risk periods is within the preset voice period number range, it is also necessary to determine whether the average value of the voice deviation coefficient of different risk voice periods is within the preset deviation coefficient range. If so, all voice investigation data are analyzed and processed, and then it is determined whether a credit investigation report needs to be generated based on the analysis results. If not, during the analysis and processing of the voice investigation data, a large language model is used to synchronously generate the credit investigation report.
2. The credit investigation report generation method based on a large language model according to claim 1, characterized in that: The credit-related keywords of the enterprise include electricity consumption data, business indicators and insurance participation indicators.
3. The credit investigation report generation method based on a large language model according to claim 1, characterized in that: The method for determining the matching voice time period during the business investigation of the enterprise is: Based on the credit-related keywords, determining the number of matches between different credit-related keywords and different time periods in the voice survey data during the business survey process; Determining matching keywords in the time period based on the number of matches with different credit-related keywords; Determine whether the time period is a matching voice time period in the business investigation process of the enterprise according to the number of the matching keywords.
4. The credit investigation report generation method based on a large language model as claimed in claim 3, characterized in that: When the number of the matching keywords is greater than a preset matching keyword number threshold, the time period is determined to be a matching voice time period in the business investigation process of the enterprise.
5. The credit investigation report generation method based on a large language model according to claim 1, characterized in that: When the credibility of the matching voice period cannot meet the requirement, all the voice investigation data are analyzed and processed, and it is determined whether a credit investigation report needs to be generated according to the analysis result.
6. The method for generating a credit investigation report based on a large language model according to claim 5, characterized in that: The analysis result includes whether there is a fraud risk or not.
7. The credit investigation report generation method based on a large language model according to claim 1, characterized in that: The emotion recognition results include fear and anxiety.
8. A computer system comprising: A memory and a processor that are communicatively connected, and a computer program stored in the memory and capable of running on the processor, characterized in that when the processor runs the computer program, a method for generating a credit investigation report based on a large language model as described in any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Interactive generation type financial survey report generation method and system based on large model
CN117952075A
Voice quality-control financial security control system and method
CN107547527A
Usage of emotion recognition in trade monitoring
US20240127335A1