Credit survey report generation method and system based on large language model

Through a method based on a large language model, combining voice period matching and financial data matching to identify and handle fraud risks, the problem of cumbersome credit investigation report generation process and difficulty in identifying fraud risks is solved, and the efficiency and accuracy of the report generation is improved.

CN120067304AActive Publication Date: 2025-05-30HANGYIN CONSUMER FINANCE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510541283.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

In the field of small and micro credit, the generation process of credit investigation reports is cumbersome, the workload is large, and it is difficult to effectively identify and deal with fraud risks, affecting the accuracy and preparation efficiency of the report.

Method used

Using a method based on a large language model, the credit association keywords are determined through the type of the enterprise, the determination of voice periods are matched, and the speech emotion recognition is determined, and the risk speech period is determined based on the matching of financial data, and a credit investigation report is generated based on this.

Benefits of technology

It realizes effective identification and processing of fraud risks, improves the efficiency and accuracy of credit investigation reports generation, and reduces server load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067304A_ABST
    Figure CN120067304A_ABST
Patent Text Reader

Abstract

The invention provides a credit survey report generation method and system based on a large language model, and belongs to the technical field of large models, and the method specifically comprises the following steps: according to matching data of different matching voice time periods and credit association keywords and distribution data in the different matching voice time periods, generating a credit survey report according to the matching data of the different matching voice time periods and the credit association keywords; when it is determined that the credibility of the matched voice time periods can meet requirements, acquiring emotion recognition results of voices in different matched voice time periods, and determining risk voice time periods in the matched voice time periods in combination with matching conditions of voice data in the different matched voice time periods and financial data of an enterprise; according to the matching condition of the distribution data of different risk voice periods and the voice data of other matched voice periods, a credit survey report generation processing mode of an enterprise by using a large language model is determined, and the credit survey report generation processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large models, and particularly relates to a method and system for generating a credit investigation report based on a large language model. Background Art

[0002] In the field of small and micro credit, in order to conduct the credit investigation process of small and micro credit reports, it is cumbersome and variable, with high requirements for the personal abilities of credit officers. The process of editing and tabulating credit reports is cumbersome and involves a large amount of work. Therefore, how to improve the accuracy and compilation efficiency of credit investigation reports has become a technical problem to be solved urgently.

[0003] To solve the above technical problems, in the invention patent application CN202410158795.2, "Method and System for Generating an Interactive Generative Financial Investigation Report Based on a Large Model", a large model is used to achieve the interactive and automated generation of due diligence investigation reports in the financial field, greatly reducing the working hours of financial practitioners and improving the efficiency and quality of due diligence investigations. However, there are the following technical defects: During the process of generating a credit report, it is necessary to comprehensively generate and process data from multiple dimensions of an enterprise. Specifically, tax data, enterprise cash flow, electricity consumption data, and voice data during the business investigation process need to be comprehensively considered. The difference in fraud risk may be reflected in the voice data during the business investigation process. Therefore, how to form a differential processing method for generating a credit investigation report based on the recognition and processing results of voice data during the business investigation process, and then improve the processing efficiency of credit investigation reports has become a technical problem to be solved urgently.

[0004] The present application provides a method and system for generating a credit investigation report based on a large language model. Summary of the Invention

[0005] To achieve the purpose of the present invention, the present invention adopts the following technical solutions: Specifically, in the first aspect, the present application provides a method for generating a credit investigation report based on a large language model, which specifically includes: S1 Based on the type of the enterprise, determine the credit-related keywords of the enterprise, and based on the credit-related keywords, determine the matching voice segments during the business investigation of the enterprise; S2 According to the matching data between different matching voice segments and the credit-related keywords, and based on the distribution data in different matching voice segments, when the credibility of the matching voice segments can meet the requirements, proceed to the next step; S3 Obtain the emotional recognition results of the voices in different matching voice segments, and combine the matching situation between the voice data in different matching voice segments and the financial data of the enterprise to determine the risk voice segments in the matching voice segments; S4 determines the generation processing method of the credit investigation report of the enterprise by using the large language model according to the distribution data of different risk voice time periods and the matching situation of the voice data of other matching voice time periods.

[0006] The beneficial effects of the present invention are as follows: Based on the emotion recognition results of the voices in different matching voice time periods and the matching situation of the voice data in different matching voice time periods with the financial data of the enterprise, the risk voice time periods in the matching voice time periods are determined. It not only single-mindedly considers the differences in the manifestations of fear or anxiety caused by fraud in the matching voice time periods, but also considers the differences in the probabilities of fraud risks caused by the deviations from the financial data, realizing the screening of the risk voice time periods with a relatively large probability of fraud risk.

[0007] According to the distribution data of different risk voice time periods and the matching situation of the voice data of other matching voice time periods, determine the generation processing method of the credit investigation report of the enterprise by using the large language model, so as to realize the determination of the generation processing method of the credit investigation report from two perspectives: the distribution discreteness of the risk voice time periods with fraud risks and the differences in fraud risks caused by the deviations of other matching voice time periods in data such as financial indicators. It not only ensures the efficiency of the generation processing of the credit investigation report in the case of relatively small fraud risks, but also reduces the impact on the server load caused by the immediate generation processing of the credit investigation report in the case of relatively large fraud risks.

[0008] A further technical solution is that the credit-related keywords of the enterprise include electricity consumption data, business indicators, and insurance participation indicators.

[0009] A further technical solution is that the method for determining the matching voice time periods in the business investigation process of the enterprise is as follows: Based on the credit-related keywords, determine the matching quantities of different time periods in the voice investigation data in the business investigation process with different credit-related keywords; Based on the matching quantities with different credit-related keywords, determine the matching keywords in the time period; According to the quantity of the matching keywords, determine whether the time period is the matching voice time period in the business investigation process of the enterprise.

[0010] A further technical solution is that the matching keywords are the credit-related keywords with the matching quantity greater than the preset matching quantity threshold.

[0011] A further technical solution is that when the number of the matching keywords is greater than a preset matching keyword number threshold, it is determined that the time period is a matching voice time period during the business investigation of the enterprise. A further technical solution is that the method for the enterprise to determine the generation processing method of the credit investigation report by using the large language model is as follows: Based on the matching situation of the voice data of different risk voice time periods and other matching voice time periods, determine the number of matching voice time periods inconsistent with the financial indicators of the risk voice time period, determine the voice deviation time period in the risk voice time period according to the number of matching voice time periods inconsistent with the financial indicators of the risk voice time period, and determine the voice deviation coefficient of different risk voice time periods according to the proportion of the voice deviation time periods in different risk voice time periods in the number of the matching voice time periods; Divide the voice investigation data into different voice time periods according to the unit duration, and determine the proportion of the number of risk voice time periods in different voice time periods according to the distribution data of the risk voice time periods; Based on the proportion of the number of risk voice time periods in different voice time periods and the voice deviation coefficients of different risk voice time periods, determine the generation processing method of the credit investigation report by the enterprise using the large language model.

[0012] A further technical solution is that based on the proportion of the number of risk voice time periods in different voice time periods and the voice deviation coefficients of different risk voice time periods, determine the generation processing method of the credit investigation report by the enterprise using the large language model, which specifically includes: Take the voice time periods with the proportion of the number of risk voice time periods greater than the preset risk time period proportion threshold as fraud risk time periods. When the number of the fraud risk time periods is not within the preset voice time period number range, there is no need to generate a credit investigation report, and it is determined that the enterprise has fraud risk; When the number of the fraud risk time periods is within the preset voice time period number range, it is also necessary to determine whether the average value of the voice deviation coefficients of different risk voice time periods is within the preset deviation coefficient range. If so, after parsing all the voice investigation data, determine whether it is necessary to generate a credit investigation report according to the parsing result. If not, during the analysis and processing of the voice investigation data, use the large language model to perform synchronous generation processing of the credit investigation report.

[0013] A further technical solution is that according to the parsing result, determine whether it is necessary to generate a credit investigation report, which specifically includes: When the proportion of the sentences with fear or anxiety emotions in the voice investigation data is greater than the preset abnormal emotion sentence number proportion, it is determined that the parsing result is that there is fraud risk and there is no need to generate a credit investigation report; When the proportion of statements with fear or anxiety emotions in the voice survey data is not greater than the preset proportion of abnormal emotion statements, it is determined that there is no fraud risk in the parsing result, and the large language model is used to generate and process the credit investigation report.

[0014] In a second aspect, the present invention provides a computer system, including: a memory and a processor connected by communication, and a computer program stored on the memory and capable of running on the processor. When the processor runs the computer program, it executes the above-mentioned method for generating a credit investigation report based on a large language model.

[0015] Other features and advantages will be described in the subsequent specification. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification and the drawings.

[0016] To make the above-mentioned objectives, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. Description of the Drawings

[0017] By referring to the drawings and describing its exemplary embodiments in detail, the above and other features and advantages of the present invention will become more obvious.

[0018] Figure 1 is a flowchart of a method for generating a credit investigation report based on a large language model; Figure 2 is a flowchart of a method for determining a matching voice period in the business investigation process of an enterprise; Figure 3 is a flowchart for determining that the credibility of the matching voice period can meet the requirements; Figure 4 is a flowchart of a method for determining a risk voice period in the matching voice period; Figure 5 is a framework diagram of a computer system. Detailed Embodiments

[0019] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.

[0020] In this application, by leveraging the fraud risks of enterprises in the surveyed voice data, a differentiated generation processing method for credit investigation reports is determined. Specifically, when the fraud risk is relatively high, after the surveyed voice data is analyzed and processed and no fraud risk is determined, the generation of the credit investigation report is carried out. When the fraud risk is relatively low, the generation of the credit investigation report is carried out during the parsing and processing of the surveyed voice data, thereby improving the efficiency of the generation processing of the credit investigation report.

[0021] Match voice time periods with more than 10 matches of credit-related keywords.

[0022] Take credit-related keywords with the number of associated matching voice time periods of more than 3 as voice-associated keywords, and determine the credibility coefficient based on the average of the total duration proportion of different matching voice time periods in the voice survey data and the quantity proportion of voice-associated keywords among the credit-related keywords. When the credibility coefficient is greater than 0.4, it is determined that the credibility degree of the matching voice time period can meet the requirements.

[0023] Determine the proportion of the number of sentences with fear or anxiety emotions in different matching voice time periods based on the emotion recognition results of the voices in different matching voice time periods, and take it as the proportion of abnormal sentence numbers. Determine the proportion of the number of financial indicators inconsistent with the financial data based on the matching situation between the voice data in the matching voice time periods and the financial data of the enterprise, and take it as the proportion of deviation indicator numbers; Based on the average of the proportion of abnormal sentence numbers and the proportion of deviation indicator numbers, determine the voice risk coefficient of the matching voice time period. When the voice risk coefficient is greater than 0.4, it is determined that the matching voice time period is a risk voice time period.

[0024] When the number of matching voice time periods inconsistent with the financial indicators of the risk voice time period is more than 2, take it as a voice deviation time period. Determine the voice deviation coefficient of different risk voice time periods according to the proportion of voice deviation time periods in different risk voice time periods among the matching voice time periods. Divide the voice survey data into different voice time periods according to the unit duration, and determine the proportion of the number of risk voice time periods in different voice time periods based on the distribution data of the risk voice time periods; Based on the proportion of the number of risk voice time periods in different voice time periods and the voice deviation coefficients of different risk voice time periods, determine the generation processing method of the credit investigation report by the enterprise using the large language model.

[0025] Using the large model to generate the credit investigation report mainly includes: 1. Record during the business investigation process; 2. Convert the language into text; 3. Perform large language model prompt engineering on the text and mark the non-compliant content. 4. Perform large language model prompt engineering on the investigation content to refine business attribute information and attribute data. 5. Follow up on the attribute information and attribute data, and generate reports in combination with the user's financial indicator data.

[0026] Embodiment 1 As Figure 1 shown, the present application provides a method for generating a credit investigation report based on a large language model, specifically including: S1. Based on the type of the enterprise, determine the credit-related keywords of the enterprise, and based on the credit-related keywords, determine the matching voice segments during the business investigation of the enterprise. Further, the credit-related keywords of the enterprise include electricity consumption data, business indicators, and insurance participation indicators.

[0027] Specifically, as Figure 2 shown, the method for determining the matching voice segments during the business investigation of the enterprise is: Based on the credit-related keywords, determine the matching quantity of different segments in the voice investigation data during the business investigation with different credit-related keywords. Based on the matching quantity with different credit-related keywords, determine the matching keywords in the segment. Determine whether the segment is a matching voice segment during the business investigation of the enterprise according to the quantity of the matching keywords.

[0028] Further, the matching keywords are credit-related keywords with a matching quantity greater than a preset matching quantity threshold.

[0029] It should be noted that when the quantity of the matching keywords is greater than the preset matching keyword quantity threshold, it is determined that the segment is a matching voice segment during the business investigation of the enterprise.

[0030] In another possible embodiment, the method for determining the matching voice segments during the business investigation of the enterprise is: Based on the credit-related keywords, determine the matching quantity of different segments in the voice investigation data during the business investigation with different credit-related keywords. Based on the matching quantity with different credit-related keywords, determine the sentences with credit-related keywords of the investigation object in the segment and use them as associated sentences. Determine whether the segment is a matching voice segment during the business investigation of the enterprise according to the quantity of the associated sentences.

[0031] Specifically, when the number of the associated statements is greater than a preset threshold of the number of associated statements, it is determined that the time period is a matching voice time period during the business investigation process of the enterprise.

[0032] In another possible embodiment, the method for determining the matching voice time period during the business investigation process of the enterprise is as follows: Based on the credit-related keywords, determine the matching quantities of different time periods in the voice investigation data during the business investigation process with different credit-related keywords. When the total matching quantity of the time period with different credit-related keywords is greater than a preset matching quantity threshold, it is determined that the time period is a matching voice time period; When the total matching quantity of the time period with different credit-related keywords is not greater than the preset matching quantity threshold: Based on the matching quantities with different credit-related keywords, determine the matching keywords in the time period. When there are no matching keywords in the time period, it is determined that the time period does not belong to the matching voice time period; When there are matching keywords in the time period: Obtain the quantity of the matching keywords in the time period. When the quantity of the matching keywords in the time period meets the requirements, it is determined that the time period belongs to the matching voice time period; When the quantity of the matching keywords in the time period does not meet the requirements: Based on the matching quantities with different credit-related keywords, determine the statements with credit-related keywords of the investigation object in the time period and use them as associated statements. When the quantity of the associated statements in the time period meets the requirements, it is determined that the time period belongs to the matching voice time period; When the quantity of the associated statements in the time period meets the requirements: Determine the correlation coefficients of different credit-related feature words based on the quantities and matching quantities of the associated statements of different credit-related feature words. When the quantity of the credit-related feature words with correlation coefficients greater than a preset correlation coefficient threshold meets the requirements, it is determined that the time period belongs to the matching voice time period; When the quantity of the credit-related feature words with correlation coefficients greater than the preset correlation coefficient threshold does not meet the requirements: Determine the feature word correlation coefficient of the time period based on the correlation coefficients of the credit-related feature words of different credit-related keywords, and determine whether the time period is a matching voice time period during the business investigation process of the enterprise according to the feature word correlation coefficient.

[0033] It should be noted that when the feature word correlation coefficient of the time period is greater than a preset feature word correlation coefficient threshold, it is determined that the time period is a matching voice time period during the business investigation process of the enterprise.

[0034] Based on the matching data of different matching voice time periods and the credit-related keywords, and based on the distribution data in different matching voice time periods, when it is determined that the credibility of the matching voice time period can meet the requirements, proceed to the next step; Specifically, as Figure 3 shown, determining that the credibility of the matching voice time period can meet the requirements specifically includes: Based on the matching data of different matching voice time periods and the credit-related keywords, determine the number of matching voice time periods of different credit-related keywords, and determine the voice-related keywords in the credit-related keywords based on the number of the matching voice time periods; Based on the distribution data of different matching voice time periods, determine the total duration proportion of different matching voice time periods in the voice survey data; Based on the total duration proportion and the quantity proportion of the voice-related keywords in the credit-related keywords, determine the credibility coefficient, and based on the credibility coefficient, determine whether the credibility of the matching voice time period can meet the requirements.

[0035] Further, the voice-related keywords are credit-related keywords with the number of matching voice time periods greater than the preset threshold of the number of matching voice time periods.

[0036] It should be noted that the credibility coefficient is determined according to the product of the total duration proportion and the quantity proportion of the voice-related keywords in the credit-related keywords.

[0037] It can be understood that the value range of the credibility coefficient is between 0 and 1. When the credibility coefficient is greater than the preset credibility coefficient threshold, it is determined that the credibility of the matching voice time period can meet the requirements.

[0038] Further, when the credibility of the matching voice time period cannot meet the requirements, then after parsing all the voice survey data, determine whether it is necessary to generate a credit investigation report according to the parsing result.

[0039] Specifically, the parsing result includes the existence of fraud risk and the non-existence of fraud risk.

[0040] In another possible embodiment, determining that the credibility of the matching voice time period can meet the requirements specifically includes: Based on the matching data of different matching voice time periods and the credit-related keywords, determine the number of matching voice time periods of different credit-related keywords; Based on the duration proportion of the matching voice time periods corresponding to different credit-related keywords in the survey voice data, determine the keyword credibility coefficients of different credit-related keywords; Determine the credibility coefficient based on the average value of the credibility coefficients of keywords associated with different credits, and determine whether the credibility degree of the matched speech period can meet the requirements based on the credibility coefficient.

[0041] In another possible embodiment, determining that the credibility degree of the matched speech period can meet the requirements specifically includes: Based on the matching data between different matched speech periods and the credit-associated keywords, determine the number of matched speech periods of different credit-associated keywords. When the number of matched speech periods of different credit-associated keywords all meets the requirements, it is determined that the credibility degree of the matched speech period can meet the requirements; When there are credit-associated keywords for which the number of matched speech periods does not meet the requirements: Obtain the number of credit-associated keywords for which the number of matched speech periods does not meet the requirements. When the proportion of the number of credit-associated keywords for which the number of matched speech periods does not meet the requirements is greater than the preset keyword number proportion threshold, it is determined that the credibility degree of the matched speech period cannot meet the requirements; When the proportion of the number of credit-associated keywords for which the number of matched speech periods does not meet the requirements is not greater than the preset keyword number proportion threshold: Obtain the number of matched speech periods in the surveyed speech data. When the number of matched speech periods meets the requirements, it is determined that the credibility degree of the matched speech period meets the requirements; When the number of matched speech periods does not meet the requirements: Based on the durations of different matched speech periods, determine the sum of the duration proportions of the matched speech periods in the surveyed speech data. When the sum of the duration proportions of the matched speech periods in the surveyed speech data meets the requirements, it is determined that the credibility degree of the matched speech period meets the requirements; When the sum of the duration proportions of the matched speech periods in the surveyed speech data does not meet the requirements: Based on the number of matched speech periods of different credit-associated keywords, the durations of different matched speech periods, and the number of matches of credit-associated keywords, determine the keyword matching coefficients of different credit-associated keywords. When there are no credit-associated keywords with keyword matching coefficients less than the preset keyword matching coefficient threshold, it is determined that the credibility degree of the matched speech period meets the requirements; When there are credit-associated keywords with keyword matching coefficients not less than the preset keyword matching coefficient threshold: Determine the credibility coefficient based on the keyword matching coefficients of different credit-associated keywords, and determine whether the credibility degree of the matched speech period can meet the requirements based on the credibility coefficient.

[0042] S3 obtains the emotion recognition results of the voices in different matched voice segments, and combines the matching situation between the voice data in different matched voice segments and the financial data of the enterprise to determine the risk voice segments in the matched voice segments; Further, the emotion recognition results include fear emotion and anxiety emotion.

[0043] Specifically, as Figure 4 shown, the method for determining the risk voice segments in the matched voice segments is as follows: Determine the proportion of the number of sentences with fear or anxiety emotions in different matched voice segments based on the emotion recognition results of the voices in different matched voice segments, and use it as the proportion of abnormal sentences; Determine the proportion of the number of financial indicators inconsistent with the financial data based on the matching situation between the voice data in the matched voice segments and the financial data of the enterprise, and use it as the proportion of deviation indicators; Based on the average value of the proportion of abnormal sentences and the proportion of deviation indicators, determine the voice risk coefficient of the matched voice segment, and use the voice risk coefficient to determine whether the matched voice segment is a risk voice segment.

[0044] Further, when the voice risk coefficient is greater than the preset voice risk coefficient threshold, it is determined that the matched voice segment is a risk voice segment.

[0045] It can be understood that when there is no risk voice segment in the matched voice segment, during the analysis and processing of voice survey data, the large language model is used to perform synchronous generation processing of credit investigation reports.

[0046] In another possible embodiment, the method for determining the risk voice segments in the matched voice segments is as follows: S31 Determine the proportion of the number of financial indicators inconsistent with the financial data based on the matching situation between the voice data in the matched voice segments and the financial data of the enterprise, and use it as the proportion of deviation indicators, and combine the number of sentences containing financial indicators to determine the data deviation coefficient of the matched voice segment; S32 Determine the proportion of the number of sentences with fear or anxiety emotions in different matched voice segments based on the emotion recognition results of the voices in different matched voice segments, and use it as the proportion of abnormal sentences, and combine the number of matches of credit-related keywords in the sentences with fear or anxiety emotions to determine the emotion abnormality coefficient of the matched voice segment; S33 determines the voice risk coefficient of the matched voice period based on the average value of the emotional anomaly coefficient and the data deviation coefficient, and uses the voice risk coefficient to determine whether the matched voice period is a risky voice period.

[0047] S4 determines the generation processing method of the credit investigation report by the enterprise using the large language model according to the distribution data of different risky voice periods and the matching situation of the voice data of other matched voice periods.

[0048] Specifically, the method for determining the generation processing method of the credit investigation report by the enterprise using the large language model is as follows: Based on the matching situation of the voice data of different risky voice periods and other matched voice periods, determine the number of matched voice periods inconsistent with the financial indicators of the risky voice period. Determine the voice deviation period in the risky voice period according to the number of matched voice periods inconsistent with the financial indicators of the risky voice period. Determine the voice deviation coefficient of different risky voice periods according to the proportion of the voice deviation periods in different risky voice periods in the number of the matched voice periods. Divide the voice survey data into different voice periods according to the unit duration, and determine the proportion of the number of risky voice periods in different voice periods according to the distribution data of the risky voice periods. Based on the proportion of the number of risky voice periods in different voice periods and the voice deviation coefficient of different risky voice periods, determine the generation processing method of the credit investigation report by the enterprise using the large language model.

[0049] Further, based on the proportion of the number of risky voice periods in different voice periods and the voice deviation coefficient of different risky voice periods, determine the generation processing method of the credit investigation report by the enterprise using the large language model, which specifically includes: Take the voice period with the proportion of the number of risky voice periods greater than the preset risk period number proportion threshold as the fraud risk period. When the number of the fraud risk periods is not within the preset voice period number interval, there is no need to generate a credit investigation report, and it is determined that the enterprise has a fraud risk. When the number of the fraud risk periods is within the preset voice period number interval, it is also necessary to determine whether the average value of the voice deviation coefficients of different risky voice periods is within the preset deviation coefficient interval. If so, after parsing all the voice survey data, determine whether it is necessary to generate a credit investigation report according to the parsing result. If not, during the analysis and processing of the voice survey data, use the large language model to generate a credit investigation report synchronously.

[0050] It should be noted that it is determined whether to generate a credit investigation report according to the parsing result, specifically including: When the proportion of statements with fear or anxiety emotions in the voice survey data is greater than the proportion of the preset abnormal emotion statement quantity, it is determined that there is a fraud risk in the parsing result, and the generation process of the credit investigation report is not required; When the proportion of statements with fear or anxiety emotions in the voice survey data is not greater than the proportion of the preset abnormal emotion statement quantity, it is determined that there is no fraud risk in the parsing result, and the generation process of the credit investigation report is carried out using a large language model.

[0051] Embodiment 2 In a second aspect, as Figure 5 shown, the present invention provides a computer system, including: a memory and a processor connected by communication, and a computer program stored on the memory and capable of running on the processor. When the processor runs the computer program, it executes the above-mentioned credit investigation report generation method based on a large language model.

[0052] Optionally, the above step S31 includes the following content: S311 determines that when there are no financial indicators inconsistent with the financial data based on the matching situation between the voice data in the matching voice period and the financial data of the enterprise, the data deviation coefficient of the matching voice period is determined to be 0, and step S32 is entered. When there are financial indicators inconsistent with the financial data, step S312 is entered; S312 determines that when the number of financial indicators inconsistent with the financial data does not meet the requirements, the matching voice period is determined as a risk voice period. When the number of financial indicators inconsistent with the financial data meets the requirements, step S313 is entered; S313 takes the financial indicators inconsistent with the financial data as deviation financial indicators, obtains the number of statements associated with different deviation financial indicators. When there are no deviation financial indicators with the number of associated statements greater than the preset data quantity threshold, step S315 is entered. When there are deviation financial indicators with the number of associated statements greater than the preset data quantity threshold, step S314 is entered; S314 determines that when the number of deviation financial indicators with the number of associated statements greater than the preset data quantity threshold does not meet the requirements, the matching voice period is determined as a risk voice period. When the number of deviation financial indicators with the number of associated statements greater than the preset data quantity threshold meets the requirements, step S315 is entered; S315 determines the data deviation coefficient of the matched voice period according to the proportion of the number of deviation indicators and the number of sentences containing financial indicators. When the data deviation coefficient of the matched voice period is greater than the preset data deviation coefficient threshold, it is determined that the matched voice period is a risk voice period. When the data deviation coefficient of the matched voice period is not greater than the preset data deviation coefficient threshold, proceed to step S32; Optionally, the above step S32 includes the following content: Based on the emotion recognition results of the voices in different matched voice periods, determine the proportion of the number of sentences with fear or anxiety emotions in different matched voice periods, and use it as the proportion of abnormal sentences. Combine the number of matches of credit-related keywords in the sentences with fear or anxiety emotions to determine the emotion abnormality coefficient of the matched voice period S321 Based on the emotion recognition results of the voices in different matched voice periods, determine the proportion of the number of sentences with fear or anxiety emotions in different matched voice periods, and use it as the proportion of abnormal sentences. When the proportion of abnormal sentences does not meet the requirements, it is determined that the matched voice period is a risk voice period. When the proportion of abnormal sentences meets the requirements, proceed to step S322; S322 Consider the sentences with credit-related keywords among the sentences with fear or anxiety emotions as risk sentences. When the number of risk sentences does not meet the requirements, it is determined that the matched voice period is a risk voice period. When the number of risk sentences meets the requirements, proceed to step S323; S323 Based on the proportion of abnormal sentences and the number of matches of credit-related keywords in the sentences with fear or anxiety emotions, determine the emotion abnormality coefficient of the matched voice period. When the emotion abnormality coefficient of the matched voice period does not meet the requirements, it is determined that the matched voice period is a risk voice period. When the emotion abnormality coefficient of the matched voice period meets the requirements, proceed to step S33.

[0053] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0054] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0055] The foregoing is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of the claims of this specification.

Claims

1. A method for generating a credit investigation report based on a large language model, characterized in that: Specifically include: Based on the type of the enterprise, determine the credit-related keywords of the enterprise, and based on the credit-related keywords, determine the matching voice time period in the business investigation process of the enterprise; When it is determined that the credibility of the matching voice period can meet the requirement based on the matching data of different matching voice periods and the credit-related keywords and the distribution data in different matching voice periods, the next step is entered; Acquire emotion recognition results of speech in different matching speech periods, and determine risky speech periods in the matching speech periods in combination with matching conditions of speech data in different matching speech periods and financial data of the enterprise; According to the distribution data of different risky speech time periods and the matching status of speech data of other matching speech time periods, the generation and processing method of the credit investigation report by the enterprise using the large language model is determined.

2. The credit investigation report generation method based on a large language model according to claim 1, characterized in that: The credit-related keywords of the enterprise include electricity consumption data, business indicators and insurance participation indicators.

3. The credit investigation report generation method based on a large language model according to claim 1, characterized in that: The method for determining the matching voice time period during the business investigation of the enterprise is: Based on the credit-related keywords, determining the number of matches between different credit-related keywords and different time periods in the voice survey data during the business survey process; Determining matching keywords in the time period based on the number of matches with different credit-related keywords; Determine whether the time period is a matching voice time period in the business investigation process of the enterprise according to the number of the matching keywords.

4. The credit investigation report generation method based on a large language model as claimed in claim 3, characterized in that: When the number of the matching keywords is greater than a preset matching keyword number threshold, the time period is determined to be a matching voice time period in the business investigation process of the enterprise.

5. The credit investigation report generation method based on a large language model according to claim 1, characterized in that: When the credibility of the matching voice period cannot meet the requirement, all the voice investigation data are analyzed and processed, and it is determined whether a credit investigation report needs to be generated according to the analysis result.

6. The method for generating a credit investigation report based on a large language model according to claim 5, characterized in that: The analysis result includes whether there is a fraud risk or not.

7. The method for generating a credit investigation report based on a large language model according to claim 1, characterized in that: The emotion recognition results include fear and anxiety.

8. The credit investigation report generation method based on a large language model according to claim 1, characterized in that: The method for determining the generation and processing method of the credit investigation report by the enterprise using the large language model is as follows: Determine the number of matching voice periods that are inconsistent with the financial indicators of the risk voice period based on the matching conditions of voice data of different risk voice periods and other matching voice periods; determine the voice deviation periods in the risk voice period based on the number of matching voice periods that are inconsistent with the financial indicators of the risk voice period; and determine the voice deviation coefficients of different risk voice periods based on the proportion of voice deviation periods in different risk voice periods in the number of matching voice periods; Divide the voice survey data into different voice time periods according to the unit duration, and determine the proportion of risky voice time periods in different voice time periods according to the distribution data of the risky voice time periods; Based on the proportion of risky speech periods in different speech periods and the speech deviation coefficients of different risky speech periods, the generation and processing method of the credit investigation report by the enterprise using a large language model is determined.

9. The credit investigation report generation method based on a large language model according to claim 8, characterized in that: Based on the proportion of risky speech periods in different speech periods and the speech deviation coefficients of different risky speech periods, the generation and processing method of the credit investigation report by the enterprise using the large language model is determined, which specifically includes: The speech periods whose proportion of risk speech periods is greater than the preset risk period proportion threshold are regarded as fraud risk periods. When the number of the fraud risk periods is not within the preset speech period number range, there is no need to generate a credit investigation report, and it is determined that the enterprise has a fraud risk. When the number of fraud risk periods is within the preset voice period number range, it is also necessary to determine whether the average value of the voice deviation coefficient of different risk voice periods is within the preset deviation coefficient range. If so, all voice investigation data are analyzed and processed, and then it is determined whether a credit investigation report needs to be generated based on the analysis results. If not, during the analysis and processing of the voice investigation data, a large language model is used to synchronously generate the credit investigation report.

10. A computer system comprising: A memory and a processor that are communicatively connected, and a computer program stored in the memory and capable of running on the processor, characterized in that when the processor runs the computer program, a method for generating a credit investigation report based on a large language model as described in any one of claims 1 to 9 is executed.

Citation Information

Patent Citations

  • Interactive generation type financial survey report generation method and system based on large model

    CN117952075A

  • Voice quality-control financial security control system and method

    CN107547527A

  • Usage of emotion recognition in trade monitoring

    US20240127335A1