Data processing method and related device

By using prompt engineering and large language models in the financial field to process credit scoring data and generate accurate risk summaries, the problem of inaccurate credit risk analysis in existing technologies is solved, and more efficient credit assessment and decision support are achieved.

CN120655403APending Publication Date: 2025-09-16HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410292380.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-12
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing text summarization algorithms are not ideal for extracting customer credit risk in the financial field, especially in the analysis of credit score data, resulting in insufficient accuracy and credibility of risk summaries.

Method used

By obtaining the customer's credit score data, using prompt engineering to generate the target prompt, and inputting it into a large language model to label the credit status output, combined with the influencing factors of the credit score and the risk analysis report, a risk summary is generated. Similarity judgment and text diversity detection are used to correct hallucination problems and generate more accurate risk analysis results.

Benefits of technology

It improves the accuracy and credibility of credit risk analysis, helps financial institutions quickly and comprehensively understand customer credit status, and supports more accurate credit granting decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655403A_ABST
    Figure CN120655403A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and a related device, and is applied to the field of artificial intelligence, in the method, first data of a first customer is acquired, and the first data comprises credit scores corresponding to a plurality of time points of the first customer; according to the first data, a target prompt prompt is generated based on prompt engineering processing, the target prompt is used for describing the credit condition indicated by the first data, and the prompt engineering processing is used for carrying out annotation output of the credit condition for the credit score; and inputting the target prompt into the large language model to obtain a risk analysis result of the first customer. According to the method, the financial institution can conveniently and comprehensively know the credit risk condition of the customer in the credit decision-making process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a data processing method and related devices. Background Art

[0002] Before granting credit to customers with credit, financing, post-payment and other needs, financial institutions need to conduct a risk analysis based on all aspects of the customer's information, including the customer's credit history, debt level, identity qualifications, transaction records, etc., and then make credit decisions, credit limit assessments, credit issuance and other operations.

[0003] To save time analyzing customer information risks and avoid missing critical information, we first extract key information from the risk summary. This information is then assembled into a risk summary and output to the risk control manager for analysis and judgment. Currently, relevant technologies typically use text summarization algorithms for summary extraction. However, these methods primarily focus on common texts such as news, and are not ideal for extracting key information in the financial sector. Summary of the Invention

[0004] The embodiments of the present application provide a data processing method and related devices, which are intended to facilitate financial institutions to fully understand the credit risk status of customers during the credit decision-making process.

[0005] In a first aspect, an embodiment of the present application provides a data processing method, the method comprising:

[0006] First data of a first customer is obtained, the first data including the first customer's credit scores at multiple time points. Specifically, the first customer's ID is first obtained and the first data is retrieved from a database. A target prompt is then generated based on the first data using prompt engineering. The target prompt describes the credit status indicated by each credit score in the first data. The prompt engineering process outputs a credit status label based on the credit score. Finally, the target prompt is input into the trained large language model to output a risk analysis result for the first customer.

[0007] In the financial sector, credit scores are a crucial assessment tool used to measure a customer's credit risk. Based on multiple factors (such as a customer's credit history, debt level, income status, and credit history), they use complex mathematical models to produce a specific score that reflects a customer's creditworthiness and the likelihood of loan default. Credit scores are typically expressed in numbers or letters, for example, a score of "800" or a grade of "A."

[0008] Using this method, we perform prompt engineering on credit scores, annotating and outputting customer credit scores at different points in time to generate prompts that can be recognized and processed by a large language model. This allows the large language model to better understand and interpret credit score data, resulting in more accurate risk analysis results.

[0009] In one possible implementation, the method further includes obtaining second data of the first customer, the second data including factors influencing a credit score; specifically, obtaining the second data from a database based on an ID of the first customer. Then, based on the second data and the risk analysis results, generating a risk summary for the first customer using a text summarization algorithm or a summary template, the risk summary is used to describe the creditworthiness of the first customer.

[0010] In this application, factors affecting credit scores include, for example, an individual's credit history, debt level, income status, etc., or a company's identity qualifications, risk performance, and ability to fulfill its obligations.

[0011] By adopting the above method, by obtaining the risk analysis report generated based on the credit score and the influencing factors of the credit score, we can have a more comprehensive understanding of the customer's credit status and generate a summary of the customer's credit risk, which is conducive to financial institutions to quickly assess the customer's credit credibility and make more accurate decisions when granting credit.

[0012] In one possible implementation, the risk summary includes at least one summary point. The method further includes: performing a similarity judgment between a first summary point and the first data to obtain first text information whose similarity is greater than a first threshold, where the first text information is included in the first data, and the first summary point is any one of the at least one summary point; and then recording a correspondence between the first text information and the first summary point, where the correspondence is used to indicate a position of the first text information in the first data.

[0013] For example, if the first summary point is "The customer's credit score in the latest accounting period is 663, which is average credit", after similarity judgment with the first data, it is concluded that the credit score corresponding to the last time node among multiple credit scores is "663". The text information in this area has the highest similarity with the first summary point, so this text information is used as the original information of the first summary point.

[0014] Using the above method, a correspondence between the first summary point and the first text information is established. This association not only enables decision makers such as risk control managers to quickly locate the original information corresponding to each summary point, but also provides a detailed explanation of the summary point, thereby improving the credibility and practicality of the risk summary.

[0015] In one possible implementation, a target prompt is input into a large language model to obtain a risk analysis result for a first customer, including: inputting the target prompt into the large language model to first obtain a first text; the large language model detects text diversity between text data of the target prompt and text data of the first text; and when the text diversity meets a preset condition, outputting the first text as the risk analysis result.

[0016] In a possible implementation, when the text diversity does not meet the preset condition, the large language model is used to update the first text until the updated text meets the preset condition.

[0017] For example, an n-gram language model (n-gram) is used for statistical analysis. First, the first text and the target prompt are segmented by n-grams to obtain their respective n-gram sets and extract the identical n-grams between the two. Then, the identical n-grams are compared with the total number of n-grams. A threshold is preset to determine whether the similarity meets the requirement. If the similarity ratio is lower than the threshold, it is considered that an illusion problem has occurred and the output needs to be re-performed; if the preset threshold is met, it is output as a risk analysis result.

[0018] By using the above method, by comparing the text diversity between the input parameters and output parameters of the large language model, it helps the large language model to promptly detect hallucination problems and correct the output.

[0019] In one possible implementation, a target prompt is generated based on prompt engineering processing according to the first data, including: obtaining a slope based on linear regression fitting calculation according to the credit scores corresponding to multiple time points, the slope being used to indicate the trend of the credit score in historical time; generating trend description information according to the slope, and the target prompt including the trend description information.

[0020] Specifically, the time point corresponding to the credit score is used as the independent variable and the credit score is used as the dependent variable. A linear regression function between the two can be calculated, where the slope is regarded as the trend of change of the credit score over time, generating trend description information.

[0021] In one possible implementation, the method further includes: extracting statistical indicators of the credit score to obtain statistical description information, where the statistical indicators include one or more of the maximum value, minimum value, mean value, and volatility of the credit score, and the target prompt includes the statistical description information.

[0022] In one possible implementation, the method further includes: matching the credit score with business knowledge information to obtain risk description information, where the business knowledge information is used to indicate the credit risk level represented by credit scores in different intervals, and the target prompt includes the risk description information.

[0023] Specifically, the business knowledge information includes, for example, a credit score range of "[750,+∞)" indicates "good credit status"; a score range of "[650,750)" indicates "average credit status"; a score range of "[0,650)" indicates "very poor credit status", etc.

[0024] Using this method, we can generate target prompts more comprehensively and objectively, helping large language models understand task requirements more quickly and generate more logical and easier-to-understand text content.

[0025] In a possible implementation, the large language model is a large language model fine-tuned using a sample set, where the sample set includes a plurality of sample prompts and a risk analysis report corresponding to each of the sample prompts.

[0026] By fine-tuning the large language model using the sample set described above, it can be better adapted to specific tasks or domains. The risk analysis report corresponding to each sample prompt provides the model with rich contextual information, enabling it to better understand the semantics and underlying intent of the prompt. This helps the large language model understand credit risk in the financial sector, thereby improving its generation accuracy and adaptability.

[0027] In a possible implementation manner, the first data is table data.

[0028] In a second aspect, an embodiment of the present application provides a data processing device, the device comprising:

[0029] an acquisition module, configured to acquire first data of a first customer, the first data including credit scores of the first customer corresponding to multiple time points;

[0030] a processing module, configured to generate a target prompt based on the first data and prompt engineering processing, wherein the target prompt is used to describe the credit status indicated by the first data, and the prompt engineering processing is used to output a credit status annotation based on the credit score;

[0031] The processing module is also used to input the target prompt into the large language model to obtain the risk analysis result of the first customer.

[0032] In one possible implementation, the acquisition module is further used to: obtain second data of the first customer, the second data including influencing factors of the credit score; the processing module is further used to generate a risk summary based on the second data and the risk analysis results, and the risk summary is used to describe the credit credibility of the first customer.

[0033] In one possible implementation, the risk summary includes at least one summary point, and the processing module is further used to: perform a similarity judgment between the first summary point and the first data to obtain first text information whose similarity is greater than a first threshold, where the first text information is included in the first data, and the first summary point is any one of the at least one summary points; and record a correspondence between the first text information and the first summary point, where the correspondence is used to indicate a position of the first text information in the first data.

[0034] In one possible implementation, the processing module is specifically used to: input the target prompt into the large language model to obtain a first text; detect the text diversity between the text data of the target prompt and the text data of the first text; when the text diversity meets a preset condition, use the first text as a risk analysis result.

[0035] In a possible implementation, when the text diversity does not meet the preset condition, the large language model is used to update the first text until the updated text meets the preset condition.

[0036] In one possible implementation, the processing module is specifically configured to: obtain a slope based on linear regression fitting calculation according to the credit scores corresponding to multiple time points, where the slope is used to indicate the trend of the credit score in historical time; generate trend description information based on the slope, where the target prompt includes the trend description information.

[0037] In one possible implementation, the processing module is further used to extract statistical indicators of the credit score to obtain statistical description information, where the statistical indicators include one or more of the maximum value, minimum value, mean value, and volatility of the credit score, and the target prompt includes the statistical description information.

[0038] In one possible implementation, the processing module is further used to match the credit score with business knowledge information to obtain risk description information, where the business knowledge information is used to indicate the credit risk level represented by credit scores in different score intervals, and the target prompt includes the risk description information.

[0039] In a possible implementation, the large language model is a large language model fine-tuned using a sample set, where the sample set includes multiple sample prompts and a risk analysis report corresponding to each sample prompt.

[0040] In a possible implementation manner, the first data is table data.

[0041] A third aspect of the present application provides a data processing device, which may include a processor coupled to a memory, wherein the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the method of the first aspect or any implementation of the first aspect is implemented. For details of the steps in each possible implementation of the first aspect executed by the processor, please refer to the first aspect and will not be repeated here.

[0042] In a fourth aspect, an embodiment of the present application provides a computing device cluster, which includes at least one computing device, each computing device including a processor and a memory, the memory storing a computer program or computer instructions, and the processor being used to call and run the computer program or computer instructions stored in the memory so that the computing device cluster executes the above-mentioned first aspect and any optional method thereof.

[0043] In a fifth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the method of any implementation manner of the first aspect.

[0044] The sixth aspect of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method of any implementation manner of the first aspect.

[0045] In a seventh aspect, the present application provides a chip system, which includes a processor for supporting a server or a data processing device to implement the functions involved in any implementation of the first aspect, for example, sending or processing the data and / or information involved in the above method. In one possible design, the chip system also includes a memory for storing program instructions and data necessary for the server or communication device. The chip system can be composed of a chip or can include a chip and other discrete devices.

[0046] The beneficial effects of the second to seventh aspects mentioned above can be referred to the introduction of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0048] Figure 1 A schematic diagram of a system architecture provided in an embodiment of the present application;

[0049] Figure 2 A flowchart of a data processing method provided in an embodiment of the present application;

[0050] Figure 3 A schematic diagram of the process of large language model training deployment provided in an embodiment of the present application;

[0051] Figure 4 A schematic diagram of the process of generating inference of a large language model provided in an embodiment of the present application;

[0052] Figure 5 A schematic diagram of the process of tracing the reference origin of a large language model provided in an embodiment of the present application;

[0053] Figure 6 An example of a data processing device provided in an embodiment of the present application is shown;

[0054] Figure 7 is a structural diagram of a computing device provided in an embodiment of the present application;

[0055] Figure 8 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0056] Figure 9 is another structural diagram of a computing device cluster provided in an embodiment of the present application;

[0057] Figure 10 A schematic diagram of the structure of the chip provided in an embodiment of the present application. DETAILED DESCRIPTION

[0058] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0059] The terms "first," "second," "third," "fourth," etc. (if any) in the specification and claims of the present application and in the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential sequence. It should be understood that the numbers used in this way are interchangeable where appropriate, so that the embodiments of the present application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or apparatus.

[0060] For ease of understanding, the following first introduces the system architecture used by the data processing method provided in the embodiments of the present application.

[0061] Figure 1 This is a schematic diagram of the system framework provided by this application. Figure 1 As shown, the structure may include a client 100, a data processing system 200, and a storage system 300. The communication connection between the client 100, the data processing system 200, and the storage system 300 may be a wired connection or a wireless connection, which is not specifically limited in this application. The number of clients 100 that establish communication connections with the data processing system 200 may be one or more, and the number of storage systems 300 that establish communication connections with the data processing system 200 may be one or more. Figure 1 A client 100 and a storage system 300 are used as an example for illustration, and this application does not make any specific limitation.

[0062] The client 100 is used to implement human-computer interaction and can be deployed on a terminal device, including a personal computer, a smart phone, a wearable device, a handheld processing device, a tablet computer, a mobile notebook, an augmented reality (AR) device, a virtual reality (VR) device, an integrated handheld game console, a wearable device, an in-vehicle device, an intelligent conference device, an intelligent advertising device, a smart home appliance, etc. The smart home appliance can be a sweeping robot, a mopping robot, etc., which are not specifically limited here. In a specific implementation, the client 100 can be a software or application running on a terminal device or computing device controlled by a user, such as a personal computer (PC) client, a browser-based World Wide Web (web) client, an application (APP) client running on a mobile terminal, or a cloud platform console, which are not specifically limited in this application.

[0063] Specifically, client 100 is used for data processing, data query, and data analysis in the financial sector, primarily for risk assessment during customer credit granting. Optionally, client 100 can be a cloud platform client, such as a cloud platform console, specifically a web-based console or an application programming interface (API)-based console, which is not specifically limited in this application. The console can provide data analysis cloud services to administrative users, and users can obtain access to the data processing system 200 provided in this application by purchasing cloud services.

[0064] The data processing system 200 and storage system 300 can be deployed on a computing device or computing device cluster, where computing devices include bare metal servers (BMS), virtual machines, containers, or edge computing devices. A BMS refers to a general-purpose physical server, such as an ARM server or an X86 server. A virtual machine refers to a complete computer system with complete hardware system functions, running in a completely isolated environment, simulated by software. Any work that can be performed on a physical computer can also be performed on a virtual machine. When creating a virtual machine on a computing device, part of the physical machine's hard disk and memory capacity is required as the virtual machine's hard disk and memory capacity. Each virtual machine has an independent basic input / output system (BIOS), hard disk, and operating system, and can be operated like a physical machine. A container is a portable software unit that can combine an application and all its dependencies into a single software package. This package is not limited by the underlying host operating system, eliminating the need to build a complex environment and simplifying the process from application development to deployment. Edge computing devices refer to devices that are closer to data sources and end users and have low latency and high bandwidth characteristics, such as smart routers, edge servers, etc. A computing device cluster can include multiple of the above computing devices, such as a data center, but this application does not specifically limit this. The storage system 300 may also be deployed in a storage array, such as a redundant array of independent disks (RAID), a network-attached storage (NAS), a storage area network (SAN), etc., which is not specifically limited in this application. Alternatively, the data processing system 200 and the storage system 300 may be deployed in the same computing device, or in different computing devices in the same computing device cluster, or in computing devices and storage arrays in the same computing device cluster, or in different computing device clusters, which is not specifically limited in this application.

[0065] Optionally, the data processing system 200 and the client 100 can be deployed in the same or different computing device clusters, for example, the client 100 is deployed on a terminal device in a first computing device cluster, and the data processing system 200 is deployed on a computing device in a second computing device cluster; or, the client 100 and the data processing system 200 are deployed in the same computing device cluster. It should be understood that the above examples are for illustration only and are not specifically limited in this application.

[0066] The functions of the large model reasoning module 210 and the induction module 220 in the data processing system 200 are described below.

[0067] The large model inference module 210 is responsible for interpreting and analyzing a customer's credit score. When the data processing system 200 receives a risk assessment request for a specific customer from the client 100, it first requests the customer's credit score data from the storage system 300. This credit score data is processed using prompt engineering techniques. The processed data is then fed into the trained large language model in the form of prompt text, generating an in-depth analysis and interpretation of the customer's credit status.

[0068] Optionally, the prompt engineering technology includes analyzing and generating overall trend description information and recent trend description information of credit score data; also includes statistical description information such as the maximum value, minimum value, mean, volatility, etc. in the score data; and also includes credit risk level information represented by score data in different intervals, etc.

[0069] The task of the induction module 220 is to integrate the credit score interpretation results output by the large model inference module 210 and combine them with other customer data stored in the storage system 300. Based on the risk summary template in the financial field, this module generates a risk summary for each customer and transmits this information to the client 100 for reference.

[0070] To facilitate understanding, the following first introduces the relevant terms and concepts mainly involved in the embodiments of this application.

[0071] (1) Large language model (LLM)

[0072] Large language models are deep learning models trained using large amounts of text data. They can generate natural language text or understand the meaning of text. Large language models can handle a variety of natural language tasks, such as text classification, question-answering, and conversation, and are an important path to artificial intelligence.

[0073] (2) Neural Network

[0074] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input. The output of the operation unit can be:

[0075]

[0076] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal of the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.

[0077] (3) Hallucinations

[0078] The hallucination problem refers to specific errors or misunderstandings that may occur when a large language model (LLM) processes and generates text. These hallucinations may arise from the model's insufficient understanding of the real world, biases in data training, or limitations of the algorithm itself.

[0079] (4) Textual diversity

[0080] Text diversity refers to the richness and variation within text. Text diversity assessment is often used to detect hallucinations between the LLM output and the input prompt. This involves comparing differences and variations in expression, vocabulary, and sentence structure between the two texts.

[0081] In this application, the text diversity of the two is measured by the distinct-n method, which calculates the number of different n-grams in the text, where n represents the order of the n-gram. The ratio of the number of non-repeated n-grams in the generated text to the total number of n-grams. If this ratio is too low, it means that there are a large number of repeated n-grams in the generated text, that is, hallucination problems occur.

[0082] (5) Parameter efficient fine tuning (PEFT)

[0083] PEFT is a parameter-efficient fine-tuning framework for large models. It aims to adapt pre-trained language models (PLMs) to various downstream applications without fine-tuning all model parameters. During fine-tuning, only a small number of parameters are optimized, resulting in relatively low training costs. PEFT supports multiple parameter fine-tuning methods, such as P-Tuning v2, P-Tuning, LoRA, Prefix Tuning, and AdaLoRA.

[0084] (6) Credit score

[0085] Credit scoring, also known as a credit risk scorecard, is a crucial component of a risk management system. It uses quantitative analysis to scientifically assess a customer's creditworthiness. This approach employs quantitative methods to assess credit risk, digitizing and standardizing a customer's credit profile. This allows for interpretable analysis of a customer's credit risk, enabling proactive risk prediction and mitigation, and driving business decisions with data. Credit risk scorecard models utilize econometric modeling based on historical data (such as an individual's credit history, debt level, and income status, or a company's identity, risk profile, and performance), to estimate the likelihood of a user defaulting due to their repayment ability. This approach creates a highly accurate and reliable credit scorecard. In the financial sector, credit scores are widely used in the approval process for personal and corporate loans and credit cards, helping to improve the objectivity and fairness of approval criteria. Credit scores also help financial institutions gain a more accurate and comprehensive understanding of the credit profile of borrowers and businesses, enabling them to make more informed lending decisions.

[0086] In this application, to interpret a customer's credit score, we extract the customer's credit score at multiple recent points in time from the database. Specifically, we obtain the customer's credit score data for each month over the past year. This continuous credit score data will form a credit time series score.

[0087] (7) Linear regression fitting

[0088] Linear regression fitting is a predictive modeling technique whose core lies in using regression analysis in mathematical statistics to determine the interdependence between two or more variables. This relationship can be approximated by a single straight line (univariate linear regression) or multiple straight lines (multivariate linear regression). In regression analysis, if there is only one independent variable and one dependent variable, and the relationship between them can be represented by a single straight line, then this regression analysis is called univariate linear regression analysis. If two or more independent variables are involved, and the relationship between the dependent and independent variables is linear, then this regression analysis is called multiple linear regression analysis.

[0089] The applicant has found through research that in current technology, when generating customer risk summaries, the common practice is to use text summarization algorithms or LLMs (such as chatglm, Pangu model, etc.) to extract text summaries of various types of customer information. However, these practices have limitations in practical applications. Most text summarization algorithms were originally designed to deal with general text scenarios such as news, so they are not ideal in specific applications in vertical fields such as finance. Similarly, those LLMs that rely on prompt engineering but have not been fine-tuned may produce logically disordered results in specific fields.

[0090] Based on this, a data processing method is provided in an embodiment of the present application. Figure 2 As shown, the data processing method provided in the embodiment of the present application includes the following steps 201-203.

[0091] Step 201: Acquire first data of a first customer.

[0092] In the embodiment of the present application, the first data of the first customer includes credit scores corresponding to multiple time points.

[0093] For example, the identification ID of the first customer is first obtained, and then the credit scores at multiple recent time points are obtained from the database based on the identification ID, for example, the credit scores corresponding to each month in the past year are obtained, as shown in Table 1. This continuous credit score can also be called a credit time series score.

[0094] time 202301 202302 202303 202304 202305 202306 202307 202308 202309 Credit score 801 801 801 801 801 801 758 748 663

[0095] Table 1

[0096] The first customer's credit score from January to June of 2023 is 801, while in July, August, and September, it is 758, 748, and 663, respectively. Furthermore, the first data may also obtain credit scores for a longer period of time, or the first data may also include descriptive information for each credit score, which is not limited here.

[0097] Step 202: Generate a target prompt based on prompt engineering processing according to the first data.

[0098] In the embodiment of the present application, prompt engineering processing is used to output a label of the credit status for the credit score, and the generated target prompt is used to define the credit status indicated by the first data.

[0099] In one possible implementation, the target prompt can include trend information describing the credit score over time. Based on linear regression fitting, time can be used as the independent variable, and the credit score corresponding to that time as the dependent variable. The resulting slope can be considered the trend of the credit score over time. The trend description information corresponding to the slope includes, for example, "customer credit is improving," "customer credit is stable," or "customer credit is deteriorating." Specifically, by calculating the credit score over time in Table 1, the trend description information can be "customer credit is deteriorating" or "customer credit score is continuously decreasing."

[0100] In one possible implementation, the target prompt may also include statistical descriptive information about the credit time series score. For example, statistical indicators may be extracted for credit time series evaluation, including the maximum, minimum, mean, and volatility of the credit score. An overall statistical analysis of the overall credit score may also be performed, as well as detailed statistical processing of the scores within a specific time period (e.g., the last three months). Specifically, by calculating the credit time series scores in Table 1, the statistical descriptive information may be "the average credit score in 2023 is 775 points," "the average credit score in the last three months is 723 points," and "the highest credit score is 801 points in January 2023."

[0101] In one possible implementation, the target prompt may also include risk description information related to the credit score. This can be combined with business knowledge, such as the credit risk level corresponding to the credit score. Specifically, a credit score range of "[750,+∞)" indicates "good credit status"; a score range of "[650,750)" indicates "average credit status"; a score range of "[0,650)" indicates "very poor credit status", etc. Then, combined with trend description information and statistical description information, the generated target prompts include, for example, "The average credit score in 2023 is 775 points, the overall credit is good, but the overall credit risk trend is deteriorating", "The average credit score in the past three months is 723 points, and the credit in the past three months is average", etc.

[0102] For example, according to the credit time series scoring in Table 1, the target prompt that can be derived is: "The current month has the lowest credit score, which is 663 points. The credit score of the current month is average, slightly lower than the overall level; the average credit score in 2023 is 775 points, and the overall credit is good, but the overall credit risk trend is getting worse; the average credit score in the past three months is 723 points, and the credit in the past three months is average; the highest credit score is 801 points in January 2023; the lowest credit score is 663 points in September 2023."

[0103] Step 203: Input the target prompt into the large language model to obtain the risk analysis result of the first customer.

[0104] In the embodiment of the present application, the large language model is a neural network model trained based on the training data in the training set, and the input parameter of the large language model is a prompt related to the customer's credit status description information.

[0105] For example, the target prompt obtained in step 202 is input into the large language model, and the inference output result obtained is: "The customer's credit score in the latest accounting period is 663, which is average credit; in the past 12 months, the customer's credit score has been declining and has become very poor, with large fluctuations; overall, the customer's credit score performance is poor and the credit risk is relatively high."

[0106] Exemplary large language models include the Chat General Language Model (ChatGLM), ChatGLM-6B, or the Pangu Large Model. The construction and application of these models often focus on common text, such as news reports. However, they are often insufficient for financial applications, particularly those involving credit risk assessment. Therefore, before deploying these large language models in real-world applications, they must be fine-tuned and trained to meet the specific needs of the financial sector to enhance their ability to understand and identify credit risk.

[0107] See also Figure 3 , Figure 3 Schematic diagram of the process for large language model training and deployment.

[0108] Credit score data is obtained from the database. Sample prompts are generated using the aforementioned prompt engineering method. The corresponding risk analysis reports are used as training samples for fine-tuning the large language model. In other words, the sample prompts and risk analysis reports help the large language model recognize prompts processed through prompt engineering, thereby enabling the large language model to recognize credit risk information.

[0109] In this embodiment, credit scores are processed using prompt engineering, with the customer's credit score annotated and output, generating prompts that can be recognized and processed by a large language model. This helps the large language model more quickly understand task requirements, enabling it to better understand and interpret credit score data, thereby generating more accurate risk analysis results. Fine-tuning the large language model through sample sets can make it more adaptable to specific tasks or domains, enabling the recognition of credit risk in the financial sector.

[0110] Next, we will introduce the usage process of the fine-tuned large language model after deployment. Figure 4 , Figure 4 A schematic diagram of the process of generating large language model inference according to an embodiment of the present application.

[0111] In one possible implementation, the large language model also includes a hallucination detection module. This module primarily addresses hallucination issues within large language models. After the inference module outputs the results, methods such as distinctN are used to detect text diversity between the output and the input prompt. If the detection result does not meet preset conditions, the system determines that a hallucination has occurred and instructs the inference module to regenerate the output. This continues until the output text meets the preset conditions. This text is then output as the risk analysis result, which can also be understood as the interpretation of the credit score.

[0112] In an embodiment of the present application, by comparing the text diversity between the input parameters and output parameters of the large language model, the large language model is helped to promptly detect hallucination problems and correct the output, making the content output by the model more accurate, reducing the inaccurate or misleading information encountered by users when using the model, and thus improving users' trust in the model.

[0113] Furthermore, after obtaining the risk analysis results for the first customer through the aforementioned data processing method, other data beyond the credit score can be further combined, such as factors influencing the credit score (such as the company's identity, risk performance, and ability to fulfill its obligations). A risk summary for the customer is generated and stored using a pre-set risk summary template. This risk summary provides reference information for decision-makers at financial institutions when granting credit to customers, helping them make more accurate decisions.

[0114] For example, if the first customer is a corporate user, obtaining other information about the first customer from the database includes:

[0115] (1) Identity and Qualifications: The company is a limited liability company registered in Beijing and operates mobile stall retail business. No shareholder information was found.

[0116] (2) Risk manifestation: The enterprise is in a continuing state. There are no abnormal situations (such as failure to publish annual reports), no cases of breach of trust (such as lawsuits), no cases of being subject to execution (details of lawsuits), and no restrictions on high consumption.

[0117] (3) Ability to perform: The company is not a listed company, with a registered capital of RMB 319,493,261. It is a limited liability company and operates in the mobile stall retail industry.

[0118] (4) Transaction records: No excessive overdue payments in that month.

[0119] (5) Behavioral performance: The ratio of the current month's bill amount to the total credit amount is 0.0, which is very low. The number of cloud services used is 1. The maximum amount paid for key products in the past six months is 0.0, which is very small.

[0120] It should be understood that the above only introduces the information queried when the first customer is an enterprise user. In actual applications, relevant information such as the person in charge of the enterprise may also be collected, which is not limited here.

[0121] Next, we combine the inference results output by the aforementioned large language model and use a risk summary template from the financial field to generate a risk summary for the customer:

[0122] 1. The customer's credit score in the latest accounting period is 663, which indicates average credit.

[0123] 2. The customer's credit score has been declining over the past 12 months and has become very poor, with significant fluctuations;

[0124] Scoring is based on the following:

[0125] (1) Identity and Qualifications: The company is a limited liability company registered in Beijing and operates mobile stall retail business. No shareholder information was found.

[0126] (2) Risk manifestation: The enterprise is in a continuing state. There are no abnormal situations (such as failure to publish annual reports), no cases of breach of trust (such as lawsuits), no cases of being subject to execution (details of lawsuits), and no restrictions on high consumption.

[0127] (3) Ability to perform: The company is not a listed company, with a registered capital of RMB 319,493,261. It is a limited liability company and operates in the mobile stall retail industry.

[0128] (4) Transaction records: No excessive overdue payments in that month.

[0129] (5) Behavioral performance: The ratio of the bill amount for the current month to the total credit amount is 0.0, which is a very low ratio.

[0130] To sum up, overall, the customer's credit score performance is poor and the credit risk is high.

[0131] Finally, the generated risk summary is stored in the database so that the risk control manager can access and review it at any time as a basis for making credit decisions for customers.

[0132] Furthermore, after generating the risk summary, the embodiment of the present application also provides a method for tracing the reference of a large language model, which associates the risk summary with the original information (i.e., the first data of the first customer) by sentence / key points to enhance the credibility of the risk summary.

[0133] See also Figure 5 , Figure 5 This is a flowchart for tracing the origin of a large language model, where each summary point of the risk summary is judged for similarity with the original information.

[0134] For example, summary point 1, "The customer's credit score for the latest accounting period is 663, indicating fair credit," is compared with the credit time series score in the original information. It can be determined that it has the highest similarity to the credit score of "663" at the time of "202309." The correspondence between the two is recorded. When the user views the risk summary, the original information is displayed in the browser indicator area through any triggering action, such as clicking or hovering the mouse over the summary point.

[0135] In the embodiments of the present application, by combining the risk analysis report generated by the credit score with the factors influencing the credit score, a more comprehensive understanding of the customer's credit status can be achieved, generating a summary of the customer's credit risk. This facilitates financial institutions to quickly assess the customer's creditworthiness and make more accurate decisions when granting credit. Furthermore, establishing a correspondence between each summary point and the original information enables decision-makers, such as risk managers, to quickly locate the original information corresponding to each summary point, providing a detailed explanation of the summary point and enhancing the credibility and practicality of the risk summary.

[0136] The above describes in detail the method provided by the embodiment of the present application. Next, the device provided by the embodiment of the present application for executing the above method will be introduced.

[0137] See also Figure 6 , Figure 6 This is a structural diagram of a data processing device provided in an embodiment of the present application. Figure 6 As shown, the data processing device provided in the embodiment of the present application includes:

[0138] An acquisition module 601 is configured to acquire first data of a first customer, the first data including credit scores of the first customer at multiple time points.

[0139] Processing module 602 is configured to generate a target prompt based on the first data and prompt engineering processing, wherein the target prompt is used to describe the credit status indicated by the first data, and the prompt engineering processing is used to output a credit status annotation based on the credit score;

[0140] The processing module 602 is further configured to input the target prompt into the large language model to obtain a risk analysis result for the first customer.

[0141] In one possible implementation, the acquisition module 601 is further used to: obtain second data of the first customer, the second data including influencing factors of the credit score; the processing module is further used to generate a risk summary based on the second data and the risk analysis results, and the risk summary is used to describe the credit credibility of the first customer.

[0142] In one possible implementation, the risk summary includes at least one summary point, and the processing module 602 is further used to: perform a similarity judgment between the first summary point and the first data to obtain first text information whose similarity is greater than a first threshold, where the first text information is included in the first data, and the first summary point is any one of the at least one summary points; and record a correspondence between the first text information and the first summary point, where the correspondence is used to indicate a position of the first text information in the first data.

[0143] In one possible implementation, the processing module 602 is specifically used to: input the target prompt into the large language model to obtain a first text; detect the text diversity between the text data of the target prompt and the text data of the first text; when the text diversity meets a preset condition, use the first text as a risk analysis result.

[0144] In a possible implementation, when the text diversity does not meet the preset condition, the large language model is used to update the first text until the updated text meets the preset condition.

[0145] In one possible implementation, processing module 602 is specifically configured to: calculate a slope based on linear regression fitting according to the credit scores corresponding to multiple time points, where the slope is used to indicate the trend of the credit score in historical time; and generate trend description information based on the slope, where the target prompt includes the trend description information.

[0146] In one possible implementation, the processing module 602 is further used to extract statistical indicators of the credit score to obtain statistical description information, where the statistical indicators include one or more of the maximum value, minimum value, mean value, and volatility of the credit score, and the target prompt includes the statistical description information.

[0147] In one possible implementation, the processing module 602 is further used to match the credit score with the business knowledge information to obtain risk description information, where the business knowledge information is used to indicate the credit risk level represented by the credit score in different intervals, and the target prompt includes the risk description information.

[0148] In a possible implementation, the large language model is a large language model fine-tuned using a sample set, where the sample set includes multiple sample prompts and a risk analysis report corresponding to each sample prompt.

[0149] In a possible implementation manner, the first data is table data.

[0150] The acquisition module and the processing module can be implemented by software or hardware. For example, the implementation of the acquisition module will be described below using the acquisition module as an example. Similarly, the implementation of the processing module can refer to the implementation of the acquisition module.

[0151] As an example of a software functional unit, the module acquisition module may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the acquisition module may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0152] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0153] As an example of a hardware functional unit, the acquisition module may include at least one computing device, such as a server. Alternatively, the acquisition module may be implemented using a central processing unit (CPU), an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system on chip (SoC), an offload card, an accelerator card, or any combination thereof.

[0154] The multiple computing devices included in the acquisition module can be distributed in the same region or in different regions. The multiple computing devices included in the acquisition module can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the acquisition module can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offload cards, accelerator cards, and other computing devices.

[0155] It should be noted that, in other embodiments, the acquisition module can be used to execute any step in the data processing method, and the processing module can be used to execute any step in the data processing method. The steps that the acquisition module and the processing module are responsible for implementing can be specified as needed. The full functions of the data processing device are realized by respectively implementing different steps in the data processing method through the acquisition module and the processing module.

[0156] The present application also provides a computing device 100. Figure 7As shown, computing device 100 includes a bus 102, a processor 104, a memory 106, and a communication interface 108. Processor 104, memory 106, and communication interface 108 communicate with each other via bus 102. Computing device 100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 100.

[0157] The bus 102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 The bus 104 may include a path for transmitting information between various components of the computing device 100 (eg, memory 106, processor 104, communication interface 108).

[0158] The processor 104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0159] The memory 106 may include volatile memory, such as random access memory (RAM). The processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0160] The memory 106 stores executable program codes, and the processor 104 executes the executable program codes to respectively implement the functions of the aforementioned acquisition module and processing module, thereby implementing the data processing method. That is, the memory 106 stores instructions for executing the data processing method.

[0161] Alternatively, the memory 106 stores executable codes, and the processor 104 executes the executable codes to respectively implement the functions of the aforementioned data processing apparatus, thereby implementing the data processing method. That is, the memory 106 stores instructions for executing the data processing method.

[0162] The communication interface 103 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or a communication network.

[0163] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0164] like Figure 8 As shown, the computing device cluster includes at least one computing device 100. The memory 106 in one or more computing devices 100 in the computing device cluster may store the same instructions for executing the data processing method.

[0165] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store some instructions for executing the data processing method. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for executing the data processing method.

[0166] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions, each for executing part of the functions of the data processing device. In other words, the instructions stored in the memory 106 in different computing devices 100 can implement the functions of one or more devices in the acquisition module and the processing module.

[0167] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 9 A possible implementation is shown. Figure 9 As shown, two computing devices 100A and 100B are connected via a network. Specifically, each computing device is connected to the network via a communication interface within the computing device. In this possible implementation, the memory 106 within computing device 100A stores instructions for executing the functions of the acquisition module. Simultaneously, the memory 106 within computing device 100B stores instructions for executing the functions of the processing module.

[0168] Figure 9The connection method between the computing device clusters shown can be based on the consideration that the data processing method provided in this application requires a large amount of storage data in the associated object storage service, so the function of executing the second SQL statement is considered to be handed over to the computing device 100B.

[0169] It should be understood that Figure 9 The functions of the computing device 100A shown in FIG. 1 may also be completed by multiple computing devices 100. Similarly, the functions of the computing device 100B may also be completed by multiple computing devices 100.

[0170] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the data processing method.

[0171] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the data processing method.

[0172] The computing device provided in the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, or a circuit. The processing unit may execute computer-executable instructions stored in the storage unit so that the chip in the execution device executes the data processing method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0173] For details, please refer to Figure 10 , Figure 10This is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 1000. NPU 1000 is mounted on the host CPU as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is arithmetic circuit 1003, which is controlled by controller 1004 to extract matrix data from memory and perform multiplication operations.

[0174] In some implementations, the arithmetic circuit 1003 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1003 is a two-dimensional systolic array. The arithmetic circuit 1003 can also be a one-dimensional systolic array or other electronic circuit capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1003 is a general-purpose matrix processor.

[0175] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1002 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1001 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1008.

[0176] Unified memory 1006 is used to store input and output data. Weight data is directly transferred to weight memory 1002 through the Direct Memory Access Controller (DMAC) 1005. Input data is also transferred to unified memory 1006 through the DMAC.

[0177] BIU stands for Bus Interface Unit, i.e., bus interface unit 1010 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1009 .

[0178] The bus interface unit 1010 (BIU) is used for the instruction fetch memory 1009 to obtain instructions from the external memory, and is also used for the storage unit access controller 1005 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0179] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1006 or transfer weight data to the weight memory 1002 or transfer input data to the input memory 1001.

[0180] The vector calculation unit 1007 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0181] In some implementations, the vector calculation unit 1007 can store the processed output vector to the unified memory 1006. For example, the vector calculation unit 1007 can apply a linear function or a nonlinear function to the output of the operation circuit 1003, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, the vector calculation unit 1007 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1003, for example, for use in subsequent layers in a neural network.

[0182] An instruction fetch buffer 1009 connected to the controller 1004 is used to store instructions used by the controller 1004;

[0183] Unified memory 1006, input memory 1001, weight memory 1002, and instruction fetch memory 1009 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0184] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0185] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0186] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A data processing method, characterized in that: include: Acquire first data of a first customer, where the first data includes credit scores of the first customer at multiple time points; generating a target prompt based on the first data and prompt engineering processing, wherein the target prompt is used to describe the credit status indicated by the first data, and the prompt engineering processing is used to output a credit status annotation for the credit score; The target prompt is input into a large language model to obtain a risk analysis result of the first customer.

2. The method according to claim 1, characterized in that The method further comprises: Acquiring second data of the first customer, where the second data includes an influencing factor of the credit score; A risk summary is generated based on the second data and the risk analysis result, where the risk summary is used to describe the credit credibility of the first customer.

3. The method according to claim 2, characterized in that The risk summary includes at least one summary point, and the method further includes: performing similarity determination on the first summary point and the first data to obtain first text information whose similarity is greater than a first threshold, wherein the first text information is included in the first data, and the first summary point is any one of the at least one summary point; A correspondence between the first text information and the first summary point is recorded, where the correspondence is used to indicate a position of the first text information in the first data.

4. The method according to any one of claims 1 to 3, characterized in that Inputting the target prompt into the large language model to obtain a risk analysis result of the first customer includes: Inputting the target prompt into the large language model to obtain a first text; detecting text diversity between the text data of the target prompt and the text data of the first text; When the text diversity meets a preset condition, the first text is used as the risk analysis result.

5. The method according to claim 4, characterized in that When the text diversity does not meet the preset condition, the large language model is used to update the first text until the updated text meets the preset condition.

6. The method according to any one of claims 1 to 5, characterized in that Generating a target prompt based on prompt engineering processing according to the first data includes: Calculating a slope based on linear regression fitting according to the credit scores corresponding to the multiple time points, wherein the slope is used to indicate the trend of the credit score in historical time; Trend description information is generated according to the slope, and the target prompt includes the trend description information.

7. The method according to claim 6, characterized in that The method further comprises: Statistical indicators are extracted from the credit score to obtain statistical description information, where the statistical indicators include one or more of the maximum value, minimum value, mean value, and volatility of the credit score, and the target prompt includes the statistical description information.

8. The method according to claim 6 or 7, characterized in that The method further comprises: The credit score is matched with the business knowledge information to obtain risk description information. The business knowledge information is used to indicate the credit risk level represented by the credit score with different score intervals. The target prompt includes the risk description information.

9. The method according to any one of claims 1 to 8, characterized in that The large language model is a large language model fine-tuned by a sample set, and the sample set includes multiple sample prompts and a risk analysis report corresponding to each sample prompt.

10. The method according to any one of claims 1 to 9, characterized in that The first data is table data.

11. A data processing device, characterized in that: include: an acquisition module, configured to acquire first data of a first customer, the first data including credit scores of the first customer corresponding to multiple time points; a processing module configured to generate a target prompt based on the first data and prompt engineering processing, wherein the target prompt is used to describe the credit status indicated by the first data, and the prompt engineering processing is used to output a credit status annotation for the credit score; The processing module is further configured to input the target prompt into a large language model to obtain a risk analysis result of the first customer.

12. The device according to claim 11, characterized in that The acquisition module is further used to: Acquiring second data of the first customer, where the second data includes an influencing factor of the credit score; The processing module is further configured to generate a risk summary based on the second data and the risk analysis result, wherein the risk summary is used to describe the credit credibility of the first customer.

13. The device according to claim 12, characterized in that The risk summary includes at least one summary point, and the processing module is further configured to: performing similarity determination on the first summary point and the first data to obtain first text information whose similarity is greater than a first threshold, wherein the first text information is included in the first data, and the first summary point is any one of the at least one summary point; A correspondence between the first text information and the first summary point is recorded, where the correspondence is used to indicate a position of the first text information in the first data.

14. The device according to any one of claims 11 to 13, characterized in that The processing module is specifically used to: Inputting the target prompt into the large language model to obtain a first text; detecting text diversity between the text data of the target prompt and the text data of the first text; When the text diversity meets a preset condition, the first text is used as the risk analysis result.

15. The device according to claim 14, characterized in that When the text diversity does not meet the preset condition, the large language model is used to update the first text until the updated text meets the preset condition.

16. The device according to any one of claims 11 to 15, characterized in that The processing module is specifically used to: Calculating a slope based on linear regression fitting according to the credit scores corresponding to the multiple time points, wherein the slope is used to indicate the trend of the credit score in historical time; Trend description information is generated according to the slope, and the target prompt includes the trend description information.

17. The device according to claim 16, characterized in that The processing module is further configured to: Statistical indicators are extracted from the credit score to obtain statistical description information, where the statistical indicators include one or more of the maximum value, minimum value, mean value, and volatility of the credit score, and the target prompt includes the statistical description information.

18. The device according to claim 16 or 17, characterized in that The processing module is further configured to: The credit score is matched with the business knowledge information to obtain risk description information. The business knowledge information is used to indicate the credit risk level represented by the credit score with different score intervals. The target prompt includes the risk description information.

19. The device according to any one of claims 11 to 18, characterized in that The large language model is a large language model fine-tuned by a sample set, and the sample set includes multiple sample prompts and a risk analysis report corresponding to each sample prompt.

20. The device according to any one of claims 11 to 19, characterized in that The first data is table data.

21. A data processing device, characterized in that: The device includes a memory and a processor; the memory stores codes, and the processor is configured to obtain the codes and execute the method according to any one of claims 1 to 10.

22. A computing device cluster, characterized in that: comprising at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so that the computing device cluster performs the operating steps of the method according to any one of claims 1 to 10.

23. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 10.

24. A computer-readable storage medium, characterized in that The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 10.