Data processing method and device, computer equipment and storage medium

By combining the trusted classification model and the confidence sub-model, the problem of large models being unable to evaluate accuracy in scenarios such as medical diagnosis and financial risk control is solved. The quantification of the confidence level of the results and the output of the degree of credibility are achieved, broadening the application scope of large models.

CN120781202APending Publication Date: 2025-10-14PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510871403.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing technologies cannot effectively evaluate the accuracy of large model prediction and classification results, which may lead to misjudgments, especially in high-risk scenarios such as medical diagnosis and financial risk control. In addition, existing methods for obtaining confidence consume computing resources or lack confidence accuracy.

Method used

A trusted classification model is used, including a classification sub-model and an external confidence sub-model. Prompt data is generated by preprocessing data, and combined with a preset confidence analysis strategy, the classification results and the corresponding first and second classification probabilities are output, and the confidence of the classification results is finally analyzed.

Benefits of technology

It is possible to quantify the credibility of the prediction results of large models without increasing too much computing resources, broaden the application scope of large models in serious scenarios, and overcome the overconfidence tendency of large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120781202A_ABST
    Figure CN120781202A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a data processing method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining to-be-processed data, and carrying out the preprocessing of the to-be-processed data, so as to obtain prompt data corresponding to the to-be-processed data; the prompt data is input into a pre-trained complete credible classification model to obtain a classification result and a corresponding first classification probability and second classification probability, and the credible classification model comprises a classification sub-model and a confidence sub-model externally hung on the classification sub-model; and calculating the confidence coefficient of the classification result based on the first classification probability and the second classification probability. The method can be applied to business system platforms of financial science and technology, medical health and the like, and solves the technical problem that the accuracy of a large model prediction classification result cannot be evaluated in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a data processing method and device, computer equipment and storage medium. BACKGROUND

[0002] Large language models have shown significant performance advantages in text classification tasks such as emotion classification and topic identification due to their strong semantic understanding and generalization capabilities, leading to a large number of efficient classification methods based on large models.

[0003] Specifically, in the field of financial risk control, such as credit card approval, personal loan risk assessment, etc., banks can use classification models to classify the credibility of users based on customer's basic information, consumption behavior, etc. data, and then determine whether to approve the credit card application and how much credit limit to give, or when assessing the solvency of an enterprise, use classification models to classify enterprises into two categories: solvency and non-solvency, based on financial ratios and other indicators, to help financial institutions decide whether to provide loans to enterprises.

[0004] In the field of medical diagnosis, efficient classification based on large models can be used in disease screening (such as diabetes, early cancer prediction), inpatient mortality risk assessment, etc. In addition, pathological images of patients can also be classified and judged, such as cancer classification (such as breast cancer benign and malignant judgment), pathological image classification, etc.

[0005] However, the existing technology has a method of missing confidence information, such as only outputting classification results (such as "positive" or "negative"), which cannot provide confidence quantification indicators associated with the results, making it difficult for users to evaluate the reliability of the results, especially in high-risk scenarios such as medical diagnosis and financial risk control, which may cause misjudgment risks. There are some methods to obtain confidence, such as the self-consistency method, which generates multiple candidate results by sampling multiple times and calculates the consistency ratio, which can indirectly reflect the confidence, but requires a large amount of computing resources. In addition, the large model can also be required to output the confidence score (such as "confidence 85%") simultaneously when generating the classification result, but due to the inherent "overconfidence" tendency of the large model and the lack of external verification mechanism, the confidence output by the large model often deviates significantly from the true accuracy. SUMMARY

[0006] The purpose of the present application is to overcome the above technical deficiencies and provide a data processing method, device, computer equipment and storage medium to solve the technical problem of the inability to evaluate the accuracy of the classification results predicted by the large model in the prior art.

[0007] To achieve the above technical purpose, the present application adopts the following technical solutions:

[0008] In a first aspect, the present invention provides a data processing method, comprising the following steps:

[0009] Acquire data to be processed, and pre-process the data to be processed to obtain prompt data corresponding to the data to be processed;

[0010] Inputting the prompt data into a pre-trained credible classification model to obtain a classification result and corresponding first classification probability and second classification probability, wherein the credible classification model includes a classification sub-model and a confidence sub-model externally mounted on the classification sub-model, the classification sub-model is used to output the classification result and the corresponding first classification probability, and the confidence sub-model is used to output the second classification probability;

[0011] Based on the first classification probability and the second classification probability, the confidence of the classification result is analyzed using a preset confidence analysis strategy.

[0012] In a second aspect, the present invention further provides a data processing device, comprising:

[0013] A data acquisition module is used to acquire data to be processed and pre-process the data to be processed to obtain prompt data corresponding to the data to be processed;

[0014] a classification probability calculation module, configured to input the prompt data into a pre-trained credible classification model to obtain a classification result and corresponding first and second classification probabilities, wherein the credible classification model includes a classification sub-model and a confidence sub-model externally mounted on the classification sub-model, the classification sub-model being configured to output the classification result and corresponding first classification probability, and the confidence sub-model being configured to output the second classification probability;

[0015] The confidence calculation module is used to analyze the confidence of the classification result based on the first classification probability and the second classification probability through a preset confidence analysis strategy.

[0016] In a third aspect, the present invention further provides a computer device comprising a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the data processing method described above when executing the computer-readable instructions.

[0017] In a fourth aspect, the present invention further provides a computer-readable storage medium having computer-readable instructions stored thereon, and the computer-readable instructions, when executed by a processor, implement the steps of the data processing method described above.

[0018] Compared with the prior art, the data processing method, device, computer equipment and storage medium provided by the present application first acquire the to-be-processed data, pre-process the to-be-processed data to obtain prompt data corresponding to the to-be-processed data; then input the prompt data into a pre-trained complete trusted classification model to obtain a classification result and corresponding first and second classification probabilities, wherein the trusted classification model includes a classification sub-model and a confidence sub-model externally connected to the classification sub-model, the classification sub-model is used to output the classification result and the corresponding first classification probability, and the confidence sub-model is used to output the second classification probability; and finally, based on the first and second classification probabilities, the confidence of the classification result is analyzed through a preset confidence analysis strategy. The present application combines the reasoning and generalization ability of a large model, and at the same time, for the downstream classification task, the confidence of the result can be obtained while the result is outputted, and the degree of confidence of the output result is measured, which to some extent, widens the applicability of the large model in some serious scenarios, such as medical diagnosis, financial risk control, etc. In addition, the present application uses the way of externally connecting the confidence network, on the one hand, the resources additionally occupied are not much, and on the other hand, in a way of approximate external verification mechanism, the inherent "overconfidence" tendency of the large model is overcome. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the schemes in the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without any creative effort.

[0020] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0021] Figure 2 is a flowchart of one embodiment of the data processing method according to the present application;

[0022] Figure 3 is Figure 2 is a flowchart of one specific embodiment of the step S100 shown in the figure

[0023] Figure 4 is Figure 2 is a flowchart of one specific embodiment of the step S200 shown in the figure

[0024] Figure 5 is Figure 2 is a flowchart of one specific embodiment of the step S300 shown in the figure

[0025] Figure 6 is a structural schematic diagram of one embodiment of the data processing device according to the present application;

[0026] Figure 7 is a schematic structural diagram of an embodiment of a computer device according to the present invention;

[0027] Figure 8 FIG. 1 is a schematic structural diagram of another embodiment of a computer device according to the present invention. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0029] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0030] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0031] The data processing method based on artificial intelligence provided by the embodiment of the present invention can be applied in Figure 1In an application environment, the client communicates with the server through a network. The server can first obtain the data to be processed through the client, pre-process the data to be processed, and obtain prompt data corresponding to the data to be processed; then input the prompt data into a pre-trained and complete trusted classification model to obtain a classification result and a corresponding first classification probability and a second classification probability, wherein the trusted classification model includes a classification sub-model and a confidence sub-model externally mounted on the classification sub-model, the classification sub-model is used to output the classification result and the corresponding first classification probability, and the confidence sub-model is used to output the second classification probability; finally, based on the first classification probability and the second classification probability, the confidence of the classification result is analyzed through a preset confidence analysis strategy. The present invention combines the reasoning and generalization capabilities of large models, and at the same time can output results for downstream classification tasks while obtaining the confidence of the results, and measuring the credibility of the output results. This has broadened the usability of large models in some serious scenarios to a certain extent, such as medical diagnosis, financial risk control, etc. In addition, the present invention uses an external confidence network, which, on the one hand, does not take up many additional resources, and on the other hand, overcomes the inherent "overconfidence" tendency of large models in a way that is similar to an external verification mechanism. Among them, the client can be, but is not limited to, various personal computers, laptops, smart phones, tablets and portable wearable devices. The server can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.

[0032] See also Figure 2 , Figure 2 A flowchart of an embodiment of a data processing method according to the present invention is shown. The data processing method is applicable to user classification scenarios in financial risk control scenarios or intelligent diagnosis in medical diagnosis scenarios, and includes steps S100 to S300.

[0033] S100: Acquire data to be processed, and pre-process the data to be processed to obtain prompt data corresponding to the data to be processed.

[0034] In this embodiment, the data to be processed must first be converted to a data format that can be recognized by the model, thereby facilitating subsequent model processing of the data. Specifically, the data to be processed can be processed through prompt engineering to generate prompt data corresponding to the data to be processed. Among them, prompt engineering is used to guide the large language model (LLM) to generate output that meets specific requirements. By designing and writing prompt text, based on natural language processing technology, the task description is embedded in the input, thereby guiding the model to generate language output that meets specific requirements. In this embodiment, zero-shot prompts, few-shot prompts, or thought chain prompts can be used to generate prompt data. Among them, zero-shot prompts refer to providing a task that has not been explicitly trained to the machine learning model to test the model's ability to generate relevant output without relying on previous examples. Few-shot prompts refer to providing the model with some example outputs to help it understand the requester's intentions. Thought chain prompts are thought chain prompts: breaking down complex tasks into intermediate steps or "reasoning chains" helps the model achieve better language understanding and create more accurate outputs.

[0035] S200. Input the prompt data into a pre-trained and complete credible classification model to obtain a classification result and a corresponding first classification probability and a second classification probability, wherein the credible classification model includes a classification sub-model and a confidence sub-model externally mounted on the classification sub-model, the classification sub-model is used to output the classification result and the corresponding first classification probability, and the confidence sub-model is used to output the second classification probability.

[0036] In this embodiment, in order to obtain the accuracy of the classification result, the classification result and the probability corresponding to the classification result are output through the pre-established nemesis classification model. Specifically, for an input x to be classified i , after the prompt project, we get , and then through the credible classification model based on the large language model, the output y of the large model is obtained respectively llm , and the probability y obtained by the external confidence sub-model con .

[0037] S300: Analyze the confidence of the classification result based on the first classification probability and the second classification probability using a preset confidence analysis strategy.

[0038] In this embodiment, after obtaining the first classification probability and the second classification probability, the confidence of the classification result can be calculated using the pre-established confidence calculation equation, and then the credibility of the output result can be measured, thereby enhancing the powerful pan-China thrust capability of the large model, and at the same time solving the problem that the large model cannot obtain confidence.

[0039] Specifically, in the field of financial technology, taking the identification of user risk level as an example, first input the user's insurance-related data, then process the user's insurance-related data through the prompt project to obtain the prompt data corresponding to the insurance-related data, and then input the prompt data into the trusted classification model. The trusted classification model will output the risk level corresponding to the user and the first classification probability and second classification probability corresponding to the risk level. Then, the confidence level of the risk level output by the model can be obtained through the first classification probability and the second classification probability.

[0040] Similarly, in the field of medical health, taking medical diagnosis as an example, the user's pathological text data is first obtained, and then the pathological text data is processed through the prompt project to obtain the prompt data corresponding to the pathological text data, and then the prompt data is input into the trusted classification model. The trusted classification model will output the diagnosis result corresponding to the user and the first classification probability and second classification probability corresponding to the diagnosis result. Then, the confidence of the diagnosis result output by the model can be obtained through the first classification probability and the second classification probability.

[0041] In an embodiment of the present invention, the data to be processed is first obtained and pre-processed to obtain prompt data corresponding to the data to be processed; the prompt data is then input into a pre-trained trustworthy classification model to obtain a classification result and a corresponding first classification probability and a second classification probability, wherein the trustworthy classification model includes a classification sub-model and a confidence sub-model externally attached to the classification sub-model, the classification sub-model is used to output the classification result and the corresponding first classification probability, and the confidence sub-model is used to output the second classification probability; finally, based on the first classification probability and the second classification probability, the confidence of the classification result is analyzed by a preset confidence analysis strategy. The present invention combines the reasoning and generalization capabilities of the large model, and at the same time, for downstream classification tasks, it can output the results while obtaining the confidence of the results and measuring the credibility of the output results. This, to a certain extent, broadens the applicability of the large model in some serious scenarios, such as medical diagnosis, financial risk control, etc. In addition, the present invention uses an external confidence network, which, on the one hand, does not occupy many additional resources, and on the other hand, overcomes the inherent "overconfidence" tendency of the large model in a way that is similar to an external verification mechanism.

[0042] In some embodiments, see Figure 3 , the step S100 specifically includes:

[0043] S110, obtaining data to be processed;

[0044] S120, performing keyword extraction on the data to be processed to extract keywords from the data to be processed;

[0045] S130: Input the keyword into a pre-built prompt model to generate prompt data corresponding to the data to be processed.

[0046] In this embodiment, the data to be processed is first obtained. The data to be processed can be of various types, such as emotional text data. Then, keyword extraction is performed on the data to be processed to facilitate the subsequent generation of prompt data. In specific implementation, natural language processing (NLP) technology, such as TF-IDF algorithm, TextRank algorithm, deep learning model, etc., can be used to extract keywords from a large amount of text, and then the keywords are input into the prompt model to obtain prompt data corresponding to the data to be processed, so that the model can clarify the task objectives, output format and logical path, thereby improving accuracy and efficiency, and reducing dependence on a large amount of labeled data.

[0047] For example, taking text sentiment classification as an example, the input text is represented as x i , and the corresponding label is represented as y i Since the present invention uses a large model for classification, it is necessary to input x i Transform it into a prompt form that the large model can understand. For example:

[0048] Please help me identify the "| <x i >|” What emotion does this text convey?

[0049] In order to improve the generalization ability of the model, the present invention will target the same x i Set multiple prompt forms, expressed as Where j represents the jth prompt form.

[0050] In some embodiments, see Figure 4 , the step S200 specifically includes:

[0051] S210: Construct a confidence sub-model, and train the confidence sub-model based on a pre-constructed training data set to obtain a fully trained confidence sub-model;

[0052] S220: constructing a credible classification model based on the fully trained confidence sub-model and the pre-established classification sub-model;

[0053] S230: Input the prompt data into the trusted classification model to obtain a classification result and corresponding first classification probability and second classification probability.

[0054] In this embodiment, the confidence sub-model is externally connected to the classification sub-model. During training, the structure of the classification sub-model remains unchanged, that is, the large language model is fixed, only the attention layer and the subsequent linear layer used for classification are trained, and then the establishment of the externally connected confidence sub-model is realized. Then, by combining the externally connected confidence sub-model and the classification sub-model, the establishment of the trusted classification model is realized. After inputting the prompt data, the classification result, the first classification probability and the second classification probability can be directly outputted by the trusted classification model.

[0055] In some embodiments, the step S210 specifically comprises:

[0056] The text data set is obtained, prompt engineering processing is performed on the text data set, a plurality of pieces of prompt data are generated, and the plurality of pieces of prompt data are used as a training data set;

[0057] A confidence sub-model is constructed, the training data set is used to train the confidence sub-model, and a trained confidence sub-model is obtained;

[0058] The confidence sub-model is optimized based on a pre-established test set, and an optimized confidence sub-model is obtained.

[0059] In this embodiment, in order to realize the establishment of the confidence sub-model, the model needs to be trained by using a plurality of pieces of prompt data. Specifically, taking text emotion classification as an example, the prompt data can be "please help me identify <x>| What emotion is this text trying to convey? | <x>|"What emotion does this text express?", "| <x>| "What's the emotion of the above text?" and other prompts exist. By constructing a training dataset with multiple prompts, the reasoning ability and applicability of the model can be enhanced.

[0060] After the model is established, it is trained using the training data set, and finally optimized using the test set. By adjusting the model or data, the model's performance on the test set is optimized, thereby increasing the accuracy of the model's output results.

[0061] In some embodiments, step S230 specifically includes:

[0062] Inputting the prompt data into the classification sub-model and calculating the attention weight of the classification sub-model;

[0063] Based on the attention weight of the classification sub-model, calculating the classification result and the corresponding first classification probability through the output layer of the classification sub-model;

[0064] The attention weight is input into the confidence sub-model to calculate the second classification probability.

[0065] In this embodiment, the input of the model is , after passing each layer of the pre-trained transformer model, the features of the last position are collected to obtain z k , where k represents the kth layer. The vectors of each layer are then compressed into a vector through an attention layer, and finally passed through a linear layer to obtain the specific classification result. Specifically, the calculation method of the attention layer is as follows:

[0066] α k =Softmax(V T ReLU(Wz k +b1)+b2)

[0067] First, for the output z of each layer k , calculate the attention weight α according to the above formula k W, b1, V, and b2 are all learnable parameters. ReLU() and Softmax() represent activation functions, and T represents matrix transpose. ReLU() is simple to calculate, alleviates gradient vanishing (the gradient in the positive interval is always 1), and has sparse activation. Softmax() can convert the input into a probability distribution and is suitable for multi-classification output layers.

[0068] In some embodiments, inputting the attention weight into the confidence sub-model to calculate the second classification probability includes:

[0069] Based on the confidence sub-model, obtaining a preset compression vector calculation equation;

[0070] Inputting the attention weight into a preset compression vector calculation equation to calculate a compression vector;

[0071] The compressed vector is input into the fully connected layer of the confidence sub-model to obtain a second classification probability.

[0072] In this embodiment, an external confidence network is constructed by extracting deep features from a large language model to predict the confidence level of the solution. Specifically, when calculating the probability of the second classification, the compression vector is first calculated using the attention weight and the features of the last position of each transformer layer. The specific formula is as follows:

[0073]

[0074] Where N represents the total number of transformer layers. After that, it passes through a fully connected layer and outputs the classification probability.

[0075] It should be noted that during the model training process, the transformer layer (i.e., the large language model) will remain fixed, and only the attention layer and subsequent linear layers used for classification will be trained. This paper recommends using weighted cross entropy as the loss function for model training.

[0076] Among them, weighted cross entropy is an improved cross entropy loss function. By assigning weights to different categories, it solves the problem of data imbalance and improves the model's attention to minority classes. It is suitable for class imbalance problems and class importance difference problems. For example, when the number of samples in some categories is far greater than that in other categories (such as medical diagnosis and fraud detection), standard cross entropy may cause the model to favor the majority class. Weighted cross entropy can balance the loss through weights. In some tasks, the error costs of different categories are different (such as misjudging cancer as healthy is more serious than vice versa), and the importance difference can be reflected through weights.

[0077] In some embodiments, see Figure 5 , the step S300 specifically includes:

[0078] S310, obtaining a pre-established confidence calculation model;

[0079] S320: Input the first classification probability and the second classification probability into the confidence calculation model to obtain a calculation result;

[0080] S330: Determine the confidence level of the classification result based on the calculation result.

[0081] In this embodiment, after the model training converges, the model can be inferred in the following way. First, for an input x to be classified i , after the prompt project, we get Then, the output y of the large model is obtained through the credible classification model based on the large language model. llm , and the probability y obtained by the external confidence module con Then, the confidence level p of the classification result is further obtained by the following method:

[0082]

[0083] in, Indicates the probability that the classification result of the large model is c, and its value is 0 or 1. Indicates the probability that the confidence model is classified as c, with a value between 0 and 1.

[0084] The data processing provided by the embodiment of the present invention has the main purpose of combining the reasoning and generalization capabilities of a large model, and at the same time being able to output results for downstream classification tasks while obtaining the confidence of the results, so as to be applicable to certain specific application scenarios, such as medical diagnosis, financial risk control, etc. The core idea of ​​the present invention is to construct an external confidence network based on the deep features extracted from the large language model to achieve the prediction of the credibility of the results. By constructing a training data set, a confidence network model is constructed according to the training data set, and finally, the classification results and their corresponding confidence can be quickly obtained through model reasoning, thereby measuring the credibility of the output results.

[0085] The technical solution provided by the present invention first obtains the data to be processed and preprocesses it to obtain prompt data corresponding to the data to be processed; then the prompt data is input into a pre-trained and fully trusted classification model to obtain a classification result and the corresponding first and second classification probabilities, wherein the trusted classification model includes a classification sub-model and a confidence sub-model externally attached to the classification sub-model, wherein the classification sub-model is used to output the classification result and the corresponding first classification probability, and the confidence sub-model is used to output the second classification probability; finally, based on the first and second classification probabilities, the confidence of the classification result is analyzed using a preset confidence analysis strategy. The present invention combines the reasoning and generalization capabilities of the large model, and can simultaneously output results for downstream classification tasks while obtaining the confidence of the results and measuring the credibility of the output results. This, to a certain extent, broadens the applicability of the large model in some serious scenarios, such as medical diagnosis and financial risk control. In addition, the present invention uses an external confidence network, which, on the one hand, consumes few additional resources, and on the other hand, overcomes the inherent "overconfidence" tendency of the large model by using a method similar to an external verification mechanism.

[0086] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0087] Another embodiment of the present invention provides a data processing device, which corresponds one-to-one to the data processing method in the above embodiment. Figure 6 The data processing device includes a data acquisition module 11, a classification probability calculation module 12, and a confidence calculation module 13. The functional modules are described in detail as follows:

[0088] The data acquisition module 11 is used to acquire data to be processed and pre-process the data to be processed to obtain prompt data corresponding to the data to be processed.

[0089] The classification probability calculation module 12 is used to input the prompt data into a pre-trained and complete credible classification model to obtain the classification result and the corresponding first classification probability and second classification probability, wherein the credible classification model includes a classification sub-model and a confidence sub-model plug-in to the classification sub-model, the classification sub-model is used to output the classification result and the corresponding first classification probability, and the confidence sub-model is used to output the second classification probability.

[0090] The confidence calculation module 13 is configured to analyze the confidence of the classification result based on the first classification probability and the second classification probability through a preset confidence analysis strategy.

[0091] In some embodiments, the data acquisition module 11 is specifically used to:

[0092] Get the data to be processed;

[0093] Performing keyword extraction on the data to be processed to extract keywords from the data to be processed;

[0094] The keywords are input into a pre-built prompt model to generate prompt data corresponding to the data to be processed.

[0095] In some embodiments, the classification probability calculation module specifically includes a confidence sub-model building unit, a credible classification model building unit and a probability calculation unit, wherein:

[0096] The confidence sub-model building unit is used to build a confidence sub-model, and train the confidence sub-model based on a pre-built training data set to obtain a fully trained confidence sub-model;

[0097] The credible classification model building unit is used to build a credible classification model based on the trained confidence sub-model and the pre-established classification sub-model;

[0098] The probability calculation unit is used to input the prompt data into the credible classification model to obtain a classification result and a corresponding first classification probability and second classification probability.

[0099] In some embodiments, the confidence sub-model building unit is specifically used to:

[0100] Acquire a text data set, perform prompt engineering processing on the text data set to generate a plurality of prompt data, and use the plurality of prompt data as a training data set;

[0101] Constructing a confidence sub-model, and training the confidence sub-model using the training data set to obtain a fully trained confidence sub-model;

[0102] The confidence sub-model is optimized based on a pre-established test set to obtain an optimized confidence sub-model.

[0103] In some embodiments, the probability calculation unit includes an attention weight calculation subunit, a first probability calculation subunit and a second probability calculation subunit, wherein,

[0104] The attention weight calculation subunit is used to input the prompt data into the classification submodel and calculate the attention weight of the classification submodel;

[0105] The first probability calculation subunit is used to calculate the classification result and the corresponding first classification probability through the output layer of the classification submodel based on the attention weight of the classification submodel;

[0106] The second probability calculation subunit is used to input the attention weight into the confidence sub-model to calculate the second classification probability.

[0107] In some embodiments, the second probability calculation subunit is specifically configured to:

[0108] Based on the confidence sub-model, obtaining a preset compression vector calculation equation;

[0109] Inputting the attention weight into a preset compression vector calculation equation to calculate a compression vector;

[0110] The compressed vector is input into the fully connected layer of the confidence sub-model to obtain a second classification probability.

[0111] In some embodiments, the confidence calculation module 13 includes a model acquisition unit, a probability input unit, and a confidence determination unit, wherein:

[0112] The model acquisition unit is used to acquire a pre-established confidence calculation model;

[0113] The probability input unit is used to input the first classification probability and the second classification probability into the confidence calculation model to obtain a calculation result;

[0114] The confidence determination unit is used to determine the confidence of the classification result based on the calculation result.

[0115] In an embodiment of the present invention, data to be processed is first obtained and preprocessed to obtain prompt data corresponding to the data to be processed; the prompt data is then input into a pre-trained trustworthy classification model to obtain a classification result and corresponding first and second classification probabilities, wherein the trustworthy classification model includes a classification sub-model and a confidence sub-model externally attached to the classification sub-model, the classification sub-model is used to output the classification result and the corresponding first classification probability, and the confidence sub-model is used to output the second classification probability; finally, based on the first and second classification probabilities, the confidence of the classification result is analyzed using a preset confidence analysis strategy. The present invention combines the reasoning and generalization capabilities of a large model, and can simultaneously output results for downstream classification tasks while obtaining the confidence of the results and measuring the credibility of the output results. This, to a certain extent, broadens the applicability of large models in some serious scenarios, such as medical diagnosis and financial risk control. In addition, the present invention uses an external confidence network, which, on the one hand, consumes few additional resources, and on the other hand, overcomes the inherent "overconfidence" tendency of large models by using a method similar to an external verification mechanism.

[0116] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0117] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0118] For the specific definition of the data processing device, please refer to the definition of the data processing method above and will not be repeated here. Each module in the above-mentioned data processing device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software so that the processor can call and execute the operations corresponding to each of the above modules.

[0119] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a data processing method based on artificial intelligence.

[0120] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the functions or steps on the client side of a data processing method.

[0121] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0122] Acquire data to be processed, and pre-process the data to be processed to obtain prompt data corresponding to the data to be processed;

[0123] Inputting the prompt data into a pre-trained credible classification model to obtain a classification result and corresponding first classification probability and second classification probability, wherein the credible classification model includes a classification sub-model and a confidence sub-model externally mounted on the classification sub-model, the classification sub-model is used to output the classification result and the corresponding first classification probability, and the confidence sub-model is used to output the second classification probability;

[0124] Based on the first classification probability and the second classification probability, the confidence of the classification result is analyzed using a preset confidence analysis strategy.

[0125] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0126] Acquire data to be processed, and pre-process the data to be processed to obtain prompt data corresponding to the data to be processed;

[0127] Inputting the prompt data into a pre-trained credible classification model to obtain a classification result and corresponding first classification probability and second classification probability, wherein the credible classification model includes a classification sub-model and a confidence sub-model externally mounted on the classification sub-model, the classification sub-model is used to output the classification result and the corresponding first classification probability, and the confidence sub-model is used to output the second classification probability;

[0128] Based on the first classification probability and the second classification probability, the confidence of the classification result is analyzed using a preset confidence analysis strategy.

[0129] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0130] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0131] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0132] In summary, the data processing method, device, computer equipment and storage medium provided by the present invention first obtain the data to be processed, pre-process the data to be processed, and obtain prompt data corresponding to the data to be processed; then input the prompt data into a pre-trained and complete trusted classification model to obtain the classification result and the corresponding first classification probability and second classification probability, wherein the trusted classification model includes a classification sub-model and a confidence sub-model plug-in to the classification sub-model, the classification sub-model is used to output the classification result and the corresponding first classification probability, and the confidence sub-model is used to output the second classification probability; finally, based on the first classification probability and the second classification probability, the confidence of the classification result is analyzed through a preset confidence analysis strategy. The present invention combines the reasoning and generalization capabilities of large models, and can output results for downstream classification tasks while obtaining the confidence of the results and measuring the credibility of the output results. This, to a certain extent, broadens the applicability of large models in some serious scenarios, such as medical diagnosis and financial risk control. In addition, the present invention uses an external confidence network, which, on the one hand, does not occupy many additional resources, and on the other hand, overcomes the "overconfidence" tendency inherent in large models in a way that is similar to an external verification mechanism.

[0133] It should be noted that if software tools or components other than those of our company appear in the embodiments of this application, they are only used for illustration and do not represent actual use.

[0134] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.< / x> < / x> < / x>

Claims

1. A data processing method, characterized in that: The steps include: Acquire data to be processed, and pre-process the data to be processed to obtain prompt data corresponding to the data to be processed; Inputting the prompt data into a pre-trained credible classification model to obtain a classification result and corresponding first classification probability and second classification probability, wherein the credible classification model includes a classification sub-model and a confidence sub-model externally mounted on the classification sub-model, the classification sub-model is used to output the classification result and the corresponding first classification probability, and the confidence sub-model is used to output the second classification probability; Based on the first classification probability and the second classification probability, the confidence of the classification result is analyzed using a preset confidence analysis strategy.

2. The data processing method according to claim 1, wherein: The acquiring of the data to be processed and preprocessing the data to be processed to obtain prompt data corresponding to the data to be processed includes: Get the data to be processed; Performing keyword extraction on the data to be processed to extract keywords from the data to be processed; The keywords are input into a pre-built prompt model to generate prompt data corresponding to the data to be processed.

3. The data processing method according to claim 1, wherein: The step of inputting the prompt data into a pre-trained trustworthy classification model to obtain a classification result and corresponding first classification probability and second classification probability includes: Constructing a confidence sub-model, and training the confidence sub-model based on a pre-constructed training data set to obtain a fully trained confidence sub-model; Building a credible classification model based on the fully trained confidence sub-model and the pre-established classification sub-model; The prompt data is input into the credible classification model to obtain a classification result and a corresponding first classification probability and second classification probability.

4. The data processing method according to claim 3, wherein: The constructing of the confidence sub-model, training the confidence sub-model based on a pre-constructed training data set to obtain a fully trained confidence sub-model, includes: Acquire a text data set, perform prompt engineering processing on the text data set to generate a plurality of prompt data, and use the plurality of prompt data as a training data set; Constructing a confidence sub-model, and training the confidence sub-model using the training data set to obtain a fully trained confidence sub-model; The confidence sub-model is optimized based on a pre-established test set to obtain an optimized confidence sub-model.

5. The data processing method according to claim 3, wherein: The step of inputting the prompt data into the trusted classification model to obtain a classification result and corresponding first classification probability and second classification probability includes: Inputting the prompt data into the classification sub-model and calculating the attention weight of the classification sub-model; Based on the attention weight of the classification sub-model, calculating the classification result and the corresponding first classification probability through the output layer of the classification sub-model; The attention weight is input into the confidence sub-model to calculate the second classification probability.

6. The data processing method according to claim 5, characterized in that: Inputting the attention weight into the confidence sub-model to calculate the second classification probability includes: Based on the confidence sub-model, obtaining a preset compression vector calculation equation; Inputting the attention weight into a preset compression vector calculation equation to calculate a compression vector; The compressed vector is input into the fully connected layer of the confidence sub-model to obtain a second classification probability.

7. The data processing method according to claim 1, wherein: The calculating the confidence level of the classification result based on the first classification probability and the second classification probability includes: Obtaining a pre-established confidence calculation model; Inputting the first classification probability and the second classification probability into the confidence calculation model to obtain a calculation result; Based on the calculation result, the confidence level of the classification result is determined.

8. A data processing device, characterized in that: include: A data acquisition module is used to acquire data to be processed and pre-process the data to be processed to obtain prompt data corresponding to the data to be processed; a classification probability calculation module, configured to input the prompt data into a pre-trained, fully trusted classification model to obtain a classification result and corresponding first and second classification probabilities, wherein the trusted classification model includes a classification sub-model and a confidence sub-model externally attached to the classification sub-model, the classification sub-model being configured to output the classification result and corresponding first classification probability, and the confidence sub-model being configured to output the second classification probability; The confidence calculation module is used to analyze the confidence of the classification result based on the first classification probability and the second classification probability through a preset confidence analysis strategy.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the data processing method according to any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the data processing method according to any one of claims 1 to 7.