Telecommunication fraud case element analysis method and device based on large language model
Through a language-based big model method, high-quality instruction data sets are constructed and fine-tuned and trained, which solves the problem of low identification efficiency and accuracy in the analysis of telecom fraud case elements, and achieves efficient and accurate case elements analysis.
Patent Information
- Application Number
- CN202510494287.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-20
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art has low identification efficiency and low generalization in the analysis of telecom fraud case elements, and the recognition accuracy needs to be improved.
The language-based big model is used to construct the initial data set by obtaining case description data and elements, and a high-parameter large model is used to pre-label and verify it with prompt engineering, and high-quality instruction data sets are selected, fine-tuned and trained on the big model, and case element analysis model is constructed, and efficient and accurate case element analysis is achieved through evaluation and result output.
It improves the efficiency and accuracy of case elements identification, has high generalization, and can effectively identify case types, traffic and connection methods, transfer methods, fraud-related websites and fraud-related software, reduces algorithm pressure, and improves the accuracy of identification.
Smart Images

Figure CN120373286A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a method and device for analyzing elements of telecommunication fraud cases based on a large language model. Background Art
[0002] With the acceleration of digital transformation, the increase in the penetration rate of smartphones, and the rapid development of network information technology, telecommunications fraud is increasingly affecting people's lives. Telecommunications fraud refers to the use of telephone, Internet or text messages to create false information and set up scams, using remote and contactless methods to commit fraud and induce victims to transfer or remit money. This type of behavior usually involves impersonating someone else or disguising as a staff member of various legal organizations or institutions. Therefore, it is of great significance to analyze the elements of historical fraud cases, summarize the case types, case induction and connection methods, case transfer methods, and extract the fraud-related APPs and fraud-related websites related to the case, so as to effectively perceive the situation of telecommunications fraud cases, which is of great significance to protecting personal information security, effectively combating criminals, preventing economic losses, and enhancing social security awareness.
[0003] At present, most of the existing methods for factor analysis based on historical fraud cases are based on expert experience summary, setting special templates, regular expression matching, keyword matching and machine learning. Specifically, there are the following types: (1) Based on expert experience, the case type, the method of induction and transfer, and the method of transfer are manually summarized from the brief case information, and the fraud-related APPs and fraud-related websites are manually extracted. This method is convenient to operate and has a high accuracy rate, but it has the problems of low efficiency and high labor costs. At the same time, it is highly subjective and may lead to problems such as incorrect identification of case types and incorrect extraction of elements.
[0004] (2) Setting up a dedicated template, that is, when police officers enter information, they follow the dedicated template. Although this method is convenient to operate, it is inefficient and has poor scalability. As the methods of committing crimes change rapidly, the template needs to be frequently expanded and updated. At the same time, it is difficult to unify the input templates of police officers in different provinces and cities.
[0005] (3) Based on regular expression matching, keyword matching and other methods, although this method can identify some elements such as URLs and transfer methods, it is difficult to effectively distinguish between normal URLs and fraudulent URLs, normal APPs and fraudulent APPs. At the same time, the ability to summarize case types and lead to fraudulent connection methods is relatively weak.
[0006] (4) Use machine learning or deep learning to convert text into vectors, and then use models such as TextCNN and Bert to classify the type of text. Although this method has the ability to summarize and classify text, the category labels need to be strictly defined in advance. Once defined, it will be difficult for the model to predict case types and drainage methods outside the preset labels, and the ability to discover new cases is weak. At the same time, the ability to extract fraud-related APPs and websites is poor.
[0007] In summary, in the process of identifying and analyzing case elements in the prior art, the identification efficiency is low, the generalization ability is low, and the identification accuracy also needs to be further improved. Summary of the Invention
[0008] Based on this, it is necessary to provide a method and device for analyzing case elements of telecommunications fraud based on a large language model with high identification efficiency, good generalization ability and high identification accuracy in the process of analyzing case elements to solve the above technical problems.
[0009] The present invention provides a method for analyzing case elements of telecommunications fraud based on a large language model, and the method includes: Obtain case description data and the corresponding case elements of the case description data, and construct an initial data set based on the case description data and the case elements; Pre-annotate the initial data set through prompt engineering to obtain a pre-annotated data set, and divide the pre-annotated data set into multiple subsets to correct the pre-annotated results in each subset to obtain a verification data set; Screen out an instruction data set from the verification data set, use the case description data in the instruction data set as the input of the large model, and use the case elements as the output of the large model to train the large model to obtain a case element analysis model; Select multiple case description data and case elements that have been marked and verified to construct an evaluation data set to evaluate the case element analysis model; When the case element analysis model passes the evaluation, input the current case description data into the case element analysis model to output the case element analysis result corresponding to the current case description data; Among them, the case elements include case types, drainage connection methods, transfer methods, fraud-related websites, and fraud-related software. The prompt engineering consists of prompt words, input examples, output examples, and real inputs. The pre-annotated data set includes prompt words, real inputs, and pre-annotated results. The instruction data set is a subset of the verification data set.
[0010] In one embodiment, pre-annotating the initial dataset through prompt engineering to obtain a pre-annotated dataset, and dividing the pre-annotated dataset into multiple subsets to correct the pre-annotation results in each subset to obtain a verification dataset, including: Based on a large model with a large number of parameters combined with the prompt engineering, pre-annotate the initial dataset, and at the same time set the output format of the large model with a large number of parameters to a JSON structure to obtain the prompt words and pre-annotation results corresponding to each real input, so as to construct the pre-annotated dataset; Extract a first subset from the pre-annotated dataset, and check and verify each real input in the first subset and the prompt words and pre-annotation results corresponding to each real input to correct the pre-annotation results and obtain the verification dataset corresponding to the first subset.
[0011] In one embodiment, screening out an instruction dataset from the verification dataset, using the case description data in the instruction dataset as the input of the large model, and using the case elements as the output of the large model to train the large model to obtain a case element analysis model, including: Infer each group of data in the prime number verification dataset through different numbers of parameters llama3 in the large model with a large number of parameters combined with the prompt engineering to obtain a first inference result and a second inference result, and calculate the inference loss of the first inference result and the second inference result based on the check and verification result; Based on the inference losses of the first inference result and the second inference result, screen out the first inference result with an inference loss lower than the first threshold and the second inference result with an inference loss exceeding the second threshold, and sort the first inference result with an inference loss lower than the first threshold and the second inference result with an inference loss exceeding the second threshold in ascending order and descending order respectively to select the intersection data with a ranking not lower than the third threshold; Among them, the instruction dataset is composed of the intersection data.
[0012] In one embodiment, the different numbers of parameters llama3 include a first number of parameters llama3 and a second number of parameters llama3, and the first number is greater than the second number; The method of screening out an instruction dataset from the verification dataset, using the case description data in the instruction dataset as the input of the large model, and using the case elements as the output of the large model to train the large model to obtain a case element analysis model further includes: Obtain the labeled JSON standardized instruction data, and use various incorrect JSON formats as the input of the large model to obtain the corrected JSON format based on the labeled JSON standardized instruction data to obtain an instruction data subset; Merge the subset of the instruction data into the instruction data set, and perform full-parameter fine-tuning training on the second quantity of parameters llama3 based on the instruction data set corresponding to the second quantity of parameters llama3, so as to set the hyperparameters of the fine-tuning training to perform fine-tuning training on the large model.
[0013] In one embodiment, the selecting a plurality of case description data and case elements that have been annotated and verified to construct an evaluation data set to evaluate the case element analysis model includes: Evaluate the case element analysis model from the dimensions of case type, drainage connection method, transfer method, fraud-related website, and fraud-related software based on the evaluation data set, and obtain a plurality of evaluation results; When the ratio between the sum of the plurality of evaluation results and the number of the plurality of evaluation results is not lower than the fourth threshold, it is determined that the evaluation of the case element analysis model passes; When the ratio between the sum of the plurality of evaluation results and the number of the plurality of evaluation results is lower than the fourth threshold, it is determined that the evaluation of the case element analysis model fails, and the instruction data set is re-screened and the large model is fine-tuned; Wherein, the plurality of evaluation results are the probabilities that the prediction results of the case type, drainage connection method, transfer method, fraud-related website, and fraud-related software coincide with their respective true results.
[0014] In one embodiment, when the evaluation of the case element analysis model passes, input the current case description data into the case element analysis model to output the case element analysis result corresponding to the current case description data, including: Obtain the current case description data and the corresponding prompt words, and use the current case description data and the corresponding prompt words as the input of the case element analysis model that has passed the evaluation; Call the case element analysis model by setting hyperparameters in different intervals to output multiple groups of case element analysis results, and vote on the multiple groups of case element analysis results to select the best case element analysis result.
[0015] In one embodiment, the method further includes: When the case element analysis result is not a JSON structure body, replace the prompt words corresponding to the current case description data with JSON instruction prompt words; Based on the JSON instruction prompt words, combine the case element analysis results of the non-JSON structure body as the input of the large model to output the case element analysis results of the standard JSON structure body.
[0016] The present invention also provides a device for analyzing elements of telecommunications fraud cases based on a large language model, the device comprising: A dataset construction module, configured to obtain case description data and the corresponding case elements of the case description data, and construct an initial dataset based on the case description data and the case elements; A data verification module, configured to perform pre-annotation on the initial dataset through prompt engineering to obtain a pre-annotated dataset, and divide the pre-annotated dataset into multiple subsets to correct the pre-annotation results in each subset to obtain a verified dataset; A large model training module, configured to screen out an instruction dataset from the verified dataset, use the case description data in the instruction dataset as the input of the large model, and use the case elements as the output of the large model to train the large model to obtain a case element analysis model; A model evaluation module, configured to select multiple case description data and case elements that have been annotated and verified to construct an evaluation dataset to evaluate the case element analysis model; A case element analysis module, configured to, when the case element analysis model passes the evaluation, input the current case description data into the case element analysis model to output the case element analysis result corresponding to the current case description data; Wherein, the case elements include case types, drainage connection methods, transfer methods, fraud-related websites, and fraud-related software, the prompt engineering is composed of prompt words, input examples, output examples, and real inputs, the pre-annotated dataset includes prompt words, real inputs, and pre-annotation results, and the instruction dataset is a subset of the verified dataset.
[0017] The present invention also provides an electronic device, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the method for analyzing elements of telecommunications fraud cases based on a large language model as described in any one of the above when executing the computer program.
[0018] The present invention also provides a computer storage medium storing a computer program, and the computer program implements the method for analyzing elements of telecommunications fraud cases based on a large language model as described in any one of the above when executed by a processor.
[0019] The above method and device for analyzing elements of telecommunications fraud cases based on a large language model have the following beneficial effects: (1) The present invention uses a large model with a large number of parameters to pre-label data, and manually checks and verifies the pre-labeled data to screen out high-quality labeled data, and produces a high-quality instruction data set for case analysis. The instruction data set used by the large model can, on the one hand, cover many subtasks such as case type classification, drainage communication method classification, transfer method classification, fraud APP extraction, and fraud website extraction, and there is no need to train multiple sets of models for combination; on the other hand, the training data used by the large model is less than the training data used by deep learning models such as TextCNN and Bert, reducing the algorithm pressure and improving the efficiency of case element recognition to a certain extent.
[0020] (2) Based on the large model combined with the high-quality instruction data set, the present invention fine-tunes and trains the large model. Due to sufficient prior knowledge and rich case analysis knowledge, the case element analysis model obtained by training has higher case element recognition accuracy and generalization compared with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 It is one of the flow diagrams of the method for analyzing telecom fraud case elements based on a language large model provided by the present invention; Figure 2 It is the schematic diagram of the telecom fraud case element analysis framework of the method for analyzing telecom fraud case elements based on a language large model in the specific embodiment provided by the present invention; Figure 3 It is the second flow diagram of the method for analyzing telecom fraud case elements based on a language large model provided by the present invention; Figure 4 It is the third flow diagram of the method for analyzing telecom fraud case elements based on a language large model provided by the present invention; Figure 5 It is the fourth flow diagram of the method for analyzing telecom fraud case elements based on a language large model provided by the present invention; Figure 6 It is the fifth flow diagram of the method for analyzing telecom fraud case elements based on a language large model provided by the present invention; Figure 7 It is the sixth flow diagram of the method for analyzing telecom fraud case elements based on a language large model provided by the present invention; Figure 8Schematic diagram VII of the method for analyzing elements of telecommunications fraud cases based on a large language model provided by the present invention; Figure 9 Schematic diagram of the structure of the device for analyzing elements of telecommunications fraud cases based on a large language model provided by the present invention; Figure 10 Internal structure diagram of the electronic device provided by the present invention. Detailed implementation manners
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] The following combines Figures 1 to 10 to describe the method and device for analyzing elements of telecommunications fraud cases based on a large language model of the present invention.
[0025] As Figure 1 shown, in one embodiment, a method for analyzing elements of telecommunications fraud cases based on a large language model includes the following steps: Step S110, obtaining case description data and case elements corresponding to the case description data, and constructing an initial data set based on the case description data and the case elements.
[0026] Among them, the case elements include case type, drainage connection method, transfer method, fraud-related website, and fraud-related software.
[0027] Specifically, the server obtains the case description data and the case elements of the case type, drainage connection method, transfer method, fraud-related website, and fraud-related software corresponding to the case description data to construct an initial data set.
[0028] Combined with Figure 2 shown, in a specific embodiment, the method for analyzing elements of telecommunications fraud cases based on a large language model provided by the present invention includes five parts: business modeling, construction of a high-quality instruction data set, model fine-tuning training, evaluation of the large model for analyzing case elements, and inference and result post-processing of the large model for analyzing case elements.
[0029] First, in the business modeling part, that is, clarifying the input (case description) and output (case type, drainage connection method, transfer method, fraud-related website, fraud-related APP) of the model, and clarifying the main ranges of each type and method, specifically including: Define the input and output of the model: The input of the model is the case description, and the output of the model is 5 types of case elements extracted from the case description, namely case type, drainage communication method, transfer method, fraud-related website, and fraud-related APP; Define the case type: The case type refers to a type abstracted by summarizing and generalizing the case description, which is a single classification, that is, a case can only belong to one type; Define the drainage communication method: The drainage communication method refers to the way the scammer contacts the victim, which is a multi-classification, that is, there may be multiple drainage communication methods in a case; Define the transfer method: The transfer method refers to the way the victim transfers wealth to the scammer, which is a multi-classification, that is, there may be multiple transfer methods in a case; Define the fraud-related website: The fraud-related website is usually the website that the scammer induces the victim to click on, which is a multi-element extraction, that is, there may be multiple fraud-related websites in a case, and the name and link of the fraud-related website need to be extracted; Define the fraud-related APP. The fraud-related APP is usually the APP that the scammer induces the victim to download, and the fraud is carried out within the fraud-related APP, which is a multi-element extraction, that is, there may be multiple fraud-related APPs in a case, and the name and download link of the fraud-related APP need to be extracted.
[0030] Step S120, pre-annotate the initial dataset through prompt engineering to obtain a pre-annotated dataset, and divide the pre-annotated dataset into multiple subsets to correct the pre-annotated results in each subset to obtain a verification dataset.
[0031] Among them, the prompt engineering is composed of prompt words, input examples, output examples, and real inputs, and the pre-annotated dataset includes prompt words, real inputs, and pre-annotated results.
[0032] Specifically, the server pre-annotates the initial dataset constructed in step S110 through the large model combined with prompt engineering to obtain a pre-annotated dataset including prompt words, real inputs, and pre-annotated results, and divides the pre-annotated dataset into multiple subsets to verify and correct the pre-annotated results in each subset to obtain a verification dataset.
[0033] Combine Figure 2As shown, in a specific embodiment, for the method for analyzing elements of telecommunications fraud cases based on a large language model provided by the present invention, in the construction part of the high-quality instruction dataset, first, 20,000 pieces of historical case description data are selected as the initial dataset. Based on a large model with a large number of parameters (llama3 with 70 billion parameters), combined with prompt engineering, the initial dataset is pre-annotated to obtain a pre-annotated dataset V1 (including three parts: prompt words, real inputs, and pre-annotated results). The pre-annotation accuracy of the large model with a large number of parameters is relatively high, which can effectively reduce the manual annotation cost. At the same time, the analysis of case elements itself integrates a variety of common natural language processing tasks, including single-classification tasks, multi-classification tasks, and multi-element information extraction tasks. To improve the accuracy of the model, the output format is limited to a standard JSON structure. The JSON structure can reduce errors in answer selection by restricting possible answers, thereby improving the performance and accuracy of classification tasks. Prompt engineering consists of prompt words, input examples, output examples, and real inputs.
[0034] In this embodiment, according to the criterion of uniform distribution of case types, 5,000 pieces of data are selected from 20,000 pieces of pre-annotated datasets (including three parts: prompt words, real inputs, and pre-annotated results), with the number of each case type maintained at about 100, as the pre-annotated dataset. A dataset with a balanced distribution is conducive to improving the generalization of the model and reducing the risk of model overfitting. Then, the 5,000 pieces of pre-annotated datasets are manually checked and verified, and the incorrect content in the pre-annotation is manually modified to improve the accuracy of the annotation results, obtaining 5,000 pieces of manually verified datasets. A model trained based on a small number of high-quality instruction datasets often has higher accuracy and better generalization than a model trained based on a large number of medium- and low-quality instruction datasets. The specific screening process is as follows: For each piece of data in the manually verified dataset, denoted as Input i , based on llama3 with 70 billion parameters and the aforementioned prompt engineering, perform inference on Input i , and obtain the inference result, denoted as ; based on llama3 with 13 billion parameters and prompt engineering, perform inference on Input i , and obtain the inference result, denoted as ; the result obtained by manual annotation verification is denoted as .
[0035] Then, the inference loss expression for llama3 with 70 billion parameters is calculated as: ; Where \(t\) represents the \(t\)-th sample in the input-output sequence, \(N\) represents the total number of samples, \(c\) represents the category in the one-hot encoding, and \(k\) represents the total number of one-hot encoding categories. represents the one-hot encoded vector of the actual label after manual annotation verification. represents the probability vector predicted by the 70 billion parameter llama3 model. Calculate the inference loss of llama3 with 13 billion parameters Similarly.
[0036] In this embodiment, the purpose of data screening is to find data that is relatively simple for large models with a large number of parameters (70 billion parameter llama3) and relatively difficult for large models with a medium number of parameters (13 billion parameter llama3), so as to ensure that the screened data is learnable and has less noise, thereby ensuring the accuracy of the model after training. Therefore, data with relatively small and relatively large should be selected as much as possible. That is, the target data is: ; In the formula, for the data obtained from the 5000 manually verified datasets, is sorted in ascending order, and is sorted in descending order. For the first 2000 intersection data, Input i and are respectively taken to obtain a high-quality instruction dataset.
[0037] In this embodiment, to solve the problem that during the inference of large models, random interference causes the output of a small number of samples to not be parsed into a standard JSON structure, 100 JSON standardization instruction data are manually annotated, with various incorrect JSON formats as the input and the corrected JSON format as the output. It is denoted as the high-quality instruction dataset.
[0038] Step S130: Screen out the instruction dataset from the verified dataset, use the case description data in the instruction dataset as the input of the large model, and use the case elements as the output of the large model to train the large model to obtain a case element analysis model.
[0039] Among them, the instruction dataset is a subset of the verified dataset.
[0040] Specifically, the server screens out a subset of part of the data from the verified dataset obtained in step S120, that is, the instruction dataset, uses the case description data in the instruction dataset as the input of the large model, and uses the case elements as the output of the large model to train the large model to obtain a case element analysis model with the function of case element analysis.
[0041] Combined withFigure 2 As shown, in a specific embodiment, for the method for analyzing elements of telecommunications fraud cases based on a large language model provided by the present invention, in the model fine-tuning training part, based on llama3 with 13 billion parameters and the high-quality instruction dataset obtained in (2), full-parameter fine-tuning training is performed on llama3 with 13 billion parameters. Fine-tuning training the large model with medium parameter quantity, llama3 with 13 billion parameters, has greater advantages than fine-tuning training the large model with high parameter quantity, llama3 with 70 billion parameters, in terms of the demand for computing power resources, the usage efficiency of the model, etc.
[0042] Based on the characteristics of the high-quality instruction dataset, set the hyperparameters for fine-tuning training as follows: --train-iters: 2000 --lr: 1e-5 --seq-length: 2048 --tensor-model-parallel-size: 8 --pipeline-model-parallel-size: 1 --micro-batch-size: 2 --global-batch-size: 128 --make-vocab-size-divisible-by: 1 --lr-decay-style: cosine --attetion-dropout: 0 --init-method-std: 0.01 --hidden-dropout: 0 --position-embedding-type: rope --normalnization: RMSNorm --min-lr: 1e-9 --weight-decay: 0.1 --lr-warmup-fraction: 0.01 --clip-grad: 1 --adam-beta1: 0.9 --adam-beta2: 0.95 --initial-loss-scale: 4096 --split: 90,10,0 --log-interval: 1 --save-interval: 1000 --eval-interval: 200 --eval-iters: 20 The model obtained after the fine-tuning training is denoted as the case element analysis large model.
[0043] Step S140: Select multiple case description data and case elements that have been annotated and verified to construct an evaluation data set to evaluate the case element analysis model.
[0044] Specifically, the server selects multiple case description data and case elements that have been annotated and verified to construct an evaluation data set to evaluate the case element analysis model trained in step S130, so as to ensure the output accuracy of the case element analysis model.
[0045] Combined with Figure 2 As shown in , in a specific embodiment, for the method for analyzing case elements of telecommunications fraud cases based on a large language model provided by the present invention, in the evaluation part of the case element analysis large model, m pieces of data that have been manually annotated and verified are selected as the evaluation data set, and the case element analysis large model is evaluated from 5 dimensions: case type, lead communication method, transfer method, fraud-related website, and fraud-related APP: 1) Case type evaluation: The case type extraction is a single-classification task. If the model prediction is exactly the same as the true answer, it is determined to be correct; otherwise, it is determined to be wrong.
[0046] P1 = sum (the model prediction is exactly the same as the true answer) / m.
[0047] 2) Lead communication method evaluation: The lead communication method extraction is a multi-classification task, and the accuracy rate is defined as: P2 = sum (the number of categories where the model prediction category is the same as the true answer category in each piece of data / max (the number of true answer categories in each piece of data, the number of model prediction categories in each piece of data)) / m.
[0048] 3) Transfer method evaluation: The transfer method extraction is a multi-classification task, and the accuracy rate is defined as: P3 = sum (the number of categories where the model prediction category is the same as the true answer category in each piece of data / max (the number of true answer categories in each piece of data, the number of model prediction categories in each piece of data)) / m.
[0049] 4) Fraud-related website evaluation: The fraud-related website extraction is a multi-element information extraction task, and the accuracy rate is defined as: P4 = sum(number of cases where the predicted names and links of fraud-related websites in each piece of data are exactly the same as the true answers / max(number of true elements in each piece of data, number of predicted elements in each piece of data)).
[0050] 5) Evaluation of fraud-related APPs: The extraction of fraud-related APPs is a multi-element information extraction task, and the accuracy rate is defined as: P5 = sum(number of cases where the predicted names and download links of fraud-related APPs in each piece of data are exactly the same as the true answers / max(number of true elements in each piece of data, number of predicted elements in each piece of data)).
[0051] Comprehensive accuracy rate: P = (P1 + P2 + P3 + P4 + P5) / 5. When the comprehensive accuracy rate meets the index of 90%, the evaluation passes; if it does not meet the index, repeat the above steps to continue optimizing the high-quality instruction dataset, re-fine-tune the training until the index is met.
[0052] Step S150, when the case element analysis model passes the evaluation, input the current case description data into the case element analysis model to output the case element analysis result corresponding to the current case description data.
[0053] Specifically, after the case element analysis model passes the evaluation, input the current case description data to be analyzed into the case element analysis model, and finally output the case element analysis result corresponding to the current case description.
[0054] Combined with Figure 2 As shown, in a specific embodiment, for the method for analyzing case elements of telecommunications fraud cases based on a large language model provided by the present invention, in the inference and result post-processing part of the case element analysis large model, when the training evaluation of the case element analysis large model is completed, the inference of the model starts.
[0055] When the model is inferring, the input includes two parts: a prompt word and a true input. The prompt word is kept the same as the prompt word in the training data, and the true input is replaced with the case description to be predicted. The multi-expert mode is used to reduce the hallucination problem of the large model, that is, by setting hyperparameters in different intervals, the model generates multiple sets of outputs, and the best result is selected through the voting method to improve the credibility of the model, which significantly improves the effect for classification tasks. Specifically as follows: Transition from output conservatism to output diversity, set 5 groups of hyperparameters in different intervals: A) temperatura: Null, top_p: Null, do_sample: False B) temperatura: 0.01, top_p: 0.1, do_sample: False C) Temperature: 0.1, top_p: 0.5, do_sample: True D) Temperature: 0.3, top_p: 0.8, do_sample: True E) Temperature: 0.5, top_p: 1, do_sample: True Through the voting method, take the majority part of the inference results of the 5 groups of parameters. When the inference results of the 5 groups of parameters are different, take the inference result of parameter A as the final result.
[0056] In this embodiment, post - processing of the incorrect JSON structure. When the result of model inference cannot be parsed into a standard JSON structure, replace the prompt with a JSON instruction prompt, replace the real input with the error output result that cannot be parsed into JSON, and let the large - model perform inference again to obtain a parsable standard JSON structure.
[0057] In this embodiment, when the model faces a new case input and the output case type, drainage communication method, and transfer method exceed the main defined range above, it proves that the large - model has emergent capabilities. Supplement this part of the data into the main range for case study and analysis.
[0058] The above - mentioned method for analyzing elements of telecom fraud cases based on a large - language model uses a large - model with a large number of parameters to pre - annotate data, manually check and verify the pre - annotated data, and filter out high - quality annotated data to produce a high - quality instruction data set for case analysis. The instruction data set used by the large - model can, on the one hand, cover many sub - tasks such as case type classification, drainage communication method classification, transfer method classification, extraction of fraud - related APPs, extraction of fraud - related websites, etc., and there is no need to train multiple sets of models for combination; on the other hand, the training data used by the large - model is less than that used by deep - learning models such as TextCNN and Bert, reducing the algorithm pressure and improving the efficiency of case element recognition to a certain extent. In addition, this method is based on the large - model combined with a high - quality instruction data set to fine - tune the large - model. Due to having sufficient prior knowledge and rich case - analysis knowledge, the trained case - element analysis model has a higher case - element recognition accuracy and generalization ability compared with the existing technology.
[0059] As Figure 3 shown, in one embodiment, the method for analyzing elements of telecom fraud cases based on a large - language model provided by the present invention, step S120 specifically includes the following steps: Step S121: Based on a large model with a large number of parameters combined with prompt engineering, pre-annotate the initial dataset. At the same time, set the output format of the large model with a large number of parameters to a JSON structure to obtain the prompt words and pre-annotation results corresponding to each real input, so as to construct a pre-annotated dataset.
[0060] Step S122: Extract a first subset from the pre-annotated dataset, and check and verify each real input in the first subset and the prompt words and pre-annotation results corresponding to each real input to correct the pre-annotation results, so as to obtain a verification dataset corresponding to the first subset.
[0061] As Figure 4 shown, in one embodiment, the method for analyzing elements of telecommunications fraud cases based on a large language model provided by the present invention, step S130 specifically includes the following steps: Step S131: Use different numbers of parameters llama3 in the large model with a large number of parameters combined with prompt engineering to infer each group of data in the prime number verification dataset, obtain a first inference result and a second inference result, and calculate the inference loss of the first inference result and the second inference result based on the check and verification results.
[0062] Step S132: Based on the inference losses of the first inference result and the second inference result, screen out the first inference result with an inference loss lower than the first threshold and the second inference result with an inference loss exceeding the second threshold, and respectively sort the first inference result with an inference loss lower than the first threshold in ascending order and the second inference result with an inference loss exceeding the second threshold in descending order to select the intersection data with a ranking not lower than the third threshold.
[0063] Among them, the instruction dataset is composed of intersection data.
[0064] As Figure 5 shown, in one embodiment, the method for analyzing elements of telecommunications fraud cases based on a large language model provided by the present invention, different numbers of parameters llama3 include a first number of parameters llama3 and a second number of parameters llama3, and the first number is greater than the second number.
[0065] Step S130 specifically further includes the following steps: Step S133: Obtain the annotated JSON standardized instruction data, and use various incorrect JSON formats as the input of the large model to obtain the corrected JSON format based on the annotated JSON standardized instruction data, so as to obtain an instruction data subset.
[0066] Step S134: Merge the instruction data subset into the instruction data set, and perform full-parameter fine-tuning training on the second quantity of parameters llama3 based on the instruction data set corresponding to the second quantity of parameters llama3, so as to set the hyperparameters of the fine-tuning training to perform fine-tuning training on the large model.
[0067] As Figure 6 shown, in one embodiment, the method for analyzing elements of telecommunications fraud cases based on a large language model provided by the present invention, step S140 specifically includes the following steps: Step S141: Evaluate the case element analysis model from the dimensions of case type, drainage connection method, transfer method, fraud-related website, and fraud-related software based on the evaluation data set to obtain multiple evaluation results.
[0068] Step S142: When the ratio between the sum of multiple evaluation results and the number of multiple evaluation results is not lower than the fourth threshold, it is determined that the evaluation of the case element analysis model passes.
[0069] Step S143: When the ratio between the sum of multiple evaluation results and the number of multiple evaluation results is lower than the fourth threshold, it is determined that the evaluation of the case element analysis model fails, and the instruction data set is re-screened and the large model is fine-tuned.
[0070] Among them, the multiple evaluation results are the probabilities that the prediction results of the case type, drainage connection method, transfer method, fraud-related website, and fraud-related software coincide with their respective true results.
[0071] As Figure 7 shown, in one embodiment, the method for analyzing elements of telecommunications fraud cases based on a large language model provided by the present invention, step S150 specifically includes the following steps: Step S151: Obtain the current case description data and the corresponding prompt words, and use the current case description data and the corresponding prompt words as the input of the case element analysis model that has passed the evaluation.
[0072] Step S152: Call the case element analysis model to output multiple groups of case element analysis results by setting hyperparameters in different intervals, and vote on the multiple groups of case element analysis results to select the best case element analysis result.
[0073] As Figure 8 shown, in one embodiment, the method for analyzing elements of telecommunications fraud cases based on a large language model provided by the present invention further includes the following steps: Step S810: When the case element analysis result is not a JSON structure, replace the prompt words corresponding to the current case description data with JSON instruction prompt words.
[0074] Step S820: Based on the JSON instruction prompt words, combined with the case element analysis results of non-JSON structures as the input of the large model, to output the case element analysis results in a standard JSON structure.
[0075] The following describes the device for analyzing case elements of telecommunications fraud based on a large language model provided by the present invention. The device for analyzing case elements of telecommunications fraud based on a large language model described below can be correspondingly referred to the method for analyzing case elements of telecommunications fraud based on a large language model described above.
[0076] As Figure 9 shown, in one embodiment, a device for analyzing case elements of telecommunications fraud based on a large language model includes a dataset construction module 910, a data verification module 920, a large model training module 930, a model evaluation module 940, and a case element analysis module 950.
[0077] The dataset construction module 910 is used to obtain case description data and the corresponding case elements, and construct an initial dataset based on the case description data and the case elements.
[0078] The data verification module 920 is used to pre-annotate the initial dataset through prompt engineering to obtain a pre-annotated dataset, and divide the pre-annotated dataset into multiple subsets to correct the pre-annotated results in each subset to obtain a verified dataset.
[0079] The large model training module 930 is used to screen out an instruction dataset from the verified dataset, use the case description data in the instruction dataset as the input of the large model, and use the case elements as the output of the large model to train the large model to obtain a case element analysis model.
[0080] The model evaluation module 940 is used to select multiple annotated and verified case description data and case elements to construct an evaluation dataset to evaluate the case element analysis model.
[0081] The case element analysis module 950 is used to input the current case description data into the case element analysis model when the case element analysis model passes the evaluation to output the case element analysis results corresponding to the current case description data.
[0082] Among them, the case elements include case types, drainage connection methods, transfer methods, fraud-related websites, and fraud-related software. The prompt engineering consists of prompt words, input examples, output examples, and real inputs. The pre-annotated dataset includes prompt words, real inputs, and pre-annotated results. The instruction dataset is a subset of the verified dataset.
[0083] In this embodiment, for the telecommunications fraud case element analysis device provided by the present invention based on a large language model, the data verification module 920 is specifically used for: Based on a large model with a large number of parameters combined with prompt engineering, pre-annotate the initial data set, and at the same time set the output format of the large model with a large number of parameters to a JSON structure to obtain the prompt words and pre-annotation results corresponding to each real input, so as to construct a pre-annotated data set.
[0084] Extract a first subset from the pre-annotated data set, and check and verify each real input in the first subset and the prompt words and pre-annotation results corresponding to each real input, so as to correct the pre-annotation results and obtain a verification data set corresponding to the first subset.
[0085] In this embodiment, for the telecommunications fraud case element analysis device provided by the present invention based on a large language model, the large model training module 930 is specifically used for: Infer each group of data in the prime number verification data set through different numbers of parameters llama3 in the large model with a large number of parameters combined with prompt engineering to obtain a first inference result and a second inference result, and calculate the inference loss of the first inference result and the second inference result based on the check and verification result.
[0086] Screen out the first inference result with an inference loss lower than the first threshold and the second inference result with an inference loss exceeding the second threshold based on the inference loss of the first inference result and the second inference result, and sort the first inference result with an inference loss lower than the first threshold and the second inference result with an inference loss exceeding the second threshold in ascending order and descending order respectively, so as to select the intersection data with a ranking not lower than the third threshold.
[0087] Among them, the instruction data set is composed of intersection data.
[0088] In this embodiment, for the telecommunications fraud case element analysis device provided by the present invention based on a large language model, different numbers of parameters llama3 include a first number of parameters llama3 and a second number of parameters llama3, and the first number is greater than the second number.
[0089] The large model training module 930 is specifically further used for: Obtain the annotated JSON standardized instruction data, and use various incorrect JSON formats as the input of the large model to obtain the corrected JSON format based on the annotated JSON standardized instruction data to obtain an instruction data subset.
[0090] Merge the instruction data subset into the instruction data set, and perform full-parameter fine-tuning training on the second number of parameters llama3 based on the instruction data set corresponding to the second number of parameters llama3, so as to set the hyperparameters of the fine-tuning training to fine-tune the large model.
[0091] In this embodiment, for the telecommunications fraud case element analysis device based on a large language model provided by the present invention, the model evaluation module 940 is specifically configured to: Evaluate the case element analysis model from the dimensions of case type, drainage connection method, transfer method, fraud-related website, and fraud-related software based on the evaluation data set, and obtain multiple evaluation results.
[0092] When the ratio between the sum of multiple evaluation results and the number of multiple evaluation results is not lower than the fourth threshold, it is determined that the evaluation of the case element analysis model passes.
[0093] When the ratio between the sum of multiple evaluation results and the number of multiple evaluation results is lower than the fourth threshold, it is determined that the evaluation of the case element analysis model fails, and the instruction data set is re-screened and the large model is fine-tuned.
[0094] Among them, the multiple evaluation results are the probabilities that the prediction results of the case type, drainage connection method, transfer method, fraud-related website, and fraud-related software coincide with their respective true results.
[0095] In this embodiment, for the telecommunications fraud case element analysis device based on a large language model provided by the present invention, the case element analysis module 950 is specifically configured to: Obtain the current case description data and the corresponding prompt words, and use the current case description data and the corresponding prompt words as the input of the case element analysis model that has passed the evaluation.
[0096] Call the case element analysis model to output multiple groups of case element analysis results by setting hyperparameters in different intervals, and vote on the multiple groups of case element analysis results to select the best case element analysis result.
[0097] In this embodiment, the telecommunications fraud case element analysis device based on a large language model provided by the present invention further includes an analysis result format correction module, which is used for: When the case element analysis result is not a JSON structure, replace the prompt words corresponding to the current case description data with JSON instruction prompt words.
[0098] Based on the JSON instruction prompt words, combine the case element analysis results of the non-JSON structure as the input of the large model to output the case element analysis results in the standard JSON structure.
[0099] Figure 10 Illustrates a schematic diagram of the physical structure of an electronic device. The electronic device can be a smart terminal, and its internal structure diagram can be as Figure 10As shown in the figure. The electronic device includes a processor, an internal memory, and a network interface connected by a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a method for analyzing elements of telecommunications fraud cases based on a language large model. The method includes: Obtain case description data and the corresponding case elements, and construct an initial data set based on the case description data and the case elements; Pre-annotate the initial data set through prompt engineering to obtain a pre-annotated data set, and divide the pre-annotated data set into multiple subsets to correct the pre-annotated results in each subset to obtain a verification data set; Screen out an instruction data set from the verification data set, use the case description data in the instruction data set as the input of the large model, and use the case elements as the output of the large model to train the large model to obtain a case element analysis model; Select multiple verified case description data and case elements to construct an evaluation data set to evaluate the case element analysis model; When the case element analysis model passes the evaluation, input the current case description data into the case element analysis model to output the case element analysis result corresponding to the current case description data; Among them, the case elements include case types, drainage connection methods, transfer methods, fraud-related websites, and fraud-related software. Prompt engineering consists of prompt words, input examples, output examples, and real inputs. The pre-annotated data set includes prompt words, real inputs, and pre-annotated results. The instruction data set is a subset of the verification data set.
[0100] Those skilled in the art can understand that Figure 10 The structure shown in the figure is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the electronic device to which the solution of the present invention is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0101] On the other hand, the present invention also provides a computer storage medium storing a computer program. When the computer program is executed by a processor, it realizes a method for analyzing elements of telecommunications fraud cases based on a language large model. The method includes: Obtain case description data and the corresponding case elements, and construct an initial data set based on the case description data and the case elements; Pre-annotate the initial dataset through prompt engineering to obtain a pre-annotated dataset, and divide the pre-annotated dataset into multiple subsets to correct the pre-annotated results in each subset to obtain a verification dataset; Screen out an instruction dataset from the verification dataset, use the case description data in the instruction dataset as the input of the large model, and use the case elements as the output of the large model to train the large model to obtain a case element analysis model; Select multiple annotated and verified case description data and case elements to construct an evaluation dataset to evaluate the case element analysis model; When the case element analysis model passes the evaluation, input the current case description data into the case element analysis model to output the case element analysis result corresponding to the current case description data; Among them, the case elements include case types, drainage connection methods, transfer methods, fraud-related websites, and fraud-related software. Prompt engineering consists of prompt words, input examples, output examples, and real inputs. The pre-annotated dataset includes prompt words, real inputs, and pre-annotated results. The instruction dataset is a subset of the verification dataset.
[0102] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium. When the processor executes the computer instructions, it implements a method for analyzing case elements of telecommunications fraud based on a language large model. The method includes: Obtain case description data and the corresponding case elements, and construct an initial dataset based on the case description data and the case elements; Pre-annotate the initial dataset through prompt engineering to obtain a pre-annotated dataset, and divide the pre-annotated dataset into multiple subsets to correct the pre-annotated results in each subset to obtain a verification dataset; Screen out an instruction dataset from the verification dataset, use the case description data in the instruction dataset as the input of the large model, and use the case elements as the output of the large model to train the large model to obtain a case element analysis model; Select multiple annotated and verified case description data and case elements to construct an evaluation dataset to evaluate the case element analysis model; When the case element analysis model passes the evaluation, input the current case description data into the case element analysis model to output the case element analysis result corresponding to the current case description data; Among them, the case elements include case type, drainage connection method, transfer method, fraud-related website, and fraud-related software. The prompting project consists of prompt words, input examples, output examples, and real inputs. The pre-annotated dataset includes prompt words, real inputs, and pre-annotation results. The instruction dataset is a subset of the verification dataset.
[0103] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory.
[0104] By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0105] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0106] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent of the present invention should be subject to the appended claims.
Claims
1. A method for analyzing elements of telecommunications fraud cases based on large language models, characterized in that, The method includes: Obtaining case description data and the case elements corresponding to the case description data, and constructing an initial data set based on the case description data and the case elements; Pre-annotating the initial data set through prompt engineering to obtain a pre-annotated data set, and dividing the pre-annotated data set into multiple subsets to correct the pre-annotated results in each subset to obtain a verification data set; Screening out an instruction data set from the verification data set, using the case description data in the instruction data set as the input of the large model, and using the case elements as the output of the large model to train the large model to obtain a case element analysis model; Selecting multiple annotated and verified case description data and case elements to construct an evaluation data set to evaluate the case element analysis model; When the case element analysis model passes the evaluation, inputting the current case description data into the case element analysis model to output the case element analysis result corresponding to the current case description data; Wherein, the case elements include case types, drainage connection methods, transfer methods, fraud-related websites, and fraud-related software. The prompt engineering consists of prompt words, input examples, output examples, and real inputs. The pre-annotated data set includes prompt words, real inputs, and pre-annotated results. The instruction data set is a subset of the verification data set.
2. The method for analyzing elements of a telecommunications fraud case based on a large language model according to claim 1, wherein The pre-annotating the initial data set through prompt engineering to obtain a pre-annotated data set, and dividing the pre-annotated data set into multiple subsets to correct the pre-annotated results in each subset to obtain a verification data set includes: Based on a large model with a large number of parameters combined with the prompt engineering, pre-annotating the initial data set, and at the same time setting the output format of the large model with a large number of parameters to a JSON structure to obtain the prompt words and pre-annotated results corresponding to each real input, so as to construct the pre-annotated data set; Extracting a first subset from the pre-annotated data set, and checking and verifying each real input and the corresponding prompt words and pre-annotated results in the first subset to correct the pre-annotated results to obtain the verification data set corresponding to the first subset.
3. The method for analyzing elements of telecommunications fraud cases based on a large language model according to claim 2, wherein The screening out an instruction data set from the verification data set, using the case description data in the instruction data set as the input of the large model, and using the case elements as the output of the large model to train the large model to obtain a case element analysis model includes: Inferring each group of data in the prime number verification data set through the prompt engineering by different numbers of parameters llama3 in the large model with a large number of parameters to obtain a first inference result and a second inference result, and calculating the inference loss of the first inference result and the second inference result based on the check and verification result; Based on the inference losses of the first inference result and the second inference result, the first inference result with an inference loss lower than the first threshold and the second inference result with an inference loss exceeding the second threshold are selected, and the first inference result with an inference loss lower than the first threshold and the second inference result with an inference loss exceeding the second threshold are sorted in ascending order and descending order respectively to select the intersection data with a ranking not lower than the third threshold; Among them, the instruction data set is composed of the intersection data.
4. The method for analyzing elements of telecommunications fraud cases based on a large language model according to claim 3, wherein The different numbers of parameters llama3 include the first number of parameters llama3 and the second number of parameters llama3, and the first number is greater than the second number; Screening out the instruction data set from the verification data set, using the case description data in the instruction data set as the input of the large model, and using the case elements as the output of the large model to train the large model to obtain a case element analysis model further includes: Obtaining the labeled JSON standardized instruction data, and using various incorrect JSON formats as the input of the large model to obtain the corrected JSON format based on the labeled JSON standardized instruction data to obtain an instruction data subset; Merging the instruction data subset into the instruction data set, and performing full-parameter fine-tuning training on the second number of parameters llama3 based on the instruction data set corresponding to the second number of parameters llama3, and setting the hyperparameters of the fine-tuning training to perform fine-tuning training on the large model.
5. The method for analyzing elements of telecommunications fraud cases based on a large language model according to claim 4, wherein Selecting a plurality of labeled and verified case description data and case elements to construct an evaluation data set to evaluate the case element analysis model, including: Evaluating the case element analysis model from the dimensions of case type, drainage connection method, transfer method, fraud-related website, and fraud-related software based on the evaluation data set to obtain a plurality of evaluation results; When the ratio between the sum of the plurality of evaluation results and the number of the plurality of evaluation results is not lower than the fourth threshold, it is determined that the evaluation of the case element analysis model passes; When the ratio between the sum of the plurality of evaluation results and the number of the plurality of evaluation results is lower than the fourth threshold, it is determined that the evaluation of the case element analysis model fails, and the instruction data set is re-screened and the large model is fine-tuned; Among them, the plurality of evaluation results are the probabilities that the prediction results of the case type, drainage connection method, transfer method, fraud-related website, and fraud-related software coincide with their respective true results.
6. The method for analyzing elements of a telecommunications fraud case based on a large language model according to claim 5, wherein When the evaluation of the case element analysis model passes, inputting the current case description data into the case element analysis model to output the case element analysis result corresponding to the current case description data, including: Obtaining the current case description data and the corresponding prompt words, and using the current case description data and the corresponding prompt words as the input of the case element analysis model that has passed the evaluation; Calling the case element analysis model by setting hyperparameters in different intervals to output multiple groups of case element analysis results, and voting on the multiple groups of case element analysis results to select the best case element analysis result.
7. The method for analyzing elements of telecommunications fraud cases based on a large language model according to claim 6, wherein The method further includes: When the analysis result of the case elements is not a JSON structure, replace the prompt word corresponding to the current case description data with a JSON instruction prompt word; Based on the JSON instruction prompt word, combine the analysis result of the case elements in the non-JSON structure as the input of the large model to output the analysis result of the case elements in the standard JSON structure.
8. An element analysis device for telecommunications fraud cases based on a large language model, characterized in that, The device includes: A dataset construction module, configured to obtain case description data and the case elements corresponding to the case description data, and construct an initial dataset based on the case description data and the case elements; A data verification module, configured to pre-annotate the initial dataset through prompt engineering to obtain a pre-annotated dataset, and divide the pre-annotated dataset into multiple subsets to correct the pre-annotated results in each subset to obtain a verified dataset; A large model training module, configured to screen out an instruction dataset from the verified dataset, use the case description data in the instruction dataset as the input of the large model, and use the case elements as the output of the large model to train the large model to obtain a case element analysis model; A model evaluation module, configured to select multiple annotated and verified case description data and case elements to construct an evaluation dataset to evaluate the case element analysis model; A case element analysis module, configured to, when the case element analysis model passes the evaluation, input the current case description data into the case element analysis model to output the analysis result of the case elements corresponding to the current case description data; Wherein, the case elements include case types, drainage connection methods, transfer methods, fraud-related websites, and fraud-related software, the prompt engineering is composed of prompt words, input examples, output examples, and real inputs, the pre-annotated dataset includes prompt words, real inputs, and pre-annotated results, and the instruction dataset is a subset of the verified dataset.
9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer storage medium stores a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.