File risk assessment method and device, equipment, storage medium and program product
The risk assessment of the statements to be evaluated in the procurement documents through a large language model, and the risk basis is selected according to the semantic similarity and correlation, which solves the problem of legal basis deviation in compliance risk assessment and improves the accuracy of legal risk assessment of procurement documents.
Patent Information
- Application Number
- CN202510315237.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, due to frequent changes in industry standards in the field of compliance risk assessment, the legal basis for outputting pre-trained large-scale language models is deviated, and the legal risk assessment results of procurement documents are relatively low.
The risk assessment is performed on multiple statements to be evaluated in the evaluation file through a large language model. According to the semantic similarity and correlation between the risk statement and the risk sample statement, the risk basis is selected and the risk assessment results are corrected.
It improves the accuracy of the legal risk assessment results of procurement documents, avoids the legal basis deviation caused by frequent changes in industry standards, and ensures the reliability of the risk assessment results.
Smart Images

Figure CN120256624A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a file risk assessment method, device, equipment, storage medium and program product. Background Art
[0002] In the related art, compliance risk assessment of procurement documents is performed through large-scale language models to obtain compliance risk assessment results of procurement documents. However, in the field of compliance risk assessment, industry standards change frequently, which may cause the legal basis output by the pre-trained large-scale language model to be biased, resulting in low accuracy of legal risk assessment results of procurement documents. Therefore, how to improve the accuracy of legal risk assessment results of procurement documents needs to be solved urgently. Summary of the invention
[0003] The main purpose of this application is to provide a document risk assessment method, device, equipment, storage medium and program product, aiming to solve the technical problem of how to improve the accuracy of the legal risk assessment results of procurement documents.
[0004] To achieve the above objectives, the present application proposes a document risk assessment method, which includes:
[0005] Get multiple statements to be evaluated in the file to be evaluated;
[0006] For each sentence to be evaluated, the corresponding first risk assessment prompt sentence is input into the first large language model for risk assessment to obtain a plurality of initial risk assessment results, wherein the first risk assessment prompt sentence is constructed by the sentence to be evaluated according to the first preset prompt word template;
[0007] For each statement to be evaluated, if the initial risk assessment result is that there is a risk, multiple first selected risk bases are screened out from multiple preset risk bases, wherein the semantic similarity between the risk statement corresponding to the initial risk assessment result and the selected risk sample statements corresponding to each first selected risk base satisfies a first preset screening condition;
[0008] For each risk statement, the first selected risk bases are ranked according to the correlation between the first selected risk bases and the risk base of the risk statement, and the target risk base is determined;
[0009] Based on all initial risk assessment results and / or all target risk evidence, the risk assessment results of the document to be assessed are obtained.
[0010] In one embodiment, the step of selecting a plurality of first selected risk criteria from a plurality of preset risk criteria includes:
[0011] Determine multiple selected risk example statements with the first vector distance meeting the first preset screening condition from the preset risk example statement library according to the first vector distance between the example statement vectors of each preset risk example statement in the preset risk example statement library and the risk statement vector of the risk statement, and obtain corresponding multiple first selected risk bases.
[0012] In one embodiment, before the step of determining the target risk basis by performing a relevance ranking on each first selected risk basis and the risk basis of the risk statement for each risk statement, the document evaluation method further includes:
[0013] For each risk statement, input the corresponding second risk assessment prompt statement into the second large language model for risk assessment, and determine the second selected risk basis from multiple first selected risk bases, where the second risk assessment prompt statement is constructed from multiple first selected risk bases and the risk statement according to the second preset prompt word template;
[0014] Screen out multiple third selected risk bases from the preset risk basis library, where the semantic similarity between the third selected risk basis and the second selected risk basis meets the second preset condition;
[0015] The step of performing a relevance ranking on each first selected risk basis according to the relevance between each first selected risk basis and the risk basis of the risk statement for each risk statement to determine the target risk basis includes:
[0016] For each risk statement, perform a relevance ranking on each third selected risk basis according to the relevance between each third selected risk basis and the second selected risk basis, and determine the target risk basis.
[0017] In one embodiment, the step of screening out multiple third selected risk bases from the preset risk basis library includes:
[0018] According to the second vector distance between the preset basis vector of each preset risk basis in the preset risk basis library and the selected basis vector of the second selected risk basis, screen out multiple third selected risk bases from the preset risk basis library whose second vector distance meets the second preset screening condition.
[0019] In one embodiment, the document risk assessment method further includes:
[0020] Construct a risk example statement sample set and an initial large language model;
[0021] Fine-tune and train the initial large language model using the risk example statement sample set to obtain the first large language model.
[0022] In one embodiment, the step of obtaining multiple statements to be evaluated in the document to be evaluated includes:
[0023] Obtain the file to be evaluated;
[0024] Split the file to be evaluated into sentences to obtain multiple sentences to be evaluated.
[0025] In addition, to achieve the above object, the present application also proposes a file risk assessment device, which includes:
[0026] An acquisition module, configured to acquire multiple sentences to be evaluated in the file to be evaluated;
[0027] A model evaluation module, configured to input the corresponding first risk assessment prompt sentence into the first large language model for risk assessment for each sentence to be evaluated, to obtain multiple initial risk assessment results, wherein the first risk assessment prompt sentence is constructed from the sentence to be evaluated according to the first preset prompt word template;
[0028] A screening module, configured to screen out multiple first selected risk bases from multiple preset risk bases for each sentence to be evaluated, wherein the semantic similarity between the risk sentence corresponding to the initial risk assessment result and the selected risk example sentences corresponding to each first selected risk base meets the first preset screening condition;
[0029] A sorting module, configured to sort the relevance of each first selected risk base according to the relevance between each first selected risk base and the risk base of the risk sentence, to determine the target risk base;
[0030] A result evaluation module, configured to obtain the risk assessment result of the file to be evaluated according to all the initial risk assessment results and / or all the target risk bases.
[0031] In addition, to achieve the above object, the present application also proposes a file risk assessment device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the file risk assessment method as described above.
[0032] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by the processor, it implements the steps of the file risk assessment method as described above.
[0033] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by the processor, it implements the steps of the file risk assessment method as described above.
[0034] One or more technical solutions proposed in this application have at least the following technical effects:
[0035] This application provides a method, device, equipment, storage medium and program product for file risk assessment. After using a large language model to perform risk assessment on multiple statements to be evaluated in a file to be evaluated, when the initial risk assessment result obtained indicates a risk, multiple selected risk example statements are determined according to the semantic similarity between the risk statement and the risk example statements, and the corresponding first selected risk basis is obtained. Then, according to the relevance between the risk basis of the risk statement and the risk basis of the risk example statements, the multiple first selected risk bases are sorted according to their relevance, and the target risk basis is determined as the risk basis of the risk statement. Thus, when it is detected that a statement to be evaluated is a risk statement, a similarity search is performed on the risk statement output by the large language model, which can ensure the semantic similarity between the risk statement and the selected risk example statements, and sort the multiple first selected risk bases corresponding to the selected risk example statements according to their relevance, thereby improving the accuracy of the target risk basis, realizing the correction of the risk basis of the risk statement in the risk assessment result, and improving the risk assessment result. When applied to the legal risk assessment of procurement documents, it can avoid the deviation of legal basis caused by frequent changes in industry standards, obtain a more accurate legal basis for procurement documents, and improve the accuracy of the legal risk assessment result of procurement documents. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0037] In order to more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or related technologies. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 Schematic flowchart of the first embodiment of the file risk assessment method of this application;
[0039] Figure 2 Schematic diagram of an exemplary first preset prompt word template provided for the file risk assessment method of this application;
[0040] Figure 3 Schematic diagram of the data format in the preset risk example statement library provided for the file risk assessment method of this application;
[0041] Figure 4 Schematic diagram of the annotation of the risk example statement samples in the risk example statement sample set provided for the file risk assessment method of this application;
[0042] Figure 5 Schematic diagram of modules of the first embodiment of the risk assessment device for the application document;
[0043] Figure 6 Schematic diagram of the device structure of the hardware operating environment involved in the document risk assessment method in the embodiment of the application.
[0044] The realization of the purpose of the application, functional features and advantages will be further described in combination with the embodiments with reference to the accompanying drawings. Specific embodiments
[0045] It should be understood that the specific embodiments described herein are only used to explain the technical solution of the application and are not used to limit the application.
[0046] To better understand the technical solution of the application, the following will be described in detail in combination with the drawings of the specification and specific embodiments.
[0047] The main solution of the embodiment of the application is: for each statement to be evaluated, input the corresponding first risk assessment prompt statement into the first large language model for risk assessment to obtain multiple initial risk assessment results, where the first risk assessment prompt statement is constructed from the statement to be evaluated according to the first preset prompt word template; for each statement to be evaluated, if the initial risk assessment result is that there is a risk, then screen out multiple first selected risk bases from multiple preset risk bases, where the semantic similarity between the risk statement corresponding to the initial risk assessment result and the selected risk example statements corresponding to each first selected risk base meets the first preset screening condition; for each risk statement, perform a relevance ranking on each first selected risk base according to the relevance between each first selected risk base and the risk base of the risk statement to determine the target risk base; obtain the risk assessment result of the document to be evaluated according to all the initial risk assessment results and / or all the target risk bases.
[0048] Different from other text files in daily life, procurement documents are a collection of a series of documents and materials created and used during the procurement process. Together, they form the legal, management, and operational basis for procurement activities. These documents are designed to ensure the transparency, fairness, compliance, and efficiency of the procurement process. Therefore, in order to ensure the legality of procurement activities, reduce potential risks, improve procurement efficiency and quality, and maintain the reputation of the enterprise, it is necessary to conduct a compliance risk analysis of procurement documents. To identify the compliance risks of procurement documents, it is first necessary to collect all relevant laws, regulations, industry norms, and internal management systems of the enterprise, such as the "Measures for Bidding and Tendering Management" and the "Procurement Management Measures". Then, each clause in the procurement documents is compared item by item with the collected laws and regulations to check whether there are any contents in the procurement documents that violate the laws and regulations. Currently, the main methods for identifying the risks of procurement documents are that the enterprise sets up a special team or designates a specific person to be responsible for the compliance analysis of procurement documents, or uses regular expressions to match the fields in the procurement documents to determine whether there are compliance risks in the procurement documents.
[0049] However, procurement documents often contain complex terms, conditions, and regulations, and the forms and contents of these documents may vary depending on the supplier, project, or industry. Although manual review of procurement documents can ensure a certain degree of meticulousness and flexibility, there are problems of low efficiency and poor accuracy. Since the requirements for compliance risk analysis of procurement documents change with the continuous update of laws, regulations, and industry standards, the solution of using regular expressions to match the fields in procurement documents to determine whether there are compliance risks in the procurement documents is difficult to adapt to new requirements, and it only focuses on the text structure and is difficult to detect semantic compliance issues, such as vague expressions or implicit violations.
[0050] In addition, it is also possible to conduct a legal risk assessment of procurement documents through large language models. However, in the field of compliance risk assessment, industry standards change frequently, which may lead to deviations in the legal basis output by pre-trained large language models, resulting in relatively low accuracy of the legal risk assessment results of procurement documents.
[0051] Therefore, the present application provides a solution. After a large language model performs risk assessment on multiple statements to be evaluated in a document to be evaluated, when the initial risk assessment result obtained indicates a risk, multiple selected risk example statements are determined according to the semantic similarity between the risk statement and the risk example statements, a corresponding first selected risk basis is obtained, and the multiple first selected risk bases are sorted according to their relevance to the risk basis of the risk statement, and the target risk basis is determined as the risk basis of the risk statement. Thus, when a statement to be evaluated is detected as a risk statement, similarity retrieval is performed on the risk statement output by the large language model, which can ensure the semantic similarity between the risk statement and the selected risk example statements, and sort the multiple first selected risk bases corresponding to the selected risk example statements according to their relevance, thereby improving the accuracy of the target risk basis, realizing the correction of the risk basis of the risk statement in the risk assessment result, and improving the risk assessment result. When applied to the legal risk assessment of procurement documents, it is possible to avoid the deviation of legal bases caused by frequent changes in industry standards, obtain more accurate legal bases for procurement documents, and improve the accuracy of the legal risk assessment result of procurement documents.
[0052] Moreover, the present application constructs a second risk assessment prompt statement according to the second preset prompt word template, in combination with multiple first selected risk bases and the risk statement, and performs risk assessment through a second large language model, providing more in-depth, accurate and valuable input for the second large language model, and improving the reliability of the second selected risk basis.
[0053] In addition, the present application integrates multiple large language models to jointly complete the document risk assessment task. In the proposed architecture of multi-model collaborative work, different large language models are assigned different tasks, enabling each large language model to play its own advantages, gradually optimizing the risk assessment result, avoiding the limitations and biases of a single model, and improving the accuracy and reliability of the risk assessment result.
[0054] It should be noted that the execution subject of this embodiment can be a computing service device with functions of document risk assessment, network communication, and program operation, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a document risk assessment device, etc. that can implement the above functions. Hereinafter, the document risk assessment device is taken as an example to illustrate this embodiment and the following embodiments.
[0055] Based on this, the embodiments of the present application provide a document risk assessment method, referring to Figures 1 to 6 , Figure 1 is a schematic flowchart of the first embodiment of the document risk assessment method of the present application, Figure 2 is a schematic diagram of a template of an exemplary first preset prompt word template provided by the document risk assessment method of the present application, Figure 3Schematic diagram of the data format in the preset risk example statement library provided for the risk assessment method of this application document Figure 4 Schematic diagram of the annotation of the risk example statement samples in the risk example statement sample set provided for the risk assessment method of this application document
[0056] In this embodiment, the document risk assessment method may include steps S100 to S500:
[0057] Step S100, obtain multiple statements to be evaluated in the document to be evaluated.
[0058] It should be noted that the document risk assessment method of this embodiment can be used to perform risk assessments such as legality, fairness, and effectiveness on the document to be evaluated. Hereinafter, the legality assessment of the procurement document will be taken as an example for illustration.
[0059] In a feasible implementation manner, step S100 may include: obtain the document to be evaluated; perform sentence splitting on the document to be evaluated to obtain multiple statements to be evaluated.
[0060] It should be noted that the document to be evaluated is a procurement document to be subjected to a legality assessment, and the document to be evaluated can be determined according to the user's selection. Performing sentence splitting on the document to be evaluated can split the text content in the document to be evaluated into multiple statements to be evaluated. Among them, using pySBD (Python Sentence Boundary Disambiguation) of the Python library, the document to be evaluated can be subjected to sentence splitting.
[0061] Step S200, for each statement to be evaluated, input the corresponding first risk assessment prompt statement into the first large language model for risk assessment to obtain multiple initial risk assessment results.
[0062] Among them, the first risk assessment prompt statement is constructed from the statement to be evaluated according to the first preset prompt word template.
[0063] It should be noted that the initial risk assessment results may include the risk results, risk type results, risk basis results, etc. of the statements to be evaluated. Among them, the risk results may include the existence of legal risks (with risks) and no legal risks (without risks); the risk type results may include risk types 1 to N, where N is the total number of risk types, which is determined according to the risk category list specified during the training of the first large language model. The risk basis results include at least one legal basis (risk basis).
[0064] The first preset prompt word template is the Prompt format of the first large language model. Before using the first large language model for risk assessment, it is necessary to first convert each statement to be evaluated into the Prompt format of the first large language model.
[0065] In one example, the risk types may include risk-free, at-risk - risk type 1... at-risk - risk type 13, corresponding to the first preset prompt template as Figure 2 shown. Among them, the sentence corresponding to the input "input" is the statement to be evaluated, and "category" is the category of the classification task, that is, the risk type. The assembled first risk assessment prompt statement is input into the first large language model risk-classify-model (risk classification model), and the output "output" corresponds to a label including one of the 14 risk types.
[0066] Step S300, for each statement to be evaluated, if the initial risk assessment result is at risk, then select multiple first selected risk bases from multiple preset risk bases.
[0067] Among them, the semantic similarity between the risk statement corresponding to the initial risk assessment result and the selected risk example statements corresponding to each first selected risk base satisfies the first preset screening condition.
[0068] It should be noted that if the initial risk assessment result is risk-free, then there is no legal risk for the corresponding statement to be evaluated, and the corresponding initial risk assessment result is directly output.
[0069] If the initial risk assessment result is at risk, then there is a legal risk for the statement to be evaluated corresponding to the initial risk assessment result, and this statement to be evaluated is a risk statement. The preset risk bases are the legal bases corresponding to multiple preset risk example statements in the pre-constructed preset risk example statement library, and the first selected risk bases are the legal bases corresponding to the selected risk example statements. The preset risk example statements can be the risk example sentences of procurement documents marked by professionals in the field of procurement document review. In the preset risk example statement library, as Figure 3 shown, the data is provided in tabular form. The table includes procurement document risk example sentences (preset risk example statements), risk types, whether there is a risk, and the involved legal bases.
[0070] The multiple selected risk example statements may include the preset risk example statements with a semantic similarity greater than the first preset similarity threshold to the risk statement, or may include the first preset number of preset risk example statements ranked first or last after sorting the multiple preset risk example statements according to the semantic similarity between the risk statement and each preset risk example statement in the preset risk example statement library. Among them, the first preset number can be 5.
[0071] Step S400: For each risk statement, based on the correlation between each first selected risk basis and the risk basis of the risk statement, perform a correlation ranking on each first selected risk basis to determine the target risk basis.
[0072] It should be noted that the target risk basis may include at least one legal basis, that is, the target risk basis includes at least one first selected risk basis. The correlation ranking can be implemented using a retrieval ranking model. The retrieval ranking model can be the BGE Re-Ranker Large model. The retrieval ranking model takes the risk basis of the risk statement and each first selected risk basis as inputs and can directly output the similarity scores between each first selected risk basis and the risk basis of the risk statement, and rank multiple first selected risk bases. Among them, the target risk basis may include the first second preset number of first selected risk bases ranked first or last. For example, the second preset number can be 3.
[0073] Step S500: Obtain the risk assessment result of the document to be evaluated based on all the initial risk assessment results and / or all the target risk bases.
[0074] It should be noted that by summarizing the initial risk assessment results and / or target risk bases of each statement to be evaluated, the risk assessment result of the document to be evaluated can be obtained. Among them, the risk assessment result of the document to be evaluated may include each statement to be evaluated and whether each statement to be evaluated has risks, the corresponding risk types, and the legal bases involved.
[0075] Thus, this embodiment provides a method for document risk assessment. After a large language model performs a risk assessment on multiple statements to be evaluated in a document to be evaluated, when the obtained initial risk assessment result indicates the existence of risks, based on the semantic similarity between the risk statement and the risk sample statement, multiple selected risk sample statements are determined, the corresponding first selected risk bases are obtained, and according to the correlation with the risk basis of the risk statement, the multiple first selected risk bases are ranked in terms of correlation to determine the target risk basis as the risk basis of the risk statement. Therefore, when a statement to be evaluated is detected as a risk statement, a similarity retrieval is performed on the risk statement output by the large language model, which can ensure the semantic similarity between the risk statement and the selected risk sample statement, and rank the first selected risk bases corresponding to the multiple selected risk sample statements in terms of correlation, thereby improving the accuracy of the target risk basis, realizing the correction of the risk basis of the risk statement in the risk assessment result, and improving the risk assessment result. When applied to the legal risk assessment of procurement documents, it can avoid the legal basis deviation caused by frequent changes in industry standards, obtain more accurate legal bases for procurement documents, and improve the accuracy of the legal risk assessment result of procurement documents.
[0076] In a feasible implementation, the document evaluation method may further include: constructing a risk example sentence sample set and an initial large language model; using the risk example sentence sample set to fine-tune and train the initial large language model to obtain a first large language model.
[0077] It should be noted that the risk example sentence sample set may include multiple risk example sentence samples, risk markers corresponding to each risk example sentence sample, risk type markers, and risk basis markers.
[0078] In one example, as Figure 4 shown, the sentence corresponding to the input "input" is a risk example sentence sample, "catagory" is the category of the classification task, that is, the risk type, and the "output" corresponding to the "label1" and "label2" are risk type markers.
[0079] In this embodiment, the risk example sentence sample set may include 5,000 risk example sentence samples, and the preset risk example sentence library may also include the 5,000 risk example sentences.
[0080] The initial large language model may be the Qwen2-1.5B-Instruct model. By using the risk example sentence sample set to perform classification fine-tuning on the Qwen2-1.5B-Instruct model, a first large language model "risk-classify-model" applicable to the risk assessment of procurement documents can be obtained. Among them, the classification fine-tuning may be Lora fine-tuning.
[0081] In specific implementation, multiple risk example sentence samples are formatted into multiple sample data according to the first preset prompt template, and the multiple sample data are divided into a training set, a validation set, and a test set according to a ratio of 8:1:1 to perform Lora classification fine-tuning on the model Qwen2-1.5B-Instruct. Among them, compared with full-parameter fine-tuning, Lora fine-tuning realizes lightweight fine-tuning of the model by introducing a low-rank matrix into the weight matrix of the model, effectively reducing the number of model parameters, reducing the computational requirements during model training and inference, and being able to quickly adapt to different application scenarios. And Qwen2-1.5B-Instruct is an open-source large language model with a parameter scale of 1.5 billion. It is one of the stronger ones among small-scale parameter language models and has strong language understanding ability. For scenarios with complex content such as compliance documents, it can more effectively understand context information. Since the model has a small number of parameters and supports Lora fine-tuning, the training of this model has low requirements for video memory, which enables effective training even in an environment with limited resources.
[0082] In addition, the fine-tuning environment of the Qwen2-1.5B-Instruct model can be based on the Ubuntu 20.04.6 LTS operating system, equipped with 1 NVIDIA V100S GPU (32G video memory). The adjustable parameters mainly relied on for LoRA fine-tuning are the LoRA rank and the LoRA alpha parameter. The LoRA rank refers to the size of the low-rank matrix in LoRA fine-tuning. Selecting a smaller rank value can significantly reduce the number of parameters to be fine-tuned, thereby improving the training efficiency and reducing the computational cost. The LoRA alpha parameter is a scaling factor used to control the impact of the low-rank matrix product on the original weight matrix. Its role is to ensure that the low-rank matrix product numerically matches the size of the original matrix to prevent numerical instability problems during training; the LoRA alpha parameter can also adjust the amplitude of fine-tuning to ensure that the impact on the original model is not too large, so as to introduce new features while maintaining the original performance of the model. Preferably, the model is repeatedly trained by arranging and combining the LoRA rank as 6, 8, 10 and the LoRA alpha parameter as 16, 32, 64 respectively, and the number of training rounds for each time is 5000 rounds. During the training process, the model weights with the minimum loss obtained on the validation set are saved. It is found that when the model is trained to about 2400 rounds, the value of the loss function on the validation set tends to be stable. Finally, the obtained multiple model weights are respectively tested on the test set, and it is found that when the LoRA rank is 8 and the LoRA alpha is 32, the classification accuracy of the model on the test set is the highest. Therefore, this model is selected as the first large language model risk-classify-model finally applicable to the risk assessment of procurement documents.
[0083] Thus, this embodiment provides a file risk assessment method, which selects a large language model as a risk classification model according to the actual usage scenario and trains the large language model according to LoRA fine-tuning, so that a highly reliable first large language model can be trained by using limited sample resources.
[0084] In a feasible implementation manner, step S300 may include: determining, from the preset risk example statement library, multiple selected risk example statements whose first vector distance satisfies the first preset screening condition according to the first vector distance between the example statement vectors of each preset risk example statement and the risk statement vector of the risk statement, and obtaining the corresponding multiple first selected risk bases.
[0085] It should be noted that the preset risk example statement library may include example statement vectors corresponding to each preset risk example statement. Before executing the file evaluation method of this embodiment, each preset risk example statement can be converted into a corresponding example statement vector by using a vector model. Among them, the vector model can be the BGE-M3 model, and the BGE-M3 model can convert the labeled preset risk example statements into 1024-dimensional example statement vectors. Storing multiple preset risk example statements, the corresponding example statement vectors, the corresponding risk categories, and the corresponding legal bases into OpenSearch can obtain the preset risk example statement library base-index.
[0086] It can be understood that the risk statement vector can be obtained by converting the risk statement using a vector model. After obtaining the risk statement vector, the risk statement vector can be stored in the preset risk example statement library to update the preset risk example statement library. The multiple selected risk example statements may include preset risk example statements with a first vector distance less than a first preset distance threshold, or may include the first preset number of preset risk example statements arranged at the front or the back after sorting multiple preset risk example statements according to multiple first vector distances.
[0087] In a specific implementation, the built-in K-Nearest Neighbors (KNN) algorithm in OpenSearch is used to retrieve in the vector retrieval library base-index and set to return the first preset number of example statement vectors most similar to the risk statement vector.
[0088] Therefore, this embodiment provides a text risk assessment method. By constructing a vector retrieval library including multiple example statement vectors and using vector retrieval to determine multiple selected risk example statements from the preset risk example statement library, the semantic similarity of the corresponding multiple first selected risk bases is ensured.
[0089] In a feasible implementation manner, before step S400, the file evaluation method may further include steps S600 to S700:
[0090] Step S600, for each risk statement, input the corresponding second risk assessment prompt statement into the second large language model for risk assessment, and determine the second selected risk basis from multiple first selected risk bases.
[0091] Among them, the second risk assessment prompt statement is constructed from multiple first selected risk bases and the risk statement according to a second preset prompt word template.
[0092] Step S700, screen out multiple third selected risk bases from the preset risk basis library.
[0093] Among them, the semantic similarity between the third selected risk basis and the second selected risk basis meets the second preset condition.
[0094] The corresponding step S400 may include: for each risk statement, according to the relevance between each third selected risk basis and the second selected risk basis, perform a relevance ranking on each third selected risk basis to determine the target risk basis.
[0095] It should be noted that in this embodiment, the second large language model can also be used to determine the second selected risk basis that is most similar to the risk statement from multiple first selected risk bases, match multiple third selected risk bases with semantic similarity to the second selected risk basis from the preset risk basis library, and determine the target risk basis based on the multiple third selected risk bases. Among them, since the second large language model has a certain degree of hallucination, multiple third selected risk bases can be used to correct the second selected risk basis. The second large language model can be the Qwen2-7B-Instruct model, and the second preset prompt template is the Prompt format of the second large language model. Before using the second large language model for risk assessment, each risk statement and the corresponding multiple first selected risk bases need to be converted into the Prompt format of the second large language model.
[0096] In an example, if the multiple first selected risk bases include 5 legal bases, the second preset prompt template can be: "Given the following 5 legal bases:\n{Legal basis 1},{Legal basis 2},{Legal basis 3},{Legal basis 4},{Legal basis 5}\n\nPlease help output the relevant legal bases involved in the following sentence: {sentence}", where sentence represents the risk statement. Input the second risk assessment prompt statement into the Qwen2-7B-Instruct model, and the model will output the target legal basis MLbasis involved in the risk statement according to its understanding.
[0097] It can be understood that the preset risk basis library may include multiple preset risk bases, and the multiple preset risk bases can be legal and regulatory bases related to multiple procurement documents. The multiple third selected risk bases may include preset risk bases with a semantic similarity greater than the second preset similarity threshold to the second selected risk basis, or may include the first or last third preset number of preset risk bases after sorting the multiple preset risk bases according to the semantic similarity between the second selected risk basis and each third selected risk basis. Among them, the third preset number can be 10.
[0098] Correspondingly, the target risk basis may include at least one third selected risk basis. The retrieval and ranking model takes the second selected risk basis and each third selected risk basis as inputs, and can directly output the similarity scores between the second selected risk basis and each third selected risk basis, and rank multiple third selected risk bases. Among them, the target risk basis may include the second preset number of third selected risk bases arranged in the front or in the back.
[0099] Thus, this embodiment provides a file risk assessment method. By constructing a preset risk basis library and using multiple third selected risk bases in the preset risk basis library to correct the second selected risk basis, it is possible to avoid the inaccuracy of the second selected risk basis caused by the hallucination of the second large language model, and further improve the accuracy of the file risk assessment result.
[0100] In a feasible implementation manner, step S700 may include: screening out multiple third selected risk bases whose second vector distances satisfy the second preset screening condition from the preset risk basis library according to the second vector distances between the preset basis vectors of each preset risk basis in the preset risk basis library and the selected basis vector of the second selected risk basis.
[0101] It should be noted that the preset risk basis library may include the preset basis vectors corresponding to each preset risk basis. Before executing the file evaluation method of this embodiment, each preset risk basis can be converted into a corresponding preset basis vector by using a vector model. Among them, the vector model can be the BGE-M3 model, and the BGE-M3 model can convert the preset risk basis into a 1024-dimensional preset basis vector. Storing multiple preset risk bases and their corresponding preset basis vectors in OpenSearch can obtain the preset risk basis library low-index.
[0102] It can be understood that the selected basis vector can be obtained by converting the second selected risk basis by using a vector model. The multiple third selected risk bases may include multiple preset risk bases whose second vector distances are less than the second preset distance threshold, or may include the third preset number of preset risk bases arranged in the front or in the back after ranking multiple preset risk bases according to multiple second vector distances.
[0103] In specific implementation, use the built-in K-Nearest Neighbors (KNN) algorithm of opensearch to retrieve in the preset risk basis library low-index and set to return the third preset number of preset risk bases that are most similar to the selected basis vector. Among them, the content corresponding to the law_base field in each return result is the third selected risk basis.
[0104] In addition, it can be understood that for the risk basis of each risk statement, the risk basis can be converted into a risk basis vector by using a vector model and stored in a preset risk basis library to update the preset risk basis library.
[0105] Therefore, this embodiment provides a file risk assessment method. By constructing a vector retrieval library including multiple preset basis vectors and using vector retrieval to determine multiple third selected risk bases from the preset risk basis library, the semantic similarity of the corresponding multiple third selected risk bases is ensured.
[0106] Moreover, this embodiment integrates multiple large language models to jointly complete the file risk assessment task. In the proposed architecture of multi-model collaborative work, different large language models are given different divisions of labor, enabling each large language model to play its own advantages, gradually optimizing the risk assessment results, avoiding the limitations and biases of a single model, and improving the accuracy and reliability of the risk assessment results.
[0107] This application provides a file risk assessment device, as Figure 5 shown. The file risk assessment device may include:
[0108] An acquisition module 10 for acquiring multiple statements to be evaluated in the file to be evaluated;
[0109] A model evaluation module 20 for, for each statement to be evaluated, inputting the corresponding first risk assessment prompt statement into the first large language model for risk assessment to obtain multiple initial risk assessment results, where the first risk assessment prompt statement is constructed from the statement to be evaluated according to the first preset prompt word template;
[0110] A screening module 30 for, for each statement to be evaluated, if the initial risk assessment result is that there is a risk, screening out multiple first selected risk bases from multiple preset risk bases, where the semantic similarity between the risk statement corresponding to the initial risk assessment result and the selected risk example statements corresponding to each first selected risk basis satisfies the first preset screening condition;
[0111] A sorting module 40 for, for each risk statement, sorting the first selected risk bases according to the correlation between each first selected risk basis and the risk basis of the risk statement to determine the target risk basis;
[0112] A result evaluation module 50 for obtaining the risk assessment result of the file to be evaluated according to all the initial risk assessment results and / or all the target risk bases.
[0113] For more implementation details in the specific implementation manner of the above file risk assessment device, refer to the description of the specific implementation manner of the file risk assessment method in the above embodiments. For the sake of brevity of the specification, it will not be repeated here.
[0114] The present application provides a file risk assessment device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the file risk assessment method in Embodiment 1 above.
[0115] Next, refer to Figure 6 , which shows a schematic structural diagram of a file risk assessment device suitable for implementing the embodiments of the present application. The file risk assessment device in the embodiments of the present application may include, but is not limited to, mobile terminals such as laptop computers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), etc., and fixed terminals such as desktop computers. Figure 6 The file risk assessment device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0116] As Figure 6As shown, the document risk assessment device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the document risk assessment device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the document risk assessment device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a document risk assessment device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be implemented or had alternatively.
[0117] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are performed.
[0118] The document risk assessment device provided by the present application adopts the document risk assessment method in the above embodiments, and can solve the technical problem of how to improve the accuracy of the legal risk assessment result of procurement documents. Compared with the related art, the beneficial effects of the document risk assessment device provided by the present application are the same as those of the document risk assessment method provided by the above embodiments, and other technical features in the document risk assessment device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0119] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0120] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0121] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the file risk assessment method in the above embodiments.
[0122] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0123] The above computer-readable storage medium can be included in the file risk assessment device; it can also exist separately without being assembled into the file risk assessment device.
[0124] The above computer-readable storage medium carries one or more programs, which, when executed by the file risk assessment device, enable the file risk assessment device to obtain multiple statements to be evaluated in the file to be evaluated; for each statement to be evaluated, input the corresponding first risk assessment prompt statement into the first large language model for risk assessment to obtain multiple initial risk assessment results, where the first risk assessment prompt statement is constructed from the statement to be evaluated according to the first preset prompt word template; for each statement to be evaluated, if the initial risk assessment result is that there is a risk, then select multiple first selected risk bases from multiple preset risk bases, where the semantic similarity between the risk statement corresponding to the initial risk assessment result and the selected risk example statements corresponding to each first selected risk basis meets the first preset screening condition; for each risk statement, sort the first selected risk bases according to the relevance between each first selected risk basis and the risk basis of the risk statement to determine the target risk basis; obtain the risk assessment result of the file to be evaluated according to all the initial risk assessment results and / or all the target risk bases.
[0125] Computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as combinations of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0127] The modules described in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0128] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned document risk assessment method, which can solve the technical problem of how to improve the accuracy of the legal risk assessment results of procurement documents. Compared with the related art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the document risk assessment method provided by the above embodiments, and will not be elaborated here.
[0129] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the document risk assessment method as described above.
[0130] The computer program product provided by the present application can solve the technical problem of how to improve the accuracy of the legal risk assessment results of procurement documents. Compared with the related art, the beneficial effects of the computer program product provided by the present application are the same as those of the document risk assessment method provided by the above embodiments, and will not be elaborated here.
[0131] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A method for file risk assessment, characterized in that The described document risk assessment method includes: Obtaining multiple statements to be evaluated in the document to be evaluated; For each of the statements to be evaluated, inputting the corresponding first risk assessment prompt statement into a first large language model for risk assessment to obtain multiple initial risk assessment results, where the first risk assessment prompt statement is constructed from the statement to be evaluated according to a first preset prompt word template; For each statement to be evaluated, if the initial risk assessment result indicates a risk, screening out multiple first selected risk bases from multiple preset risk bases, where the semantic similarity between the risk statement corresponding to the initial risk assessment result and the selected risk example statements corresponding to each of the first selected risk bases meets a first preset screening condition; For each of the risk statements, sorting the first selected risk bases according to the relevance between each of the first selected risk bases and the risk basis of the risk statement to determine the target risk basis; Obtaining the risk assessment result of the document to be evaluated based on all the initial risk assessment results and / or all the target risk bases.
2. The document risk assessment method according to claim 1, wherein The step of screening out multiple first selected risk bases from multiple preset risk bases includes: Determining, from the preset risk example statement library, multiple selected risk example statements whose first vector distance meets the first preset screening condition according to the first vector distance between the example statement vectors of each preset risk example statement in the preset risk example statement library and the risk statement vector of the risk statement, to obtain multiple corresponding first selected risk bases.
3. The document risk assessment method according to claim 1, characterized in that, Before the step of, for each of the risk statements, sorting the first selected risk bases according to the relevance between each of the first selected risk bases and the risk basis of the risk statement to determine the target risk basis, the document assessment method further includes: For each of the risk statements, inputting the corresponding second risk assessment prompt statement into a second large language model for risk assessment to determine second selected risk bases from multiple first selected risk bases, where the second risk assessment prompt statement is constructed from multiple first selected risk bases and the risk statement according to a second preset prompt word template; Screening out multiple third selected risk bases from a preset risk basis library, where the semantic similarity between the third selected risk bases and the second selected risk bases meets a second preset condition; The step of, for each of the risk statements, sorting the first selected risk bases according to the relevance between each of the first selected risk bases and the risk basis of the risk statement to determine the target risk basis includes: For each of the risk statements, sorting the third selected risk bases according to the relevance between each of the third selected risk bases and the second selected risk bases to determine the target risk basis.
4. The document risk assessment method according to claim 3, wherein The step of screening out multiple third selected risk bases from the preset risk basis library includes: Based on the second vector distance between the preset basis vectors of each preset risk basis in the preset risk basis library and the selected basis vector of the second selected risk basis, multiple third selected risk bases whose second vector distance meets the second preset screening condition are screened out from the preset risk basis library.
5. The document risk assessment method according to any one of claims 1 to 4, characterized in that, The document risk assessment method further includes: Constructing a risk example sentence sample set and an initial large language model; Fine-tuning and training the initial large language model using the risk example sentence sample set to obtain the first large language model.
6. The document risk assessment method according to any one of claims 1 to 4, characterized in that The step of obtaining multiple to-be-evaluated sentences in the to-be-evaluated document includes: Obtaining the to-be-evaluated document; Splitting the to-be-evaluated document into sentences to obtain multiple to-be-evaluated sentences.
7. A file risk assessment device, characterized in that, The document risk assessment device includes: An acquisition module, configured to acquire multiple to-be-evaluated sentences in the to-be-evaluated document; A model evaluation module, configured to, for each of the to-be-evaluated sentences, input a corresponding first risk assessment prompt sentence into the first large language model for risk assessment to obtain multiple initial risk assessment results, where the first risk assessment prompt sentence is constructed from the to-be-evaluated sentence according to a first preset prompt word template; A screening module, configured to, for each to-be-evaluated sentence, if the initial risk assessment result is that there is a risk, screen out multiple first selected risk bases from multiple preset risk bases, where the semantic similarity between the risk sentence corresponding to the initial risk assessment result and the selected risk example sentences corresponding to each of the first selected risk bases meets the first preset screening condition; A sorting module, configured to, for each of the risk sentences, sort the first selected risk bases according to the correlation between each of the first selected risk bases and the risk basis of the risk sentence to determine the target risk basis; A result evaluation module, configured to obtain the risk assessment result of the to-be-evaluated document according to all the initial risk assessment results and / or all the target risk bases.
8. A file risk assessment device, characterized in that, The document risk assessment device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and is configured by the computer program to implement the steps of the document risk assessment method according to any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the document risk assessment method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the document risk assessment method according to any one of claims 1 to 6.