A contract risk identification method, device and equipment and storage medium

By segmenting and expanding contracts and combining them with a risk ranking model to identify contract risks, the problem of low accuracy in existing technologies has been solved, achieving efficient and accurate contract risk identification.

CN114943217BActive Publication Date: 2025-12-30SHENZHEN HIVE BOX NETWORK TECH LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210602862.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-30
Publication Date
2025-12-30
Estimated Expiration
2042-05-30

AI Technical Summary

Technical Problem

Existing technologies have low accuracy and limited scope in identifying contract risks, making it difficult to identify risks efficiently and accurately, which can easily lead to abnormal losses for one or both parties to the contract.

Method used

The process involves segmenting the contract to be processed to generate multiple contract sentences, receiving business risk information input by the user, expanding it according to a preset pattern, confirming contract sentences suspected of being risky based on the expanded information from multiple dimensions, and using a risk ranking model to calculate similarity and rank them to determine the sentence with the greatest risk.

Benefits of technology

It improves the accuracy and efficiency of contract risk identification, can cope with various contract risk scenarios, has good scalability and generalization ability, and maximizes the accuracy of risk identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114943217B_ABST
    Figure CN114943217B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a contract risk identification method, device, equipment and storage medium. The method comprises: performing segmentation processing on a to-be-processed contract to generate a plurality of contract sentences; receiving user input business attention risk information; performing preset mode expansion on the business attention risk information; based on the expanded business attention risk information, using a preset algorithm to confirm a suspected attention risk contract sentence from the plurality of contract sentences in multiple dimensions; using a risk ranking model to calculate the similarity between the business attention risk information and each suspected attention risk contract sentence; and according to the similarity, ranking the suspected attention risk contract sentences to determine a maximum risk contract sentence from the suspected attention risk contract sentences. The technical solution of the embodiments of the present application can cope with variable contract risk scenarios, has good scalability and generalization ability, and improves the risk identification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of computer software, and in particular to a method, apparatus, device and storage medium for identifying contract risks. Background Technology

[0002] A crucial aspect of contract risk identification for both parties is risk assessment. Contract risks often lead to abnormal losses for one or both parties, and may even result in fraud, causing even more severe losses. Current technologies for contract risk identification typically rely on keyword recognition or semantic vector matching. The former identifies risky keywords, but its accuracy is low and it's prone to missing key risks. The latter uses word vectors and word weights for automatic risk identification, but its accuracy is also low, failing to efficiently and accurately address the problem of contract risk identification. Summary of the Invention

[0003] This invention provides a method, apparatus, device, and storage medium for identifying contract risks, which can improve the efficiency and accuracy of contract risk identification.

[0004] According to one aspect of the present invention, an embodiment of the present invention provides a method for identifying contract risks, comprising:

[0005] The contract to be processed is split into multiple contract sentences;

[0006] Receive business risk information input by the user;

[0007] The preset mode is expanded to include the risk information that the business focuses on;

[0008] Based on the expanded business risk information, a preset algorithm is used to identify contract sentences suspected of being risky from multiple contract sentences from multiple dimensions.

[0009] The similarity between the business-related risk information and each contract sentence suspected of being a risk concern is calculated using a risk ranking model.

[0010] Contract sentences suspected of posing a risk are ranked according to the similarity score in order to identify the contract sentence with the greatest risk from among the suspected risk-posing contract sentences.

[0011] According to another aspect of the present invention, a contract risk identification device is provided, comprising:

[0012] The contract processing module is used to segment contracts to be processed to generate multiple contract sentences;

[0013] The risk expansion module is used to receive business-related risk information input by the user and to expand the business-related risk information in a preset mode.

[0014] The suspected risk confirmation module is used to identify contract sentences suspected of being of concern from multiple contract sentences from multiple dimensions based on the expanded business concern risk information and a preset algorithm.

[0015] The similarity calculation module is used to calculate the similarity between the business-related risk information and each contract sentence suspected of being a risk of concern using a risk ranking model;

[0016] The risk confirmation module is used to sort contract sentences suspected of posing risks based on the similarity, so as to identify the contract sentence with the greatest risk from the contract sentences suspected of posing risks.

[0017] According to another aspect of the present invention, a computer device is provided, the computer device comprising:

[0018] At least one processor; and

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform a contract risk identification method according to any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement a contract risk identification method according to any embodiment of the present invention.

[0022] The contract risk identification method provided in this invention involves: segmenting the contract to be processed to generate multiple contract sentences; receiving business-related risk information input by the user; expanding the business-related risk information using a preset pattern; identifying contract sentences suspected of having risks from multiple contract sentences using a preset algorithm based on the expanded business-related risk information; calculating the similarity between the business-related risk information and each contract sentence suspected of having risks using a risk ranking model; and ranking the contract sentences suspected of having risks according to the similarity to determine the contract sentence with the greatest risk from among the contract sentences suspected of having risks. This application embodiment, as an efficient contract risk identification and matching mechanism, solves the technical problems of low risk identification accuracy and a limited range of identified contract risks in existing technologies. It can cope with diverse contract risk scenarios, has good scalability and generalization ability, maximizes the accuracy of risk identification, and improves the efficiency of contract risk identification. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A flowchart illustrating a contract risk identification method provided in an embodiment of the present invention;

[0025] Figure 2 A flowchart illustrating another contract risk identification method provided in an embodiment of the present invention;

[0026] Figure 3 This is a schematic diagram of the structure of a contract risk identification device provided in an embodiment of the present invention;

[0027] Figure 4 This is a structural block diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] Example 1

[0030] Figure 1 This is a schematic diagram of a contract risk identification method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where contract risks are determined. The method can be executed by a contract risk identification device, which can be implemented by software and / or hardware, and is generally integrated into a computer device. Figure 1 As shown, the method includes the following steps:

[0031] Step S101: Segment the contract to be processed to generate multiple contract sentences.

[0032] In this embodiment of the invention, an input contract to be processed is obtained, and the contract content in the contract to be processed is segmented to generate multiple segmented contract sentences.

[0033] Optionally, the step of segmenting the contract to be processed to generate multiple contract sentences includes: segmenting the contract to be processed into sentences according to the sentence segmentation rules to generate a set of contract sentences; segmenting the contract sentences into words according to the word segmentation rules to generate contract words; and obtaining a dictionary to train word vectors to generate a word vector model.

[0034] The sentence segmentation rule can be any kind of sentence segmentation rule; for example, a segmentation rule using a special sentence delimiter can be used. The word segmentation rule can also be any kind of word segmentation rule; word vector training can be any word vector training method, for example, collecting data from open-source Chinese encyclopedias to train a word vector model. This embodiment of the invention does not limit this.

[0035] In this embodiment of the invention, the contract to be processed is segmented into a set of contract sentences according to the sentence segmentation rules. Then, the contract sentences are segmented into contract words according to the word segmentation rules to generate contract words for the contract sentences. A word vector model is trained using a dictionary.

[0036] Step S102: Receive business risk information input by the user; expand the business risk information using a preset mode.

[0037] Among them, the business-related risk information can be risk statements in the contract related to business, and the preset mode can be a pre-set risk expansion mode.

[0038] In this embodiment of the invention, business-related risk information input by a user is received, and a pre-set risk expansion mode is used to expand the business-related risk information. Optionally, the expansion of the business-related risk information using the pre-set mode includes: segmenting the business-related risk information into words according to word segmentation rules to generate risk words; calculating the keyword score of each risk word based on its word frequency and inverse document frequency; sorting the keyword scores of all risk words and extracting multiple risk keywords from the business-related risk information; inputting the multiple risk keywords into a word vector model to obtain the most similar synonyms corresponding to the risk keywords, and then replacing all risk keywords with synonyms to expand the business-related risk information.

[0039] In this embodiment of the invention, risk information of concern to the business is segmented into risk words. The word frequency and inverse document frequency of each risk word are statistically analyzed. The word frequency and inverse document frequency of each risk word are multiplied, and the product is used as the keyword score of the risk word. Then, the keyword scores of all risk words are sorted from top to bottom, and multiple risk keywords with keyword scores within a preset ranking are selected. The selected multiple risk keywords are input into a word vector model to obtain the most similar synonyms corresponding to the risk keywords. Then, all risk keywords are replaced with the corresponding synonyms to expand the risk information of concern to the business.

[0040] For example, by segmenting the risk information of business concern into words, calculating the word frequency and inverse document frequency of each risk word, and using the product of the word frequency and inverse document frequency of each risk word as the keyword score, the top n risk keywords with the highest scores are selected. The n risk keywords are then input into a word vector model to obtain the most similar synonyms for each risk keyword. Finally, all risk keywords are replaced with synonyms to expand the risk information of business concern to n.

[0041] Optionally, in another embodiment of the present invention, the expansion of the business-related risk information in a preset mode further includes: querying a historical business risk database based on the business-related risk information to determine the similarity between the business-related risk information and stored historical business risks; sorting the similarity between the business-related risk information and stored historical business risks to determine historical business risks similar to the business-related risk information; and expanding the business-related risk information based on the similar historical business risks.

[0042] The historical business risk database stores historical business risks, and can be either a non-relational database or a relational database. Historical business risks can be business risk information arising from contracts.

[0043] In this embodiment of the invention, after a user inputs business risk information, the user queries the historical business risk database based on the business risk information. The historical business risk database generates stored historical business risks based on the query results, and performs a similarity analysis between the business risk information and the stored historical business risks to generate a similarity score between the business risk information and each stored historical business risk. Then, based on the similarity scores between the business risk information and each stored historical business risk, the data is sorted from top to bottom, and historical business risks with similarity scores within a preset ranking are selected as similar historical business risks. These similar historical business risks are then used as business risk information to expand the business risk information database.

[0044] For example, the system inputs business risk information into the historical business risk database for querying, retrieves n historical business risks, calculates the similarity between each historical business risk and the business risk information, sorts the n historical business risks in descending order according to the similarity, selects the m historical business risks ranked mth as similar historical business risks, and expands the business risk information by m items.

[0045] Step S103: Based on the expanded business risk information, use a preset algorithm to identify contract sentences suspected of being risky from multiple contract sentences from multiple dimensions.

[0046] Optionally, the step of using a preset algorithm to identify contract sentences suspected of being of concern to risk from multiple contract sentences from multiple dimensions based on the expanded business concern risk information includes:

[0047] The word frequencies of contract segmentation and risk segmentation in multiple contract sentences are calculated separately to obtain the word frequencies of contract segmentation and risk segmentation. Then, the weights of contract segmentation and risk segmentation are determined based on the word frequencies of contract segmentation and risk segmentation, respectively.

[0048] The word frequency similarity between the contract to be processed and the business-related risk information is calculated based on the word frequency of contract segmentation, the word frequency of risk segmentation, the word weight of contract segmentation, and the word weight of risk segmentation.

[0049] The contract word segmentation and risk word segmentation are statistically analyzed separately to determine the contract word segmentation set and the risk word segmentation set. The similarity is calculated based on the contract word segmentation set and the risk word segmentation set to obtain the intersection and union similarity between the contract to be processed and the business-related risk.

[0050] In this embodiment of the invention, the word frequencies of contract segmentation and risk segmentation in multiple contract clauses are calculated to determine the word frequencies of contract segmentation and risk segmentation. Then, the weights of contract segmentation and risk segmentation in the contract clauses are determined. Based on the word frequencies and weights of contract segmentation and risk segmentation, the word frequency similarity between the contract to be processed and the business-related risk information is calculated. The word frequency IDF(qi) and the Score(Q,d) of the risk segmentation weight in the word frequency similarity are calculated as follows:

[0051]

[0052]

[0053] Where N is the total number of contract sentences, df i The number of sentences containing word segmentation; Q and d represent the two sentences whose similarity is to be calculated, respectively; k1 and b are adjustment factors, which can be set to k1 = 2, b = 0.75, f i For word segmentation in q i The frequency of occurrence in document d, where dl is the length of the current sentence. avg This represents the average sentence length.

[0054] Furthermore, all contract and risk word segments are statistically analyzed to determine the contract word segment set and risk word segment set. Then, the co-occurring words of risk and contract word segments are calculated, and the proportion of co-occurring words in all words is determined as the intersection-union similarity Jaccard(A,B) between the contract to be processed and the business-related risk. The calculation method for Jaccard(A,B) is as follows:

[0055]

[0056] Where A and B are the risk segmentation set and the contract segmentation set, respectively.

[0057] Optionally, the step of using a preset algorithm to identify contract sentences suspected of having risks from multiple contract sentences based on the expanded business risk information further includes: obtaining word vectors and word weights of the business risk information and contract sentences respectively according to the word vector model; obtaining the corresponding sentence vectors of the business risk information and contract sentences by weighting the word vectors and word weights of the business risk information and contract sentences; and then calculating the similarity between the sentence vectors of the business risk information and contract sentences to obtain the semantic similarity between the contract to be processed and the business risk information.

[0058] In this embodiment of the invention, word vectors and word weights of business-related risk information and contract sentences are obtained through a word vector model. The word vectors and word weights of the business-related risk information and contract sentences are then weighted to generate sentence vectors for the corresponding business-related risk information and contract sentences. Furthermore, the similarity between the sentence vectors of the business-related risk information and contract sentences is calculated to determine the semantic similarity between the contract to be processed and the business-related risk information. The semantic similarity cos(a,b) is calculated as follows:

[0059]

[0060] Where a and b are the sentence vectors of the business-related risk items and the contract sentences, respectively. The sentence vectors can be obtained by weighting word vectors and the IDF(qi) formula.

[0061] Optionally, the step of using a preset algorithm to identify contract sentences suspected of being of concern to risk from multiple contract sentences from multiple dimensions based on the expanded business concern risk information further includes:

[0062] Based on the word frequency similarity, intersection-union similarity, and semantic similarity between the contract to be processed and the business risk information, contract sentences suspected of being of risk concern are identified from multiple contract sentences from multiple dimensions.

[0063] Optionally, contract sentences suspected of posing a risk concern can be initially identified using word frequency similarity, intersection-union similarity, and semantic similarity of the risk information. Then, all initially identified contract sentences suspected of posing a risk concern are merged and deduplicated to confirm the contract sentences with confirmed risk concerns. Furthermore, to reduce computational complexity, the initially identified contract sentences suspected of posing a risk concern can be sorted, and the higher-ranking sentences can be retained.

[0064] For example, 30 contract sentences suspected of being risk-related are initially identified by word frequency similarity, intersection-union similarity, and semantic similarity of business risk information. Then, the 30 contract sentences suspected of being risk-related are sorted according to each similarity. Finally, the 90 initially identified contract sentences suspected of being risk-related are merged and deduplicated, and the 30 deduplicated contract sentences suspected of being risk-related are retained as the final identified contract sentences suspected of being risk-related.

[0065] Step S104: Use a risk ranking model to calculate the similarity between the business-related risk information and each contract sentence suspected of being a risk concern.

[0066] The risk ranking model is used to calculate the similarity between information on risks of concern to the business and each contract sentence suspected of representing a risk of concern. The risk ranking model may include, but is not limited to, at least one of bidirectional recurrent neural networks, convolutional neural network models, recurrent neural networks, and structural recurrent neural networks.

[0067] Specifically, before using the risk ranking model to calculate the similarity between the business-related risk information and each contract sentence suspected of representing a risk of concern, the process also includes:

[0068] Obtain sample business risk information, and establish a training set and a test set of sample business risk information based on this information. Obtain sample contract sentences that may indicate potential risks, and establish a training set and a test set of sample contract sentences that may indicate potential risks based on these sample contract sentences.

[0069] The risk ranking model is generated by inputting sample business risk information and suspected risk-related contract sentences from the business risk training set into a neural network model. Specifically, the model generates training similarities by inputting these samples into the neural network model. The training error is determined based on the generated training similarities and the expected similarities between the sample business risk information and the suspected risk-related contract sentences. The parameters of the neural network model are then adjusted based on this training error to obtain the risk ranking model.

[0070] Furthermore, the similarity between business-related risk information and each contract sentence suspected of being related to a risk is calculated using a risk ranking model.

[0071] Step S105: Sort the contract sentences suspected of posing risks based on the similarity to determine the contract sentence with the greatest risk from among the contract sentences suspected of posing risks.

[0072] The contract risk identification method provided in this invention involves: segmenting the contract to be processed to generate multiple contract sentences; receiving business-related risk information input by the user; expanding the business-related risk information using a preset pattern; identifying contract sentences suspected of having risks from multiple contract sentences using a preset algorithm based on the expanded business-related risk information; calculating the similarity between the business-related risk information and each contract sentence suspected of having risks using a risk ranking model; and ranking the contract sentences suspected of having risks according to the similarity to determine the contract sentence with the greatest risk from among the contract sentences suspected of having risks. This embodiment of the application, as an efficient contract risk identification and matching mechanism, solves the technical problems of low risk identification accuracy and a limited range of identified contract risks in existing technologies. It can cope with diverse contract risk scenarios, has good scalability and generalization ability, maximizes the accuracy of risk identification, and improves the efficiency of contract risk identification.

[0073] Example 2

[0074] Figure 2 This is a flowchart illustrating another contract risk identification method provided by an embodiment of the present invention. Based on the aforementioned optional embodiments, this embodiment optimizes and provides a specific optional implementation method for calculating the similarity between the business-related risk information and each contract sentence suspected of representing a risk of concern using a risk ranking model. Specifically, as shown... Figure 2 As shown, the method includes the following steps:

[0075] Step S201: Segment the contract to be processed to generate multiple contract sentences.

[0076] Step S202: Receive business risk information input by the user.

[0077] Step S203: Expand the preset mode of the business-related risk information.

[0078] Step S204: Based on the expanded business risk information, use a preset algorithm to identify contract sentences suspected of being risky from multiple contract sentences from multiple dimensions.

[0079] Step S205: Input the business risk information and the contract sentences that are suspected of being risk-related into the risk ranking model to calculate the matching degree between the two, and obtain the similarity between the business risk information and the contract sentences that are suspected of being risk-related.

[0080] Optionally, a risk ranking model can be used to calculate the similarity between business-related risk information and contract sentences that may indicate risk concerns, and then the similarity scores can be used to rank the risk levels. Specifically, business-related risk information is used as input 'a' to the risk ranking model, and contract sentences that may indicate risk concerns are used as input 'b'.

[0081] For example, the business-related risk information could be "the site needs to be restored to its original state upon withdrawal," and the risk level could be ranked as follows: 1. "When the client withdraws, the site and facilities shall not be damaged. If the site or facilities are damaged, the client shall be responsible for restoring the site to its original state," 0.998 (similarity score).

[0082] 2. "If the parties fail to reach an agreement on continuing the contract after the expiration of the contract, Party A shall remove the items in the leased premises and restore them to their original condition within 7 days from the date of expiration of this contract", 0.985 (similarity score);

[0083] Optionally, the steps for calculating the similarity between contract sentences related to business-concerned risks and those suspected of being risk-concerned, using the risk ranking model, are as follows:

[0084] (1) The risk ranking model first uses a word embedding layer to vectorize contract sentences that represent business-related risk information and suspected risk information;

[0085] (2) The two obtained vectors are bidirectionally encoded using a bidirectional recurrent neural network. First, the forget gate weight W is used. f and deviation b f For the current state x t and historical state h t-1 After calculation, through The probability of forgetting historical information is obtained by compressing it to 0-1, i.e., formula (1); similarly, the input probability of the information at the current moment is obtained through formula (2); the output probability of the information at the current moment is obtained through formula (5); and the candidate memory cell weight W at the current moment is used to obtain the probability of forgetting the information at the current moment. c and deviation b c For the current state x t and historical state h t-1 After calculation, through Obtain the current state of candidate memory cells That is, formula (3); through the forgetting probability f t and input probability i t Based on the current state of candidate memory cells and the state of historical memory cells Get the current state of the memory cell c t That is, formula (4); finally, output probability o tThe tanh function stores the cell state c at the current time step. t After calculation, the historical state h at the current time is output. t That is, formula (6); where These are the historical states h at each time step output after being encoded by a bidirectional recurrent neural network. t That is, formulas (8) and (9);

[0086] f t =sigmoid(W f *[h t-1 x t ]+b f (1)

[0087] i t =sigmoid(W i *[h t-1 x t ]+b i (2)

[0088]

[0089]

[0090] o t =sigmoid(W o *[h t-1 x t ]+b o (5)

[0091] h t =o t *tan h(C t (6)

[0092]

[0093]

[0094]

[0095] (3) Then, the deep interaction between the two vectors is realized through the self-attention layer mechanism, and the input vector is calculated by formula (10). At time i, relative to the vector The importance of time j is e ij ;pass Calculated Throughout attention score That is, formula (11). Similarly, from formula (12), we can obtain... Throughout Attention score during sequence encoding This process introduces a representation of each party's contribution to the other;

[0096]

[0097]

[0098]

[0099] (4) Finally, the results of the interaction are concatenated to obtain m. a and m b That is, formulas (13) and (14); after being encoded again through a bidirectional recurrent neural network, we obtain and That is, formulas (15) and (16); after dimensionality reduction through max pooling and average pooling, p is obtained. a and p b That is, formulas (17) and (18); finally, through the weight W of the fully connected layer ffn and bias b ffn The similarity between the two is calculated, i.e., formula (19);

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106] similar = softmax(W ffn *[p a p b ]+b ffn (19)

[0107] (5) Based on the similarity between business-related risk information and suspected contract risk sentences, sort them to obtain the contract sentence with the highest risk, which is the corresponding contract risk.

[0108] Step S206: Sort the contract sentences suspected of posing risks based on the similarity to determine the contract sentence with the greatest risk from among the contract sentences suspected of posing risks.

[0109] The contract risk identification method provided in this invention involves: segmenting the contract to be processed to generate multiple contract sentences; receiving business-related risk information input by the user; expanding the business-related risk information using a preset pattern; identifying contract sentences suspected of having risks from multiple contract sentences using a preset algorithm based on the expanded business-related risk information; calculating the similarity between the business-related risk information and each contract sentence suspected of having risks using a risk ranking model; and ranking the contract sentences suspected of having risks based on the similarity to determine the contract sentence with the greatest risk from among the contract sentences suspected of having risks. This application, as an efficient contract risk identification and matching mechanism, solves the technical problems of low risk identification accuracy and a limited range of identified contract risks in existing technologies. It can cope with diverse contract risk scenarios, has good scalability and generalization ability, maximizes the accuracy of risk identification, and improves the efficiency of contract risk identification while having low requirements for new training data.

[0110] Example 3

[0111] Figure 3 This is a schematic diagram of a contract risk identification device provided in Embodiment 4 of the present invention, as shown below. Figure 3 As shown, the device includes: a contract processing module 310, a risk expansion module 320, a suspected risk confirmation module 330, a similarity calculation module 340, and a risk confirmation module 350, wherein:

[0112] Contract processing module 310 is used to segment the contract to be processed to generate multiple contract sentences;

[0113] The risk expansion module 320 is used to receive business-related risk information input by the user and to expand the business-related risk information in a preset mode.

[0114] The suspected risk confirmation module 330 is used to confirm suspected risk-related contract sentences from multiple contract sentences from multiple dimensions based on the expanded business-concerned risk information and a preset algorithm.

[0115] The similarity calculation module 340 is used to calculate the similarity between the business-related risk information and each contract sentence suspected of being a risk of concern using a risk ranking model;

[0116] The risk confirmation module 350 is used to sort contract sentences suspected of posing risks based on the similarity, so as to determine the contract sentence with the greatest risk from the contract sentences suspected of posing risks.

[0117] The contract risk identification device provided in this invention segmentes the contract to be processed to generate multiple contract sentences; receives business-related risk information input by the user; expands the business-related risk information using a preset pattern; based on the expanded business-related risk information, uses a preset algorithm to identify contract sentences suspected of having risks from multiple dimensions among the multiple contract sentences; calculates the similarity between the business-related risk information and each contract sentence suspected of having risks using a risk ranking model; and ranks the contract sentences suspected of having risks according to the similarity to determine the contract sentence with the greatest risk from among the contract sentences suspected of having risks. This embodiment of the application, as an efficient contract risk identification and matching mechanism, solves the technical problems of low risk identification accuracy and a limited range of identified contract risks in existing technologies. It can cope with diverse contract risk scenarios, has good scalability and generalization ability, maximizes the accuracy of risk identification, and improves the efficiency of contract risk identification.

[0118] Optionally, the contract processing module 310 is specifically used to: segment the contract to be processed into sentences according to the sentence segmentation rules, and generate a set of contract sentences;

[0119] The contract clauses are segmented according to the word segmentation rules to generate contract word segments;

[0120] Obtain the dictionary to train word vectors and generate a word vector model.

[0121] Optionally, the risk expansion module 320 is specifically used to: segment the business-related risk information according to the word segmentation rules to generate risk words;

[0122] The keyword score of each risk word is calculated based on its term frequency and inverse document frequency; the keyword scores of all risk words are sorted, and multiple risk keywords related to the business-related risk information are extracted.

[0123] The risk keywords related to the business's risk information are input into a word vector model to obtain the most similar synonyms corresponding to the risk keywords. Then, all risk keywords are replaced with synonyms to expand the risk information related to the business's risk information.

[0124] Optionally, the risk expansion module 320 is also specifically used for:

[0125] Based on the business risk information, query the historical business risk database to determine the similarity between the business risk information and the stored historical business risks;

[0126] The similarity between the business-related risk information and the stored historical business risks is sorted to determine the historical business risks that are similar to the business-related risk information.

[0127] The business risk information is expanded based on similar historical business risks.

[0128] Optionally, the suspected risk confirmation module 330 is specifically used to: calculate the word frequency of contract segmentation and risk segmentation in multiple contract sentences respectively, obtain the word frequency of contract segmentation and risk segmentation, and then determine the contract segmentation weight and risk segmentation weight respectively based on the word frequency of contract segmentation and risk segmentation.

[0129] The word frequency similarity between the contract to be processed and the business-related risk information is calculated based on the word frequency of contract segmentation, the word frequency of risk segmentation, the word weight of contract segmentation, and the word weight of risk segmentation.

[0130] The contract word segmentation and risk word segmentation are statistically analyzed separately to determine the contract word segmentation set and the risk word segmentation set. The similarity is calculated based on the contract word segmentation set and the risk word segmentation set to obtain the intersection and union similarity between the contract to be processed and the business-related risk.

[0131] Furthermore, the suspected risk confirmation module 330 is specifically used to: obtain the word vectors and word weights of the business-related risk information and the contract sentence according to the word vector model;

[0132] By weighting the word vectors and word weights of the business-related risk information and contract sentences, the corresponding sentence vectors of the business-related risk information and contract sentences are obtained. Then, similarity is calculated based on the sentence vectors of the business-related risk information and contract sentences to obtain the semantic similarity between the contract to be processed and the business-related risk information.

[0133] Furthermore, the suspected risk confirmation module 330 is specifically used to: confirm suspected risk-related contract sentences from multiple contract sentences from multiple dimensions based on the word frequency similarity, intersection-union similarity, and semantic similarity between the contract to be processed and the business-related risk information.

[0134] Optionally, the similarity calculation module 340 is specifically used for:

[0135] The risk information concerning business concerns and the contract sentences that are suspected of being risk-concerned are respectively input into the risk ranking model to calculate the degree of matching between the two, thereby obtaining the similarity between the risk information concerning business concerns and the contract sentences that are suspected of being risk-concerned.

[0136] Since the contract risk identification device described above is an apparatus capable of executing the contract risk identification method in the embodiments of the present invention, those skilled in the art can understand the specific implementation methods and various variations of the contract risk identification device in this embodiment based on the contract risk identification method described in the embodiments of the present invention. Therefore, how the contract risk identification device implements the contract risk identification method in the embodiments of the present invention will not be described in detail here. Any apparatus used by those skilled in the art to implement the contract risk identification method in the embodiments of the present invention falls within the scope of protection of this application.

[0137] Example 4

[0138] Figure 4 A schematic diagram of a computer device 10 that can be used to implement embodiments of the present invention is shown. The computer device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The computer device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0139] like Figure 4 As shown, the computer device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer programs stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the computer device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0140] Multiple components in computer device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows computer device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0141] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as contract risk identification methods.

[0142] In some embodiments, the contract risk identification method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on computer device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the contract risk identification method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the contract risk identification method by any other suitable means (e.g., by means of firmware).

[0143] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0144] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0145] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0146] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0147] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0148] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0149] Example 5

[0150] Embodiment 5 of the present invention also provides a computer storage medium for storing a computer program. When executed by a computer processor, the computer program is used to perform the contract risk identification method described in any of the above embodiments of the present invention: segmenting the contract to be processed to generate multiple contract sentences; receiving business concern risk information input by a user; expanding the business concern risk information according to a preset pattern; based on the expanded business concern risk information, using a preset algorithm to confirm contract sentences suspected of having concerns about risks from multiple dimensions among the multiple contract sentences; using a risk ranking model to calculate the similarity between the business concern risk information and each contract sentence suspected of having concerns about risks; and ranking the contract sentences suspected of having concerns about risks according to the similarity to determine the contract sentence with the greatest risk from the contract sentences suspected of having concerns about risks.

[0151] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0152] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0153] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, radio frequency (RF), or any suitable combination thereof.

[0154] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0155] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0156] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A contract risk identification method, characterized by, The method comprises the following steps: segmenting a to-be-processed contract to generate a plurality of contract sentences; receiving user input business concern risk information; wherein the business concern risk information is a risk sentence associated with the business in the contract; performing preset mode expansion on the business concern risk information; based on the expanded business concern risk information, using a preset algorithm to confirm contract sentences with suspected concern risks from the plurality of contract sentences in multiple dimensions; using a risk ranking model to calculate the similarity between the business concern risk information and each contract sentence with suspected concern risks; ranking the contract sentences with suspected concern risks according to the similarity to determine the contract sentence with the greatest risk from the contract sentences with suspected concern risks; using a risk ranking model to calculate the similarity between the business concern risk information and each contract sentence with suspected concern risks, comprising: respectively inputting the business concern risk information and the contract sentence with suspected concern risks into the risk ranking model to calculate the matching degree of the two, thereby obtaining the similarity between the business concern risk information and the contract sentence with suspected concern risks; the respective inputting of the business concern risk information and the contract sentence with suspected concern risks into the risk ranking model to calculate the matching degree of the two, thereby obtaining the similarity between the business concern risk information and the contract sentence with suspected concern risks, comprising: the risk ranking model vectorizes the business concern risk information and the contract sentence with suspected concern risks through a word embedding layer to obtain a vector corresponding to the business concern risk information and a vector corresponding to the contract sentence with suspected concern risks; The vector corresponding to the service attention risk information obtained and the vector corresponding to the contract sentence suspected of the attention risk are respectively subjected to bidirectional sequence coding through a bidirectional recurrent neural network to obtain and ; wherein, , are historical states of each time step output after bidirectional recurrent neural network coding ; The output of the bidirectional recurrent neural network after bidirectional sequence encoding is deep-interacted through a self-attention layer mechanism; and two semantic vectors are obtained respectively 、 ; wherein is the semantic vector calculated by the self-attention layer mechanism of the risk ranking model; is the semantic vector calculated by the self-attention layer mechanism of the risk ranking model; the semantic vector calculated by the self-attention layer mechanism of the risk ranking model;​ The results of the deep interaction are spliced to obtain and ; wherein ; ; Will And Through the bidirectional recurrent neural network coding, get And ; wherein, Through the bidirectional recurrent neural network coding result; Through the bidirectional recurrent neural network coding result;​​ The coded results are dimensionally reduced by max-pooling and average-pooling to obtain and ; wherein, ; ; Will and The similarity of the contract sentences of the business attention risk information and the suspected attention risk is calculated by the full connection layer weight and bias.

2. The method of claim 1, wherein, the segmentation processing of the to-be-processed contract to generate a plurality of contract sentences, comprising: performing sentence segmentation on the to-be-processed contract according to a sentence segmentation rule to generate a contract sentence set; performing word segmentation on the contract sentence according to a word segmentation rule to generate a contract word segmentation; obtaining a dictionary to perform word vector training to generate a word vector model.

3. The method of claim 1, wherein, the preset mode expansion of the business concern risk information, comprising: performing word segmentation on the business concern risk information according to a word segmentation rule to generate a risk word segmentation; calculating a keyword score of each risk word in the risk word segmentation according to the word frequency and the inverse document frequency of each risk word; sorting the keyword scores of all risk words to extract a plurality of risk keywords of the business concern risk information; inputting the plurality of risk keywords of the business concern risk information into the word vector model to obtain the most similar synonyms corresponding to the risk keywords, and then replacing all risk keywords with synonyms to expand the business concern risk information.

4. The method of claim 1, wherein, the preset mode expansion of the business concern risk information, further comprising: querying a historical business risk database according to the business concern risk information to determine the similarity between the business concern risk information and the stored historical business risks; sorting the similarity between the business concern risk information and the stored historical business risks to determine the historical business risks similar to the business concern risk information; expanding the business concern risk information according to the similar historical business risks.

5. The method of claim 2, wherein, The contract sentence suspected of the risk of attention is confirmed from the plurality of contract sentences based on the expanded business attention risk information in multiple dimensions according to a preset algorithm, including: The word frequencies of the contract word segmentation and the risk word segmentation in the plurality of contract clauses are calculated respectively to obtain the contract word segmentation frequency and the risk word segmentation frequency, and then the contract word segmentation weight and the risk word segmentation weight are determined according to the contract word segmentation frequency and the risk word segmentation frequency respectively; The word frequency similarity between the to-be-processed contract and the business attention risk information is calculated according to the contract word segmentation frequency, the risk word segmentation frequency, the contract word segmentation weight and the risk word segmentation weight. The contract word segmentation set and the risk word segmentation set are respectively determined by the contract word segmentation and the risk word segmentation, and the intersection-union set similarity between the to-be-processed contract and the business attention risk information is calculated according to the contract word segmentation set and the risk word segmentation set.

6. The method of claim 2, wherein, The contract sentence suspected of the risk of attention is confirmed from the plurality of contract sentences based on the expanded business attention risk information in multiple dimensions according to a preset algorithm, including: The word vectors and word weights of the business attention risk information and the contract sentence are obtained according to the word vector model respectively; The sentence vectors of the business attention risk information and the contract sentence are obtained by a weighted manner according to the word vectors and word weights of the business attention risk information and the contract sentence, and then the semantic similarity between the to-be-processed contract and the business attention risk information is calculated according to the sentence vectors of the business attention risk information and the contract sentence.

7. The method of claim 6, wherein, The contract sentence suspected of the risk of attention is confirmed from the plurality of contract sentences based on the expanded business attention risk information in multiple dimensions according to a preset algorithm, including: The contract sentence suspected of the risk of attention is confirmed from the plurality of contract sentences based on the expanded business attention risk information in multiple dimensions according to a preset algorithm, including:

8. A contract risk identification apparatus characterized by comprising: Including: The contract processing module is used for segmenting the to-be-processed contract to generate a plurality of contract sentences; The risk expansion module is used for receiving the business attention risk information input by a user; The business attention risk information is expanded in a preset mode; wherein the business attention risk information is a risk sentence related to the business in the contract; The suspected risk confirmation module is used for confirming the contract sentence suspected of the risk of attention from the plurality of contract sentences based on the expanded business attention risk information in multiple dimensions according to a preset algorithm; The similarity calculation module is used for calculating the similarity between the business attention risk information and each contract sentence suspected of the risk of attention by using a risk ranking model; The risk confirmation module is used for ranking the contract sentence suspected of the risk of attention according to the similarity, so as to determine the contract sentence with the maximum risk from the contract sentence suspected of the risk of attention. The similarity calculation module is specifically used for: The risk ranking model obtains the vector corresponding to the business attention risk information and the vector corresponding to the contract sentence suspected of the risk of attention by vectorizing the business attention risk information and the contract sentence suspected of the risk of attention through a word embedding layer. The vector corresponding to the service attention risk information obtained and the vector corresponding to the contract sentence suspected of the attention risk are respectively subjected to bidirectional sequence coding through a bidirectional recurrent neural network to obtain and ; wherein, , are historical states of each time step output after bidirectional recurrent neural network coding ; The output of the bidirectional recurrent neural network after bidirectional sequence encoding is deep-interacted through a self-attention layer mechanism; and two semantic vectors are obtained respectively 、 ; wherein is a semantic vector obtained through a self-attention layer mechanism of the risk ranking model , is a semantic vector obtained through a self-attention layer mechanism of the risk ranking model , is a semantic vector obtained through a self-attention layer mechanism of the risk ranking model . The results of the deep interaction are spliced to obtain and ; wherein ; ; Will And Through the bidirectional recurrent neural network coding, get And ; wherein, Is Through the bidirectional recurrent neural network coding result; Is Through the bidirectional recurrent neural network coding result; The coded results are dimensionally reduced by max-pooling and average-pooling to obtain and ; wherein, ; ; Will and The similarity of the contract sentences of the business attention risk information and the suspected attention risk is calculated by the full connection layer weight and bias.

9. A computer device, comprising: The computer device includes: At least one processor; and The memory is connected in communication with the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the contract risk identification method in any one of claims 1-7.

10. A computer storage medium, characterized in that The computer storage medium stores computer instructions for enabling the processor to implement the contract risk identification method in any one of claims 1-7 when executed.

Citation Information

Patent Citations

  • Computer-executed text risk prediction method and apparatus

    CN109299228A

  • Contract term risk examination method and device

    CN110163478A

  • Intention classification method and device based on voting decision, equipment and storage medium

    CN111444722A