A legal risk early warning method and system based on neural network

Through a phased training method based on neural networks, the problem of low efficiency in traditional methods has been solved, and efficient and accurate text information extraction and legal risk warning from images or voices have been achieved, thereby improving the efficiency and accuracy of legal risk management.

CN119990766BActive Publication Date: 2025-09-16WUHAN PKU HIGH-TECH SOFT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510086579.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-09-16
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The traditional manual extraction and proofreading of legal and contractual documents is inefficient and error-prone. Existing technologies make it difficult to efficiently and accurately extract text information from images or voice and provide risk warnings.

Method used

A neural network-based method is used to train the text extraction model and risk prediction model in stages. By clarifying the image and voice information, the neural network model is used for training and optimization to construct the text extraction and risk prediction models.

Benefits of technology

It realizes the automatic and efficient extraction of text information from images or voices, and accurately predicts potential legal risks, thereby improving the efficiency and accuracy of legal risk management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990766B_ABST
    Figure CN119990766B_ABST
Patent Text Reader

Abstract

The present invention provides a legal risk warning method and system based on a neural network, which relates to the field of data processing technology, including obtaining multiple first information, as well as second information and third information corresponding to each first information; inputting a preset neural network model based on the first information and the corresponding second information for training to obtain a text extraction model; training the neural network model based on the second information and the corresponding third information to obtain a risk prediction model; obtaining target information, and inputting the target information into the prediction model to obtain target risk warning information, wherein the prediction model is composed of a text extraction model and a risk prediction model. By training the text extraction model and the risk prediction model in stages, the present invention can automatically and efficiently extract text information from images or voices, and predict potential legal risks based on the text information, thereby reducing manual intervention and improving the efficiency and accuracy of legal risk management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a legal risk early warning method and system based on a neural network. Background Art

[0002] With the advancement of information technology, businesses and organizations face the challenge of extracting massive amounts of data and information when processing legal and contract-related documents. Traditional manual extraction and proofreading methods are inefficient and error-prone. Therefore, efficiently and accurately extracting textual information from images or speech, and further predicting potential risks, has become a pressing technical challenge. Existing technologies often rely on rules or keyword matching, making them difficult to handle the variability and complexity of legal documents and ineffective in providing risk warnings and predictions. Summary of the Invention

[0003] The purpose of the present invention is to provide a legal risk early warning method and system based on a neural network to improve the above-mentioned problems. To achieve the above-mentioned purpose, the technical solution adopted by the present invention is as follows:

[0004] In a first aspect, the present application provides a legal risk early warning method based on a neural network, comprising:

[0005] Acquire multiple first information, as well as second information and third information corresponding to each first information, wherein the first information is image information or voice information carrying legal and contract-related information, the second information is text information obtained by extracting and manually correcting the first information, and the third information is risk warning information corresponding to the first information;

[0006] Inputting a preset neural network model based on the first information and the corresponding second information for training, and stopping the training when a preset first loss function satisfies a first set condition to obtain a text extraction model;

[0007] Training the neural network model based on the second information and the corresponding third information, and stopping the training when a preset second loss function satisfies a second set condition to obtain a risk prediction model;

[0008] Target information is acquired and input into a prediction model to obtain target risk warning information, wherein the prediction model is composed of the text extraction model and the risk prediction model.

[0009] Secondly, this application also provides a legal risk early warning system based on a neural network, including:

[0010] a first acquisition unit, configured to acquire a plurality of first information, and second information and third information corresponding to each first information, wherein the first information is image information or voice information carrying legal and contract-related information, the second information is text information obtained by extracting and manually correcting the first information, and the third information is risk warning information corresponding to the first information;

[0011] A first input unit is configured to input a preset neural network model for training based on the first information and the corresponding second information, and to stop training when a preset first loss function satisfies a first set condition to obtain a text extraction model;

[0012] a training unit, configured to train the neural network model based on the second information and the corresponding third information, and stop training when a preset second loss function satisfies a second set condition to obtain a risk prediction model;

[0013] The second acquisition unit is used to acquire target information and input the target information into a prediction model to obtain target risk warning information. The prediction model is composed of the text extraction model and the risk prediction model.

[0014] The beneficial effects of the present invention are:

[0015] By training the text extraction model and risk prediction model in stages, the present invention can automatically and efficiently extract text information from images or voices, and predict potential legal risks based on the text information, thereby reducing manual intervention and improving the efficiency and accuracy of legal risk management.

[0016] Other features and advantages of the present invention will be set forth in the following description, and in part will be apparent from the description, or may be learned by practicing embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 Schematic diagram of the process of the legal risk early warning method based on neural network in the present invention;

[0019] Figure 2 This is a schematic diagram of the structure of the legal risk early warning system based on neural network in the present invention;

[0020] Markings in the figure: 10, first acquisition unit; 20, first input unit; 30, training unit; 40, second acquisition unit. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0022] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.

[0023] Example 1:

[0024] This embodiment provides a legal risk early warning method based on neural network. Figure 1 , including step S10, step S20, step S30 and step S40.

[0025] Step S10. Acquire multiple first information, as well as second and third information corresponding to each first information, where the first information is image information or voice information carrying legal and contract-related information, the second information is text information obtained by extracting and manually correcting the first information, and the third information is risk warning information corresponding to the first information;

[0026] Step S20: inputting a preset neural network model based on the first information and the corresponding second information for training, and stopping the training when the preset first loss function satisfies a first set condition to obtain a text extraction model;

[0027] The transmission methods for contract and legal information are becoming increasingly diverse, especially with the increasing prevalence of image and audio formats. The information contained in these images and audio often needs to be converted into text for further analysis and processing. This process is not only time-consuming and labor-intensive, but also susceptible to noise, blur, or non-standardized formats, resulting in inaccurate extracted text information. Therefore, in this embodiment, the acquired image or audio information needs to be clarified before text extraction to ensure accurate text extraction.

[0028] Specifically, the step S20 specifically includes steps S21 to S26:

[0029] Step S21. Obtain interference information, where the interference information is interference image information or interference voice information;

[0030] The interfering image information is an image with blurred text due to shooting conditions, changes in ambient light or other factors; the interfering voice information is a voice with background noise, echo or unclear sound quality.

[0031] Step S22. Preprocess the first information and the interference information to obtain preprocessed first processed information and second processed information respectively;

[0032] Step S23: Inputting the first processed information and the second processed information into the trained clarity enhancement model to obtain enhanced processed information;

[0033] Considering that both image information and voice information may involve the existence of some interference information, it is necessary to clean and enhance the image information and voice information so that the corresponding text images and text voices are clearer, which facilitates the subsequent accurate extraction and recognition of text.

[0034] Step S24: Input the enhanced voice information into the neural network model to obtain predicted text information, which helps to accurately convert key information in the voice signal into text information and reduce the impact of interference on voice recognition.

[0035] Step S25. Calculating a first loss function based on the predicted text information and the corresponding text information. When the first loss function does not satisfy the first set condition, adjusting the weight threshold vector combination to obtain an adjusted neural network model, where the weight threshold vector combination includes initial weight vectors between neurons in each layer of the neural network model and an initial threshold matrix vector for each neuron.

[0036] Step S26. When the first loss function satisfies the first set condition, the corresponding neural network model is used as the text extraction model;

[0037] Therefore, the prediction ability of the neural network model can be continuously optimized by calculating the first loss function and adjusting the weights and thresholds according to the results. The specific first setting conditions can be set accordingly according to the model accuracy and are not particularly limited here.

[0038] Step S30: training the neural network model based on the second information and the corresponding third information, and stopping the training when the preset second loss function satisfies the second set condition to obtain a risk prediction model;

[0039] The constructed neural network model consists of a processing layer, a hidden layer, and an output layer. The processing layer consists of an input layer and an optimization layer. Its primary function is to convert input text into a data type that the neural network can process and generate the input data required by the hidden layer. Simultaneously, the optimization layer within the processing layer optimizes the initial weights and thresholds generated by the neural network model to obtain preliminary optimal weights and thresholds. The hidden layer processes the output data from the processing layer, using a multi-level abstraction approach to linearly partition the nonlinear input data. It then adjusts the input data's characteristics through weighted correction calculations to ultimately generate the input data required by the output layer. The output layer then performs weighted correction calculations on the output data from the hidden layer to produce the initial risk prediction results of the neural network model. This initial risk prediction result is then evaluated using a loss function, calculating the gradient of the loss values ​​at each layer and performing a backpropagation to optimize the neural network prediction model, ultimately achieving accurate risk prediction.

[0040] Step S40. Obtain target information and input the target information into the prediction model to obtain target risk warning information. The prediction model is composed of a text extraction model and a risk prediction model. Ultimately, the text extraction model and the risk prediction model are used to effectively identify and predict target risks.

[0041] Example 2:

[0042] The only difference between this embodiment and embodiment 1 is that step S23 inputs the first processed information and the second processed information into the trained sharp enhancement model to obtain enhanced processed information, which specifically includes the following steps:

[0043] Step S231: Map the first processed information into the frequency domain and extract the frequency domain features to obtain first feature information;

[0044] Step S232: Map the second processed information into the frequency domain and extract the frequency domain features to obtain second feature information;

[0045] Step S233. Process the first feature information and the second feature information based on a non-negative matrix factorization algorithm to obtain processed third feature information and fourth feature information respectively;

[0046] Step S234: Perform cleaning and strengthening processing on the first information based on the third characteristic information and the fourth characteristic information.

[0047] Therefore, this embodiment can effectively capture the key characteristics hidden in the data by mapping the first processed information and the second processed information to the frequency domain and extracting frequency domain features, and then use the non-negative matrix decomposition algorithm to process the extracted features, further separate and optimize the first feature information and the second feature information, and generate purer third feature information and fourth feature information, thereby greatly reducing the impact of interference on the data.

[0048] Further cleaning and strengthening processing significantly improves the clarity and quality of the first information, making it easier for subsequent models to accurately extract target text information in highly ambiguous and noisy environments, thereby improving the accuracy and robustness of subsequent prediction models.

[0049] In addition, step S25. calculating a first loss function based on the predicted text information and the corresponding text information, and when the first loss function does not meet the first set condition, adjusting the weight threshold vector combination to obtain an adjusted neural network model, includes the following steps:

[0050] Step S251. Extracting a plurality of target words from the predicted text information and the corresponding text information, wherein the target words are words whose frequency of appearance in both the predicted text information and the text information is greater than a set threshold;

[0051] Step S252: Randomly select a number of target words from the plurality of target words as reference words, and calculate the difference between the occurrence frequencies of the remaining target words in the predicted text information and the occurrence frequencies of the reference words, and the difference between the occurrence frequencies of the remaining target words in the text information and the occurrence frequencies of the reference words, to obtain a plurality of first differences and a plurality of second differences.

[0052] Step S253: Construct a first difference vector and a second difference vector based on the plurality of first differences and the plurality of second differences respectively;

[0053] In this embodiment, the words with the highest appearance frequency in the top 15 are used as the benchmark words, and the difference between the appearance frequency of the remaining target words in the predicted text information and the appearance frequency of the first benchmark word is used as the first difference vector:

[0054]

[0055] in, is the first difference vector corresponding to the first benchmark vocabulary; is the difference between the frequency of occurrence of the zth target word and the first benchmark word in the predicted text information; z is the number of target words in the predicted text information.

[0056] The difference between the frequency of occurrence of the remaining target words in the text information and the frequency of occurrence of the first benchmark word is used as the second difference vector:

[0057]

[0058] in, is the second difference vector corresponding to the first benchmark vocabulary; is the difference between the frequency of occurrence of the zth target word and the first benchmark word in the text information; z is the number of target words in the text information; thus, a total of fifteen first difference vectors and fifteen second difference vectors can be obtained.

[0059] Step S254: constructing the first loss function based on the cosine distance of the angle between the first difference vector and the second difference vector;

[0060] Specifically, the first loss function calculation formula includes:

[0061]

[0062] in, is the first difference vector corresponding to the j-th benchmark vocabulary; is the second difference vector corresponding to the jth benchmark word; m is the number of benchmark words, j∈m, and m is taken as 15 in this application.

[0063] Example 3:

[0064] The difference between this embodiment and embodiment 1 or 2 is that step S30: training the neural network model based on the second information and the corresponding third information, and stopping the training when the preset second loss function meets the second set condition, thereby obtaining a risk prediction model, including the following steps:

[0065] Step S31. Input multiple second information into the processing layer for classification and fusion to obtain multiple overall feature data;

[0066] Given the large amount of text in legal and contractual information, which can also contain a high level of noise and redundancy, these issues can affect data validity. If used directly in model training without processing, this can lead to overfitting—over-reliance on the details of the training data, reducing the ability to generalize to new information. Therefore, effective feature extraction from this information is necessary to reduce the adverse effects of noise and redundancy on model training, thereby improving model stability and performance.

[0067] Step S32: Input the plurality of overall feature data into the hidden layer, and obtain a plurality of feature vector data through multi-layer nonlinear transformation or the like.

[0068] Step S33. Input multiple feature vector data into the output layer to obtain multiple prediction risk warning information;

[0069] The feature vectors extracted through the hidden layer usually have strong recognition. After being input into the output layer, they can more accurately predict risk warning information and improve the accuracy and robustness of the prediction model.

[0070] And, step S34. calculating a second loss function based on the predicted risk warning information and the corresponding risk warning information, and when the second loss function does not meet the second set condition, adjusting the neural network model parameters until the first loss function meets the first set condition, thereby obtaining a risk prediction model;

[0071] Among them, risk warning information includes risk level, risk probability and risk category. Among them, the risk level indicates the severity of the risk, which is divided into low risk, medium risk and high risk; the risk probability indicates the possibility of predicting the occurrence of a risk event, and the probability value is usually between 0 and 1; the risk category classifies risks according to different standards, such as legal risk, financial risk, operational risk, etc.

[0072] Example 4:

[0073] The only difference between this embodiment and embodiment 3 is that step S31. Multiple second information input processing layers are classified and integrated to obtain multiple overall feature data, which specifically includes the following steps:

[0074] Step S311: Perform classification processing based on the content contained in the second information to obtain multiple target classification clusters and corresponding warning risk probabilities;

[0075] In this embodiment, the content can be divided into multiple categories based on the text in the second information. For example, for general contracts, it can be divided into multiple classification clusters such as main information, main terms, additional terms, and breach of contract risks based on the content. There are also differences in the categories divided for different contract contents, and no special restrictions are made here.

[0076] Step S312: Encode the content in each target classification cluster to obtain target feature data and multiple feature data, where the target feature data represents the content in the target classification cluster that is most likely to have a warning risk;

[0077] Step S313. Calculate the correlation degree between the target feature data and the multiple feature data in each target classification cluster based on the grey correlation analysis algorithm, and treat the target feature data and the multiple feature data as a node, and treat the correlation degree values ​​between the target feature data and the multiple feature data as edges between the nodes, where the length of the edge is related to the correlation degree value;

[0078] Step S314: updating the plurality of feature data based on the correlation degree value, the warning risk probability and the target feature data to obtain the updated plurality of feature data;

[0079] Step S315: Fusing the target feature data in each target classification cluster with the updated plurality of feature data to obtain a plurality of overall feature data;

[0080] Considering that the probability of early warning risks in different classification clusters is not the same, taking the above classification cluster as an example: the contract subject information usually includes the basic information of the contracting parties, such as company name, legal representative, contact information, etc. The above information usually does not contain significant risk factors, so the probability of early warning risks is relatively low; while the main terms in the contract (such as contract subject, price, delivery time, etc.) and breach of contract risks (such as payment delay, breach of contract liability, etc.) often hide potential early warning risks. Changes or execution issues of these terms may directly affect the performance of the contract or trigger legal disputes, so the probability of early warning risks is relatively high.

[0081] For the content in a classification cluster, the correlation degree values ​​between the target feature data and multiple feature data in each classification cluster are different. In this application, the target feature data represents the content in the target classification cluster that is most likely to have a warning risk. Feature data with a low correlation degree value with the target feature data usually indicates that this data is less important in risk warning or contributes less to the model's prediction. Therefore, only feature data with a correlation degree value greater than a certain value is updated and applied to subsequent model training, thereby reducing unnecessary data processing during model training and improving training efficiency.

[0082] Therefore, for different classification clusters and the feature data in each classification cluster, updates include:

[0083] t' iq =σ(P q *t iq *S niq *t nq )

[0084] Where t′ iq is the updated data of the i-th feature data in the q-th classification cluster; t iq is the i-th feature data in the q-th classification cluster; S niq is the correlation degree between the i-th feature data and the target feature data in the q-th classification cluster; t nq is the target feature data in the qth classification cluster; P q is the warning risk probability corresponding to q classification clusters; σ is the activation function.

[0085] The feature data in different classification clusters and each classification cluster are updated. The updated feature data can better emphasize similar features and highlight key information related to risk warning, thereby improving the model's ability to identify potential risks.

[0086] Example 5:

[0087] This embodiment differs from Embodiment 3 only in that, in step S34, a second loss function is calculated based on the predicted risk warning information and the corresponding risk warning information. When the second loss function does not satisfy the second set condition, the neural network model parameters are adjusted until the first loss function satisfies the first set condition, thereby obtaining a risk prediction model. The steps include:

[0088] Step S341. When the risk category in the predicted risk warning information is consistent with the risk category in the corresponding risk warning information, a second loss function is constructed based on the cross entropy loss function of the risk level and risk occurrence probability in the predicted risk warning information and the risk warning information;

[0089] When the risk categories of the predicted risk warning information and the corresponding actual risk warning information are consistent, an indicator for measuring the accuracy of the prediction can be constructed by simply quantifying the differences in risk levels and risk occurrence probabilities.

[0090] Step S342. When the risk categories in the predicted risk warning information and the corresponding risk warning information are inconsistent, calculate the ratio of the preset joint occurrence probability of the risk category to the risk occurrence probability in the predicted risk warning information and the risk warning information to obtain a first ratio and a second ratio, respectively;

[0091] In this embodiment, when the risk categories are inconsistent, the joint occurrence probability of the preset risk categories is calculated, and the ratio of the risk occurrence probability in the predicted value to the actual value is obtained. The two ratios respectively reflect the relative deviation between the two, providing a quantitative basis for subsequent loss calculation.

[0092] Step S343: Calculate a first loss based on the mean square error of the first ratio and the second ratio; for example, in this embodiment, the first loss may be calculated using the mean square error (MSE).

[0093] Step S344. Calculate a second loss based on the cross entropy loss function of the predicted risk warning information and the risk level in the risk warning information;

[0094] Therefore, considering the predicted and actual values ​​of the risk level, this embodiment calculates the second loss through the cross entropy loss function to help the model optimize in the risk level dimension.

[0095] Step S345. Calculate a third loss based on the cross entropy loss function of the predicted risk warning information and the risk occurrence probability in the risk warning information;

[0096] Specifically, for the probability of risk occurrence, the third loss is calculated through the cross entropy loss function to optimize the model's ability to predict the possibility of event occurrence.

[0097] Step S346. Construct a second loss function based on the first loss, the second loss, and the third loss to achieve all-round optimization of the model, weigh the deviations in different dimensions, and ensure that the model achieves high-precision prediction effects in key aspects such as risk category, risk level, and probability of occurrence.

[0098] Example 6:

[0099] like Figure 2 As shown, this embodiment provides a legal risk warning system based on a neural network, which can implement the legal risk warning method described in any one of Examples 1-5. The legal risk warning system includes:

[0100] A first acquisition unit 10 is configured to acquire a plurality of first information, as well as second information and third information corresponding to each first information, wherein the first information is image information or voice information carrying legal and contract-related information, the second information is text information obtained by extracting and manually correcting the first information, and the third information is risk warning information corresponding to the first information;

[0101] A first input unit 20 is used to input a preset neural network model for training based on the first information and the corresponding second information, and when a preset first loss function satisfies a first set condition, the training is stopped to obtain a text extraction model;

[0102] A training unit 30 is configured to train the neural network model based on the second information and the corresponding third information, and to stop training when a preset second loss function satisfies a second set condition to obtain a risk prediction model;

[0103] The second acquisition unit 40 is used to acquire target information and input the target information into the prediction model to obtain target risk warning information. The prediction model is composed of a text extraction model and a risk prediction model.

[0104] It should be noted that, regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.

[0105] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

[0106] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A legal risk early warning method based on neural network, characterized in that: include: Acquire multiple first information, as well as second information and third information corresponding to each first information, wherein the first information is image information or voice information carrying legal and contract-related information, the second information is text information obtained by extracting and manually correcting the first information, and the third information is risk warning information corresponding to the first information; Inputting a preset neural network model based on the first information and the corresponding second information for training, and stopping the training when a preset first loss function satisfies a first set condition to obtain a text extraction model; Training the neural network model based on the second information and the corresponding third information, and stopping the training when a preset second loss function satisfies a second set condition to obtain a risk prediction model; Acquire target information and input the target information into a prediction model to obtain target risk warning information, wherein the prediction model is composed of the text extraction model and the risk prediction model; The method includes inputting a preset neural network model based on the first information and the corresponding second information for training, and stopping the training when the preset first loss function satisfies a first set condition to obtain a text extraction model, including: Obtaining interference information, where the interference information is interference image information or interference voice information; Preprocessing the first information and the interference information to obtain preprocessed first processed information and second processed information, respectively; inputting the first processed information and the second processed information into a trained clarity enhancement model to obtain enhanced processed information; Inputting the enhanced speech information into the neural network model to obtain predicted text information; Calculating a first loss function based on the predicted text information and the corresponding text information, and when the first loss function does not meet a first set condition, adjusting a weight threshold vector combination to obtain an adjusted neural network model, wherein the weight threshold vector combination includes an initial weight vector between neurons in each layer of the neural network model and an initial threshold matrix vector for each neuron; When the first loss function satisfies a first set condition, the corresponding neural network model serves as the text extraction model; Calculating a first loss function based on the predicted text information and the corresponding text information includes: Extracting a plurality of target words from the predicted text information and the corresponding text information, wherein the target words are words whose frequencies of appearance in both the predicted text information and the text information are greater than a set threshold; Randomly selecting a number of target words from a plurality of target words as reference words, and calculating the difference between the appearance frequency of the remaining target words in the predicted text information and the appearance frequency of the reference words, and the difference between the appearance frequency of the remaining target words in the text information and the appearance frequency of the reference words, to obtain a plurality of first differences and a plurality of second differences; Constructing a first difference vector and a second difference vector based on the plurality of first differences and the plurality of second differences respectively; The first loss function is constructed based on the cosine distance of the angle between the first difference vector and the second difference vector.

2. The legal risk early warning method according to claim 1 is characterized in that The neural network model is trained based on the second information and the corresponding third information. When the preset second loss function meets the second set condition, the training is stopped to obtain a risk prediction model. The neural network model includes a processing layer, a hidden layer and an output layer, including: Inputting the plurality of the second information into the processing layer for classification and fusion to obtain a plurality of overall feature data; Inputting a plurality of overall feature data into the hidden layer to obtain a plurality of feature vector data; Inputting multiple feature vector data into the output layer to obtain multiple prediction risk warning information; A second loss function is calculated based on the predicted risk warning information and the corresponding risk warning information. When the second loss function does not meet the second setting condition, the neural network model parameters are adjusted until the first loss function meets the first setting condition, thereby obtaining the risk prediction model.

3. The legal risk early warning method according to claim 2 is characterized in that , input multiple pieces of the second information into the processing layer for classification and fusion, and obtain multiple overall feature data, including: Performing classification processing based on the content contained in the second information to obtain multiple target classification clusters and corresponding warning risk probabilities; Encoding the content in each target classification cluster to obtain target feature data and multiple feature data, wherein the target feature data represents the content in the target classification cluster that is most likely to have a warning risk; Calculate the correlation degree between the target feature data and multiple feature data in each target classification cluster based on the grey correlation analysis algorithm, and take the target feature data and multiple feature data as a node, and take the correlation degree values ​​between the target feature data and the multiple feature data as edges between the nodes, and the length of the edge is related to the correlation degree value; updating the plurality of feature data based on the correlation degree value, the warning risk probability and the target feature data to obtain the updated plurality of feature data; The target feature data in each target classification cluster and the updated multiple feature data are fused to obtain multiple overall feature data.

4. The legal risk early warning method according to claim 2 is characterized in that , calculating a second loss function based on the predicted risk warning information and the corresponding risk warning information, wherein the risk warning information includes the risk level, the risk occurrence probability and the risk category, including: When the risk categories in the predicted risk warning information and the corresponding risk warning information are consistent, constructing the second loss function based on the cross entropy loss function of the risk level and the risk occurrence probability in the predicted risk warning information and the risk warning information; When the risk categories in the predicted risk warning information and the corresponding risk warning information are inconsistent, calculating the ratio of the preset joint occurrence probability of the risk category and the risk occurrence probability in the predicted risk warning information and the risk warning information to obtain a first ratio and a second ratio respectively; Calculating a first loss based on a mean square error of the first ratio and the second ratio; Calculating a second loss based on a cross entropy loss function of the predicted risk warning information and the risk level in the risk warning information; Calculating a third loss based on a cross entropy loss function of the predicted risk warning information and the risk occurrence probability in the risk warning information; The second loss function is constructed based on the first loss, the second loss and the third loss.

5. The legal risk early warning method according to claim 1 is characterized in that Inputting the first processed information and the second processed information into the trained clarity enhancement model to obtain enhanced processed information includes: Mapping the first processed information to the frequency domain and extracting frequency domain features to obtain first feature information; Mapping the second processed information to the frequency domain and extracting frequency domain features to obtain second feature information; Processing the first feature information and the second feature information based on a non-negative matrix decomposition algorithm to obtain processed third feature information and fourth feature information respectively; The first information is subjected to a cleaning enhancement process based on the third characteristic information and the fourth characteristic information.

6. The neural network-based legal risk early warning method according to claim 3 is characterized in that , based on the correlation degree value, the warning risk probability and the target feature data, multiple feature data are updated to obtain multiple updated feature data, including: in, For the Updated feature data; For the Characteristic data; For the The correlation degree between the feature data and the target feature data; is the target feature data; is the activation function.

7. The legal risk early warning method according to claim 1 is characterized in that ,The first loss function calculation formula includes: in, is the first difference vector corresponding to the first benchmark vocabulary; is the second difference vector corresponding to the first benchmark vocabulary; is the baseline vocabulary size.

8. A legal risk early warning system based on neural network, characterized by: include: a first acquisition unit, configured to acquire a plurality of first information, and second information and third information corresponding to each first information, wherein the first information is image information or voice information carrying legal and contract-related information, the second information is text information obtained by extracting and manually correcting the first information, and the third information is risk warning information corresponding to the first information; A first input unit is configured to input a preset neural network model for training based on the first information and the corresponding second information, and to stop training when a preset first loss function satisfies a first set condition to obtain a text extraction model; a training unit, configured to train the neural network model based on the second information and the corresponding third information, and stop training when a preset second loss function satisfies a second set condition to obtain a risk prediction model; a second acquisition unit, configured to acquire target information and input the target information into a prediction model to obtain target risk warning information, wherein the prediction model is composed of the text extraction model and the risk prediction model; The first input unit further includes: Obtaining interference information, where the interference information is interference image information or interference voice information; Preprocessing the first information and the interference information to obtain preprocessed first processed information and second processed information, respectively; inputting the first processed information and the second processed information into a trained clarity enhancement model to obtain enhanced processed information; Inputting the enhanced speech information into the neural network model to obtain predicted text information; Calculating a first loss function based on the predicted text information and the corresponding text information, and when the first loss function does not meet a first set condition, adjusting a weight threshold vector combination to obtain an adjusted neural network model, wherein the weight threshold vector combination includes an initial weight vector between neurons in each layer of the neural network model and an initial threshold matrix vector for each neuron; When the first loss function satisfies a first set condition, the corresponding neural network model serves as the text extraction model; Calculating a first loss function based on the predicted text information and the corresponding text information includes: Extracting a plurality of target words from the predicted text information and the corresponding text information, wherein the target words are words whose frequencies of appearance in both the predicted text information and the text information are greater than a set threshold; Randomly selecting a number of target words from a plurality of target words as reference words, and calculating the difference between the appearance frequency of the remaining target words in the predicted text information and the appearance frequency of the reference words, and the difference between the appearance frequency of the remaining target words in the text information and the appearance frequency of the reference words, to obtain a plurality of first differences and a plurality of second differences; Constructing a first difference vector and a second difference vector based on the plurality of first differences and the plurality of second differences respectively; The first loss function is constructed based on the cosine distance of the angle between the first difference vector and the second difference vector.

Citation Information

Patent Citations

  • Contract term risk identification method and device

    CN110705265A

  • Prompting method, device and equipment for risk information in contract text

    CN112632989A