Training method for a model for screening judgment documents of controversial focus
By adopting the positive unlabeling strategy and support vector machine model in the screening of dispute focus judgment documents, the problem of poor identification results in the lack of labeled data in the prior art is solved, and a more accurate and efficient screening effect is achieved.
Patent Information
- Application Number
- CN202111632891.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-28
AI Technical Summary
The prior art lacks sufficient labeling data in the screening of dispute focus judgment documents, resulting in poor model identification results, making it difficult to effectively screen out documents containing dispute focus.
The positive unlabeling strategy is adopted to construct the dispute focus characteristics at the document level from a small number of labeled documents, and the support vector machine model is used to identify the judgment documents with dispute focus from a large number of labeled documents.
Through the improved screening scheme, judicial documents containing dispute focus can be more accurately identified, improving the efficiency and accuracy of screening.
Smart Images

Figure CN114428847B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of text processing. More specifically, it relates to a training method for a model for screening judgment documents of disputed focuses, a screening method for judgment documents of disputed focuses, a device, and an electronic device. Background Art
[0002] With the continuous deepening of the informatization construction in the legal industry, not only the collection and governance of massive data are required, but also the automatic processing of legal documents by computers. In the processing of legal documents, the disputed focus is the key information summarized by legal professionals based on the statements of both parties in the complaint and the answer. The disputed focus can reflect the divergence points of the case and the issues it focuses on. The induction and summary of the disputed focus usually require strong professional knowledge and domain experience, and also consume a lot of energy. Moreover, since not all judgment documents contain disputed focuses, when studying the issue of disputed focuses, it is first necessary to screen out the judgment documents containing disputed focuses from a large number of documents.
[0003] In Chinese Patent CN107291688A, a document screening method for legal judgment documents is proposed. The invention discloses a method for analyzing the similarity of judgment documents based on a topic model. This method uses the LDA (Latent Dirichlet Allocation) topic model in machine learning and proposes a semantic-based, semi-automatic, and general similarity analysis method for judgment documents. Based on the general similarity analysis method, this method fully considers the characteristics of rich professional vocabulary and complex semantics in the content of judgment documents and utilizes the semi-structured characteristics of judgment documents, thereby improving the accuracy and applicability of the similarity analysis of judgment documents.
[0004] Chinese Patent CN113282726A discloses a data processing method, system, device, medium, and data analysis method, which relates to the field of natural language processing. It includes: screening out judgment document data containing a preset type of crime from a judgment document library; extracting case fact data from the information data of the judgment document; establishing a graph neural network and inputting the case fact data into the graph neural network; extracting entity information data, relationship data between entities, and event sequence information data at the node positions of the graph neural network and converting them into result data for output. The result data includes (entity 1, relationship, entity 2) triples and the crime type to which the judgment document belongs. The analysis and statistical results obtained by this invention are based on the full amount of judgment document data of the preset crime type in the judgment document library. Through a deep learning method for heuristic information extraction, the statistical dimensions are more comprehensive and less susceptible to human interference.
[0005] The screening of adjudication documents with controversial focuses is essentially a text classification problem, that is, to determine whether an adjudication document has controversial focuses through a model. Currently, the text classification of adjudication documents is mainly identified through neural network models in natural language processing technology. This method usually requires a large amount of labeled data, and the effect of model recognition depends on the quantity and quality of the labeled data. However, in the problem of screening adjudication documents with controversial focuses, due to the fact that the annotation of controversial focuses requires a great deal of effort, only a small number of labeled positive samples and a large number of unlabeled samples are available.
[0006] Therefore, it is desired to provide an improved screening scheme for adjudication documents with controversial focuses. Summary of the Invention
[0007] To solve the above technical problems, the present application is proposed. Embodiments of the present application provide a training method for a model for screening adjudication documents with controversial focuses, a screening method for adjudication documents with controversial focuses, a device, and an electronic device. Based on the positive-unlabeled strategy, it constructs document-level controversial focus features from a small number of existing labeled documents, and uses a support vector machine model to identify adjudication documents with controversial focuses from a large number of unlabeled documents, thereby improving the screening scheme for adjudication documents with controversial focuses.
[0008] According to one aspect of the present application, there is provided a training method for a model for screening adjudication documents with controversial focuses, including: obtaining training adjudication documents, where the training adjudication documents include a small part of labeled samples and a large part of unlabeled samples; performing text preprocessing and paragraph extraction on the training adjudication documents; constructing a first feature vector of the labeled samples and a second feature vector of the unlabeled samples from the training adjudication documents; and using the positive-unlabeled learning strategy to train a support vector machine model with the labeled samples as positive samples and the unlabeled samples as negative samples, where the support vector machine model is used to classify adjudication documents as whether they contain controversial focuses.
[0009] In the above training method for a model for screening adjudication documents with controversial focuses, the text preprocessing includes digital unified conversion and / or punctuation mark unification.
[0010] In the above training method for a model for screening adjudication documents with controversial focuses, the paragraph extraction includes extracting corresponding adjudication analysis process segments and defense segments from the training adjudication documents through a document segmentation model.
[0011] In the above training method of the model for screening adjudication documents of controversial issues, constructing the first feature vector of the labeled samples and the second feature vector of the unlabeled samples from the training adjudication documents includes: determining document paragraphs with a relevance greater than a predetermined threshold to the controversial issues based on the paragraph type and paragraph length of whether the adjudication document contains controversial issues; and constructing the first feature vector of the labeled samples and the second feature vector of the unlabeled samples based on the keywords obtained by analyzing and counting the sentences of the controversial issues and the document paragraphs.
[0012] In the above training method of the model for screening adjudication documents of controversial issues, the first feature vector and the second feature vector include the following features and their weights: the number of paragraphs of the adjudication document, the length of the adjudication analysis process paragraph, the length of the defense paragraph, the number of reasoning feature words in the adjudication analysis process paragraph, and the number of reasoning feature words in the defense paragraph
[0013] In the above training method of the model for screening adjudication documents of controversial issues, using the positive-unlabeled learning strategy to train a support vector machine model with the labeled samples as positive samples and the unlabeled samples as negative samples includes: Step 1, taking the first feature vector as the positive feature vector and the second feature vector as the negative feature vector; Step 2, determining the separating hyperplane with the largest geometric margin of the support vector machine model based on the positive feature vector and the negative feature vector; Step 3, determining candidate negative feature vectors on the same side of the separating hyperplane as the positive feature vector; Step 4, converting candidate negative feature vectors among the candidate negative feature vectors that are at a distance greater than a predetermined distance threshold from the separating hyperplane into positive feature vectors; and Step 5, iterating Steps 2 to 4 until no more positive feature vectors are converted.
[0014] In the above training method of the model for screening adjudication documents of controversial issues, before using the positive-unlabeled learning strategy to train a support vector machine model with the labeled samples as positive samples and the unlabeled samples as negative samples, it further includes: selecting samples with a feature distribution similar to the first feature vector from the unlabeled samples based on the first feature vector of the labeled text as the labeled samples.
[0015] According to another aspect of the present application, a method for screening adjudication documents of controversial issues is provided, including: obtaining unlabeled adjudication documents to be screened; obtaining the distance value between the feature vector corresponding to the unlabeled adjudication documents to be screened and the separating hyperplane of the support vector machine using the above-mentioned model for screening adjudication documents of controversial issues; and screening whether the adjudication documents contain controversial issues based on the distance value as the confidence level that the unlabeled adjudication documents to be screened contain controversial issues.
[0016] According to another aspect of the present application, there is provided a training device for a model for screening adjudication documents of controversial focuses, including: a sample acquisition unit for acquiring training adjudication documents, where the training adjudication documents include a small number of labeled samples and a large number of unlabeled samples; a preprocessing unit for performing text preprocessing and paragraph extraction on the training adjudication documents; a feature extraction unit for constructing a first feature vector of the labeled samples and a second feature vector of the unlabeled samples from the training adjudication documents; and a model training unit for using a positive-unlabeled learning strategy to train a support vector machine model with the labeled samples as positive samples and the unlabeled samples as negative samples, where the support vector machine model is used to classify adjudication documents as whether they contain controversial focuses.
[0017] According to yet another aspect of the present application, there is provided a screening device for adjudication documents of controversial focuses, including: a document acquisition unit for acquiring unlabeled adjudication documents to be screened; a distance determination unit for obtaining the distance value between the feature vector corresponding to the unlabeled adjudication documents to be screened and the separation hyperplane of the support vector machine using the model for screening adjudication documents of controversial focuses as described above; and a document screening unit for screening whether the adjudication documents contain controversial focuses based on the distance value as the confidence level that the unlabeled adjudication documents to be screened contain controversial focuses.
[0018] According to still another aspect of the present application, there is provided an electronic device, including: a processor; and a memory, where computer program instructions are stored in the memory, and when the computer program instructions run on the processor, the processor is caused to execute the training method for the model for screening adjudication documents of controversial focuses as described above or the screening method for adjudication documents of controversial focuses as described above.
[0019] According to yet another aspect of the present application, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a computing device, they can be operable to execute the training method for the model for screening adjudication documents of controversial focuses as described above or the screening method for adjudication documents of controversial focuses as described above.
[0020] The training method for the model for screening adjudication documents of controversial focuses, the screening method for adjudication documents of controversial focuses, the device, and the electronic device provided by the embodiments of the present application can construct document-level controversial focus features from a small number of existing labeled documents based on the positive-unlabeled strategy, and use a support vector machine model to identify adjudication documents with controversial focuses from a large number of unlabeled documents, thereby improving the screening scheme for adjudication documents of controversial focuses. Description of the Drawings
[0021] By reading the detailed description in the following preferred specific embodiments, various other advantages and benefits of the present application will become clear to those of ordinary skill in the art. The accompanying drawings of the specification are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Obviously, the following described drawings are only some embodiments of the present application, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.
[0022] Figure 1 The flowchart of the training method of the model for screening the adjudication documents of the focus of disputes according to an embodiment of the present application is illustrated;
[0023] Figure 2 The flowchart of the screening method of the adjudication documents of the focus of disputes according to an embodiment of the present application is illustrated;
[0024] Figure 3 The block diagram of the training device of the model for screening the adjudication documents of the focus of disputes according to an embodiment of the present application is illustrated;
[0025] Figure 4 The block diagram of the screening device of the adjudication documents of the focus of disputes according to an embodiment of the present application is illustrated;
[0026] Figure 5 The block diagram of the electronic device according to an embodiment of the present application is illustrated. Specific Embodiments
[0027] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0028] Exemplary method
[0029] Figure 1 The flowchart of the training method of the model for screening the adjudication documents of the focus of disputes according to an embodiment of the present application is illustrated.
[0030] As Figure 1 shown, the training method of the model for screening the adjudication documents of the focus of disputes according to an embodiment of the present application includes the following steps.
[0031] S110. Obtain training judgment documents, where the training judgment documents include a small number of labeled samples and a large number of unlabeled samples. That is, as described above, since the annotation of the dispute focus requires a great deal of effort, only a small number of labeled samples and a large number of unlabeled samples are available. Therefore, in the embodiments of the present application, the training judgment documents include a small number of labeled samples and a large number of unlabeled samples. For example, the proportion of labeled samples in the total samples is less than 5%. And those skilled in the art can understand that the labeled samples are judgment documents that have been annotated with the dispute focus.
[0032] Here, based on the professional knowledge in the field of legal judgment document processing, the dispute focus is the key issue summarized by the judge regarding the disputes over evidence facts and legal application, which is both the main content of the court trial and the main line for making judgment documents. The dispute focus is related to the focus of reasoning in the judgment document. If the dispute focus of an article is summarized accurately, the reasoning of the judgment document will have a clear goal and direction; on the contrary, if the dispute focus is summarized inaccurately or incompletely, the judgment document may be biased in reasoning and aimless.
[0033] The dispute focus is the main problem that needs to be resolved after a dispute occurs between the parties. First of all, it is a problem. Specifically, it includes the main problems in aspects such as the facts, evidence, legal provisions, and responsibilities that give rise to the dispute. Since it is a problem, it can be described in languages such as "whether" and "how", such as "whether the contract is effective", "whether it constitutes infringement", "how to determine the liability", etc. This is also a common expression in judgment documents in judicial practice and an important reference for automatically identifying whether there is a dispute focus in judgment documents by machines.
[0034] S120. Perform text preprocessing and paragraph extraction on the training judgment documents. Here, text preprocessing mainly performs anomaly processing on the text, including unified conversion of numbers and unified punctuation. And paragraph extraction is to extract the corresponding judgment analysis process paragraphs and defense paragraphs from the judgment documents through a document segmentation model.
[0035] Therefore, in the training method of the model for screening dispute focus judgment documents according to the embodiments of the present application, the text preprocessing includes unified conversion of numbers and / or unified punctuation.
[0036] And, in the training method of the model for screening dispute focus judgment documents according to the embodiments of the present application, the paragraph extraction includes extracting the corresponding judgment analysis process paragraphs and defense paragraphs from the training judgment documents through a document segmentation model.
[0037] S130. Construct the first feature vector of the labeled samples and the second feature vector of the unlabeled samples from the training judgment documents. That is, based on the professional knowledge in the field of legal judgment document processing, not all cases necessarily need to summarize the controversial issues. Generally speaking, except for the cases where the defendant fully acknowledges the facts and litigation requests claimed by the plaintiff, there will inevitably be controversial differences between the plaintiff and the defendant when there are pleadings, and thus there will inevitably be controversial issues. However, in the judgment documents, not all controversies need to be summarized as controversial issues for discussion. Therefore, the main features of controversial issues generally appear in the corresponding defense paragraphs and judgment analysis paragraphs. By observing and analyzing the judgment documents with labeled controversial issues, it is found that the paragraph type and paragraph length of certain paragraphs in the judgment documents determine whether there are controversial issues in the judgment documents. Specifically, if a document has a defense paragraph, it indicates that the defendant has put forward reasons for defense, and this document is likely to have an expression of controversial issues. For the judgment analysis process paragraph, if the paragraph length is long, it indicates that the case is more complex and the document has a longer description of the case analysis process, so the possibility of having controversial issues is also greater. Therefore, the existence of a defense paragraph and paragraphs with a length greater than a predetermined threshold can be considered as the characteristics of judgment documents with controversial issues.
[0038] In addition, through the analysis and statistics of sentences with controversial issues, multiple reasoning key words are summarized, mainly including some transition words and connecting words often appearing in judgment documents such as "although, but, even if, therefore", etc. That is, the number of these key words in a specific paragraph determines to a certain extent whether there are relevant expressions of controversial issues in the judgment document. Therefore, these key words can be used as the characteristics of judgment documents with controversial issues.
[0039] In this way, through the paragraph characteristics and key word characteristics, in the embodiments of this application, for judgment documents of different case types, the constructed characteristics are independent of the case types. Therefore, based on a small amount of labeling, through the semi-supervised learning method, examples with higher scores in the unlabeled data can be found as positive examples.
[0040] In an example, the controversial issue characteristics at the document level mainly include: the number of paragraphs in the judgment document, the length of the judgment analysis process paragraph, the length of the defense paragraph, the number of reasoning characteristic words in the judgment analysis process paragraph, and the number of reasoning characteristic words in the defense paragraph.
[0041] Moreover, for these features, in the method for training a model for screening controversial focus judgment documents according to an embodiment of the present application, the feature weights can also be learned through the model. That is, in the process of model learning, first, the unlabeled text is regarded as a negative example, and the labeled text is regarded as a positive example, features are extracted, and the model is trained in the first stage. After the model converges, the model will automatically learn the feature weights. In this way, through the features extracted above and the corresponding weights of the features, the feature vectors for classification of each sample data can be constructed.
[0042] Therefore, in the method for training a model for screening controversial focus judgment documents according to an embodiment of the present application, constructing the first feature vector of the labeled sample and the second feature vector of the unlabeled sample from the training judgment documents includes: determining document paragraphs with a relevance greater than a predetermined threshold to the controversial focus based on the paragraph type and paragraph length of whether the judgment document contains the controversial focus; and constructing the first feature vector of the labeled sample and the second feature vector of the unlabeled sample based on the keywords obtained by analyzing and counting the sentences of the controversial focus and the document paragraphs.
[0043] Moreover, in the method for training a model for screening controversial focus judgment documents according to an embodiment of the present application, the first feature vector and the second feature vector include the following features and their weights: the number of paragraphs of the judgment document, the length of the judgment analysis process paragraph, the length of the defense paragraph, the number of reasoning feature words in the judgment analysis process paragraph, and the number of reasoning feature words in the defense paragraph.
[0044] S140. Using the positive-unlabeled learning strategy, training a support vector machine model with the labeled sample as the positive sample and the unlabeled sample as the negative sample, where the support vector machine model is used to classify judgment documents as whether they contain controversial foci.
[0045] Specifically, positive-unlabeled learning (PU Learning) is a research direction of semi-supervised learning, which refers to training a classifier in the case of only positive classes and unlabeled data, and has wide applications in fields such as data retrieval, anomaly detection, and sequence data monitoring. In the method for training a model for screening controversial focus judgment documents according to an embodiment of the present application, based on the PU-learning strategy, the positive samples and the unlabeled samples are respectively regarded as positive samples and negative samples, and then a support vector machine is trained using these data. The support vector machine can assign a score to each judgment document. Usually, the score of the positive sample is higher than that of the negative sample. Therefore, for those unlabeled samples, the ones with higher scores are most likely to be positive samples.
[0046] That is, the support vector machine (SVM) is essentially a binary classification model. Its basic model is a linear classifier with the maximum margin defined in the feature space. It also includes the kernel trick, which makes it a substantially non-linear classifier. The learning strategy of SVM is to maximize the margin, which can be formulated as a problem of solving a convex quadratic programming, and is also equivalent to the minimization problem of the regularized hinge loss function. The learning algorithm of SVM is the optimization algorithm for solving convex quadratic programming. The basic idea of SVM learning is to solve the separating hyperplane that can correctly divide the training data set and has the maximum geometric margin. For a linearly separable data set, there are infinitely many such hyperplanes (i.e., perceptrons), but the separating hyperplane with the maximum geometric margin is unique.
[0047] Therefore, in the embodiment of the present application, during the training process, the feature vectors of the unlabeled judgment documents whose feature vectors are on the same side of the separating hyperplane as the feature vectors of the labeled judgment documents with controversial focuses are mainly concerned. The judgment documents that are far from the hyperplane among these judgment documents are added to the labeled data, and then the next round of training is continued in the same way. After repeating many times, when there are no newly added judgment documents with controversial focuses, the screening model tends to be stable and can be used to predict whether other similar judgment documents contain controversial focuses.
[0048] That is, in the training method of the model for screening judgment documents with controversial focuses according to the embodiment of the present application, using the positive-unlabeled learning strategy, training the support vector machine model with the labeled samples as positive samples and the unlabeled samples as negative samples includes: Step 1, taking the first feature vector as the positive feature vector and the second feature vector as the negative feature vector; Step 2, determining the separating hyperplane with the maximum geometric margin of the support vector machine model based on the positive feature vector and the negative feature vector; Step 3, determining the candidate negative feature vectors on the same side of the separating hyperplane as the positive feature vector; Step 4, converting the candidate negative feature vectors whose distance from the separating hyperplane is greater than a predetermined distance threshold among the candidate negative feature vectors into positive feature vectors; and iterating Step 2 to Step 4 until no more positive feature vectors are converted.
[0049] In addition, when training the model for screening judgment documents with controversial focuses according to the embodiment of the present application, it is also possible to train only using unlabeled text. That is, since there are actually some positive examples in the unlabeled text, during the previous model training process, the model can identify some positive examples existing in the unlabeled text through the unique feature distribution of the positive examples learned from the positive examples. Then these positive examples are added to the labeled text and the next stage of training is carried out. This process is repeated many times. After a certain number of rounds of training, a stable state is reached. When the model can no longer find data similar to the positive examples from the unlabeled data, the model for identifying controversial focuses is obtained.
[0050] Therefore, in the training method of the model for screening judgment documents of controversial focuses according to the embodiments of the present application, before using the positive unlabeled learning strategy to train the support vector machine model with the labeled samples as positive samples and the unlabeled samples as negative samples, it further includes: based on the first feature vector of the labeled text, selecting samples with a feature distribution similar to the first feature vector from the unlabeled samples as the labeled samples.
[0051] Moreover, based on the above-mentioned training method of the model for screening judgment documents of controversial focuses, the embodiments of the present application provide a method for screening judgment documents of controversial focuses.
[0052] Figure 2 The flowchart of the method for screening judgment documents of controversial focuses according to the embodiments of the present application is illustrated.
[0053] As Figure 2 shown, the method for screening judgment documents of controversial focuses according to the embodiments of the present application includes the following steps.
[0054] S210, obtaining unlabeled judgment documents to be screened.
[0055] S220, using the model for screening judgment documents of controversial focuses as described above to obtain the distance value between the feature vector corresponding to the unlabeled judgment documents to be screened and the separation hyperplane of the support vector machine; and
[0056] S230, based on the distance value as the confidence level that the unlabeled judgment documents to be screened contain controversial focuses, screening whether the judgment documents contain controversial focuses.
[0057] Exemplary apparatus
[0058] Figure 3 The block diagram of the training device of the model for screening judgment documents of controversial focuses according to the embodiments of the present application is illustrated.
[0059] As Figure 3As shown, the training device 300 for the model for screening judgment documents of controversial focuses according to an embodiment of the present application includes: a sample acquisition unit 310 for acquiring training judgment documents, where the training judgment documents include a small part of labeled samples and a large part of unlabeled samples; a preprocessing unit 320 for performing text preprocessing and paragraph extraction on the training judgment documents; a feature extraction unit 330 for constructing a first feature vector of the labeled samples and a second feature vector of the unlabeled samples from the training judgment documents; and a model training unit 340 for using the positive-unlabeled learning strategy to train a support vector machine model with the labeled samples as positive samples and the unlabeled samples as negative samples, where the support vector machine model is used to classify judgment documents as whether they contain controversial focuses.
[0060] In one example, in the above training device 200 for the model for screening judgment documents of controversial focuses, the preprocessing unit 320 is used to perform unified digital conversion and / or unified punctuation.
[0061] In one example, in the above training device 200 for the model for screening judgment documents of controversial focuses, the preprocessing unit 320 is used to extract corresponding judgment analysis process segments and defense segments from the training judgment documents through a document segmentation model.
[0062] In one example, in the above training device 200 for the model for screening judgment documents of controversial focuses, the feature extraction unit 330 is used to: determine document paragraphs with a relevance greater than a predetermined threshold to the controversial focus based on the paragraph type and paragraph length of whether the judgment document contains a controversial focus; and construct a first feature vector of the labeled samples and a second feature vector of the unlabeled samples based on the keywords obtained by analyzing and counting the sentences of the controversial focus and the document paragraphs.
[0063] In one example, in the above training device 200 for the model for screening judgment documents of controversial focuses, the first feature vector and the second feature vector include the following features and their weights: the number of paragraphs of the judgment document, the length of the judgment analysis process segment, the length of the defense paragraph, the number of reasoning feature words in the judgment analysis process segment, the number of reasoning feature words in the defense paragraph
[0064] In one example, in the training apparatus 200 of the above-mentioned model for screening adjudication documents of dispute focus, the model training unit 340 is configured to perform the following steps: Step 1, taking the first feature vector as the positive feature vector and the second feature vector as the negative feature vector; Step 2, determining the separating hyperplane with the largest geometric margin of the support vector machine model based on the positive feature vector and the negative feature vector; Step 3, determining candidate negative feature vectors on the same side of the separating hyperplane as the positive feature vector; Step 4, converting candidate negative feature vectors among the candidate negative feature vectors that are at a distance greater than a predetermined distance threshold from the separating hyperplane into positive feature vectors; and Step 5, iterating Steps 2 to 4 until no more positive feature vectors are converted.
[0065] In one example, in the training apparatus 200 of the above-mentioned model for screening adjudication documents of dispute focus, the model training unit 340 is further configured to: before training the support vector machine model using the positive unlabeled learning strategy with the labeled samples as positive samples and the unlabeled samples as negative samples, based on the first feature vector of the labeled text, select samples with a feature distribution similar to the first feature vector from the unlabeled samples as the labeled samples.
[0066] Figure 4 The block diagram of the screening apparatus for adjudication documents of dispute focus according to an embodiment of the present application is illustrated.
[0067] As Figure 4 shown, the screening apparatus 400 for adjudication documents of dispute focus according to an embodiment of the present application includes: a document acquisition unit 410, configured to acquire unlabeled adjudication documents to be screened; a distance determination unit 420, configured to obtain the distance value between the feature vector corresponding to the unlabeled adjudication documents to be screened and the separating hyperplane of the support vector machine using the above-mentioned model for screening adjudication documents of dispute focus; and a document screening unit 430, configured to screen whether the adjudication documents contain dispute focus based on the distance value as the confidence level that the unlabeled adjudication documents to be screened contain dispute focus.
[0068] Here, those skilled in the art can understand that the specific functions and operations of each unit and module in the above-mentioned training apparatus 300 for the model for screening adjudication documents of dispute focus and the screening apparatus 400 for adjudication documents of dispute focus have been introduced in detail in the above-mentioned training method for the model for screening adjudication documents of dispute focus and the screening method for adjudication documents of dispute focus, and thus, the repeated description thereof will be omitted. Figure 1 and Figure 2 described, and therefore, the repeated description thereof will be omitted.
[0069] As described above, the training device 300 for the model for screening controversial focus judgment documents and the screening device 400 for controversial focus judgment documents according to the embodiments of the present application can be implemented in various terminal devices, such as a server for processing legal judgment documents. In one example, the training device 300 for the model for screening controversial focus judgment documents and the screening device 400 for controversial focus judgment documents according to the embodiments of the present application can be integrated into the terminal device as a software module and / or a hardware module. For example, the training device 300 for the model for screening controversial focus judgment documents and the screening device 400 for controversial focus judgment documents can be a software module in the operating system of the terminal device, or can be an application developed for the terminal device; of course, the training device 300 for the model for screening controversial focus judgment documents and the screening device 400 for controversial focus judgment documents can also be one of the numerous hardware modules of the terminal device.
[0070] Alternatively, in another example, the training device 300 for the model for screening controversial focus judgment documents and the screening device 400 for controversial focus judgment documents and the terminal device can also be separate devices, and the training device 300 for the model for screening controversial focus judgment documents and the screening device 400 for controversial focus judgment documents can be connected to the terminal device through a wired and / or wireless network, and transmit interaction information in accordance with a predefined data format.
[0071] Exemplary electronic device
[0072] Next, reference will be made to Figure 5 to describe the electronic device according to the embodiments of the present application.
[0073] Figure 5 The block diagram of the electronic device according to the embodiments of the present application is illustrated.
[0074] As Figure 5 shown, the electronic device 10 includes one or more processors 11 and a memory 12.
[0075] The processor 11 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 10 to perform desired functions.
[0076] The memory 12 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 11 may run the program instructions to implement the training method of the model for screening the judgment documents of the focus of dispute and the screening method of the judgment documents of the focus of dispute described above in various embodiments of the present application and / or other desired functions. Various contents such as judgment documents and extracted features may also be stored in the computer-readable storage media.
[0077] In one example, the electronic device 10 may further include: an input device 13 and an output device 14, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0078] For example, the input device 13 may be, for example, a keyboard, a mouse, etc.
[0079] The output device 14 may output various information to the outside, such as the screening result of the judgment document, etc. The output device 14 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0080] Of course, for simplicity, Figure 5 only some of the components related to the present application in the electronic device 10 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 10 may further include any other appropriate components.
[0081] Exemplary computer program product and computer-readable storage medium
[0082] In addition to the above methods and devices, the embodiments of the present application may also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the training method of the model for screening the judgment documents of the focus of dispute and the screening method of the judgment documents of the focus of dispute described in the "Exemplary Method" section above of this specification according to various embodiments of the present application.
[0083] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0084] In addition, an embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the training method of the model for screening dispute focus judgment documents and the screening method of dispute focus judgment documents according to various embodiments of the present application described in the above "Exemplary Method" section of this specification.
[0085] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0086] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. In addition, the above-disclosed specific details are only for illustrative purposes and for ease of understanding, and are not limitations. The above details do not limit the present application to necessarily implement using the above specific details.
[0087] The block diagrams of the devices, apparatuses, equipment, and systems involved in this application are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms meaning "including but not limited to" and can be used interchangeably with each other. The word "or" and "and" used herein refer to the phrase "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The phrase "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.
[0088] It should also be noted that in the devices, equipment, and methods of this application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this application.
[0089] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0090] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A training method for a model for screening judgment documents of controversial focuses, characterized in that, it includes: Obtain training judgment documents, where the training judgment documents include a small number of labeled samples and a large number of unlabeled samples; Perform text preprocessing and paragraph extraction on the training judgment documents; Construct a first feature vector of the labeled samples and a second feature vector of the unlabeled samples from the training judgment documents; and Use the positive unlabeled learning strategy, use the labeled samples as positive samples and the unlabeled samples as negative samples to train a support vector machine model, and the support vector machine model is used to classify judgment documents as whether they contain controversial focuses; The constructing the first feature vector of the labeled samples and the second feature vector of the unlabeled samples from the training judgment documents includes: Based on the paragraph type and paragraph length of whether the judgment document contains a controversial focus, determine the document paragraphs with a relevance greater than a predetermined threshold to the controversial focus; and, Construct the first feature vector of the labeled samples and the second feature vector of the unlabeled samples based on the keywords obtained by analyzing and counting the sentences of the controversial focus and the document paragraphs; The paragraph types containing controversial focuses include defense paragraphs and judgment analysis process paragraphs; The using the positive unlabeled learning strategy, using the labeled samples as positive samples and the unlabeled samples as negative samples to train a support vector machine model includes: Step 1, use the first feature vector as the positive feature vector and the second feature vector as the negative feature vector; Step 2, determine the separation hyperplane with the largest geometric margin of the support vector machine model based on the positive feature vector and the negative feature vector; Step 3, determine the candidate negative feature vectors on the same side of the separation hyperplane as the positive feature vector; Step 4, convert the candidate negative feature vectors among the candidate negative feature vectors that are at a distance greater than a predetermined distance threshold from the separation hyperplane into positive feature vectors; and, Step 5, iterate steps 2 to 4 until no positive feature vectors are converted anymore.
2. The training method for a model for screening judgment documents of controversial focuses according to claim 1, characterized in that, The text preprocessing includes digital unified conversion and / or punctuation unified conversion.
3. The training method for a model for screening judgment documents of controversial focuses according to claim 1, characterized in that, The paragraph extraction includes extracting the corresponding judgment analysis process paragraphs and defense paragraphs from the training judgment documents through a document segmentation model.
4. The training method for a model for screening judgment documents of controversial focuses according to claim 1, characterized in that, The first feature vector and the second feature vector include the following features and their weights: the number of paragraphs of the judgment document, the length of the judgment analysis process paragraphs, the length of the defense paragraphs, the number of reasoning feature words in the judgment analysis process paragraphs, the number of reasoning feature words in the defense paragraphs.
5. The training method for a model for screening judgment documents of controversial focuses according to claim 1, characterized in that, Before using the positive-unlabeled learning strategy to train a support vector machine model with the labeled samples as positive samples and the unlabeled samples as negative samples, it further includes: Based on the first feature vector of the labeled samples, select samples from the unlabeled samples that have a feature distribution similar to the first feature vector as the labeled samples.
6. A method for screening adjudicative documents of controversial issues Characterized in that It includes: Obtain unlabeled adjudicative documents to be screened; Use the model for screening adjudicative documents of controversial issues according to any one of claims 1 to 5 to obtain the distance value between the feature vector corresponding to the unlabeled adjudicative document to be screened and the separation hyperplane of the support vector machine; And Based on the distance value as the confidence level that the unlabeled adjudicative document to be screened contains controversial issues, screen whether the adjudicative document contains controversial issues.
7. A training device for a model for screening adjudicative documents of controversial issues Characterized in that It includes: A sample acquisition unit for acquiring training adjudicative documents, where the training adjudicative documents include a small number of labeled samples and a large number of unlabeled samples; A preprocessing unit for performing text preprocessing and paragraph extraction on the training adjudicative documents; A feature extraction unit for constructing the first feature vector of the labeled samples and the second feature vector of the unlabeled samples from the training adjudicative documents; And A model training unit for using the positive-unlabeled learning strategy to train a support vector machine model with the labeled samples as positive samples and the unlabeled samples as negative samples, where the support vector machine model is used to classify adjudicative documents as whether they contain controversial issues; The feature extraction unit constructs the first feature vector of the labeled samples and the second feature vector of the unlabeled samples from the training adjudicative documents, including: Based on the paragraph type and paragraph length of whether the adjudicative document contains controversial issues, determine the document paragraphs with a relevance greater than a predetermined threshold to the controversial issues; and, Construct the first feature vector of the labeled samples and the second feature vector of the unlabeled samples based on the keywords obtained by analyzing and counting the sentences of the controversial issues and the document paragraphs; The paragraph types containing controversial issues include defense paragraphs and adjudication analysis process paragraphs; The model training unit uses the positive-unlabeled learning strategy to train a support vector machine model with the labeled samples as positive samples and the unlabeled samples as negative samples, including: Step 1, use the first feature vector as the positive feature vector and the second feature vector as the negative feature vector; Step 2, determine the separation hyperplane with the largest geometric margin of the support vector machine model based on the positive feature vector and the negative feature vector; Step 3, determine the candidate negative feature vectors on the same side of the separation hyperplane as the positive feature vector; Step 4, convert the candidate negative feature vectors among the candidate negative feature vectors that are at a distance greater than a predetermined distance threshold from the separation hyperplane into positive feature vectors; and, Step 5, iterate steps 2 to 4 until no positive feature vectors are converted anymore.
8. A screening device for adjudication documents of controversial focus, characterized in that, it includes: a document acquisition unit for acquiring unannotated adjudication documents to be screened; a distance determination unit for obtaining the distance value between the feature vector corresponding to the unannotated adjudication document to be screened and the separation hyperplane of the support vector machine by using the model for screening adjudication documents of controversial focus described in any one of claims 1 to 5; and a document screening unit for screening whether the adjudication document contains a controversial focus based on the distance value as the confidence level that the unannotated adjudication document to be screened contains a controversial focus.
9. An electronic device, characterized in that, it includes: a processor; and a memory in which computer program instructions are stored, and when the computer program instructions run on the processor, the processor is caused to execute the training method of the model for screening adjudication documents of controversial focus described in any one of claims 1 to 5 or the screening method of adjudication documents of controversial focus described in claim 6.
10. A computer-readable storage medium, characterized in that, computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a computing device, they can be operated to execute the training method of the model for screening adjudication documents of controversial focus described in any one of claims 1 to 5 or the screening method of adjudication documents of controversial focus described in claim 6.
Citation Information
Patent Citations
Topic model-based judgment document similarity analysis method
CN107291688A
Data processing method, system, device, medium and data analysis method
CN113282726A
Dispute focus discovery method and device based on dispute focus entity, and terminal
CN111814477A
System and method for electronic text classification
US20200050618A1