Method and system for handling small-amount unhealthy asset litigation cases

By constructing a dynamic knowledge graph and introducing an external legal knowledge base and setting up legal relationship constraints, the problem of difficult to take into account compliance and similarity in handling small non-performing asset litigation cases is solved, and the legal interpretability and compliance of case grouping are achieved.

CN120256528APending Publication Date: 2025-07-04BEIJING LUSHU TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510317229.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In handling small non-performing assets litigation cases, it is difficult for the existing technology to take into account the legal compliance and similarity of the cases, resulting in inconsistent grouping results and lack of explanatory ability.

Method used

By constructing a dynamic knowledge graph, introducing an external legal knowledge base, setting legal relationships as constraints, designing a similarity calculation function that includes legal relationships and case characteristics ontology, and using a hierarchical clustering algorithm to cluster cases to generate case materials.

Benefits of technology

It improves the legal interpretability and compliance of case grouping, ensures that clustering results are in line with legal relationships, and enhances the credibility and interpretability of grouping results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256528A_ABST
    Figure CN120256528A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for processing a small-amount unhealthy asset litigation case, and relates to the field of knowledge maps, and the method comprises the steps: obtaining the data of the small-amount unhealthy asset litigation case; preprocessing the case data to obtain a structured case attribute table; a dynamic knowledge graph is constructed, the dynamic knowledge graph is interconnected with an external legal knowledge base, and knowledge reasoning is introduced for dynamic expansion of knowledge; designing a similarity calculation function including legal relationship constraints and case feature ontologies; embedding the similarity calculation function into a hierarchical clustering algorithm, and performing case clustering and grouping by using the embedded hierarchical clustering algorithm; and according to a grouping result, generating a case filing material by adopting knowledge reasoning. Aiming at the problem that law compliance and case similarity are difficult to consider in existing small-amount unhealthy asset litigation case grouping, an external law knowledge base is introduced to update a domain knowledge graph, and a law relationship is set as a constraint, so that the law interpretability of case grouping is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of knowledge graphs, and particularly to a method and system for handling small-amount non-performing asset litigation cases. Background Art

[0002] In the process of recovering small-amount non-performing assets, litigation is one of the most common and effective means. However, there are some problems that need to be solved urgently in the traditional way of handling small-amount non-performing asset litigation cases. First, the number of cases is large and scattered, and the efficiency of manual classification and processing is low. Second, there is no unified standard and specification for case grouping, which easily leads to problems such as unreasonable grouping and inconsistent processing results. Moreover, case grouping often only considers the surface similarity of cases, ignoring the relevance and constraints of legal relationships between cases, resulting in the grouping results may be contrary to legal provisions, lacking compliance and interpretability.

[0003] In response to the above problems, some attempts have been made in the prior art to automatically group cases using artificial intelligence and big data technologies. These methods learn and cluster case features through machine learning algorithms, which improve the efficiency and accuracy of case grouping to a certain extent. However, these methods still have some limitations. On the one hand, they mainly rely on the features of the case data itself, lacking consideration of domain knowledge and legal rules, and it is difficult to ensure the legal compliance of the grouping results. On the other hand, these methods usually map case features to a high-dimensional space for similarity calculation and clustering, lacking clear legal explanations, and the interpretability and credibility of the grouping results are relatively low.

[0004] For example, patent document CN111241241B discloses a case retrieval method, device, equipment and storage medium based on a knowledge graph, which improves the usability of the case retrieval system. The method of this application includes: constructing a legal case knowledge graph according to text information, performing random walk sampling on the node set data constructed according to the legal case knowledge graph to obtain multiple sequence data, training the model based on the multiple sequence data through a word conversion vector algorithm to obtain an updated target model, obtaining target text information, and analyzing the target text information through the target model to construct a knowledge graph to be retrieved, retrieving in the legal case knowledge graph according to the knowledge graph to be retrieved to obtain case information associated with the knowledge graph to be retrieved, and obtaining the output case information according to the first similarity and the second similarity of the case information. However, when calculating the case similarity, this solution mainly relies on the structural similarity and semantic similarity of nodes in the knowledge graph, lacking consideration of the legal relationship between cases, and it is difficult to ensure the legal compliance of case grouping. Summary of the Invention

[0005] In view of the difficulty in balancing legal compliance and case similarity in the grouping of small-value non-performing asset litigation cases in the existing technology, the present application provides a method and system for processing small-value non-performing asset litigation cases. By introducing an external legal knowledge base to update the domain knowledge graph and setting legal relationships as constraints, the legal interpretability of case grouping is improved.

[0006] The object of the present application is achieved by the following technical solutions.

[0007] One aspect of the present application provides a method for processing small-value non-performing asset litigation cases, including: S1, obtaining small-value non-performing asset litigation case data; S2, preprocessing the case data to obtain a structured case attribute table, where the case attribute table includes debt elements and party information; S3, constructing a dynamic knowledge graph according to the case attribute table, and the dynamic knowledge graph is interconnected with an external legal knowledge base to introduce knowledge reasoning for dynamic expansion of knowledge; S4, designing a similarity calculation function including legal relationship constraints and case feature ontologies according to the dynamic knowledge graph; S5, embedding the similarity calculation function into a hierarchical clustering algorithm, and using the embedded hierarchical clustering algorithm for case clustering and grouping; S6, generating case-filing materials by using knowledge reasoning according to the grouping result, and the case-filing materials include a statement of claim and an evidence list.

[0008] Further, S2, preprocessing the case data to obtain a structured case attribute table, where the case attribute table includes debt elements and party information, includes: S21, parsing and feature extracting the case text data by using regular expressions. To improve the flexibility and robustness of matching, we introduce a fuzzy matching operator in the regular expressions. The fuzzy matching operator uses the edit distance algorithm to calculate the similarity between the matching string and the pattern rule. When the similarity exceeds the set threshold, it is considered a successful match. This fuzzy matching mechanism allows for a certain literal difference between the keywords in the pattern rule and the case text, improving the recall rate of feature extraction. For example, the keyword "loan amount" in the pattern rule can match synonymous expressions such as "loan limit" and "loan quantity" in the case text. By predefined a series of pattern rules, we can accurately extract debt information features such as debt amount and debt term from the case text.

[0009] S22. Semantically annotate and perform entity recognition on the case features extracted in S21. We use the Named Entity Recognition (NER) method to label the entity categories of the key information in the case features. Common NER methods include rule-based methods, statistical machine learning-based methods (such as Conditional Random Field, CRF), and deep learning-based methods (such as BiLSTM-CRF). We can select an appropriate NER method according to the characteristics of the case domain and the data scale. Through NER, we can label the key information in the case features as predefined entity categories, such as "debt amount", "debt term", "debtor information", etc. The annotated case features form a structured representation of the case features.

[0010] S23. Based on the structured case features obtained in S22, we construct a case attribute table. The case attribute table uses a relational model (such as a relational database) to organize and store the case features. In the case attribute table, each row corresponds to a case, and each column corresponds to a case attribute (such as debt amount, debt term, etc.). The values of the case attributes come from the specific values of the corresponding entity categories in the structured case features. Through the case attribute table, we transform the unstructured case text into a structured form that is convenient for computer processing. The case attribute table provides a data basis for subsequent tasks such as case similarity calculation and clustering analysis.

[0011] Traditional regular expressions are very sensitive to changes in text format and grammar. Spelling mistakes, abbreviations, synonyms, etc. in the case data will all lead to matching failures. Regular expressions lack flexibility and fault tolerance, and it is difficult to handle the diversity and uncertainty of case data. In this application, by introducing fuzzy matching operators and using the edit distance algorithm to calculate the similarity between the matching string and the pattern rule, it is possible to tolerate changes in the format and grammar of case data to a certain extent, and improve the success rate and accuracy of matching. By introducing fuzzy matching operators, the fault tolerance of the regular expression matching algorithm is enhanced.

[0012] Further, in S3, a dynamic knowledge graph is constructed according to the case attribute table. S31: Construct a static knowledge graph. Using the case attribute association information contained in the case attribute table, we first construct a static knowledge graph of the case. The static knowledge graph uses a graph database (such as Neo4j) for storage and management. In the static knowledge graph, the semantic associations between case attributes are represented by RDF (Resource Description Framework) triples. An RDF triple consists of three parts: a subject, a predicate, and an object, corresponding to the case attribute entity, the attribute relationship, and the associated attribute entity respectively. For example, the association between the case attributes "loan amount" and "loan term" can be represented as the triple (loan amount, associated with, loan term). Through RDF triples, we can formally depict the semantic relationship network between case attributes.

[0013] S32, Link and map with an external knowledge base. To enhance the semantic expression ability of the static knowledge graph, we need to link and map it with an external legal knowledge base. The external legal knowledge base contains structured knowledge (such as legal concepts, legal provisions, etc.) and unstructured knowledge (such as case text) in related legal fields. We use the method of ontology alignment to achieve the semantic association between the static knowledge graph and the external knowledge base. Ontology alignment aims to discover semantically equivalent or similar concepts in two ontologies (i.e., knowledge bases) and establish mapping relationships. Common ontology alignment techniques include methods based on string similarity, methods based on structural similarity, methods based on external resources, etc. By comprehensively applying these techniques, we can establish semantic links between the static knowledge graph and the external legal knowledge base to achieve knowledge complementarity and expansion.

[0014] S33. Generate a dynamic knowledge graph. Based on the link mapping between the static knowledge graph and the external knowledge base, we further utilize knowledge fusion technology to integrate legal knowledge in the external knowledge base into the static knowledge graph to generate a dynamic knowledge graph. This process mainly includes the following steps: Legal concept embedding: Use pre-trained semantic representation models (such as Word2Vec, BERT, etc.) to perform semantic embedding on legal concepts (such as legal terms, legal articles) in the external knowledge base to obtain their semantic vector representations. Semantic embedding can depict the semantic similarity between legal concepts. Case attribute concept representation learning: Utilize the RDF triple information in the static knowledge graph and, through knowledge representation learning methods (such as TransE, TransR, etc.), perform representation learning on the case attribute concepts in the graph to obtain their context semantic vectors. Semantic similarity calculation: Calculate the similarity (such as cosine similarity) between the semantic vectors of legal concepts and the context vectors of case attribute concepts to construct a semantic similarity matrix between the two types of concepts. Concept matching: According to the semantic similarity matrix, use a heuristic algorithm (such as the Hungarian algorithm) to match legal concepts with case attribute concepts to obtain the optimal concept alignment result. The Hungarian algorithm can achieve a one-to-one mapping between two concept sets in the sense of global optimality through maximum weight matching. Knowledge fusion: According to the concept matching results, fuse legal knowledge (such as legal articles, case precedents) in the external legal knowledge base that matches the case attribute concepts into the static knowledge graph, expand the association relationships and semantic information between case attributes, and finally form a dynamic knowledge graph containing rich legal knowledge.

[0015] Furthermore, S4. Design a similarity calculation function that includes legal relationship constraints and case feature ontology based on the dynamic knowledge graph, including: S41. Organize the case attribute concepts in the dynamic knowledge graph into concept categories according to semantic relevance to form case feature concepts; and use ontology language to describe the association constraints between case feature concepts to form a case feature ontology.

[0016] S42. Convert the matching relationship between legal concepts in the external legal knowledge base and case attribute concepts in the dynamic knowledge graph into legal relationship constraints; extract the matching relationship between legal concepts and case attribute concepts contained in the dynamic knowledge graph generated in S3 into formal legal relationship constraints. The constraints describe the legal conditions that case attributes need to meet. This solution makes implicit legal knowledge explicit as constraint rules by mining the concept matching relationships in the graph, which are used to guide case similarity assessment. This helps improve the legal compliance of similarity calculation.

[0017] S43. Set up a sub-function for calculating the semantic similarity of concepts according to the semantic description of the case feature concepts to measure the semantic similarity between case feature concepts; set up a sub-function for calculating the similarity of constraint relationships according to the legal relationship constraints to evaluate the relevance of case feature concepts under legal constraints; construct a sub-function for calculating the semantic similarity between concepts based on the concept semantics defined in the case feature ontology. At the same time, construct a sub-function for calculating the similarity of constraint relationships that reflects the legal relevance between case features according to the legal relationship constraints extracted in S42.

[0018] S44. Combine the sub-function for calculating semantic similarity and the sub-function for calculating similarity of constraint relationships with weights to obtain a similarity calculation function. This solution combines semantic similarity and constraint similarity to describe case similarity from two dimensions: semantic relevance and legal consistency.

[0019] Further, in S5, embed the similarity calculation function into the hierarchical clustering algorithm and use the embedded hierarchical clustering algorithm to perform case clustering and grouping, including: S51. Use the similarity calculation function to calculate the similarity value between each pair of cases in the dynamic knowledge graph, and construct a similarity matrix according to the similarity value; S52. According to the similarity matrix, calculate the distance metric between cases; use the hierarchical clustering algorithm to perform layer-by-layer aggregation or splitting of cases according to the distance metric; S53. Perform a penalty for violating constraints on the clustering operations that do not meet the legal relationship constraints during the aggregation or splitting of cases; S54. Generate a case clustering tree, where each node of the clustering tree corresponds to a case sub-cluster and the leaf node corresponds to a single case.

[0020] Specifically, this application introduces a penalty mechanism for violating constraints to ensure that the clustering process meets the legal relationship constraints. That is, during hierarchical clustering in S52, each time an attempt is made to merge two cases (or case sub-clusters) into a new cluster. Before merging, first determine whether the merge operation meets the legal relationship constraints. Specifically, for the two cases (or sub-clusters) to be merged, check whether the legal relationship constraints of the corresponding nodes in the dynamic knowledge graph are consistent. For example, extract all the legal relationship constraints associated with the two cases (or sub-clusters) in the graph to form constraint sets A and B. Calculate the similarity between constraint sets A and B, and measure it with the Jaccard similarity coefficient: Sim(A, B) = |A ∩ B| / |A ∪ B|, where |A ∩ B| represents the number of intersection elements of A and B, and |A ∪ B| represents the number of union elements of A and B. Set a similarity threshold θ. If Sim(A, B) < θ, it is considered that the legal constraints of the two cases (or sub-clusters) are inconsistent and the merge operation violates the constraints.

[0021] When it is determined that the merge operation violates the legal constraints, the operation needs to be punished. Common penalty methods include: Penalty coefficient method: Design a penalty coefficient α (0<α<1) to attenuate the similarity value obtained by the merge operation that violates the constraint. If the similarity of the original merge operation is sim, the similarity after the penalty is sim'=α*sim. The value of α reflects the intensity of the penalty for the violation of the constraint. The smaller α is, the more severe the penalty is. Penalty function method: Design a penalty function f(sim) with the same scale as the similarity value to impose a penalty on the merge operation that violates the constraint. The similarity value after the penalty is sim'=sim-f(sim). The penalty function f(sim) can be a subtracting function of the similarity value, such as f(sim)=β*sim, where β is the penalty intensity parameter. Direct elimination method: For the merge operation that violates the constraint, directly set its similarity value to negative infinity or 0, so that it cannot participate in the subsequent clustering merge. This method has the strongest penalty and can completely prohibit clustering behaviors that violate the constraints. By reducing the similarity value, the possibility of clustering and merging cases that violate legal constraints is reduced, so that the clustering results are more consistent with legal relationship constraints. When generating the case clustering tree in S54, by embedding the constraint violation penalty mechanism, the cases under the same node in the clustering tree can be made more consistent in legal relationship constraints, thereby enhancing the legal compliance and interpretability of the clustering results.

[0022] Further, S52, according to the similarity matrix, the distance metric between the cases is calculated; according to the distance metric, the hierarchical clustering algorithm is used to aggregate or split the cases layer by layer, including: all cases are used as initial clusters; the average distance between each case in the initial cluster and the cases in the corresponding neighborhood is calculated using a density-based anomaly detection algorithm, and the cases with an average distance greater than a threshold are identified as noise points, and the noise points in the initial cluster are removed; the diameter of the cases in the initial cluster after the noise points are removed is calculated as the maximum distance in the cluster; the cluster with the largest diameter is selected for splitting, and the cluster with a distance greater than the threshold δ[i] is selected by calculating the local density ρ[i] and distance δ[i] of the cases in the cluster. min And the local density is greater than the threshold ρ min The cases are taken as seed cases; with the seed cases as the center, the cases are divided into corresponding sub-clusters according to the distance between the cases and the seed cases; the sub-clusters are split recursively until each cluster contains only one case or the preset cluster number threshold is reached.

[0023] Specifically, the cluster splitting of the traditional DIANA algorithm depends on the seed case corresponding to the maximum distance within the cluster. When there are noise or abnormal cases in the data, these cases may be located at the end points of the cluster diameter and are selected as seeds. Noise cases as seeds may disrupt the normal cluster splitting and lead to unstable clustering results.

[0024] In this application, first, a density-based outlier detection algorithm is used to detect noise in the cases of the initial clusters before cluster splitting. By calculating the average distance between each case and the cases within its neighborhood, the cases with an average distance greater than the threshold are identified as noise points and removed from the initial clusters. Second, when selecting seed cases, instead of simply selecting the cases at the endpoints of the cluster diameter, the local density ρ[i] and distance δ[i] of the cases are calculated, and the density and distance characteristics of the cases are considered comprehensively. Only the cases that simultaneously satisfy the distance being greater than the threshold δ min and the local density being greater than the threshold ρ min are selected as seed cases. This selection strategy can effectively prevent noise cases or marginal cases from becoming seeds and improve the representativeness and reliability of the seed cases.

[0025] Furthermore, the local density ρ[i] and distance δ[i] of the cases within the cluster are calculated, including: The calculation formula for the local density ρ[i] is: where dist(i,j) represents the distance between case i and case j, n is the total number of cases within the cluster, and d c [i] is the adaptive neighborhood radius of case i; d c [i] = μ i + α × σ i , where μ i and σ i are the mean and standard deviation of the distances between case i and other cases respectively, and α is a coefficient for controlling the neighborhood range, which can be adjusted according to the compactness of the cluster; The calculation formula for the distance δ[i] is: If the local density of case i is the largest within the cluster, then i.e., the distance between case i and the case with the farthest distance within the cluster; The calculation formula for the local density threshold ρ min is: ρ min = μ ρ - β × σ ρ , where μ ρ and σ ρ are the mean and standard deviation of the local densities of all cases within the cluster respectively, and β is a coefficient for controlling the threshold, which can be adjusted according to the density distribution characteristics of the cluster. The calculation formula for the distance threshold δ min is: δ min = μ δ + γ × σ δ , where μ δ and σ δ are the mean and standard deviation of the distances of all cases within the cluster respectively, and γ is a coefficient for controlling the threshold, which can be adjusted according to the distance distribution characteristics of the cluster.

[0026] Specifically, in the clustering application scenario of small - amount non - performing asset litigation cases, the distribution of cases often shows unevenness and complexity. The similarity and relevance between cases may be affected by various factors, such as the credit status of debtors, the amount of arrears, the overdue time, etc. Traditional fixed neighborhood and threshold settings are difficult to effectively capture the local characteristics and overall features of the case distribution, and are prone to unstable and low - quality clustering results.

[0027] In this application, on the one hand, an adaptive neighborhood radius d c [i] is introduced, which improves the adaptability of local density calculation. Traditional density calculation usually uses a fixed neighborhood radius, ignoring the unevenness of case distribution. In this application, the neighborhood radius is dynamically adjusted according to the mean and standard deviation of the distances between case i and other cases, which can better adapt to the local characteristics of case distribution. This design of the adaptive neighborhood makes the calculation of local density more accurate and can effectively reflect the closeness of cases in the local area. On the other hand, a relative density threshold ρ min and a relative distance threshold δ min are adopted, which improves the robustness of seed case selection. Traditional threshold settings usually rely on manual experience or fixed values and are difficult to adapt to the characteristics of different case distributions. In this application, the mean and standard deviation of the local densities and distances of all cases within the cluster are used to adaptively determine the threshold, which can better adapt to the overall characteristics of case distribution. By controlling the coefficients β and γ, the strictness of the threshold can be flexibly adjusted, making the selection of seed cases more robust and reliable.

[0028] Another aspect of this application also provides a small - amount non - performing asset litigation case processing system for implementing a small - amount non - performing asset litigation case processing method of this application.

[0029] Compared with the prior art, the advantages of this application are as follows:

[0030] On the one hand, traditional static knowledge graphs are usually constructed based on fixed knowledge sources, and their knowledge representation and association relationships are difficult to change once generated and cannot reflect the dynamic changes of case data in a timely manner. In this application, through the interconnection with an external legal knowledge base, knowledge reasoning technology is used to dynamically expand the knowledge graph, enabling it to continuously absorb new legal knowledge and case information and realizing real - time update of knowledge representation. This dynamic expansion mechanism solves the problem of the solidification of static knowledge graphs, enables the knowledge graph to adapt to the dynamic changes of cases, and improves its adaptability and expressive ability for cases.

[0031] Secondly, based on the dynamic knowledge graph, the present application designs a similarity calculation function that includes legal relationship constraints and case feature ontologies, endowing the measurement of case similarity with legal basis and interpretability. Traditional static knowledge graphs usually have difficulty accurately depicting the legal relationships and constraints between cases, making the calculation of case similarity lack legal interpretability. The dynamic knowledge graph, by integrating external legal knowledge, introducing legal relationship constraints, and using ontology technology to semantically represent case features, enables the similarity calculation function to fully consider the legal attributes and features of cases, improving the accuracy and interpretability of similarity measurement.

[0032] On the other hand, the traditional rule- and graph-based knowledge reasoning process is usually a black-box process, and the reasoning results lack interpretability and traceability. When new case attribute associations are inferred, it is impossible to clearly explain the reasoning process and basis, and it is difficult to review and verify the reasoning results. This may lead to the uncertainty and risk of the reasoning results. In the present application, first, through ontology alignment, the legal concepts in the external knowledge base are semantically mapped to the case attribute concepts in the static knowledge graph, establishing a corresponding relationship between the two. This semantic-level association provides a basis for subsequent knowledge fusion and also makes the external knowledge introduced in the reasoning process have a clear semantic source and basis. Secondly, by using a pre-trained semantic representation model to embed legal concepts, their semantic vectors are obtained; at the same time, using the RDF triple information of the static knowledge graph, the semantic vectors of case attribute concepts are obtained through context representation learning. Mapping concept semantics to the vector space makes the calculation of semantic similarity more accurate and efficient. Finally, by calculating the semantic similarity matrix between concepts and using heuristic optimization strategies such as the Hungarian algorithm, a globally optimal concept matching scheme is obtained. This matching strategy based on similarity optimization ensures the accuracy and reliability of knowledge fusion, making the introduced external knowledge have a strong semantic association with case attributes. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The present application will be further described in the form of exemplary embodiments, and these exemplary embodiments will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:

[0034] Figure 1 is an exemplary flowchart of a method for handling small-amount non-performing asset litigation cases according to some embodiments of the present application;

[0035] Figure 2 is an exemplary flowchart of constructing a structured case attribute table according to some embodiments of the present application;

[0036] Figure 3is an exemplary flowchart for constructing a dynamic knowledge graph according to some embodiments of the present application;

[0037] Figure 4 is an exemplary flowchart for establishing a similarity calculation function according to some embodiments of the present application. Detailed implementation manners

[0038] The methods and systems provided in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0039] As Figure 1 shown, a method for handling small-amount non-performing asset litigation cases includes: obtaining small-amount non-performing asset litigation case data; preprocessing the case data to obtain a structured case attribute table, where the case attribute table includes debt elements and party information; constructing a dynamic knowledge graph according to the case attribute table, and the dynamic knowledge graph is interconnected with an external legal knowledge base to introduce knowledge reasoning for dynamic expansion of knowledge; designing a similarity calculation function including legal relationship constraints and case feature ontologies according to the dynamic knowledge graph; embedding the similarity calculation function into a hierarchical clustering algorithm, and using the embedded hierarchical clustering algorithm to perform case clustering and grouping; generating case-filing materials by using knowledge reasoning according to the grouping results, where the case-filing materials include a complaint and an evidence list.

[0040] Specifically, S1, obtaining small-amount non-performing asset litigation case data; obtaining the case data of small-amount non-performing asset litigation from a public database, and the case data usually exists in an unstructured or semi-structured form such as text documents, spreadsheets, XML files, etc., and includes the basic information of the case, party information, debt description, etc.

[0041] As Figure 2As shown in the figure, in S2, preprocess the case data to obtain a structured case attribute table, including: S21, collect regular expressions introducing fuzzy matching operators to perform text parsing on the case data, and introduce the edit distance algorithm as a measure of fuzzy matching to calculate the similarity between two strings. Common edit distance algorithms include Levenshtein distance, Damerau-Levenshtein distance, etc., and appropriate algorithms can be selected according to requirements. Define fuzzy matching operators, such as "~" or "≈", indicating that a certain difference is allowed between strings. According to the characteristics of the case data and the key information to be extracted, design a series of regular expression pattern rules. The pattern rules can be designed for different parts such as the case title, text, debt description, etc. to cover all aspects of the case data. In the pattern rules, use the introduced fuzzy matching operators to allow a certain difference between the keywords and the actual text. For example, for the debt amount, the pattern rule "owed amount [≈:]\d+(.\d+)? yuan" can be designed, indicating to match the number and unit "yuan" following "owed amount", and allowing a certain difference between "owed amount" and the actual text. Traverse the case data and apply the designed regular expression pattern rules to each case text for matching. For each pattern rule, calculate the edit distance between the keywords in it and the corresponding part in the case text. If the edit distance is less than the preset threshold (such as 2 or 3), it is considered a successful match, and the corresponding key information is extracted. For the successfully matched key information, such as debt amount, debt term, debtor name, etc., save them as case features. For each case text, the key information extracted through pattern matching forms the feature set of the case. The case feature set can be represented by data structures such as dictionaries and lists for subsequent processing and analysis. For the key information that cannot be successfully matched, default values or special marks can be used for filling to ensure the integrity of the feature set. By introducing regular expressions with fuzzy matching operators, the flexibility and robustness of text parsing can be improved. It allows a certain difference between the keywords and the actual text and can handle situations such as typos, synonyms, abbreviations, etc. in the case data.

[0042] S22, use the named entity recognition method to perform semantic annotation on the case features, and use a pre-trained named entity recognition model to perform semantic annotation on the text in the case feature set. The named entity recognition model can be trained based on machine learning algorithms such as conditional random fields (CRF) and recurrent neural networks (RNN). Through named entity recognition, identify the entity categories in the case features, such as debt amount, debt term, debtor name, etc. Use the identified entity categories as labels to form a structured case feature representation with the corresponding case feature text.

[0043] S23. Construct a structured case attribute table according to the structured case features, organize and store the structured case features in the form of a relational model to form a case attribute table. The case attribute table is in the form of a two-dimensional table, with each row corresponding to a case and each column corresponding to a case attribute. The case attributes include case number, case title, debt amount, debt term, debtor name, creditor information, etc. For missing or incomplete attribute values, default values or special marks can be used for filling.

[0044] As Figure 3 shown, S3. Construct a dynamic knowledge graph according to the case attribute table. The dynamic knowledge graph is interconnected with an external legal knowledge base and introduces knowledge reasoning for dynamic expansion of knowledge, including: S31. Construct a static knowledge graph according to the case attribute table. The case attribute table contains the following attributes: case number, debtor name, debt amount, debt term, creditor name, etc. By analyzing the logical relationships between the attributes, the following association relationships can be extracted: There is an "owing" relationship between the debtor and the debt amount. There is a "term" relationship between the debtor and the debt term. There is a "lending" relationship between the debtor and the creditor. These association relationships reflect the semantic connections between the case attributes and are the basis for constructing the knowledge graph. Represent the semantic associations between the case attributes in the form of RDF (Resource Description Framework) triples. An RDF triple consists of a subject, a predicate, and an object, corresponding to the case attributes and their relationships respectively. Use a graph database (such as Neo4j) to store the RDF triples and construct a static knowledge graph. The graph database can efficiently represent and query complex association relationships.

[0045] S32. Link and map the static knowledge graph with an external legal knowledge base, select a suitable external legal knowledge base, such as a laws and regulations database, a case database, etc., to obtain relevant structured and unstructured legal knowledge. Through ontology alignment technology, semantically associate the case attribute concepts in the static knowledge graph with the relevant concepts in the external legal knowledge base. Ontology alignment can find semantically equivalent or similar concepts in the two knowledge bases based on indicators such as semantic similarity and structural similarity of the concepts.

[0046] Based on the results of ontology alignment, establish links between the static knowledge graph and the external legal knowledge base. The links can be achieved in the following ways: URI link: Assign unique URIs to the case attribute concepts in the static knowledge graph and the relevant concepts in the external legal knowledge base, and establish links between them through URIs. Alignment relation triple: Use RDF triples to represent the alignment relations between concepts, such as <Concept A, owl:sameAs, Concept B>, indicating that Concept A and Concept B are semantically equivalent. Link attribute: Add link attributes to the case attribute concepts in the static knowledge graph, such as "reference law articles", "relevant case precedents", etc., and the attribute values point to the relevant concepts in the external legal knowledge base. Through the establishment of links, the knowledge interoperability and sharing between the static knowledge graph and the external legal knowledge base are realized, and the knowledge coverage of the knowledge graph is extended.

[0047] Link mapping between the static knowledge graph and the external legal knowledge base realizes the interoperability and sharing of knowledge. First, select a suitable external legal knowledge base, such as a legal regulations database, a case database, etc., to obtain relevant structured and unstructured legal knowledge. Then, through ontology alignment technology, based on indicators such as semantic similarity and structural similarity, find the semantic associations between the case attribute concepts in the static knowledge graph and the relevant concepts in the external legal knowledge base. Finally, construct links between the static knowledge graph and the external legal knowledge base, and through methods such as URI links, alignment relation triples, and link attributes, realize the interoperability and sharing of knowledge.

[0048] S33. On the basis of link mapping, use the knowledge fusion method to expand the static knowledge graph with the external legal knowledge base. First, select a suitable external legal knowledge base, such as a legal regulations database, a case database, etc. Perform text preprocessing on the legal concepts (such as legal terms, law articles, etc.) in the legal knowledge base, such as word segmentation, stop word removal, etc. Use pre-trained semantic representation models, such as Word2Vec or BERT, to transform legal concepts into low-dimensional dense semantic embedding vectors. The semantic embedding vectors can capture the semantic similarity between legal concepts, and similar concepts are closer in the vector space. Use the RDF triple information in the static knowledge graph to extract the context information of the case attribute concepts. The context information includes structural information such as the neighbor nodes and connection relationships of the concepts in the knowledge graph. Use technologies such as graph neural networks (such as Graph Convolutional Networks, GCN) to learn the context representation of the case attribute concepts. By aggregating the neighbor information of the concepts, the graph neural network can generate concept representation vectors that consider the graph structure.

[0049] Secondly, using the semantic embedding vectors of legal concepts and the context vectors of case attribute concepts, calculate the similarity between them. Similarity can be measured using metrics such as cosine similarity, Euclidean distance, etc. Construct a semantic similarity matrix, where the rows of the matrix represent legal concepts, the columns represent case attribute concepts, and the matrix elements represent the similarity between them.

[0050] Then, transform the concept matching problem into a bipartite graph matching problem, with legal concepts and case attribute concepts as the two vertex sets of the bipartite graph respectively. Use heuristic algorithms such as the Hungarian algorithm to find the optimal matching between legal concepts and case attribute concepts based on the semantic similarity matrix. The Hungarian algorithm continuously adjusts the matching relationship and finally finds a matching scheme with the maximum total similarity.

[0051] Finally, according to the results of concept matching, fuse the matching legal concepts and their associated knowledge in the external legal knowledge base into the static knowledge graph. The fusion process can be achieved by adding new nodes, relationships, etc. to the knowledge graph. For example, if the legal concept "debt offset" matches the case attribute concept "debt amount", a new node "debt offset" can be added to the knowledge graph and an association relationship can be established with "debt amount". Through knowledge fusion, the static knowledge graph is expanded to include richer legal knowledge. The static knowledge graph fused with external legal knowledge is called a dynamic knowledge graph. The dynamic knowledge graph integrates the knowledge of the case attribute table and the external legal knowledge base, with a wider knowledge coverage and deeper semantic associations.

[0052] As Figure 4 shown in S4, according to the dynamic knowledge graph, design a similarity calculation function that includes legal relationship constraints and case feature ontology, including: In S41, we use natural language processing techniques, such as the word vector (Word Embedding) model, to semantically model the case attribute concepts in the dynamic knowledge graph. By calculating the semantic similarity between concepts, aggregate semantically related attribute concepts into higher-level case feature concepts. Clustering algorithms such as k-means can be used for concept clustering. Then, use ontology description languages (such as OWL, RDF) to define the association constraints between case feature concepts. For example, use owl:Class to define concept categories, use owl:ObjectProperty to define relationships between concepts, and use owl:Restriction to define association constraints. An example constraint is: "The concept of repayment method depends on the concept of debt amount, and the debt amount needs to be within a certain range." In this way, the case feature ontology is formed.

[0053] In S42, legal relation constraints are extracted by using the mapping relationship between legal concepts and case attributes obtained when constructing a dynamic knowledge graph in S3. For example, if the case attribute "loan amount" is mapped to the legal concept "private small-amount lending", and there is such a legal provision in the legal knowledge base: "The amount of private lending generally shall not exceed 500,000 yuan", then the constraint "loan amount <= 500,000 yuan" can be extracted. Constraint extraction can be achieved through natural language processing technologies such as template matching and dependency analysis.

[0054] Proceed to S43. First, use the concept semantic information defined in the case feature ontology to construct a semantic vector representation of the concept for calculating concept semantic similarity. Common semantic vector learning models include Word2Vec, Glove, etc. Then design a concept semantic similarity sub-function. Common similarity metrics include cosine similarity, Jaccard similarity, etc. For the constraint similarity sub-function, it can be designed to measure the degree to which the attribute values of two cases satisfy the association constraints extracted from the ontology and the external knowledge base. For example, calculate the degree to which the loan amount of a case satisfies the constraint "loan amount <= 500,000 yuan" at the same time. The degree of satisfaction can be quantified and represented in the range of 0-1.

[0055] In S44, a weighted sum of the semantic similarity and constraint similarity sub-functions is performed to form a case similarity function. That is: CaseSim(C1, C2) = w1 * SemanticSim(C1, C2) + w2 * ConstraintSim(C1, C2); where w1 and w2 are the weight coefficients of the two sub-functions and can be adjusted according to domain knowledge. SemanticSim and ConstraintSim are the semantic similarity sub-function and constraint similarity sub-function constructed in S43 respectively. The CaseSim function can be used for case clustering analysis in S5.

[0056] In the similarity calculation of this application, by setting a concept semantic similarity calculation sub-function and a constraint relation similarity sub-function, the semantic similarity of case feature concepts and the legal constraint relevance are measured respectively. Finally, the results of the two sub-functions are fused in a weighted combination manner to obtain a similarity calculation function that comprehensively considers semantic similarity and legal constraints. This similarity calculation function makes full use of the semantic information of the knowledge graph and the ontology, and at the same time considers the influence of legal factors, providing the interpretability of reasoning.

[0057] S5, embed the similarity calculation function into the hierarchical clustering algorithm, and use the embedded hierarchical clustering algorithm to cluster and group cases, including: S51, use the similarity calculation function to calculate the similarity value of each pair of cases in the dynamic knowledge graph, and construct a similarity matrix based on the similarity value; for each pair of cases in the dynamic knowledge graph, use the designed similarity calculation function to calculate the similarity value between them. The similarity calculation function comprehensively considers the semantic similarity of the case feature concepts and the relevance of the legal relationship constraints. The calculated similarity values ​​are organized into a similarity matrix, the rows and columns of the matrix correspond to the cases respectively, and the matrix elements represent the similarity values ​​of the corresponding case pairs.

[0058] S52, according to the similarity matrix, calculate the distance measurement between the cases; according to the distance measurement, use the hierarchical clustering algorithm to aggregate or split the cases layer by layer; including: converting the similarity matrix into a distance matrix, and the distance measurement can use Euclidean distance, Manhattan distance, etc. Use the hierarchical clustering algorithm to aggregate or split the cases layer by layer, and the commonly used hierarchical clustering algorithms include AGNES (AGglomerative NESting) and DIANA (DIvisive ANAlysis).

[0059] Taking the DIANA algorithm as an example, all cases are used as initial clusters, and each case is a separate cluster. The average distance between each case in the initial cluster and the cases in the corresponding neighborhood is calculated using a density-based anomaly detection algorithm (such as LOF). Cases with an average distance greater than a threshold are identified as noise points, and the noise points in the initial cluster are removed. The diameter of the cases in the initial cluster after removing the noise points (i.e., the maximum distance within the cluster) is calculated, and the cluster with the largest diameter is selected for splitting.

[0060] Within the selected cluster, by calculating the local density ρ[i] and distance δ[i] of the cases, select the clusters whose distance is greater than the threshold δ min And the local density is greater than the threshold ρ min The case of is taken as the seed case. min =μ ρ -β×σ ρ , where μ ρ and σ ρ are the mean and standard deviation of the local density of all cases in the cluster, and β is the coefficient of the control threshold, which can be adjusted according to the density distribution characteristics of the cluster. min The calculation formula is: min =μ δ +γ×σ δ , where μ δ and σ δ are the mean and standard deviation of the distances of all cases in the cluster, respectively; γ is the coefficient of the control threshold, which can be adjusted according to the distance distribution characteristics of the cluster.

[0061] The calculation of the local density ρ[i] takes into account the distances between case i and the cases within its adaptive neighborhood, reflecting the degree of closeness of the cases within the local area. The calculation formula for the local density ρ[i] is as follows: where dist(i,j) represents the distance between case i and case j, n is the total number of cases within the cluster, and d c [i] is the radius of the adaptive neighborhood of case i;

[0062] d c [i] = μ i + α × σ i where μ i and σ i are respectively the mean and standard deviation of the distances between case i and other cases, and α is a coefficient for controlling the neighborhood range, which can be adjusted according to the compactness of the cluster.

[0063] The calculation of the distance δ[i] takes into account the distance between case i and the nearest case with a local density greater than it, reflecting the distance between the case and the high-density area. The calculation formula for the distance δ[i] is as follows: If the local density of case i is the largest within the cluster, then that is, the distance between case i and the case with the farthest distance within the cluster.

[0064] The local density threshold ρ min and the distance threshold δ min are calculated based on the mean and standard deviation of the local density and distance of the cases within the cluster, and can be adjusted according to the density and distance distribution characteristics of the cluster. Centered on the seed cases, the cases are divided into corresponding sub-clusters according to the distances between the cases and the seed cases. The sub-clusters are recursively split until each cluster contains only one case or reaches the preset cluster quantity threshold.

[0065] S53. For the clustering operations that do not satisfy the legal relationship constraints during the aggregation or splitting of cases, impose penalties for violating the constraints. Specifically, according to the case feature ontology and legal relationship constraints, determine whether the clustering operation meets the constraint conditions. For example, if two cases have mutually exclusive legal relationships (such as the plaintiff and the defendant), they should not be aggregated into the same cluster. Another example is that if the legal attribute of a certain case is incompatible with the legal attribute of the cluster (such as the case belongs to a criminal case while the cluster belongs to a civil case), then this case should not be divided into this cluster.

[0066] If the clustering operation violates the legal relationship constraints, punish the operation. The purpose of the punishment is to prompt the clustering algorithm to avoid generating clustering results that do not meet the constraint conditions. The punishment can be achieved in the following two ways: Increase the cost of the clustering operation: In the hierarchical clustering algorithm, the cost of the clustering operation is usually determined by the distance metric between clusters. For clustering operations that violate the constraints, the cost value can be artificially increased so that they are excluded or postponed during the clustering process. For example, the cost value of a clustering operation that violates the constraints can be multiplied by a penalty coefficient greater than 1 to significantly increase its cost. Reduce the evaluation index of the clustering: Clustering algorithms usually use certain evaluation indexes to measure the quality of the clustering results, such as the silhouette coefficient, Davies-Bouldin index, etc. For clustering results that contain clustering operations that violate the constraints, the corresponding evaluation index value can be reduced. For example, when calculating the clustering evaluation index, a penalty term can be introduced for clustering operations that violate the constraints to reduce their contribution to the evaluation index. During the process of case aggregation or splitting, perform constraint violation punishment on clustering operations that do not meet the legal relationship constraints. Through constraint satisfaction checking, identify clustering operations that violate the constraints and punish them. The punishment can be achieved by increasing the cost of the clustering operation or reducing the evaluation index of the clustering, prompting the clustering algorithm to avoid generating clustering results that do not meet the constraint conditions. Through the constraint violation punishment mechanism, it is ensured that the case clustering results comply with the legal relationship constraints, improving the compliance and interpretability of the clustering results. This is of great significance for applying case clustering and grouping in the legal field, helping to generate case grouping results that conform to legal logic and constraint conditions

[0067] S54. Generate a case clustering tree: Organize the clustering results generated by the hierarchical clustering algorithm into a case clustering tree. Each node of the clustering tree corresponds to a case sub-cluster, the root node corresponds to the initial cluster, and the leaf node corresponds to a single case. The clustering tree shows the hierarchical structure and clustering process of the case clustering, facilitating the analysis and understanding of the case grouping situation.

[0068] S6. According to the grouping results, use knowledge reasoning to generate case-filing materials. The case-filing materials include a complaint and an evidence list. Specifically, extract the relevant attribute information of the case from the dynamic knowledge graph, such as the case type, parties, cause of action, case description, etc. Through knowledge reasoning, extract key elements from the case attribute information, such as litigation requests, factual basis, legal basis, etc. A complaint is a legal document filed by the plaintiff with the court, containing content such as litigation requests, factual basis, legal basis, etc.

[0069] Using natural language generation technology, according to the case attribute information and key elements, automatically generate each part of the complaint: Litigation request part: Generate standardized litigation request statements according to the case type and cause of action, such as "Request the court to order the defendant to compensate the plaintiff for economic losses", etc. Factual basis part: Generate a statement of the case facts according to the case description and relevant evidence, such as "The plaintiff borrowed money from the defendant on a certain date, and the defendant has not returned it yet", etc. Legal basis part: Quote relevant legal provisions according to the case type and cause of action. Combine the generated content of each part into a complete complaint, and perform formatting and editing to ensure the standardization and readability of the complaint.

[0070] The evidence list is an enumeration and description of the evidence materials involved in the case, including the name of the evidence, the content to be proved, the source of the evidence, etc. Extract the relevant evidence information of this case from the dynamic knowledge graph, such as contracts, invoices, bank statements, etc. For each piece of evidence, generate a standardized evidence description according to its type and attributes, such as "A copy of the loan contract, proving the lending relationship between the plaintiff and the defendant", etc. Organize the generated evidence descriptions into an evidence list in a certain order and format, and correspond it to the complaint. According to the legal knowledge base and the case feature ontology, reason and verify the content of the complaint and the evidence list. Check whether the litigation request in the complaint conforms to the case type and cause of action, whether the factual basis is consistent with the case description, whether the legal basis is appropriate, etc. Check whether the evidence in the evidence list is relevant to the case facts, whether the evidence description is accurate, whether the source of the evidence is legal, etc. Mark and prompt the non-compliant content, and give modification suggestions to ensure the compliance of the complaint and the evidence list. Output the complaint and the evidence list according to the specified format and requirements to form a complete case-filing material. Associate the case-filing material with the case attribute information and store it in the case management system or knowledge base for subsequent query and invocation.

Claims

1. A method for handling small-amount non-performing asset litigation cases, characterized in that, Including: S1. Obtain data of small - amount non - performing asset litigation cases; S2. Pre - process the case data to obtain a structured case attribute table, where the case attribute table contains debt elements and party information; S3. Construct a dynamic knowledge graph according to the case attribute table. The dynamic knowledge graph is interconnected with an external legal knowledge base and introduces knowledge reasoning for dynamic expansion of knowledge; S4. Design a similarity calculation function that includes legal - relationship constraints and case - feature ontology according to the dynamic knowledge graph; S5. Embed the similarity calculation function into the hierarchical clustering algorithm, and use the embedded hierarchical clustering algorithm to group the cases; S6. Generate case - filing materials using knowledge reasoning according to the grouping results. The case - filing materials include a statement of claim and an evidence list.

2. The method for handling small - amount non - performing asset litigation cases according to claim 1, wherein: S2. Pre - process the case data to obtain a structured case attribute table, including: S21. Use a regular expression with a fuzzy matching operator to perform text parsing on the case data. Through predefined pattern rules, extract case features in the case data, where the case features include debt information; the fuzzy matching operator calculates the similarity between the matching string and the pattern rule using the edit - distance algorithm; S22. Use a named - entity recognition method to perform semantic annotation on the case features, identify the entity categories of the case features, and obtain structured case features; the entity categories include debt amount, debt term, and debtor information; S23. Construct a structured case attribute table according to the structured case features. The case attribute table organizes case features using a relational model, with each row corresponding to a case and each column corresponding to a case attribute.

3. The method for handling small - amount non - performing asset litigation cases according to claim 1, wherein: S3. Construct a dynamic knowledge graph according to the case attribute table, including: S31. According to the case attribute table, extract the association relationships between case attributes and construct a static knowledge graph based on a graph database; the static knowledge graph uses RDF triples to represent the semantic associations between case attributes and reflects case attributes and relationships in the form of subject, predicate, and object; S32. Link and map the static knowledge graph with an external legal knowledge base, and realize the semantic association of relevant concepts in the static knowledge graph and the external legal knowledge base through ontology alignment; the external legal knowledge base contains structured and unstructured knowledge of relevant case precedents; S33. On the basis of the link mapping, use a knowledge - fusion method based on ontology alignment to expand the static knowledge graph using the external legal knowledge base to generate a dynamic knowledge graph.

4. The method for handling small - amount non - performing asset litigation cases according to claim 3, wherein: S33. Generate a dynamic knowledge graph, including: Use a pre - trained semantic representation model to perform semantic embedding on legal concepts in the external legal knowledge base to obtain semantic embedding vectors of legal concepts; the legal concepts include legal terms and legal articles; Use the RDF triple information in the static knowledge graph to perform context representation learning on the case attribute concepts in the static knowledge graph, and obtain the context vectors of the case attribute concepts; Calculate the similarity between the semantic embedding vectors of legal concepts and the context vectors of case attribute concepts, and construct a semantic similarity matrix between legal concepts and case attribute concepts; According to the semantic similarity matrix, use a heuristic algorithm for concept matching to obtain the optimal matching result between legal concepts and case attribute concepts; According to the optimal matching result, fuse the matching legal knowledge in the external legal knowledge base into the static knowledge graph to expand the association relationship between case attributes and generate a dynamic knowledge graph containing legal knowledge.

5. The method for handling small-amount non-performing asset litigation cases according to claim 4, wherein: The heuristic algorithm uses the Hungarian algorithm.

6. The method for handling small-amount non-performing asset litigation cases according to claim 5, wherein: S4. According to the dynamic knowledge graph, design a similarity calculation function including legal relationship constraints and case feature ontology, including: S41. Organize the case attribute concepts in the dynamic knowledge graph into concept categories according to semantic relevance to form case feature concepts; and use ontology language to describe the association constraints between case feature concepts to form a case feature ontology; S42. Convert the matching relationship between legal concepts in the external legal knowledge base and case attribute concepts in the dynamic knowledge graph into legal relationship constraints; S43. According to the semantic description of case feature concepts, set a concept semantic similarity calculation sub-function to measure the semantic similarity between case feature concepts; According to the legal relationship constraints, set a constraint relationship similarity sub-function to evaluate the relevance of case feature concepts under legal constraints; S44. Perform weighted combination on the semantic similarity calculation sub-function and the constraint relationship similarity sub-function to obtain a similarity calculation function.

7. The method for handling small-amount non-performing asset litigation cases according to any one of claims 1 to 5, wherein: S5. Embed the similarity calculation function into the hierarchical clustering algorithm, and use the embedded hierarchical clustering algorithm to perform case clustering grouping, including: S51. Use the similarity calculation function to calculate the similarity values of each pair of cases in the dynamic knowledge graph, and construct a similarity matrix according to the similarity values; S52. According to the similarity matrix, calculate the distance metric between cases; use the hierarchical clustering algorithm to perform layer-by-layer aggregation or splitting of cases according to the distance metric; S53. Perform a violation constraint penalty on the clustering operations that do not meet the legal relationship constraints during the case aggregation or splitting; S54. Generate a case clustering tree, where each node of the clustering tree corresponds to a case sub-cluster, and the leaf node corresponds to a single case.

8. The method for handling small-amount non-performing asset litigation cases according to claim 7, wherein: S52. Use the hierarchical clustering algorithm to perform layer-by-layer aggregation or splitting of cases according to the distance metric, including: Take all cases as the initial clustering; Calculate the average distance between each case in the initial cluster and the cases in its corresponding neighborhood using a density-based outlier detection algorithm, identify the cases with an average distance greater than the threshold as noise points, and remove the noise points from the initial cluster; Calculate the diameter of the cases in the initial cluster after removing the noise points as the maximum distance within the cluster; Select the cluster with the largest diameter for splitting. By calculating the local density ρ[i] and distance δ[i] of the cases within the cluster, select the cases where the distance is greater than the threshold δ min and the local density is greater than the threshold ρ min as the seed cases; Centering on the seed case, divide the cases into corresponding sub-clusters according to the distance between the cases and the seed case; Recursively split the sub-clusters until each cluster contains only one case or reaches the preset cluster quantity threshold.

9. The method for processing small-amount non-performing asset litigation cases according to claim 8, characterized in that: Calculate the local density ρ[i] and distance δ[i] of the cases within the cluster, including: The calculation formula for the local density ρ[i] is: Among them, dist(i, j) represents the distance between case i and case j, n is the total number of cases within the cluster, and d c [i] is the adaptive neighborhood radius of case i; d c [i] = μ i + α × σ i Among them, μ i and σ i are the mean and standard deviation of the distances between case i and other cases respectively, and α is the coefficient for controlling the neighborhood range; The calculation formula for the distance δ[i] is: If the local density of case i is the largest within the cluster, then That is, the distance between case i and the case with the farthest distance within the cluster; Local density threshold ρ min The calculation formula is as follows: ρ min = μ ρ - β × σ ρ Among them, μ ρ and σ ρ are the mean and standard deviation of the local densities of all cases within the cluster respectively, and β is the coefficient for controlling the threshold; Distance threshold δ min The calculation formula is as follows: δ min = μ δ + γ × σ δ Among them, μ δ and σ δ are the mean and standard deviation of the distances of all cases within the cluster respectively, and γ is the coefficient of the control threshold.

10. A system for processing small-amount non-performing asset litigation cases, characterized in that it includes: At least one processing unit; configured to execute instructions to implement the method for processing small-amount non-performing asset litigation cases according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Case retrieval method, device, equipment and storage medium based on knowledge graph

    CN111241241B

Cited By

  • Intelligent litigation case identification method and system based on RAG and knowledge graph

    CN120542404A