Intelligent legal text analysis method based on legal concept pedigree
By constructing a knowledge graph of legal concept genealogy and integrating natural language processing technology, the shortcomings of the existing legal text analytical system in dealing with complex legal concepts and logical relationships are solved, and efficient, accurate analysis of legal texts and support for legal reasoning are achieved.
Patent Information
- Application Number
- CN202510053560.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The existing legal text analytic system has shortcomings in understanding and handling complex legal concepts and logical relationships, and it is difficult to adapt to the frequent changes in legal provisions and the diversity of text scenarios, resulting in poor performance in the face of new legal problems.
By integrating advanced data processing technology and natural language processing technology, we can build a knowledge graph of legal concept genealogy, identify key legal concepts and their logical relationships in legal texts, and provide legal provision matching and case search services through intelligent algorithms, and continuously optimize them in combination with user feedback.
It realizes efficient and accurate analysis of legal texts, improves the accuracy of legal concept recognition, deeply explores the logical relationship between legal articles, provides strong support for legal reasoning, and maintains the timeliness and accuracy of the system.
Smart Images

Figure CN119990290A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text processing, and in particular relates to a legal text intelligent parsing method based on the legal concept pedigree. Background Art
[0002] In the legal field, with the surge in the number of legal provisions and the complexity of legal practice, legal workers are facing unprecedented challenges. Traditional legal text analysis methods rely on manual reading and research, which is time-consuming and prone to deviations. Although the development of natural language processing (NLP) technology has brought new possibilities for legal text analysis, existing legal text parsing systems still have deficiencies in understanding and processing complex legal concepts and logical relationships. In particular, for the concept spectrum unique to the legal system, existing systems can often only perform simple association analysis and lack deep-level concept construction and application. In addition, these systems often rely on predefined rules and templates when processing legal texts, which makes it difficult to adapt to the frequent changes in legal provisions and the diversity of text scenarios, resulting in poor performance when facing new legal issues.
[0003] Existing technical solutions are not capable of processing complex legal texts, especially in identifying legal concepts and parsing the logical relationship of legal provisions, and cannot meet the growing requirements for accuracy and efficiency of legal services. In addition, these systems have defects in feedback mechanisms and continuous optimization, and fail to effectively use user feedback and actual usage to improve system performance, affecting the reliability and practicality of legal analysis systems. Summary of the invention
[0004] The purpose of the present invention is to provide a method for intelligent parsing of legal texts based on the legal concept genealogy. By integrating advanced data processing technology and natural language processing technology, a legal concept genealogy knowledge graph is constructed to achieve efficient and accurate parsing of legal texts to solve the problems raised in the above background technology.
[0005] To achieve the above-mentioned purpose, the present invention adopts the following technical solution: a legal text intelligent parsing method based on the legal concept pedigree, the method comprising the following steps:
[0006] a) preprocessing the input legal text, including but not limited to text cleaning, word segmentation and part-of-speech tagging, to obtain preprocessed text data; b) based on step (a), using natural language processing technology to identify key legal concepts in the preprocessed text data; c) constructing a legal concept genealogy knowledge graph based on the key legal concepts identified in step (b), wherein the graph at least includes the name, definition, translation and its upper and lower relationships of the concept; d) based on the legal concept genealogy knowledge graph constructed in step (c), deeply analyzing the preprocessed text data to identify the logical relationship in the legal text; e) based on the analysis result of step (d), providing users with legal text matching and case retrieval services through intelligent algorithms; f) in combination with the services provided in step (e), receiving user feedback and continuously optimizing and updating the legal concept genealogy knowledge graph.
[0007] Preferably, the step (a) includes the following sub-steps: a1) using a regular expression tool to remove irrelevant characters and punctuation marks in the text; a2) using the jieba word segmentation tool to segment the processed text into meaningful vocabulary units to form text data after word segmentation; a3) using the part-of-speech tagging function of jieba to perform part-of-speech tagging on the text data after word segmentation, and identify the grammatical role of each word; a4) performing semantic analysis, determining the dependency relationship between words through the dependency syntax analysis function of HanLP, and identifying key entities in the text.
[0008] Preferably, the step (a) further comprises: a5) performing named entity recognition on the preprocessed text data using natural language processing technology to identify important entities involved in the text; a6) filtering and normalizing the results of the named entity recognition.
[0009] Preferably, the step (b) includes the following sub-steps: b1) installing the PyTorch or TensorFlow deep learning framework and the Transformers library on the computer; b2) loading the LegalBERT model of HuggingFace through the Transformers library, which is a pre-trained model optimized for legal texts; b3) obtaining the text data preprocessed by step (a), encoding it according to the requirements of the LegalBERT model, and converting it into tokenids and attentionmasks that can be processed by the model; b4) inputting the encoded data into the LegalBERT model, using the model to classify each token, and identifying and extracting the legal concept categories in the text.
[0010] Preferably, the step (b) also includes: b5) mapping the identified key legal concepts to corresponding nodes in the legal concept genealogy knowledge graph constructed in step (c) to enhance the correlation between concepts; b6) semantically expanding the identified legal concepts according to the context, identifying related superordinate concepts or subordinate concepts, and enriching the parsing level of the legal text.
[0011] Preferably, the step (c) includes the following sub-steps: c1) designing a graph structure, defining legal concept nodes, each node containing the name, definition, and translation attributes of the concept, and creating term nodes to contain term translations and related explanations; c2) defining the relationship between legal concepts, using the IS_A relationship to represent the hierarchical relationship of concepts, and the HAS_TYPE relationship to represent the genus-species relationship of concepts; c3) importing the compiled legal terms and their definitions, translations, and logical relationship data into the graph database, and creating corresponding nodes and relationships; c4) adding the legal concepts identified in step (b) as nodes to the graph through batch operations or scripting, and connecting them using defined relationships to form a concept genealogy.
[0012] Preferably, the step (d) includes the following sub-steps: d1) analyzing the pre-processed text data based on the legal concept genealogy knowledge graph constructed in step (c) to identify the legal concepts in the text and their logical relationships; d2) using the concept nodes and their relationships in the graph to identify the reference, supplement or conflicting logical relationships of the legal provisions in the text; d3) applying the Cypher query language to perform complex legal concept searches and path queries to extract logical associations in the legal text; d4) post-processing the query results to convert the legal concepts and their logical relationships output by the model into an easily understandable form, laying the foundation for providing legal provision matching and case retrieval services in the subsequent step (e).
[0013] Preferably, step (d) also includes: d5) using a combination of rule-based methods and machine learning technology to identify entities in the legal text; d6) automatically retrieving and displaying the relationship network between entities in the legal text through intelligent retrieval technology to help users more intuitively understand the intrinsic connection between legal provisions and cases.
[0014] Preferably, step (e) includes the following sub-steps: e1) using the legal concepts and their logical relationships parsed in step (d), automatically retrieving legal provisions related to the case facts or legal issues input by the user through an intelligent algorithm; e2) intelligently matching relevant legal clauses based on the specific content of the user's query, and providing detailed explanations and applicable scenarios of the clauses; e3) using the cosine similarity calculation method, vectorizing the case text and related cases based on the LegalBERT model, calculating the cosine similarity between these vectors, and obtaining a text similarity score; e4) sorting the cases according to the similarity score, giving priority to displaying cases with a higher similarity to the current legal issues and factual situations, and providing users with corresponding legal basis and support.
[0015] Preferably, step (f) includes the following sub-steps: f1) designing a user feedback mechanism to allow users to evaluate the effectiveness of the legal text matching and case retrieval results provided in step (e); f2) collecting feedback information submitted by users, including but not limited to user satisfaction scores for the system recommendation results, correction suggestions, and new cases or legal clauses manually added by users; f3) analyzing the collected feedback information, and adjusting the node attributes or relationships in the legal concept genealogy knowledge graph based on user feedback to optimize the graph structure; f4) regularly crawling the latest legal concepts and provisions from the data source, and using automated scripts to update the nodes and relationships in the Neo4j graph to ensure the timeliness and accuracy of the knowledge graph.
[0016] Technical effects and advantages of the present invention: Compared with the prior art, the present invention proposes a legal text intelligent parsing method based on the legal concept pedigree, which has the following advantages:
[0017] The present invention integrates advanced data processing technology and natural language processing technology to construct a knowledge graph of legal concept genealogy, thereby achieving efficient and accurate analysis of legal texts. This method not only improves the accuracy of legal concept recognition, but also deeply explores the logical relationship between legal provisions by constructing a knowledge graph of legal concept genealogy, providing strong support for legal reasoning. In addition, through continuous optimization and feedback mechanisms, the present invention can timely adjust and optimize model parameters and algorithm logic based on user feedback, ensuring that the system can keep up with the pace of legal development and maintain a high degree of timeliness and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flow chart of the intelligent analysis method of legal text based on the legal concept pedigree of the present invention;
[0019] Figure 2 This is a schematic diagram of case relationships in this embodiment;
[0020] Figure 3This is a topological diagram for explaining the case results according to the standard implementation in this embodiment;
[0021] Figure 4 Flowchart for generating conclusions in this embodiment. DETAILED DESCRIPTION
[0022] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0023] The present invention provides Figure 1 A legal text intelligent parsing method based on legal concept pedigree is shown, the method comprising the following steps:
[0024] a) Preprocessing the input legal text, including but not limited to text cleaning, word segmentation and part-of-speech tagging, to obtain preprocessed text data; specifically including the following sub-steps:
[0025] a1) Use regular expression tools to remove irrelevant characters and punctuation marks in the text; a2) Use the jieba word segmentation tool to segment the processed text into meaningful vocabulary units to form text data after word segmentation; a3) Use the jieba part-of-speech tagging function to tag the text data after word segmentation and identify the grammatical role of each word; a4) Perform semantic analysis, determine the dependency relationship between words through the dependency syntax analysis function of HanLP, and identify the key entities in the text. a5) Use natural language processing technology to perform named entity recognition on the pre-processed text data and identify the important entities involved in the text; a6) Filter and normalize the results of named entity recognition.
[0026] By carefully preprocessing the input legal text, including but not limited to text cleaning, word segmentation and part-of-speech tagging, and further performing semantic analysis and named entity recognition, this method can effectively improve the quality and efficiency of legal text parsing. Specifically, regular expressions are used to remove irrelevant characters and punctuation marks to ensure the cleanliness of the text; the jieba word segmentation tool is used to segment the text into meaningful vocabulary units for subsequent processing; the jieba part-of-speech tagging function is used to identify the grammatical role of the vocabulary, which enhances the structuring of the text; the dependency relationship between vocabulary is determined through the dependency syntactic analysis function of HanLP, and the key entities in the text are identified, so that the system can better understand the deep meaning of the text; further, through the named entity recognition technology, the system can identify the important entities involved in the text, filter and normalize them, reduce noise data, and improve the accuracy and reliability of subsequent processing steps. Overall, these preprocessing steps provide a solid foundation for the subsequent legal concept recognition and extraction, the construction of the legal concept genealogy knowledge graph, and intelligent parsing and decision support, ensuring the efficiency and accuracy of the entire legal text intelligent parsing process.
[0027] b) Based on step (a), using natural language processing technology to identify key legal concepts in the pre-processed text data; specifically comprising the following sub-steps:
[0028] b1) Install the PyTorch or TensorFlow deep learning framework and Transformers library on the computer; b2) Load the LegalBERT model of HuggingFace through the Transformers library, which is a pre-trained model optimized for legal texts; b3) Obtain the text data preprocessed by step (a), encode it according to the requirements of the LegalBERT model, and convert it into tokenids and attentionmasks that the model can process; b4) Input the encoded data into the LegalBERT model, use the model to classify each token, identify and extract the legal concept categories in the text. b5) Map the identified key legal concepts to the corresponding nodes in the legal concept genealogy knowledge graph constructed in step (c) to enhance the correlation between concepts; b6) Perform semantic expansion on the identified legal concepts according to the context, identify related superordinate or subordinate concepts, and enrich the parsing level of the legal text.
[0029] By utilizing deep learning technology, especially the LegalBERT model optimized for legal texts, this method can accurately identify key legal concepts in preprocessed legal text data. Specifically, by installing the necessary deep learning framework and model library in the computing environment, the LegalBERT model is ensured to run smoothly. After loading the LegalBERT model, it is used to encode and classify the preprocessed text data, and legal concepts in the text, such as legal clauses, case names, etc., can be efficiently identified. Mapping the identified legal concepts to the legal concept genealogy knowledge graph not only enhances the correlation between concepts, but also identifies related superordinate or subordinate concepts through the semantic extension of the context, greatly enriching the parsing level of the legal text. In this way, this method not only improves the accuracy of legal concept recognition, but also provides a solid foundation for the subsequent deep parsing of legal texts and identification of logical relationships, ensuring the efficiency and accuracy of the intelligent parsing process of legal texts.
[0030] c) constructing a legal concept genealogy knowledge graph based on the key legal concepts identified in step (b), wherein the graph at least includes the name, definition, translation and upper and lower relationships of the concept; specifically comprising the following sub-steps:
[0031] c1) Design the graph structure and define legal concept nodes, each of which contains the name, definition, and translation attributes of the concept, and create term nodes to contain term translations and related explanations; c2) Define the relationships between legal concepts, using the IS_A relationship to represent the hierarchical relationship of concepts and the HAS_TYPE relationship to represent the genus-species relationship of concepts; c3) Import the compiled legal terms and their definitions, translations, and logical relationship data into the graph database, and create corresponding nodes and relationships; c4) Add the legal concepts identified in step (b) as nodes to the graph through batch operations or scripting, and connect them using the defined relationships to form a concept genealogy.
[0032] By constructing a knowledge graph of legal concept genealogy, this method can systematically organize and manage key concepts in legal texts, thereby achieving a deeper understanding and analysis of legal texts. Specifically:
[0033] Structured knowledge representation: Designing the graph structure and defining the legal concept nodes and their attributes (such as name, definition, translation) helps to clearly represent each legal concept and ensure that each concept has a detailed description. This structured representation facilitates subsequent data query and analysis.
[0034] Rich semantic relationships: Define the relationships between legal concepts, such as using the IS_A relationship to represent the hierarchical relationship of concepts and the HAS_TYPE relationship to represent the genus-species relationship of concepts. This enables the graph to reflect the logical associations between legal concepts and enhances the semantic richness and hierarchy of the graph.
[0035] Efficient graph construction and updating: By importing the compiled legal terms and their definitions, translations, and logical relationship data into the graph database and creating corresponding nodes and relationships, this method achieves efficient construction of the legal concept genealogy knowledge graph. At the same time, through batch operations or scripting, the identified legal concepts are added to the graph and connected according to the defined relationships, forming a complete and dynamically updated concept genealogy, ensuring the real-time and accuracy of the graph.
[0036] Enhanced legal text parsing capabilities: The constructed legal concept genealogy knowledge graph can not only help the system understand the concepts and their interrelationships in legal texts more accurately, but also provide strong support for subsequent services such as legal text matching and case retrieval. With the rich information in the graph, the system can provide more accurate and comprehensive legal services.
[0037] In summary, by constructing a legal concept genealogy knowledge graph, this method significantly improves the accuracy and efficiency of intelligent parsing of legal texts and provides more powerful technical support for legal workers.
[0038] d) According to the legal concept pedigree knowledge graph constructed in step (c), based on the database of the constructed legal concept pedigree knowledge graph, the input case data is analyzed, the identified extracted legal concepts are linked to the knowledge graph, legal concept search and path query are performed, the logical associations in the legal text are refined, the pre-processed text data is deeply analyzed, and the logical relationships in the legal text are identified; specifically, the following sub-steps are included:
[0039] d1) Based on the legal concept genealogy knowledge graph constructed in step (c), analyze the pre-processed text data to identify the legal concepts in the text and their logical relationships; d2) Use the concept nodes and their relationships in the graph to identify the reference, supplement or conflicting logical relationships of legal provisions in the text; d3) Apply the Cypher query language to perform complex legal concept searches and path queries to extract logical associations in legal texts; d4) Post-process the query results to convert the legal concepts and their logical relationships output by the model into an easy-to-understand form, laying the foundation for providing legal text matching and case retrieval services in the subsequent step (e). d5) Use a combination of rule-based methods and machine learning technology to identify entities in legal texts; d6) Use intelligent retrieval technology to automatically retrieve and display the relationship network between entities in legal texts, helping users to more intuitively understand the internal connection between legal texts and cases.
[0040] By deeply analyzing the preprocessed text data based on the legal concept genealogy knowledge graph, this method can effectively identify the key legal concepts in the legal text and their logical relationships. Specifically:
[0041] Detailed legal concept analysis: Based on the legal concept genealogy knowledge graph constructed in step (c), this method can conduct in-depth analysis of the preprocessed text data and identify the legal concepts in the text and their logical relationships. This helps the system understand the connotation of the legal text more accurately and provides a solid foundation for subsequent services.
[0042] Identification of logical relationships: Using the concept nodes and their relationships in the graph, this method can identify logical relationships such as references, supplements or conflicts in legal provisions in the text. This ability is crucial for legal workers because it can help them discover and solve potential problems more quickly when dealing with complex legal matters.
[0043] Efficient query and retrieval: By applying the Cypher query language to perform complex legal concept searches and path queries, this method can extract logical associations in legal texts from the graph. Cypher is a query language designed specifically for graph databases that can efficiently handle complex query tasks, allowing the system to quickly and accurately provide the required information.
[0044] Post-processing and presentation of results: Post-process the query results to convert the legal concepts and logical relationships output by the model into an easy-to-understand form, providing users with intuitive legal clause matching and case retrieval services. This process not only improves the availability of information, but also enhances the user experience.
[0045] Entity recognition and intelligent retrieval: Combining rule-based methods and machine learning technology, this method can accurately identify entities in legal texts and automatically retrieve and display the relationship network between entities through intelligent retrieval technology. This approach not only improves the accuracy of entity recognition, but also enables users to more intuitively understand the inherent connection between legal provisions and cases, thereby better supporting legal decision-making.
[0046] In summary, this method greatly improves the accuracy and efficiency of legal text parsing by deeply analyzing legal texts, identifying key legal concepts and their logical relationships, and combining intelligent retrieval technology, providing legal workers with a powerful and easy-to-use tool.
[0047] e) Based on the analysis results of step (d), a case database is constructed to provide users with legal text matching and case retrieval services through intelligent algorithms; specifically, the following sub-steps are included:
[0048] e1) Using the legal concepts and their logical relationships parsed in step (d), the intelligent algorithm is used to automatically retrieve legal provisions related to the case facts or legal issues input by the user; e2) According to the specific content of the user's query, the relevant legal clauses are intelligently matched, and a detailed explanation and applicable scenarios of the clauses are given; e3) Using the cosine similarity calculation method, the case text and related cases are vectorized based on the LegalBERT model, and the cosine similarity between these vectors is calculated to obtain the text similarity score; e4) The cases are sorted according to the similarity score, and the cases with higher similarity to the current legal issues and factual situations are displayed first, providing users with corresponding legal basis and support.
[0049] Based on the parsing results of step (d), this method can provide users with efficient and accurate legal text matching and case retrieval services. The specific technical effects are as follows:
[0050] Intelligent matching of legal provisions: Using the legal concepts and their logical relationships analyzed in step (d), the intelligent algorithm automatically retrieves legal provisions related to the case facts or legal issues entered by the user. This enables users to quickly find the legal basis related to their cases, improving the efficiency and accuracy of legal work.
[0051] Detailed explanation and applicable scenarios: According to the specific content of the user's query, intelligently match the relevant legal terms and provide detailed explanations and applicable scenarios of the terms. This not only helps users understand the specific meaning of the legal provisions, but also provides guidance on actual application scenarios, enhancing the actual operability of the legal provisions.
[0052] Case retrieval based on cosine similarity: Using the cosine similarity calculation method, the case text and related cases are vectorized based on the LegalBERT model, and the cosine similarity between these vectors is calculated to obtain the similarity score of the text. This method can quantify the similarity between texts and provide a scientific basis for subsequent case sorting.
[0053] Case sorting and display: Cases are sorted according to the calculated similarity scores, with priority given to cases with high similarity to the current legal issues and factual situations. This sorting mechanism ensures that the most relevant cases are presented to users first, allowing users to quickly find the most valuable case materials.
[0054] Enhanced user experience: Through the above steps, the system can provide customized legal services, allowing users to easily find legal provisions and cases that are highly relevant to their cases, thereby greatly improving user satisfaction and the practicality of the system. In addition, this intelligent service model reduces the user's learning cost, making it easier for non-professionals to understand and apply legal knowledge.
[0055] In summary, this method significantly improves the accuracy and efficiency of legal text matching and case retrieval through the application of intelligent algorithms and deep learning technologies, providing users with an efficient, convenient and accurate legal information service platform.
[0056] f) Combined with the services provided in step (e), receive user feedback and continuously optimize and update the legal concept genealogy knowledge graph. Specifically, it includes the following sub-steps:
[0057] f1) Design a user feedback mechanism to allow users to evaluate the effectiveness of the legal text matching and case retrieval results provided in step (e); f2) Collect feedback information submitted by users, including but not limited to user satisfaction scores for the system recommendation results, correction suggestions, and new cases or legal clauses manually added by users; f3) Analyze the collected feedback information, and adjust the node attributes or relationships in the legal concept genealogy knowledge graph based on user feedback to optimize the graph structure; f4) Regularly capture the latest legal concepts and provisions from the data source, and use automated scripts to update the nodes and relationships in the Neo4j graph to ensure the timeliness and accuracy of the knowledge graph.
[0058] By combining the services provided in step (e), this method can receive user feedback and continuously optimize and update the legal concept genealogy knowledge graph. The specific technical effects are as follows:
[0059] Improve user participation: Design a user feedback mechanism to allow users to evaluate the effectiveness of the legal text matching and case retrieval results provided in step (e). This mechanism enhances user participation, enables users to directly participate in the system improvement process, and improves the system's interactivity and user experience.
[0060] Feedback information collection and analysis: Collect feedback information submitted by users, including but not limited to user satisfaction ratings of system recommendation results, correction suggestions, and new cases or legal clauses manually added by users. In this way, the system can timely understand the actual needs and problems encountered by users, providing valuable data support for subsequent optimization work.
[0061] Continuous optimization of the knowledge graph: Analyze the collected feedback information, and adjust the node attributes or relationships in the legal concept genealogy knowledge graph according to user feedback to optimize the graph structure. This means that the system can continuously adjust and improve the content of the knowledge graph according to the actual usage of users, making it more in line with actual needs and improving the accuracy and effectiveness of analysis and services.
[0062] Real-time data update: The latest legal concepts and provisions are captured from the data source regularly, and the nodes and relationships in the Neo4j graph are updated using automated scripts to ensure the timeliness and accuracy of the knowledge graph. This regular update mechanism ensures that the system can always provide the latest and most accurate legal information, avoiding information lags caused by changes in legal provisions.
[0063] Enhance system reliability and practicality: Through the establishment of a user feedback mechanism and continuous data updates, the system can continuously improve itself and enhance its reliability and practicality. User feedback information is promptly applied to the optimization of the knowledge graph, enabling the system to better serve users and meet legal needs in different scenarios.
[0064] In summary, by receiving user feedback and continuously optimizing the legal concept genealogy knowledge graph, this method not only improves the accuracy and timeliness of the system, but also enhances user satisfaction and the practicality of the system, providing legal workers with more efficient and reliable legal information retrieval and analysis tools.
[0065] On the other hand, this embodiment also provides a legal text intelligent parsing system based on the legal concept genealogy, which is mainly composed of four modules: data preprocessing module, legal concept identification and extraction module, legal concept genealogy knowledge graph construction and update module, and intelligent parsing and decision support module.
[0066] Data Preprocessing Module: Responsible for basic processing tasks such as cleaning, tokenizing, and part-of-speech tagging of legal texts, providing standardized input data for subsequent processing. This process usually includes basic processing tasks such as cleaning, tokenizing, and part-of-speech tagging. When the user inputs their situation and demands, natural language processing (NLP) techniques can be used to tokenize, perform part-of-speech tagging, and semantic analysis on the text input by the user, so as to understand the user's intentions and needs. This includes identifying key entities in the input (such as case types, legal provisions, dates, etc.) and analyzing the user's specific demands. The specific steps and techniques are as follows:
[0067] Text Preprocessing
[0068] Clean the text: Use regular expressions to remove irrelevant characters and punctuation marks. For Chinese texts, tools such as re (regular expression library) can effectively remove unnecessary characters and ensure the text is clean and tidy.
[0069] Tokenization: Use the jieba tokenization tool to split a continuous character sequence into meaningful words. Jieba is a popular Chinese tokenization tool, and by using jieba.lcut(text), the text can be decomposed into words. For example, "我想离婚" (I want a divorce) will be split into three words: "我" (I), "想" (want), and "离婚" (divorce). Tokenization is the basis of Chinese text processing because there are no spaces as word separators in Chinese texts.
[0070] Part-of-Speech Tagging
[0071] Use jieba for part-of-speech tagging. Jieba provides the jieba.posseg.lcut(text) method, which can perform part-of-speech tagging on each word. In the sentence "我想离婚" (I want a divorce), the part-of-speech tagging results may be:
[0072] "我" (I) is a pronoun (marked as r)
[0073] "想" (want) is a verb (marked as v)
[0074] "离婚" (divorce) is a noun (marked as n)
[0075] Semantic Analysis
[0076] Parse the sentence structure: This is done through the dependency parsing function of HanLP. This step no longer involves word segmentation, but directly conducts in-depth syntactic structure analysis on the already segmented text. The core of dependency parsing is to determine the dependency relationships between words in a sentence. Dependency relationships are usually represented as a pair of relationships, where one word (referred to as the head word or central word) depends on another word (referred to as the modifier or dependent word). The result of the analysis usually appears as a dependency tree, where each word node has a relationship of depending on another word. The root node of the tree is usually the predicate verb of the sentence, and other nodes are organized according to their dependency relationships. HanLP processes the already segmented data (such as ["I", "want", "divorce"]), identifies "want" as the root node, which is the predicate verb, "I" as the subject, and "divorce" as the object, thus understanding the overall structure and semantic meaning of the sentence.
[0077] Identify key entities: Use the named entity recognition (NER) function of HanLP to extract important entities in the text. When processing already segmented data, the NER model can identify and classify key entities in the text, such as person names, locations, dates, etc. Although there may be no significant entities in a simple segmented sentence (such as ["I", "want", "divorce"]), when dealing with more complex texts, HanLP can effectively identify and label various entities, providing key information for text analysis.
[0078] Legal concept recognition and extraction module: Use NLP technologies such as deep learning, combined with a legal-specific term library and rule set, to automatically identify and extract key information such as legal concepts, legal provisions, and case elements in the text. The LegalBERT model of HuggingFace can be used. This model is optimized for legal texts. By calling the corresponding API or loading the model, it conducts the recognition and extraction of legal concepts on the text that has been cleaned and segmented in the "data preprocessing module". LegalBERT is a variant of BERT that has been optimized and pre-trained for legal texts. It can better understand the vocabulary and context in the legal field, and thus performs well in the task of legal concept recognition and extraction. The following are the steps for using the LegalBERT model to conduct legal concept recognition and extraction on the keywords extracted in the "data preprocessing module":
[0079] 1) Prepare the environment:
[0080] a) Ensure that the computing environment is equipped with appropriate hardware (such as GPU) and software environment to run the LegalBERT model.
[0081] b) Install deep learning frameworks such as PyTorch or TensorFlow. The LegalBERT model is usually built based on these frameworks.
[0082] c) Install the Transformers library, a library provided by HuggingFace for loading and using various pre-trained models, including LegalBERT.
[0083] 2) Load the LegalBERT model:
[0084] a) Use AutoModelForTokenClassification (for named entity recognition tasks) or AutoModelForSequenceClassification (for classification tasks, if applicable) from the Transformers library to load the LegalBERT model.
[0085] b) Determine and specify the model name, e.g.
[0086] 'bert-base-uncased-legal-bert-finetuned-caselaw' or other fine-tuned LegalBERT model name (the specific name needs to look up the exact name of the LegalBERT model provided in HuggingFaceModelHub).
[0087] 3) Prepare input data:
[0088] a) Obtain the cleaned, segmented and part-of-speech tagged text data from the "Data Preprocessing Module".
[0089] b) Encode the text data as required by the LegalBERT model, which usually includes converting the text into tokenids and attention masks. Make sure the input data is in the correct format so that the model can process it accurately.
[0090] 4) Identify legal concepts:
[0091] a) Pass the prepared input data to the LegalBERT model.
[0092] b) The model will classify each input token and output a category prediction for each token, which may include different legal concept categories (such as legal terms, case names, legal terms, etc.).
[0093] c) Based on the output of the model, identify and extract tokens or token sequences related to legal concepts.
[0094] 5) Post-processing:
[0095] a) Convert the token sequence output by the model back to the original text format.
[0096] As needed, the extracted legal concepts are further processed or validated to ensure their accuracy and relevance.
[0097] Legal knowledge graph construction and update module: Based on the extracted legal concepts and their relationships, a legal knowledge graph is constructed, covering multi-dimensional information such as legal provisions, cases, and judicial interpretations. At the same time, a dynamic update mechanism is supported to ensure the timeliness and accuracy of the graph content.
[0098] 1) Construction of the legal concept genealogy corpus. The corpus follows the principles of univocity, scientificity, and systematicity, and includes elements such as definition, term translation, concept symmetry, and logical relationship of concept genealogy. Since the introduction of Western learning to the East, the legal concept system has integrated the introduction of Western missionaries, the efforts of local translators, and Japanese translation. All foreign words have also been systematically adjusted and reorganized in the process of integration and acceptance, and each noun has become an element in the language system. The legal concept system has typical local and contemporary characteristics. This invention compiles a legal concept corpus consisting of 10,000 legal terms based on the principles and methods for the examination of scientific and technological terms proposed in my country's "Principles and Methods for the Examination of Scientific and Technological Terms" (revised draft) (including: univocity principle; scientific principle; systematic, concise, national, international and conventional principles; coordination and consistency principle, etc.), while also having the normative, behavioral, and technical characteristics of legal concepts.
[0099] 2) The construction of the legal concept genealogy corpus is based on the genus-species relationship and the hierarchical relationship. Genus-species relationship: This refers to the relationship in which one concept contains another concept, also known as the inclusion relationship or classification relationship. In the legal field, this relationship is very common. For example, "rights" is a broad concept, which can be further subdivided into "civil rights", "criminal rights", etc.; "civil rights" can be further subdivided into "property rights", "creditor's rights", etc.; "property rights" are further subdivided into "self-property rights" (such as ownership), "other-property rights" (such as usufruct rights, security rights), etc. Hierarchical relationship: This is similar to the genus-species relationship, but more emphasis is placed on the relative position in the hierarchical structure. In the hierarchical relationship, the "superordinate concept" is a more general and abstract concept, while the "subordinate concept" is a more specific and special concept. For example, "rights" is the superordinate concept of "civil rights", "civil rights" is the superordinate concept of "property rights", and "self-property rights" is the subordinate concept of "property rights". The specific process is:
[0100] Graphic Design
[0101] Node definition: Legal concept node: Each legal concept (such as "rights", "civil rights", "property rights", etc.) is treated as an independent node. Each node contains attributes such as: name, definition, translation, and other features related to the concept.
[0102] Terminology nodes: For terms used in legal concepts, you can also create separate nodes, including term translations and related explanations.
[0103] Relationship definition:
[0104] Superordinate and subordinate relationships (IS_A): used to indicate the hierarchical relationship of concepts. For example, "rights" is a superordinate concept of "civil rights", and "civil rights" is a superordinate concept of "property rights".
[0105] Genus-species relationship (HAS_TYPE): used to represent the genus-species relationship of concepts. For example, "property rights" is a type of "civil rights".
[0106] Data structure design
[0107] Concept attributes:
[0108] Univocity principle: ensure that each concept node corresponds to only one unique definition to avoid confusion. For example, the node of "property rights" should clearly point to the definition of "property rights" and not be confused with other concepts.
[0109] Principles of systematicity and scientificity: Establish systematic hierarchical and genus-species relationships for each concept node to ensure the integrity of the concept genealogy.
[0110] Principle of nationality and internationality: Translation and internationally accepted interpretations are added to nodes to ensure the comprehensibility of concepts in an international legal environment.
[0111] Data import and graph construction
[0112] Data import: Import the compiled legal terms and their definitions, translations, logical relationships, etc. into Neo4j. You can use Neo4j's CSV import tool or API to import data in batches.
[0113] Create nodes and relationships:
[0114] Through batch operations or scripting, each legal concept, terminology and other data is created as a node in Neo4j.
[0115] These nodes are connected using defined relationships (hypernym and genus / species) to form a conceptual spectrum. For example, the "rights" node is connected to the "civil rights" node through the IS_A relationship, and then to the "property rights" node through the same relationship.
[0116] Management and updating of knowledge graph
[0117] Dynamic update: After the graph is established, it needs to be updated regularly to reflect changes in legal concepts or new legal provisions. This can be achieved through script automation, which regularly captures the latest legal concepts and provisions from the data source and updates the nodes and relationships in the Neo4j graph.
[0118] Index and query optimization: Create indexes for key node attributes (such as concept names) to speed up queries. At the same time, Cypher query statements can be used to provide users with complex legal concept searches and path queries to facilitate the retrieval of legal information.
[0119] Intelligent analysis and decision support module: Under the framework of Luima (a hypothetical system name with legal unstructured information management as the core architecture), combined with the support of the legal concept corpus, we can build an efficient and intelligent system for in-depth analysis of legal texts, and then support key tasks such as similar case retrieval, legal clause matching, and compliance inspection. The analysis results are displayed to users through a visual interface, and intelligent decision-making suggestions are provided.
[0120] System architecture:
[0121] Data Layer
[0122] Legal concept corpus: Establish and maintain a legal concept corpus containing rich legal terms, definitions, relationships and cases, and ensure that it supports fast retrieval and dynamic updates. The corpus should cover multiple levels of legal concept relationships to support advanced legal reasoning and query functions.
[0123] Legal text data source: Collect and integrate various legal text data sources, including laws and regulations, judicial precedents, legal papers, etc., to provide a rich data foundation for the system. Data sources should be stored and managed in categories to meet the needs of different legal fields.
[0124] Processing Layer
[0125] Unstructured information processing: Use natural language processing technology (NLP) to perform word segmentation, part-of-speech tagging, named entity recognition (NER) and other processing on legal texts to convert unstructured texts into structured or semi-structured data.
[0126] Application Layer
[0127] Functional application module
[0128] Similar case retrieval: Based on the case facts or legal issues entered by the user, similar historical cases are automatically retrieved in the knowledge graph, and the key information and legal basis of the relevant cases are displayed.
[0129] Legal clause matching: Automatically match relevant legal clauses based on the specific content of the user's query, and provide detailed explanations and applicable scenarios of the clauses.
[0130] System application module: login registration, user management, access control and system monitoring, responsible for user identity authentication and system operation management
[0131] Interface Layer
[0132] Web service: Provides HTTP / HTTPS interface through RESTful or SOAP protocol, handles client requests and interacts with the service layer.
[0133] Service interface API: defines the communication specifications between modules within the system to ensure consistent data exchange and function calls.
[0134] User Interface (UI): Provides a front-end interface that allows users to interact with the system, perform operations and view results through web services and APIs.
[0135] The system needs the following functions:
[0136] Legal concepts can be broken down into three main components: (1) the connotation of the concept, which is the unchanging part that reflects the essential attributes or unique properties of the object; (2) the extension of the concept, which provides a series of examples that meet the requirements of the connotation and defines the scope of application of the concept; (3) other influencing factors related to the concept. Legal information retrieval aims to help users automatically retrieve relevant concepts and roles required to solve legal problems. This retrieval method not only focuses on the matching of concepts in the text, but also emphasizes the connotation of the concept and its role in the document, simulating the user's information needs to solve the problem. For example, when users conduct legal arguments, the accurate identification and application of concepts and their roles are crucial.
[0137] The main steps are as follows:
[0138] Word segmentation and part-of-speech tagging
[0139] Tool: Jieba
[0140] Word segmentation: Word segmentation is performed based on the Forward Maximum Matching (FMM) algorithm.
[0141] Part-of-speech tagging: Use part-of-speech tagging based on Hidden Markov Model (HMM).
[0142] Key concept identification: BiLSTM-CRF
[0143] BiLSTM-CRF (Bidirectional Long Short-Term Memory Network-Conditional Random Field) is a deep learning model commonly used in sequence labeling tasks, especially for named entity recognition (NER). BiLSTM-CRF combines BiLSTM and CRF models to better process contextual information in sequence data.
[0144] BiLSTM: Bidirectional long short-term memory network, which can capture contextual information. t , calculate its hidden layer representation h t :
[0145] The forward LSTM hidden layer representation is:
[0146] The forward LSTM hidden state at the current time step t is the hidden state of the previous time step through the forward pass and the current input x t Calculated.
[0147] The backward LSTM hidden layer representation is:
[0148] The backward LSTM hidden state at the current time step t is the hidden state of the next time step through the back propagation and the current input x t Calculated.
[0149] Bidirectional LSTM output representation:
[0150] h of bidirectional LSTM t Is to forward the LSTM hidden state and the backward LSTM hidden state Obtained by concatenation.
[0151] CRF: Conditional Random Field is used to decode the optimal label sequence for sequence annotation. For a given label sequence y and word sequence x, calculate the conditional probability:
[0152]
[0153] Among them, score(x,y) is the score function calculated by the CRF layer.
[0154] The core algorithm is shown in Algorithm 1:
[0155]
[0156]
[0157] Conceptual Connotation Extraction: HanLP
[0158] HanLP provides a variety of natural language processing functions, among which dependency syntactic analysis is a key technology that can be used to extract the connotation of concepts. By performing dependency syntactic analysis on text, the dependencies between words can be identified, thereby extracting the core attributes of the concept. The generation of dependency syntactic trees usually uses algorithms such as Eisner's Algorithm, which builds a tree structure to represent the dependencies between words in a sentence. In the dependency relationship, ROOT represents the root node of the sentence. Each word has a dependency relationship with another word, and this related word is called the "dependency parent node" of the word. It helps to accurately extract key attributes that reflect the connotation of the concept from the text.
[0159] The core algorithm is shown in Algorithm 2:
[0160]
[0161]
[0162] Concept extension extraction: LegalBERT
[0163] LegalBERT is a BERT pre-trained model optimized for legal texts, capable of processing language structures and expressions unique to the legal field. Through model prediction, LegalBERT performs semantic analysis and classification on the input legal concepts to extract their extension information. The model can not only understand the core meaning of legal concepts, but also identify the scope of application of concepts. Using the legal corpus, LegalBERT extracts specific examples that conform to the connotation of the concept, helping to establish the extension of the concept in different contexts.
[0164] The core algorithm is shown in Algorithm 3:
[0165]
[0166] Related concept association: Neo4j
[0167] Neo4j is a graph database that can build and query legal knowledge graphs and support association and correlation analysis between concepts. Through graph query, the Cypher query language is used to find nodes and their relationships related to the target concept in the knowledge graph. Neo4j can efficiently extract the associations between concepts, help identify other legal clauses, precedents or concepts related to the current concept, and further expand the scope of legal information retrieval.
[0168] The core algorithm is shown in Algorithm 4:
[0169]
[0170] Second, by extracting key legal concepts from the case facts, relevant judicial precedents can be retrieved. We call the input question the current factual situation, and then process the retrieved legal sources at three levels: first, analogize the current factual situation with the case supporting the plaintiff; second, distinguish the argument of the cited case from the current factual situation, which represents the defendant and its supporting counter-example; finally, distinguish the rebuttal of the counter-example case from the current factual situation and the hypothetical proposed facts, and strengthen the plaintiff's argument in the current factual situation when possible, such as Figure 2 shown.
[0171] The analogy between the present factual situation and the cited case is intended to illustrate legally relevant similarities that explain why they should be decided in the same way. These similarities reflect common legal factors between the present factual situation and the cited case. If at least one of these legal factors supports one party's argument, then the case can serve as a potential reason to support that party's argument, thereby assigning the same result to the present factual situation.
[0172] The specific steps are as follows:
[0173] Text vectorization: Use LegalBERT to convert the input legal concepts (i.e. key legal concepts) into vector representations. The core algorithm is as follows:
[0174]
[0175]
[0176] The vector of the current factual situation is compared with the vector of the case supporting the plaintiff to calculate the similarity between them. By using cosine similarity to measure the similarity of the two text vectors, it is possible to quantify the degree of similarity between the current factual situation and the case supporting the plaintiff. This process provides a quantitative basis for analogical argumentation and helps to determine whether there is a legally relevant similarity.
[0177] Cosine similarity is used to measure the similarity between two vectors and is often used to compare text vectors.
[0178] The calculation formula is:
[0179]
[0180] Where: A·B is the dot product (inner product) of vector A and vector B.
[0181] ||A|| and ||B|| are the Euclidean norms (i.e. the lengths of the vectors) of vector A and vector B respectively.
[0182] The value range of chord similarity is between [-1, 1], where 1 means exactly the same, 0 means irrelevant, and -1 means completely opposite. In text similarity calculation, it is usually expected that the closer the value is to 1, the higher the similarity.
[0183] Cited Case Analysis: Identify and distinguish the arguments of the cited case from the current factual context, and evaluate whether these arguments support the defendant and its counterexample. By comparing the vector of the cited case with the current factual context, it is possible to quantify the support of the cited case for the defendant's position and its counterexample. This process helps to clarify whether the arguments in the cited case effectively support the defendant or can be used as evidence to refute the plaintiff. This step uses the cosine similarity calculation method to measure the similarity between the two to complete the argument differentiation.
[0184] Counter-example case analysis: Extract and analyze counter-example cases from the current factual situation and the proposed facts of the hypothesis to clarify how the counter-examples affect the plaintiff's argument. During the analysis, the counter-examples are compared using similarity calculation methods to evaluate their relevance to the current facts. By identifying the key arguments of the counter-examples, these arguments can be refuted more effectively and, if possible, further strengthened the plaintiff's argument in the current factual situation. The cosine similarity method is also used.
[0185] Third, use the semantic web for legal reasoning. The semantic web is a graphical structure in which nodes represent concepts. Arcs represent the relationship between concepts. By building a semantic web, it is possible to show how courts interpret case outcomes based on standard facts. Courts compare key facts with legal concepts in regulations to determine whether these facts meet specific legal standards and thus draw conclusions, such as Figure 3 shown.
[0186] The specific steps are as follows:
[0187] Building the Semantic Web
[0188] The semantic web displays legal concepts and their relationships in a graphical structure. In the semantic web, nodes represent legal concepts and arcs represent relationships between concepts. This structure allows us to explain case outcomes based on standard facts. By building a semantic web, we can clearly represent how legal concepts are related to other concepts and how they work together to support or refute arguments in a case.
[0189] Node definition: Identify and model legal concepts as nodes in the semantic web. For example, "domestic violence", "cheating", and "bigamy" can be different nodes. These nodes represent specific legal concepts or facts in the case.
[0190] Relationship modeling: Define the relationships between nodes and represent these relationships as arcs. For example, the node "domestic violence" can have a "support" relationship with the node "divorce", indicating that domestic violence may support a divorce decision. The definition of relationships should meet the needs of legal reasoning and be able to clearly represent the legal connection between different concepts.
[0191] Data import: Import the defined legal concepts and their relationships into the Neo4j graph database. The data import process includes storing the node and arc information in Neo4j in a suitable format so that it can be effectively queried and operated.
[0192] Comparison of standard facts and regulations
[0193] Evaluate whether a case meets certain legal standards by comparing the facts of the case with concepts in laws and regulations. This process helps determine how legal concepts apply to a specific case, thereby concluding whether the relevant legal requirements have been met.
[0194] Node matching: Match the facts in the case (e.g., “domestic violence”) with legal concept nodes in the semantic web. This step involves finding concept nodes related to the facts of the case in the graph database.
[0195] Relationship analysis: Analyze the relationship between case facts and legal standards through graph traversal algorithms. Two commonly used graph traversal algorithms include:
[0196] Breadth-first search (BFS): Starting from the starting node, explore all nodes in the graph layer by layer until the target node is found or all paths are traversed. It is suitable for finding the shortest path and understanding the extensive connection of nodes.
[0197] Depth-first search (DFS): Start from the starting node and follow a path until you reach the target node or there are no further paths. It is suitable for discovering all possible paths in a graph or analyzing the hierarchy of nodes.
[0198] Graph query: Use the query language of the graph database (such as Cypher) to extract qualified nodes and relationships from Neo4j. The purpose of graph query is to find out the legal standard nodes related to the facts of the case and check whether these nodes support the specific requirements of the case. Cypher is a query language used in Neo4j to operate graph data. It allows users to describe nodes and relationships in the graph structure through declarative syntax and perform complex queries and data operations.
[0199] The core algorithm is as follows:
[0200]
[0201]
[0202] Legal concept verification and conclusion deduction
[0203] Verify the matching and relevance of legal concepts, and deduce the case results by analyzing the nodes and relationships in the semantic web. This process includes judging whether the case meets the legal standards through the reasoning rules and relationships in the graph database and generating the final conclusion.
[0204] Inference rules apply:
[0205] Define rules: Define the reasoning rules applicable to the case based on legal norms and standards. These rules can include reasoning logic extracted from legal provisions, case law, and other legal literature.
[0206] Apply rules: Use the reasoning engine to apply these rules through custom Cypher queries in Neo4j. The reasoning engine can handle complex logical relationships, combine the relationships between nodes with legal rules, and deduce the conclusion of the case.
[0207] Conclusion generation:
[0208] Graph query analysis: Through the Cypher query language, key nodes and relationships in the graph data are extracted to analyze whether the case meets legal standards. Graph query can help confirm which legal concepts and relationships support or refute the arguments in the case.
[0209] Derive conclusion: Based on the graph query results and reasoning rules, generate the final case conclusion. This conclusion will indicate whether the case meets specific legal standards and provide support for the court's decision.
[0210] Construct a theory constructor. The constructor should integrate relevant cases, legal rules, preferences in cases, and value preference rules in the Anli Library. With this information, the theory constructor can make legal arguments based on existing cases and legal rules, and predict the possible outcomes of new cases, such as Figure 4 shown.
[0211] IRRAC logical structure sorting
[0212] First, use natural language processing (NLP) tools to analyze the text. Choose appropriate NLP tools, such as HanLP, for basic text processing such as word segmentation, part-of-speech tagging, and named entity recognition. Next, use the pre-trained LegalBERT model to identify the IRRAC structure of the text. In specific operations, the text will first be input into the LegalBERT model, and the model will analyze the text according to the patterns learned during pre-training. The model will identify different parts of the text and label them as one of the five parts of IRRAC: Issue, Rule, Application, Argument, and Conclusion. For example, the model may label one section of the text as "Rule" and another section as "Application". To extract this information, the algorithm will segment the case text into different parts and input each part into the classification model for identification. After the identification is completed, the output of the model will be organized into a structured data format for further analysis and presentation. Ultimately, storing this structured data will allow for systematic analysis of cases and use of the IRRAC structure as a component of legal consulting reports so that users can better understand and use case information.
[0213]
[0214]
[0215]
[0216] Ranking of effectiveness of articles
[0217] When implementing the sorting of legal provisions, the retrieved legal provisions are first graded according to their effectiveness, nodes are created for each provision in the Neo4j graph database, and the effectiveness level field is marked in the node attributes, such as the Constitution, laws, administrative regulations, departmental regulations, and local regulations. Next, in order to deal with the conflict between legal provisions, a conflict rule field is defined for each provision in the database, and relevant conflict resolution strategies are formulated, such as time priority, special priority, etc. Then, the Cypher query language is used to sort the provisions, first sorting them from high to low according to the effectiveness level, and then applying the conflict rules to adjust the sorting order to ensure that in the case of a conflict between provisions, the provisions with the appropriate priority are displayed. Finally, the provisions with a higher effectiveness level are displayed first, and the conflict rules are considered to ensure that the most applicable legal basis is provided.
[0218] Case relevance ranking
[0219] Use the similarity analysis between the case and the current case to sort related cases, and prioritize the cases that are most similar to the current factual situation. The similarity calculation is achieved through cosine similarity, which is a common text similarity measurement method. The specific implementation steps include: First, vectorize the case text and related cases based on a model such as LegalBERT, and then calculate the cosine similarity between these vectors to obtain the text similarity score. Then, sort the cases according to the similarity score, and prioritize the cases with higher similarity to the current legal issue and factual situation.
[0220] Comprehensive ranking and report generation
[0221] First, extract the effectiveness level of legal provisions, IRRAC logic structure, and case relevance data from Neo4j. Calculate the comprehensive score of each provision and case based on the preset weight distribution (for example, 40% for the effectiveness level of the provision, 30% for the IRRAC logic, and 30% for the case relevance). Then, sort the legal provisions and cases according to these comprehensive scores. Next, use automated report generation tools (such as Jinja2 or LaTeX) to design the structure of the legal consulting report, fill the sorting results and analysis data into the report template, and generate the final report file. This process ensures that the report contains a statement of legal issues, interpretation of applicable laws, case analysis, and final legal advice, so that users can obtain comprehensive and systematic legal consulting information.
[0222] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A legal text intelligent parsing method based on legal concept pedigree, characterized by: The method comprises the following steps: a) Preprocessing the input legal text, including but not limited to text cleaning, word segmentation and part-of-speech tagging, to obtain preprocessed text data; b) based on step (a), using natural language processing technology to identify key legal concepts in the pre-processed text data; c) constructing a legal concept genealogy knowledge graph based on the key legal concepts identified in step (b), wherein the graph at least includes the name, definition, translation and upper and lower relationships of the concepts; d) Analyzing the input case data based on the legal concept pedigree knowledge graph constructed in step (c), linking the identified extracted legal concepts to the knowledge graph, performing legal concept search and path query, extracting logical associations in the legal text, deeply analyzing the preprocessed text data, and identifying logical relationships in the legal text; e) Based on the analysis results of step (d), a case database is constructed to provide users with legal text matching and case retrieval services through intelligent algorithms; f) In combination with the services provided in step (e), receive user feedback and continuously optimize and update the legal concept genealogy knowledge graph.
2. According to claim 1, a legal text intelligent parsing method based on legal concept pedigree is characterized by: The step (a) comprises the following sub-steps: a1) Use regular expression tools to remove irrelevant characters and punctuation marks in the text; a2) Use the Jieba word segmentation tool to segment the processed text into meaningful vocabulary units to form segmented text data; a3) Use jieba's part-of-speech tagging function to tag the text data after word segmentation and identify the grammatical role of each word; a4) Perform semantic analysis, determine the dependency relationship between words through the dependency syntactic analysis function of HanLP, and identify the key entities in the text.
3. The intelligent analysis method of legal text based on legal concept pedigree according to claim 2 is characterized by: The step (a) further comprises: a5) Using natural language processing technology to perform named entity recognition on the pre-processed text data to identify important entities involved in the text; a6) Filter and normalize the results of named entity recognition.
4. According to claim 1, a legal text intelligent parsing method based on legal concept pedigree is characterized by: The step (b) comprises the following sub-steps: b1) Install PyTorch or TensorFlow deep learning framework and Transformers library on your computer; b2) Load the LegalBERT model of HuggingFace through the Transformers library. This model is a pre-trained model optimized for legal texts. b3) Obtain the text data preprocessed in step (a), encode it according to the requirements of the LegalBERT model, and convert it into tokenids and attention masks that can be processed by the model; b4) Input the encoded data into the LegalBERT model, use the model to classify each token, and identify and extract the legal concept categories in the text.
5. The intelligent analysis method of legal text based on legal concept pedigree according to claim 4 is characterized by: The step (b) further comprises: b5) Mapping the identified key legal concepts to the corresponding nodes in the legal concept genealogy knowledge graph constructed in step (c) to enhance the correlation between concepts; b6) Perform semantic expansion on the identified legal concepts according to the context, identify related superordinate or subordinate concepts, and enrich the parsing level of the legal text.
6. The intelligent analysis method of legal text based on legal concept pedigree according to claim 1 is characterized by: The step (c) comprises the following sub-steps: c1) Design the graph structure, define legal concept nodes, each node contains the name, definition, translation attributes of the concept, and create term nodes to contain term translations and related explanations; c2) Define the relationship between legal concepts, using the IS_A relationship to represent the hierarchical relationship of concepts and the HAS_TYPE relationship to represent the genus-species relationship of concepts; c3) Import the compiled legal terms and their definitions, translations, and logical relationship data into the graph database and create corresponding nodes and relationships; c4) Add the legal concepts identified in step (b) as nodes into the graph through batch operations or scripting, and connect them using defined relationships to form a concept genealogy.
7. According to claim 1, a legal text intelligent analysis method based on legal concept pedigree is characterized by: The step (d) comprises the following sub-steps: d1) Based on the legal concept genealogy knowledge graph constructed in step (c), analyze the preprocessed text data to identify the legal concepts in the text and their logical relationships; d2) Using the concept nodes and their relationships in the graph, identify the logical relationships of reference, supplement or conflict of legal provisions in the text; d3) Apply the Cypher query language to perform complex legal concept searches and path queries to extract logical associations in legal texts; d4) Post-process the query results to convert the legal concepts and their logical relationships output by the model into an easily understandable form, laying the foundation for providing legal text matching and case retrieval services in the subsequent step (e).
8. The intelligent analysis method of legal text based on legal concept pedigree according to claim 1 is characterized by: The step (d) further comprises: d5) Identify entities in legal texts using a combination of rule-based methods and machine learning techniques; d6) Through intelligent retrieval technology, the relationship network between entities in legal texts is automatically retrieved and displayed, helping users to more intuitively understand the internal connection between legal provisions and cases.
9. The intelligent analysis method of legal text based on legal concept pedigree according to claim 1 is characterized by: The step (e) comprises the following sub-steps: e1) using the legal concepts and their logical relationships parsed in step (d), automatically searching for legal provisions related to the case facts or legal issues input by the user through an intelligent algorithm; e2) Intelligently match relevant legal terms based on the specific content of the user's query, and provide detailed explanations and applicable scenarios of the terms; e3) Using the cosine similarity calculation method, the case text and related cases are vectorized based on the LegalBERT model, and the cosine similarity between these vectors is calculated to obtain the text similarity score; e4) Sort cases according to similarity scores, prioritize cases with higher similarity to current legal issues and factual situations, and provide users with corresponding legal basis and support.
10. The intelligent analysis method of legal text based on legal concept pedigree according to claim 1 is characterized by: The step (f) comprises the following sub-steps: f1) Designing a user feedback mechanism to allow users to evaluate the effectiveness of the legal text matching and case retrieval results provided in step (e); f2) Collect feedback information submitted by users, including but not limited to user satisfaction ratings of the system's recommended results, correction suggestions, and new cases or legal clauses manually added by users; f3) Analyze the collected feedback information and adjust the node attributes or relationships in the legal concept genealogy knowledge graph according to user feedback to optimize the graph structure; f4) Regularly capture the latest legal concepts and provisions from the data source, and use automated scripts to update the nodes and relationships in the Neo4j graph to ensure the timeliness and accuracy of the knowledge graph.
Citation Information
Cited By
Class case recommendation method based on deep understanding
CN120492612A
A case-based recommendation method based on deep understanding
CN120492612B
Intelligent processing system for enhanced training of legal text small samples
CN120706434A
Intelligent processing system for small sample enhancement training of legal text
CN120706434B
Legal clause recommendation method, electronic device and program product
CN120744092A