A method and system for intelligently generating legal dictionaries

By constructing a syntactic parser and legal term extraction rule library based on context-independent law, legal texts are analyzed and term extraction are solved, and the problem of inaccurate output when the existing legal AI system is used to process legal texts is achieved, and efficient and accurate legal text processing and term extraction are achieved.

CN119886121BActive Publication Date: 2025-06-06SHANGHAI ZHENLING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510369311.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-06
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

When processing legal texts, the existing legal AI system has poor quality of input content, resulting in inaccurate output content and cannot effectively meet the professional requirements of legality, which limits the value of legal AI in practice.

Method used

An intelligent generation method of legal dictionary is proposed. By constructing a syntax parser and legal term extraction rule database based on context-independent laws, the legal text is parsed grammatical structure and key legal term extraction, and multi-dimensional interpretation is generated to improve the accuracy and efficiency of legal text processing.

Benefits of technology

It realizes efficient structured analysis of legal texts and precise extraction of terminology. The generated legal dictionary can greatly improve the accuracy and field adaptability of legal AI, optimize the output content, and make it more in line with legal professional requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119886121B_ABST
    Figure CN119886121B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of natural language data processing, and specifically to a method and system for intelligently generating a legal dictionary. The present application performs syntactic analysis on legal texts and extracts target syntactic components by constructing a parser that conforms to a context-free grammar and is suitable for processing legal texts, thereby generating a phrase structure grammar tree, and further extracts key legal terms based on a preset legal term extraction rule base and generates corresponding explanations. The present application scheme is capable of efficient structured analysis of legal texts and precise extraction of terms, thereby generating a legal dictionary intelligently, efficiently, and accurately, and also performs well in cross-document consistency and domain new words, greatly improving the accuracy, domain adaptability, computational efficiency, and interpretability of legal dictionaries and their subsequent applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language data processing, and in particular to a method and system for intelligently generating a legal dictionary. Background Art

[0002] In today's era of rapid development of digitalization and intelligence, the legal field actively introduces artificial intelligence technology, aiming to create an efficient and accurate legal AI system to assist legal practice and promote legal research to a new level. Against the backdrop of the continuous deepening of informatization construction, the application of various legal retrieval systems, intelligent legal consulting platforms, etc. is becoming more and more widespread. However, the operation effect of existing platforms is highly dependent on the accuracy and effectiveness of legal AI. They need to rely on accurate understanding of input content and output high-quality responses to meet users' different legal needs. Moreover, in actual applications, the response effect of existing legal AI is often greatly reduced. This is largely due to the many problems with the content submitted to it, such as inaccurate expression of legal issues input when providing content, incorrect use of legal terms, or typos, which makes key information unclear; for example, some content is translated by other AI, resulting in semantic deviations and logical confusion.

[0003] The above defects make it difficult for legal AI to accurately grasp the user's intentions, and thus cannot give a satisfactory and accurate response, which greatly limits the value of legal AI in legal practice and services. If a high-quality legal dictionary can be built as a basis, the content submitted to or processed by legal AI and the response content generated can be corrected and standardized, then the problem of poor input content quality can be solved from the source, and the output content can also be optimized to make it more in line with legal professional requirements, thereby greatly improving accuracy and allowing legal AI to better serve the general public.

[0004] However, there is no efficient and intelligent solution for constructing legal dictionaries. Traditional legal dictionaries need to be manually generated and updated, which is not only inaccurate and inefficient, but also has a single meaning and can only be applied to some professionals. It cannot be widely popularized and cannot be well integrated with platforms or AI applications. In addition, the existing general text processing solutions are not suitable for legal term extraction, which is manifested in the following aspects:

[0005] 1. Structural analysis limitations: When parsing the grammatical structure of the input content, it is often not comprehensive and in-depth enough. It is difficult to effectively handle complex legal sentence structures based only on simple part-of-speech tagging or conventional lexical analysis, especially sentences involving multi-layer nesting and complex modification relationships.

[0006] 2. Insufficient grasp of accuracy and relevance: It is difficult to fully consider the relationship between them. For some compound nouns or legal phrases formed by specific collocations, it is impossible to correctly identify legal phrases by identifying individual words in isolation or simply judging their importance based on word frequency statistics. It is also difficult to retain low-frequency words or merge them through contextual semantic dependencies, making it difficult to extract accurate legal terms.

[0007] 3. Insufficient interpretation: Over-reliance on a certain interpretation leads to problems or ambiguity in the accuracy of the entries. Some entries may have incorrect interpretations, vague concepts, or be inconsistent with current legal provisions. Summary of the invention

[0008] Based on this, this application proposes a method and system for intelligently generating a legal dictionary to address the above-mentioned problems, aiming to intelligently, efficiently and accurately parse legal texts and generate legal dictionaries, and use this to process texts to generate texts or inputs that better meet the requirements of the legal profession, thereby greatly improving the intelligence of the platform and the accuracy of AI applications.

[0009] On one hand, the present application provides a method for intelligently generating a legal dictionary, the method comprising:

[0010] Build a context-free parser and legal term extraction rule base;

[0011] Calling the parser to perform grammatical structure analysis on the legal text, identifying and extracting target syntactic components containing legal terminology structures;

[0012] generating a phrase structure grammar tree based on the target syntactic component;

[0013] Traversing the phrase structure grammar tree, and extracting key legal terms based on the legal term extraction rule base;

[0014] Merging preset adjacent terms according to the combined information degree, and / or splitting preset core semantic term components based on the legal term extraction rule base;

[0015] Generate multi-dimensional explanations corresponding to each legal term.

[0016] Furthermore, the method further comprises:

[0017] Read legal text and process it in chunks;

[0018] Use the word segmentation model to segment the block text into word sequences and mark the part-of-speech tags of each word;

[0019] Perform syntactic analysis on the block text to obtain the grammatical relationship between words and the word index;

[0020] Based on the words, part-of-speech tags, grammatical relations and word indexes, syntactic analysis is performed and a phrase structure grammar tree is generated.

[0021] Furthermore, the method further comprises:

[0022] Acquire the target syntactic structure of the context-free method;

[0023] Matching the words, part-of-speech tags, grammatical relations and word indices with target syntactic components that conform to the target syntactic structure based on the right-branching principle;

[0024] Taking the sentence as the root node, generating corresponding child nodes layer by layer with the target syntactic components;

[0025] Add each word and part-of-speech tag of each target syntactic component to the corresponding child node.

[0026] Preferably, extracting key legal terms based on the legal term extraction rule base includes:

[0027] Based on word tags and / or tag combination recognition, one or more of the following target objects are extracted:

[0028] Extracting a sequence of consecutive noun phrases containing only nouns or only nouns and adjectives and no punctuation or only nouns and preset symbol tags, and / or,

[0029] Extracting independent verb nodes or verb phrases without nested verb phrases and without content, and / or,

[0030] The core semantic units of the logically focused noun phrases and verb phrases are extracted, and it is determined whether the noun phrases contain pre-set orientation and / or belonging relationship modifying components.

[0031] Furthermore, the method further comprises:

[0032] Calculate the combined information of consecutive adjacent nouns or phrases:

[0033] ,

[0034] Among them, Mi (N 1 ,N 2 ) is a noun or phrase N 1 、N 2 The combined information degree, p(N 1 )、p (N 2 ) is the word N 1 、N 2 The probability of appearing in the preset legal expectation library, p (N 1 ,N 2 ) is the probability of the two occurring together, and α is the adjustment coefficient;

[0035] If the combined information degree is greater than a preset threshold, the consecutive adjacent nouns are merged.

[0036] Preferably, the method further comprises:

[0037] Obtain the first and second interpretations of each legal term;

[0038] Extracting the content, source, region and / or timeliness element information of the first interpretation, and extracting the content, source, region, timeliness and / or conflict resolution element information of the second interpretation;

[0039] The first interpretation element information and the second interpretation element information are stored in association with the corresponding legal terms according to a spatiotemporal multi-dimensional hierarchical model, and / or,

[0040] The first interpretation content and the second interpretation content are vectorized and stored in a preset vector database, and associations with corresponding legal terms are established.

[0041] A second aspect of the present application provides a text processing method, which is executed based on a legal dictionary generated by any of the above methods, and the text processing method comprises:

[0042] Obtain one or more terms to be processed from the target text;

[0043] Matching the term to be processed with the legal term in the legal dictionary, and / or calculating the similarity between the term to be processed and the explanation of each legal term;

[0044] Obtaining the best matching one or more target interpretations and / or corresponding target legal terms;

[0045] The term to be processed is interpreted, verified or prompted based on the target legal term and / or target interpretation.

[0046] A third aspect of the present application provides a legal dictionary intelligent generation system, the system comprising:

[0047] A rule building unit, used to build a syntactic parser and a legal term extraction rule base based on the context-free method that complies with legal texts;

[0048] A syntactic parsing unit, used to call the parser to perform grammatical structure analysis on the legal text, identify and extract preset target syntactic components;

[0049] A structure generation unit, used for generating a phrase structure grammar tree based on the target syntactic component;

[0050] A term extraction unit, configured to traverse the phrase structure grammar tree, extract key legal terms based on the legal term extraction rule base, merge preset adjacent key terms according to the combined information degree, and / or split preset core semantic key term components based on the legal term extraction rule base;

[0051] The explanation generation unit is used to generate multi-dimensional explanations corresponding to each legal term.

[0052] A fourth aspect of the present application provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of any one of the above methods.

[0053] A fifth aspect of the present application provides a computer terminal device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of any one of the above methods.

[0054] The technical solution provided in the above application performs syntactic analysis on legal texts and extracts target syntactic components by constructing a parser that conforms to a context-free grammar and is suitable for processing legal texts, thereby generating a phrase structure grammar tree, and further extracting key legal terms based on a preset legal term extraction rule base and generating corresponding interpretations. The solution of the present application can efficiently perform structured analysis and precise term extraction on legal texts, thereby generating a legal dictionary intelligently, efficiently, and accurately, and also performs well in cross-document consistency and domain new words, greatly improving the accuracy, domain adaptability, computational efficiency, and interpretability of legal dictionaries and their subsequent platforms or AI applications.

[0055] Furthermore, the present application adopts a multi-layer structure and combines right-branch matching to meet the different legal definitions of the same term that are common in the legal field, further improving the accuracy and efficiency of legal terminology. In addition, the present application also highly structuredly integrates the scattered and complex legal terms and their interpretation data, through the term hierarchical model under the spatiotemporal dimension, and the orderly storage of vector data in the vector database, so that the originally disordered legal information becomes clear and easy to manage and query, laying a good data foundation for various subsequent applications (such as conflict resolution, retrieval, AI, etc.). BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0057] in:

[0058] Figure 1 is a flow chart of a method for intelligently generating a legal dictionary in one embodiment;

[0059] Figure 2 A schematic diagram of generating a phrase structure grammar tree in one embodiment;

[0060] Figure 3 A schematic diagram of a legal term extraction result in one embodiment;

[0061] Figure 4 is a structural block diagram of a legal dictionary intelligent generation system in one embodiment;

[0062] Figure 5 FIG. 4 is a structural block diagram of a computer device in one embodiment. DETAILED DESCRIPTION

[0063] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0064] The terms "include", "comprising" and "having" and any variations thereof in the specification and claims of the present application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices. In the claims, specification and drawings of the present application, relational terms such as "first" and "second" are merely used to distinguish one entity / operation / object from another entity / operation / object, and do not necessarily require or imply any such actual relationship or order between these entities / operations / objects.

[0065] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The occurrence of the phrase at various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0066] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0067] In one embodiment, if Figure 1 The flowchart of a method for intelligently generating a legal dictionary in the present application is shown, and the method comprises:

[0068] S10. Build a syntactic parser and legal terminology extraction rule base based on context-free method.

[0069] Legal texts are usually rigorously structured, using a large number of noun phrases and specific structures, and the sentence structure of legal texts is complex, with multiple nested noun phrases and prepositional phrases. This application scheme starts with syntactic structure analysis, and builds a syntactic parser and a corresponding legal term extraction rule base based on the context-free method (CFG), aiming to analyze the syntactic structure of legal texts through formalized grammatical rules, and then accurately extract the legal terms therein. Its goal is to achieve an automated, efficient and highly accurate way to separate the professional terms contained in legal texts, and provide basic data support for subsequent applications such as legal term dictionary construction, legal information retrieval, and knowledge graph construction.

[0070] Specifically, by analyzing a large number of legal texts, typical legal sentence patterns are summarized to obtain a representation in the form of rules, and then a set of syntactic rules that conform to legal texts is constructed, such as "[NP[subject, noun phrase]]→[VP[behavior, verb phrase]]→[NP[result, noun phrase]]" or "S→NP VP, VP → V NP | VP, NP →Det N | NP" that conform to CFG, which are used to describe sentences related to legal behavior. Then, a context-free method (CFG) algorithm such as NLTK, StanfordParser or custom CYK / Earley is used to construct the parser based on the defined grammatical rules. The parser can also be constructed by calling or integrating HanLP or other grammar analysis APIs through interfaces, such as building a parser by calling or integrating the parseDependency method of HanLP through interfaces.

[0071] Furthermore, through lexical terminology analysis of a large number of legal terms, we summarize the lexical patterns that can represent legal terms and formulate corresponding extraction rules, and dynamically add or update rules through error analysis (such as missed terms), so as to form a legal terminology extraction rule base.

[0072] S11. Calling the parser to perform grammatical structure analysis on the legal text, identifying and extracting target syntactic components containing legal terminology structures.

[0073] Specifically, in order to make the constructed legal dictionary more universal and comprehensive, a legal text library is pre-constructed in one embodiment of the present application to obtain any legal text for parsing and term extraction, including: collecting various legal text resources, including but not limited to legal texts, court judgments, standard contracts, legal opinions and other texts from different sources and types to ensure that the data has sufficient coverage and representativeness. At the same time, legal texts from different regions and different legal systems can be combined to enrich the diversity of the corpus. After collecting the text, professional data conversion tools are further used to uniformly convert it into a plain text format, and irrelevant formatting marks in the text are removed, such as extra spaces, tabs, line breaks, etc., to ensure the consistency and processability of the data, thereby forming a legal text library.

[0074] S111, reading the legal text from the pre-built legal text library and processing it in blocks.

[0075] Specifically, the legal text of the pre-constructed legal text library is read through a preset API call, and is processed in blocks, such as dividing the read legal text into blocks according to paragraphs or sentences to form a block-processed text.

[0076] S112. Use the word segmentation model to segment the block text into word sequences and mark the part-of-speech tag of each word.

[0077] Specifically, by using a word segmentation model (such as a dictionary-based CRF or a pre-trained deep learning model), the sentences of the block text are segmented into word sequences, and the part of speech of each word in the word sequence is marked (such as "force majeure / NP"). In addition, the words can be converted into pre-trained word vectors (such as Word2Vec, BERT, etc.), and finally the word sequence with part of speech tags and word vectors can be output.

[0078] S113, performing syntactic analysis on the divided text to obtain word grammatical associations and word indexes.

[0079] Specifically, by performing syntactic analysis on the current sentence of the block text, the grammatical relationship between each word (such as subject-predicate, verb-object, etc.) is extracted, and the grammar of all words is obtained according to the association relationship between each word, which specifically includes:

[0080] In one embodiment, the grammatical association relationship of each word is gradually generated by using state transition actions in combination with using a support vector machine (SVM) or a neural network (such as MLP) to predict the best transition action, including:

[0081] Establish feature engineering to extract features of the current state, including: word parts of speech, word vectors, and positions; buffer caches the word parts of speech and word vectors to be processed;

[0082] Predict the best transfer action using a support vector machine (SVM) or a neural network such as MLP to minimize the prediction error of the sequence of transfer actions;

[0083] Output the grammatical association type and its central word index for each word.

[0084] And / or, in one embodiment, the dependency analysis is modeled as a graph structure, all possible dependency edges are scored through the model, and a maximum spanning tree (MST) is selected, including:

[0085] Encode sentence context using, for example, bidirectional LSTM or Transformer;

[0086] Calculate the dependency score (such as subject-predicate relationship score, verb-object relationship score) for each pair of words, and generate a legal dependency tree through a dynamic programming algorithm (such as the Eisner algorithm);

[0087] By maximizing the score of correct dependency edges while suppressing incorrect ones;

[0088] Output the dependency type of each word and its center word index.

[0089] S114, performing syntactic analysis based on the words, part-of-speech tags, grammatical relations and word indexes and generating a phrase structure grammar tree.

[0090] S12. Generate a phrase structure syntax tree based on the target syntactic component.

[0091] Specifically, the acquired words, corresponding part-of-speech tags, grammatical relations, and word index data are sorted and integrated together to facilitate the rapid and accurate calling of each element when constructing the syntax tree, and to clarify the relevant attributes of each word and its position in the sentence structure. Then, a phrase structure syntax tree is generated, including:

[0092] S121. Obtain the target syntactic structure of the context-free method.

[0093] Specifically, the target syntactic structure may be based on one or more syntactic structures in the aforementioned syntactic rule set that conforms to the legal text, and other target legal syntactic structures may be added as needed.

[0094] S122, matching the words, part-of-speech tags, grammatical relations and word indexes with target syntactic components that conform to the target syntactic structure based on the right-branching principle.

[0095] Specifically, legal texts often construct complex and long sentences through right-branching due to the need for rigor. When there are nested modifications, legal terms often nest multiple limiting conditions through right-branching. Therefore, the present application adopts the right-branching matching principle to match the target syntactic components in the words, part-of-speech tags, grammatical relations and word indices that conform to the target syntactic structure. For example, in one embodiment, the target syntactic components are extracted by matching the main sentence by performing the following expansion:

[0096] Plain Text: [CP[conditional clause]] → [NP[subject]] → [VP[behavior]] → [NP[result]].

[0097] S123, taking the sentence as the root node, and generating corresponding child nodes layer by layer with the target syntactic components.

[0098] Specifically, first, the entire sentence is set as the root node and marked as "S or IP", and then according to the target syntactic components determined above, corresponding child nodes are generated under the root node "S", such as how to generate a node marked as "NP" to represent a noun phrase, and a node marked as "VP" to represent a verb phrase, etc. In this way, the branch structure of the first layer of the grammar tree is constructed, and two child nodes representing the main syntactic components of the sentence are extended downward from the root node. The corresponding child nodes are further generated by decomposing the target syntactic components layer by layer, such as further generating child nodes for the components of child nodes such as "NP" or "VP" in combination with the association relationship.

[0099] S124, adding each word and part-of-speech tag of each target syntactic component to the corresponding child node.

[0100] Based on the corresponding sub-nodes generated layer by layer based on the target syntactic components, and adding each word and part-of-speech tag to the corresponding sub-node, a phrase structure grammar tree that can clearly reflect the syntactic structure of the sentence is finally constructed. The phrase structure grammar tree can be generated and displayed in the form of a list field or a multi-branch tree. Figure 2 As shown, it is a schematic diagram of the phrase structure grammar tree generated for the legal clause "When a limited liability company is established, the shareholders shall bear the legal consequences of the civil activities they engage in for the purpose of establishing the company" in one embodiment of the present application.

[0101] S13, traversing the phrase structure grammar tree, extracting key legal terms based on the legal term extraction rule base; and merging preset adjacent terms according to the combined information degree, and / or, splitting preset core semantic term components based on the legal term extraction rule base.

[0102] Specifically, the solution of this application uses recursive traversal of the phrase structure grammar tree to merge eligible language components. For those containing nested nodes or multi-way tree structures, their sub-tree nodes are recursively expanded, and the components to be merged are merged layer by layer. When extracting target terms, they are identified and extracted by pre-constructing a legal term extraction rule base. Preferably, the key legal terms are extracted based on the legal term extraction rule base, and the key terms and phrases are extracted by identifying the following tags or combinations:

[0103] NP (noun phrase): Extract a continuous noun phrase sequence that contains only nouns or nouns and adjectives without punctuation or only contains nouns and preset tags, such as "[NP [JJ significant][NN negligence]], [NP [informed-consent], [NP [NN insurance][NN contract]]", etc.;

[0104] VP (verb phrase): Extract a verb phrase that does not contain nested verb phrases and does not contain quantifiers, such as "[VP [VV perform][NN contract]]", etc.;

[0105] VV (verb): Extract an independent verb node, such as "[VV establish]]", etc.;

[0106] Centrifugal core unit: Extract the core semantic unit that logically focuses on NP (noun phrase) and VP (verb phrase), and determine whether the noun phrase contains preset locational and / or possessive relationship modifying components. If the noun phrase internally contains preset locational and / or possessive relationship modifying components, such as DNP / LCP, etc., the noun phrase is further split based on the aforementioned rule base.

[0107] It should be noted that the term extraction by the legal term extraction rule base in the above embodiments of this application is a preferred implementation manner. Those skilled in the art can also extract terms by combining other extraction rules according to needs. That is, the legal term extraction rule base constructed in this application can also add, delete, and modify and update the corresponding rules according to subsequent needs to adapt to the continuously changing and updated legal term characteristics or application scenarios. For example, with the introduction of new laws and regulations, some new legal concepts and terms appear, or in a specific legal business field (such as emerging network legal affairs), it is found that there are term extraction situations that cannot be covered by the existing rules, and then corresponding new rules need to be added or updated.

[0108] Preferably, in an embodiment of this application, for the consecutive adjacent nouns for extracting the key legal terms, the following further processing is included:

[0109] a. Calculate the combined information measure of consecutive adjacent nouns or phrase words:

[0110] ,

[0111] Among them, Mi (N 1 ,N 2 ) is a noun or phrase N 1 、N 2 The combined information degree, p(N 1 )、p (N 2 ) is the word N 1 、N 2 The probability of appearing in the preset legal expectation library, p (N 1 ,N 2 ) is the probability of the two occurring together, and α is the adjustment coefficient;

[0112] b. If the combined information degree is greater than a preset threshold, the consecutive adjacent nouns or phrases are merged; otherwise, the consecutive adjacent nouns or phrases are split.

[0113] By calculating the combination degree of each adjacent noun through the above-mentioned scheme of the present application, the scheme of the present application can accurately identify and extract relevant complex legal terms, such as "insurance contract, legal consequences", etc.

[0114] In one embodiment, the above embodiments of the present application are directed to Figure 2 The phrase structure grammar tree is used for legal terminology, and the extracted results are as follows Figure 3 As shown, it can be seen from the results that the present application solution can accurately identify and extract various legal terms.

[0115] Through the above-mentioned implementation mode, the present application can accurately extract the phrase "shareholders at the time of establishment of a limited liability company" as a legal term, whose overall function is jointly determined by "shareholders" and "at the time of establishment", and further split the core phrases. If only "shareholders" are extracted as the core word, the key time information of "at the time of establishment" may be lost, affecting the accuracy of legal interpretation and the accuracy of subsequent application retrieval matching.

[0116] The above implementation scheme of the present application constructs a parser that conforms to a context-free grammar and is suitable for processing legal texts to perform syntactic parsing on the legal text and extract target syntactic components, thereby generating a phrase structure grammar tree, and further extracts key legal terms based on a preset legal term extraction rule base and generates corresponding interpretations. By extracting the target legal syntactic structure and then re-establishing the tree structure, the amount of data processing for non-legal syntactic structures can be reduced, recognition accuracy can be improved, and data processing efficiency can be improved, thereby enabling efficient structured parsing of legal texts and precise extraction of terms, and intelligently, efficiently, and accurately generating legal dictionaries, thereby greatly improving the accuracy, domain adaptability, computational efficiency, and interpretability of legal dictionaries and their subsequent applications.

[0117] S14: Store the legal terms and generate multi-dimensional interpretations corresponding to the legal terms.

[0118] Specifically, the extracted results are checked for duplicates in the database and stored in a preset database, and explanations of each term are generated, including:

[0119] S141. Obtain the first interpretation and the second interpretation of each legal term.

[0120] Specifically, for each term (including phrases) stored in the database, its first interpretation (basic interpretation) is obtained. The basic interpretation is mainly an interpretation that conforms to the principle of the original meaning of the legal text and is generally used as the main interpretation. The first interpretation is obtained by obtaining the definitions of laws, regulations, standards, mainstream theories, etc. The first interpretation is mainly used for forward retrieval to provide the most accurate interpretation.

[0121] Furthermore, for each term stored in the database, its second interpretation (extended interpretation) is obtained. The extended interpretation corresponds to the purpose interpretation and system interpretation. The extended interpretation also corresponds to country and time.

[0122] S142. Extract the content, source, region and / or timeliness element information of the first interpretation, and extract the content, source, region, timeliness and / or conflict resolution element information of the second interpretation.

[0123] Specifically, by setting up a corresponding data model, the element information of the first interpretation and the second interpretation is extracted for associated storage. The data model includes at least a four-dimensional data model. Each term includes a basic interpretation and an extended interpretation. The first interpretation includes multi-dimensional interpretation information such as the content, source, region, and time limit of the interpretation. The second interpretation includes multi-dimensional interpretation information such as the content, source, region, time limit, and conflict resolution of different second interpretations. For the interpretation of each term, obtain and record its applicable time range (time limit), geographical range (such as applicable area), and source authority level (such as legal level, administrative regulations, etc.), and set up a corresponding conflict resolution mechanism to resolve conflicts when multiple interpretations conflict. Specifically including:

[0124] Time (limit of effect) dimension: By recording the scope of limitation of each interpretation, it clearly presents the changes in the interpretation of legal terms over time, and understands the evolution of its connotation and application requirements in different historical stages. It helps to accurately select the interpretation content that was effective at the time when studying past legal practices or handling legal affairs across different time periods.

[0125] Spatial (regional) dimension: Different regions have different legal cultures, legislative needs and judicial practice characteristics, so there will be differences in the interpretation of the same legal term. Clarifying regional information will facilitate accurate application of law and comparative law research in cross-border and cross-regional legal affairs, and avoid errors in legal understanding and application caused by ignoring regional differences.

[0126] Interpretation source dimension: Noting that the interpretation comes from specific legal documents, judicial interpretations, typical cases, etc., helps to trace the authority and legal basis of the interpretation. When encountering disputes or needing further in-depth research, it is easy to find the original text for more detailed study, and it can also judge its weight and priority in the application of the law based on the different levels of the source.

[0127] Furthermore, the present application also resolves conflicts in cases of ambiguous interpretations by setting up a conflict resolution mechanism, including:

[0128] 1. Categorize the interpretation conflicts, including relevance conflicts, time conflicts, regional conflicts, and hierarchical conflicts;

[0129] 2. Obtain relevant data for the current scenario, including text type, time, region, etc.;

[0130] 3. According to the correlation data, adjust the text type weight, timeliness weight, region weight, and rank weight respectively to calculate the comprehensive weight;

[0131] 4. Re-sort by weight and return the Top-K results.

[0132] S143. The first interpretation element information and the second interpretation element information are stored in association with the corresponding legal terms according to a spatiotemporal multi-dimensional hierarchical model.

[0133] Preferably, in one embodiment, the present application stores the interpretation element information through a spatiotemporal hierarchical model, and pre-builds a spatiotemporal hierarchical storage model to hierarchize terms according to the spatiotemporal dimension. The present application considers multiple dimensions such as time, region, source level, conflict resolution, etc. through a multidimensional data spatiotemporal hierarchical model. The data model includes these dimensions as attribute fields, and the interpretation of each term needs to be stored hierarchically on these dimensions to facilitate retrieval according to different conditions during query.

[0134] When storing relational databases, you can use databases such as PostgreSQL that are suitable for structured storage, or graph databases such as Neo4j that process the relationship between terms, or document databases such as MongoDB that are suitable for storing semi-structured databases, etc. Further, by developing corresponding query interfaces or API interfaces, it is convenient for users (such as legal practitioners, legal researchers, etc.) to enter the name of legal terms and other optional screening conditions (such as region, time range, interpretation level, etc.), and quickly retrieve the interpretation content that meets the requirements from the database.

[0135] and / or,

[0136] S144. Vectorize the first interpretation content and the second interpretation content and store them in a preset vector database, and establish associations with corresponding legal terms.

[0137] Specifically, the solution of the present application also includes converting each first interpretation and second interpretation into a vector representation through a preset vectorization method. For example, if the BERT model is used, the vector representation of the interpretation text can be obtained by calling the corresponding API.

[0138] Then select a suitable vector database, such as Milvus, Faiss, etc., to facilitate efficient storage and query of vector data. Then configure it according to the requirements of the selected database, including creating indexes, setting storage parameters, etc., to improve the storage and query efficiency of data. Finally, store the vectors of the first interpretation and the second interpretation in the vector database respectively. While storing the vectors, record the corresponding interpretation content and related term information of each vector, and establish the association between the vector and the term. This association can be achieved through the metadata function of the database or an additional mapping table.

[0139] The above implementation scheme of the present application adopts a multi-layer structure and combines right-branch matching to meet the different legal definitions of the same term that are common in the legal field, further improving the accuracy and efficiency of legal terminology. In addition, the present application also highly structuredly integrates the scattered and complex legal terms and their interpretation data. Through the term hierarchical model under the time and space dimensions, and the orderly storage of vector data in the vector database, the originally disordered legal information becomes clear and easy to manage and query, laying a good data foundation for various subsequent applications (such as conflict resolution, retrieval, etc.).

[0140] A second aspect of the present application provides a text processing method, which is executed based on a legal dictionary generated by any of the above methods, and the text processing method comprises:

[0141] S20, obtaining one or more terms to be processed from the target text.

[0142] Specifically, the target text may be text currently input by the user, such as text currently edited, written, or drafted by the user, or text currently input and submitted by the user for retrieval purposes, or text uploaded and submitted to the system for verification or query.

[0143] S21. Match the term to be processed with the legal term in the legal dictionary, and / or calculate the similarity between the term to be processed and the explanation of each legal term.

[0144] Specifically, if the legal interpretation of the target term needs to be retrieved, the terms in the input text can be matched with the terms in the legal dictionary through a matching algorithm. If the corresponding terms and / or interpretations need to be returned based on semantics, the similarity between the term to be processed and the interpretations of each legal term can be calculated, such as by converting vectorized text into a vector representation and calculating the cosine similarity between the first interpretation and the second interpretation.

[0145] S22. Obtain the most matching one or more target interpretations and / or corresponding target legal terms.

[0146] Specifically, the matching results or similarities may be sorted through the above matching or calculation to obtain the most matching one or more target interpretations and / or corresponding target legal terms.

[0147] S23. Interpret, verify or prompt the term to be processed based on the target legal term and / or target interpretation.

[0148] Through the above-mentioned text processing method of the present application, it is possible to more accurately identify or locate words or terms in the input text or submitted text that do not meet the requirements of legal regulations, and to quickly and accurately search based on the terms submitted by the user. For example, in a certain clause of the submitted draft contract, "Party B shall pay Party A punitive damages equivalent to 30% of the contract price", here "punitive damages" should refer to "liquidated damages" to provide users with appropriate descriptions or prompts.

[0149] In one embodiment, Figure 4 As shown, it is a structural block diagram of a legal dictionary intelligent generation system provided by the present application, and the system includes:

[0150] A rule building unit, used to build a syntactic parser and a legal term extraction rule base based on the context-free method that complies with legal texts;

[0151] A syntactic parsing unit, used to call the parser to perform grammatical structure parsing on the legal text, and to identify and extract target syntactic components containing legal terminology structures;

[0152] A structure generation unit, used for generating a phrase structure grammar tree based on the target syntactic component;

[0153] A term extraction unit, configured to traverse the phrase structure grammar tree, extract key legal terms based on the legal term extraction rule base, merge preset adjacent terms according to the combined information degree, and / or split preset core semantic term components based on the legal term extraction rule base;

[0154] The explanation generation unit is used to generate multi-dimensional explanations corresponding to each legal term.

[0155] In one embodiment, Figure 5 As shown, the present application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:

[0156] Build a context-free parser and legal term extraction rule base;

[0157] Calling the parser to perform grammatical structure analysis on the legal text, identifying and extracting target syntactic components containing legal terminology structures;

[0158] generating a phrase structure grammar tree based on the target syntactic component;

[0159] Traversing the phrase structure grammar tree, and extracting key legal terms based on the legal term extraction rule base;

[0160] Merging preset adjacent terms according to the combined information degree, and / or splitting preset core semantic term components based on the legal term extraction rule base;

[0161] The legal terms are stored and multi-dimensional interpretations corresponding to the legal terms are generated.

[0162] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0163] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0164] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for intelligently generating a legal dictionary, characterized in that: The method comprises: Build a syntactic parser and legal term extraction rule base based on context-free method that complies with legal text; Calling the parser to perform grammatical structure analysis on the legal text, identifying and extracting target syntactic components containing legal terminology structures; generating a phrase structure grammar tree based on the target syntactic component; Traversing the phrase structure grammar tree, and extracting key legal terms based on the legal term extraction rule base; Merging preset adjacent legal terms according to the combined information degree, and / or splitting the legal term components of the preset core semantic units based on the legal term extraction rule base; Generate a multi-dimensional interpretation corresponding to each legal term, wherein the multi-dimensional interpretation includes content, source, region and / or timeliness; The step of extracting key legal terms based on the legal term extraction rule base includes: Extracting the core semantic units of the logically focused noun phrase and verb phrase, and determining whether the noun phrase contains a preset orientation and / or belonging relationship modifying component; and / or, Extracting a sequence of consecutive noun phrases containing only nouns, only nouns and adjectives without punctuation, or only nouns and pre-set tags; and / or, Extract independent verb nodes or verb phrases without nested verb phrases and without content.

2. The method according to claim 1, characterized in that The method further comprises: Read legal text and process it in chunks; Use the word segmentation model to segment the block text into word sequences and mark the part-of-speech tags of each word; Perform syntactic analysis on the block text to obtain the grammatical relationship between words and the word index.

3. The method according to claim 2, characterized in that The method further comprises: Acquire the target syntactic structure of the context-free method; Matching the words, part-of-speech tags, grammatical relations and word indices with target syntactic components that conform to the target syntactic structure based on the right-branching principle; Taking the sentence as the root node, generating corresponding child nodes layer by layer with the target syntactic components; Add each word and part-of-speech tag of each target syntactic component to the corresponding child node.

4. The method according to claim 1, characterized in that: The method further comprises: Calculate the combined information of consecutive adjacent nouns or phrases: , Among them, Mi (N1, N2) is the combined information degree of nouns or phrases N1 and N2, p(N1) and p(N2) are the probabilities of N1 and N2 appearing in the preset legal prediction library, p(N1, N2) is the probability of the two appearing together, and α is the adjustment coefficient; If the combined information degree is greater than a preset threshold, the consecutive adjacent nouns or phrases are merged.

5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Obtain the first and second interpretations of each legal term; Extracting the content, source, region and / or timeliness element information of the first interpretation, and extracting the content, source, region, timeliness and / or conflict resolution element information of the second interpretation; The first interpretation element information and the second interpretation element information are stored in association with the corresponding legal terms according to a spatiotemporal multi-dimensional hierarchical model, and / or, The first interpretation content and the second interpretation content are vectorized and stored in a preset vector database, and associations with corresponding legal terms are established.

6. A text processing method, executed based on a legal dictionary generated by the method according to any one of claims 1 to 5, characterized in that: The text processing method comprises: Obtain one or more terms to be processed from the target text; Matching the term to be processed with the legal term in the legal dictionary, and / or calculating the similarity between the term to be processed and the explanation of each legal term; Obtaining the best matching one or more target interpretations and / or corresponding target legal terms; The term to be processed is interpreted, verified or prompted based on the target legal term and / or target interpretation.

7. A legal dictionary intelligent generation system, characterized in that: The system comprises: A rule building unit, used to build a syntactic parser and a legal term extraction rule base based on the context-free method that complies with legal texts; A syntactic parsing unit, used to call the parser to perform grammatical structure parsing on the legal text, and to identify and extract target syntactic components containing legal terminology structures; A structure generation unit, used for generating a phrase structure grammar tree based on the target syntactic component; A term extraction unit, configured to traverse the phrase structure grammar tree, extract key legal terms based on the legal term extraction rule base, merge preset adjacent legal terms according to the combined information degree, and / or split the legal term components of the preset core semantic units based on the legal term extraction rule base; An explanation generation unit, used to generate a multi-dimensional explanation corresponding to each legal term, wherein the multi-dimensional explanation includes content, source, region and / or timeliness; The step of extracting key legal terms based on the legal term extraction rule base includes: Extracting the core semantic units of the logically focused noun phrase and verb phrase, and determining whether the noun phrase contains a preset orientation and / or belonging relationship modifying component; and / or, Extracting a sequence of consecutive noun phrases containing only nouns, only nouns and adjectives without punctuation, or only nouns and pre-set tags; and / or, Extract independent verb nodes or verb phrases without nested verb phrases and without content.

8. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the legal dictionary intelligent generation method according to any one of claims 1 to 5 or the text processing method according to claim 6.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the legal dictionary intelligent generation method according to any one of claims 1 to 5 or the text processing method according to claim 6.

Citation Information

Patent Citations

  • Legal cognition method and device based on multi-level multi-dimension semantic comprehension and medium

    CN108073569A

  • Construction method and device of field knowledge library, computer equipment and storage medium

    CN108664595A