Legal text processing method and device based on knowledge graph, equipment and medium

By constructing a legal text processing method based on knowledge graph, the problem of insufficient correlation of legal text in the prior art is solved, and the effect of efficient acquisition and updating of legal texts is achieved.

CN120030994AInactive Publication Date: 2025-05-23CHONGQING IND POLYTECHNIC COLLEGE +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510109828.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When building a knowledge graph of legal texts, the prior art lacks methods to establish correlation between different legal texts, making it difficult for users to obtain relevant legal texts efficiently.

Method used

By constructing a legal text processing method based on knowledge graphs, we can obtain processing requests for legal texts in real time, parse the requests and obtain pending types based on the pre-constructed knowledge graphs, perform text segmentation and extract segmentation words, and update or store legal texts to establish correlation between different legal texts.

Benefits of technology

It has achieved the establishment of correlation between different legal texts, improved the efficiency and accuracy of users' acquisition of legal texts, and can timely update and store legal provisions and judicial interpretations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030994A_ABST
    Figure CN120030994A_ABST
Patent Text Reader

Abstract

The invention provides a legal text processing method and device based on a knowledge graph, equipment and a medium. The method comprises the steps of obtaining a text processing request based on a legal text in real time; analyzing the text processing request and obtaining a to-be-processed type of the legal text according to a pre-constructed knowledge graph; if the legal text is the first type, extracting an abstract of the legal text and storing or deleting the abstract; if the legal text is of the second type, format conversion is carried out on the legal text, text segmentation is carried out based on the converted format, segmentation words are extracted, if the legal text is a first segmentation word, updating processing and storage are carried out on the legal text according to the first segmentation word, and if the legal text is a second segmentation word, updating processing is carried out on the legal text; if yes, executing extraction processing of a third segmented word and a fourth segmented word of the legal text according to the second segmented word, generating the legal text and storing the legal text. Through feature recognition of different types of legal texts and construction of the knowledge graph, the way and efficiency of obtaining legal contents by a user are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph construction, and specifically to a legal text processing method, device, equipment and medium based on knowledge graph. Background Art

[0002] With the development of information technology, the application of artificial intelligence in the field of legal texts has become more and more extensive, including in evidence collection, case analysis, legal document reading and analysis, etc. Knowledge graph construction includes steps such as named entity recognition, relationship extraction, entity alignment, and knowledge completion. In order to improve construction efficiency, existing technologies have attempted to pre-label training data to train small models to complete the various process steps in the above graph construction tasks, such as pre-training named entity recognition models, relationship extraction models, entity alignment models, etc.

[0003] In the prior art, knowledge graphs have also been applied to the construction of legal texts, such as the invention patent with publication number CN118350456B, entitled a method for constructing a knowledge graph based on intellectual property legal documents. It is applied to the construction process of intellectual property graphs for intellectual property legal documents, which is limited to intellectual property books and only considers the triple construction process of legal documents as a whole. However, in the existing legal search process, it is not only necessary to obtain relevant cases of the law, but also to directly obtain the content of the legal provisions when viewing the cases. There is a lack of connection between cases and legal provisions. Secondly, when entering cases and legal provisions into the database, new legal provisions or judicial interpretations need to be updated in a timely manner, which requires the construction of new entities and relationship foundations based on the existing knowledge graph. However, the prior art only extracts characters or words from the text, and does not consider the correlation between different legal texts as a whole, which makes it impossible for users to obtain relevant legal texts more effectively. Summary of the invention

[0004] In order to solve the above technical problems, the present invention proposes a legal text processing method, device, equipment and medium based on knowledge graph, so as to enable users to efficiently and accurately obtain legal texts by establishing associations between different legal texts.

[0005] In a first aspect, the present invention proposes a legal text processing method based on a knowledge graph, the method comprising the following steps:

[0006] Obtain text processing requests based on legal texts in real time;

[0007] Parsing the text processing request and obtaining the type of the legal text to be processed according to a pre-built knowledge graph;

[0008] If it is the first type, extracting a summary of the legal text and storing or deleting it;

[0009] If it is the second type, the format of the legal text is converted, the text is segmented based on the converted format, and the segmentation words are extracted. If it is the first segmentation word, the legal text is updated and stored according to the first segmentation word. If it is the second segmentation word, the third and fourth segmentation words of the legal text are extracted according to the second segmentation word, and then a legal text based on the second segmentation word, the third segmentation word and the fourth segmentation word is generated and stored.

[0010] As a further improvement of the present invention, the step of obtaining a text processing request based on a legal text in real time also includes the step of constructing a knowledge graph, specifically:

[0011] Obtain source documents for many types of legal texts;

[0012] Extracting the category features of the source document to generate legal sub-texts of different categories;

[0013] Extract summary features of the legal sub-text, schedule the first legal sub-text based on the first summary feature to the first container, schedule the second legal sub-text based on the second summary feature to the second container, and the legal text ontology in the first container and the second container constitute the ontology of the knowledge graph.

[0014] As a further improvement of the present invention, the step of scheduling the first legal subtext based on the first abstract feature to the first container further includes:

[0015] The first legal sub-text is preprocessed in the first container to generate a provision text based on the legal provision and an interpretation text based on the judicial interpretation, and the text headers of the provision text and the interpretation text are respectively extracted to obtain the first time information and the legal attributes, and the provision text and the interpretation text are respectively sorted according to the first time information, and then the thread association of the provision text and the interpretation text is established based on the legal attributes and stored in the first container.

[0016] As a further improvement of the present invention, the step of dispatching the second legal subtext based on the second abstract feature to the second container further includes:

[0017] Preprocessing the second legal subtext in a second container to generate review texts based on different legal review types;

[0018] Convert the format of the audit text into an image format, generate a plurality of images corresponding to the page numbers of the audit text, and arrange the plurality of images from top to bottom;

[0019] Scanning a first image among the multiple images from top to bottom to generate a first scanning result based on the personnel information;

[0020] Scheduling the last image of the multiple images and performing a scan from bottom to top, and continuing to scan upward after obtaining the second time information, to generate a second scan result based on the second time information and the legal review result;

[0021] extracting the second time information and the legal review result from the second scanning result;

[0022] After splitting the feature categories based on the legal review results, one or more main labels based on one or more different review result categories are generated and used as the main labels of one or more second sub-containers;

[0023] The second sub-container includes a plurality of sub-containers of sub-labels formed by the first scanning result and the second time information, and the review text corresponding to the first scanning result and the second time information is stored in the sub-container.

[0024] As a further improvement of the present invention, the step of scheduling the second legal sub-text based on the second summary feature to the second container also includes: extracting the legal article information in the second scanning result, generating thread associations between multiple legal articles and the article text in the first container for directly retrieving the specific content of the legal article.

[0025] As a further improvement of the present invention, the step of parsing the text processing request and obtaining the type of the legal text to be processed according to the pre-constructed knowledge graph includes: parsing the text processing request and performing feature extraction on the legal text, if it is a first summary feature, scheduling it to be matched in the first container of the knowledge graph, and if it is a second summary feature, scheduling it to be matched in the second container of the knowledge graph; if no match is found in the first container or the same legal text is found in the second container, the type to be processed is the first type, and if a similar legal text is found in the first container or the same legal text is not found in the second container, the type to be processed is the second type;

[0026] If it is the first type, the summary of the extracted legal text is stored in the first container, or the legal text for which the text processing request is sent is deleted after the same legal text is matched in the second container.

[0027] As a further improvement of the present invention, if it is the second type, the legal text is converted into a format, the text is segmented based on the converted format, and segmentation words are extracted; if it is a first segmentation word, the legal text is updated and stored according to the first segmentation word; if it is a second segmentation word, the third segmentation word and the fourth segmentation word of the legal text are extracted according to the second segmentation word, and then the steps of generating a legal text based on the second segmentation word, the third segmentation word and the fourth segmentation word for storage include:

[0028] If it is the second type, the legal text is converted into a format, the text is segmented based on the converted image format, and the segmentation words are extracted. If the segmentation word is a first segmentation word based on the first summary feature, similar legal texts are retrieved for feature matching, and update annotations are performed on similar legal texts and legal texts based on text processing requests and then stored in the first container; if the segmentation word is a second segmentation word based on the second summary feature, addressing is performed in the second container to locate the second sub-container where the main label matching the second summary feature is located, and then the third segmentation word based on the second time information and the fourth segmentation word based on the first scanning result are obtained based on the image scanning, and then the legal text is stored.

[0029] In a second aspect, the present invention further proposes a processing device for the legal text processing method based on the knowledge graph according to the first aspect, the device comprising:

[0030] An acquisition module, used to acquire text processing requests based on legal texts in real time;

[0031] A parsing module, used for parsing the text processing request and obtaining the type of the legal text to be processed according to a pre-built knowledge graph;

[0032] A processing module is used to extract the summary of the legal text and store or delete it if the type to be processed is the first type; if it is the second type, convert the format of the legal text, segment the text based on the converted format, and extract segmentation words; if it is the first segmentation word, update the legal text according to the first segmentation word and store it; if it is the second segmentation word, extract the third and fourth segmentation words of the legal text according to the second segmentation word, and then generate a legal text storage based on the second segmentation word, the third segmentation word and the fourth segmentation word.

[0033] In a third aspect, the present invention proposes an electronic device comprising a processor and a memory storing a computer program, wherein when the processor executes the computer program, the steps of the legal text processing method based on the knowledge graph as described in the first aspect are implemented.

[0034] In a fourth aspect, the present invention proposes a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the knowledge graph-based legal text processing method as described in the first aspect.

[0035] The present invention proposes a legal text processing method, device, equipment and medium based on knowledge graph, including: obtaining a text processing request based on a legal text in real time; parsing the text processing request and obtaining the type of the legal text to be processed according to a pre-built knowledge graph; if it is the first type, extracting the summary of the legal text and storing or deleting it; if it is the second type, converting the format of the legal text, segmenting the text based on the converted format, and extracting segmentation words; if it is the first segmentation word, updating the legal text according to the first segmentation word and storing it; if it is the second segmentation word, extracting the third segmentation word and the fourth segmentation word of the legal text according to the second segmentation word, and generating and storing the legal text based on the second segmentation word, the third segmentation word and the fourth segmentation word. By identifying the features of different types of legal texts and building a knowledge graph, the way and efficiency of users obtaining legal content are effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a diagram of an implementation scenario of the legal text processing method based on knowledge graph proposed in the present invention.

[0037] Figure 2 This is a flow chart of the legal text processing method based on knowledge graph proposed in the present invention.

[0038] Figure 3 Framework diagram of the knowledge graph-based legal text processing device proposed for the invention.

[0039] Figure 4 Schematic diagram of an electronic device. DETAILED DESCRIPTION

[0040] For ease of understanding, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0041] The terms used herein are only used for the purpose of describing specific example embodiments and are not restrictive. For example, unless the context clearly indicates otherwise, as used herein, the singular forms "a", "an" and "the" may also include plural forms. When used in this specification, the terms "include", "comprise" and / or "contain" mean that the associated integers, steps, operations, elements and / or components exist, but do not exclude the existence of one or more other features, integers, steps, operations, elements, components and / or groups or that other features, integers, steps, operations, elements, components and / or groups may be added in the system / method.

[0042] In view of the following description, these and other features of the present specification, as well as the operation and function of the related elements of the structure, and the economy of the combination and manufacture of the parts can be significantly improved. Reference is made to the accompanying drawings, all of which form a part of this specification. However, it should be clearly understood that the drawings are for illustration and description purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.

[0043] The flowcharts used in this specification illustrate the operations implemented by the system according to some embodiments in this specification. It should be clearly understood that the operations of the flowcharts may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.

[0044] Figure 1 The following is a diagram showing an application scenario of the legal text processing method based on knowledge graph proposed in the present invention. Figure 1 As shown, the present invention proposes a legal text processing method based on knowledge graph, which is specifically applied to the interaction scenario between the user terminal 101 and the server terminal 102. The user terminal 101 can send a request to obtain the legal text to the server terminal 102, or the legal text can be entered into the server terminal 102 through the user terminal 101.

[0045] In this embodiment, there are one or more user terminals 101, and they interact with the server terminal 102 in a wired or wireless manner. It is understandable that the user terminal 101 can be a fixed device such as a desktop computer, or a mobile device such as a mobile phone or a tablet. The purpose is to facilitate users to realize processing based on legal texts through the communication link established with the server terminal 102 anytime and anywhere. The application environment of the user terminal 101 can be in the environment of general users, professional users and special users. For example, general users can be set as non-legal workers. Such users query relevant legal texts as their own legal knowledge acquisition or simple understanding of similar cases. General users usually only care about cases or legal knowledge related to themselves, so the content they need to acquire is only in a local scope; professional users can be considered as non-special users of legal workers, such as lawyers and other professional legal workers. Due to the particularity of their work and the different working scenes of lawyers in different fields, the content they need to acquire will be expanded to the field scope or the whole field; special users in this embodiment are personnel who are professionally engaged in legal text entry. Such personnel can be technicians who develop data technology or personnel who enter legal texts in courts. This makes it easier to understand the scenarios in which there are legal text request interactions and legal text entry interactions when the user end 101 interacts with the server end 102 .

[0046] In this embodiment, the server 102 serves as the main body for providing content, and is configured to have multiple container storage modes based on the knowledge graph, and thread associations are established between different containers. It can be understood that the knowledge graph of this embodiment takes the legal text content as an entity and the thread associations between different containers as relationships to form triples. The triples of this embodiment are not limited to different words in a single text as entities, but also include the format conversion of the legal text and the scanned text after cutting as entities, realizing entity construction from the local to the whole, thereby forming a multi-modal triple knowledge graph model.

[0047] It should be noted that a first container 1021 and a second container 1022 are set at the server 102 to store different text contents. In the first container, contents based on articles and judicial interpretations are formed based on different legal types such as criminal law, civil law, company law, etc. It can also be understood that the first container includes an article sub-container and an interpretation sub-container. The article is placed in the article sub-container, and the judicial interpretation is placed in the interpretation sub-container. The article sub-container and the interpretation sub-container are further constructed into multiple sub-containers based on the label of the legal type, and the association between different containers is established at the same time. For example, the criminal law article is under a sub-container of the article container, and the criminal law amendment is under a sub-container of the second sub-container. Both belong to the criminal law content. When the relevant content of the amendment needs to be obtained when querying the article, the corresponding content is directly obtained through the thread association. This is the explanation for establishing the association between the article container and the interpretation container.

[0048] The above content serves as an explanation of the application scenario of the legal text processing method based on the knowledge graph of this embodiment. It can be understood that when the user terminal 101 sends a legal text acquisition request to the server terminal 102, the user can enter the set keywords through the input interface of the user terminal 101 and send them to the server terminal 102, and the server terminal 102 matches them according to the input keywords; at the same time, when the user terminal 101 sends a text storage request to the server terminal 102, the server terminal 102 can also perform corresponding processing according to the text content entered by the user through the interface, which will be specifically explained below.

[0049] Figure 2 A flow chart of the legal text processing method based on knowledge graph is shown, Figure 2 As shown, this embodiment proposes a legal text processing method based on knowledge graph, which specifically includes the following steps:

[0050] Step S201. Obtaining a text processing request based on a legal text in real time;

[0051] Step S202. Parse the text processing request and obtain the type of the legal text to be processed according to the pre-built knowledge graph. If it is the first type, execute step S203; if it is the second type, execute step S204;

[0052] Step S203. If it is the first type, extract the summary of the legal text and store or delete it;

[0053] Step S204. If it is the second type, the format of the legal text is converted, the text is segmented based on the converted format, and the segmentation words are extracted. If it is the first segmentation word, the legal text is updated and stored according to the first segmentation word. If it is the second segmentation word, the third and fourth segmentation words of the legal text are extracted according to the second segmentation word, and then a legal text storage based on the second segmentation word, the third segmentation word and the fourth segmentation word is generated.

[0054] In this embodiment, as described above, the processing method of the legal text is realized by constructing a communication interaction method between the user end and the server end. In order to achieve the purpose of the present invention, in the embodiment, it is necessary to pre-construct a knowledge graph based on the legal text, which specifically includes:

[0055] Obtain source documents for various types of legal texts,

[0056] Extracting the category features of the source document to generate legal sub-texts of different categories;

[0057] Extract summary features of the legal sub-text, schedule the first legal sub-text based on the first summary feature to the first container, schedule the second legal sub-text based on the second summary feature to the second container, and the legal text ontology in the first container and the second container constitute the ontology of the knowledge graph.

[0058] It should be noted that the source files of legal texts include published legal provisions, judicial interpretations, and review texts of different legal review types. Specifically, legal provisions may include but are not limited to specific provisions of criminal law, civil law, company law, administrative law, etc. Judicial interpretations include but are not limited to various judicial interpretations. Review texts of different legal review types include but are not limited to mediation documents, arbitration documents, judgments, etc. Legal provisions and judicial interpretations may be updated at different times, and review texts of review types are constantly being added. Therefore, in this embodiment, processing is performed according to different text features. For example, if the first summary feature is set to a provision or interpretation, it is considered to be the first legal sub-provision. If the second summary feature is based on the content of a mediation document, arbitration document, or judgment document, it is considered to be the second legal sub-provision, and is stored in different containers respectively.

[0059] In this embodiment, the step of scheduling the first legal sub-text based on the first summary feature to the first container also includes: preprocessing the first legal sub-text in the first container, generating an article text based on the legal article and an interpretation text based on the judicial interpretation, respectively extracting the text headers of the article text and the interpretation text to obtain the first time information and the legal attributes, respectively sorting the article text and the interpretation text according to the first time information, and then establishing the thread association between the article text and the interpretation text based on the legal attributes and storing them in the first container.

[0060] It should be noted that in order to distinguish between the article text and the interpretation text, the article sub-container and the interpretation sub-container can be set in the first container, and both include sub-containers based on different legal attributes (such as criminal law, civil law, etc.). In addition, the time information of any legal document is very important, such as the promulgation time of the article and interpretation, the final review time of the review text, etc., so that the article and interpretation can be updated in time. Therefore, the first time information is obtained when constructing the knowledge graph to facilitate the judgment of whether the update operation needs to be carried out in time.

[0061] In this embodiment, the step of scheduling the second legal sub-text based on the second summary feature to the second container also includes: preprocessing the second legal sub-text in the second container to generate a review text based on different legal review types; converting the format of the review text into an image format to generate multiple images corresponding to the page number of the review text, and the multiple images are arranged from top to bottom; performing a top-to-bottom scan on the first image of the multiple images until the repeated personnel information appears for the first time, and then stopping the scanning to generate a first scan result based on the personnel information; scheduling the last image of the multiple images and performing a bottom-to-top scan, and continuing to scan upward after obtaining the second time information, until the features of all legal review results are obtained, and then stopping the scanning to generate a second scan result based on the second time information and the legal review result; extracting the second time information and the legal review result in the second scan result; performing feature category splitting based on the legal review result to generate one or more main tags based on one or more different review result categories, and use them as the main tags of one or more second sub-containers; the second sub-container includes multiple sub-containers of sub-tags consisting of the first scan result and the second time information, and the review text corresponding to the first scan result and the second time information is stored in the sub-container.

[0062] It should be noted that, in this embodiment, the construction process of the knowledge graph involves the processing of the second legal sub-text in the second container. As mentioned above, the second legal sub-text can be review texts of different legal review types, including but not limited to mediation documents, arbitration documents, judgment documents, etc. The type of this type of text is basically the type marked on the top of the first page, such as mediation, arbitration or judgment. In addition, the first few paragraphs of the first page are descriptions of basic information, including but not limited to information such as the client and the trustee. Therefore, when establishing a relationship, only the text of the first page needs to be scanned. Of course, it is also possible that the first page cannot fully display all personnel information. When the text in this embodiment is converted into an image and scanned from top to bottom, it is performed in a commonly used way of extracting words for personnel names. As long as the personnel information is extracted, the present invention does not make specific limitations. The final time of different review texts, such as the time of the judgment is on the last page. When obtaining the second time information, in order to indicate the time of the review text, the text feature extraction after scanning from bottom to top will obtain the crime, applicable clauses and other contents of the judgment. The extraction of these crimes and other different review results is to classify and extract cases according to the review results. At the same time, these articles do not actually have specific content after extraction. Therefore, it is necessary to establish a thread association with the article text in the first container, so as to directly locate the specific content of the legal article.

[0063] In this embodiment, parsing the text processing request and obtaining the type of legal text to be processed according to the pre-built knowledge graph in step S102 specifically includes:

[0064] Parse the text processing request and perform feature extraction on the legal text; if it is a first summary feature, schedule it to the first container of the knowledge graph for matching; if it is a second summary feature, schedule it to the second container of the knowledge graph for matching; if no match is found in the first container or the same legal text is found in the second container, the type to be processed is the first type; if a similar legal text is found in the first container or the same legal text is not found in the second container, the type to be processed is the second type.

[0065] It should be noted that the legal text carried in the text processing request sent by the user end may be an article to be updated, a new legal case, or a repeated legal case, article, etc. Different situations are handled differently. When the parsed legal text carries the first summary feature, that is, the text content indicating the article or interpretation, it is dispatched to the first container for matching. If the text content of the article or interpretation is not matched, it is a new article content, and it is the first type at this time. If it carries the second summary feature, that is, the feature of the review text, it is matched in the second container. If the same review text is matched, it is also considered to be the first type. And perform the first type of storage or deletion. For the aforementioned, if not, if the text content of the article or interpretation is not matched, it is a new article content. At this time, you only need to extract the summary content of the article to identify the type of the article and perform the storage operation of a sub-container in the first container. At this time, it is also necessary to store according to the storage rules set in the first container, such as which law the article belongs to or which specific law it is related to, etc., which will not be repeated here. Similarly, if similar legal texts are matched in the first container, it should be noted that similar legal texts refer to texts with the same name of articles or interpretations, but only the content of some articles has changed, indicating that the articles and interpretations have been updated. At this time, the old content needs to be replaced by the updated articles and interpretations. At this time, it is identified as the second type. In addition, if the same legal text is not matched in the second container, it is a new legal text based on the review text, which needs to be stored in the second container and is also identified as the second type. Thus, the first type and the second type are clearly stated. It can be understood that the processing method of the first type includes storage or deletion, and the processing method of the second type includes storage based on update and feature extraction. Both types of operations can be executed as an expansion method of the knowledge graph to achieve comprehensive coverage of legal texts based on the knowledge graph.

[0066] In this embodiment, if it is the second type, the legal text is format converted, the text is segmented based on the converted image format, and the segmented words are extracted. If the segmented word is a first segmented word based on the first summary feature, similar legal texts are retrieved for feature matching, and update annotations are performed on similar legal texts and legal texts based on text processing requests and then stored in the first container; if the segmented word is a second segmented word based on the second summary feature, addressing is performed in the second container to locate the second sub-container where the main label matching the second summary feature is located, and then the third segmented word based on the second time information and the fourth segmented word based on the first scanning result are obtained based on the image scan, and then the legal text is stored.

[0067] It should be noted that the first summary feature indicates the text feature of the article or interpretation, and its first segmented word is used as the indicator word of the article or interpretation, so as to retrieve the similar legal text content in the first container and perform the matching of the legal text content. Since the legal text at this time is actually a new article or interpretation text, the parsed legal text does not need to be processed, but the similar article or interpretation in the original first container is directly processed, and the image format conversion is performed and the features are compared. Different features, such as deleted content or added content, are updated and marked in similar legal texts, and stacked storage is performed in chronological order, such as the latest article and the old article are all timestamped to indicate the previous promulgation time of the article. For the text of the second summary feature, it is the review text. Because it is a new text, it is necessary to extract the second segmented word based on the second summary feature. After matching the sub-container, it is stored in a sub-container in the second container according to the third segmented word of the second time letter and the fourth segmented word of the first scanning result. This storage is also based on the same method as the aforementioned second container storage rule, which will not be repeated here.

[0068] In this embodiment, the step of obtaining a text processing request based on a legal text in real time may also include executing a user terminal sending a processing request based on a certain legal text to a server terminal, and after parsing the processing request, identifying that the content of the processing request is an acquisition request based on a certain legal text, then performing a legal text matching operation in the first container or the second container. It should be noted that in the interface-based input request, the user terminal can execute by entering feature keywords of a certain legal text, or can extract relevant legal texts after feature matching based on an image of a certain text.

[0069] Thus, the foregoing has explained the legal text processing process based on the knowledge graph. It should be noted that the knowledge graph construction process is not limited to the above content, but actually also includes feature extraction of keywords in articles, interpretations or review texts, such as keyword searches after accurate identification of "self-defense" and "fraud". These contents do not belong to the improved content of the present invention, but can be used as auxiliary means to achieve the purpose of the present invention. These means will help to locate specific content. The purpose of the present invention is to construct a knowledge graph in a summary manner, and to realize legal text processing based on the knowledge graph through the configuration of summary entities and relationships based on thread associations.

[0070] Figure 3 The framework diagram of the legal text processing device based on the knowledge graph is shown as follows: Figure 3 As shown, the present invention also proposes a legal text processing device 300 based on knowledge graph, and the device 300 includes:

[0071] An acquisition module 301 is used to acquire a text processing request based on a legal text in real time;

[0072] A parsing module 302, used to parse the text processing request and obtain the type of the legal text to be processed according to a pre-built knowledge graph;

[0073] The processing module 303 is used to extract the summary of the legal text and store or delete it if the type to be processed is the first type; if it is the second type, convert the format of the legal text, segment the text based on the converted format, and extract the segmentation words; if it is the first segmentation word, update the legal text according to the first segmentation word and store it; if it is the second segmentation word, extract the third and fourth segmentation words of the legal text according to the second segmentation word, and then generate a legal text storage based on the second segmentation word, the third segmentation word and the fourth segmentation word.

[0074] Figure 4 An example of a physical structure diagram of an electronic device is shown. Figure 4 As shown, the electronic device 400 may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the computer program in the memory 430 to execute the steps of the legal text processing method based on the knowledge graph.

[0075] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0076] On the other hand, an embodiment of the present application also provides a processor-readable storage medium, which stores a computer program, and the computer program is used to enable the processor to execute the steps of the knowledge graph-based legal text processing method provided in the above embodiments.

[0077] The processor-readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NANDFLASH), solid-state drive (SSD)), etc.

[0078] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0079] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A legal text processing method based on knowledge graph, characterized in that: The method comprises the following steps: Obtain text processing requests based on legal texts in real time; Parsing the text processing request and obtaining the type of the legal text to be processed according to a pre-built knowledge graph; If it is the first type, extracting a summary of the legal text and storing or deleting it; If it is the second type, the format of the legal text is converted, the text is segmented based on the converted format, and the segmentation words are extracted. If it is the first segmentation word, the legal text is updated and stored according to the first segmentation word. If it is the second segmentation word, the third and fourth segmentation words of the legal text are extracted according to the second segmentation word, and then a legal text storage based on the second segmentation word, the third segmentation word and the fourth segmentation word is generated.

2. The legal text processing method according to claim 1, characterized in that: The step of obtaining a text processing request based on a legal text in real time also includes the step of building a knowledge graph, which is as follows: Obtain source documents for many types of legal texts; Extracting the category features of the source document to generate legal sub-texts of different categories; Extract summary features of the legal sub-text, schedule the first legal sub-text based on the first summary feature to the first container, schedule the second legal sub-text based on the second summary feature to the second container, and the legal text ontology in the first container and the second container constitute the ontology of the knowledge graph.

3. The legal text processing method according to claim 2, characterized in that: The step of dispatching the first legal subtext based on the first summary feature to the first container also includes: The first legal sub-text is preprocessed in the first container to generate a provision text based on the legal provision and an interpretation text based on the judicial interpretation, and the text headers of the provision text and the interpretation text are respectively extracted to obtain the first time information and the legal attributes, and the provision text and the interpretation text are respectively sorted according to the first time information, and then the thread association of the provision text and the interpretation text is established based on the legal attributes and stored in the first container.

4. The legal text processing method according to claim 3, characterized in that: The step of dispatching the second legal subtext based on the second summary feature to the second container further includes: Preprocessing the second legal subtext in the second container to generate review texts based on different legal review types; converting the format of the review text into an image format to generate a plurality of images corresponding to the page numbers of the review text, wherein the plurality of images are arranged from top to bottom; Scanning a first image among the multiple images from top to bottom to generate a first scanning result based on the personnel information; Scheduling the last image of the multiple images and performing a scan from bottom to top, and continuing to scan upward after obtaining the second time information, to generate a second scan result based on the second time information and the legal review result; extracting the second time information and the legal review result from the second scanning result; After splitting the feature categories based on the legal review results, one or more main labels based on one or more different review result categories are generated and used as the main labels of one or more second sub-containers; The second sub-container includes a plurality of sub-containers of sub-labels formed by the first scanning result and the second time information, and the review text corresponding to the first scanning result and the second time information is stored in the sub-container.

5. The legal text processing method according to claim 4, characterized in that: The step of scheduling the second legal sub-text based on the second summary feature to the second container also includes: extracting the legal article information in the second scanning result, generating thread associations between multiple legal articles and the article text in the first container for directly retrieving the specific content of the legal article.

6. The legal document processing method according to claim 5, characterized in that: The step of parsing the text processing request and obtaining the type of the legal text to be processed according to the pre-built knowledge graph includes: parsing the text processing request and performing feature extraction on the legal text, if it is a first summary feature, dispatching it to the first container of the knowledge graph for matching, if it is a second summary feature, dispatching it to the second container of the knowledge graph for matching; if no match is found in the first container or the same legal text is found in the second container, the type to be processed is the first type, if a similar legal text is found in the first container or the same legal text is not found in the second container, the type to be processed is the second type; If it is the first type, the summary of the extracted legal text is stored in the first container, or the legal text for which the text processing request is sent is deleted after the same legal text is matched in the second container.

7. The legal document processing method according to claim 6, characterized in that: If it is the second type, the legal text is formatted, the text is segmented based on the converted format, and segmentation words are extracted; if it is the first segmentation word, the legal text is updated and stored according to the first segmentation word; if it is the second segmentation word, the third segmentation word and the fourth segmentation word of the legal text are extracted according to the second segmentation word, and the steps of generating a legal text based on the second segmentation word, the third segmentation word and the fourth segmentation word for storage include: If it is the second type, the legal text is converted into a format, the text is segmented based on the converted image format, and the segmentation words are extracted. If the segmentation word is a first segmentation word based on the first summary feature, similar legal texts are retrieved for feature matching, and update annotations are performed on similar legal texts and legal texts based on text processing requests and then stored in the first container; if the segmentation word is a second segmentation word based on the second summary feature, addressing is performed in the second container to locate the second sub-container where the main label matching the second summary feature is located, and then the third segmentation word based on the second time information and the fourth segmentation word based on the first scanning result are obtained based on the image scanning, and then the legal text is stored.

8. The processing device of the legal text processing method based on knowledge graph according to any one of claims 1 to 7, characterized in that: The device comprises: An acquisition module, used to acquire text processing requests based on legal texts in real time; A parsing module, used for parsing the text processing request and obtaining the type of the legal text to be processed according to a pre-built knowledge graph; A processing module is used to extract the summary of the legal text and store or delete it if the type to be processed is the first type; if it is the second type, convert the format of the legal text, segment the text based on the converted format, and extract segmentation words; if it is the first segmentation word, update the legal text according to the first segmentation word and store it; if it is the second segmentation word, extract the third and fourth segmentation words of the legal text according to the second segmentation word, and then generate a legal text storage based on the second segmentation word, the third segmentation word and the fourth segmentation word.

9. An electronic device comprising a processor and a memory storing a computer program, characterized in that: When the processor executes the computer program, the steps of the legal text processing method based on the knowledge graph are implemented as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the legal text processing method based on knowledge graph as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • A knowledge graph construction method based on intellectual property legal documents

    CN118350456B