Law query method and equipment based on large model and knowledge graph, and medium

By constructing a legal query method based on large models and knowledge graphs, the problems of inefficient and insufficient accuracy of traditional legal consulting services are solved, and efficient, accurate query and flexible retrieval of legal data are achieved.

CN119940539APending Publication Date: 2025-05-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510010667.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The traditional legal consulting service model relies on the expertise and intuition of lawyers or legal experts, resulting in inefficiency, high cost and difficult to avoid errors or omissions caused by human factors, especially when facing massive legal information resources.

Method used

The legal query method based on large models and knowledge graphs is adopted, and legal data is obtained through pre-set collection channels, word segmentation processing and named entity recognition are carried out, detailed knowledge graphs are built, and search and search using search engines and graph databases.

Benefits of technology

It realizes efficient and accurate query of legal-related data, improves query efficiency and accuracy, enhances the comprehensiveness and flexibility of query, and reduces human error.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940539A_ABST
    Figure CN119940539A_ABST
Patent Text Reader

Abstract

The invention discloses a law query method and device based on a large model and a knowledge graph and a medium, and the method comprises the steps: obtaining related data of a law through a preset collection channel, and storing the related data in a preset database; performing word segmentation processing on the related data, performing named entity recognition on the related data after word segmentation processing to determine a plurality of entities, and determining an association relationship among the plurality of entities; taking the plurality of entities as nodes of the knowledge graph, taking the association relationship as edges of the knowledge graph, constructing the knowledge graph according to the nodes and the edges, and visually displaying the knowledge graph; determining a weight value corresponding to the node, and determining a query statement template according to the node, the edge and the weight value; and determining query content of the user, rewriting the query content according to the query statement template, and retrieving the database according to the rewritten query content to obtain a query result. According to the method, the query efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a legal query method, device and medium based on a big model and knowledge graph. Background Art

[0002] With the development of information technology, the legal field has ushered in a wave of digital and networked transformation. Massive amounts of legal documents, cases and regulations have been carefully sorted and converted into digital formats and stored in the cloud for quick access by users around the world. However, in the face of a large amount of legal information resources, the traditional legal consulting service model is highly dependent on the professional knowledge, experience and intuition of lawyers or legal experts to screen, search and interpret information. This process is not only inefficient and costly, but also difficult to avoid errors or omissions caused by human factors. Summary of the invention

[0003] In order to solve the above problems, the present application proposes a legal query method based on a big model and a knowledge graph, including: obtaining relevant legal data through a pre-set collection channel, and storing the relevant data in a pre-set database; performing word segmentation on the relevant data, and performing named entity recognition on the relevant data after word segmentation to determine multiple entities and determine the association relationship between the multiple entities; using the multiple entities as nodes of the knowledge graph, and using the association relationship as the edge of the knowledge graph, so as to construct the knowledge graph according to the nodes and the edges, and visualize the knowledge graph; determining the weight value corresponding to the node, and determining a query statement template according to the node, the edge and the weight value; determining the user's query content, rewriting the query content according to the query statement template, and searching the database according to the rewritten query content to obtain a query result.

[0004] In one example, searching the database according to the rewritten query content specifically includes: determining a pre-set search engine to obtain a query statement corresponding to the rewritten query content through the search engine; searching the database according to the query statement to obtain cases and articles corresponding to the query content.

[0005] In one example, the method further includes: determining all records stored in the database, converting all records to determine embedding vectors corresponding to all records, and storing the embedding vectors in the search engine; determining the cosine similarity corresponding to the embedding vector through the search engine, determining similar results corresponding to the query content according to the cosine similarity, and thereby supplementing the query results according to the similar results.

[0006] In one example, before constructing the knowledge graph according to the nodes and the edges, the method further includes: determining a pre-set graph database, determining a pattern of the knowledge graph through the graph database, and determining the nodes and the edges according to the pattern.

[0007] In one example, determining a query statement template based on the nodes, the edges, and the weight values ​​specifically includes: determining a pre-set initial query statement template, filling the initial query statement template according to the nodes and the edges; and supplementing the filled initial query statement template according to the weight values ​​to form a multi-dimensional query statement template.

[0008] In one example, the method further includes: determining a query intent corresponding to the query content, sorting and integrating the query results according to the query intent, and sending the sorted and integrated query results to a user.

[0009] In one example, the collection channels include manual import, web crawlers, and API interface docking.

[0010] In one example, the method also includes: if the collection channel is manual import, format conversion is performed based on the file uploaded by the user to determine the relevant data; if the collection channel is a web crawler, multiple relevant websites are crawled through a pre-set crawler framework, the crawled data is label cleaned, and the data after label cleaning is rule matched to remove irrelevant characters in the data; if the collection channel is an API interface docking, the corresponding relevant agency is determined according to the pre-set API interface, and the relevant agencies are retrieved through the API interface to determine the relevant data.

[0011] On the other hand, the present application also proposes a legal query device based on a big model and a knowledge graph, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the legal query device based on the big model and the knowledge graph can execute: obtaining relevant legal data through a preset collection channel, and storing the relevant data in a preset database; performing word segmentation processing on the relevant data, and performing named entity recognition on the relevant data after word segmentation processing to determine multiple entities, and determine the association relationship between the multiple entities; using the multiple entities as nodes of the knowledge graph, and using the association relationship as the edge of the knowledge graph, so as to construct the knowledge graph according to the nodes and the edges, and visualize the knowledge graph; determining the weight value corresponding to the node, and determining a query statement template according to the node, the edge and the weight value; determining the user's query content, rewriting the query content according to the query statement template, and searching the database according to the rewritten query content to obtain the query result.

[0012] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to: obtain legal data through a preset collection channel, and store the relevant data in a preset database; perform word segmentation on the relevant data, and perform named entity recognition on the relevant data after word segmentation to determine multiple entities and determine the association relationship between the multiple entities; use the multiple entities as nodes of a knowledge graph, and use the association relationship as an edge of the knowledge graph, so as to construct the knowledge graph based on the nodes and the edges, and visualize the knowledge graph; determine the weight value corresponding to the node, and determine a query statement template based on the node, the edge and the weight value; determine the user's query content, rewrite the query content according to the query statement template, and search the database based on the rewritten query content to obtain a query result.

[0013] This application achieves efficient and accurate query of legal-related data by constructing a legal query method based on a large model and knowledge graph. Obtain legal data through pre-set collection channels, and perform word segmentation, named entity recognition and other steps to construct a detailed knowledge graph, making the query process more intuitive and convenient. Using a search engine to retrieve the rewritten query content, it is possible to quickly locate relevant cases and clauses, thereby improving query efficiency. At the same time, by embedding vectors for all records in the database and using cosine similarity to determine similar results, the query results are further supplemented, enhancing the comprehensiveness and accuracy of the query. When determining the query statement template, this application considers multiple dimensions such as nodes, edges, and weight values, forming a multi-dimensional query statement template, thereby improving the flexibility and adaptability of the query. At the same time, the query results are sorted and integrated according to the query intent, so that users can obtain the most needed information faster. The collection channels of this application are diverse, including manual import, web crawlers, API interface docking, etc., which can flexibly adapt to different data sources and improve the richness and diversity of data. This application also performs corresponding processing on data from different collection channels, such as format conversion, label cleaning, rule matching, etc., to ensure the accuracy and consistency of the data. This application achieves efficient and accurate query of legal-related data by constructing a legal query method based on a large model and knowledge graph, improves the efficiency and accuracy of the query, and enhances the comprehensiveness and flexibility of the query. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0015] Figure 1 A flowchart of a legal query method based on a big model and a knowledge graph in an embodiment of the present application;

[0016] Figure 2 This is a schematic diagram of a legal query device based on a big model and a knowledge graph in an embodiment of the present application. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0018] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0019] With the increasing maturity of natural language processing (NLP) and knowledge graph technology, legal consulting services have ushered in changes. NLP technology enables computers to "understand" and process human language, so that they can automatically extract key information from massive legal texts and conduct preliminary legal analysis. Knowledge graphs can build a legal knowledge system, intuitively display complex legal relationships in a graphical way, and help users better understand the internal connections between legal provisions.

[0020] At present, the existing legal consultation system faces multiple challenges, which are manifested as follows: the data collection method is too single and relies on fixed data sources, making it difficult to capture and update the latest legal provisions and cases in a timely manner, resulting in information lag; the accuracy of key information extraction from legal texts is insufficient, which directly affects the quality of subsequent knowledge graph construction, making the information in the graph possibly not comprehensive and accurate enough; in the process of knowledge graph construction, the means of entity recognition and relationship extraction are not perfect, further weakening the integrity and practicality of the graph; the retrieval mechanism relies too much on keyword matching, making it difficult to flexibly respond to users' fuzzy queries or multi-dimensional query needs; finally, the result feedback mechanism lacks effective user participation, and the system cannot be continuously optimized according to the user's actual experience, limiting the further improvement of its service quality.

[0021] like Figure 1 As shown, in order to solve the above problems, the embodiment of the present application provides a legal query method based on a large model and a knowledge graph, the method comprising:

[0022] S101. Obtain legal data related to the law through a preset collection channel, and store the legal data in a preset database.

[0023] The pre-set data collection and storage module is responsible for data collection and processing. The functions of this module cover the following three data collection channels: manual import, web crawler, and API interface docking. Manual import users can upload serialized Excel tables or JSON files through the interface to directly import data. Using the Scrapy crawler framework, crawl the latest judgments and legal provisions from multiple legal websites. The crawler will grab relevant web pages and save them for subsequent processing. The API interface cooperates with legal authorities to obtain relevant provisions and case data through API interface calls.

[0024] For data obtained through web crawlers, data cleaning is required. This includes removing HTML tags and using rule matching technology to remove extra spaces, punctuation marks, and other irrelevant characters. The cleaned data is fed into the BERT model for key information extraction. The BERT model is a pre-trained model based on the Transformer architecture. It can understand the meaning of words based on context and accurately extract key information such as time, location, case type, and judgment results. This information is crucial for the subsequent construction of knowledge graphs and structured storage of data.

[0025] The BERT model learns deep representations of language through pre-training tasks. After pre-training, the BERT model is adapted to these tasks by adding an additional output layer and fine-tuning. The bidirectionality of the BERT model enables it to consider the contextual information before and after each word in the text at the same time. This ability allows BERT to more accurately understand the meaning of words in a specific context. During the fine-tuning stage, the BERT model can learn how to extract key information from the text, which is usually achieved by adding a classifier to the output layer of the model, which can predict whether each word in the text belongs to a certain key information category, such as time, place, case type, judgment result, etc. After extracting key information, post-processing steps may be required to further refine and organize this information. For example, using rule matching or heuristic methods to merge or filter out redundant or irrelevant information. In the legal field, when processing a judgment, the BERT model can accurately identify information such as the time, place, parties involved, case type, and final judgment result of the case. This information can be used to build a legal knowledge graph to provide convenient query and reference services for users such as lawyers and judges.

[0026] After the information is extracted, it is stored in the database. Choosing Elasticsearch as a full-text search engine not only enables fast text content retrieval, but also supports complex query conditions. Each stored record is converted into a 1024-dimensional embedding vector using the bge-m3 model, and this vector is stored as an additional field in Elasticsearch for more efficient retrieval and analysis.

[0027] S102: performing word segmentation processing on the relevant data, and performing named entity recognition on the relevant data after the word segmentation processing to determine multiple entities and determine the association relationship between the multiple entities.

[0028] Because Chinese is different from English, it does not have natural separators. Therefore, Jieba word segmentation is used to segment Chinese text into meaningful vocabulary units. After word segmentation, the Bi-LSTM-CRF model is used for named entity recognition (NER). This model can accurately identify various entities in the document, such as people, organizations, and places. In order to have a deeper understanding of the characters mentioned in the document and their relationships, a large-scale language model based on the Transformer architecture is further introduced. This model has powerful language understanding and generation capabilities. By designing specific prompts, such as: "Please list all the characters and their relationships in the text in detail. The characters include..., and the content is... Example: (Xiao Ming, Xiao Hua, brotherly relationship)", the model can generate corresponding answers based on these prompts.

[0029] S103. Use the multiple entities as nodes of the knowledge graph, and use the association relationships as edges of the knowledge graph, so as to construct the knowledge graph according to the nodes and the edges, and visualize the knowledge graph.

[0030] Neo4j is pre-selected as the graph database. Neo4j excels at storing and querying data with complex structures and complicated relationships. During the construction process, the specific mode of Neo4j is determined, which mainly includes nodes (representing each entity) and edges (representing the relationship between these entities). For example, a specific person can be set as a node, and social relationships such as "friends" and "colleagues" can be used as edges connecting these nodes to show the relationship between people. In this way, complex relationships in the knowledge graph can be clearly and intuitively presented and queried.

[0031] S104: Determine a weight value corresponding to the node, and determine a query statement template according to the node, the edge, and the weight value.

[0032] In order to achieve effective query of knowledge graph, query statement templates are written using Cypher language. Users can select specific characters to fill in these templates according to actual needs, so as to support query from multiple dimensions. At the same time, in order to deeply analyze the relationship between the subjects in the graph, the PageRank graph analysis algorithm is used to evaluate the importance of each entity and display the relationship between them. The PageRank algorithm plays a key role in evaluating the importance of entities in the graph. By running the PageRank algorithm, the weight value of each node is calculated, and then it is determined which entities occupy a more core position in the graph. In order to display these relationships more intuitively, the knowledge graph is presented in the form of a graph with the help of visualization tools such as Graphviz. Users can not only clearly see the connections between different entities, but also explore these relationships in depth through an interactive interface to obtain richer and deeper information.

[0033] S105: Determine the query content of the user, rewrite the query content according to the query statement template, and search the database according to the rewritten query content to obtain a query result.

[0034] In the process of case analysis, you can refer to relevant laws, past cases and the suggestions of the big model. For this purpose, a case term retrieval and citation module is introduced. When the user enters the query content, the question is rewritten first, that is, a prompt is designed, such as "Please convert the user's question into a more precise legal professional terminology." The prompt is designed to guide the model to generate a more specific and targeted version of the question. Subsequently, the question entered by the user will be submitted to the big model for processing, and the model will return a rewritten version of the question. This version will incorporate more detailed and specific information about time, place, legal cases, etc., thereby providing a more solid foundation for subsequent case retrieval and analysis.

[0035] In one embodiment, a hybrid search strategy is used to perform content search tasks in the database. According to the key attributes such as time and place included in the rewritten question, an exact match is performed in Elasticsearch to screen out cases and legal provisions that meet these conditions. In addition to relying on exact matching, the vector encoding generated by the bge-m3 model is also used to further calculate the cosine similarity in Elasticsearch to find the closest results at the semantic level. This comprehensive search method can effectively make up for the shortcomings of simple keyword search, especially when dealing with fuzzy queries or queries with unclear intentions, and can show its unique advantages.

[0036] In one embodiment, after successfully retrieving relevant result information, it is necessary to conduct a comprehensive and integrated analysis of this information and sort it according to its importance. Specifically, the first 20 relevant results retrieved, including 10 cases and 10 legal provisions, are input into the big model, which integrates this information according to the user's query intent. Then, the big model will give sorting suggestions based on the integrated information, sort the legal provisions according to these suggestions, and return the sorted results to the user. In this way, users can clearly see which legal provisions are most relevant to their queries.

[0037] In one embodiment, in order to ensure the accuracy and reliability of the query, a reinforcement learning model is constructed using the PPO algorithm. Every time a user queries, he or she will be asked to score the query result. When a sufficient amount of scoring data is collected, a reward model will be trained using this data. After that, the model will generate three sets of content each time, and the reward model will score them, and finally the set of content with the highest score will be selected to present to the user. In addition, the PPO algorithm will regularly use newly collected data for continuous training to ensure that the reward model can continue to adapt and meet the needs of users.

[0038] In one embodiment, a scheduled task is pre-set to periodically check whether there are updates to the data source. Once new data is detected, the data collection and processing process will be automatically started to integrate the new data. The update mechanism is incremental, which means that all data will not be downloaded again for each update, but only some of the data that has been added or changed since the last update will be downloaded. This approach significantly shortens the data processing time. In some specific cases, data updates may need to be performed immediately, and a manual update function is provided for this purpose. Users can trigger data update operations at any time by simply clicking a button on the pre-set interface. Manually triggered update requests will be given a higher processing priority to ensure that users can quickly obtain the latest data content.

[0039] like Figure 2 As shown, the embodiment of the present application also provides a legal query device based on a large model and a knowledge graph, including:

[0040] at least one processor; and,

[0041] a memory communicatively connected to at least one processor; wherein,

[0042] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor to enable a legal query device based on a big model and a knowledge graph to perform:

[0043] Acquire legal data through a preset collection channel, and store the relevant data in a preset database;

[0044] Performing word segmentation processing on the relevant data, and performing named entity recognition on the relevant data after the word segmentation processing to determine multiple entities and determine the association relationship between the multiple entities;

[0045] The multiple entities are used as nodes of a knowledge graph, and the association relationships are used as edges of the knowledge graph, so as to construct the knowledge graph according to the nodes and the edges, and visualize the knowledge graph;

[0046] Determine a weight value corresponding to the node, and determine a query statement template according to the node, the edge and the weight value;

[0047] The query content of the user is determined, the query content is rewritten according to the query statement template, and the database is searched according to the rewritten query content to obtain a query result.

[0048] The embodiment of the present application further provides a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured as follows:

[0049] Acquire legal data through a preset collection channel, and store the relevant data in a preset database;

[0050] Performing word segmentation processing on the relevant data, and performing named entity recognition on the relevant data after the word segmentation processing to determine multiple entities and determine the association relationship between the multiple entities;

[0051] The multiple entities are used as nodes of a knowledge graph, and the association relationships are used as edges of the knowledge graph, so as to construct the knowledge graph according to the nodes and the edges, and visualize the knowledge graph;

[0052] Determine a weight value corresponding to the node, and determine a query statement template according to the node, the edge and the weight value;

[0053] The query content of the user is determined, the query content is rewritten according to the query statement template, and the database is searched according to the rewritten query content to obtain a query result.

[0054] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0055] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.

[0056] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0057] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0058] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0059] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0060] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0061] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0062] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0063] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0064] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0065] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0066] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0067] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0068] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A legal query method based on a big model and knowledge graph, characterized in that: include: Acquire legal data through a preset collection channel, and store the relevant data in a preset database; Performing word segmentation processing on the relevant data, and performing named entity recognition on the relevant data after the word segmentation processing to determine multiple entities and determine the association relationship between the multiple entities; The multiple entities are used as nodes of a knowledge graph, and the association relationships are used as edges of the knowledge graph, so as to construct the knowledge graph according to the nodes and the edges, and visualize the knowledge graph; Determine a weight value corresponding to the node, and determine a query statement template according to the node, the edge and the weight value; The query content of the user is determined, the query content is rewritten according to the query statement template, and the database is searched according to the rewritten query content to obtain a query result.

2. The method according to claim 1, characterized in that Searching the database according to the rewritten query content specifically includes: Determining a preset search engine to obtain a query statement corresponding to the rewritten query content through the search engine; The database is searched according to the query statement to obtain cases and articles corresponding to the query content.

3. The method according to claim 2, characterized in that The method further comprises: Determining all records stored in the database, converting all records to determine embedding vectors corresponding to all records, and storing the embedding vectors in the search engine; The cosine similarity corresponding to the embedded vector is determined by the search engine, so as to determine similar results corresponding to the query content according to the cosine similarity, thereby supplementing the query result according to the similar results.

4. The method according to claim 1, characterized in that: Before constructing the knowledge graph according to the nodes and the edges, the method further includes: A pre-set graph database is determined, and a pattern of the knowledge graph is determined through the graph database to determine the nodes and the edges according to the pattern.

5. The method according to claim 1, characterized in that: Determining a query statement template according to the node, the edge, and the weight value specifically includes: Determine a preset initial query statement template, and fill the initial query statement template according to the node and the edge; The filled initial query statement template is supplemented according to the weight value to form a multi-dimensional query statement template.

6. The method according to claim 1, characterized in that The method further comprises: Determine the query intent corresponding to the query content, sort and integrate the query results according to the query intent, and send the sorted and integrated query results to the user.

7. The method according to claim 1, characterized in that The collection channels include manual import, web crawlers, and API interface docking.

8. The method according to claim 7, characterized in that The method further comprises: If the collection channel is manual import, the format conversion is performed according to the file uploaded by the user to determine the relevant data; If the collection channel is a web crawler, multiple relevant websites are crawled through a pre-set crawler framework, the crawled data is label-cleaned, and the label-cleaned data is rule-matched to remove irrelevant characters in the data; If the collection channel is an API interface connection, the corresponding relevant agency is determined according to the preset API interface, and the relevant agency is retrieved through the API interface to determine the relevant data.

9. A legal query device based on a large model and knowledge graph, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the legal query device based on a large model and a knowledge graph to perform: Acquire legal data through a preset collection channel, and store the relevant data in a preset database; Performing word segmentation processing on the relevant data, and performing named entity recognition on the relevant data after the word segmentation processing to determine multiple entities and determine the association relationship between the multiple entities; The multiple entities are used as nodes of a knowledge graph, and the association relationships are used as edges of the knowledge graph, so as to construct the knowledge graph according to the nodes and the edges, and visualize the knowledge graph; Determine a weight value corresponding to the node, and determine a query statement template according to the node, the edge and the weight value; The query content of the user is determined, the query content is rewritten according to the query statement template, and the database is searched according to the rewritten query content to obtain a query result.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Acquire legal data through a preset collection channel, and store the relevant data in a preset database; Performing word segmentation processing on the relevant data, and performing named entity recognition on the relevant data after the word segmentation processing to determine multiple entities and determine the association relationship between the multiple entities; The multiple entities are used as nodes of a knowledge graph, and the association relationships are used as edges of the knowledge graph, so as to construct the knowledge graph according to the nodes and the edges, and visualize the knowledge graph; Determine a weight value corresponding to the node, and determine a query statement template according to the node, the edge and the weight value; The query content of the user is determined, the query content is rewritten according to the query statement template, and the database is searched according to the rewritten query content to obtain a query result.

Citation Information

Cited By

  • Medical text feature extraction method and system combined with NPL and large model, and medium

    CN120910255A