Consumer right protection multi-source legal knowledge graph construction and intelligent retrieval method
By constructing a multi-source legal knowledge graph, unified representation and entity alignment of laws, regulations, industry norms and judicial cases are achieved. By combining semantic vectors with keyword retrieval, high-coverage candidate sets are generated and multi-hop path reasoning is performed. This solves the problems of legal knowledge fragmentation and insufficient regional adaptability in the existing system, and improves the knowledge base integrity and retrieval accuracy of the legal consultation system.
Patent Information
- Application Number
- CN202511265100.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing intelligent legal consulting systems in the field of consumer rights protection face technical challenges in integrating and accurately retrieving multi-source heterogeneous legal knowledge, including isolated storage of regulations and case data, lack of entity alignment and semantic fusion across data sources, low precision and recall resulting from over-reliance on keyword matching, lack of reasoning about the association between legal clauses and cases, and lack of perception mechanisms for regional and industry contexts, resulting in insufficient knowledge base completeness, retrieval accuracy, and timeliness.
By constructing a multi-source legal knowledge graph, using entity alignment and semantic mapping mechanisms to uniformly represent regulations, industry norms and judicial cases as a knowledge graph with regional and industry source labels, combining legal dictionaries and synonym libraries to parse user natural language, using semantic vector retrieval and keyword retrieval in parallel strategies to generate high-coverage candidate sets, and generating legal basis chains through multi-hop path reasoning, embedding legal clauses and case source labels, to achieve real-time updating and regional adaptation of the knowledge base.
It achieves deep integration of multi-level legal knowledge, improves the integrity and relevance of the knowledge base, ensures high recall and high-precision retrieval results, enhances the professional depth and explainability of answers, has cross-regional adaptability, and maintains system stability in fault scenarios.
Smart Images

Figure CN120743931A_ABST
Abstract
Description
Background Art
[0001] The present invention relates to the field of legal knowledge graphs and retrieval, and in particular to a multi-source legal knowledge graph construction and intelligent retrieval method for consumer rights protection.
[0002] The current intelligent legal consultation system in the field of consumer rights protection faces the technical challenges of integrating and accurately retrieving multi-source heterogeneous legal knowledge. Relevant legal knowledge is scattered across heterogeneous data sources such as laws, regulations, industry norms, and judicial cases. Its structure is low and its semantic expressions vary significantly, resulting in serious fragmentation and siloing of the knowledge base. Traditional systems mainly rely on retrieval mechanisms based on keyword matching, which cannot effectively bridge the semantic gap between users' natural language descriptions and legal terminology, and lack the ability to dynamically adapt to regional differences and industry characteristics. At the same time, due to the lack of a unified structured representation of legal knowledge, the system has difficulty in realizing associative reasoning across clauses and cases, resulting in a lack of legal basis chain support for the answers to complex multi-factor problems. Existing solutions have systemic defects in knowledge coverage, retrieval accuracy, and result interpretability. There is an urgent need to achieve deep integration of multi-source knowledge and semantic-driven intelligent retrieval through technological innovation.
[0003] The core defects of existing methods mainly include: 1. Regulations and case data are stored in an isolated form, lacking entity alignment and semantic fusion mechanisms across data sources, making it impossible to build a knowledge network covering multi-level legal systems, resulting in limited knowledge base integrity; 2. Over-reliance on keyword literal matching or single semantic vector retrieval. The former is prone to missing semantically related clauses due to terminology differences, and the latter is difficult to ensure full recall of key elements explicitly mentioned by users. Both cannot achieve both high precision and high recall; 3. No explicit association path between legal clauses and cases is established. The system cannot perform multi-hop reasoning to generate conclusions based on multiple legal bases, resulting in a lack of professional depth and logical explainability in answers to complex consultations; 4. There is a lack of active perception mechanism for regional and industry contexts, and the knowledge base update relies on manual batch operations, which cannot respond to regulatory revisions and new cases in real time, resulting in insufficient timeliness and regional adaptability of answers. Therefore, based on the above difficulties, the present invention proposes a multi-source legal knowledge graph construction and intelligent retrieval method for consumer rights protection. Summary of the Invention
[0004] Purpose of the Invention In order to solve the above problems, the purpose of the present invention is to provide a multi-source legal knowledge graph construction and intelligent retrieval method for consumer rights protection. By constructing a unified knowledge representation and intelligent retrieval framework, it can achieve deep integration of multi-level legal knowledge, accurate analysis of user intentions, and adaptive retrieval in cross-regional scenarios, and enhance the authority of answers and transparency of reasoning, ultimately improving the coverage, accuracy and credibility of legal consulting services.
[0005] Technical Solution To achieve the above objectives, the present invention provides a multi-source legal knowledge graph construction and intelligent retrieval method for consumer rights protection. This method, through entity alignment and semantic mapping mechanisms, uniformly represents heterogeneous data such as industry regulations and judicial cases into a knowledge graph with regional and industry source tags, addressing the fragmentation of legal knowledge. It combines legal dictionaries and thesaurus to parse user natural language questions, extracting legal entities, dispute elements, and regional and industry contextual constraints to generate structured queries. It employs a parallel strategy of semantic vector retrieval and keyword retrieval, dynamically limiting the search scope based on contextual pre-filtering, and generating a high-coverage candidate set through adjustable weighted fusion scoring. It concatenates rule clauses and case nodes on the knowledge graph to form a reasoning chain, calculates path scores based on node matching and topological distance decay, and generates a legal basis chain with contribution weights. It embeds machine-readable reference tags for legal clauses and cases in answers, and achieves real-time synchronization of the knowledge base through incremental extraction and model collaborative optimization. This solution systematically addresses the problems of fragmented legal knowledge, inefficient retrieval, and weak reasoning, achieving a coordinated improvement in coverage, authority, and timeliness.
[0006] In the first aspect, the present invention provides a multi-source legal knowledge graph construction and intelligent retrieval method for consumer rights protection, including: Construct a multi-source legal knowledge graph that integrates laws, regulations, industry norms, and case texts. Use an entity alignment mechanism to unify semantically equivalent legal concept nodes, and attach regional and industry source labels to the nodes to achieve integrated representation and traceability of cross-source knowledge. Parse user natural language questions, extract the legal entities, dispute elements, and regional and industry context constraints involved, and generate structured queries that can adapt to regional differences to bridge the semantic gap. A hybrid search strategy combining semantic vector search and keyword search is adopted, and the search scope is pre-filtered and limited in combination with the contextual constraints, to generate a set of candidate legal clauses that achieves both high recall and high precision; Based on the knowledge graph, multi-hop path reasoning is performed on candidate legal clauses to generate associated paths with node contribution weights, forming a traceable legal basis chain to enhance the professional depth of the answer. The multi-hop path reasoning connects the regulatory clause nodes and the associated case nodes in the knowledge graph to form an inference chain, and calculates the comprehensive path score based on the path node matching degree and topological distance, where the starting node has a higher weight than the jump node, to ensure the completeness of the answer to complex questions with multiple factors. Output traceable answers with legal clause citations and case sources to ensure the authority and verifiability of the conclusions.
[0007] Furthermore, during the construction of the legal knowledge graph, joint entity recognition and relationship extraction are performed on structured regulatory texts and unstructured case texts to generate triple knowledge units across data sources, and a synonymous concept mapping mechanism is used to map legal terminology specifications with different expressions but semantically equivalent to the same node of the knowledge graph, eliminating the problem of term fragmentation and achieving knowledge association and coherence.
[0008] Furthermore, based on a pre-built adaptive dictionary for the legal field, user spoken expressions are mapped into legal terms, and the extracted geographical and industry contexts are used as mandatory constraints in the retrieval stage to dynamically limit the subspace scope of regulatory retrieval.
[0009] Furthermore, in the hybrid retrieval strategy, there are two parallel sub-modules: a semantic vector retrieval engine and a keyword retrieval engine. The semantic vector retrieval engine uses a deep language model to calculate the implicit semantic similarity between the question and the legal terms, and the keyword retrieval engine matches legal terms and term identifiers through an inverted index, and performs weighted fusion sorting on the two types of results according to an adjustable weight coefficient, thereby achieving a dynamic balance between semantic relevance and term coverage to optimize the overall retrieval efficiency.
[0010] Furthermore, the pre-filtering and limiting of the search scope is performed before the hybrid search begins, and the search authority of the corresponding regional legal and regulatory database or industry specification database is activated based on the context constraints obtained by the analysis, thereby improving the accuracy of regional adaptation and avoiding interference from invalid regulations.
[0011] Furthermore, the comprehensive path score is obtained by accumulating the weighted matching scores of each node on the path, and the weight assigned to a node decreases monotonically as the jump distance between the node and the user question increases.
[0012] Furthermore, in the process of generating traceable answers, the name of the regulation, article number and case source are embedded in the natural language answer text in the form of machine-readable tags; in pure search mode, results containing the original text of the legal provisions corresponding to the question are returned, and the keywords hit therein are visually highlighted.
[0013] Furthermore, when constructing the legal knowledge graph, a pre-calculated confidence score is assigned to each node, wherein the confidence score is determined based on the data source type, data source publishing, and adjudication authority level corresponding to the node; In the hybrid search strategy, the confidence score is used as an adjustment factor for the search priority, and legal clauses corresponding to data sources with high confidence scores are returned first; In the multi-hop path reasoning, the confidence score is combined with the matching degree and topological distance of the nodes in the path to calculate the path comprehensive score, and a higher weight contribution is given to the nodes with high confidence scores.
[0014] Furthermore, the method continuously monitors changes in legal data sources, updates graph nodes and relationships through incremental extraction, and optimizes semantic vector models and keyword dictionaries based on newly added data features, continuously maintaining the timeliness of the knowledge base and synchronously optimizing the retrieval model.
[0015] In a second aspect, the present invention further provides a multi-source legal knowledge graph construction and intelligent retrieval system for consumer rights protection. The system, according to the method described in the first aspect, comprises: A knowledge graph construction module, configured to perform entity recognition, relationship extraction, and semantic alignment processing on multi-source heterogeneous data to construct the legal knowledge graph; The user question parsing module integrates legal domain dictionaries and synonym libraries to parse user natural language questions into structured queries with regional and industry constraints; A hybrid retrieval module, comprising a semantic vector retrieval submodule and a keyword retrieval submodule running in parallel, for performing semantic retrieval and keyword retrieval on the structured query and fusing the results; A path reasoning fusion module performs multi-hop association path search and score calculation on the retrieval candidate results based on the graph database to generate the legal basis link and determine the optimal reasoning path; A legal result output module, configured to generate a natural language answer including a legal clause reference and a case source identifier based on the optimal path reasoning result; The resource dynamic scheduling module is used to automatically switch the semantic model computing mode according to the hardware resource status, dynamically manage the graph database connection pool, and implement the retrieval retry mechanism in abnormal situations; The service fault-tolerant control module is used to interrupt the reasoning process and roll back the output of keyword search results when an exception occurs in the path reasoning function, and switch to pure keyword search mode when the semantic retrieval model fails.
[0016] In a third aspect, the present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run by a processor, the multi-source legal knowledge graph construction and intelligent retrieval method for consumer rights protection is executed.
[0017] The present invention integrates heterogeneous data sources such as laws and regulations, industry standards and judicial cases, generates cross-source triples through entity recognition and relationship extraction, and realizes semantic unified representation of knowledge nodes based on entity alignment and synonym concept mapping mechanism, while adding regional and industry source labels to nodes to support knowledge traceability; combines legal field dictionaries and synonym libraries to parse user natural language questions, extracts legal entities, dispute elements and regional industry context constraints, and dynamically limits the search subspace based on context pre-filtering; adopts a hybrid search strategy that executes semantic vector retrieval and keyword retrieval in parallel, and fuses semantic relevance and term coverage scores through adjustable weight coefficients to generate candidate result sets with high recall and high precision; connects regulatory clauses and case nodes in series on the knowledge graph to form an inference chain, calculates the comprehensive score of the path based on node matching and topological distance decay weighting, and generates a legal basis chain with contribution weight annotation; embeds machine-readable labels of regulatory names, clause numbers and case references in the output, and constructs a collaborative optimization closed-loop update mechanism to maintain the timeliness of the knowledge base. This solution is implemented through a modular system architecture, supporting dynamic adaptation of heterogeneous computing resources and service degradation strategies to ensure robustness in high-concurrency scenarios.
[0018] The present invention eliminates the fragmentation and isolation of legal knowledge through the fusion of multi-source heterogeneous knowledge and entity alignment mechanism, constructs a unified semantic network covering multi-level legal systems, and significantly improves the completeness and relevance of the knowledge base; the dual-engine hybrid retrieval strategy takes into account both semantic generalization capabilities and precise term recall, and combined with the context pre-filtering mechanism, it significantly expands the retrieval coverage while ensuring high relevance, effectively bridging the semantic gap between user natural language and legal professional terminology; the path reasoning mechanism generates a legal basis chain with contribution weights, and combines the clause references and case traceability tags embedded in the answer to give the conclusion strong legal basis support and logical traceability, greatly improving the professional depth of the answer and user credibility; the closed-loop update system based on incremental extraction synchronizes legal dynamics in real time, ensuring that the knowledge base and retrieval model continue to adapt to the latest regulations, while the regional context awareness mechanism guarantees the effectiveness adaptability of answers in cross-regional scenarios; the dynamic scheduling of heterogeneous computing resources and the service degradation strategy support the preservation of core functions in failure scenarios and maintain the response stability of high concurrent access. This solution achieves a systematic breakthrough in existing technologies in terms of the breadth of knowledge integration, comprehensive retrieval performance, professional depth of answers and system adaptability.
[0019] Beneficial effects By implementing the multi-source legal knowledge graph construction and intelligent retrieval method for consumer rights protection provided by the present invention, the following technical effects are achieved: (1) Through entity alignment and cross-source semantic mapping, heterogeneous data such as laws, regulations, industry norms, and judicial cases are uniformly represented as a homogeneous graph structure, with additional regional and industry source labels. This architecture enables deep association and source traceability of multi-level legal knowledge, significantly improving the integrity and relevance of the knowledge base. This framework solves the problem of incomplete coverage caused by knowledge silos in traditional systems and provides global knowledge support for accurate retrieval.
[0020] (2) It provides intelligent question parsing and context pre-filtering strategies for legal consultation Q&A, combining a "dual-engine" hybrid search solution that combines semantic vector search with keyword search, and dynamically pre-filters the search scope based on the user's region and industry context. This mechanism ensures high relevance while also taking into account the completeness of terminology coverage and semantic generalization capabilities, breaking through the inherent limitations of a single search model in terms of recall and precision, and effectively bridging the semantic gap between users' natural language expressions and legal terminology.
[0021] (3) The multi-hop path reasoning fusion mechanism based on the knowledge graph can connect multiple relevant legal clause nodes and case nodes to form a reasoning chain, and calculate the path score based on the node matching degree and topological distance weight, and select the optimal link to generate a conclusion. Compared with the existing question-answering system that only relies on a single legal provision to directly answer questions, this chain reasoning mechanism gives the system the ability to comprehensively analyze the cross-legal basis for complex issues, ensuring the generation of legal conclusions and case sources with a logical chain, greatly improving the professional depth and explainability of the answers, and solving the defect of existing technologies that fragmented knowledge cannot support chain reasoning.
[0022] (4) An adaptive update mechanism for the legal knowledge base is proposed. Through a closed-loop mechanism of regulatory monitoring, incremental extraction, and model collaborative optimization, it realizes the automatic discovery, extraction, and index update of newly enacted legal provisions and cases. This mechanism ensures that the knowledge base content is synchronized with legal dynamics in real time, continuously maintains the timeliness and authority of search results, overcomes the problem of knowledge obsolescence caused by delayed manual updates in traditional systems, and significantly reduces long-term operation and maintenance costs.
[0023] (5) Integrate regional context information for search pre-filtering and have the ability to adapt to cross-regional regulations. The system can automatically limit the search scope based on the user's region and prioritize the return of regulations in the corresponding region, ensuring that the answers comply with local legal and policy requirements. Thanks to the regional label integration and context constraint strategy, the situation of citing invalid regulations in different regions is effectively avoided, and the practical applicability of the answers is improved.
[0024] (6) A comprehensive exception handling and fault-tolerance mechanism has been established. By combining dynamic scheduling of heterogeneous computing resources with service degradation strategies, the system can ensure that it can still operate reliably and provide basic services when failures occur in various modules or when the number of concurrent requests is too high. For example, when the semantic model fails to load or the knowledge graph database connection is abnormal, the system can automatically switch to a simplified model or pure keyword search mode to continue service. When the performance of some modules degrades, time-consuming steps can be temporarily skipped to ensure timely response. The above fault-tolerance scheme significantly improves the robustness of the system in fault scenarios, avoids system crashes and shutdowns due to single point failures, and ensures long-term stability and service continuity. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to make the above-mentioned multi-source legal knowledge graph construction and intelligent retrieval method for consumer rights protection of the present invention more obvious and easy to understand, the following is a brief introduction to the drawings required for use in the specific implementation of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0026] Figure 1 It represents a flow chart of the present application method; Figure 2 A diagram showing the system composition and technical implementation principle of this application; Figure 3 Represents the intelligent retrieval flow chart. DETAILED DESCRIPTION
[0027] Example 1: Provides a multi-source legal knowledge graph construction and intelligent retrieval method for consumer rights protection. The method process is as follows Figure 1 As shown, the details are as follows.
[0028] A method for constructing a legal knowledge graph by integrating multi-source, heterogeneous data is provided. Data sources include structured legal and regulatory texts, industry agreement clauses, and unstructured texts such as consumer complaints and mediation outcomes. By preprocessing and extracting information from these multi-source data, key information is converted into a unified knowledge graph representation. For example, legal provisions are parsed into "triples" of legal knowledge, and typical cases are associated with corresponding legal article nodes, constructing a semantic network across regulations and cases. Specifically, natural language processing techniques are used to automatically identify entities and relationships within legal texts. Key entities within legal texts are extracted using word segmentation and named entity recognition models, and legal relationships between entities are identified using a relationship extraction algorithm. The extracted entities and relationships are inserted into a graph database as triples, forming a legal knowledge graph. For data from different sources, an entity alignment and fusion mechanism is implemented to merge semantically equivalent or related nodes. Furthermore, a source label is added to each node in the graph, indicating its origin as a regulation, industry standard, or case. This enables the knowledge graph to accommodate heterogeneous legal knowledge while maintaining traceability. Through the above method, a unified knowledge graph covering laws and regulations related to consumer rights protection and typical cases is constructed, which opens up the connection between various knowledge sources and eliminates the problems of fragmentation and isolation of legal knowledge in the existing system.
[0029] The system composition and technical implementation principle diagram of this application are as follows Figure 2 shown.
[0030] A dedicated semantic parsing module is designed to identify user intent and extract key elements from natural language user-entered questions. First, using word segmentation, part-of-speech tagging, and named entity recognition (NER) techniques, key elements are identified and extracted from the user's question, including key information such as the name of the goods or services involved, the focus of the dispute, the amount, and the timeframe, as well as any legal entities, regions, or industry attributes. Legal-specific dictionaries and thesaurus are incorporated to address inconsistencies between user terms and legal terminology. For example, the module can identify the colloquial expression "return" as corresponding to the formal legal term "terminate a contract," and recognize "refund" as a synonym for "refund." Second, the parsing module incorporates contextual information extraction. If a user's question contains a reference to a region or industry, this information is recorded as a contextual constraint for use in subsequent search steps to filter out irrelevant regulatory provisions. Through this semantic understanding process, the user's unstructured question is converted into a structured representation of the query intent, encompassing key legal semantics, the entities involved, and the contextual context, providing accurate and rich input for subsequent searches.
[0031] For structured queries obtained through parsing, a "dual-engine" hybrid search strategy combining semantic vector search and keyword search is employed. First, the semantic vector engine performs semantic search, using a deep language model trained or fine-tuned on legal corpora to encode user questions into semantic vectors. Approximate nearest neighbor search is then performed on a pre-built regulatory text vector index to identify the top-N semantically similar legal provisions or knowledge nodes. Semantic vector search transcends superficial differences between user terms and regulatory wording, capturing deeper matches between user questions and legal provisions based on semantic similarity. Second, the keyword engine performs exact match search, extracting significant legal keywords from the query. Searches are then performed within the knowledge base using an inverted index or Boolean search to ensure that legal provisions directly relevant to the user's question are not missed. In practical implementations, a contextual filtering strategy is preferably applied first. Based on the contextual information obtained from the question parsing, the search scope is prioritized to the corresponding regional regulatory database or industry standard sub-database, thereby enhancing the relevance of search results and avoiding interference from irrelevant regional regulations. Semantic vector search and keyword search are then performed in parallel, with both engines running asynchronously and concurrently to improve efficiency. After the candidate result sets obtained by the two approaches are merged, they are sorted and output according to the relevance. During the fusion matching process, a comprehensive matching score is calculated for each candidate result, for example, using a linear weighted method to calculate Where, Score the overall match; Candidate entry Semantic matching score with user questions; Indicates the exact match score of the keyword; and It is a weight coefficient, ranging from 0 to 1, used to balance the contribution of semantic matching and keyword matching to the final score. and Based on the validation set results, the initial value was set to 0.5 to give equal importance to semantic and keyword searches. This dual-engine fusion retrieval method overcomes the limitations of a single search method. It not only expands the recall range through vector semantic search, but also ensures the relevance and accuracy of the results through precise keyword matching, thus achieving legal knowledge retrieval that balances high recall and high precision.
[0032] For the retrieved legal knowledge candidate results, a path fusion and reasoning mechanism is provided to support multi-hop reasoning on the knowledge graph and the interpretability of the results. First, when the user's question involves multiple legal points or requires cross-knowledge association, the system will perform a multi-hop path search on the knowledge graph to find relevant nodes and connect them in series to form a solution link. For example, for complex consumer dispute issues, the system may first locate the relevant legal clause node, and then find the typical case node connected to the legal clause through the graph, and finally reason out the answer based on the key points of the case's ruling. The path fusion module regards the associated paths formed by jumping along multiple nodes of the graph as different reasoning schemes, and calculates a path score for each candidate reasoning path to evaluate the relevance and credibility of the path to answering the current question. The path scoring formula is set based on factors such as the matching degree of each node on the path and the path length, for example, in the form of step-by-step cumulative weighting: ; Where, path Contains sequentially connected knowledge units The candidate entry matching score calculated in ; For the path The matching score of each node relative to the user's question; is the weight coefficient of the node in the path. The weight is determined by the direct relevance of the node to the question or its position in the reasoning chain. For example, the closer the node is to the user's question, the greater the weight is. An optional solution is to make the starting node The weight is attenuated by a certain proportion with each additional hop, and the specific attenuation coefficient is tuned through experiments. By scoring the matching degree of each link on the path and summarizing the weighted scores, the system selects the reasoning path with the highest score to generate the final answer. At the same time, the proportion of each node score in the total path score can also be used for explanation, that is, the contribution of each legal basis to the conclusion is output to the user, explaining how the system reached the answer through multi-step reasoning. This module gives the system the ability to chain reasoning for complex questions, and can integrate multiple laws and evidence for reasoning, improving the completeness, professionalism and explainability of the answers.
[0033] When generating responses to users, the system outputs the relevant legal provisions and their sources, clearly annotating the underlying regulations and article numbers in the answers, and including references to corresponding cases where necessary. Specifically, when integrated into a question-and-answer dialogue system, a dialogue generation model or pre-set template fuses search results to generate natural language answers, embedding legal reference tags or annotations within the answer sentences, allowing users to clearly identify the legal text underlying the conclusions. This legal reference mechanism ensures that the answers are well-documented, significantly enhancing the authority and credibility of the results. In pure search scenarios without a dialogue interface, the system can also directly return the original text of the matching legal provisions and highlight the key words, allowing users to quickly locate key information. By incorporating clear legal basis and case references into the results, this effectively overcomes the lack of supporting evidence in traditional intelligent question-and-answer systems, making the source of the answers clear to users and further enhancing their trust in the system's responses.
[0034] To ensure the timeliness and integrity of the knowledge graph, a knowledge base adaptive update mechanism is also provided. This mechanism runs in the background and includes three sub-processes: regulation monitoring, case collection, and incremental updates. First, regulation monitoring: This regularly monitors legal and regulatory databases, government announcements, and publicly available information sources from local consumer associations to automatically acquire newly enacted or revised laws, regulations, and policy documents, as well as newly added representative cases of consumer rights protection. Second, change detection: Newly acquired regulations are compared with older versions in the knowledge base to detect new, modified, or repealed clauses. Newly altered clauses or entries are marked as pending updates. Third, incremental knowledge extraction: The same information extraction process as for graph construction is repeated for modified or newly added regulations and case texts, generating newly added entity and relationship triples and inserting them into the knowledge graph. Node source labels and relationships are also updated. If changes affect existing nodes, node attributes are updated or relationships are adjusted to maintain graph consistency. To improve the accuracy of automatic extraction, semi-supervised learning is used to gradually optimize the extraction model. Updates to important legal clauses are also manually reviewed and confirmed before being stored. Finally, when the knowledge base content is significantly expanded, the retrieval model is updated, and the updated data is used to continue training or fine-tuning the semantic vector model so that it can adapt to the newly added terms and knowledge. If necessary, the keyword matching dictionary can also be retrained to ensure that new regulatory terms are also included in the search. Through the above-mentioned continuous update mechanism, the legal knowledge base of this system can keep up with changes in laws and policies, ensuring that search results are synchronized with the latest regulations, and overcoming the problems of outdated knowledge and incomplete coverage in traditional systems. Even when new local regulations are frequently introduced or new cases appear in the consumer field, the knowledge graph can be expanded in a timely manner to maintain high-quality support for user questions.
[0035] The system is deployed in an engineering manner according to a modular architecture, and computing resources and cross-platform compatibility are fully considered. In terms of model loading order, when the system starts, it is preferred to load various models and indexes in sequence according to the frequency of use and dependencies. For example, the NLP model and dictionary resources required for user question parsing are loaded first, and then the semantic vector retrieval model and its vector index are loaded. Then, the inverted index structure required for keyword retrieval is initialized, and finally the connection to the knowledge graph database is started. A reasonable loading order ensures that the system enters a serviceable state as soon as possible, avoiding delaying the loading of the model when the user asks a question for the first time. In terms of GPU and CPU resource judgment, the system automatically detects the GPU acceleration resources available in the operating environment at startup. If an available GPU is detected, the GPU version of the deep model is loaded and the GPU is used for semantic encoding and ANN search operations to accelerate vector retrieval and inference calculations. If the GPU resources are insufficient or unavailable, it automatically switches to CPU mode to load the streamlined model or reduce the batch size to ensure service stability. Regarding index loading strategies, the system uses memory mapping or lazy loading to pre-load necessary vector index segments into memory at startup, reducing real-time query read overhead. Large index files are also loaded on demand to avoid excessive memory usage. Structures such as inverted indexes can be serialized and stored and rebuilt or loaded into memory cache at startup, improving search speed. Regarding graph database connection pool management, since the retrieval and inference modules frequently access the graph database, connection pooling technology is used to maintain a certain number of database connections for reuse. The system dynamically adjusts the connection pool size based on concurrent requests, with the upper limit limited by database performance and hardware resources. If a connection is idle for a long time or an error occurs, the connection pool manager automatically recycles and recreates it to ensure healthy and available connections. Furthermore, for cross-module data exchange, asynchronous message queues or shared caches are used to reduce inter-module coupling and improve the system's concurrent processing capabilities. Regarding cross-platform compatibility, the system supports containerization and distributed deployment, integrating all components in Docker containers for easy deployment across different operating systems and server environments. Through container orchestration, the system's service modules can be elastically scaled to accommodate varying workloads. For example, during peak consulting periods, retrieval and inference module replicas can be automatically expanded to handle high-concurrency requests, while resources can be scaled back during off-peak hours to save costs. The overall architectural design adheres to the principles of loose coupling and microservices. Functional modules can be centrally deployed on a single server or distributed across multiple servers or even remotely in the cloud, meeting the performance requirements of large knowledge bases and high-concurrency scenarios.
[0036] Each module works together through a clear data interface and calling sequence to form a complete intelligent retrieval process. Figure 3As shown, when the user asks a question, the system calls the relevant modules in a predetermined order for processing. Some steps can be accelerated in parallel. First, the user question parsing module synchronously receives the user input and parses the question intention, and outputs a structured query to pass it to the retrieval module; then the hybrid retrieval module starts two engines to perform concurrent retrieval of the query: one engine calls the semantic vector model to calculate the query vector and retrieves similar terms in the regulatory vector index, and the other engine performs keyword exact matching retrieval at the same time. The retrieval process of the two engines is relatively independent and can be completed asynchronously and in parallel. After the retrieval engine is completed, the hybrid retrieval module summarizes the semantic and keyword candidate results, and The scoring mechanism is integrated and sorted, and the top-N candidates are passed to the reasoning module. Next, the path reasoning fusion module queries the relevant nodes from the knowledge graph database based on the candidate result set, performs multi-hop association search to expand the candidate answer path, and calculates the path for each path. Score and select the best path solution. This process requires access to the knowledge graph database to obtain the relationship between nodes, which is regarded as a serial synchronization process with the retrieval stage. If there are many candidate results, you can also consider evaluating multiple paths in parallel to reduce time. When the optimal path is determined, the result output module assembles the answer, generates a reply text based on the content of the legal terms and reasoning conclusions in the path, attaches the reference basis and returns it to the user. In this process, each module transmits information through a predefined data structure interface. The modules are mainly called synchronously in the order mentioned above, but an asynchronous concurrency mechanism is used within the retrieval to improve efficiency. The knowledge base update module runs independently in the background and does not affect the foreground query process. It performs updates asynchronously according to the set time interval or trigger conditions, and writes the update results to the knowledge graph database and index for the next retrieval by the foreground module. By Figure 3 The process shown realizes the full-link processing from user questions to precise matching of legal basis and then to reasoning and interpretation output.
[0037] The system also features a comprehensive exception handling mechanism, ensuring reliable operation and providing degraded service when module failures occur. If a core model fails to load during system startup, the system will detect the loading exception and immediately initiate a fallback plan, such as attempting to load a streamlined fallback model or switching to keyword-only search mode to ensure basic search functionality. Errors are logged and operations personnel are notified to investigate the issue. When a connection timeout or query failure occurs in the knowledge graph database, the connection pool management module activates a retry mechanism, retrying the connection or query after a short wait. If multiple retries fail, a circuit breaker policy is triggered, temporarily suspending requests to the database. During the circuit breaker period, the system implements contingency plans, such as reusing pre-cached knowledge data or limiting functions to only returning retrieved legal results without performing path reasoning, to ensure that the system remains partially responsive to user requests. For time-consuming modules such as search and reasoning, the system has an automatic fallback mechanism. When high concurrent request volume causes response delays, the system can temporarily streamline the process, such as narrowing the top-N vector search range or skipping complex multi-hop reasoning, prioritizing timely responses for core question-answering functionality. When a submodule repeatedly experiences anomalies, the system uses the service governance module to mark it as unavailable and isolate it, while also enabling a backup strategy to maintain uninterrupted external services. The system records detailed logs for all types of anomalies and sets monitoring alarms, promptly notifying maintenance personnel to intervene when module anomalies, response timeouts, and other situations occur. Through these mechanisms, the intelligent retrieval system is highly robust and fault-tolerant. Even when some models are unavailable or connectivity to external knowledge bases is limited, it can still provide basic services through automatic degradation, avoiding system crashes and improving long-term operational stability.
Claims
1. Construction of multi-source legal knowledge graph and intelligent retrieval method for consumer rights protection, characterized by: include: Construct a multi-source legal knowledge graph that integrates laws, regulations, industry norms, and case texts. Use an entity alignment mechanism to unify semantically equivalent legal concept nodes, and attach regional and industry source labels to the nodes to achieve integrated representation and traceability of cross-source knowledge. Parse user natural language questions to extract the legal entities, dispute elements, and regional and industry context constraints involved; A hybrid search strategy combining semantic vector search and keyword search is adopted, and the search scope is pre-filtered and limited in combination with the contextual constraints to generate a set of candidate legal clauses; Based on the knowledge graph, multi-hop path reasoning is performed on the candidate legal terms to generate an associated path with node contribution weights; the multi-hop path reasoning connects the legal terms node and the associated case node in the knowledge graph to form an inference chain, and calculates the comprehensive score of the path based on the path node matching degree and topological distance, where the starting node has a higher weight than the jump node; Output traceable answers with legal clause citations and case sources.
2. The method according to claim 1, wherein: During the construction of the legal knowledge graph, joint entity recognition and relationship extraction are performed on structured regulatory texts and unstructured case texts to generate triple knowledge units across data sources, and legal terminology specifications with different expressions but semantically equivalent are mapped to the same node of the knowledge graph through a synonym concept mapping mechanism.
3. The method according to claim 1, wherein: Based on a pre-built adaptive dictionary for the legal field, user spoken expressions are mapped into legal terms, and the extracted geographical and industry contexts are used as mandatory constraints in the retrieval stage to dynamically limit the subspace scope of regulatory retrieval.
4. The method according to claim 1, wherein: In the hybrid search strategy, there are two parallel sub-modules: a semantic vector search engine and a keyword search engine. The semantic vector search engine uses a deep language model to calculate the implicit semantic similarity between questions and legal terms, and the keyword search engine matches legal terms and term identifiers through an inverted index, and performs weighted fusion sorting on the two types of results according to an adjustable weight coefficient.
5. The method according to claim 4, characterized in that: The pre-filtering and limiting of the search scope is performed before the hybrid search begins, and the search authority of the corresponding regional legal and regulatory database or industry specification database is activated based on the context constraints obtained by the analysis.
6. The method according to claim 1, wherein: The comprehensive path score is obtained by accumulating the weighted matching scores of each node on the path, and the weight assigned to a node decreases monotonically as the jump distance between it and the user's question increases.
7. The method according to claim 1, wherein: In the process of generating traceable answers, the name of the regulation, article number and case source are embedded in the natural language answer text in the form of machine-readable tags; in pure search mode, results containing the original text of the legal provisions corresponding to the question are returned, and the keywords hit therein are visually highlighted.
8. The method according to claim 1, wherein: When constructing the legal knowledge graph, each node is assigned a pre-calculated confidence score, the confidence score being determined based on the type of data source corresponding to the node, the data source publishing, and the hierarchy of the adjudication body; In the hybrid search strategy, the confidence score is used as an adjustment factor for the search priority, and legal clauses corresponding to data sources with high confidence scores are returned first; In the multi-hop path reasoning, the confidence score is combined with the matching degree and topological distance of the nodes in the path to calculate the path comprehensive score, and a higher weight contribution is given to the nodes with high confidence scores.
9. The method according to claim 1, wherein: The method continuously monitors changes in legal data sources, updates graph nodes and relationships through incremental extraction, and optimizes semantic vector models and keyword dictionaries based on newly added data features.
10. A multi-source legal knowledge graph and intelligent retrieval system for consumer rights protection, characterized by: The implementation of the system is based on the method according to any one of claims 1 to 9, including: A knowledge graph construction module, configured to perform entity recognition, relationship extraction, and semantic alignment processing on multi-source heterogeneous data to construct the legal knowledge graph; The user question parsing module integrates legal domain dictionaries and synonym libraries to parse user natural language questions into structured queries with regional and industry constraints; A hybrid retrieval module, comprising a semantic vector retrieval submodule and a keyword retrieval submodule running in parallel, for performing semantic retrieval and keyword retrieval on the structured query and fusing the results; A path reasoning fusion module performs multi-hop association path search and score calculation on the retrieval candidate results based on the graph database to generate the legal basis link and determine the optimal reasoning path; A legal result output module, configured to generate a natural language answer including a legal clause reference and a case source identifier based on the optimal path reasoning result; The resource dynamic scheduling module is used to automatically switch the semantic model computing mode according to the hardware resource status, dynamically manage the graph database connection pool, and implement the retrieval retry mechanism in abnormal situations; The service fault-tolerant control module is used to interrupt the reasoning process and roll back the output of keyword search results when an exception occurs in the path reasoning function, and switch to pure keyword search mode when the semantic retrieval model fails.
Citation Information
Patent Citations
Inference method and device for multi-hop question and answer of knowledge graph and storage medium
CN118503366A
Method and system for enhancing RAG questions and answers through mixed retrieval method
CN118627625A
Information retrieval query method for legal data service platform
CN118981512A
Non-performing asset cross-scene question and answer framework based on knowledge graph
CN120216706A
Traffic event analysis method, device and equipment based on multi-hop causal path exploration
CN120372304A
Cited By
Retrieval system for bidding law document large language model
CN120994814A
A retrieval system for a large language model of legal documents for bidding
CN120994814B
Model reasoning method, device and equipment and computer readable storage medium
CN121072697A
A model inference method, device, equipment and computer readable storage medium
CN121072697B
Model training method and device, equipment, storage medium and product
CN121327197A