A legal service method and system based on a large language model and related devices

By constructing a pre-defined intelligent body database and using RAG retrieval technology, the shortcomings of large legal models in terms of complexity and professionalism have been addressed, enabling flexible responses to personalized legal issues and improving the intelligence and accuracy of legal services.

CN120407818BActive Publication Date: 2025-11-07EAST CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510921151.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-11-07
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

Existing legal models struggle to meet the complexities, specializations, diversity, and rigor of legal language, lacking the depth of legal knowledge and reasoning ability required to flexibly address individual needs, and their outputs lack accuracy and interpretability.

Method used

By constructing a pre-defined intelligent agent library, selecting pre-defined intelligent agents that meet category and domain information, and combining RAG retrieval and visualization technologies, the system generates target results for user requests, including multi-source heterogeneous big data collection, construction of multi-domain and multi-modal vector knowledge bases, application of multi-intelligent agents in legal scenarios, and data privacy and security protection.

Benefits of technology

It enables efficient and flexible legal services to address different legal issues, enhances the intelligence and accessibility of legal services, and ensures the accuracy and interpretability of the output results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407818B_ABST
    Figure CN120407818B_ABST
Patent Text Reader

Abstract

The application provides a legal service method and system based on a large language model and related equipment, and the method comprises the following steps: obtaining user request information, and identifying category information and field information of the user request information; selecting a preset intelligent agent meeting the category information and the field information from a preset intelligent agent library according to the category information and the field information, taking the preset intelligent agent meeting the category information and the field information as a target intelligent agent, and the target intelligent agent comprising a preset large language model meeting the field information and a preset function matrix meeting the category information; inputting the user request information into the target intelligent agent to obtain an initial result; querying enhanced information from a preset vector knowledge base through RAG retrieval technology based on the user request information; generating a target result for the user request information based on the enhanced information and the initial result, and displaying the target result through a visualization technology. The application can flexibly cope with different legal problems in the face of the diversity and individualized demand of legal services.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and legal technology, in particular to a legal service method and system based on a large language model and related equipment. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, large language models (LLM) are increasingly widely used in various fields. In the legal field, as a highly specialized and standardized industry, it has also begun to explore the use of large language models to improve the level of legal service intelligence. Traditional legal service models mainly rely on manual processing, which has low efficiency, high cost, limited service coverage and other problems, and is difficult to meet the growing demand for legal services.

[0003] In recent years, domestic and foreign research institutions and enterprises have invested in the development of legal large models.

[0004] Although there have been many achievements, the current legal large model still faces some challenges. First, the complexity and professionalism of legal language require the model to have deep legal knowledge and reasoning ability. Second, the diversity and individualized needs of legal services require the model to be able to flexibly respond to different legal issues. Finally, the rigor and authority of legal services require the model's output to have a high degree of accuracy and explainability. SUMMARY

[0005] In order to solve the problems in the prior art, the present application provides a legal service method and system based on a large language model and related equipment to solve the problems of complexity and professionalism of legal language requiring the model to have deep legal knowledge and reasoning ability, diversity and individualized needs of legal services requiring the model to be able to flexibly respond to different legal issues, and rigor and authority of legal services requiring the model's output to have a high degree of accuracy and explainability.

[0006] To solve the above technical problems, the present application adopts the following technical solutions.

[0007] In a first aspect, the present application provides a legal service method based on a large language model, comprising the following steps:

[0008] Obtain user request information and identify the category information and domain information of the user request information;

[0009] Select a preset agent that meets the category information and the domain information from a preset agent library according to the category information and the domain information, and use the preset agent that meets the category information and the domain information as a target agent, the target agent including a preset large language model that meets the domain information and a preset function matrix that meets the category information;

[0010] inputting the user request information into the target agent to obtain an initial result;

[0011] based on the user request information, querying enhanced information in a preset vector knowledge base through RAG retrieval technology;

[0012] generating a target result for the user request information based on the enhanced information and the initial result, and displaying the target result through visualization technology.

[0013] Optionally, the step of identifying the category information and the field information of the user request information comprises:

[0014] identifying the user request information through intent recognition technology to obtain intent information of the user request information, and determining the category information and the field information of the user request information according to the intent information;

[0015] Alternatively,

[0016] obtaining source information of the user request information, and determining the category information and the field information of the user request information based on the source information.

[0017] Optionally, before the step of selecting a preset agent satisfying the category information and the field information from a preset agent library according to the category information and the field information, the method further comprises:

[0018] based on a preset distributed data collection network, collecting original data in different legal fields through data collection technology, and labeling the original data by taking source information of the original data as label information to obtain an initial data set, wherein the data collection technology at least includes text analysis technology, image feature extraction technology and video frame analysis technology;

[0019] performing data analysis on each original data in the initial data set to obtain an analysis result, and filtering each original data in the initial data set according to the analysis result to obtain target data satisfying a preset requirement, wherein the data analysis includes comparative analysis, time series analysis, structure analysis, RFM analysis, grouping analysis and regression analysis;

[0020] performing format conversion on the target data to obtain a JSON-LD format data set containing three dimensions of fact elements, legal application and judgment result, wherein when the target data is unstructured data, the target data is converted into a JSON-LD format data containing three dimensions of fact elements, legal application and judgment result through entity recognition technology, relationship extraction technology and semantic disambiguation technology;

[0021] The initial large language model for different legal fields is trained according to a preset tuning algorithm based on the JSON-LD format data set, to obtain a preset large language model for different legal fields, and the preset tuning algorithm includes at least one of a LoRA algorithm, an SFT algorithm, a DPO algorithm, a P-Tuning algorithm, an RAG-FT algorithm, an RLHF algorithm, and a KTO algorithm.

[0022] A preset agent library is constructed according to the preset large language model for different legal fields and a preset function matrix for different categories of information, and the preset function matrix includes a first matrix based on a semantic analysis engine of BERT-Judicial version, a second matrix based on a similarity calculation model integrating space-time elements, a third matrix based on GPT-4 and a legal format verifier, a fourth matrix integrating adversarial debate and evidence chain analysis algorithm, and a fifth matrix based on multi-language legal alignment technology, and the initial large language model includes a Qwen-14B model for civil dispute field, a DeepSeek-legal model for commercial dispute field, a ChatGLM3-6B model for criminal defense field, and a LawyerLlama model for administrative dispute field.

[0023] Optionally, the step of training the initial large language model for different legal fields according to a preset tuning algorithm based on the JSON-LD format data set to obtain a preset large language model for different legal fields includes:

[0024] The proportioning of data in the JSON-LD format data set is optimized by domain term injection to obtain a target JSON-LD format data set;

[0025] The initial large language model for different legal fields is trained according to a preset tuning algorithm based on the target JSON-LD format data set to obtain a secondary large language model for different legal fields;

[0026] The secondary large language model is dimensionally calibrated with a cosine similarity greater than 0.85 as a standard to obtain a preset large language model for different legal fields.

[0027] Optionally, before the step of querying enhanced information in a preset vector knowledge base through RAG retrieval technology based on the user request information, the method further includes:

[0028] The original data is integrated into a statute library, a judge library, a lawyer library, and a case library according to the source of the original data;

[0029] According to the legal element-based association index system, the data in the legal provision library, the judge library, the lawyer library and the case library are associated to obtain multi-dimensional semantic association data with an association density greater than 85%;

[0030] A private embedding model is adopted to construct a unified vector space mapping system of text, image and video, and a feature vector of the multi-dimensional semantic association data is extracted according to the unified vector space mapping system, and a preset vector knowledge base is established based on the feature vector.

[0031] Optionally, the step of querying the enhanced information in the preset vector knowledge base based on the user request information through the RAG retrieval technology comprises:

[0032] The user request information is parsed to obtain a query feature vector;

[0033] The cosine similarity of the query feature vector and the feature vector in the preset vector knowledge base is calculated, and the feature vector with a cosine similarity greater than or equal to a cosine similarity threshold is identified, the identified feature vectors are sorted according to the cosine similarity to obtain a feature vector queue, and the first ten feature vectors in the feature vector queue are selected as target feature vectors;

[0034] Document data corresponding to the target feature vectors is selected from the legal provision library, the judge library, the lawyer library and the case library;

[0035] The document data is subjected to legal logic verification to exclude invalid legal provision references to obtain target document data;

[0036] A retrieval report with judgment gist annotation is generated based on the target document data, and the retrieval report with judgment gist annotation is taken as enhanced information.

[0037] Optionally, the method further comprises: when there are at least two preset agents in the preset agent library that satisfy the category information and the field information, establishing a data transmission channel between the preset agents according to a preset rule to enable evidence chain sharing between the preset agents; when there are at least two preset agents in the preset agent library that satisfy the category information and the field information, and the output results of the preset agents conflict, optimizing the output results through a federated learning algorithm to obtain a target result for the user request information.

[0038] In another aspect, the present application also provides a legal service system based on a large language model, the system comprising:

[0039] A request receiving module is configured to obtain user request information and identify the category information and the field information of the user request information;

[0040] The selecting module is configured to select a preset agent satisfying the category information and the field information from a preset agent library according to the category information and the field information, and take the preset agent satisfying the category information and the field information as a target agent, wherein the target agent comprises a preset large language model satisfying the field information and a preset function matrix satisfying the category information.

[0041] The executing module is configured to input the user request information into the target agent to obtain an initial result.

[0042] The enhancing module is configured to query enhancement information from a preset vector knowledge base based on the user request information by using a RAG retrieval technology.

[0043] The displaying module is configured to generate a target result for the user request information based on the enhancement information and the initial result, and display the target result by using a visualization technology.

[0044] In a third aspect, the present application also provides an electronic device, which comprises a memory, a processor, and a first computer program stored in the memory and executable on the processor, and the processor implements the large language model-based legal service method when executing the first computer program.

[0045] In a fourth aspect, the present application also provides a computer readable storage medium, which stores a second computer program, and the second computer program implements the large language model-based legal service method when executed by a processor.

[0046] Advantages:

[0047] The present application can obtain user request information, identify category information of the user request information, select a preset agent corresponding to the category information from a preset agent library according to the category information to obtain a target agent, input the user request information into the target agent to obtain an initial result, query enhancement information from a preset vector knowledge base based on the user request information by using a RAG retrieval technology, generate a target result for the user request information based on the enhancement information and the initial result, and display the target result by using a visualization technology. The present application can flexibly cope with different legal problems in the face of the diversity and individualized demand of legal services. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 is a running flowchart of a large language model-based legal service method of the present application;

[0049] Figure 2 is an architecture diagram of a legal service method based on a large language model in the present application;

[0050] Figure 3 is a multi-source heterogeneous big data collection and unified analysis technology roadmap in the present application;

[0051] Figure 4 is a large model training technology architecture diagram based on high-quality sample data in the present application;

[0052] Figure 5 is a multi-domain multi-modal vector knowledge base and RAG enhanced retrieval technology roadmap in the present application;

[0053] Figure 6 is a legal full-scene multi-agent application technology roadmap in the present application;

[0054] Figure 7 is a legal digital visualization technology roadmap based on a large model in the present application;

[0055] Figure 8 is a data privacy and security guarantee technology roadmap for a legal service large model base in the present application;

[0056] Figure 9 is a structure schematic diagram of a legal service system based on a large language model in an embodiment of the present application;

[0057] Figure 10 is a structure schematic diagram of an electronic device provided by an embodiment of the present application;

[0058] Figure 11 is a structure schematic diagram of a computer readable storage medium provided by an embodiment of the present application. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application.

[0060] The embodiments described in the present application are only embodiments of part of the present application, not all embodiments. All other embodiments obtained by those skilled in the art without creative labor based on the spirit of the present application are within the protection scope of the present application.

[0061] Through the integration of the above functions, the legal large model system will effectively improve the efficiency and quality of legal services, and promote the intelligentization and universalization development of legal services.

[0062] This invention, based on the collaborative work of multiple key steps, is used for the collection and unified analysis of multi-source heterogeneous big data, for large-scale model training with high-quality sample data, for multi-domain and multi-modal vector knowledge bases and RAG-enhanced retrieval, for multi-agent applications in all legal scenarios, for legal digital visualization, and for data privacy and security. It realizes a comprehensive legal big data system integrating functions such as judicial big data, legal question answering, case retrieval, document generation, foreign-related rule of law, intelligent courts, case visualization, and security compliance. The aim is to address issues such as the complexity and professionalism of legal language requiring models to possess profound legal knowledge and reasoning abilities, the diversity and personalized needs of legal services requiring models to flexibly respond to different legal issues, and the rigor and authority of legal services requiring models to have highly accurate and interpretable output results.

[0063] See Figure 1 As shown, the present invention provides a legal service method based on a large language model, comprising the following steps:

[0064] S110. Obtain user request information and identify the category information and domain information of the user request information;

[0065] S120. Select a preset intelligent agent that satisfies the category information and the domain information from the preset intelligent agent library according to the category information and the domain information, and take the preset intelligent agent that satisfies the category information and the domain information as the target intelligent agent. The target intelligent agent includes a preset large language model that satisfies the domain information and a preset function matrix that satisfies the category information.

[0066] S130. Input the user request information into the target intelligent agent to obtain an initial result;

[0067] S140. Based on the user request information, retrieve enhanced information from a preset vector knowledge base using RAG retrieval technology;

[0068] S150. Generate a target result for the user request information based on the enhanced information and the initial result, and display the target result using visualization technology.

[0069] For example, such as Figure 2 As shown, the system architecture for implementing the legal service method based on a large language model provided by this invention includes:

[0070] Multi-source heterogeneous big data collection and unified analysis; large model training based on high-quality sample data; multi-domain and multimodal vector knowledge base and RAG-enhanced retrieval; intelligent agent applications; legal digital visualization based on large models; data privacy and security assurance.

[0071] In a possible implementation, the step of identifying the category information and the field information of the user request information comprises:

[0072] The user request information is identified through an intent recognition technology to obtain intent information of the user request information, and the category information and the field information of the user request information are determined according to the intent information.

[0073] Alternatively,

[0074] The source information of the user request information is obtained, and the category information and the field information of the user request information are determined based on the source information.

[0075] For example, when the user inputs the user request information through a specific button on a specific page, if the user selects a “civil dispute consultation” button on the page, the field information of the user request information is determined to be the civil field, and the type information of the user request information is determined to be the legal question and answer type according to the source information (button) of the user request information.

[0076] For example, when the user inputs a request through voice, the user voice information is converted into text information through a voice conversion technology, and then the category information and the field information of the user request information are determined according to the text information through an intent recognition technology.

[0077] In a possible implementation, before the step of selecting the preset agent satisfying the category information and the field information from the preset agent library according to the category information and the field information, the method further comprises:

[0078] Based on a preset distributed data collection network, raw data in different legal fields are collected through a data collection technology, and the source information of the raw data is used as label information to label the raw data to obtain an initial data set, and the data collection technology at least includes a text analysis technology, an image feature extraction technology and a video frame analysis technology.

[0079] The raw data in the initial data set are analyzed to obtain an analysis result, and the raw data in the initial data set are filtered according to the analysis result to obtain target data satisfying a preset requirement, and the data analysis includes comparative analysis, time series analysis, structure analysis, RFM analysis, grouping analysis and regression analysis.

[0080] The target data is format-converted to obtain a JSON-LD format data set containing three dimensions of fact elements, legal application, and judgment results, wherein, when the target data is unstructured data, the target data is converted into a JSON-LD format data set containing three dimensions of fact elements, legal application, and judgment results through entity recognition technology, relationship extraction technology, and semantic disambiguation technology.

[0081] Based on the JSON-LD format data set, an initial large language model for different legal fields is trained according to a preset tuning algorithm to obtain a preset large language model for different legal fields, and the preset tuning algorithm includes at least one of LoRA algorithm, SFT algorithm, DPO algorithm, P-Tuning algorithm, RAG-FT algorithm, RLHF algorithm, and KTO algorithm.

[0082] A preset agent library is constructed according to the preset large language model for different legal fields and a preset function matrix for different categories of information, the preset function matrix includes a first matrix based on a semantic analysis engine of BERT-judicial version, a second matrix based on a similarity calculation model integrating time and space elements, a third matrix based on GPT-4 and a legal format checker, a fourth matrix integrating adversarial argumentation and evidence chain analysis algorithm, and a fifth matrix based on multi-language legal alignment technology, and the initial large language model includes a Qwen-14B model for civil dispute field, a DeepSeek-legal model for commercial dispute field, a ChatGLM3-6B model for criminal defense field, and a LawyerLlama model for administrative dispute field.

[0083] As shown in the example, Figure 3 When data collection is performed, a distributed data collection network covering four fields of civil dispute, commercial dispute, criminal defense, and administrative dispute is built, and technical modules such as text analysis, image feature extraction, and video frame analysis are deployed to realize full-form data collection of structured and unstructured data in legal practice scenarios, and a data source labeling and quality preliminary screening mechanism is simultaneously established. A field data fusion system is constructed. A raw data pool of four fields of civil, commercial, criminal, and administrative is established, a distributed crawler system is deployed to realize automatic collection of 15+ data sources such as judgment documents network and law firm case library, a data blood relationship tracking mechanism is designed to record 40+ metadata tags such as case cause, trial level, and region of each case, and a field-specific data reserve library with a daily incremental capacity of 2TB is formed.

[0084] The integrated law library (laws and regulations and judicial interpretations), the judge library (judgment rules and trial experience), the lawyer library (agent cases and practical strategies), the case library (judicial precedents and case characteristics), the foreign-related library (foreign laws and regulations and cross-border cases), the document library (litigation documents and contract document templates), and the enterprise library (identity information and activity data) seven databases are designed to associate rules of knowledge graphs, and complete standardization cleaning and metadata labeling of raw data.

[0085] In data analysis, six types of core algorithm modules, including comparative analysis (standard conflict detection), time series analysis (judicial decision prediction), structural analysis (case element deconstruction), RFM analysis (litigation risk stratification assessment), grouping analysis (judge's decision preference clustering), and regression analysis (compensation amount prediction model), are configured in the unified analysis platform to establish a multi-dimensional analysis index system and a visual interactive interface.

[0086] As shown in Figure 3 The multi-source heterogeneous data collection and unified analysis process is as follows: the bottom layer collects data, which is quality checked and then fed to the middle layer database; the middle layer structured data is pushed to the top layer analysis engine through the API interface; the analysis results generated by the top layer, such as case feature vectors and judgment prediction coefficients, are fed back to the database to optimize the data tag system, forming a closed-loop optimization link of "data cleaning-knowledge precipitation-intelligent feedback".

[0087] The bottom-up data integration and intelligent analysis lay the foundation for the legal large model system. The bottom layer relies on a multi-source heterogeneous data collection system, covering text, images, video, and other diversified data forms in the four legal fields of civil, commercial, criminal, and administrative, achieving full-dimensional coverage of legal practice scenarios; the middle layer deploys a structured data collection platform, integrating seven core databases of law library, judge library, lawyer library, case library, foreign-related library, document library, and enterprise library, and converting raw data into computable resources through standardized processing; the top layer builds a unified analysis platform, providing six types of intelligent analysis methods, including comparative analysis, grouping analysis, structural analysis, RFM analysis, time series analysis, and regression analysis, forming a technical closed loop from data collection, governance to multi-modal analysis. The three-layer architecture realizes dynamic interaction through bidirectional data flow arrows, supporting the upward convergence and cleaning of bottom-layer heterogeneous data, and ensuring the feedback of analysis results to optimize data quality, ultimately providing the legal large model system with multi-agent collaboration with broad and deep data support and quantitative analysis capabilities.

[0088] Exemplarily, by adopting a text structuring engine based on a BERT-law version, through a three-layer processing architecture of entity recognition (parties, citation of laws, etc.), relationship extraction (legal fact determination logic), and semantic disambiguation (standardization of professional terms), an unstructured legal document is converted into a JSON-LD format standard data set containing three dimensions of fact elements, legal application, and judgment results.

[0089] In a possible implementation, the step of training, based on the JSON-LD format data set, an initial large language model for different legal fields according to a preset tuning algorithm to obtain a preset large language model for different legal fields includes:

[0090] The data ratio in the JSON-LD format data set is optimized by domain term injection to obtain a target JSON-LD format data set.

[0091] Based on the target JSON-LD format data set, an initial large language model for different legal fields is trained according to a preset tuning algorithm to obtain a secondary large language model for different legal fields.

[0092] The secondary large language model is calibrated in a dimensional space with a cosine similarity greater than 0.85 as a standard to obtain a preset large language model for different legal fields.

[0093] Exemplarily, as shown in Figure 4 Qwen-14B model is selected for civil disputes, DeepSeek-legal model is selected for commercial disputes, ChatGLM3-6B model is selected for criminal defense, and LawyerLlama model is selected for administrative disputes. The above models are subjected to three-stage adaptation and modification (more than 2 million legal word tables are loaded through domain term injection, the proportion of pre-training data is optimized to increase the proportion of domain data to 65%, and the cosine similarity is greater than 0.85 through dimensional space calibration), and after training, a preset large language model for different legal fields is obtained.

[0094] Exemplarily, when facing different user request information, that is, when different preset large language models are running, LoRA algorithm for low-rank adaptation, SFT algorithm for supervised fine-tuning, DPO algorithm for direct preference optimization, P-Tuning algorithm for prompt tuning, RAG-FT algorithm for retrieval enhancement fine-tuning, RLHF algorithm for human feedback reinforcement learning, and KTO algorithm for knowledge distillation are dynamically combined for tuning.

[0095] The precise adaptation of the large model in the legal field is realized through the whole-link optimization from the data source to the algorithm application. Relying on the historical case processing data in the four legal fields of civil, commercial, criminal and administrative, original data resources covering all business scenarios are provided. The data layer converts massive unstructured data into standard data sets through a selection and cleaning mechanism to ensure the quality benchmark of training samples. The model layer uses general large models such as Qwen or Deepseek as the base to generate four special large models in the fields of civil, commercial, criminal and administrative. The training layer innovatively integrates seven algorithm modules such as LoRA, SFT and Pre-Training to form a composite optimization system, and realizes the deep injection of legal professional ability through four directional optimization channels such as civil dispute fine-tuning. Through the technical transmission path between each level, a bidirectional enhancement mechanism is formed, which not only maintains the general ability of the basic model but also strengthens the knowledge density of the professional field, finally building an intelligent body collaborative training paradigm with both legal reasoning accuracy and cross-field adaptability, providing scalable model support for multi-task legal services.

[0096] In one possible implementation, before the step of querying the enhanced information in the preset vector knowledge base through the RAG retrieval technology based on the user request information, the method further includes:

[0097] According to the source of the original data, the original data is integrated into a law library, a judge library, a lawyer library, and a case library;

[0098] According to the correlation index system based on legal elements, the data in the law library, the judge library, the lawyer library, and the case library are correlated to obtain multi-dimensional semantic correlation data with a correlation density greater than 85%;

[0099] A private embedding model is used to construct a unified vector space mapping system for text, images, and videos, and feature vectors of the multi-dimensional semantic correlation data are extracted according to the unified vector space mapping system, and a preset vector knowledge base is established based on the feature vectors.

[0100] For example, Figure 5As shown, the integrated law library (containing more than 4.5 million current laws and regulations), the judge library (covering the analysis of the judicial tendency of more than 400,000 judges), the lawyer library (more than 700,000 case strategy), and the case library (more than 150 million judicial documents) four core databases, establish an association index system based on legal elements (case-by-law-judgment result triplets), realize multi-dimensional semantic association of knowledge nodes (association density > 85%), and obtain multi-dimensional semantic association data; using Voyage-large-2-instruct (text semantic coding), SFR-Embedding-Mistral (cross-modal alignment), and BGE-M3 (multilingual adaptation) three types of private embedding models, a unified vector space mapping system of text (TXT / DOCX format), image (scanned archives / signature samples), and video (court trial video) is constructed, and efficient extraction (coding delay < 150ms) of 1280-dimensional vector of legal semantic features is realized, and a vector knowledge base of multi-field and multi-modal is obtained.

[0101] For example, for civil and commercial (contract clause deconstruction), criminal (electronic evidence chain analysis), etc. fields, special data processing modules are deployed: structured data uses regular template extraction (accuracy > 95%), and unstructured data is analyzed by LayoutLMv3 for document layout analysis, and a field-specific feature-enhanced vector space is constructed (in-class similarity improved by 23%).

[0102] In one possible implementation, the step of querying the enhanced information in the preset vector knowledge base based on the user request information through the RAG retrieval technology comprises:

[0103] The user request information is parsed to obtain a query feature vector;

[0104] The cosine similarity of the query feature vector and the feature vectors in the preset vector knowledge base is calculated, and the feature vectors with cosine similarity greater than or equal to a cosine similarity threshold are identified, the identified feature vectors are sorted according to the cosine similarity, a feature vector queue is obtained, and the first ten feature vectors in the feature vector queue are selected as target feature vectors;

[0105] Selecting document data corresponding to the target feature vectors in the law library, the judge library, the lawyer library, and the case library;

[0106] Performing legal logic verification on the document data to exclude invalid law references, and obtaining target document data;

[0107] Generating a retrieval report with judgment gist annotation based on the target document data, and taking the retrieval report with judgment gist annotation as the enhanced information.

[0108] For example, the user input is parsed using a search enhancement generation pipeline, and the parsing process supports OCR recognition of handwritten complaints;

[0109] Multi-modal query vectorization is performed to find query results (target feature vectors) with a cosine similarity threshold of ≥ 0.78;

[0110] Knowledge base candidate set recall, specifically, selecting cases / statutes (document data) corresponding to the target feature vector from the statute library, the judge library, the lawyer library, and the case library;

[0111] The CONTEXT module is used for legal logic verification to exclude invalid statute references;

[0112] Based on the target document data, a search report with judgment gist annotations is generated.

[0113] Specifically, the blue (data layer), purple (algorithm layer), and yellow (application layer) color blocks are used to distinguish functional modules: the data layer deploys MySQL clusters (structured data) and MinIO object storage (unstructured data); the algorithm layer runs Faiss vector indexing (10 billion-level vector retrieval) and BERT-legal version similarity calculation; the application layer provides an API gateway (QPS ≥ 500) and a visual search interface (supports three-dimensional knowledge graph navigation).

[0114] By integrating legal professional databases and advanced vectorization technology, precise information matching is achieved. The system integrates four core databases, including the statute library, the judge library, the lawyer library, and the case library, to form a structured knowledge base. Users can input text or images to trigger the search process. The system uses Voyage-large-2-instruct, SFR-Embedding-Mistral, and BGE-M3 private embedding models to perform multi-modal vector encoding on the input content. It simultaneously converts heterogeneous data such as text, images, and videos in the knowledge base into a unified vector space representation. Through comparative learning and similarity calculation, dynamic matching of input features and knowledge base resources is achieved. The CONTEXT module uses RAG technology to enhance the search results, outputting high-precision search content that meets legal professional needs. The technology foundation is deployed in layers to handle multi-domain data processing modules, covering structured and unstructured data in vertical fields such as civil disputes, commercial disputes, criminal defense, and administrative disputes, forming a complete technology chain from raw data cleaning to multi-modal vectorization and intelligent search. The system uses a modular architecture and color block coding to clearly show the synergy between the data layer, algorithm layer, and application layer. Blue, purple, and yellow (colors are not limited to specific colors, but are used to distinguish functions) color blocks distinguish the work processes of different functional modules, ensuring deep integration of legal professional data and strengthening cross-modal information association retrieval capabilities, providing accurate and reliable knowledge support for multi-agent collaboration.

[0115] In a possible implementation, the method further includes: when there are at least two preset agents in the preset agent library that meet the category information and the field information, establishing a data transmission channel between the preset agents according to a preset rule to enable evidence chain sharing between the preset agents; and when there are at least two preset agents in the preset agent library that meet the category information and the field information, and output results of the preset agents conflict, optimizing the output results by using a federated learning algorithm to obtain a target result for the user request information.

[0116] As shown in the example, Figure 6 An intelligent scheduling hub of the overall Agent is constructed, which is used to ensure that the retrieval path calls the RAG enhancement system in real time to obtain the latest legal basis; the intelligent scheduling hub is also used to dynamically load a field fine-tuning model in a training path, for example, to update the parameter weight every week; through preset task routing rules of 23 types of cases such as civil, criminal, and administrative cases, the intelligent distribution with a request recognition accuracy of >95% is realized, for example, a lending dispute is automatically routed to a similar case retrieval Agent and a document generation Agent for cooperative processing.

[0117] As an example, a bidirectional interaction channel is built: a vertical channel realizes real-time data exchange between the Agent and the knowledge base; a horizontal channel establishes an evidence chain sharing mechanism between Agents, for example, a dispute focus found by a legal question and answer Agent is automatically synchronized to an intelligent court Agent; at the same time, when the conclusions of multiple Agents conflict, a federated learning mechanism is started to optimize the decision.

[0118] Specifically, a hierarchical architecture design is adopted: a knowledge layer deploys a graph database (Neo4j) and a vector database (Milvus) for mixed storage; an Agent layer manages an Agent cluster through Kubernetes orchestration (application layer development of Web, mobile terminal, and API three terminal entrances, design of a legal consultation service matrix board.

[0119] As an example, through microcirculation, user feedback data (adoption rate, correction record) is used to update the Agent strategy network every day; through macrocirculation, cross-field knowledge distillation (criminal → commercial strategy migration success rate >65%) is implemented every month, and the A / B test is used to optimize the cooperative strategy.

[0120] The organic linkage of the agent cluster and the knowledge base is realized through the overall Agent. The core of the system is composed of six special knowledge bases, including statute base, case base, document base, judge base, lawyer base and foreign-related base, which are connected with five functional Agent modules, such as legal Q&A, case search, document generation, intelligent court and foreign-related legal system, to form a service matrix covering the whole scene of legal practice. When the user initiates a service request through the five entrances of legal Q&A and case search on the application end, the overall Agent performs task scheduling based on the customized search and training double path, and drives the corresponding intelligent agent to call the associated knowledge base resources to perform professional operations: the legal Q&A Agent analyzes the statutes and judicial interpretations in the statute base in real time, the case search Agent accurately matches the historical cases in the case base, the document generation Agent intelligently calls the format specifications in the document base, the intelligent court Agent integrates the practical experience of the judge and lawyer base, and the foreign-related legal system Agent connects the international legal resources in the foreign-related base. Through the two-way interaction channel, each intelligent agent maintains dynamic cooperation with the overall Agent, not only realizing the accurate calling of knowledge base resources, but also completing the feedback optimization of task execution results. The system adopts a modular architecture to clearly show the technical transmission path of the knowledge layer, the intelligent agent layer and the application layer, and through the continuous iteration of the customized training mechanism, the multi-agent cooperation efficiency is improved, and finally a legal whole-process solution with professional depth and scene coverage is formed.

[0121] The method further includes: Figure 7 As shown, the electronic evidence and restoration module (supporting FTK / EnCase forensic tool chain), storage medium fixation and analysis device (including hard disk hash verification and read-only interface), mass data acquisition system (adopting script to download public data, covering 20+ data sources for real-time crawling), program reverse analysis sandbox (dynamic analysis of Windows / Linux binary files), multi-modal data processing engine (EXIF / document metadata stripping) and digital simulation interface are deployed to form a full-chain technical support covering data tracing, evidence fixation and cleaning, and are connected with the lower layer system through a secure communication protocol.

[0122] Based on the distributed storage architecture, general data (Internet public information) and legal industry data (8 types of judicial documents and court trial videos) are integrated to build a fusion analysis network containing natural language processing (BERT-judicial version), intelligent voice (Wav2Vec 2.0 legal speech recognition) and computer vision (YOLOv5 evidence positioning). Through machine learning (XGBoost evidence chain evaluation) and search enhancement technology (RAG legal knowledge injection), multi-modal data joint analysis is realized, and 12-dimensional quantitative indicators such as legal requirement matching degree (F1>0.88) and behavior pattern correlation degree are output.

[0123] A three-dimensional legal scene mirror is developed using Unreal Engine 5, a digital twin module (1:1 restoration of hardware devices involved in the case), an image and video analyzer (reconstruction of the space-time trajectory of electronic evidence), a hardware acceleration rendering interface (NVIDIA Omniverse real-time collaboration), and a communication simulation protocol to reproduce network attack links. An interactive evidence report is generated by combining natural language processing (supports VR / AR multi-end viewing), forming a closed-loop process of data collection, intelligent analysis, and scene reconstruction, and achieving a visual decision support capability with an electronic evidence dynamic restoration error rate of less than 3% and a court trial deduction efficiency improvement of 55%.

[0124] The deep analysis and three-dimensional presentation of legal digital evidence are achieved through the cooperation of security agents, AI engines, and visualization platforms. A security agent cluster is deployed at the top level, integrating six core modules: electronic forensics and restoration, storage medium fixation and analysis, mass data collection, program reverse analysis, multi-modal data processing, and digital simulation, forming a full-chain technical support covering data tracing, evidence fixation, and cleaning. The middle layer is driven by an AI engine, which builds an intelligent analysis hub based on natural language processing, neural networks, intelligent speech, computer vision, and machine question answering. Through distributed storage and computing architecture, it integrates general data, legal industry data, and intelligent cloud platform resources, and realizes the fusion analysis of multi-modal data through machine learning, deep learning, and search enhancement technology. The bottom layer builds a visualization digital simulation platform, which reconstructs the digital mirror of legal scenes based on binary reverse, electronic forensics, digital twinning, image and video analysis, and other technologies, and realizes the dynamic restoration of electronic evidence through hardware simulation and communication simulation. With the help of natural language processing and hardware acceleration technology, the efficiency of three-dimensional modeling of legal data is improved. The system adopts a blue and white modular architecture to clearly divide the functional levels. The security agent and the AI engine realize intelligent filtering and feature extraction of legal data through bidirectional data flow. The visualization platform converts the analysis results into an interactive simulation model through multi-dimensional rendering technology, forming a closed-loop process from data collection, intelligent analysis to scene reconstruction, and providing a visual analysis solution for legal practice with technical credibility and decision support value.

[0125] The method further includes: Figure 8 As shown, an irreversible ciphertext interaction system based on multi-party secure computation (MPC) is deployed to establish a four-stage closed loop covering encrypted model training (homomorphic encryption parameter fine-tuning), secure deployment (TEE trusted execution environment), ciphertext calculation request (access control based on SM9 algorithm), and ciphertext result output (differential privacy perturbation). A distributed secure computing node cluster (1000+ core parallel computing capability) is designed to realize secure interaction between data providers and users through a federated learning architecture (sensitive information leakage rate <0.01%), supporting 500,000 legal data call requests per day.

[0126] The LLM privacy protection system is established, including: an adaptive privacy perception module (based on BiLSTM sensitivity grading, with an identification accuracy of >95%); an intelligent cleaning engine (regular expression + deep learning joint desensitization); a privacy attribute modeling (Word2Vec + knowledge graph construction of 1300-dimensional legal privacy feature space); a leakage risk assessment model (sensitive exposure index calculation based on Monte Carlo simulation); a secure data publishing gateway (supporting K-anonymity and L-diversity policy dynamic configuration); and an audit tracking system (blockchain storage of key operation logs), realizing real-time desensitization processing of legal texts (desensitization efficiency >500 / second) and individualized protection policy configuration (supporting 32 kinds of combination templates).

[0127] In one possible implementation, the training phase privacy computing framework injects encryption parameters (model privacy budget ε<1.0), and the inference phase privacy protection system monitors data flow in real time (detection delay <50ms); a risk transmission blocking module is developed (based on Causal Inference to locate the leakage path), which automatically triggers data isolation and model hot update (response time <3s) when sensitive information exposure risk (CVSS score ≥7.0) is detected, ultimately forming a full-chain security closed loop covering data collection (GB / T 35273 compliance), processing (ISO / IEC 27001 certification), and application (GDPR cross-border transmission protection), supporting the compliance operation of the "Zhi Zhe Qian Lu" large model handling 100+ legal consultation services per day.

[0128] Through the synergistic effect of privacy computing and adaptive security mechanism, the whole-process data security management and control is realized. The left side of the system establishes a privacy computing framework based on irreversible ciphertext interaction, which runs through the four-stage core processes of encrypted model training, deployment, computing request and operation result, and the ciphertext interaction of data providers and users is coordinated by the security computing node, which blocks the risk of original sensitive information leakage while ensuring data availability; the right side builds a large language model dedicated privacy protection system, dynamically identifies data privacy sensitivity through an adaptive privacy perception module, desensitizes the input information through sensitive data intelligent cleaning technology, establishes a privacy protection modeling method relying on a privacy attribute discovery algorithm, and quantifies the sensitive information exposure index by using a privacy data leakage risk assessment model, finally forming a personalized and customizable secure data publishing mechanism. The dual-track system realizes deep integration in the whole life cycle of data, the privacy computing framework provides an encrypted operation environment for model training and inference, and the large model security system optimizes data quality through dynamic evaluation and cleaning. The two systems work together to drive the value flow of legal data within the security boundary, and jointly support the compliance operation of the "Zhi Zhe Qian Lu" legal service large language model, building a privacy and security protection closed loop throughout the whole chain of data collection, processing and application for the multi-agent collaborative system.

[0129] The application builds a full-process intelligent service system in diversified legal scenes such as civil disputes, commercial disputes, criminal defense, administrative disputes, etc. Through the fusion of multi-agent collaborative architecture and multi-modal technology, the precision and scene adaptability of legal services are significantly improved. The system relies on a three-layer data governance architecture and a four-layer training system to realize the deep analysis of judicial big data and the dynamic injection of legal knowledge. Combined with RAG enhanced search and three-dimensional visualization technology, the accuracy of legal reasoning is improved by more than 25%, the efficiency of case matching is improved by 40%, and the system supports intelligent alignment of cross-border legal resources and digital twin deduction of the trial process. The modular design of the system breaks through the technical limitations of traditional single models and can be flexibly extended to legal subfields such as administrative compliance and international arbitration.

[0130] The application builds a legal large model system integrating judicial big data processing, legal question and answer, case search, document generation, foreign-related law, intelligent courtroom, case visualization, security compliance, and other functions, becoming a key to improving the intelligent level of legal services. The system should be able to collect, clean, and structure judicial big data, and build a knowledge base suitable for legal large model training and reasoning; based on the knowledge base, provide multi-functional intelligent question and answer services, including legal question and answer, case search, and document generation; support multi-lingual legal corpus understanding and cross-domain matching and translation prompts for legal provisions in different countries / regions; build a simulated courtroom function, simulate the trial process using a large model and generate complete dialogue records through preset courtroom scenarios and legal roles; realize case visualization, display case facts, legal relationships, procedural processes, and related cases in a visual chart on the user interface.

[0131] A "data-model-application" three-axis linkage legal intelligence ecosystem is built: through a two-way governance mechanism of multi-source heterogeneous data, the contradiction between legal professionalism and data complexity is solved; a composite optimization algorithm and a domain-specific model matrix are used to balance legal semantic understanding and cross-scene migration ability; based on a multi-agent dynamic collaboration framework, legal question and answer, case search, document generation, intelligent courtroom, foreign-related law, and case visualization are upgraded to full-process closed-loop services, and provide basic data for subsequent win rate prediction functions. Especially through private embedded models and a dual-track privacy protection system, under the premise of ensuring data security, the technical bottlenecks of legal multi-modal data processing and cross-border compliance application are overcome, providing a reliable intelligent infrastructure for judicial digitization, which can promote the improvement of legal service efficiency and user adoption rate.

[0132] On the other hand, as Figure 9 indicated, the application also provides a coiler control system, the system comprising:

[0133] The request receiving module 201 is configured to acquire user request information and identify category information and field information of the user request information.

[0134] The selection module 202 is configured to select a preset intelligent agent satisfying the category information and the field information from a preset intelligent agent library according to the category information and the field information, and take the preset intelligent agent satisfying the category information and the field information as a target intelligent agent, wherein the target intelligent agent comprises a preset large language model satisfying the field information and a preset function matrix satisfying the category information.

[0135] The execution module 203 is configured to input the user request information into the target intelligent agent to obtain an initial result.

[0136] The enhancement module 204 is configured to query enhancement information from a preset vector knowledge base based on the user request information by using a RAG retrieval technology.

[0137] The display module 205 is configured to generate a target result for the user request information based on the enhancement information and the initial result, and display the target result by using a visualization technology.

[0138] In a possible implementation, as shown in Figure 10 The terminal device 300 comprises a memory 310, a processor 320, and a first computer program 311 stored in the memory 310 and capable of running on the processor 320. When the processor 320 executes the first computer program 311, the following steps are implemented: acquiring user request information and identifying category information and field information of the user request information; selecting a preset intelligent agent satisfying the category information and the field information from a preset intelligent agent library according to the category information and the field information, and taking the preset intelligent agent satisfying the category information and the field information as a target intelligent agent, wherein the target intelligent agent comprises a preset large language model satisfying the field information and a preset function matrix satisfying the category information; inputting the user request information into the target intelligent agent to obtain an initial result; querying enhancement information from a preset vector knowledge base based on the user request information by using a RAG retrieval technology; and generating a target result for the user request information based on the enhancement information and the initial result, and displaying the target result by using a visualization technology.

[0139] In a possible implementation, as shown in Figure 11As shown, the embodiment of the present application provides a computer readable storage medium 400, which stores a second computer program 411, and the second computer program 411 is executed by a processor to realize: obtaining user request information, and identifying category information and field information of the user request information; selecting a preset agent meeting the category information and the field information from a preset agent library according to the category information and the field information, taking the preset agent meeting the category information and the field information as a target agent, the target agent including a preset large language model meeting the field information and a preset function matrix meeting the category information; inputting the user request information into the target agent to obtain an initial result; querying enhanced information in a preset vector knowledge base through RAG retrieval technology based on the user request information; generating a target result for the user request information based on the enhanced information and the initial result, and displaying the target result through a visualization technology.

[0140] It should be noted that the information interaction, execution process and the like between the above apparatuses / units are based on the same concept as the method embodiments of the present application, and the specific functions and the technical effects brought by the same can be referred to the method embodiments part, which will not be repeated here.

[0141] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of each functional unit and module are only for convenient distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0142] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the present application can implement all or part of the processes in the above-mentioned embodiment methods through a computer program to instruct relevant hardware to complete, and the computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium can at least include any entity or device capable of carrying the computer program code to the photographing device / terminal equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer readable medium can not be an electrical carrier signal and a telecommunication signal.

[0143] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0144] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0145] In the embodiments provided by the present application, it should be understood that the disclosed apparatus / network device and method can be implemented in other ways. For example, the above-described apparatus / network device embodiments are merely schematic, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed each other can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0146] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected to achieve the purpose of the embodiment scheme according to actual needs.

[0147] The above embodiments are only used to illustrate the technical solutions of the present application, not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

[0148] The above disclosure is only the preferred embodiment of the present application, of course, cannot limit the scope of the right of the present application, therefore, the equivalent changes made by the claims of the present application, still belong to the scope covered by the present application.

Claims

1. A method for legal services based on a large language model, characterized by, The method comprises the following steps: obtaining user request information through a dual-track privacy protection system, and identifying category information and field information of the user request information, the dual-track privacy protection system comprising an irreversible ciphertext interaction system based on multi-party secure computation and an LLM privacy protection system; based on a preset distributed data collection network, collecting original data in different legal fields through data collection technology, and labeling the original data by taking the source information of the original data as label information to obtain an initial data set, the data collection technology at least comprising text analysis technology, image feature extraction technology and video frame analysis technology; performing data analysis on each original data in the initial data set to obtain an analysis result, and performing screening on each original data in the initial data set according to the analysis result to obtain target data meeting preset requirements, the data analysis comprising comparative analysis, time series analysis, structure analysis, RFM analysis, grouping analysis and regression analysis; performing format conversion on the target data to obtain a JSON-LD format data set comprising three dimensions of fact elements, legal application and judgment results, wherein when the target data is unstructured data, the target data is converted into a JSON-LD format data set comprising three dimensions of fact elements, legal application and judgment results through entity recognition technology, relationship extraction technology and semantic disambiguation technology; based on the JSON-LD format data set, training an initial large language model for different legal fields according to a preset tuning algorithm to obtain a preset large language model for different legal fields, the preset tuning algorithm comprising at least one of LoRA algorithm, SFT algorithm, DPO algorithm, P-Tuning algorithm, RAG-FT algorithm, RLHF algorithm and KTO algorithm; constructing a preset agent library according to the preset large language model for different legal fields and a preset function matrix for different category information, the preset function matrix comprising a first matrix based on a semantic analysis engine of BERT-judicial version, a second matrix based on a similarity calculation model integrating space-time elements, a third matrix based on GPT-4 and a legal format verifier, a fourth matrix integrating adversarial debate and evidence chain analysis algorithm, and a fifth matrix based on multi-language legal alignment technology, the initial large language model comprising a Qwen-14B model for civil dispute field, a DeepSeek-legal model for commercial dispute field, a ChatGLM3-6B model for criminal defense field, and a LawyerLlama model for administrative dispute field; selecting a preset agent meeting the category information and the field information from the preset agent library according to the category information and the field information, and taking the preset agent meeting the category information and the field information as a target agent, the target agent comprising a preset large language model meeting the field information and a preset function matrix meeting the category information; inputting the user request information into the target agent to obtain an initial result; Integrating the original data into a statute library, a judge library, a lawyer library and a case library according to the sources of the original data; Correlating the data in the statute library, the judge library, the lawyer library and the case library according to a correlation index system based on legal elements to obtain multi-dimensional semantic correlation data with a correlation density greater than 85%; Using a privatization embedding model to construct a unified vector space mapping system of text, image and video, extracting feature vectors of the multi-dimensional semantic correlation data according to the unified vector space mapping system, and establishing a preset vector knowledge base based on the feature vectors; Based on the user request information, querying enhanced information in the preset vector knowledge base through RAG retrieval technology; Based on the enhanced information and the initial result, generating a target result for the user request information, and displaying the target result through visualization technology.

2. The legal service method based on a large language model according to claim 1, characterized in that, The step of identifying the category information and the field information of the user request information comprises: Identifying the user request information through intent recognition technology to obtain intent information of the user request information, and determining the category information and the field information of the user request information according to the intent information; Or, Obtaining the source information of the user request information, and determining the category information and the field information of the user request information based on the source information.

3. The legal service method based on a large language model according to claim 1, characterized in that, The step of training an initial large language model for different legal fields according to a preset tuning algorithm based on the JSON-LD format data set to obtain a preset large language model for different legal fields comprises: Optimizing the matching of data in the JSON-LD format data set through domain term injection to obtain a target JSON-LD format data set; Training an initial large language model for different legal fields according to a preset tuning algorithm based on the target JSON-LD format data set to obtain a secondary large language model for different legal fields; Calibrating the dimension space of the secondary large language model to obtain a preset large language model for different legal fields with a cosine similarity greater than 0.85 as the standard.

4. The legal service method based on a large language model according to claim 1, characterized in that, The step of querying enhanced information in the preset vector knowledge base through RAG retrieval technology based on the user request information comprises: Analyzing the user request information to obtain a query feature vector; Calculating the cosine similarity of the query feature vector with the feature vectors in the preset vector knowledge base, and identifying the feature vectors with a cosine similarity greater than or equal to a cosine similarity threshold value, sorting the identified feature vectors according to the cosine similarity to obtain a feature vector queue, and selecting the first ten feature vectors in the feature vector queue as target feature vectors; Selecting document data corresponding to the target feature vectors in the statute library, the judge library, the lawyer library and the case library; Performing legal logic verification on the document data to exclude invalid statute references to obtain target document data; Generating a retrieval report with judgment gist annotation based on the target document data, and taking the retrieval report with judgment gist annotation as enhanced information.

5. The legal service method based on a large language model according to claim 1, characterized in that, The method further comprises: When there are at least two preset agents in the preset agent library that meet the category information and the field information, a data transmission channel is established between the preset agents according to a preset rule, so that evidence chains between the preset agents are shared. When there are at least two preset agents in the preset agent library that meet the category information and the field information, and output results of the preset agents conflict, the output results are optimized through a federated learning algorithm to obtain a target result for the user request information. 6.A legal service system based on a large language model, characterized by, The system comprises: The request receiving module is configured to obtain user request information through a dual-track privacy protection system, and identify category information and field information of the user request information, wherein the dual-track privacy protection system includes an irreversible ciphertext interaction system based on multi-party secure computation and an LLM privacy protection system; based on a preset distributed data collection network, raw data in different legal fields is collected through a data collection technology, and source information of the raw data is taken as labeling information to label the raw data, so as to obtain an initial data set, wherein the data collection technology at least includes a text analysis technology, an image feature extraction technology and a video frame analysis technology; data analysis is performed on each raw data in the initial data set to obtain an analysis result, and each raw data in the initial data set is filtered according to the analysis result to obtain target data meeting a preset requirement, wherein the data analysis includes comparative analysis, time series analysis, structure analysis, RFM analysis, grouping analysis and regression analysis; the target data is format-converted to obtain a JSON-LD format data set including three dimensions of fact elements, legal application and judgment results, wherein when the target data is unstructured data, the target data is converted into the JSON-LD format data including the three dimensions of fact elements, legal application and judgment results through entity recognition technology, relationship extraction technology and semantic disambiguation technology; based on the JSON-LD format data set, an initial large language model for different legal fields is trained according to a preset tuning algorithm to obtain a preset large language model for different legal fields, wherein the preset tuning algorithm includes at least one of a LoRA algorithm, an SFT algorithm, a DPO algorithm, a P-Tuning algorithm, an RAG-FT algorithm, an RLHF algorithm and a KTO algorithm; a preset agent library is constructed according to the preset large language model for different legal fields and a preset function matrix for different category information, wherein the preset function matrix includes a first matrix of a semantic analysis engine based on BERT-justice version, a second matrix of a similarity calculation model based on fusion of space-time elements, a third matrix based on GPT-4 and a legal format verifier, a fourth matrix integrating adversarial argumentation and evidence chain analysis algorithm, and a fifth matrix based on multi-language legal alignment technology, and the initial large language model includes a Qwen-14B model for civil dispute field, a DeepSeek-legal model for commercial dispute field, a ChatGLM3-6B model for criminal defense field, and a LawyerLlama model for administrative dispute field; The selection module constructs a preset agent library according to a preset large language model for different legal fields and a preset function matrix for different category information, the preset function matrix including a first matrix based on a semantic analysis engine of BERT-judicial version, a second matrix based on a similarity calculation model integrating space-time elements, a third matrix based on GPT-4 and a legal format verifier, a fourth matrix integrating adversarial argumentation and evidence chain analysis algorithm, and a fifth matrix based on multilingual legal alignment technology, the initial large language model including a Qwen-14B model for civil dispute field, a DeepSeek-legal model for commercial dispute field, a ChatGLM3-6B model for criminal defense field, and a LawyerLlama model for administrative dispute field; the preset agent library is selected according to the category information and the field information to obtain a preset agent satisfying the category information and the field information, the preset agent satisfying the category information and the field information is taken as a target agent, and the target agent includes a preset large language model satisfying the field information and a preset function matrix satisfying the category information; The execution module is configured to input the user request information into the target agent to obtain an initial result. The enhancement module is configured to integrate the original data into a statute library, a judge library, a lawyer library, and a case library according to sources of the original data; associate data in the statute library, the judge library, the lawyer library, and the case library based on a correlation index system based on legal elements to obtain multi-dimensional semantic correlation data with an association density greater than 85%; construct a unified vector space mapping system of text, image, and video by using a privatization embedding model, extract feature vectors of the multi-dimensional semantic correlation data according to the unified vector space mapping system, and establish a preset vector knowledge base based on the feature vectors; query enhanced information in the preset vector knowledge base through RAG retrieval technology based on the user request information; and the display module is configured to generate a target result for the user request information based on the enhanced information and the initial result, and display the target result through a visualization technology.

7. An electronic device comprising a memory, a processor, and a first computer program stored in the memory and executable on the processor, characterized in that, The first computer program is executed by the processor to implement the large language model-based legal service method of any one of claims 1 to 5.

8. A computer-readable storage medium storing a second computer program, wherein the second computer program comprises instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7. The second computer program is executed by the processor to implement the large language model-based legal service method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Intellectual property and academic assistant system and implementation method and device thereof, electronic equipment and storage medium

    CN119623619A

  • Question and answer method and device based on large language model, electronic equipment and computer program product

    CN119886332A