Dynamic two-stage policy and regulation conflict detection method and system
Through the dynamic two-stage policy and regulation conflict detection method, combined with the NLI and LLM models, efficient and accurate conflict detection between enterprise regulations and legal provisions is achieved, and the problem of insufficient detection efficiency and accuracy in the existing technology is solved.
Patent Information
- Application Number
- CN202510302275.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to achieve efficient and accurate conflict detection by enterprise provisions relative to legal provisions, especially in resource-constrained environments.
The dynamic two-stage policy and regulation conflict detection method is adopted to detect semantic contradictions through the NLI model, and uncertainty estimation is used to decide whether to hand over the detection results to the LLM model for in-depth analysis to achieve more accurate compliance detection.
It achieves the improvement of system response efficiency while maintaining high accuracy, reduces the demand for computing resources, and is suitable for real-time detection applications, which can help enterprises timely discover and reduce the risk of policy and regulation conflicts.
Smart Images

Figure CN120218080A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and in particular, to a dynamic two-stage policy and regulation conflict detection method and system. Background Art
[0002] The Natural Language Inference (NLI) model is a model for identifying semantic relationships between two texts and is commonly used in text reasoning tasks. As shown in Table 1, the NLI model can efficiently detect the contradictory relationship between text pairs and is an effective tool for handling semantic conflicts in enterprise compliance detection. The application of the NLI model in compliance detection can help enterprises quickly identify potential regulatory conflicts, thereby reducing legal risks. As an example of the NLI task in Table 1, the model inputs a premise sentence and a hypothesis sentence and determines the entailment relationship between these two sentences. The application of the NLI model is mainly limited to classification tasks, suffering from the problem of insufficient interpretability of detection results and also showing limitations when dealing with complex legal logical relationships.
[0003] Table 1
[0004]
[0005] Large language models (LLMs) such as GPT have powerful natural language processing capabilities and can generate fluent and contextually consistent text. Through the training of a large number of legal texts, these models can identify the relationships between complex legal logics and clauses, thus providing detailed and natural explanations in compliance detection. The application of the LLM model enables compliance detection not only to identify potential compliance risks but also to generate detailed explanations to assist in decision-making. However, with the continuous expansion of the model scale, the demand for computing resources grows exponentially, which limits the practical application of the LLM. At the same time, small models (NLI models) for specific tasks maintain high performance while significantly reducing computational complexity, making them more adaptable and efficient in resource-constrained environments. How to establish an efficient dynamic cooperation mechanism between these two types of models has become a problem to be solved.
[0006] Uncertainty Estimation (UE) can dynamically adjust the task allocation between the LLM and NLI by evaluating the uncertainty of the model's prediction results, thereby optimizing resource utilization and enhancing system performance. Specifically, when the uncertainty estimation indicates that the task is simple and the small model has sufficient confidence, the small model will be preferentially used to save computing resources; while in cases of high uncertainty, the large model will be invoked to improve the accuracy of the results. This collaborative interaction framework can not only reduce the overall computational burden but also improve the system's response efficiency while maintaining high accuracy, thus enabling the complementary advantages of large and small models. Uncertainty estimation is used to dynamically evaluate the difficulty of the task and the confidence of the model, and thus intelligently allocate tasks between large and small models.
[0007] After retrieval, Chinese Patent Application Publication No. CN108595415A discloses a legal differentiation determination method, device, computer device, and storage medium. The method includes: obtaining the legal and regulatory text information to be determined; performing word segmentation processing on the legal and regulatory text information to be determined; performing matching calculations on the word-segmented legal and regulatory text information with a legal and regulatory clause corpus to obtain the same part of the vocabulary and the different part of the vocabulary respectively; performing similarity combination calculations on the same part of the vocabulary and the different part of the vocabulary to obtain the differentiation result between the legal and regulatory text information to be determined and the existing laws and regulations. This existing patent application has the problem that after finding the relevant legal articles, it does not perform compliance detection on the legal and regulatory text to be determined.
[0008] How to achieve efficient and accurate conflict detection of enterprise regulations relative to legal clauses has become a technical problem to be solved. Summary of the Invention
[0009] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a dynamic two-stage policy and regulation conflict detection method and system.
[0010] The purpose of the present invention can be achieved by the following technical solutions:
[0011] According to one aspect of the present invention, a dynamic two-stage policy and regulation conflict detection method is provided. The method includes the following steps:
[0012] Step 1, extract relevant legal text pairs from a legal database and enterprise regulations, and perform vectorization processing;
[0013] Step 2, retrieve the K legal clauses most relevant to the enterprise regulations, where K≥1;
[0014] Step 3, perform semantic contradiction detection in the first stage: input the K legal clauses obtained in Step 2 into an NLI model for semantic contradiction detection, and screen out the text pairs with semantic contradictions;
[0015] Step 4, based on the uncertainty U output by the uncertainty estimation NLI model, if a semantic contradiction detection result is considered inaccurate, then execute Step 5; otherwise, consider the semantic contradiction detection result accurate and end its detection process;
[0016] Step 5, perform compliance detection and analysis in the second stage: Input the text pairs with semantic contradictions into the LLM model for in-depth analysis to obtain potential compliance issues of the enterprise and provide detailed explanations.
[0017] Preferably, compare the uncertainty U with a preset uncertainty threshold T. If U > T, then consider a text pair with semantic contradictions screened by the NLI model as an uncertain problem, which will be processed in the second stage by the LLM model; otherwise, consider a text pair with semantic contradictions screened by the NLI model as an uncertain problem and do not allocate it to the LLM model for in-depth analysis.
[0018] Preferably, in Step 3, the process of semantic contradiction detection is specifically as follows: Match the input enterprise regulations with the retrieved legal regulations in multiple Step 2 and then input them into the NLI model for semantic analysis to output classification results, including semantic contradiction, semantic neutrality, and semantic entailment.
[0019] Preferably, in Step 2, the process of retrieving the legal regulations most relevant to the enterprise regulations includes:
[0020] Use the enterprise regulations as the input query, which is converted into a dense vector form after being processed by the query vectorization module;
[0021] Calculate the similarity between the dense vector converted from the enterprise regulations and the vector of the legal documents that have been quantified and indexed, and use the retrieval model based on dense vectors to retrieve the legal regulations most relevant to the enterprise regulations.
[0022] More preferably, the retrieval model based on dense vectors is the M3e-base model.
[0023] Preferably, in Step 4, the process of compliance detection and analysis includes:
[0024] Obtain the corresponding text pairs with semantic contradictions by querying with the input enterprise regulation clauses;
[0025] Input the text pairs with semantic contradictions into the LLM model for compliance detection and analysis, and according to the expert prompt professional interpretation results, divide the output results into two categories: enterprise compliance analysis and enterprise non-compliance analysis, and the enterprise non-compliance analysis includes explanations of non-compliance.
[0026] Preferably, the NLI model is the Erlangshen-MegatronBert-1.3B model; the LLM model uses the ChatGLM2-6B model.
[0027] Preferably, the method further includes step 6: manually further evaluate the accuracy and interpretability of the analysis results of the potential compliance issues of the enterprise in step 5.
[0028] According to another aspect of the present invention, there is provided a two-stage policy and regulation conflict detection system, including:
[0029] Legal database: a legal database obtained by integrating legal data sets and government public data, through data cleaning and processing, and classified according to legal types and regions;
[0030] Legal text pair extraction module: using a retrieval model to extract relevant legal text pairs from the legal database and enterprise regulations, and vectorize these texts;
[0031] Legal clause retrieval module: converting the input enterprise regulations into a dense vector form, and calculating the similarity with the legal document vectors that have been quantified and indexed to obtain the most relevant legal clauses;
[0032] Semantic contradiction retrieval module: This module is used in the first stage of policy and regulation conflict detection, using the NLI model to identify semantic contradictions between the legal clause text pairs most relevant to the enterprise regulations, so as to screen out the clauses that may have conflicts;
[0033] Compliance detection and analysis module: This module is used in the second stage of policy and regulation conflict detection. The LLM model conducts a more in-depth analysis of the text pairs detected and screened for semantic contradictions to finally determine whether the enterprise regulations comply with legal requirements and detect potential compliance issues.
[0034] Preferably, the system further includes an uncertainty estimation module, which estimates the uncertainty of the output result of the semantic contradiction retrieval module, and judges whether it is necessary for the compliance detection and analysis module to continue in-depth analysis according to the uncertainty estimation result.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1) The present invention connects the NLI and LLM models through uncertainty estimation. In the first stage, after using the NLI model to screen out text pairs with semantic contradictions, uncertainty estimation is performed on the detection results. If a certain detection result is considered a deterministic problem, it is not assigned to the LLM model for in-depth analysis; otherwise, this detection result is assigned to the LLM model in the next second stage for compliance detection and analysis, realizing dynamic two-stage detection, taking into account both the efficiency and accuracy of compliance detection, so as to help enterprises promptly discover and reduce the risk of potential policy and regulatory conflicts.
[0037] 2) The NLI model of the present invention uses the Erlangshen-MegatronBert-1.3B model, which is relatively small in scale, fast in inference speed, and its efficient context capture ability and enhanced prediction performance enable it to maintain high accuracy when processing complex legal texts, helping to improve detection efficiency and being suitable for real-time detection applications.
[0038] 3) The LLM model of the present invention is good at processing implicit semantics and complex contexts, and is used for compliance detection and analysis of text pairs with semantic contradictions. It interprets the results professionally according to expert prompts and has strong interpretability, thus improving the accuracy of detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of two-stage policy and regulatory conflict detection in the present invention;
[0040] Figure 2 It is a schematic diagram of the structure of the retrieval model in the present invention;
[0041] Figure 3 It is a schematic diagram of the process of semantic contradiction detection in the first stage of the present invention;
[0042] Figure 4 It is a schematic diagram of the process of compliance detection and analysis in the second stage of the present invention;
[0043] Figure 5 It is a schematic diagram of the distribution of legal types in the legal database of the present invention;
[0044] Figure 5 (a) is Figure 5 a schematic diagram of the detailed distribution of legal types in other parts of
[0045] Figure 6 It is a schematic diagram of the relationship between the retrieval Top_K value, average similarity, and efficiency improvement ratio in the present invention;
[0046] Figure 7 It is a schematic diagram of the comparison of the accuracy rate and the percentage of risk reduction of multiple data sets in the present invention;
[0047] Figure 8 This is a schematic diagram of the conflict detection system in the present invention. Specific implementation manners
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] Embodiment 1
[0050] This embodiment relates to a two-stage policy and regulation conflict detection method. This method is based on natural language inference (NLI) and large language model (LLM), integrating the advantages of natural language inference and large language model. By introducing the large language model (LLM) in the second stage, the interpretability of the detection results is enhanced, and the analysis ability of complex semantic relationships is improved, so as to be able to efficiently and accurately detect potential conflicts between policy and regulation texts and help enterprises reduce legal compliance risks.
[0051] The two-stage policy and regulation compliance detection method of this method is as Figure 1 , including the following four steps:
[0052] Step 1, construction of a legal database;
[0053] Step 2, legal provision retrieval;
[0054] Step 3, perform semantic contradiction detection in the first stage;
[0055] Step 4, perform compliance detection and analysis in the second stage to obtain suspected non-compliant clauses;
[0056] Step 5, the suspected non-compliant clauses screened by the model will be further evaluated manually to ensure the accuracy and interpretability of the analysis results.
[0057] In Step 1, first, a legal database is constructed in a semi-automated manner. While screening data through the model, manual inspection and feedback on data quality are used to produce the legal database. Then, relevant legal text pairs are extracted from the legal database and enterprise regulations by using a retrieval model, and these texts are vectorized, and relevant legal text pairs are retrieved through the retrieval model.
[0058] By integrating legal datasets from multiple open-source platforms and manually obtaining government public data, mainly the reformatted Chinese legal dataset, a database containing 22,552 legal documents was established. These data were strictly screened, only retaining legal documents with the status of "valid", and classified according to legal types and regions. The distribution of legal document types is as Figure 5 shown, Figure 5 The detailed distribution of legal documents under the "Other" category in Figure 5 (a) is shown.
[0059] To ensure the structuring and cleanliness of the data, we used Python scripts to uniformly convert legal documents in various formats into JSON format, and performed sentence splitting and filtering of invalid text. At the same time, the data cleaning rules were continuously optimized during the project process to ensure data consistency and accuracy.
[0060] The process of legal article retrieval in Step 2 is as Figure 2 shown, including the following steps:
[0061] First, the enterprise regulations are used as input queries. After being processed by the query vectorization module, they are converted into dense vector form.
[0062] Subsequently, the dense vectors converted from the enterprise regulations are used to calculate the similarity with the vectors of legal documents that have been quantified and indexed to determine the most relevant legal articles. The retrieval results of the K (K≥1) most relevant legal articles retrieved will be further sent to the semantic contradiction detection module to identify possible semantic conflicts.
[0063] Traditional word-frequency-based retrieval models such as TF-IDF and BM25 models are suitable for keyword matching and text relevance retrieval, but they perform poorly in dealing with complex semantic relationships. In contrast, dense vector-based retrieval models such as Word2Vec, BERT, and S-BERT can better capture semantic relationships in text, but have high computational resource requirements. The M3e-base model is further optimized on this basis and can support accurate retrieval of Chinese, homogeneous, and heterogeneous texts. Therefore, the M3e-base model is used in this embodiment for legal article retrieval. Through this retrieval process, we can efficiently locate legal articles related to enterprise regulations in a large-scale legal database, providing a basis for subsequent compliance analysis.
[0064] Step 3, that is, semantic contradiction detection is carried out in the first stage. The goal of this stage is to identify semantic conflicts between enterprise regulations and legal texts through a natural language inference (NLI) model. This process is as Figure 3As shown, first, the input query is matched with multiple retrieved results, and then the NLI model is used to perform semantic analysis on these text pairs, classifying them as semantically contradictory, semantically neutral, or semantically entailed.
[0065] The NLI model is used to identify semantic contradictions between text pairs, thereby screening out potentially conflicting clauses. Then, an uncertainty estimation method is used to dynamically allocate the texts that require in-depth analysis. Specifically, the following three uncertainty estimation methods are used: confidence-based, distribution-based, and sampling-based methods. The confidence-based method is used by default.
[0066] Confidence-based method:
[0067] Least Confidence (LC): For a C-classification problem, given the input problem s and the prediction y ∈ {1,..., C}, the uncertainty is described as:
[0068] u LC (s) = 1 - max c∈C p(y = c|s) (1)
[0069] where p(y = c|s) is the Softmax score for class c.
[0070] Distribution-based method:
[0071] Prediction Entropy (PE): Prediction Entropy (PE) is a direct and effective method for estimating uncertainty. The maximum entropy occurs when all results have the same probability. The uncertainty u PE (s) of the input s can be quantified as its prediction entropy:
[0072]
[0073] Sampling-based method:
[0074] Monte Carlo Dropout (MCD): Monte Carlo Dropout (MCD) is a typical sampling method. Specifically, the MCD method performs M statistical random forward passes with dropout activation:
[0075]
[0076] For classification problems, the entropy value is then obtained through Monte Carlo integration to get MCD. For regression problems, the mean of the predicted variances can be used as the uncertainty. All three methods use the cross-entropy loss function to fine-tune the model.
[0077] Through uncertainty estimation, we can obtain the uncertainty (U) of the NLI model, which will be compared with the uncertainty threshold (T) preset by human experts. If U>T is obtained through uncertainty estimation, we consider this question to be an uncertain question for the NLI model, so it will be processed in the second stage by the large language model (LLM); otherwise, we consider this question to be a definite question for the NLI model, and the detection ends, and the second-stage compliance detection and detailed analysis will no longer be carried out.
[0078] For the detected text pairs with semantic contradictions, the corresponding uncertainties are obtained using the three uncertainty estimation methods mentioned above. If U>T, it will enter the second stage for compliance detection and detailed analysis. In this paper, the Erlangshen-MegatronBert-1.3B model, which can balance accuracy and efficiency, is selected as the core tool for semantic contradiction detection. Its efficient context capture ability and enhanced prediction performance enable it to maintain high accuracy when processing complex legal texts while meeting the project's requirements for speed and computing resources.
[0079] Through uncertainty estimation, intelligence realizes task allocation between the large language model (LLM) and the small model (i.e., the natural language inference (NLI) model), which enables the entire system to improve the response efficiency while maintaining high accuracy.
[0080] After semantic contradiction detection, for the detected text pairs with semantic contradictions, if they are determined to be uncertain questions after uncertainty evaluation, they will enter the second stage for compliance detection and detailed analysis.
[0081] Step 4, that is, the compliance detection and detailed analysis in the second stage, the large language model (LLM) conducts a more in-depth analysis of the text detected and screened for semantic contradictions to finally determine whether the enterprise regulations comply with legal requirements and detect potential compliance issues. The process of the second stage is as Figure 4 shown. First, the text pairs that have passed semantic contradiction detection are obtained through the input query, and then the large language model (LLM) is used to conduct compliance detection and analysis on these texts. In this embodiment, the ChatGLM2-6B model is used as the main LLM tool.
[0082] During the compliance detection process, first, legal experts provide prompts to ensure that the model can accurately understand and handle the complexity of legal texts. Subsequently, the large language model (LLM) analyzes the input text, which is divided into two categories: enterprise compliance analysis and enterprise non-compliance analysis. Through this process, the model can effectively identify potential compliance issues and provide detailed explanations to help enterprises better understand and handle legal risks.
[0083] Example 2
[0084] This embodiment also involves the evaluation of a two-stage policy and regulation conflict detection method for evaluating the efficiency improvement of the method of the present application in enterprise compliance detection.
[0085] In the experiment, we used a part of enterprise documents as input and conducted experimental analysis for different Top_K values. By setting different Top_K values, we evaluated the efficiency performance of the model in semantic contradiction detection and suspected non-compliance analysis and quantified it as the percentage index of efficiency improvement as follows:
[0086]
[0087] During the retrieval process, we adopted the average retrieval similarity as the main evaluation index to evaluate the performance of the model under different retrieval conditions. The average retrieval similarity measures the accuracy of the model by calculating the similarity between the input query text and the retrieved text. Specifically, given a query text Query text and a retrieved text Retrievaltext, the similarity between the two is expressed as:
[0088]
[0089] where q represents the vector representation of the input query text and d represents the vector representation of the retrieved text.
[0090]
[0091] where K represents the top Top_K texts retrieved and N represents the number of input queries in a document. This index measures the average similarity between each query and its corresponding top K retrieval results to evaluate the performance of the model at different retrieval depths.
[0092] By analyzing the retrieval results under different Top_K values, we quantified the improvements of the model in reducing the workload of manual review and accelerating the detection process. Table 2 shows the efficiency improvement ratios of the two-stage method under different K values.
[0093] Table 2
[0094]
[0095] Such as Figure 6, as the Top_K value increases, the number of retrieved text pairs and the number of detected semantic contradictions increase significantly, while the average similarity decreases slightly. Specifically, the efficiency improvement ratio 1 decreases slightly in the preliminary screening stage, from 94.58% to 93.63%, indicating that as the data volume increases, the screening efficiency of the model decreases slightly but remains high. In the efficiency improvement ratio 2, the efficiency of the model in the further screening and compliance analysis stages is stable, remaining between 94.56% and 96.24%, showing high accuracy and stability.
[0096] We used three large natural language inference datasets (CMNLI, OCNLI, SNLI) to analyze the performance of the NLI model in reducing enterprise compliance risks. By evaluating the accuracy of the model on these datasets, we calculated the percentage of risk reduction by the model in compliance detection, and the calculation method is as follows:
[0097]
[0098] The calculation results are shown in Table 3 and Figure 7 as follows. In terms of detection efficiency, the two-stage framework performs excellently. The NLI model can quickly screen out semantically contradictory texts during the preliminary screening. Using the uncertainty estimation method can improve the reliability of the model while ensuring the detection accuracy of the model. Under different Top_K values in the experiment, the efficiency improvement in the preliminary screening exceeds 93%, and it also stabilizes above 94% in the subsequent analysis stage, significantly shortening the detection cycle. In terms of accuracy, the combination of NLI and the LLM model is powerful, and the accuracy exceeds 84% on multiple datasets, and it can reliably judge compliance or non-compliance. In terms of interpretability, it has strong interpretability. The LLM model interprets the results professionally according to the expert prompt, helping enterprises clarify compliance issues and countermeasures. And the framework is flexible and extensible, and each module can be adjusted and optimized as needed to adapt to different industries, and can be customized according to the actual situation of the enterprise to ensure the detection timeliness and accuracy, and effectively assist enterprise compliance management. Therefore, it can be seen that the present invention application has significant advantages in the field of policy and regulation compliance detection.
[0099] Table 3
[0100]
[0101]
[0102] To adapt to the particularity of the legal dataset, we introduced an adjustment coefficient α = 0.85 to correct the risk reduction percentage of the model. After adjustment, the average risk reduction percentage decreased from 92.25% to 78.41%. This result indicates that despite the differences between datasets, the application of the NLI model in the legal field can still significantly reduce the compliance risks of enterprises, providing a reliable compliance management tool for enterprises.
[0103] In the current complex and ever-changing regulatory environment, traditional manual analysis methods can no longer meet the enterprise's demand for efficient and accurate compliance detection. In contrast, an automated analysis framework based on large language model technology significantly improves the processing efficiency of regulatory assessment and policy consistency. By combining a natural language inference model with a large language model, the two-stage detection framework proposed in this study effectively solves the problems of semantic contradiction and legal conflict identification, greatly enhancing the speed and reliability of enterprise compliance detection.
[0104] Example 3
[0105] It also relates to a two-stage policy and regulation conflict detection system, as Figure 8 shown, the system includes:
[0106] Legal database: Integrate legal data sets from multiple open-source platforms and manually obtained government public data. After strict screening, only legal documents with the status of "valid" are retained. The data is structured and cleaned, and classified according to legal types and regions.
[0107] Legal text pair extraction module: Use a retrieval model to extract relevant legal text pairs from the legal database and enterprise regulations, and vectorize these texts.
[0108] Legal clause retrieval module: Take enterprise regulations as input queries. After being processed by the query vectorization module, they are converted into dense vector form. The dense vectors converted from enterprise regulations are calculated for similarity with the vectorized and indexed legal document vectors to determine the most relevant legal clauses. The retrieval results of the K most relevant legal clauses retrieved will be further sent to the semantic contradiction detection module to identify possible semantic conflicts.
[0109] Semantic contradiction retrieval module: This module is used in the first stage of policy and regulation conflict detection. It uses the NLI model to identify semantic contradictions between the legal clause text pairs most relevant to enterprise regulations, thereby screening out potentially conflicting clauses, and obtaining the corresponding uncertainty based on the uncertainty estimation method. If U>T, it will enter the second stage for compliance detection and detailed analysis; otherwise, it will not perform the second stage of compliance detection and detailed analysis.
[0110] Uncertainty estimation module, which performs uncertainty estimation on the output results of the semantic contradiction retrieval module, and judges whether the compliance detection and analysis module needs to continue in-depth analysis according to the uncertainty estimation results.
[0111] Compliance Detection and Analysis Module: This module is used in the second stage of policy and regulation conflict detection. The LLM model conducts a more in-depth analysis of the text after semantic contradiction detection and screening to ultimately determine whether the enterprise regulations comply with legal requirements and detect potential compliance issues.
[0112] Embodiment 4
[0113] The electronic device of the present invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0114] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0115] The processing unit executes the various methods and processes described above. For example, in some embodiments, the method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute the method by any other suitable means (e.g., by means of firmware).
[0116] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.
[0117] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0118] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0119] As described above, only the specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of various equivalent modifications or substitutions, and all such modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A dynamic two-stage policy and regulation conflict detection method, characterized in that: The method comprises the following steps: Step 1: extract relevant legal text pairs from the legal database and corporate regulations and perform vectorization processing; Step 2: retrieve the K legal clauses most relevant to the enterprise regulations, where K ≥ 1; Step 3: The first stage is to perform semantic contradiction detection: the K legal clauses obtained in step 2 are input into the NLI model for semantic contradiction detection, and semantically contradictory text pairs are screened out; Step 4: estimate the uncertainty U output by the NLI model through uncertainty. If a semantic contradiction detection result is considered inaccurate, execute step 5; otherwise, the semantic contradiction detection result is considered accurate and the detection process ends. Step 5, the second phase of compliance detection and analysis: The semantically contradictory text is input into the LLM model for analysis to obtain the company's potential compliance issues and provide detailed explanations.
2. A dynamic two-stage policy and regulation conflict detection method according to claim 1, characterized in that: Compare the uncertainty U with the preset uncertainty threshold T. If U>T, the semantically contradictory text pair selected by the NLI model is considered to be an uncertain issue and will be processed by the LLM model in the second stage. Otherwise, a semantically contradictory text pair screened out by the NLI model is considered to be an uncertain issue and is not assigned to the LLM model for in-depth analysis.
3. A dynamic two-stage policy and regulation conflict detection method according to claim 1, characterized in that: In step 3, the process of semantic contradiction detection is specifically as follows: matching the input enterprise regulations with the multiple legal terms retrieved in step 2, and then inputting them into the NLI model for semantic analysis, and outputting classification results, including semantic contradiction, semantic neutrality and semantic implication.
4. A dynamic two-stage policy and regulation conflict detection method according to claim 1, characterized in that: In step 2, the process of retrieving the legal provisions most relevant to the enterprise's regulations includes: The enterprise regulations are used as input queries, which are processed by the query vectorization module and converted into dense vector form; The dense vector converted from the enterprise regulations is used to calculate the similarity with the quantified and indexed legal document vector, and the legal clauses most relevant to the enterprise regulations are retrieved using a dense vector-based retrieval model.
5. A dynamic two-stage policy and regulation conflict detection method according to claim 4, characterized in that: The dense vector-based retrieval model is the M3e-base model.
6. A dynamic two-stage policy and regulation conflict detection method according to claim 1, characterized in that: In step 4, the compliance detection and analysis process includes: Obtain corresponding semantically contradictory text pairs through querying the input enterprise regulations; The semantically contradictory text is input into the LLM model for compliance detection and analysis, and the output results are divided into two categories based on the expert prompt professional interpretation results: enterprise compliance analysis and enterprise non-compliance analysis. Enterprise non-compliance analysis includes unqualified explanations.
7. A dynamic two-stage policy and regulation conflict detection method according to claim 1, characterized in that: The NLI model is the Erlangshen-MegatronBert-1.3B model; the LLM model adopts the ChatGLM2-6B model.
8. A dynamic two-stage policy and regulation conflict detection method according to claim 1, characterized in that: The method further includes step 6: manually further evaluating the accuracy and interpretability of the analysis results of the potential compliance issues of the enterprise in step 5.
9. A two-stage policy and regulation conflict detection system, used to execute the dynamic two-stage policy and regulation conflict detection method according to any one of claims 1 to 8, characterized in that: The system comprises: Legal database: A legal database obtained by integrating legal data sets and government public data, and after data cleaning and classification according to legal type and region; Legal text pair extraction module: uses the retrieval model to extract relevant legal text pairs from the legal database and corporate regulations, and vectorizes these texts; Legal terms retrieval module: converts the input corporate regulations into dense vector form, and calculates the similarity with the quantified and indexed legal document vectors to obtain the most relevant legal terms; Semantic contradiction retrieval module: This module is used in the first stage of policy and regulation conflict detection. It uses the NLI model to identify the semantic contradictions between the legal clause text pairs that are most relevant to the enterprise regulations, thereby screening out the clauses that may be in conflict; Compliance detection and analysis module: This module is used in the second stage of policy and regulatory conflict detection. The LLM model conducts a deeper analysis of the text pairs after semantic contradiction detection and screening to ultimately determine whether the enterprise regulations are consistent with legal requirements and detect potential compliance issues.
10. The system according to claim 9, characterized in that The system also includes an uncertainty estimation module, which performs uncertainty estimation on the output results of the semantic contradiction retrieval module and determines whether the compliance detection and analysis module needs to continue to perform in-depth analysis based on the uncertainty estimation results.
Citation Information
Patent Citations
Law differentiation determining method and apparatus, computer device, and storage medium
CN108595415A
Cited By
Conflict-aware legal case judgment prediction method and system
CN121235859A
Text similarity calculation method and related product
CN121278409A
Content comparison and analysis method and system based on convergence policy file and large model
CN122838612A