Real-time compliance conflict detection method and system for large language models of legal documents

CN122819213APending Publication Date: 2026-09-25SHENZHEN MINGXIN DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611257821.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-19
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]基于此,有必要针对现有技术中小语种法律术语多义性导致语义提取错误,以及无法区分不同区域宗教规则适用差异的技术问题,提供一种法律文件的大语言模型实时合规冲突检测方法及系统

Benefits of technology

[0009]本申请实施方式提供的法律文件的大语言模型实时合规冲突检测方法,通过针对特定语种法律术语文化特性设置的目标语义解析模型对待检测法律文件进行解析,使同一术语在世俗法律语境与宗教义务语境下被赋予不同语义,所提取的法律要素和法律语义信息能够反映文本的真实法律含义,避免了后续合规评估建立在错误语义基础之上。通过根据法律要素中的地域信息和法律行为类型确定区域性规则库,并以此调节各维度的权重系数,使合规冲突量化评分过程能够区分不同区域在宗教规则适用上的差异,合规冲突量化评分与法律文本在该区域实际面临的合规风险之间的偏离缩小,提高了对涉及区域宗教规则的法律文本进行冲突风险程度评估的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819213A_ABST
    Figure CN122819213A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of large language model, and provides a large language model real-time compliance conflict detection method and system for legal documents, which comprises the following steps: analyzing a to-be-detected legal document based on a target semantic analysis model, and extracting legal elements and legal semantic information; evaluating the legal elements and the legal semantic information based on a scoring framework in multiple dimensions of a specific language, so as to determine initial scores of the dimensions; determining a regional rule base of regional religious rules and commercial custom rules corresponding to the legal text according to regional information and legal behavior types in the legal elements; and determining a dynamic weight distribution result according to the regional rule base of the regional religious rules and the commercial custom rules corresponding to the legal text. The application ensures that the extracted legal elements and legal semantic information can reflect the true legal meaning of the text, and improves the accuracy of the conflict risk degree evaluation of the legal text related to regional religious rules.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large language model technology, and in particular to a method and system for real-time compliance conflict detection of legal documents using large language models. Background Technology

[0002] In related technologies, compliance checks on legal documents in less commonly spoken languages ​​typically involve first translating the text into a common language, and then performing compliance verification through keyword matching or a fixed rule base. Legal terminology in less commonly spoken languages ​​often carries specific religious and cultural connotations; the same word can express secular legal concepts as well as religious obligations in different contexts. General translation models cannot distinguish the precise legal semantics of such terms based on cultural context. Existing methods assess the risk of legal conflict using uniform evaluation criteria, treating all legal texts in the same way.

[0003] When ambiguous terms in legal texts are not semantically disambiguated according to cultural context, the extracted legal elements and semantic information fail to reflect the true legal meaning of the text, leading to subsequent compliance assessments based on flawed semantic foundations. When legal texts involve religious rules of a specific region, uniform assessment standards cannot distinguish the differences in the application of religious rules across different regions. This results in a deviation between the quantitative score of compliance conflicts and the actual compliance risks faced by the legal text in that region, leading to inaccurate assessments of the conflict risk level of legal texts involving regional religious rules. Summary of the Invention

[0004] Therefore, it is necessary to provide a real-time compliance conflict detection method and system for legal documents based on a large language model, addressing the technical problems of semantic extraction errors caused by the polysemy of legal terms in minority languages ​​and the inability to distinguish the differences in the application of religious rules in different regions.

[0005] Firstly, a real-time compliance conflict detection method for legal documents using a large language model is provided, comprising: parsing the legal document to be detected based on a target semantic parsing model to extract legal elements and legal semantic information, wherein the target semantic parsing model is set for the cultural characteristics of legal terminology in a specific language within the legal document; evaluating the legal elements and legal semantic information based on a multi-dimensional scoring framework for the specific language to determine initial scores for each dimension; determining a regional rule base of regional religious rules and commercial practice rules corresponding to the legal text based on the regional information and legal behavior type in the legal elements; adjusting the weight coefficients of each dimension based on the regional rule base of regional religious rules and commercial practice rules corresponding to the legal text to determine a dynamic weight allocation result; and weighting the weight coefficients corresponding to each dimension in the dynamic weight allocation result with the initial score of each dimension to determine a comprehensive conflict quantification score, wherein the comprehensive conflict quantification score is used to assess the degree of conflict risk of the legal text.

[0006] Secondly, a real-time compliance conflict detection system for a large language model of legal documents is provided. The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the aforementioned real-time compliance conflict detection method for a large language model of legal documents.

[0007] Thirdly, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-mentioned method for real-time compliance conflict detection of large language models of legal documents.

[0008] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for real-time compliance conflict detection of large language models of legal documents.

[0009] The real-time compliance conflict detection method for legal documents provided in this application uses a target semantic parsing model designed for the cultural characteristics of legal terminology in specific languages ​​to analyze the legal documents under test. This allows the same term to be assigned different meanings in secular legal contexts and religious obligation contexts. The extracted legal elements and legal semantic information can reflect the true legal meaning of the text, avoiding subsequent compliance assessments based on erroneous semantics. By determining a regional rule base based on the geographical information and legal behavior type in the legal elements, and adjusting the weight coefficients of each dimension accordingly, the compliance conflict quantification scoring process can distinguish the differences in the application of religious rules in different regions. The deviation between the compliance conflict quantification score and the actual compliance risks faced by the legal text in that region is reduced, improving the accuracy of assessing the degree of conflict risk for legal texts involving regional religious rules. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] in: Figure 1 This is a flowchart of the real-time compliance conflict detection method for a large language model of legal documents in an embodiment of the present invention; Figure 2 This is a schematic diagram of the modules of the real-time compliance conflict detection system for a large language model of legal documents in an embodiment of the present invention; Figure 3 This is a structural block diagram of the computer device server in an embodiment of the present invention; Figure 4 This is a structural block diagram of the computer device client in an embodiment of the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] Please see Figure 1 As shown, Figure 1 A flowchart illustrating the real-time compliance conflict detection method for a large language model of legal documents provided in this embodiment of the invention includes the following steps: Step 01: Based on the target semantic parsing model, the legal document to be detected is parsed to extract legal elements and legal semantic information. The target semantic parsing model is set according to the cultural characteristics of legal terminology in specific languages ​​in the legal document; Step 02: Based on a multi-dimensional scoring framework for a specific language, evaluate the legal elements and legal semantic information to determine the initial scores for each dimension; Step 03: Based on the geographical information and type of legal act in the legal elements, determine the regional rule base of regional religious rules and commercial customary rules corresponding to the legal text; Step 04: Based on the regional rule base of regional religious rules and commercial customary rules corresponding to the legal text, adjust the weight coefficients of each dimension to determine the dynamic weight allocation result; Step 05: The weight coefficients corresponding to each dimension in the dynamic weight allocation results are weighted and calculated with the initial score of each dimension to determine the comprehensive conflict quantification score. The comprehensive conflict quantification score is used to assess the degree of conflict risk of legal texts.

[0014] Specifically, in step 01, the target semantic parsing model is a multilingual pre-trained model fine-tuned from legal corpora in a specific language. The multilingual pre-trained model employs a Cross-lingual Language Model - Robustly Optimized Bidirectional Encoder Representations from Transformers approach (XLM-RoBERTa). XLM-RoBERTa is pre-trained on corpora covering multiple languages. During fine-tuning, a total of X tens of thousands of legal texts in the relevant language are collected. Legal experts annotate polysemous terms influenced by religious and cultural backgrounds, with annotation categories including secular legal context and religious obligation context. The annotated corpora are divided into training, validation, and test sets in an 8:1:1 ratio. The fine-tuning process uses a cross-entropy loss function, a learning rate of 2e-5, three training epochs, and a batch size of 16. The fine-tuned model achieves a semantic disambiguation accuracy of over 92% on the validation set, demonstrating cross-lingual semantic encoding capabilities.

[0015] XLM-RoBERTa is pre-trained on corpora containing multiple languages, possessing cross-linguistic semantic encoding capabilities. The specific method for setting up legal terminology cultural characteristics for specific languages ​​involves collecting legal text corpora in that language and training samples labeled with the meanings of polysemous terms in different contexts. The legal text corpora can be obtained from publicly available legal databases (such as the United Nations Legal Library, the World Intellectual Property Organization legal database, etc.), and the labeling work is completed by professionals with legal backgrounds in that language. For less commonly used languages ​​with scarce labeling resources, an active learning strategy can be adopted, first training an initial model with a small number of labeled samples, then using the model to assist in labeling more samples, iteratively optimizing the model. Fine-tuning is then performed on the XLM-RoBERTa model, resulting in the target semantic parsing model. The fine-tuned target semantic parsing model learns the semantic distinction boundaries of polysemous legal terms in different contexts within that language.

[0016] Specific languages ​​refer to less commonly spoken languages ​​with religious or cultural backgrounds used in the legal documents being examined, such as Arabic and Thai. Some terms in legal texts carry different meanings in secular legal contexts and religious / cultural contexts. For example, the Arabic term "..." "In general commercial contracts, it signifies agency, while in documents involving religious obligations, it also implies religious fiduciary duty."

[0017] When parsing legal documents, the target semantic parsing model first identifies the language type of the document and invokes the cultural feature configuration learned during fine-tuning for that language. The cultural feature configuration defines the set of polysemous terms influenced by religious and cultural backgrounds in that language, as well as the semantic distinction rules for each polysemous term in different contexts. The target semantic parsing model performs word segmentation and syntactic analysis on the legal document to locate words belonging to the polysemous term set. Then, based on the sentence context and paragraph theme of the word, combined with the contextual judgment conditions set in the cultural feature configuration, the target semantic parsing model determines the specific referential meaning of the word in the current legal document. For example, when " "When the terms involve descriptions of the trustee's obligations and contain religious terminology, the target semantic parsing model will..." The meaning of "" is defined as religious fiduciary responsibility rather than agency.

[0018] Through the above process, the target semantic parsing model extracts legal elements from legal documents. These legal elements include contract parties, rights and obligations clauses, geographical identifiers, and industry type identifiers. At the same time, it extracts disambiguated legal semantic information, which is the accurate legal meaning corresponding to each legal element.

[0019] Because the target semantic parsing model completes the cultural context disambiguation of polysemous terms during the parsing phase, the extracted legal elements and legal semantic information can reflect the true legal meaning of the text, avoiding semantic errors caused by confusing secular legal concepts with religious obligation concepts.

[0020] In this application, the target semantic parsing model refers to a multilingual pre-trained model fine-tuned by legal text corpora in a specific language. This model can distinguish the different semantics of the same term in secular legal context and religious cultural context based on the context of the legal term, and output the semantic parsing result after cultural context disambiguation.

[0021] In step 02, based on a multi-dimensional scoring framework for a specific language, legal elements and legal semantic information are evaluated to determine the initial scores for each dimension. The multi-dimensional scoring framework consists of dimensions corresponding to different types of legal rules, with each dimension corresponding to a set of legal rules.

[0022] During the evaluation, identifying information associated with each dimension is obtained from the extracted legal elements and legal semantic information. This identifying information is compared with the set of legal rules corresponding to that dimension. The initial score for that dimension is determined based on the degree of coverage between the identifying information and the rules in the legal rule set. Coverage indicates the extent to which the content involved in the legal elements and legal semantic information is encompassed by the legal rule set for that dimension. A larger coverage results in a higher initial score for that dimension, and a smaller coverage results in a lower initial score.

[0023] The initial scores for each dimension are used in subsequent steps to dynamically adjust the weights in conjunction with the regional rule base. The initial scores themselves do not involve adjusting the priority between dimensions.

[0024] In step 03, the regional rule base is composed of two parts: regional religious rules and commercial practice rules. The regional religious rules originate from a regional religious code knowledge base, which stores religious commercial rules from multiple regions, each marked with a corresponding regional identifier. The commercial practice rules originate from a commercial practice database, which stores industry practices for various types of legal acts, each marked with a corresponding legal act type identifier.

[0025] The following explanation uses a petroleum equipment procurement contract from country X as an example, but the technical solution of this invention is not limited to this and can be applied to any legal document with regional religious rules. Taking a petroleum equipment procurement contract from Country X as an example, the legal elements are extracted, identifying the geographical information as "Country X" and the legal act type as "international sale of goods." Using "Country X" as the search condition, the database retrieves religious business rules with the geographical identifier "Country X," as well as relevant provisions for Country X in commercial clause annotations and specific religious financial compliance standards (such as Islamic financial compliance standards). Using "international sale of goods" as the search condition, the database retrieves industry practices with the legal act type identified as "international sale of goods." The retrieved religious business rules and industry practices are then combined to form a regional rules database.

[0026] In this application, a regional rule base refers to a knowledge base composed of religious business rules retrieved from a regional religious code knowledge base and industry practices retrieved from a commercial practice database, based on the geographical information and type of legal acts in legal documents. It is used to characterize religious compliance requirements and commercial practice constraints related to specific types of legal acts within a specific region.

[0027] In step 04, the legal elements and legal semantic information are compared one by one with the rules in the regional rule base. When the comparison results show that the legal elements and legal semantic information are related to the religious rules in the regional rule base, it indicates that the legal document to be tested involves religious compliance requirements in a specific region. At this time, it is necessary to adjust the weight allocation to reflect the priority of religious rules in the evaluation.

[0028] Taking the oil equipment procurement contract of Country X as an example, in the regional rule library, the provisions on prohibiting unfair transactions in the commercial terms notes are related to the price and acceptance clauses in the contract, and the requirements on prohibiting interest in the specific religious financial compliance standards are related to the payment clauses in the contract.

[0029] Due to the aforementioned correlation, the weight coefficient of the industry rules dimension is increased accordingly, while the initial score of the international law dimension and the weight coefficient of the national law dimension are decreased accordingly, resulting in a dynamic weight allocation.

[0030] In step 05, the legal elements and legal semantic information are compared with the corresponding legal rule base one by one, the number of inconsistent articles is counted, and the proportion of inconsistent articles to the total number of relevant articles is subtracted by 1 to obtain the initial score of the corresponding dimension.

[0031] Taking the oil equipment procurement contract of Country X as an example, the initial scores for the three dimensions are 0.80, 0.67 and 0.38, respectively, and the adjusted weight coefficients are 0.20, 0.15 and 0.65, respectively. The weighted summation yields a comprehensive conflict quantification score of approximately 0.51.

[0032] The comprehensive conflict quantification score is compared with a preset risk level threshold to assess the degree of conflict risk in the legal text. Within the preset risk level threshold, a score below 0.40 is considered high risk, 0.40 to 0.70 is medium risk, and a score above 0.70 is low risk. A score of 0.51 falls into the medium risk range, indicating a potential conflict risk regarding religious compliance in a specific region, requiring compliance adjustments.

[0033] In this application, the comprehensive conflict quantification score refers to the value obtained by weighting and summing the initial scores of each dimension with the weight coefficients in the dynamic weight allocation result. The value ranges from 0 to 1 and is used to comprehensively evaluate the degree of compliance conflict risk of legal texts under multiple dimensions. The closer the value is to 1, the higher the compliance level and the lower the conflict risk.

[0034] The real-time compliance conflict detection method for legal documents provided in this application uses a target semantic parsing model designed for the cultural characteristics of legal terminology in specific languages ​​to analyze the legal documents under test. This allows the same term to be assigned different meanings in secular legal contexts and religious obligation contexts. The extracted legal elements and legal semantic information can reflect the true legal meaning of the text, avoiding subsequent compliance assessments based on erroneous semantics. By determining a regional rule base based on the geographical information and legal behavior type in the legal elements, and adjusting the weight coefficients of each dimension accordingly, the compliance conflict quantification scoring process can distinguish the differences in the application of religious rules in different regions. The deviation between the compliance conflict quantification score and the actual compliance risks faced by the legal text in the corresponding region is reduced, improving the accuracy of assessing the degree of conflict risk for legal texts involving regional religious rules.

[0035] Figure 2This is a schematic diagram of the modules of the real-time compliance conflict detection system for legal documents using a large language model, as described in this embodiment of the invention. The system includes an acquisition module, a large language model parsing module, and an output module. The acquisition module is configured to acquire the legal documents to be detected. The large language model parsing module is configured to parse the legal documents based on a target semantic parsing model, extracting legal elements and legal semantic information. The target semantic parsing model is set to account for the cultural characteristics of legal terminology in specific languages ​​within the legal documents. Based on a multi-dimensional scoring framework for specific languages, the system evaluates the legal elements and legal semantic information to determine the initial scores for each dimension. Based on the regional information and legal behavior type in the legal elements, the system determines the regional rule base of regional religious rules and commercial practice rules corresponding to the legal text. Based on the regional rule base of regional religious rules and commercial practice rules corresponding to the legal text, the system adjusts the weight coefficients of each dimension to determine the dynamic weight allocation result. The weight coefficients corresponding to each dimension in the dynamic weight allocation result are weighted and calculated with the initial score of each dimension to determine the comprehensive conflict quantification score. The output module is configured to output the comprehensive conflict quantification score.

[0036] In some embodiments, step 01 above includes: Using a multilingual pre-trained model as the target semantic parsing model, the legal documents are semantically parsed to obtain preliminary semantic parsing results; The semantic vectors of each word in the preliminary semantic analysis results are matched with the semantic vectors of each entry in the regional religious code knowledge base; Based on the matching results, semantic disambiguation is performed on polysemous legal terms in legal documents that are influenced by religious and cultural backgrounds, in order to determine the accurate legal semantics of polysemous legal terms in the corresponding cultural context.

[0037] Specifically, when parsing the legal documents to be tested, XLM-RoBERTa performs word segmentation on the text, dividing it into token sequences. Then, it performs semantic encoding on the token sequences, outputting the semantic vector of each token as the preliminary semantic parsing result.

[0038] The semantic vector of each word in the preliminary semantic analysis results is compared with the semantic vector of each entry in the regional religious code knowledge base using cosine similarity calculation.

[0039] The cosine similarity is calculated as follows: the dot product of the word semantic vector and the item semantic vector is divided by the product of the magnitudes of the word semantic vector and the item semantic vector. The word semantic vector is denoted as A, the item semantic vector as B, and the components of each dimension of the word semantic vector are denoted as A1, A2 to An, and the components of each dimension of the item semantic vector are denoted as B1, B2 to Bn, where n is the vector dimension.

[0040] Multiply A1 by B1, add A2 by B2, and so on up to An by Bn to get the dot product. Add the squares of A1 and A2, and so on up to An, then take the square root to get the magnitude of vector A. Add the squares of B1 and B2, and so on up to Bn, then take the square root to get the magnitude of vector B. Divide the dot product by the product of the magnitudes A and B to get the cosine similarity.

[0041] The cosine similarity value ranges from -1 to 1, with values ​​closer to 1 indicating greater semantic similarity between the word and the entry. A cosine similarity greater than or equal to 0.8 indicates a successful match between the word and the entry.

[0042] When a polysemous legal term is matched to an entry in a regional religious code knowledge base, the matching result includes religious annotations or cultural context information corresponding to that entry. Based on the religious annotations in the matching result, the specific meaning of the polysemous legal term in the corresponding cultural context is determined. For example, the Arabic term "..." For example, the religious annotations in the matching results indicate that the term refers to religious fiduciary duty in contexts involving fiduciary duties and religious terminology, and to agency in secular commercial contract contexts.

[0043] In some implementations, the regional religious code knowledge base stores religious annotations, and the matching result is the corresponding entry in the regional religious code knowledge base that matches the polysemous legal terms in the preliminary semantic analysis results.

[0044] Based on the matching results, semantic disambiguation of polysemous legal terms in legal documents that are influenced by religious and cultural backgrounds includes: Based on the corresponding entries in the matching results, determine the religious annotations in the regional religious code knowledge base corresponding to the preliminary semantic analysis results; Furthermore, based on the corresponding entries in the matching results, the cultural context information corresponding to the preliminary semantic analysis results in the regional religious code knowledge base is determined.

[0045] Specifically, the regional religious code knowledge base stores religious annotations, which are stored as structured fields in each entry of the regional religious code knowledge base. The matching result is the corresponding entry in the regional religious code knowledge base for the polysemous legal terms in the preliminary semantic analysis results. The corresponding entry contains a religious annotation field and a cultural context information field.

[0046] Taking a legal document involving a commercial agency contract in Southeast Asia as an example, a certain term in the preliminary semantic analysis results can express both a general commercial agency relationship and an agency relationship based on specific religious customs in different contexts within the same language. After semantic vector matching, the corresponding entries are obtained. The religious annotation field of the corresponding entries records that "under specific religious customs, the agent in the agency relationship, in addition to bearing the obligations of the commercial contract, also bears religious moral responsibilities, including the obligation to safeguard the entrusted property, which is not automatically terminated upon the expiration of the contract." The cultural context information field records that "in the commercial practices of a specific region, this type of agency relationship must be registered with a religious institution before it can take effect."

[0047] Taking a legal document involving a land lease agreement in South Asia as an example, a certain term in the preliminary semantic analysis results expresses an ordinary lease relationship in a secular legal context, but a lease relationship with specific religious restrictions in a religious legal context. After matching the corresponding entry, the religious annotation field records that "under a specific religious legal system, the above term means that the lessee's use of the leased property must comply with religious doctrines, and the lessor has the right to unilaterally terminate the lease relationship if the lessee violates the religious restrictions." The cultural context information field records that "under a specific religious legal system, clauses involving religious uses in land lease relationships take precedence over secular lease regulations."

[0048] In some implementations, the legal document to be detected includes an image to be detected. Step 01 above also includes: Extract image features from the image to be detected. Image features include signature features or symbol features. By using a cross-modal alignment model, textual clauses in legal documents are semantically aligned with image features; Based on the alignment results, it is determined whether the compliance attributes expressed in the text clauses are consistent with the religious or cultural attributes represented by the image features. When the compliance attributes expressed in the text clauses are inconsistent with the religious or cultural attributes represented by the image features, the inconsistency is identified as an implicit compliance conflict.

[0049] Specifically, the legal documents to be inspected include images to be inspected, such as images of seals on contract signing pages or images of signatures on notarized documents.

[0050] When extracting image features from an image to be detected, the image is input into a convolutional neural network for multi-layer convolution and pooling processing, and the output is an image feature vector. Image features include signature features and symbol features. Signature features are visual patterns such as the shape, texture, and text layout of a seal or signature, while symbol features are visual patterns of graphic elements such as religious symbols and certification marks.

[0051] A cross-modal alignment model is used to semantically align text clauses in legal documents with image features. The cross-modal alignment model employs a Contrastive Language-Image Pre-training (CLIP) model. CLIP is trained on a large-scale image-text pairing dataset. For specific image types in legal documents, such as seals, signatures, and religious symbols, the CLIP model is fine-tuned using image-text pairing data from legal documents. This fine-tuning dataset includes images of seals, signatures, and religious symbols from various types of legal documents, along with their corresponding text descriptions, to improve the accuracy of the cross-modal alignment model in encoding image features of legal documents, enabling the mapping of text and images to the same semantic space. During semantic alignment, the text clauses are input into CLIP's text encoder to obtain text semantic vectors, and the image features are input into CLIP's image encoder to obtain image semantic vectors. The cosine similarity between the two semantic vectors is then calculated.

[0052] Based on the alignment results, when the cosine similarity between the text semantic vector and the image semantic vector is lower than the preset alignment threshold, it is determined that the compliance attributes expressed by the text clauses are inconsistent with the religious or cultural attributes represented by the image features, and the inconsistency is identified as a hidden compliance conflict.

[0053] The preset alignment threshold is determined based on the validation set performance of the cross-modal alignment model, and is set to 0.65 in this embodiment. When the cosine similarity is lower than 0.65, the semantic alignment between the text clause and the image feature is deemed to have failed.

[0054] In this application, implicit compliance conflict refers to a situation where the compliance attributes expressed by the textual clauses of a legal document are inconsistent with the religious or cultural attributes represented by image features (such as signatures, symbols, etc.). Such conflict cannot be detected by simple textual analysis and requires cross-modal semantic alignment to be discovered.

[0055] Taking a commercial contract from a specific religious region as an example, the text does not contain the phrase "specific religious certification," indicating a general commercial compliance attribute. However, after processing with the CLIP image encoder, the semantic vector of the seal image on the contract's signature page is highly similar to the semantic vector of the text related to "specific religious bank," representing a specific religious financial attribute. The inconsistency between the compliance attribute of the text terms and the religious attribute of the seal image is identified as an implicit compliance conflict.

[0056] In some embodiments, step 02 above includes: Compare legal elements and legal semantic information with regional rule bases; Based on the comparison results, determine whether the legal elements and legal semantic information are related to the religious rules in the regional rule base; If a correlation exists, adjust the weight of the industry rules dimension so that the adjusted weight of the industry rules dimension is greater than the weight of the national law dimension, and also greater than the weight of the international law dimension.

[0057] Specifically, the national law dimension corresponds to a national legal rules database, which stores legal provisions of various countries or regions, such as the provisions on agency terms and exclusive agency clauses in Saudi Arabia's Commercial Agency Law. The international law dimension corresponds to an international legal rules database, which stores international treaties and cross-border regulations applicable to cross-border commercial activities, such as the cross-border payment restrictions in the U.S. Foreign Corrupt Practices Act. The industry rules dimension corresponds to an industry rules database, which stores compliance standards and practices specific to certain industries, such as the prohibition of interest and risk-sharing requirements in specific religious financial compliance standards.

[0058] Using a petroleum equipment procurement contract from Country X as the legal document to be examined, with the contract language being Arabic, the legal elements are identified as follows: the geographical identifier is "Country X," the legal act type is "international sale of goods," the industry type is "energy equipment," and the cross-border clause is "cross-border payment."

[0059] There are 15 articles related to "international sale of goods" in the national legal rules database. By comparing the geographical identifier "Country X" with the applicable geographical field of each of the 15 articles, 13 articles contain "Country X" in their applicable geographical field. The number of matching articles is 13, which is approximately 0.87 of the 15 articles. Therefore, the initial score for the national law dimension is 0.87.

[0060] There are 12 articles related to "international sale of goods" in the international legal rules database. By comparing the cross-border clause identifier "cross-border payment" with the jurisdiction conditions field of each of the 12 articles, 7 articles contain "cross-border payment" in their jurisdiction conditions field. The number of matching articles is 7, representing approximately 0.58 of the 12 articles. Therefore, the initial score for the international law dimension is 0.58.

[0061] There are 8 articles related to "international goods sales" in the industry rule base. The industry type identifier "energy equipment" is compared one by one with the applicable industry fields of the 8 articles. Five articles contain "energy equipment" in their applicable industry fields, resulting in a match of 5 articles, which is approximately 0.63 of the 8 articles. Therefore, the initial score for the industry rule dimension is 0.63.

[0062] The weight adjustment is performed as follows: Let the initial weight for the industry rule dimension be w_industry, the initial weight for the national law dimension be w_national, and the initial weight for the international law dimension be w_international. When a correlation is detected between a legal element and a religious rule in the regional rule base, the weights for the industry rule dimension are adjusted to w_industry' = w_industry × α, the weights for the national law dimension are adjusted to w_national' = w_national × β, and the weights for the international law dimension are adjusted to w_international' = w_international × β, where α > 1, 0 < β < 1, and α and β satisfy w_industry' > w_national' and w_industry' > w_international'. In this embodiment, α = 2.0, and β = 0.5.

[0063] In some embodiments, step 03 above includes: Using geographical information as the primary search criterion, the system retrieves religious business rules associated with geographical information from the regional religious code knowledge base. Using the type of legal act as the second search criterion, search the commercial practice database for industry practices that match the type of legal act. Rules that satisfy the first and second search criteria are merged to form a regional rule base.

[0064] Specifically, the regional religious legal code knowledge base is stored in a relational database format. This knowledge base is constructed by extracting provisions from official religious legal documents of various countries or regions, business guidelines issued by religious institutions, and religious business rules compiled from academic research. Legal experts then annotate each rule with a regional identifier and the applicable type of legal action. The knowledge base employs a regular update mechanism, with a team of experts reviewing and adding newly released religious legal documents and business guidelines quarterly. Each religious business rule record includes a regional identifier field and a rule content field. The regional identifier field stores the name of the country or region to which the religious business rule applies, and the rule content field stores the specific provisions of the rule. During retrieval, regional information is used as a query condition. Records matching the regional identifier field with the regional information are selected from all records in the regional religious legal code knowledge base, and the rule content field of the matching records is used as the retrieved religious business rule.

[0065] The commercial practice database is stored in a relational database format. Each industry practice record contains a legal act type identifier field and a practice content field. The legal act type identifier field stores the name of the transaction type to which the industry practice applies, and the practice content field stores the specific clauses of the practice. During retrieval, the legal act type is used as the query condition. Records in the commercial practice database that match the legal act type identifier field are selected, and the practice content field of the matching records is used as the retrieved industry practice.

[0066] Taking an international sales contract involving Country X as an example, the geographical information is "Country X," and the legal act type is "international sales of goods." In the regional religious code knowledge base, filtering by "Country X" yields religious business rules associated with Country X, including provisions on transaction restrictions in specific religious financial compliance standards and provisions on transaction fairness in specific religious texts.

[0067] By filtering the database of commercial practices using "international sale of goods" as the search criterion, industry practices matching international sale of goods were found, including practices regarding the delivery of goods in the Incoterms and practices regarding payment methods in trade in specific regions.

[0068] The rule content fields of religious business rules that meet the first search condition and the custom content fields of industry practices that meet the second search condition are stored in the same data table, which is the regional rule base. Each record in the regional rule base contains a rule source identifier, which is used to distinguish whether the rule belongs to religious business rules or industry practices.

[0069] The rules in the merged regional rule base include both religious compliance requirements and commercial customary constraints. They can provide restrictions and compliance requirements for transactions related to religious doctrines within a specific region, as well as common practices and standards for specific types of transactions in business practice.

[0070] In some implementations, step 05 above includes: The initial score for each dimension is obtained by comparing the legal elements and legal semantic information with the national legal rule base corresponding to the national law dimension, the international legal rule base corresponding to the international law dimension, and the industry rule base corresponding to the industry rule dimension, one by one. The comprehensive conflict quantification score is compared with the preset risk level threshold, and the risk level corresponding to the legal text is determined based on the comparison results.

[0071] Specifically, the weight coefficient of each dimension in the dynamic weight allocation result is multiplied by the initial score of the corresponding dimension, and the products of each dimension are summed to obtain the comprehensive conflict quantitative score. The comprehensive conflict quantitative score ranges from 0 to 1, with the value closer to 1 indicating a higher degree of compliance and a lower risk of conflict.

[0072] The comprehensive conflict quantification score is compared with preset risk level thresholds. These thresholds include a first threshold of 0.40 and a second threshold of 0.70. When the comprehensive conflict quantification score is less than the first threshold, the legal text is classified as high-risk. When the comprehensive conflict quantification score is greater than or equal to the first threshold and less than or equal to the second threshold, the legal text is classified as medium-risk. When the comprehensive conflict quantification score is greater than the second threshold, the legal text is classified as low-risk.

[0073] Taking a petroleum equipment procurement contract from Country X as an example, the dynamic weight allocation results show that the weight coefficient for the national law dimension is 0.20, the weight coefficient for the international law dimension is 0.15, and the weight coefficient for the industry rules dimension is 0.65. The initial score for the national law dimension is 0.80, the initial score for the international law dimension is 0.67, and the initial score for the industry rules dimension is 0.38. The weighted calculation is 0.20 multiplied by 0.80 plus 0.15 multiplied by 0.67 plus 0.65 multiplied by 0.38, resulting in a comprehensive conflict quantification score of 0.51. 0.51 is greater than the first threshold of 0.40 and less than the second threshold of 0.70, thus determining the risk level corresponding to this legal text as medium risk.

[0074] In some implementations, following step 05 above, the real-time compliance conflict detection method for large language models of legal documents further includes: When the risk level corresponding to the legal text is greater than the preset risk level, based on the initial score of each dimension, the clauses whose initial scores are lower than the preset compliance threshold are identified as conflict clauses during the clause-by-clause comparison process. The conflict clause text is input into the large language model, which then uses compliance requirements from the regional rule base as constraints to generate compliance alternative clauses expressed in a specific language based on the original semantics of the conflict clause text.

[0075] Specifically, the method for identifying conflicting clause texts is as follows: During the clause-by-clause comparison process in each dimension when calculating the initial score, the comparison results of each rule with legal elements and legal semantic information are recorded. Each comparison record includes a rule identifier, comparison result, and clause score, whereby the clause score represents the degree of matching between the clause and the rule. When a clause score is lower than a preset compliance threshold, the clause is marked as a conflicting clause. The texts marked as conflicting clauses across all dimensions are then aggregated to obtain the conflicting clause text.

[0076] The conflict clause text is input into the Large Language Model (LLM). The LLM adopts a Generative Pre-trained Transformer (GPT) architecture, specifically GPT-4 (Generative Pre-trained Transformer 4, the fourth generation of generative pre-trained transformer models). GPT-4 has been trained on a large-scale multilingual corpus and has the ability to generate multilingual text.

[0077] When driving the large language model to generate compliance alternative clauses, compliance requirements from regional rule bases are used as constraints and input into the large language model along with the conflicting clause texts. These constraints include religious and business rule provisions and industry practice provisions from the regional rule bases that correspond to the conflicting clauses.

[0078] Under constraints, the large language model rewrites the conflicting clauses based on their original semantics, generating alternative clauses that comply with compliance requirements and maintain the original business intent. The generated compliant alternative clauses are expressed in a specific language that matches the language of the legal document being tested.

[0079] In this embodiment, the preset compliance threshold is set to 0.60, meaning that if a clause score is lower than 0.60, the clause is considered to be in conflict with the corresponding rule.

[0080] Taking a petroleum equipment procurement contract from country X as an example, the risk level is medium, higher than the preset risk level. During a clause-by-clause comparison at the industry rule level, the payment clause in the contract scores 0.25 when compared with the rule prohibiting interest in a specific religious financial compliance standard, which is below the preset compliance threshold of 0.60. Therefore, the payment clause is marked as a conflict clause. The text of the payment clause, "The buyer shall pay five percent of the total contract price as financing interest," is input into the large language model as the conflict clause text. The compliance requirement corresponding to the payment clause in the regional rule base is "Interest is prohibited, and transactions should be based on the principle of risk sharing." Under these constraints, the large language model rewrites the conflict clause text as "The buyer and seller complete the transaction settlement in a risk-sharing manner, and both parties share profits and bear losses according to their investment ratio," generating a compliant alternative clause expressed in a specific language.

[0081] After generating compliance alternative clauses, the generated clauses are compared again with the compliance requirements in the regional rule base to verify whether the generated clauses meet all relevant compliance requirements. If the verification fails, the large language model is re-driven to generate new clauses until the verification passes or the maximum number of retries is reached (3 times in this example). Only compliance alternative clauses that pass the verification are used as the final output.

[0082] In some implementations, the driving large language model uses compliance requirements from a regional rule base as constraints to generate compliance alternatives expressed in a specific language based on the original semantics of the conflicting clause texts. This also includes: Based on the compliance requirements upon which the compliance alternative clauses were generated, retrieve the legal provisions corresponding to the compliance requirements from the regional rule base; Compliance alternatives and legal provisions are used as outputs.

[0083] Specifically, the legal provision field records the article number and original text of the national or international law cited by the rule. During retrieval, the compliance requirements upon which the generated compliance alternative clause is based are used as the query criteria. Records in the regional rule base whose rule content fields match the compliance requirements are then matched, and the legal provision field is extracted from the matched records to obtain the legal provision information.

[0084] Taking a petroleum equipment procurement contract from country X as an example, the compliance requirement upon which the generated compliance alternative clause is based is "prohibition of charging interest, and transactions should be based on the principle of risk-sharing." Records in the regional rule base that match this compliance requirement are identified, and the corresponding legal provision field is extracted. This legal provision field contains "Article X of the Specific Religious Financial Compliance Standard" and the original text of the provision. The compliance alternative clause and the legal provision information are combined as the output, resulting in a final output that includes both the modified contract clause text and the original legal provision upon which the clause is based.

[0085] The output results are organized in a structured format. Compliance alternative clauses are stored in the clause text field, and legal information is stored in the legal basis field. The two fields together form an output record.

[0086] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the server-side functions or steps of a real-time compliance conflict detection method for a large language model of legal documents.

[0087] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a real-time compliance conflict detection method for a large language model of legal documents.

[0088] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0089] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0090] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0091] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A real-time compliance conflict detection method for a large language model of legal documents, characterized in that, Includes the following steps: The target semantic parsing model is used to parse the legal document to be tested, extract legal elements and legal semantic information. The target semantic parsing model is set for the cultural characteristics of legal terminology in a specific language in the legal document. Based on the scoring framework of the specific language in multiple dimensions, the legal elements and the legal semantic information are evaluated to determine the initial score for each dimension. The scoring framework of multiple dimensions includes national law dimension, international law dimension and industry rule dimension. Based on the geographical information and legal behavior type in the legal elements, determine the regional rule base of regional religious rules and commercial practice rules corresponding to the legal text; Based on the regional rule base corresponding to the legal text, which includes regional religious rules and commercial practice rules, the weight coefficients of each dimension are adjusted to determine the dynamic weight allocation result. If the legal elements and the legal semantic information are related to the religious rules in the regional rule base, the weight of the industry rule dimension is greater than the weight of the national law dimension, and also greater than the weight of the international law dimension. The weight coefficients corresponding to each dimension in the dynamic weight allocation result are weighted and calculated with the initial score of each dimension to determine the comprehensive conflict quantification score, which is used to assess the degree of conflict risk of the legal text.

2. The real-time compliance conflict detection method for legal documents using a large language model according to claim 1, characterized in that, The target semantic parsing model is used to parse the legal document to be detected and extract legal elements and legal semantic information, including: Using a multilingual pre-trained model as the target semantic parsing model, the legal document is semantically parsed to obtain preliminary semantic parsing results; The semantic vectors of each word in the preliminary semantic analysis results are matched with the semantic vectors of each entry in the regional religious code knowledge base; Based on the matching results, semantic disambiguation is performed on the polysemous legal terms in the legal documents that are influenced by religious and cultural backgrounds, so as to determine the accurate legal semantics of the polysemous legal terms in the corresponding cultural context.

3. The real-time compliance conflict detection method for legal documents using a large language model according to claim 2, characterized in that, The regional religious code knowledge base stores religious annotations, and the matching result is the corresponding entry in the regional religious code knowledge base that matches the polysemous legal terms in the preliminary semantic analysis result. The semantic disambiguation of polysemous legal terms in the legal documents influenced by religious and cultural backgrounds, based on the matching results, includes: Based on the corresponding entries in the matching results, determine the religious annotations corresponding to the preliminary semantic analysis results in the regional religious code knowledge base; Furthermore, based on the corresponding entries in the matching results, the cultural context information corresponding to the preliminary semantic parsing results in the regional religious code knowledge base is determined.

4. The real-time compliance conflict detection method for legal documents using a large language model according to claim 1, characterized in that, The legal documents to be detected include the images to be detected; The process of parsing the legal document to be detected based on the target semantic parsing model and extracting legal elements and legal semantic information also includes: Extract image features from the image to be detected, including signature features or symbol features; Using a cross-modal alignment model, the textual clauses in the legal documents are semantically aligned with the image features; Based on the alignment results, it is determined whether the compliance attributes expressed by the text clauses are consistent with the religious or cultural attributes represented by the image features. When the compliance attributes expressed by the text clauses are inconsistent with the religious or cultural attributes represented by the image features, the inconsistency is identified as an implicit compliance conflict.

5. The real-time compliance conflict detection method for legal documents using a large language model according to claim 1, characterized in that, The multi-dimensional scoring framework includes national law, international law, and industry rules dimensions. The process of adjusting the weight coefficients of each dimension based on a regional rule base containing regional religious and commercial customary rules corresponding to the legal text to determine the dynamic weight allocation result includes: The legal elements and legal semantic information are compared with the regional rule base; Based on the comparison results, determine whether the legal elements and the legal semantic information are related to the religious rules in the regional rule base; If a correlation exists, adjust the weight of the industry rule dimension so that the adjusted weight of the industry rule dimension is greater than the weight of the national law dimension and greater than the weight of the international law dimension.

6. The real-time compliance conflict detection method for legal documents using a large language model according to claim 1, characterized in that, The regional rule base for determining the regional religious rules and commercial practice rules corresponding to the legal text based on the geographical information and legal act type in the legal elements includes: Using the aforementioned geographical information as the first search condition, religious business rules associated with the geographical information are retrieved from the regional religious code knowledge base. Using the legal act type as the second search condition, search the commercial practice database for industry practices that match the legal act type; Rules that satisfy the first search condition and the second search condition are merged to form the regional rule base.

7. The real-time compliance conflict detection method for legal documents using a large language model according to claim 1, characterized in that, The step of weighting the weight coefficients corresponding to each dimension in the dynamic weight allocation result with the initial score of each dimension to determine the comprehensive conflict quantification score includes: The initial score for each dimension is obtained by comparing the legal elements and the legal semantic information with the national legal rule base corresponding to the national law dimension, the international legal rule base corresponding to the international law dimension, and the industry rule base corresponding to the industry rule dimension, one by one. The comprehensive conflict quantification score is compared with a preset risk level threshold, and the risk level corresponding to the legal text is determined based on the comparison result.

8. The real-time compliance conflict detection method for legal documents using a large language model according to claim 7, characterized in that, After comparing the comprehensive conflict quantification score with a preset risk level threshold and determining the risk level corresponding to the legal text based on the comparison result, the method further includes: When the risk level corresponding to the legal text is greater than the preset risk level, based on the initial score of each dimension, the clauses whose initial scores are lower than the preset compliance threshold are identified as conflict clauses during the clause-by-clause comparison process. The conflicting clause text is input into a large language model, which is then driven to generate a compliance alternative clause expressed in the specific language, based on the original semantics of the conflicting clause text and using compliance requirements in the regional rule base as constraints.

9. The real-time compliance conflict detection method for legal documents using a large language model according to claim 8, characterized in that, The method of driving the large language model, using compliance requirements in the regional rule base as constraints, and generating compliance alternative clauses expressed in the specific language based on the original semantics of the conflict clause text, further includes: Based on the compliance requirements upon which the compliance alternative clauses were generated, legal provisions corresponding to the compliance requirements are retrieved from the regional rule base. The aforementioned compliance alternatives and the aforementioned legal provisions are used as output results.

10. A real-time compliance conflict detection system for a large language model of legal documents, characterized in that, It includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the real-time compliance conflict detection method for a large language model of legal documents as described in any one of claims 1 to 9.