A consumer rights protection intelligent review method and system based on a large language model
By employing an intelligent review method based on a large language model, which combines preprocessing, semantic understanding, and joint retrieval of knowledge graphs and rule bases, the inefficiency and inconsistent results of existing technologies are resolved, achieving efficient, interpretable, and traceable consumer rights protection compliance review.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI RONGSHU INFORMATION TECH CO LTD
- Filing Date
- 2026-01-26
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies are inefficient, produce inconsistent results, have a high rate of misjudgment, lack interpretability and traceability, cannot understand the semantic context of text, and lack the ability to understand and relate structured knowledge in the vertical field of consumer protection.
An intelligent review method based on a large language model is adopted. Through preprocessing, semantic understanding, knowledge graph and rule base joint retrieval and matching, risk factors are identified, compliance judgment is made, and interpretable review results are generated through multi-dimensional calibration and logical consistency verification.
It has enabled efficient, interpretable, and traceable consumer rights protection compliance reviews, improved the efficiency and accuracy of data processing, and enhanced the reliability and interpretability of the results.
Smart Images

Figure CN121563447B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and natural language processing technology, and in particular to a consumer rights protection intelligent review method and system based on a large language model. Background Technology
[0002] In today's digital business environment, conducting consumer protection (hereinafter referred to as consumer protection) compliance reviews of marketing texts, product descriptions, and terms of service is a crucial aspect of platform risk management. Existing technologies have significant and specific shortcomings in data processing, primarily manifested in the following two modes:
[0003] First, the purely manual data processing model is inefficient when handling massive, real-time text data streams, leading to task backlogs. Its processing standards heavily rely on the subjective experience and immediate state of the reviewers, resulting in inconsistent judgments of the same text and a lack of objectivity and repeatability in the data processing results. Furthermore, maintaining a large-scale professional review team incurs high human and management costs. Second, the automated data processing model based on keywords and rule engines has more fundamental technical limitations: its data processing logic is rigid, relying on string matching and failing to understand the semantic context of the text, potentially leading to a high misjudgment rate. For example, it is difficult to distinguish between "historical lowest price" and "historical lowest price information." Faced with variations and homophones in potential risk statements, its feature generalization ability is poor, resulting in a high risk of missed judgments. More importantly, this model lacks the understanding and ability to connect structured knowledge within the consumer protection vertical, making it unable to perform in-depth risk pattern reasoning. In addition, its processing lacks interpretability, typically only outputting binary conclusions and failing to provide review reasons based on evidence fragments, potentially leading to unexplainable results and difficulties in manual review and issue tracing. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a consumer rights protection intelligent review method and system based on a large language model, which automatically transforms the original text into a highly credible review conclusion with interpretability through intelligent processing throughout the entire process.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0006] Firstly, a consumer rights protection intelligent review method based on a large language model, the method comprising:
[0007] The text to be reviewed is preprocessed to obtain structured text feature data;
[0008] The structured text feature data is semantically understood using a pre-trained large language model to obtain a text semantic vector containing domain semantic information;
[0009] By combining knowledge graphs and rule bases, the semantic vectors of the text are jointly retrieved and matched to identify and output a set of risk factors.
[0010] Based on the aforementioned set of risk factors, a compliance assessment is conducted to obtain a preliminary review conclusion that includes risk level, sufficiency of evidence, rule matching degree, and baseline confidence level.
[0011] The risk level, evidence sufficiency, and rule matching degree in the preliminary review conclusion are standardized and quantified to obtain the corresponding numerical feature vector; the numerical feature vector is mapped to the coordinate points in the three-dimensional evaluation feature space composed of risk level, evidence sufficiency, and rule matching degree to obtain a multi-dimensional comprehensive status.
[0012] Based on the coordinates of the overall state, a confidence adjustment strategy is determined, and the preliminary review conclusion is calibrated to obtain an intermediate review result with calibrated confidence.
[0013] The intermediate audit results are subjected to logical consistency verification to obtain the verification confidence level; the verification confidence level and the calibration confidence level are weighted and fused to obtain the final audit result with interpretable audit reasons.
[0014] Secondly, a consumer rights protection intelligent review system based on a large language model includes:
[0015] The preprocessing module is used to preprocess the text to be reviewed to obtain structured text feature data;
[0016] The semantic understanding module is used to perform semantic understanding on the structured text feature data using a pre-trained large language model to obtain a text semantic vector containing domain semantic information.
[0017] The identification module is used to combine the knowledge graph and the rule base to jointly retrieve and match the semantic vector of the text, identify and output a set of risk elements;
[0018] The compliance assessment module is used to make compliance assessments based on the set of risk factors and obtain preliminary review conclusions including risk level, sufficiency of evidence, rule matching degree and benchmark confidence level.
[0019] The state mapping module is used to standardize and quantify the risk level, evidence sufficiency, and rule matching degree in the preliminary review conclusion to obtain the corresponding numerical feature vector; and to map the numerical feature vector to coordinate points in the three-dimensional evaluation feature space composed of risk level, evidence sufficiency, and rule matching degree to obtain a multi-dimensional comprehensive state.
[0020] The calibration module is used to determine the confidence adjustment strategy based on the coordinates of the overall state, calibrate the preliminary audit conclusion, and obtain an intermediate audit result with calibration confidence.
[0021] The fusion module is used to perform logical consistency verification on the intermediate audit results to obtain the verification confidence level; and to perform weighted fusion of the verification confidence level and the calibration confidence level to obtain the final audit result with interpretable audit reasons.
[0022] Thirdly, a computing device, comprising:
[0023] One or more processors;
[0024] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.
[0025] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0026] The above-described solution of the present invention has at least the following beneficial effects:
[0027] The process involves: standardizing and preprocessing the text to be reviewed, transforming it into structured text feature data; transforming fragmented text information into a standardized data format, providing highly adaptable foundational material for subsequent semantic understanding, retrieval, and matching processes, and improving the smoothness of data processing; leveraging a pre-trained large language model to mine the deep semantics of the structured text; converting text features into vector forms carrying domain-specific semantic information, accurately capturing professional semantic relationships within the consumer protection vertical domain, and providing strong semantic support for risk element identification; integrating domain knowledge from knowledge graphs with the judgment criteria of rule bases for joint retrieval and matching; integrating the advantages of both resources to broaden the coverage of risk identification, accurately locating and extracting risk elements, forming a structured set of risk elements, and strengthening the comprehensiveness and relevance of risk identification; conducting systematic compliance judgments based on the complete set of risk elements; and outputting data including risk level and evidence. The preliminary conclusions, encompassing sufficiency, rule matching, and baseline confidence, provide rich and core judgment data for subsequent calibration and verification. Key evaluation indicators in the preliminary conclusions are standardized, quantified, and mapped in three-dimensional space. This eliminates quantitative differences between indicators, concretizing abstract multi-dimensional evaluation data into spatial coordinates, clearly presenting the overall state and providing intuitive data for confidence adjustment. Targeted confidence adjustment strategies are developed based on the three-dimensional spatial coordinates. The preliminary review conclusions are dynamically calibrated to ensure a high degree of compatibility between confidence and overall state, improving the reliability and adaptability of intermediate review results. Logical consistency verification is conducted, and the two types of confidence are integrated. The internal logical reliability verification of the results is supplemented by weighted fusion of multi-dimensional credibility evidence, along with interpretable review reasons, ensuring the final results possess both completeness and traceability, thus enhancing practical application value. Attached Figure Description
[0028] Figure 1 This is a flowchart illustrating an intelligent review method for consumer rights protection based on a large language model, provided by an embodiment of the present invention.
[0029] Figure 2 This is a schematic diagram of a consumer rights protection intelligent review system based on a large language model, provided by an embodiment of the present invention. Detailed Implementation
[0030] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0031] like Figure 1As shown, embodiments of the present invention propose a smart review method for consumer rights protection based on a large language model, the method comprising the following steps:
[0032] Step 100: Preprocess the text to be reviewed to obtain structured text feature data;
[0033] Step 200: Use a pre-trained large language model to perform semantic understanding on the structured text feature data to obtain a text semantic vector containing domain semantic information;
[0034] Step 300: Combine the knowledge graph and rule base to perform joint retrieval and matching of the text semantic vector, identify and output the risk element set;
[0035] Step 400: Based on the set of risk factors, a compliance determination is made to obtain a preliminary audit conclusion that includes risk level, sufficiency of evidence, rule matching degree and benchmark confidence level;
[0036] Step 500: Standardize and quantify the risk level, evidence sufficiency, and rule matching degree in the preliminary review conclusion to obtain the corresponding numerical feature vector; map the numerical feature vector to coordinate points in the three-dimensional evaluation feature space composed of risk level, evidence sufficiency, and rule matching degree to obtain a multi-dimensional comprehensive status.
[0037] Step 600: Determine the confidence adjustment strategy based on the coordinates of the comprehensive state, calibrate the preliminary review conclusion, and obtain an intermediate review result with calibrated confidence.
[0038] Step 700: Perform logical consistency verification on the intermediate audit results to obtain the verification confidence level; perform weighted fusion of the verification confidence level and the calibration confidence level to obtain the final audit result with interpretable audit reasons.
[0039] In this embodiment of the invention, the format of the text data to be reviewed is standardized to reduce redundant information interference and improve the efficiency and standardization of subsequent data processing; domain-related semantic information in the text is deeply mined to allow the data to carry more accurate business-related content, providing high-quality data support for subsequent matching; the related data of the knowledge graph and the standard data of the rule base are integrated to broaden the scope of risk element extraction, making the data basis for risk identification more comprehensive and improving the completeness of element extraction; risk elements are transformed into concrete judgment dimension data, making the data support for the review conclusion clearer and enhancing the traceability of the conclusion; abstract review dimensions are transformed into quantifiable numerical data to achieve unified measurement of multi-dimensional data and make the description of the comprehensive state more accurate; adjustment strategies are formulated based on the coordinate characteristics of quantitative data to provide data basis for confidence calibration and improve the consistency of intermediate results; and verification data and calibration confidence data are integrated to achieve complementary integration of multi-source data, making the data support for the final result more sufficient and enhancing the reliability of the result.
[0040] In a preferred embodiment of the present invention, step 100 above, which preprocesses the text to be reviewed to obtain structured text feature data, includes:
[0041] Step 101 involves cleaning the text to be reviewed, removing irrelevant formatting and noise characters to obtain purified text. Specifically, this includes: first, comprehensively traversing the obtained text to be reviewed to identify irrelevant formatting marks and noise characters. Irrelevant formatting marks include, but are not limited to, line breaks, tabs, special delimiters, and other formatting symbols that do not carry semantic information; noise characters include meaningless garbled characters, repetitive and redundant characters, etc. Then, according to preset character filtering rules, each character obtained through the traversal is verified one by one, retaining valid characters that reflect the core semantics of the text, and removing the identified irrelevant formatting marks and noise characters. Through this filtering and verification process, a purified text with clear semantics and no redundant interference is finally obtained. This purified text lays a standardized text foundation for subsequent word segmentation and part-of-speech tagging processes.
[0042] Step 102 involves performing word segmentation and part-of-speech tagging on the purified text to obtain word sequences and corresponding grammatical tags. Specifically, this includes: based on the purified text, firstly, according to the lexical structure rules, semantic pause features, and common word segmentation rules of Chinese text, the purified text is decomposed into continuous and non-overlapping word sequences, ensuring that each word is a basic linguistic unit capable of independently expressing a certain semantic meaning; on this basis, a preset part-of-speech tagging resource library is retrieved. The preset process for this resource library is as follows: firstly, professional corpus in the field of consumer rights protection is collected, including: compliance review standard texts, descriptions of non-compliance cases, marketing activity rules, product and service terms, etc., while supplementing commonly used corpus in general language scenarios to form a comprehensive corpus set; the comprehensive corpus set is cleaned, removing irrelevant formats, noise characters, and redundant information, and then word segmentation is performed to obtain standardized word units; subsequently, based on Chinese grammar norms and combined with consumer... The usage characteristics of terms in the field of rights protection are analyzed. Each standardized word unit is labeled with its corresponding grammatical category, determining grammatical attributes such as nouns, verbs, adjectives, adverbs, and quantifiers. Next, a fixed mapping relationship is established between the labeled word units and their corresponding grammatical categories, forming a basic resource library. Finally, the basic resource library is iteratively optimized by continuously adding new vocabulary in the field, correcting labeling biases, and supplementing grammatical attribute mappings in special contexts, ultimately resulting in a pre-defined part-of-speech tagging resource library. This resource library stores a large number of mapping relationships between commonly used words and their corresponding grammatical categories. By comparing and matching each word obtained from word segmentation with the entries in the resource library one by one, the grammatical attributes corresponding to each word are determined, and each word is assigned a corresponding grammatical tag. This grammatical tag can determine the grammatical function of the word in a sentence, completing both the structured decomposition of the text and providing basic data support with grammatical attributes for subsequent word vector conversion processes.
[0043] Step 103: Based on the word sequence and its corresponding grammatical tags, a pre-constructed word vector matrix is used for searching and linear transformation to convert each word and its grammatical role into a high-dimensional value in a continuous vector space to obtain the corresponding distributed feature vector. Specifically, this includes: first, retrieving a pre-generated and stored word vector matrix. The pre-construction process of this word vector matrix is as follows: First, collect professional corpora in the field of consumer rights protection, including but not limited to compliance review standard texts, non-compliance case description texts, marketing activity rule texts, and product service terms texts, while supplementing with massive corpora in general language scenarios to form a comprehensive corpus; then, clean the corpora in the comprehensive corpus, removing irrelevant formats, noise characters, and redundant information, and then perform word segmentation and part-of-speech tagging to obtain standardized word sequences and their corresponding grammatical tags; subsequently, based on the standardized corpus, using... The vector training method maps each word to a continuous vector space. Simultaneously, it adjusts the vector values based on the functional attributes of different grammatical roles of the words, enabling the vectors to simultaneously represent the semantic and grammatical features of the words. Finally, through multiple iterations and optimizations, fixed-dimensional vector values are determined, forming a word vector matrix containing a massive amount of initial vector data corresponding to different grammatical roles. This matrix is then stored for later use. The vector data in this matrix are presented in fixed-dimensional numerical form. Next, based on the word sequence and corresponding grammatical tags obtained in step 102, a precise search is performed in the aforementioned word vector matrix to obtain the initial vector corresponding to each word and its grammatical role. Subsequently, according to a preset linear transformation rule, the obtained initial vectors are numerically adjusted and their dimensions optimized. Through linear transformation, the semantic information of the words themselves and the structural information of their grammatical roles are initially fused to obtain a preliminary fused vector.
[0044] In another embodiment, the initial fused vectors are quantitatively adjusted by calculating the length of the arc path in the continuous vector space, so that the semantic distance representation between vectors is more in line with the semantic logical association of natural language. This requires first determining key calculation parameters such as the arc radius and central angle based on the semantic association strength between each initial fused vector and other associated vectors in the same context. The specific parameter determination logic is as follows:
[0045] Semantic association strength is quantified by the co-occurrence frequency ratio of words in the same context and semantic similarity score, with a value range from 0 to 1. The higher the association strength, the smaller the central angle. For example, an association strength of 0.9 corresponds to a central angle of 30 degrees, an association strength of 0.5 corresponds to a central angle of 90 degrees, and an association strength of 0.3 corresponds to a central angle of 150 degrees. The radius of the arc is set according to the coreness of the words in the context. The radius of the core keyword is set to 5, the radius of the auxiliary modifier is set to 3, and the radius of the functional word is set to 1. The length of the arc path is then calculated based on the numerical correlation between the radius and the central angle. In the calculation, the central angle is first converted to radians, and then the specific length is obtained by multiplying the radius and the radian value. At the same time, a specific implementation example is provided in the context of the consumer protection field. If the words corresponding to the initial fusion vector are "unmarked original price" and the words corresponding to the associated vector in the same context are "pricing specification price labeling requirements", the semantic association strength between the two is calculated to be 0.85, corresponding to a central angle of 45 degrees. "Unmarked original price" is used as the core keyword. With the radius set to 5, the final calculated arc length is approximately 3.93. Next, the initial fused vector values are corrected and calibrated based on this arc length. If the initial values of the initial fused vectors are 2.1, 3.5, 4.2..., then the values of each dimension are fine-tuned according to the ratio of the arc length to the radius. Dimensions with larger deviations are shrunk proportionally, while dimensions with smaller deviations remain relatively stable. The corrected vector values are 2.0, 3.3, 4.0..., making the semantic distance representation between vectors more closely match the semantic logical connections of natural language. Finally, each word and its grammatical role after correction and calibration are converted into a high-dimensional value with fixed dimensions in a continuous vector space. This high-dimensional value is the corresponding distributed feature vector. This process not only preserves the core semantic information of the words but also achieves precise correction of the vector values through arc length calculation, further optimizing the numerical conversion effect of text information and providing a more reliable and computable numerical basis for subsequent feature enhancement and selection.
[0046] Step 104 involves enhancing and filtering the distributed feature vectors using a consumer rights protection domain dictionary to obtain structured text feature data. Specifically, this includes: first, loading a pre-built consumer rights protection domain dictionary, which covers professional terminology, common expressions, risk-related vocabulary, and compliance specifications within the consumer rights protection domain, with each term corresponding to a domain attribute identifier; then, comparing the similarity between the distributed feature vectors and the feature vectors corresponding to the terms in the domain dictionary, strengthening the numerical representation strength of feature vectors highly correlated with domain attributes while weakening the numerical influence of feature vectors unrelated to the consumer rights protection domain; further filtering the enhanced distributed feature vectors based on a preset feature filtering threshold, removing non-core feature vectors with numerical representation strength below the threshold, and retaining core feature vectors that accurately reflect the semantics related to the consumer rights protection domain. Through this feature enhancement and filtering operation, structured, domain-specific text feature data is finally obtained, providing high-quality feature support tailored to business needs for subsequent joint retrieval and matching using knowledge graphs and rule bases.
[0047] In this embodiment of the invention, redundant formatting and interfering characters in the text are removed to simplify the data format and improve the purity of the text data; the text is split into ordered word sequences to determine the grammatical attributes of words, giving the text data structured features and providing a clear foundation for subsequent feature extraction; discrete words and grammatical roles are transformed into continuous high-dimensional values, preserving the potential semantic relationships between words, giving the text data quantifiable and calculable characteristics; vertical domain professional vocabulary resources are integrated to strengthen the expression of domain-related features and filter non-core features, making the structured data more suitable for specific business scenarios.
[0048] In a preferred embodiment of the present invention, step 200 above, which utilizes a pre-trained large language model to perform semantic understanding on the structured text feature data to obtain a text semantic vector containing domain semantic information, includes:
[0049] Step 201 involves inputting the structured text feature data as an input sequence into the encoder layer of a pre-trained large language model. Specifically, this includes: first, format normalization and dimensional calibration of the acquired structured text feature data. This structured text feature data is a pre-processed sequence of fixed-dimensional feature vectors containing core feature information related to consumer rights protection. The construction, training, and implementation process of the pre-trained large language model is as follows: first, a massive amount of text corpus in general scenarios is collected, while supplementing it with professional corpus in the field of consumer rights protection, including: compliance review standard text, marketing activity rule text, product and service terms text, and relevant case description text, forming a comprehensive training corpus; the comprehensive training corpus is cleaned to remove irrelevant formats, noise characters, and redundant information, and then subjected to word segmentation, part-of-speech tagging, and sequence normalization to obtain standardized training corpus; subsequently, the model's basic architecture is constructed, including an encoder layer, a multi-head self-attention mechanism module, a multi-layer perceptron, and a domain adaptation layer. The layered network structure determines the functional division of each layer and the data transmission path. Training objectives are set based on standardized training corpora, and model parameters are adjusted through multiple rounds of iterative training. Attention weight allocation and feature transformation logic are optimized to improve the model's ability to capture semantic associations and feature extraction accuracy. After training, the model undergoes performance testing and tuning to ensure stable processing of various text data, ultimately forming a pre-trained large language model adapted to the semantic understanding needs of the consumer rights protection domain, and is then deployed. Further, according to the input format requirements preset by the encoder layer of this pre-trained large language model, the feature vector sequence is length-adapted and data type-converted to ensure that the dimension and numerical range of each feature vector match the input interface of the encoder layer. Through the above format adaptation and dimension calibration operations, the structured text feature data is transformed into an input sequence that meets the model's processing requirements. This input sequence is then completely input into the encoder layer of the pre-trained large language model, laying the data foundation for subsequent feature extraction and interaction within the sequence.
[0050] Step 202 involves extracting and interacting with the multi-head self-attention mechanism in the encoder layer and the multilayer perceptron to obtain a hidden state sequence containing local semantic features. Specifically, after receiving the input sequence, the encoder layer first activates the multi-head self-attention calculation module. This module calculates the association weights of feature vectors at different positions in the input sequence using multiple preset attention heads, assigning different attention weights based on the semantic association between features. This focuses on key features within the sequence while capturing long-distance dependencies between features at different positions. After completing the multi-head self-attention calculation, the feature results output by each attention head are summarized and integrated to obtain a pre-associated feature sequence. Then, the pre-associated feature sequence is input into the multilayer perceptron, where deep processing of the features is performed through multilayer linear transformation and nonlinear mapping to enhance the discriminative and expressive power of the features, while filtering out redundant secondary features. Through the association calculation of the multi-head self-attention mechanism and the deep transformation of the multilayer perceptron, a hidden state sequence containing rich local semantic features is finally obtained. This sequence completely preserves the semantic association information of each local feature in the input sequence.
[0051] Step 203 involves performing context aggregation and pooling operations on the hidden state sequence to obtain document-level features that aggregate contextual semantic information. Specifically, this includes: based on the hidden state sequence, firstly, performing context association analysis on the hidden states at each position in the sequence according to the semantic logical order of the text; then, based on the semantic roles and contextual dependencies corresponding to each hidden state, performing context aggregation using a weighted summation method to integrate local semantic features scattered in different positions into overall semantic information with global association; after completing context aggregation, further performing pooling operations, using preset pooling rules to simplify the dimensions and extract core features from the aggregated global semantic information. These pooling rules can select mean pooling, max pooling, or hybrid pooling methods according to the importance of the text semantics, eliminating irrelevant redundant information and condensing the core semantic content of the text; through the above coherent operations of context aggregation and pooling, a document-level feature that aggregates complete contextual semantic information is finally obtained. This feature is presented in a fixed-dimensional numerical form and can comprehensively represent the overall semantic connotation of the text.
[0052] Step 204 involves inputting the document-level features into the domain adaptation layer of the large language model. Combined with a consumer rights protection domain dictionary, domain semantic enhancement is performed to obtain a text semantic vector containing domain semantic information. Specifically, this includes: First, inputting the document-level features into the pre-trained large language model's built-in domain adaptation layer. This layer pre-loads relevant feature data from a consumer rights protection domain dictionary, which covers professional terminology, compliant expression standards, commonly used vocabulary, and related semantic association rules in the consumer rights protection field. Based on this, the domain adaptation layer compares and matches the document-level features with the vocabulary features in the domain dictionary one by one, calculating their semantic similarity. This strengthens the expression intensity of features highly relevant to the domain dictionary while weakening the influence of features unrelated to the domain. Through the above feature comparison and intensity adjustment domain semantic enhancement operations, the document-level features are fully integrated with the professional semantic attributes of the consumer rights protection field, ultimately transforming into a text semantic vector containing domain semantic information. This vector retains the overall semantic connotation of the text while possessing distinct vertical domain attributes, providing semantic data support for subsequent risk factor identification and matching.
[0053] In this embodiment of the invention, the data input format is standardized to enable efficient adaptation of structured feature data to the model encoder layer, providing a regular input foundation for subsequent semantic extraction; multi-dimensional correlations of features within a sequence are mined to strengthen the interaction and fusion between different features, making the representation of local semantic features richer and more comprehensive; scattered local semantic information is integrated to condense the core semantic content of the text, making the data present global document-level features and improving the overall integrity of the data; vertical domain professional semantic resources are incorporated to strengthen the domain attribute correlation of the data, making the text semantic vector more suitable for specific business scenarios and improving the domain adaptability of the data.
[0054] In a preferred embodiment of the present invention, step 300, which combines a knowledge graph and a rule base to jointly retrieve and match the text semantic vector, identifies and outputs a set of risk elements, includes:
[0055] Step 301 involves calculating the similarity between the semantic vector of the text and the vectorized entities and relationships in the knowledge graph to obtain a set of candidate knowledge fragments. Specifically, this includes: First, determining the foundation of the knowledge graph. This knowledge graph is pre-constructed based on in-depth professional knowledge in the field of consumer rights protection, specifically covering core entities, entity attributes, and inter-entity relationships within this field. Core entities include: promotional activities, price tags, service commitments, standard terms and conditions, product efficacy descriptions, and other objects directly related to compliance review. Entity attributes include: the way promotional content is expressed, the form of price labeling, the specific content of commitments, the structure of clauses, and other detailed information. Inter-entity relationships include: promotion... The system identifies the correspondence between content and actual goods and services, the relationship between price promises and implementation standards, and the binding relationship between standard terms and consumer rights. All entities and relationships have been vectorized using a unified vector encoding method, forming fixed-dimensional vector data to ensure consistency across dimensions and direct comparison calculations. This data is then centrally stored in a graph database. Next, all entity and relationship vectors in the graph database are retrieved and compared one-by-one with the text semantic vectors. A pre-defined cosine similarity calculation rule is used to measure the closeness of the two vectors in the vector space. By comprehensively traversing all entity and relationship vectors in the graph database, the system ensures that no potentially related knowledge content is overlooked.
[0056] Subsequently, the similarity calculation results are sorted and filtered from high to low. The filtering threshold is based on the lowest similarity value of effective knowledge fragments in historical compliance audit cases, and is set in combination with the actual needs of current consumer protection audits. Knowledge content corresponding to entities and relation vectors with similarity reaching the preset threshold is selected. This knowledge content includes knowledge fragments involving specific scenarios such as the use of absolute terms in promotional statements, price tags not including original price descriptions, service commitments not clearly specifying the performance period, and standard terms excluding consumers' main rights. These fragments are integrated to form a candidate knowledge fragment set, which provides rich domain knowledge support that fits the actual audit scenario for the accurate identification of subsequent risk factors.
[0057] Step 302: Based on the candidate knowledge fragment set and the compliance rules in the rule base, perform multimodal matching, calculate semantic relevance and rule fit, and obtain a matching metric. Specifically, this includes: first, determining the core structure of the rule base, which comprehensively integrates various normative requirements related to consumer rights protection. Legal requirements include specific clauses related to text compliance in the Consumer Rights Protection Law, Advertising Law, Price Law, and E-commerce Law; industry compliance standards include e-commerce industry marketing and promotion norms, financial service information disclosure requirements, and life service industry commitment fulfillment standards; platform review norms include a list of prohibited expressions in product promotion, unified price labeling standards, and service... The guidelines outline essential elements of the terms of service, and then break down these requirements into quantifiable and matchable specific compliance rules. Each rule includes its applicable scenario, judgment conditions, and associated semantic features. For example, the rule prohibiting false advertising applies to texts describing product efficacy and service quality, and the judgment condition is that the advertised content is inconsistent with the actual efficacy of the product or the actual standard of the service. The associated semantic feature is the semantic pattern of exaggerated statements and fictitious efficacy descriptions. Similarly, the rule on price labeling applies to texts displaying product prices, and the judgment condition is that the original price and the current price are clearly labeled and the original price is verifiable. The associated semantic feature is the semantic expression pattern of the price components.
[0058] Based on this, multimodal matching is performed on the candidate knowledge fragment set and the compliance rule entries in the rule base. On the one hand, the semantic orientation of each candidate knowledge fragment and each compliance rule entry is analyzed from a semantic perspective, and the semantic relevance is calculated. For example, it is determined whether the best expression of the product's efficacy in the candidate knowledge fragment is closely related to the semantics of the rule entry prohibiting the use of absolute terms in advertising. On the other hand, from the perspective of rule clause adaptation, each candidate knowledge fragment is checked to see if it meets the specific judgment conditions of each compliance rule entry, and the rule fit is calculated. For example, it is checked whether the price expression in the candidate knowledge fragment meets the judgment condition of including the original price and current price.
[0059] Then, semantic relevance and rule fit are integrated according to a preset weight ratio. This weight ratio is determined based on the importance level of the rules and the strength of semantic relevance. For example, the weight of legal and regulatory rules is higher than that of industry standards and platform internal regulations. The weight ratio of semantic relevance and rule fit is set to 4:6. Finally, a matching metric that comprehensively reflects the matching effect is obtained. This metric provides a quantitative basis for the subsequent screening of key risk factors.
[0060] Step 303: Based on the matching metric, the candidate knowledge fragment set is screened and fused to extract key structured risk elements. Specifically, this includes: First, based on the matching metric, a reasonable screening threshold is set. This threshold is determined by referring to the lowest matching metric corresponding to effective risk elements in past audits, combined with the compliance rigor of the current consumer protection audit scenario and historical data verification results. For example, a matching metric of 0.7 is set as the screening threshold. Then, the matching metric corresponding to each candidate knowledge fragment is precisely compared with the screening threshold one by one to ensure unbiased judgment. Candidate knowledge fragments with metric values higher than the threshold are retained, as these fragments are all effective information highly related to compliance risks. At the same time, irrelevant or weakly relevant fragments with metric values lower than the threshold are removed to avoid irrelevant information interfering with the extraction of risk elements.
[0061] Subsequently, the selected candidate knowledge fragments undergo systematic fusion processing. For multiple fragments involving the same entity, relationship, or compliance rule, they are integrated into complete knowledge units through information complementarity and redundancy removal. For example, if multiple candidate knowledge fragments involve the absence of original price information, some fragments mention only the current price, and some fragments mention the lack of historical original price records, these fragments are integrated into complete knowledge units where the price description lacks original price information and historical record descriptions. Duplicate description details are removed, and relevant supporting information from different fragments is added. Through the above screening and fusion operations, key information that reflects text compliance risks is extracted and transformed into structured risk elements. These elements contain core risk-related information, supporting details, and relevant evidence, ensuring the core nature and completeness of the risk elements.
[0062] Step 304 involves merging and formatting the structured risk elements according to risk type, associated entities, and confidence level to obtain a risk element set. Specifically, after extracting the key structured risk elements, they are first precisely categorized according to a preset risk classification standard. This standard comprehensively covers common risk types in the field of consumer rights protection, including: advertising-related risks such as false advertising, absolute language advertising, and misleading advertising; price-related risks such as non-compliant pricing, non-standard price labeling, and unclear price promises; clause-related risks such as unfair standard clauses, vague clause wording, and exclusion of consumers' main rights; and commitment-related risks such as unfulfilled promises, unclear commitment periods, and unenforceable commitment content. Simultaneously, the associated entities corresponding to each structured risk element are analyzed one by one to determine the core objects involved in the risk and other related entities. For example, the associated entities for advertising-related risks include product promotional pages, the entity that wrote the marketing copy, and the actual efficacy parameters of the product; the associated entities for price-related risks include product price tags, pricing decision-making departments, and price implementation standards.
[0063] In addition, the confidence level data corresponding to each risk element in the previous similarity calculation and rule fit integration process is retained. This data directly reflects the reliability of the risk element, and its value is derived from the comprehensive deduction of the previous calculation results. Finally, in accordance with the unified data format requirements, the classified risk types, related entity information, confidence level data and detailed risk descriptions are merged and organized to standardize the presentation format and arrangement order of the data. For example, the data is arranged in order of risk type priority from high to low and confidence level from high to low. The data fields of each risk element are determined to include: risk type code, related entity name, confidence level value, risk description text, and related rule number, forming a set of risk elements with a clear structure and complete information, providing standardized data support that can be directly used for subsequent compliance judgments.
[0064] In this embodiment of the invention, the vector association between text data and domain knowledge is established, broadening the sources of risk-related data and providing a candidate basis for subsequent matching. The dual data dimensions of knowledge level and rule level are integrated to make the measurement of matching results more comprehensive and enhance the effectiveness of data association. The core relevant data is focused on and redundant candidate information is eliminated to make the extraction of risk elements more targeted and improve the refinement of data. The data format of risk elements is standardized to realize the structured presentation of risk-related data, which facilitates the efficient use of subsequent compliance judgment and improves the usability of data.
[0065] In a preferred embodiment of the present invention, step 400 above, based on the set of risk factors, performs a compliance determination to obtain a preliminary audit conclusion including risk level, sufficiency of evidence, rule matching degree, and benchmark confidence level, including:
[0066] Step 401 involves parsing the risk element set and extracting the type, associated entities, and initial confidence level information for each risk element. Specifically, this includes: First, determining the structured data format of the risk element set. This set contains categorized risk types, clear associated entity information, and initially calculated confidence level data, all stored in standardized field formats. Next, a data parsing program is initiated, breaking down the risk element set item by item according to preset field extraction rules, extracting the core information dimensions corresponding to each risk element. Specifically, risk types cover common categories in consumer protection fields such as advertising-related, price-related, terms-related, and commitment-related. Associated entities are identified as specific objects involved in the risk, such as products, marketing materials, service terms, and platform entities. The initial confidence level is a quantitative value generated during the risk element extraction process, reflecting the reliability of the elements. Through the above targeted parsing and extraction operations, the risk element set is transformed into basic data with clear information in each dimension, directly usable for subsequent calculations, providing data support for each stage of compliance determination.
[0067] Step 402 involves matching the risk elements with a pre-set compliance rule library item by item. Through rule template matching and semantic similarity calculation, the rule matching degree for each risk element is obtained. Specifically, this includes: first, retrieving the pre-set compliance rule library. The pre-setting process for this compliance rule library is as follows: First, comprehensively collect various regulatory bases related to consumer rights protection, including specific clauses from laws and regulations such as the Consumer Rights Protection Law, Advertising Law, Price Law, and E-commerce Law; industry compliance standards such as e-commerce industry marketing and promotion compliance standards, financial service information disclosure requirements, and life service industry commitment fulfillment standards; and platform review standards such as the list of prohibited expressions in product promotion, unified price labeling rules, and guidelines for essential elements of service terms. Then, systematically sort and classify these collected regulatory bases, dividing them into dedicated modules according to risk types such as promotional, price, terms, and commitment, ensuring clear rule classification and alignment with review scenarios. Next, each... The specifications under the module are broken down into structured rule templates. Each template contains the core elements of the rule, namely the key information required for judgment, the description of the applicable scenario (i.e., the text type corresponding to the rule, such as marketing text, price labeling text, etc.), and the judgment logic expression (i.e., the definition standard for compliance and non-compliance). This ensures that the information in each template is complete and has the feasibility of direct matching. Then, a vector encoding method consistent with the text semantic vector is used to vectorize each structured rule template, so that the rule template vector and the risk element vector to be matched have the same dimension and comparability. Finally, by continuously collecting newly added legal and regulatory clauses, industry standard updates, and platform review specification adjustments, the judgment logic deviations in the templates are corrected, rule templates for special business scenarios are added, and the rule library is iteratively optimized to form a preset compliance rule library. All compliance rules in this rule library have been converted into structured rule templates and have been vectorized for semantic comparison.
[0068] Based on this, each risk element is matched item by item with the rule templates in the compliance rule base. First, the rule template matching is used to complete the preliminary adaptation at the formal level, and to check whether the applicable scenarios and core information of the risk element are consistent with the preset requirements of the rule template. Then, the preset semantic similarity calculation rules are used to compare the semantic consistency between the risk element and the rule template, and to quantify the degree of semantic fit between the two. Subsequently, the formal matching results and the semantic similarity calculation results are integrated according to a preset ratio to obtain the rule matching degree corresponding to each risk element. This matching degree is presented as a quantitative value within a fixed range, which intuitively reflects the degree of fit between the risk element and the compliance rules.
[0069] Step 403: Based on the rule matching degree and the initial confidence level of the risk element, calculate the compliance deviation assessment value for each risk element. Specifically, this includes: First, determining the core roles of the rule matching degree and the initial confidence level. The rule matching degree directly reflects the degree of fit between the risk element and compliance requirements, while the initial confidence level reflects the reliability of the risk element itself. Together, they constitute the core assessment basis for the compliance deviation. Then, according to a preset calculation logic, the two types of data are integrated. First, the rule matching degree and the initial confidence level are standardized according to a unified numerical range to ensure that they are on the same quantitative dimension and can be directly applied. The calculation process involves setting a weighting scheme based on the core needs of consumer protection compliance audits. The rule matching degree has a higher weight than the initial confidence level, as rule matching is the core basis for determining compliance deviations. The weighting percentages are determined by referencing the influence of the two types of indicators on compliance judgments in historical audit data. For example, the rule matching degree weight is set at 0.6, and the initial confidence level weight is set at 0.4. A weighted summation method is used to calculate the compliance deviation assessment value for each risk element. This assessment value quantitatively reflects the specific degree to which the risk element deviates from compliance requirements, providing a quantitative basis for subsequent risk level determination.
[0070] Step 404: For each risk element, analyze its associated set of evidence fragments. Calculate the sufficiency of evidence by assessing the credibility of the evidence source and semantic relevance. This includes: First, retrieving the set of evidence fragments associated with each risk element. These fragments originate from key sentences extracted from the original text to be reviewed, relevant knowledge entries matched in the knowledge graph, and corresponding compliance clauses in the rule base, with each fragment labeled with its specific source type. Next, a dual-dimensional assessment is performed on each evidence fragment. Firstly, the credibility of the evidence source is analyzed, setting a credibility grading standard based on the source type. For example, fragments from the original text to be reviewed have the highest credibility, followed by knowledge graph related entries, and then rule base fragments. The terms of the document further explain that, according to this standard, each piece of evidence is assigned a corresponding quantitative score for credibility. On the other hand, the semantic relevance of the evidence is calculated by using a preset semantic comparison method to analyze the degree of semantic connection between the evidence piece and the corresponding risk element, quantifying the degree of consistency between the two in terms of core information orientation and logical expression, and obtaining a semantic relevance score. Subsequently, the credibility scores and semantic relevance scores of all evidence pieces corresponding to each risk element are averaged separately, and then weighted and integrated according to preset weight ratios (such as source credibility weight 0.4 and semantic relevance weight 0.6) to obtain the evidence sufficiency assessment value of the risk element. This value directly reflects the strength of the evidence' support for the risk element.
[0071] Step 405: Based on the compliance deviation assessment value and the evidence sufficiency assessment value, determine the risk level of each risk element through a preset risk level judgment matrix. Specifically, this includes: First, retrieving the preset risk level judgment matrix. The preset process of this matrix is as follows: First, define the core assessment dimensions of consumer rights protection compliance audit, and determine the compliance deviation assessment value and the evidence sufficiency assessment value as the core indicators of the two-dimensional matrix, ensuring that the indicators can comprehensively cover the key elements of risk judgment; then, collect a large number of historical compliance audit cases in the field of consumer rights protection, extract compliance deviation-related data and evidence support-related data of the judgment results in the cases, determine the distribution range and key critical points of the two types of data through statistical analysis, divide the compliance deviation assessment value into multiple continuous intervals from low to high, and at the same time divide the evidence sufficiency assessment value into continuous intervals with the same number from low to high, ensuring that the interval division can accurately distinguish the differences in different risk levels.
[0072] Subsequently, based on the severity standards for compliance risks in the field of consumer rights protection, four risk level gradients—low, medium, high, and extremely high—were established, and the core characteristics corresponding to each gradient were determined. For example, extremely high risk corresponds to a significant degree of compliance deviation with sufficient supporting evidence; high risk corresponds to a relatively high degree of compliance deviation with fairly sufficient supporting evidence; medium risk corresponds to a moderate degree of compliance deviation with moderate supporting evidence; and low risk corresponds to a slight degree of compliance deviation with insufficient supporting evidence. Then, through reverse verification using historical audit cases and incorporating industry experts' experience and suggestions on risk assessment, the correspondence between the boundary values of each interval and the risk level in the matrix was adjusted to ensure that each intersecting cell... The corresponding risk levels align with the judgment logic of actual audit scenarios. Finally, by continuously incorporating new audit case data, the interval division and level correspondence rules are dynamically revised to complete the iterative optimization of the matrix, ultimately forming a preset risk level judgment matrix. This matrix is a two-dimensional quantitative assessment matrix, with the horizontal axis representing the range of compliance deviation assessment values and the vertical axis representing the range of evidence sufficiency assessment values. Each intersecting cell in the matrix corresponds to a clear risk level, and each gradient has a clear judgment standard. For example, intersecting cells with high compliance deviation and high evidence sufficiency correspond to extremely high risk, while intersecting cells with low compliance deviation and low evidence sufficiency correspond to low risk.
[0073] Next, the compliance deviation assessment value and the evidence sufficiency assessment value are mapped to the horizontal and vertical axis intervals of the matrix, respectively. The corresponding target cell in the matrix is found through bidirectional positioning. Then, based on the preset risk level definition of the target cell, the risk level corresponding to the risk element is determined to ensure that the level determination of each risk element is based on a unified quantitative standard and logical rule, thus ensuring the consistency and standardization of risk level determination.
[0074] Step 406: Based on the risk level, evidence sufficiency assessment value, and rule matching degree of each risk element, calculate the baseline confidence level using a weighted fusion function; aggregate the risk level, evidence sufficiency assessment value, rule matching degree, and baseline confidence level of all risk elements to obtain a preliminary review conclusion. Specifically, this includes: First, setting the core parameters and weight allocation rules of the weighted fusion function. The weight allocation is determined based on the priority of each indicator's impact on the baseline confidence level, with the risk level having the highest weight because it directly reflects the severity of the risk; the evidence sufficiency assessment value has the second highest weight; and the rule matching degree has the third highest weight. The weighting percentages are determined through a combination of historical audit data verification and expert evaluation. For example, risk level has a weight of 0.4, evidence sufficiency assessment value has a weight of 0.3, and rule matching degree has a weight of 0.3. Based on this, a dot product weighting algorithm is introduced to optimize the weighted fusion effect. This algorithm is a method for weighted integration of multiple indicators through vector operations. The core is to form indicator vectors and weight vectors by combining standardized indicator values with their corresponding weights. The weighted result is obtained by calculating the sum of the products of corresponding elements of the two vectors. This accurately highlights the correlation strength between each indicator and its weight, strengthening the influence of core indicators on the final result. First, the risk level quantification score, evidence sufficiency assessment value, and rule matching degree value corresponding to each risk element are standardized. The specific standardization logic is as follows: the risk level is divided into five levels, with original scores of 1 to 5 points, corresponding to 0.2 to 1.0 after standardization; the evidence sufficiency assessment value and rule matching degree are originally 0 to 100 points, and are converted into standardized values of 0 to 1.0 according to the numerical proportion, ensuring that the values of each indicator are in the same quantitative range and meet the requirements of dot product operation; taking the risk element of not indicating the original price as an example, its risk level is extremely high, with an original quantification score of 5 points, which is 1.0 after standardization. The original evidence sufficiency assessment score was 82, which was standardized to 0.82; the original rule matching score was 86, which was standardized to 0.86. These scores were combined into an indicator vector of 0.86, 0.82, and 1.0. Simultaneously, the preset weights of each indicator (0.3, 0.3, and 0.4) were used to form a weight vector. The sum of the corresponding element-wise products of the two vectors was calculated using dot product operations, i.e., the sum of 1.0 multiplied by 0.4, 0.82 multiplied by 0.3, and 0.86 multiplied by 0.3, resulting in a preliminary weighted result of 0.904. This calculation process accurately reflects the correlation strength between each indicator and its weight, strengthening the influence of core indicators on the final result.
[0075] The core function of this weighted fusion function is to systematically integrate indicators across three key dimensions. By pre-setting weights, it highlights the impact of core indicators while considering the correlation between various indicators, avoiding biases caused by a single indicator dominating the results. The generated value serves as the benchmark confidence level, used to quantify the reliability of the judgment results related to each risk factor, providing a core quantitative basis for the aggregation and calibration of subsequent review conclusions. Next, the preliminary weighted result is substituted into the weighted fusion function and further integrated and optimized using pre-set operational logic. This includes range calibration and reasonableness verification of the preliminary weighted result, such as correcting values exceeding the 0-1 range to the nearest boundary, and addressing obvious... Results deviating from the average of similar risk elements undergo secondary verification to ensure they meet the quantitative standards for confidence level, ultimately yielding a baseline confidence level of 0.90 for each risk element. Subsequently, the judgment data for all risk elements are aggregated and categorized according to risk type, risk level, and other dimensions. The risk level, evidence sufficiency assessment value, rule matching degree, and baseline confidence level of each risk element are summarized in a standardized data format to form a clear and complete preliminary review conclusion. This conclusion fully retains the core judgment criteria and quantitative indicators for each risk element, providing comprehensive data support for subsequent confidence level calibration and logical verification.
[0076] In this embodiment of the invention, the core information dimensions of risk elements are broken down to determine key data items, allowing risk-related data to present a clear and structured form, providing a well-organized foundation for subsequent compliance judgments. A precise correlation is established between risk elements and compliance rules, strengthening the data correspondence through dual matching logic, making rule adaptation results more data-supported and improving the comprehensiveness of the matching process. Two key data categories, rule matching degree and initial confidence level, are integrated to quantify the deviation between risk and compliance requirements, providing measurable numerical evidence for the assessment results and enhancing the objectivity of the assessment. The supporting value of evidence fragments is analyzed from multiple dimensions, integrating source credibility and semantic relevance data, making the supporting strength of evidence quantifiable and improving the data dimensions of the assessment. A correspondence between quantitative data and risk levels is established based on a preset matrix, ensuring that risk level judgments follow a unified standard and guaranteeing the consistency and standardization of the judgment results. A comprehensive confidence index is generated by integrating multi-dimensional assessment data, aggregating key judgment information of all risk elements, making the data dimensions of the preliminary review conclusion complete and logically coherent, providing a comprehensive data foundation for subsequent calibration.
[0077] In a preferred embodiment of the present invention, step 500 involves standardizing and quantifying the risk level, sufficiency of evidence, and rule matching degree in the preliminary review conclusion to obtain corresponding numerical feature vectors; mapping the numerical feature vectors to coordinate points in a three-dimensional evaluation feature space composed of risk level, sufficiency of evidence, and rule matching degree to obtain a multi-dimensional comprehensive state, including:
[0078] Step 501: Extract the values corresponding to risk level, evidence sufficiency, and rule matching degree from the assessment data included in the preliminary review conclusion. Specifically, this includes: First, determining the structured data composition of the preliminary review conclusion. This conclusion has aggregated the core judgment data of all risk elements, where the risk level corresponds to four gradients of quantitative scores: low, medium, high, and extremely high. Evidence sufficiency and rule matching degree are both quantitative assessment values within fixed ranges, and all data are stored with clear field labels. Based on this, the data extraction program is started. According to the preset field retrieval rules, the risk level value, evidence sufficiency value, and rule matching degree value corresponding to each risk element are located one by one from the assessment data of the preliminary review conclusion, ensuring that the value extraction of each indicator is accurate and without confusion or omission. Through the above targeted extraction operation, the values of the three core assessment indicators are separated from the multivariate judgment data to form an independent single-dimensional value set, providing a clear and pure basic material for subsequent standardized calculations.
[0079] Step 502: Based on the preset numerical ranges of each indicator, normalize the values corresponding to the risk level, sufficiency of evidence, and rule matching degree to obtain standardized risk level assessment values, sufficiency of evidence assessment values, and rule matching degree assessment values. Specifically, this includes: first, retrieving the preset numerical range standards for each indicator. The preset process for these standards is as follows: First, for the risk level indicator, combining the risk definition specifications for consumer rights protection compliance review, and referring to the actual judgment of low, medium, high, and extremely high risk gradients in historical review cases, statistically analyze the distribution of the original quantitative scores corresponding to each gradient to determine the numerical boundaries of each gradient. For example, low risk corresponds to 0 to 25 points, medium risk to 26 to 50 points, high risk to 51 to 75 points, and extremely high risk to 76 to 100 points, forming the original quantitative score range for the risk level; then, for the sufficiency of evidence indicator, collect... The initial calculation results of all evidence sufficiency assessments during the preliminary review process are statistically analyzed, and their numerical distribution ranges are determined. Combined with industry experts' grading suggestions on the strength of evidence support, outliers are eliminated to determine a reasonable numerical range and upper and lower boundary values. Then, for the rule matching degree indicator, based on the matching logic of the compliance rule library and historical matching result data, the distribution pattern of the original rule matching degree values is analyzed. Combined with the correlation verification between the matching results and actual compliance judgments, a fixed quantitative range and upper and lower boundaries for the rule matching degree are determined. Finally, a pre-set numerical range standard for each indicator is formed. This standard is determined based on historical data statistical analysis of consumer protection compliance reviews. The original quantitative score range corresponding to the risk level is set according to the definition standards of four gradients. The original numerical ranges for evidence sufficiency and rule matching degree are the fixed quantitative ranges determined during the preliminary calculation process, and the upper and lower boundary values for the numerical range of each indicator are clearly marked.
[0080] Furthermore, for each extracted indicator value, a preset normalization calculation rule is applied for processing. By comparing the individual indicator value with the upper and lower boundaries of the indicator's value range, the differences in the original value ranges between different indicators are eliminated, and all indicator values are uniformly mapped to the standard quantitative range of 0 to 1. Through the above normalization processing, standardized risk level assessment values, evidence sufficiency assessment values, and rule matching degree assessment values are obtained, ensuring that the assessment data of the three dimensions are on the same quantitative benchmark and have the conditions for direct integration and comparative analysis.
[0081] Step 503 involves concatenating the standardized risk level assessment value, evidence sufficiency assessment value, and rule matching degree assessment value according to a preset dimensional order to obtain a numerical feature vector. Specifically, this includes: First, determining the preset dimensional order rules. This order is set in conjunction with the core priorities of consumer protection compliance review. The preset process is as follows: First, analyze the core objectives and assessment logic of consumer protection compliance review to determine that the risk level directly reflects the severity of compliance risk and is the core basis for judging the review result, thus having the highest priority; evidence sufficiency relates to the credibility of the risk level determination and is a key element supporting the risk conclusion, thus having the second highest priority; rule matching degree reflects the degree of fit between risk elements and compliance standards and is a key factor in risk judgment. First, the basic criteria are established, with priority given to the next step. Then, by referencing numerous historical compliance audit cases, the impact of different dimension orders on the accuracy and consistency of audit conclusions is statistically analyzed to verify the rationality of this priority ranking. Next, combined with assessment suggestions from consumer protection experts, the logical connections between dimension orders are adjusted to ensure that the ranking conforms to the business audit process and facilitates subsequent data integration and spatial mapping. Finally, by continuously incorporating new audit scenarios and case data, the dimension order rules are dynamically optimized to ensure they adapt to various consumer protection compliance audit needs. Ultimately, data is concatenated according to the order of risk level assessment value, evidence sufficiency assessment value, and rule matching degree assessment value, and this order applies to the vector construction of all risk elements, ensuring data format consistency.
[0082] Based on this, the three standardized assessment values corresponding to each risk element are sequentially linked and integrated in the above-preset order to form a fixed-length one-dimensional data sequence. The value at each position in this sequence corresponds to a clear assessment dimension without any errors or inversions. Through the above-mentioned orderly splicing operation, the three scattered single-dimensional standardized data are transformed into a well-structured and clearly dimensional numerical feature vector. Each vector can fully carry the multi-dimensional assessment information of a single risk element, providing a standardized data carrier for subsequent three-dimensional spatial mapping.
[0083] Step 504: Using the three standardized assessment dimensions of risk level, sufficiency of evidence, and rule matching degree as the baseline axes, a three-dimensional assessment feature space is constructed. The score values of each dimension in the numerical feature vector are mapped to a coordinate point in the three-dimensional assessment feature space to obtain the comprehensive state of the preliminary review conclusion across multiple dimensions. Specifically, this includes: First, based on the multi-dimensional assessment requirements of consumer protection compliance review, a three-dimensional assessment feature space is constructed based on the standardized assessment dimensions. The horizontal axis is set as the standardized risk level assessment axis, the vertical axis as the standardized sufficiency of evidence assessment axis, and the vertical axis as the standardized rule matching degree assessment axis. The numerical range of all three axes is consistent with the normalized 0 to 1 interval, and each scale on the axis is... The numerical values are clearly quantified. Further, for the numerical feature vector obtained in step 503, the specific values corresponding to the risk level assessment value, evidence sufficiency assessment value, and rule matching degree assessment value are extracted and mapped to the horizontal, vertical, and axial axes of the three-dimensional assessment feature space, respectively, to determine the precise position of each value on the corresponding axis. Through the collaborative positioning of the three-dimensional values, a unique coordinate point is formed in the three-dimensional assessment feature space. The position of this coordinate point intuitively reflects the comprehensive status of a single risk element in the three dimensions of risk severity, evidence support strength, and rule conformity. This transforms the abstract multi-dimensional assessment data into concrete spatial location information, providing intuitive and comprehensive data support for subsequent confidence calibration and the formation of the final review conclusion.
[0084] In this embodiment of the invention, the focus is on core assessment indicators. The numerical values corresponding to risk level, sufficiency of evidence, and rule matching degree are precisely separated from the multi-dimensional data of the preliminary review conclusions, allowing key data to be presented independently and providing targeted and clear foundational data for subsequent standardization processing. The original numerical range differences between different indicators are eliminated, and the values of each indicator are unified to the same quantitative range through normalization calculations, enabling direct comparison and integration of multi-dimensional assessment data and improving data consistency and comparability. The standardized assessment values are integrated according to a preset logic to form a structurally unified numerical feature vector, transforming scattered single-dimensional data into a regular comprehensive data carrier, providing a standardized data form for subsequent spatial mapping. A three-dimensional assessment framework is constructed to concretize abstract data, intuitively presenting the correlation and overall state of assessment values across dimensions through coordinate point mapping, making the multi-dimensional comprehensive assessment results easier to interpret and enhancing the global presentation effect of the data.
[0085] In a preferred embodiment of the present invention, step 600, which involves determining a confidence adjustment strategy based on the coordinates of the overall state and calibrating the preliminary review conclusion to obtain an intermediate review result with calibrated confidence, includes:
[0086] Step 601: Based on the coordinates of the comprehensive state in each dimension of the three-dimensional evaluation feature space, and according to the preset regional division threshold, determine the evaluation region to which the coordinates belong. Specifically, this includes: First, determining the three reference axes of the three-dimensional evaluation feature space, corresponding to the standardized risk level evaluation value, evidence sufficiency evaluation value, and rule matching degree evaluation value, respectively. The numerical range of each axis is a standard interval of 0 to 1. On this basis, retrieving the preset regional division threshold. The specific process of setting this threshold is as follows: First, deeply integrating the core requirements of consumer protection compliance review, determining the three core bases for threshold setting: historical data statistics, industry expert experience demonstration, and adaptation requirements for different risk scenarios, ensuring that the threshold conforms to the actual review data patterns and fits the actual business scenario; then, extensively collecting consumer rights protection data... The historical compliance audit data in the consumer protection field covers complete assessment data for various risk scenarios, including those related to advertising, pricing, terms, and commitments. The distribution of three standardized indicators was extracted, and statistical analysis was used to determine the concentration range and key critical points of each indicator, providing data support for segmentation. Then, considering the refined requirements for confidence level assessment in consumer protection compliance audits, segmentation logic was planned for each of the three baseline axes. Each axis was divided into three continuous numerical segments based on the severity of risk and the reliability of the audit. For example, the risk level assessment axis was divided into segments of 0 to 0.3, 0.3 to 0.6, and 0.6 to 1.0. The evidence sufficiency assessment axis and the rule matching degree assessment axis used the same segmentation range to ensure consistency in the assessment logic. The boundary values of each segment were determined by combining historical data critical points and expert recommendations.
[0087] Subsequently, through multiple rounds of reverse verification using historical cases, the degree of consistency between the assessment areas corresponding to different segment combinations and the actual audit results was compared. Based on the verification results, the segment boundary values were fine-tuned to ensure that each segment could accurately distinguish the comprehensive state of different confidence levels. Finally, the segment boundary values of the three benchmark axes were integrated to form a unified and standardized regional division threshold system. This threshold was set based on historical data statistics of consumer protection compliance audits, industry expert experience demonstrations, and adaptation needs of different risk scenarios. Multiple continuous numerical segments were defined for each of the three benchmark axes, and the boundary values of each segment together constitute the regional division threshold system.
[0088] Furthermore, the specific values of the comprehensive state coordinate points on the three reference axes are extracted, and each value on each axis is compared with the corresponding region division threshold to determine the numerical segment to which the value belongs. Through the combination of numerical segments on the three axes, the unique evaluation region corresponding to the coordinate point in the three-dimensional evaluation feature space is defined. For example, the combination of high risk level segment, high evidence sufficiency segment, and high rule matching degree segment forms a high-confidence stable region, while the combination of medium risk level segment, low evidence sufficiency segment, and medium rule matching degree segment forms a medium-confidence region to be verified, ensuring that the region classification of each coordinate point has a quantitative basis.
[0089] Step 602: Based on the assessment region to which the coordinate point belongs and the preset weight coefficient bound to that region, calculate the confidence adjustment coefficient for the preliminary review conclusion. Specifically, this includes: first, determining the preset weight coefficient system. The process of setting up this system is as follows: First, comprehensively review the core risk characteristics of all assessment regions in the three-dimensional assessment feature space, including: the segmented combination characteristics of risk level, evidence sufficiency, and rule matching degree corresponding to each region, and define the differences in review reliability reflected by different regions. For example, high-confidence stable regions correspond to clear risks, sufficient evidence, and tight rule matching; medium-confidence unverified regions correspond to medium risks and shortcomings in evidence or rule matching; and low-confidence uncertain regions correspond to ambiguous risks and insufficient evidence support. The system identifies weaknesses in rule matching; it then extensively collects historical calibration data from consumer protection compliance audits, including data on the consistency between preliminary audit conclusions and manual review results in different assessment areas, and data on the impact of different adjustment coefficient values on the accuracy of confidence level calibration. Statistical analysis is used to determine the correlation between coefficient values and calibration effectiveness. Furthermore, combined with the experience and expertise of consumer protection experts, the system determines the base range of coefficients for each region's audit reliability level. For example, in high-confidence stable regions, confidence stability needs to be strengthened, with a coefficient range of 0.9 to 1.1; in medium-confidence regions requiring verification, the impact of shortcomings in evidence or rules needs to be reflected, with a coefficient range of 0.7 to 0.9; and in low-confidence uncertain regions, key reviews need to be emphasized, with a coefficient range of 0.5 to 0.7.
[0090] Subsequently, through multiple rounds of reverse verification using historical cases, different coefficient values were substituted into actual review scenarios. The consistency between the calibrated confidence levels and the actual review results was compared, and the boundary values of the coefficient intervals for each region were fine-tuned to ensure that the coefficient adjustments accurately matched the regional characteristics and review requirements. At the same time, special cases at the regional boundaries were considered, and fine-tuning space based on axis value bias was reserved. Finally, a complete preset weight coefficient system was formed. This system binds a unique confidence adjustment coefficient to each evaluation region in the three-dimensional evaluation feature space. The coefficient values are determined based on the risk characteristics, review reliability, and historical calibration effects of the region. For example, the adjustment coefficient for high-confidence stable regions is set to 0.9 to 1.1, aiming to fine-tune the baseline confidence level to enhance its stability; the adjustment coefficient for medium-confidence regions to be verified is set to 0.7 to 0.9, used to moderately reduce the baseline confidence level to reflect the impact of insufficient evidence support or inadequate rule matching; and the adjustment coefficient for low-confidence uncertain regions is set to 0.5 to 0.7, indicating that the conclusion needs to be reviewed more closely.
[0091] Based on this, according to the evaluation area to which the coordinate point belongs as determined in step 601, the preset weight coefficient bound to that area is accurately retrieved. If the coordinate point falls near the boundary of the area, the coefficient is fine-tuned based on its value bias on the three axes. For example, if it is biased towards the high confidence segment, the upper limit of the coefficient interval of that area is taken; if it is biased towards the low confidence segment, the lower limit of the interval is taken. Through the above area matching and coefficient fine-tuning operations, the confidence adjustment coefficient for the current preliminary review conclusion is calculated to ensure the adaptability of the coefficient to the overall status.
[0092] Step 603: Using the confidence adjustment coefficient, dynamically weight the baseline confidence level of the preliminary review conclusion to obtain the calibrated confidence level value. Specifically, this includes: First, extracting the baseline confidence level corresponding to each risk element in the preliminary review conclusion. This value has been calculated using a weighted fusion function and quantifies the reliability of the judgment result before calibration. Based on this, dynamically weighting the confidence adjustment coefficient obtained in step 602 with the corresponding baseline confidence level, the calculation process follows a preset operational logic, i.e., obtaining the confidence level by multiplying the baseline confidence level by the adjustment coefficient. If the initial calibration result exceeds the reasonable confidence range of 0 to 1, range constraint processing is performed to correct the result to the boundary value of the range; if the initial calibration result is within the reasonable range, the result is directly retained as the calibrated confidence value. Through the above dynamic weighting and range calibration operations, the benchmark confidence can be adjusted in a targeted manner according to the regional characteristics of the overall state. This not only retains the core basis of the original judgment, but also makes up for the limitations of a single benchmark confidence through the regional adaptation coefficient, ensuring that the calibrated confidence value is more in line with the reliability requirements of the actual audit scenario.
[0093] Step 604 involves integrating the calibrated confidence level value with the preliminary audit conclusion to obtain an intermediate audit result with calibrated confidence level. Specifically, this includes: First, determining the structured data format of the preliminary audit conclusion, which includes core fields such as risk level, evidence sufficiency assessment value, rule matching degree, and baseline confidence level for each risk element. Based on this, the calibrated confidence level value is added as a new field and linked to the corresponding preliminary audit conclusion data for each risk element, ensuring a one-to-one correspondence between the original judgment information and the calibrated confidence level, without misalignment or omission. Further, according to preset data integration rules, the associated data of all risk elements is uniformly formatted, standardizing the field arrangement order, numerical precision, and expression format. For example, the calibrated confidence level value is retained to two decimal places and arranged in a fixed order with fields such as risk level and evidence sufficiency assessment value. Through the above-mentioned linking and formatting integration operations, a structurally complete and information-rich intermediate audit result is formed. This result includes both the core judgment content of the preliminary audit conclusion and the calibrated confidence level indicator, providing reliable data support for subsequent logical verification and the final audit conclusion output.
[0094] In this embodiment of the invention, the evaluation area is defined based on three-dimensional spatial values and preset thresholds, providing data support for the classification of the overall state and offering directional guidance for confidence adjustment, thus enhancing the pertinence of the adjustment strategy. The evaluation area is correlated with preset weight coefficients to calculate adjustment coefficients, ensuring that the coefficient values are highly compatible with the overall state, guaranteeing the rationality and relevance of the adjustment basis, and improving the accuracy of the coefficients. Dynamic weighted calculation of the calibration benchmark confidence level allows the confidence level value to be dynamically adjusted with the overall state, avoiding the limitations of fixed calculation modes and enhancing the dynamic adaptability of the confidence level. Integrating the calibration confidence level with the preliminary review conclusions enriches the data dimensions of the review results, ensuring that the intermediate results contain both core judgment information and calibrated credibility indicators, thereby improving the completeness and usability of the results.
[0095] In a preferred embodiment of the present invention, step 700 involves performing a logical consistency check on the intermediate audit results to obtain a verification confidence level; and then weighting and fusing the verification confidence level with the calibration confidence level to obtain a final audit result with interpretable audit reasons, including:
[0096] Step 701 involves performing logical conflict detection on the intermediate audit results and their associated evidence fragment sets, identifying and outputting a set of logical conflicts between different risk elements. Specifically, this includes: First, determining the core data composition of the intermediate audit results, which includes the risk level, calibration confidence level, associated entities, and corresponding evidence fragment sets for each risk element. The evidence fragment sets encompass key supporting information such as original text excerpts from the audited text, knowledge graph matching entries, and references to compliance clauses in the rule base. Based on this, a logical conflict detection procedure is initiated, conducting a comprehensive review around the core dimensions of the risk elements. This includes detection dimensions such as the logical compatibility of risk types, the consistency of associated entities, and the logical contradictions in evidence support. Logical compatibility of risk types includes situations where a single product promotional text contains both false advertising and compliant advertising, which are contradictory risk types. The consistency of associated entities includes whether the product entity associated with the risk element is consistent with the product entity pointed to by the evidence fragments. Logical contradictions in evidence support include contradictory conclusions from different evidence fragments regarding the same risk element.
[0097] By comparing each risk element pairwise and cross-validating the core semantics of its associated evidence fragments, combinations of risk elements with logical contradictions and their corresponding points of conflict are identified. For example, if the associated evidence for a non-compliant price risk element shows that the original price was not marked, while the associated evidence for another compliant price risk element shows that the original price was marked, this situation is determined to be a logical conflict. All identified conflict information is organized according to fields such as conflict type, involved risk elements, and source of contradictory evidence to form a structured set of logical conflicts, providing a clear target for subsequent consistency assessment.
[0098] Step 702: Based on the conflict types and quantities in the set of logical conflicts, calculate the logical consistency assessment value of the intermediate audit results using a preset consistency measurement rule. Specifically, this includes: first, determining the preset consistency measurement rule system. The preset process for this system is as follows: First, comprehensively collect historical logical conflict cases in the field of consumer rights protection compliance audits, covering various risk scenarios such as advertising-related, price-related, clause-related, and commitment-related issues. Systematically sort out all conflict types that appear, screen out the core conflict types that significantly affect the reliability of the audit results, and finally determine direct contradictions in risk types, conflicts involving related entities, and contradictory evidence support as the core conflict categories. Then, through statistical analysis of the probability of different conflict types causing deviations between the audit results and the actual compliance situation in historical cases, and combined with the argumentation and assessment of the severity of various conflicts by experts in the consumer protection field, determine the weight levels of various conflicts. Among them, direct contradictions in risk types have the greatest impact on the reliability of the results because they directly negate the core judgment logic, thus having the highest weight. Conflicts involving related entities have the second highest weight because they involve the consistency of the judgment object, and contradictory evidence support has the lowest weight because it can be corrected by supplementary evidence verification.
[0099] Subsequently, to address the impact of the number of conflicts, the consistency differences between intermediate review results and manual review conclusions under different conflict numbers were statistically analyzed. A tiered, incremental impact coefficient was established based on the degree of difference; for each increment in the number of conflicts, the impact coefficient increased by one level, ensuring that the greater the number of conflicts, the more fully the disruption to logical consistency was reflected. Then, through multiple rounds of reverse verification using historical cases, the initially set weights and impact coefficients were substituted into actual review data to compare whether the calculated logical consistency assessment value matched the logical consistency determined by humans. Based on the verification results, the weight values and coefficient tier ranges were fine-tuned to optimize the accuracy of the rules. Finally, new review scenarios were continuously incorporated. The system dynamically updates conflict classifications, weight assignments, and impact coefficient tiers based on new conflict types and manual review feedback data, forming a complete and adaptable pre-defined consistency measurement rule system. This system pre-classifies common logical conflict types in consumer protection compliance audits, with core conflict types including direct contradictions in risk types, conflicts between related entities, and contradictory evidence support. Among these, direct contradictions in risk types have the greatest impact on the reliability of the results and are assigned the highest weight, followed by conflicts between related entities, and then contradictory evidence support. At the same time, an impact coefficient is set for the number of conflicts; the more conflicts there are, the greater the degree of damage to logical consistency, with the impact coefficient increasing in a tiered manner.
[0100] Based on this, the set of logical conflicts is decomposed and analyzed. First, the total number of conflicts is counted, and then each conflict is classified according to the preset conflict type classification standard to determine the specific number of each type of conflict. Next, the number of each type of conflict is multiplied by the corresponding weight, and then multiplied by the conflict quantity influence coefficient to obtain the weighted conflict value of each type of conflict. Then, all weighted conflict values are summed, and the summation result is subtracted from 1 to ensure that the logical consistency assessment value is negatively correlated with the degree of conflict, thus obtaining the logical consistency assessment value of the intermediate review result. This value ranges from 0 to 1. The higher the value, the more logically consistent the intermediate review result is, and vice versa.
[0101] Step 703 converts the logical consistency assessment value into a verification confidence level. Specifically, this includes: First, retrieving a preset conversion rule between assessment values and confidence levels. This rule is based on historical data verification results from consumer protection compliance audits, determining the correspondence between the logical consistency assessment value and the verification confidence level, and ensuring that the converted verification confidence level and the calibration confidence level are in the same quantitative dimension (0 to 1 range). Based on this, the logical consistency assessment value obtained in step 702 is substituted into this conversion rule. If the assessment value is in the range of 0.8 to 1.0, it is directly mapped to a verification confidence level of 0.9 to 1.0, corresponding to a logical height. Consistency is achieved when the evaluation value is between 0.5 and 0.8, which is mapped to a verification confidence level of 0.6 to 0.9, indicating that the logic is basically consistent. When the evaluation value is between 0 and 0.5, which is mapped to a verification confidence level of 0 to 0.6, there is a significant contradiction in the logic. During the conversion process, if the evaluation value falls on the boundary of the interval, it is fine-tuned according to the severity of the conflict type. For example, if the evaluation value is 0.8 and there is no core conflict type, the upper limit of the corresponding confidence level interval is taken. Through the above mapping and fine-tuning operations, the logical consistency evaluation value is accurately converted into the verification confidence level, providing a unified standard quantitative indicator for subsequent multi-dimensional confidence level fusion.
[0102] Step 704: Based on the preset fusion weights, the verification confidence and calibration confidence are weighted and calculated to obtain the comprehensive confidence. Specifically, this includes: First, setting a preset fusion weight system. The preset process of this system is as follows: First, the core positioning and role boundaries of the two types of confidence in consumer protection compliance review are determined. The calibration confidence is generated based on the comprehensive state of the three-dimensional evaluation feature space, which is directly related to the adaptability of the preliminary review conclusion and the actual review scenario, and the reliability of the result judgment. It is the core foundation supporting the final confidence. The verification confidence is generated based on logical consistency verification, focusing on the self-consistency of the internal logic of the review result. It is a supplementary verification to the core foundation. The two work together to ensure the comprehensiveness of the comprehensive confidence.
[0103] Next, historical audit case data for various risk scenarios in the consumer protection field were extensively collected, covering core scenarios such as advertising-related, price-related, terms-related, and commitment-related scenarios. The specific values of calibration confidence and verification confidence in the cases, along with the corresponding accuracy feedback of the final audit results, were extracted. Subsequently, through statistical analysis of historical data, the impact of the two types of confidence on the accuracy of the final results was quantified. It was determined that calibration confidence, because it directly determines the core of the fit between the result and the scenario, has a higher weight on accuracy, while verification confidence, playing a supplementary verification role, has a relatively lower weight. Based on the impact analysis, the weight of calibration confidence was initially set at 0.6 to 0.7, and the weight of verification confidence at 0.3 to 0.4. Then, through multiple rounds of reverse verification using historical cases, different weight combinations were substituted into actual data, and the fused results were compared. By comprehensively considering the consistency between the confidence level and the results of manual review, the boundaries of the weight intervals are fine-tuned to ensure that the weight ratio maximizes the reliability of the final result. Finally, based on feedback from the platform's internal test data regarding improved review accuracy and optimized result consistency, the weight allocation system is dynamically improved, ultimately forming a pre-set fusion weight system. The weight allocation of this system is determined based on the core roles of the two types of confidence levels. The calibration confidence level is generated based on the comprehensive state of the three-dimensional evaluation feature space, directly reflecting the suitability and reliability of the preliminary review conclusion, with a weight ratio set at 0.6 to 0.7. The verification confidence level is generated based on logical consistency verification and serves as a supplementary verification of the inherent reliability of the result, with a weight ratio set at 0.3 to 0.4. This weight ratio is determined by verifying the degree of influence of the two types of confidence levels on the accuracy of the final result in historical review cases.
[0104] Based on this, the verification confidence and calibration confidence are extracted and weighted and summed according to a preset weight ratio, for example, the verification confidence weight is 0.3 and the calibration confidence weight is 0.7. The preliminary fusion result is obtained through the corresponding calculation logic. If the preliminary fusion result exceeds the reasonable range of 0 to 1, boundary constraint processing is performed to correct the result to the nearest boundary value within the range. Through the above weighted fusion and range calibration operations, a comprehensive confidence level with both adaptability and logical reliability is obtained, which fully reflects the reliability of the final audit result.
[0105] Step 705: Based on the comprehensive confidence level, intermediate audit results, and evidence fragment set, obtain the final audit result with interpretable audit reasons. Specifically, this includes: First, determining the structured output requirements for the final audit result. This result must simultaneously include a quantitative confidence level indicator, core judgment information, interpretable audit reasons, and derivation path tracing to address the shortcomings of existing technologies that only output binary conclusions. Simultaneously, a dedicated decision tree model for consumer protection compliance audits is pre-constructed. This model uses a hybrid training logic employing over 100,000 labeled audit cases in the consumer protection field and industry compliance rules. Branch division is optimized using the ID3 algorithm, and node levels are determined with the goal of maximizing information gain. First-level branches must satisfy an information gain value greater than or equal to 0.7. Second-level and lower-level sub-nodes are divided according to quantitative thresholds, such as price-related risk sub-nodes corresponding to a price description ratio greater than... The decision tree is defined as follows: A high-match sub-node has a rule matching similarity of ≥85% and contains core keywords such as pricing and discounts. The confidence factor of the terminal node is assigned based on the historical review accuracy rate, with 1.2 for extremely high-risk conclusions and 0.8 for low-risk conclusions. The root node of the decision tree is the core objective of consumer protection compliance review. First-level branch nodes correspond to key review dimensions such as risk type determination, rule matching degree verification, and evidence sufficiency verification. Second-level and lower-level branch nodes correspond to specific judgment criteria for each dimension. For example, the risk type branch includes sub-nodes related to advertising, price, terms, and commitments; the rule matching degree branch includes sub-nodes such as high matching, medium matching, and low matching; and the evidence sufficiency branch includes sub-nodes such as sufficient, average, and insufficient. Each terminal node corresponds to a clear review conclusion and confidence factor.
[0106] Based on this, data is processed according to a pre-set integration logic combined with a decision tree path algorithm. This algorithm is a traversal algorithm used in a decision tree model specifically for consumer protection compliance audits to trace the derivation trajectory of risk elements. Its core is to locate the starting branch node through feature matching, and then traverse the lower-level nodes according to pre-set rules, ultimately forming a complete flow path from the root node to the terminal node, providing a traceable basis for the audit conclusions. First, the comprehensive confidence level is placed first as the core credibility indicator in the results. Then, details of each risk element in the intermediate audit results are presented, including key information such as risk type, related entities, and risk level. Next, the complete derivation path of each risk element in the decision tree is traced using a depth-first traversal algorithm. First, the core features of the risk element are matched with the first-level branch judgment criteria through keyword hash mapping. Then, the second-level and lower-level nodes are traversed according to the rule that the feature vector similarity is greater than or equal to 90%, recording data in real time. Valid nodes with a score of 75 or higher are matched to form a complete flow path. The branch flow process from the root node to the terminal node is determined and marked. Taking the unmarked original price risk element as an example, the price-related risk first-level branch is matched first. Then, due to the rule matching similarity of 92%, it flows to the high-matching sub-node. Then, due to the evidence supporting 3 or more compliance clauses being cited, it is judged as sufficient. Finally, it flows to the extremely high-risk terminal node. The complete path is: compliance review target price-related risk judgment rule matching degree high, evidence sufficiency sufficient, extremely high risk conclusion. Each branch node has a corresponding matching judgment basis. Then, each risk element is bound to a corresponding set of evidence fragments. The core features of the evidence and the feature vector of the node judgment standard are extracted through a two-way feature alignment algorithm. The matching is confirmed to be valid with a similarity of 80% or higher. The matching dimensions are marked, such as keyword matching and clause matching, to ensure that the evidence fragments correspond one-to-one with the judgment standards of each branch node of the decision tree.
[0107] For identified logical conflicts, an XOR comparison algorithm is used to compare the judgment results of conflict risk elements node by node. A conflict point is identified when the XOR value is 1. The conflict intensity is calculated with a weight of 1.0 for first-level nodes, 0.8 for second-level nodes, and 0.5 for terminal nodes. A core conflict is defined as an intensity greater than or equal to 0.8. Supplementary conflict descriptions are provided, along with the path conflict point in the decision tree. For example, if two risk elements are judged as sufficient and insufficient at the evidence sufficiency verification node, the conflict handling results are also included, such as adjusting the confidence impact factor of the corresponding branch to 0.6 to reduce the confidence of the corresponding risk element. Finally, the conflict is handled according to risk... All information is structured and arranged in descending order of level to form a complete audit rationale, including comprehensive confidence level, details of risk factors, decision tree derivation path, supporting evidence, and conflict explanations. This ensures that the derivation process of each judgment is traceable, the basis for each decision node is verifiable, and the handling logic for each logical conflict is clear. Through the above integration, path tracing, and arrangement operations, a final audit result with a complete structure, comprehensive information, clear derivation logic, and strong interpretability is obtained. This not only meets the judgment requirements of compliance audits but also provides more sufficient derivation basis and supporting materials for manual review and issue tracing.
[0108] In this embodiment of the invention, logical contradictions among risk factors are investigated to identify core conflict information, providing targeted data for subsequent consistency assessment and enhancing the logical comprehensiveness of the audit results. The logical consistency status is quantified by generating an assessment value through comprehensive analysis of conflict types and quantities, providing measurable numerical support for consistency judgments and improving the objectivity of the assessment. The logical consistency assessment results are transformed into a unified-dimensional confidence index, achieving integrability with calibration confidence and laying the foundation for multi-dimensional credibility integration. The dual credibility data from logical verification and prior calibration are integrated, and the two types of core indicators are fused through weighted calculation, allowing the comprehensive confidence to cover multi-dimensional reliability evidence and expanding the data support dimensions of confidence. Finally, the comprehensive confidence, core audit data, and evidence fragments are integrated, ensuring the final result possesses both quantitative credibility and audit basis, enhancing the interpretability and practical value of the result.
[0109] like Figure 2 As shown, embodiments of the present invention also provide a consumer rights protection intelligent review system based on a large language model, comprising:
[0110] The preprocessing module is used to preprocess the text to be reviewed to obtain structured text feature data;
[0111] The semantic understanding module is used to perform semantic understanding on the structured text feature data using a pre-trained large language model to obtain a text semantic vector containing domain semantic information.
[0112] The identification module is used to combine the knowledge graph and the rule base to jointly retrieve and match the semantic vector of the text, identify and output a set of risk elements;
[0113] The compliance assessment module is used to make compliance assessments based on the set of risk factors and obtain preliminary review conclusions including risk level, sufficiency of evidence, rule matching degree and benchmark confidence level.
[0114] The state mapping module is used to standardize and quantify the risk level, evidence sufficiency, and rule matching degree in the preliminary review conclusion to obtain the corresponding numerical feature vector; and to map the numerical feature vector to coordinate points in the three-dimensional evaluation feature space composed of risk level, evidence sufficiency, and rule matching degree to obtain a multi-dimensional comprehensive state.
[0115] The calibration module is used to determine the confidence adjustment strategy based on the coordinates of the overall state, calibrate the preliminary audit conclusion, and obtain an intermediate audit result with calibration confidence.
[0116] The fusion module is used to perform logical consistency verification on the intermediate audit results to obtain the verification confidence level; and to perform weighted fusion of the verification confidence level and the calibration confidence level to obtain the final audit result with interpretable audit reasons.
[0117] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0118] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0119] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A consumer rights protection intelligent review method based on a large language model, characterized in that, The method includes: Step 100: Preprocess the text to be reviewed to obtain structured text feature data; Step 200: Use a pre-trained large language model to perform semantic understanding on the structured text feature data to obtain a text semantic vector containing domain semantic information; Step 300: Combine the knowledge graph and rule base to perform joint retrieval and matching of the text semantic vector, identify and output the risk element set; Step 400: Based on the set of risk elements, a compliance assessment is performed to obtain a preliminary audit conclusion including risk level, evidence sufficiency, rule matching degree, and baseline confidence level. This includes: parsing the set of risk elements to extract the type, associated entities, and initial confidence level information of each risk element; matching each risk element with a preset compliance rule base item by item, and obtaining the rule matching degree corresponding to each risk element through rule template matching and semantic similarity calculation; calculating the compliance deviation assessment value of each risk element based on the rule matching degree and the initial confidence level of the risk element; for each risk element, analyzing its associated set of evidence fragments, and obtaining the evidence sufficiency assessment value through evidence source credibility and semantic relevance calculation; determining the risk level of each risk element based on the compliance deviation assessment value and the evidence sufficiency assessment value through a preset risk level judgment matrix; calculating the baseline confidence level based on the risk level, evidence sufficiency assessment value, and rule matching degree of each risk element through a weighted fusion function; and aggregating the risk level, evidence sufficiency assessment value, rule matching degree, and baseline confidence level of all risk elements to obtain a preliminary audit conclusion. Step 500: Standardize and quantify the risk level, evidence sufficiency, and rule matching degree in the preliminary review conclusion to obtain the corresponding numerical feature vector; map the numerical feature vector to coordinate points in the three-dimensional evaluation feature space composed of risk level, evidence sufficiency, and rule matching degree to obtain a multi-dimensional comprehensive status. Step 600: Determine the confidence adjustment strategy based on the coordinates of the comprehensive state, calibrate the preliminary review conclusion, and obtain an intermediate review result with calibrated confidence. Step 700: Perform logical consistency verification on the intermediate audit results to obtain the verification confidence level; perform weighted fusion of the verification confidence level and the calibration confidence level to obtain the final audit result with interpretable audit reasons.
2. The intelligent review method for consumer rights protection based on a large language model according to claim 1, characterized in that, Step 100 includes: The text to be reviewed is cleaned to remove irrelevant formatting and noisy characters, resulting in cleaned text. The purified text is segmented and part-of-speech tagged to obtain word sequences and corresponding grammatical tags; Based on the word sequence and its corresponding grammatical tags, each word and its grammatical role are transformed into a high-dimensional value in a continuous vector space by searching and linear transformation through a pre-constructed word vector matrix, so as to obtain the corresponding distributed feature vector. The distributed feature vectors are enhanced and filtered by combining a dictionary in the field of consumer rights protection to obtain structured text feature data.
3. The intelligent review method for consumer rights protection based on a large language model according to claim 2, characterized in that, Step 200 includes: The structured text feature data is used as an input sequence and fed into the encoder layer of a pre-trained large language model. The multi-head self-attention mechanism in the encoder layer is used to perform intra-sequence feature extraction and interaction with the multilayer perceptron to obtain a hidden state sequence containing local semantic features. Perform context aggregation and pooling operations on the hidden state sequence to obtain document-level features that aggregate contextual semantic information; The document-level features are input into the domain adaptation layer of the large language model, and combined with the consumer rights protection domain dictionary, domain semantic enhancement is performed to obtain a text semantic vector containing domain semantic information.
4. The intelligent review method for consumer rights protection based on a large language model according to claim 3, characterized in that, Step 300 includes: The similarity between the semantic vector of the text and the vectorized entities and relations in the knowledge graph is calculated to obtain a set of candidate knowledge fragments; Multimodal matching is performed based on the candidate knowledge fragment set and the compliance rules in the rule base to calculate the semantic relevance and rule fit, and obtain the matching metric value. Based on the matching metric, the candidate knowledge fragment set is screened and fused to extract key structured risk elements; The structured risk elements are grouped and formatted according to risk type, related entities, and confidence level to obtain a risk element set.
5. The intelligent review method for consumer rights protection based on a large language model according to claim 4, characterized in that, Step 500 includes: From the assessment data included in the preliminary review conclusion, extract the values corresponding to risk level, sufficiency of evidence, and rule matching degree respectively; Based on the preset numerical range of each indicator, the numerical values corresponding to the risk level, evidence sufficiency and rule matching degree are normalized and calculated respectively to obtain the standardized risk level assessment value, evidence sufficiency assessment value and rule matching degree assessment value. The standardized risk level assessment value, evidence sufficiency assessment value, and rule matching degree assessment value are concatenated in a preset dimensional order to obtain a numerical feature vector. Using risk level, sufficiency of evidence, and rule matching degree as the three standardized evaluation dimensions as the reference axis, a three-dimensional evaluation feature space is constructed, and the score values of each dimension in the numerical feature vector are mapped to a coordinate point in the three-dimensional evaluation feature space to obtain the comprehensive status of the preliminary review conclusion in multiple dimensions.
6. The intelligent review method for consumer rights protection based on a large language model according to claim 5, characterized in that, Step 600 includes: Based on the coordinates of the comprehensive state in each dimension of the three-dimensional evaluation feature space, and according to the preset region division threshold, the evaluation region to which the coordinates belong is determined. Based on the evaluation region to which the coordinate point belongs and the preset weight coefficient bound to that region, calculate the confidence adjustment coefficient for the preliminary review conclusion; Using the confidence adjustment coefficient, the baseline confidence level of the preliminary review conclusion is dynamically weighted and calculated to obtain the calibrated confidence level value; The calibrated confidence level value is integrated with the preliminary audit conclusion to obtain an intermediate audit result with calibrated confidence level.
7. The intelligent review method for consumer rights protection based on a large language model according to claim 6, characterized in that, Step 700 includes: Logical conflict detection is performed on the intermediate audit results and their associated evidence fragments to identify and output the set of logical conflicts between different risk factors; Based on the conflict types and quantities in the set of logical conflicts, the logical consistency evaluation value of the intermediate audit results is calculated using preset consistency measurement rules. Convert the logical consistency evaluation value into a verification confidence level; Based on the preset fusion weights, the verification confidence and calibration confidence are weighted and calculated to obtain the comprehensive confidence. Based on the overall confidence level, intermediate audit results, and set of evidence fragments, a final audit result with explanatory audit reasons is obtained.
8. A consumer rights protection intelligent review system based on a large language model, wherein the system implements the method as described in any one of claims 1 to 7, characterized in that, include: The preprocessing module is used to preprocess the text to be reviewed to obtain structured text feature data; The semantic understanding module is used to perform semantic understanding on the structured text feature data using a pre-trained large language model to obtain a text semantic vector containing domain semantic information. The identification module is used to combine the knowledge graph and the rule base to jointly retrieve and match the semantic vector of the text, identify and output a set of risk elements; The compliance assessment module is used to make compliance assessments based on the set of risk factors and obtain preliminary review conclusions including risk level, sufficiency of evidence, rule matching degree and benchmark confidence level. The state mapping module is used to standardize and quantify the risk level, evidence sufficiency, and rule matching degree in the preliminary review conclusion to obtain the corresponding numerical feature vector; and to map the numerical feature vector to coordinate points in the three-dimensional evaluation feature space composed of risk level, evidence sufficiency, and rule matching degree to obtain a multi-dimensional comprehensive state. The calibration module is used to determine the confidence adjustment strategy based on the coordinates of the overall state, calibrate the preliminary audit conclusion, and obtain an intermediate audit result with calibration confidence. The fusion module is used to perform logical consistency verification on the intermediate audit results to obtain the verification confidence level. The verification confidence level and the calibration confidence level are weighted and fused to obtain the final audit result with interpretable audit reasons.
9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Supply chain contract intelligent review system and method based on large language model
CN120672147A
AI generation content detection and review method and device, equipment and storage medium
CN121093102A
Cited By
A method and system for risk assessment of entry and exit articles at a port based on framework mining
CN122311890A