Judgment document content induction generation method based on image recognition

By employing multimodal information extraction and dynamic rule reasoning methods, the problems of low efficiency and poor accuracy in the traditional process of generating judgment documents have been solved. This has enabled the automated and intelligent generation of judgment documents, adapting to the format requirements of different courts and improving the efficiency and quality of judicial case handling.

CN121189290APending Publication Date: 2025-12-23NANJING TONGDAHAI INFORMATION TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511309383.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

The traditional process of generating court judgments is cumbersome, time-consuming, and susceptible to human factors. Existing technologies are insufficient in recognizing non-standard texts and logical connections, making it difficult to adapt to the personalized needs of different court formats and complex cases. Furthermore, they lack the ability to verify the accuracy of legal provisions.

Method used

By employing multimodal information extraction, dynamic rule reasoning, and intelligent template generation, and using OCR recognition technology to obtain judicial documents, combined with proprietary cognitive models and deep learning algorithms, key elements are extracted and logically linked to generate judicial documents that conform to the format specifications.

Benefits of technology

It has improved the efficiency and accuracy of generating judgment documents, reduced human error, enhanced judicial efficiency, met the format requirements of different courts, and ensured the compliance and logical coherence of the documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121189290A_ABST
    Figure CN121189290A_ABST
Patent Text Reader

Abstract

The invention discloses a judgment document content induction and generation method based on image recognition, relates to the field of legal informatization and artificial intelligence, and aims to solve the problems of many cases and few people of judicial, low processing precision of traditional documents, poor adaptability and the like. The method comprises the following four steps: 1, element extraction: preprocessing and identifying a multi-modal case material through an improved OCR (CNN-Transform fusion model), and extracting key elements in combination with semantic error correction and a proprietary cognitive model (BiLSTM-CRF + knowledge graph); 2, element induction: loading rules according to the action cause, initializing elements and deducing associated elements, and carrying out threshold verification; 3, rule reasoning: circularly deducing core information of the constructed document, referring to current legal provisions, and recording a deduction log; and 4, generating a document, automatically matching court template filling contents, and outputting a compliant document after multi-dimensional verification. According to the method, the document generation efficiency and quality are greatly improved, various judicial requirements are met, and construction of a smart court is assisted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of legal informatization and artificial intelligence fusion, and particularly relates to a judgment document content induction generation method based on image recognition. BACKGROUND

[0002] The traditional judgment document generation process relies on the manual arrangement of case materials by judges, induction of dispute focuses, citation of legal provisions and writing of documents, and has problems such as complicated process, long time consumption, format inconsistency or logical omissions caused by human factors, which not only reduces the judicial case handling efficiency, but also may affect the judicial public credibility.

[0003] The existing legal document automatic processing technology is mostly limited to single text extraction or fixed template filling, and has the following deficiencies: first, the OCR recognition is only for standard printed text, and the recognition accuracy of handwritten notes, blurred signatures, table data and picture evidence (such as accident scene photos and injury identification report drawings) in evidence materials is insufficient, resulting in missing of key information; second, the element extraction relies on static rules, cannot dynamically adjust the extraction dimension according to the case cause change, and lacks deep mining of the logical association between elements, such as the inability to automatically identify the corresponding relationship between “injury grade” and “disability compensation calculation standard”; third, the document generation template is fixed, and it is difficult to adapt to the format specifications of different courts and the individualized expression needs of complex cases, and lacks a verification mechanism for the accuracy of legal provision citation, which is prone to outdated or incorrect citation problems.

[0004] To solve the above technical problems, the present application proposes a judgment document content induction generation method which integrates multi-modal information extraction, dynamic rule reasoning and intelligent template generation, breaks through the limitations of existing solutions through technical innovation, realizes the full-process automation, intelligence and precision of judgment document generation, and provides efficient technical support for judicial practice. SUMMARY

[0005] The present application aims to provide a judgment document content induction generation method based on image recognition to solve the problems raised in the background art.

[0006] To achieve the above-mentioned purpose, the present application provides the following technical solution: a judgment document content induction generation method based on image recognition, comprising the following steps:

[0007] Step 1, element extraction: using OCR recognition technology to obtain the contents of complaint, defense, evidence material list and court record type case materials, processing the contents through a special cognitive model, extracting key element information of the case and storing; the key element information at least includes party information, dispute information and compensation information;

[0008] Step 2, element induction: according to the case cause, obtain the elements and derivation rules that need to be induced from the configuration, initialize the elements according to the element attributes, and then use the extracted elements to derive other related elements according to the derivation rules; The initialization process at least includes date formatting, amount formatting, single / multiple selection content uniform code conversion;

[0009] Step 3, rule reasoning: cyclically deduce the induced elements to construct the dispute information, trial and investigation information, court believes information, judgment result information and legal provisions information of the judicial document, and persistently store the above information to the database;

[0010] Step 4, document generation: obtain the dispute information, trial and investigation information, court believes information, judgment result information and legal provisions information from the database, fill them into the judicial document, and generate the judicial document.

[0011] Preferably, the content in step 1 is processed by a special cognitive model to extract key element information of the case and the specific implementation steps are as follows:

[0012] Step 1.1, case material preprocessing and classification loading: the content data recognized by OCR recognition technology is preprocessed, and the specific processing steps are as follows:

[0013] Material format adaptation and batch import: format analysis: call file analysis interface (such as Apache PDFBox for PDF, POI for DOCX), separate the text layer, image layer and annotation layer in the material; For pure scanning (no text layer), automatically mark as "OCR processing file"; For mixed format files (such as PDF with handwritten annotations), extract the annotation layer separately; Material classification: based on keyword matching (such as "complaint" "defense" "evidence list" title keywords) and text features (such as evidence list table structure, court record "question and answer" expression), automatically classify the materials into "complaint class" "defense class" "evidence material class" "court record class", the classification accuracy rate needs to reach more than 98%, and the materials that do not match successfully are prompted for manual classification assistance;

[0014] Image preprocessing (for scanning / photographing materials): pre-process the image data in the content data, first perform tilt correction, that is, use Hough transform to detect the text line direction in the image, calculate the tilt angle And then the picture is corrected by rotation, and then the noise is removed, and the adaptive median filtering algorithm is used to process the spot noise and paper texture noise in the scanned document, and the filter window size is dynamically adjusted according to the noise density of the pixel neighborhood to avoid the blurring of the traditional median filtering to the text details; Finally, the image data is enhanced by contrast, and the Retinex algorithm is used to enhance the contrast of text and background, and the text blur problem caused by yellowing paper and uneven lighting is solved

[0015] Step 1.2, improved OCR recognition technology to realize text extraction: adopt CNN-Transformer fusion OCR model to realize unified recognition of multi-modal information such as printed text, handwritten notes, table data and picture annotation, and the model architecture is divided into 3 layers:

[0016] Feature extraction layer: improved ResNet-50 network is used as the backbone network, and deformable convolution is added in the network to enhance the feature capture ability of inclined and curved text; That is, the deformable convolution adds an offset to the standard convolution kernel , , is the number of convolution kernel sampling points, so that the sampling points adapt to the shape change of the text, and the offset calculation formula is: , is the offset prediction value of the convolution layer output, is the scaling factor, and the final output dimension of the feature extraction layer is text feature map, wherein , is the height / width of the feature map;

[0017] Text sequence modeling layer: introduce the Transformer encoder (including 6 encoding blocks) to convert the 2D feature map output by the feature extraction layer into 1D text sequence; First, the text position information is injected through position encoding, and the position encoding formula is: ; , is the sequence position; is the dimension index; , that is, the input dimension of the Transformer, and then the self-attention mechanism is used to capture the text context association, and the self-attention weight calculation formula is: , is the query matrix, is the key matrix, is the value matrix, are all mapped from the feature map, is the attention head dimension;

[0018] Text decoding layer: adopt CTC decoding algorithm to realize alignment-free text recognition: By introducing "whitespace symbol" ", with a length of The model output sequence is mapped to a length of The text sequence has the following loss function: ,in For the sample size, For the input image, For real text labels, For model prediction The probability, These are the model parameters; by minimizing this loss function, the model learns the mapping relationship between "image features and text sequences".

[0019] Then, specific optimization strategies are adopted for different types of information to improve recognition accuracy;

[0020] Step 1.3: Semantic error correction algorithm to correct OCR recognition errors: If "similar character errors" or "semantic logic errors" exist after OCR recognition, they are corrected using a semantic error correction algorithm that combines rules and deep learning. The specific steps are as follows:

[0021] Rule-based error correction: Based on the professionalism and standardization of legal texts, three types of rule bases are constructed: Legal terminology dictionary: containing 50,000+ legal professional terms, identifying and correcting terminology errors through string matching; Format rule base: formulating validation rules for formatted information; Contextual logic rules: correcting contradictory errors based on the semantic logic of legal texts.

[0022] Deep learning error correction: For semantic errors that rules cannot cover, a BERT-law pre-trained model is used for contextual semantic correction. The model input is a windowed sequence of OCR-recognized text, and the output is the corrected text sequence. Specifically, a legal text error correction dataset is constructed, and some characters in the correct text are replaced with incorrect characters to form "incorrect text-correct text" training pairs. The LawBERT model is fine-tuned to learn the contextual dependencies between "incorrect characters" and "correct characters." The OCR-recognized text is then processed... Sliding window to extract subsequence Input the LawBERT model, and the model outputs... Corrected probability distribution Choose the character with the highest probability. As a result of the correction, the formula is: ,in For character sets, this method is used to correct "context-dependent errors";

[0023] Step 1.4: Extraction of Key Element Information Using a Proprietary Cognitive Model: The proprietary cognitive model uses a "knowledge graph in the legal domain" as a constraint, integrates a BiLSTM-CRF model with rule-based reasoning, and achieves full-dimensional extraction of "explicit elements + implicit elements." The model architecture consists of three layers:

[0024] Knowledge Graph Constraint Layer: Constructs a knowledge graph of case elements, defines element types and relationships between elements, and provides semantic constraints for element extraction;

[0025] Feature recognition layer: The BiLSTM-CRF model is used to identify explicit feature entities in the text. The input of the model is the semantically corrected text sequence, and the output is the feature entity label (using the BIO labeling system: B-XXX indicates the start of the feature, I-XXX indicates the middle of the feature, and O indicates a non-feature).

[0026] BiLSTM layer: Converts the text sequence into word vectors (using the legal domain word embedding model Word2Vec, with a dimension of 300), inputs them into the BiLSTM network, captures the bidirectional semantic features of the text, and outputs the hidden state. (Dimension is 256), the formula is: ; ; ,in for Word vectors at time points, For positive LSTM output, For inverted LSTM output, The spliced ​​bidirectional features;

[0027] CRF layer: A CRF layer is added after the BiLSTM output layer to learn the transition probabilities between feature labels (e.g., "B-Name of the person concerned" can only be followed by "I-Name of the person concerned" or "O"). Its loss function is: ,in The actual label sequence, For all possible label sequences, For model prediction The probability of element entity recognition; improving the continuity and accuracy of element entity recognition through CRF layer;

[0028] Implicit Element Reasoning Layer: Based on knowledge graphs and rule-based reasoning, implicit element information is derived from explicit elements.

[0029] Preferably, the specific content of image preprocessing in step 1.1 is as follows:

[0030] Tilt correction: The Hough transform maps lines in image space to parameter space. ,in The perpendicular distance from the origin to the line is... is the angle between the straight line and the axis, the formula is ; the text tilt angle is determined by counting the peak value corresponding to the statistical parameter space , and the coordinate transformation of the image pixel is carried out through the rotation matrix to realize the tilt correction;

[0031] Noise removal: use adaptive median filtering algorithm to process the spot noise and paper texture noise in the scanned document; the filter window is defined as , the window size range is , and the processing steps are: calculating the minimum value , maximum value , median value and center pixel value of the window; if , go to the second stage judgment: if , output , otherwise output ; if or , expand the window to , repeat the above steps until the window reaches , and finally output ;

[0032] Contrast enhancement: adopt Retinex algorithm to enhance the contrast of text and background, specifically, the image is decomposed into "reflection component " and "illumination component ", that is , the illumination component is separated through logarithmic transformation and Gaussian filtering, the formula is , wherein is the Gaussian filter kernel, denotes convolution operation.

[0033] Preferably, in step 1.2, different types of information are targeted to improve recognition accuracy through special optimization strategies, and the specific content is:

[0034] Handwritten annotation recognition: build a handwritten text dataset, fine-tune the CNN-Transformer model, adjust and optimize the recognition of connected characters and variant characters; introduce stroke feature constraints, calculate the curvature and direction change of the handwriting trajectory to assist in correcting recognition errors;

[0035] ​Table data recognition: adopt the two-step method of "table structure detection + cell content recognition", first detect the row and column boundaries of the table through the Mask R-CNN model to generate cell coordinates, and then perform OCR recognition on each cell separately, and match the cell content with the table header through coordinate association;

[0036] Picture annotation recognition: for the text annotations in the accident scene photos and injury photos, first detect the text area through the YOLOv8 model, input the cropped image into the OCR model for recognition, and simultaneously verify the rationality of the annotation content by combining the image semantics.

[0037] Preferably, in step 2, after deriving other associated elements, an element verification step is further included: comparing the information of the same element in different case materials, if the information difference exceeds the preset threshold, it is marked as a suspicious element and manual review is prompted.

[0038] Preferably, in step 3, when constructing various information of the judicial document, the corresponding legal provision database needs to be referred to, to ensure that the legal provision information is the current effective and matched legal provision with the case facts; in the persistent storage process, the element derivation process and rule reference log are recorded, to facilitate the subsequent tracing and checking of the judicial document generation process.

[0039] Preferably, in step 4, the judicial document has multiple format templates, and the appropriate template is automatically selected for content filling according to the format requirements of the court of jurisdiction.

[0040] Preferably, in step 4, after generating the judicial document, a document verification step is further included: checking the document format, text expression, and logical coherence, and if there are problems, the modification is prompted until the judicial document that meets the judicial norms is generated.

[0041] Compared with the prior art, the beneficial effects of the present application are:

[0042] Efficiently cracking the pressure of judicial case handling: through the automatic extraction of multi-modal information by OCR, the extraction of elements by special cognitive models, and the automatic construction of document content by rule reasoning, the efficiency of the generation cycle of single case judicial documents can be effectively improved; more than 90% of repetitive links are automatically processed, and a judge who handles 500 cases a year can reduce about 1200 hours of document work each year, focusing on core judicial decision-making, and alleviating the contradiction between "more cases and fewer people".

[0043] Accurate guarantee of document quality: through pre-processing such as Hough transform and adaptive median filtering, combined with a CNN-Transformer fusion OCR model, the multi-modal information recognition accuracy is above 95%, the handwritten annotation recognition accuracy is 92%, the table data extraction error rate is less than 3%, and the information omission problem is solved; the "element + document" double check makes the format compliance rate 100%, the article citation accuracy rate 98%, and the logical contradiction rate below 0.5%, reducing the risk of objection and re-examination caused by document quality.

[0044] Flexible adaptation to judicial practice: a "case cause exclusive rule base + court template base" is constructed, automatically adapting to the element derivation rules of more than 99% of common case causes such as civil and criminal cases, and matching the format specifications of courts across the country; complex cases are generated by LawBERT to generate personalized expressions, and the human editing interface is reserved, the personalized adaptation efficiency is improved, and the diversified needs are met.

[0045] Supporting judicial digitization and traceability: full digital processing reduces the printing cost of paper materials; automatically recording "element source - rule reference - review record" in all links to form an unalterable log, realizing the "checkable, traceable and certifiable" of document generation, and supporting judicial transparency and clean government prevention and control.

[0046] Technological innovation leads the industry: the "multi-modal extraction + knowledge graph + reinforcement learning" architecture is first created, which integrates computer vision, natural language processing and legal knowledge engineering, and the related algorithms can be reused in legal document scenarios such as prosecution and arbitration, promoting the iteration and upgrading of legal information technology. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 The method flowchart of the present application is shown. DETAILED DESCRIPTION

[0048] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0049] Please refer to Figure 1 The present application provides a technical solution: a judgment document content induction generation method based on image recognition, comprising the following steps:

[0050] Step 1, element extraction: the contents of case materials such as indictments, defenses, evidence material lists and court records are obtained by using OCR recognition technology, the contents are processed by a special cognitive model, and the key element information of the case is extracted and stored; the key element information at least includes party information, dispute information and compensation information;

[0051] Element extraction, as the core pre-process of automatic generation of judicial documents, needs to realize the accurate conversion from "unstructured case materials" to "structured key elements". This process takes improved OCR recognition technology as the information input entrance, takes multi-modal fusion special cognitive model as the core of element extraction, and is equipped with semantic error correction and structured storage modules to form a full-process closed loop of "collection-identification-correction-understanding-extraction". The specific implementation steps are divided into five stages, which are closely linked in logic, and the accuracy and efficiency of element extraction are improved through algorithm optimization and model improvement;

[0052] The specific implementation steps for extracting key element information of a case by using a special cognitive model are as follows:

[0053] Step 1.1, case material preprocessing and classification loading: the content data recognized by OCR recognition technology is preprocessed, and the specific processing steps are as follows:

[0054] Material format adaptation and batch import: format analysis: call file analysis interface (such as Apache PDFBox for PDF, POI for DOCX), separate text layer, image layer and annotation layer in the material; for pure scanned materials (without text layer), automatically mark as "OCR processing file"; for mixed format files (such as PDF with handwritten annotations), extract the annotation layer separately; material classification: based on keyword matching (such as "complaint" "defense" "evidence list" and other title keywords) and text features (such as the table structure of evidence list and the "question and answer" expression of court records), the materials are automatically classified into "complaint class" "defense class" "evidence material class" "court record class", and the classification accuracy rate needs to reach more than 98%, and the materials that fail to match are prompted for manual classification assistance;

[0055] Image preprocessing (for scanned / photographed materials): to solve the problem of low recognition accuracy of traditional OCR for blurred, inclined and noisy materials, the image materials need to be preprocessed, and the specific steps are as follows:

[0056] Firstly, the image data in the content data is preprocessed, and the inclination correction is performed, that is, the text line direction in the image is detected by using Hough transformation, and the inclination angle is calculated and rotation correction is performed; then, the noise of the picture is removed, and the adaptive median filtering algorithm is used to process the spot noise and paper texture noise in the scanned materials, and the filter window size is dynamically adjusted according to the noise density of the pixel neighborhood to avoid the blurring of text details caused by traditional median filtering; finally, the contrast of the image data is enhanced, and the Retinex algorithm is used to enhance the contrast between text and background, solving the problem of text blurring caused by yellowing paper and uneven light;

[0057] The specific implementation of image preprocessing is as follows:

[0058] Skew correction: the Hough transform maps a straight line in the image space to a parameter space , where is the perpendicular distance from the origin to the straight line, is the included angle between the straight line and the axis, and the formula is ; the skew angle of the text is determined by counting the value corresponding to the peak in the parameter space, and the coordinate transformation of the image pixel is performed through a rotation matrix to achieve skew correction;

[0059] Noise removal: adaptive median filtering algorithm is used to process the speckle noise and paper texture noise in the scanned document; the filter window is defined as , and the window size ranges from ; the processing steps are as follows: calculate the minimum value , the maximum value , the median value , and the center pixel value in the window; if , go to the second stage: if , output , otherwise output ; if or , expand the window to , repeat the above steps until the window reaches , and finally output ;

[0060] Contrast enhancement: Retinex algorithm is used to enhance the contrast between the text and the background, which is specifically to decompose the image into "reflection component " and "illumination component ", i.e. , separate the illumination component through logarithmic transformation and Gaussian filtering, and the formula is , where is the Gaussian filter kernel, denotes convolution operation.

[0061] Step 1.2, improved OCR recognition technology to realize text extraction: adopt CNN-Transformer fusion OCR model to realize unified recognition of multi-modal information such as "printed text + handwritten notes + table data + picture annotation", and the model architecture is divided into 3 layers:

[0062] Feature extraction layer: an improved ResNet-50 network is used as the backbone network, and a deformable convolution is added to the network to enhance the feature capturing ability of the inclined and curved text; that is, the deformable convolution adds an offset to the standard convolution kernel wherein , is the number of convolution kernel sampling points, and the sampling points adapt to the text shape change, and the offset calculation formula is: wherein is the offset prediction value of the convolution layer output, is a scaling factor, and the final output dimension of the feature extraction layer is text feature map, wherein , is the height / width of the feature map;

[0063] Text sequence modeling layer: a Transformer encoder (containing 6 encoding blocks) is introduced to convert the 2D feature map output by the feature extraction layer into a 1D text sequence; first, the text position information is injected through position encoding, and the position encoding formula is: ; wherein is the sequence position; is the dimension index; , that is, the input dimension of the Transformer, and then the self-attention mechanism is used to capture the text context association, and the self-attention weight calculation formula is: wherein is the query matrix, is the key matrix, is the value matrix, are all obtained by mapping the feature map, is the attention head dimension;

[0064] Text decoding layer: the CTC decoding algorithm is used to realize the alignment-free text recognition: By introducing a "blank symbol ", the model output sequence with a length of is mapped to a text sequence with a length of , and the loss function is: wherein is the number of samples, is the input image, is the real text label, is the probability predicted by the model , and is the model parameter; by minimizing the loss function, the model learns the mapping relationship between "image features and text sequences";

[0065] Then for different types of information, adopt special optimization strategy to improve the recognition accuracy, the specific content is:

[0066] Handwritten annotation recognition: build handwritten text dataset, fine-tune CNN-Transformer model, adjust and optimize the recognition of connected script and variant characters; introduce stroke feature constraint, calculate the curvature and direction change of handwritten trajectory to assist in correcting recognition errors;

[0067] Table data recognition: adopt "table structure detection + cell content recognition" two-step method, first detect the row and column boundaries of the table through Mask R-CNN model to generate cell coordinates; then perform OCR recognition on each cell separately, and match the cell content with the table header through coordinate association;

[0068] Picture annotation recognition: for the text annotations in accident scene photos and injury photos, first detect the text area through YOLOv8 model, then input the cropped image into the OCR model for recognition, and verify the rationality of the annotation content by combining image semantics

[0069] Step 1.3, semantic error correction algorithm to correct OCR recognition errors: after OCR recognition, there are "similar character errors" and "semantic logic errors", which can be corrected by a semantic error correction algorithm combining rules and deep learning, the specific steps are as follows:

[0070] Rule-based error correction: based on the professionalism and standardization of legal texts, build 3 types of rule libraries: legal terminology dictionary: include 50,000+ legal professional terms (such as "joint and several liability", "disability grade", "litigation period"), identify and correct terminology errors (such as "joint and several liability" to "joint and several liability") through string matching (edit distance ≤1);

[0071] Format rule library: develop verification rules for formatted information; such as: date format: if the recognition result is "2024.5.10", automatically correct it to "2024-05-10" (consistent with the format of judicial documents); amount format: if the recognition result is "compensation 50,000", convert it to "compensation 50,000 yuan" through the number mapping rule ("ten thousand" corresponds to ); ID number: check the 18-digit code rule (the first 6 digits are administrative division code, the middle 8 digits are birth date, the last 3 digits are sequence code, and the last 1 digit is check code), if the verification fails, prompt manual review.

[0072] Context logic rules: based on the semantic logic of legal texts, correct contradictory errors, such as "plaintiff claims 12,000 yuan of medical expenses from defendant, evidence list shows 2,000 yuan of medical expenses", through the logic rule of "claimed amount ≥ evidence amount", mark "evidence list medical expenses 2,000 yuan" as suspicious information and prompt review;

[0073] Deep learning correction: For semantic errors that cannot be covered by rules, use the BERT-law pre-training model for context semantic correction. The model input is the windowed sequence of OCR recognized text, and the output is the corrected text sequence. The specific implementation is as follows:

[0074] Model training: Construct a legal text correction dataset, replace part of the characters in the correct text with error characters (e.g., "injury" → "injury"), form "error text-correct text" training pair; Fine-tune the LawBERT model to learn the context dependency relationship between "error character-correct character";

[0075] Correction reasoning: For OCR recognized text , sliding window takes subsequence , input LawBERT model, model output corrected probability distribution , select the character with the highest probability as the correction result, the formula is: , where is the character set, and this way realizes the correction of "context-dependent errors" (e.g., "plaintiff requires compensation" is corrected to "plaintiff requires compensation");

[0076] Step 1.4, key element information extraction by specialized cognitive model: The specialized cognitive model is constrained by the "legal domain knowledge graph", and combines BiLSTM-CRF model and rule reasoning to realize full-dimensional extraction of "explicit elements + implicit elements". The model architecture is divided into 3 layers:

[0077] Knowledge graph constraint layer: Construct a case element knowledge graph, define element types (e.g., "party information" includes "name, ID number, address", "compensation information" includes "compensation item, amount, payment method") and element relationship (e.g., "injury grade" is associated with "disability compensation"), providing semantic constraints for element extraction;

[0078] Element recognition layer: Use BiLSTM-CRF model to identify explicit element entities in text. The model input is the text sequence after semantic correction, and the output is the element entity label (use BIO tagging system: B-XXX indicates the start of the element, I-XXX indicates the middle of the element, and O indicates non-element);

[0079] BiLSTM layer: Convert text sequence to word vector (use legal domain word embedding model Word2Vec, dimension 300), input BiLSTM network, capture bidirectional semantic features of text, output hidden state (dimension 256), formula: ; ; where is the word vector at time t, is the forward LSTM output, is the backward LSTM output, is the concatenated bidirectional feature;

[0080] CRF layer: a CRF layer is added after the BiLSTM output layer to learn the transition probability between element labels (e.g., “B-Party Name” can only be followed by “I-Party Name” or “O”), and the loss function is: where is the true label sequence, is all possible label sequences, is the probability of the model predicting ; the CRF layer improves the continuity and accuracy of element entity recognition (e.g., avoids interruptions in “Party Name” labels);

[0081] Implicit element reasoning layer: based on knowledge graphs and rule-based reasoning, implicit element information is derived from explicit elements. For example:

[0082] Rule 1: if the explicit element contains “suffering from an eighth-grade injury”, then derive the implicit element “victim situation: eighth-grade injury”;

[0083] Rule 2: if the explicit element contains “medical expenses 12000 yuan, lost wages 8000 yuan”, then derive the implicit element “compensation items: medical expenses, lost wages” “total compensation amount: 20000 yuan”;

[0084] Mathematical rule representation: using first-order predicate logic, rule 1 can be represented as where is the case entity, which is matched and executed by the rule engine for reasoning.

[0085] The following takes the “Motor Vehicle Traffic Accident Liability Dispute” case as an example, and the element extraction process is as follows:

[0086] Explicit element recognition: input the semantic error-corrected complaint text into the BiLSTM-CRF model, and the model outputs the explicit element entity, such as:

[0087] “Plaintiff: Zhang San (ID number: 1101011990XXXX1234, address: Beijing Chaoyang District XX Road XX) ”→ extract “party information (plaintiff: Zhang San, ID number: 1101011990XXXX1234, address: Beijing Chaoyang District XX Road XX)”;

[0088] "Plaintiff claims medical expenses of 12000 yuan, lost wages of 8000 yuan" → extract "compensation information (medical expenses: 12000 yuan, lost wages: 8000 yuan)".

[0089] Implicit element reasoning: the rule engine loads the "motor vehicle traffic accident liability dispute" exclusive rule base, matches the explicit elements to perform reasoning:

[0090] Matching explicit elements "medical expenses of 12000 yuan, lost wages of 8000 yuan" → triggering rule 2 → deriving "total compensation amount: 20000 yuan";

[0091] If the explicit element contains "the plaintiff constitutes eight-level disability according to the disability identification report" → trigger rule 1 → derive "victim situation: eight-level disability" "disability compensation: local annual per capita disposable income × 20 years × 30%" (note: 30% is the compensation coefficient for eight-level disability).

[0092] Element standardization: standardize the extracted elements according to the preset format, such as:

[0093] Date element: "2024-05-10" → "2024-05-10";

[0094] Amount element: "12000 yuan" → "¥12,000.00";

[0095] Code conversion: "full responsibility" → "01".

[0096] Step 2, element induction: according to the case cause, obtain the elements and derivation rules that need to be induced from the configuration, initialize the elements according to the element attributes, and then use the extracted elements to derive other associated elements according to the derivation rules; the initialization process at least includes date formatting, amount formatting, and single / checkbox content uniform code conversion;

[0097] In step 2, after deriving other associated elements, it also includes a step of element verification: comparing the information of the same element in different case materials, if the information difference exceeds the preset threshold, it is marked as a suspicious element and prompts manual review

[0098] Step 3, rule reasoning: cyclically deduce the induced elements to construct the dispute information, trial findings information, court believes information, judgment result information, and legal provisions information of the judgment document, and persistently store the above information in the database;

[0099] When constructing various types of information of the judgment document, the corresponding legal provisions database needs to be referred to, to ensure that the legal provisions information is the current effective and matched with the case facts; during the persistent storage process, the element derivation process and rule reference log are recorded, which facilitates the subsequent tracing and checking of the judgment document generation process

[0100] Step 4, document generation: obtaining the dispute information, trial finding information, court opinion information, judgment result information and legal provision information from the database, filling into the judgment document, and generating the judgment document;

[0101] The judgment document has multiple format templates, and the template that is suitable for content filling is automatically selected according to the format requirement of the court of jurisdiction; after the judgment document is generated, a document checking step is further included: checking the document format, word expression and logical coherence, and if there is a problem, the modification is prompted until the judgment document that meets the judicial standard is generated.

[0102] The application discloses a judgment document content induction generation method based on image recognition, and relates to the technical field of legal informatization and artificial intelligence. The method innovatively integrates OCR image recognition, multi-modal information fusion, dynamic rule engine and large model reinforcement learning technology, realizes the automation and intelligentization of the whole process of judgment document generation. First, through the improved OCR recognition technology and the multi-modal information extraction module, the text, table, seal and handwritten annotation information of legal documents such as the complaint, the defense statement, the evidence material list and the court trial record are accurately extracted; second, the dynamic rule engine is introduced, combined with the case cause exclusive knowledge graph and the deep learning model, the extraction, formatting and correlation deduction of the key elements of the case are completed; third, the rule reasoning module optimized by reinforcement learning is used to build the logical correlation system of the dispute focus, the trial finding, the court opinion and the judgment result, and realize the intelligent matching and citation of the legal provisions; finally, based on the adaptive document template generation module, the reasoning result is automatically filled, and the judgment document that meets the format standard and judicial logic is generated. The application effectively improves the generation efficiency and accuracy of the judgment document, relieves the pressure of the judicial system "more cases and fewer people", promotes the digital transformation of judicial office, and at the same time, guarantees the legality and standardization of the document through the multi-dimensional quality checking mechanism.

[0103] Although the embodiments of the present application have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. A method for summarizing and generating the content of court judgments based on image recognition, characterized in that, Includes the following steps: Step 1, Element Extraction: OCR recognition technology is used to obtain the content of case materials such as complaints, answers, lists of evidence, and court transcripts. The content is then processed using a proprietary cognitive model to extract and store key element information of the case. The key element information includes at least party information, dispute information, and compensation information. Step 2, Element Summarization: Based on the case's cause of action, obtain the elements to be summarized and deduced from the configuration, along with the deduction rules. Initialize the elements according to their attributes, and then, based on the deduction rules, use the extracted elements to generate other related elements. The initialization process includes at least date formatting, amount formatting, and uniform conversion of single / multiple selection content to code. Step 3, Rule Reasoning: Iteratively deduce the summarized elements to construct the disputed information, the information ascertained during the trial, the information held by the court, the information of the judgment result, and the legal provisions in the judgment document, and persistently store the above information in the database; Step 4: Document Generation: Retrieve the disputed information, investigation findings, the court's opinion, the judgment result, and the legal provisions from the database, fill them into the judgment document, and generate the judgment document.

2. The method for summarizing and generating judgment document content based on image recognition according to claim 1, characterized in that: In step 1, the content is processed using a proprietary cognitive model to extract and store key elements of the case. The specific implementation steps are as follows: Step 1.1, Case Material Preprocessing and Classification Loading: The content data identified using OCR recognition technology will be preprocessed. The specific processing steps are as follows: Material Format Adaptation and Batch Import: Format Parsing: Calls the file parsing interface to separate the text layer, image layer, and annotation layer in the material; for pure scanned documents, they are automatically marked as "files to be processed by OCR"; for mixed format files, the annotation layer is extracted separately; Material Classification: Based on keyword matching and text features, materials are automatically classified into "indictment", "response", "evidence materials", and "court transcripts". Materials that do not match are prompted for manual classification assistance. Image preprocessing: The image data in the content data is preprocessed. First, tilt correction is performed, that is, the Hough transform is used to detect the text line direction in the image and the tilt angle is calculated. Rotation correction was then performed. Subsequently, noise removal was applied to the image, employing an adaptive median filtering algorithm to address speckle noise and paper texture noise in the scanned document. Specifically, the filtering window size was dynamically adjusted based on the noise density of the pixel neighborhood to avoid blurring of text details caused by traditional median filtering. Finally, contrast enhancement was performed on the image data using the Retinex algorithm to enhance the contrast between text and background, resolving text blurring issues caused by yellowing paper and uneven lighting. Step 1.2, Improved OCR Recognition Technology for Text Extraction: A CNN-Transformer fusion OCR model is used to achieve unified recognition of multimodal information including "printed text + handwritten annotations + tabular data + image annotations." The model architecture consists of three layers: Feature extraction layer: An improved ResNet-50 network is used as the backbone network, and deformable convolutions are added to the network to enhance the feature capture capability of tilted and curved text; Deformable convolution adds an offset to the standard convolution kernel. ,in , The offset is calculated using the following formula: (This formula represents the number of sampling points for the convolution kernel, allowing the sampling points to adapt to changes in text shape.) ,in The offset prediction value is the output of the convolutional layer. The scaling factor is used to determine the final output dimension of the feature extraction layer. The text feature map, where , The feature map height / width; Text sequence modeling layer: Introduces a Transformer encoder to convert the 2D feature map output from the feature extraction layer into a 1D text sequence; firstly, it injects text position information through positional encoding, the positional encoding formula being: ; ,in For sequence position; For dimension indexing; This refers to the input dimension of the Transformer, which is then used to capture the contextual relationships within the text through a self-attention mechanism. The formula for calculating the self-attention weights is as follows: ,in For querying the matrix, The key matrix, For value matrices, All are obtained by feature map mapping. For attention head dimension; Text decoding layer: Employs the CTC decoding algorithm to achieve unaligned text recognition. By introducing "whitespace symbol" ", with a length of The model output sequence is mapped to a length of The text sequence has the following loss function: ,in For the sample size, For the input image, For real text labels, For model prediction The probability, These are the model parameters; by minimizing this loss function, the model learns the mapping relationship between "image features and text sequences". Then, specific optimization strategies are adopted for different types of information to improve recognition accuracy; Step 1.3: Semantic error correction algorithm to correct OCR recognition errors: After OCR recognition, there are "similar character errors" and "semantic logic errors". These are corrected using a semantic error correction algorithm that combines rules and deep learning. The specific steps are as follows: Rule-based error correction: Based on the professionalism and standardization of legal texts, three types of rule bases are constructed: Legal terminology dictionary: containing 50,000+ legal professional terms, identifying and correcting terminology errors through string matching; Formatting rule base: Defines validation rules for formatted information; Context logic rules: Corrects inconsistencies and errors based on the semantic logic of legal texts; Deep learning error correction: For semantic errors that cannot be covered by rules, a BERT-law pre-trained model is used for contextual semantic error correction. The model input is a windowed sequence of OCR-recognized text, and the output is the corrected text sequence. Specifically, a legal text error correction dataset is constructed, and some characters in the correct text are replaced with incorrect characters to form training pairs of "incorrect text-correct text". The LawBERT model is fine-tuned to enable it to learn the contextual dependencies of "incorrect character-correct character". OCR recognition of text Sliding window to extract subsequence Input the LawBERT model, and the model outputs... Corrected probability distribution Choose the character with the highest probability. As a result of the correction, the formula is: ,in For character sets; Step 1.4: Extraction of Key Element Information Using a Proprietary Cognitive Model: The proprietary cognitive model uses a "knowledge graph in the legal domain" as a constraint, integrates a BiLSTM-CRF model with rule-based reasoning, and achieves full-dimensional extraction of "explicit elements + implicit elements." The model architecture consists of three layers: Knowledge Graph Constraint Layer: Constructs a knowledge graph of case elements, defines element types and relationships between elements, and provides semantic constraints for element extraction; Feature recognition layer: The BiLSTM-CRF model is used to identify explicit feature entities in the text. The input of the model is the semantically corrected text sequence, and the output is the feature entity label. BiLSTM layer: Converts the text sequence into word vectors, inputs them into the BiLSTM network, captures the bidirectional semantic features of the text, and outputs the hidden state. The formula is: ; ; ,in for Word vectors at time points, For positive LSTM output, For inverted LSTM output, The spliced ​​bidirectional features; CRF layer: A CRF layer is added after the BiLSTM output layer to learn the transition probabilities between feature labels. Its loss function is: ,in The actual label sequence, For all possible label sequences, For model prediction The probability of element entity recognition; improving the continuity and accuracy of element entity recognition through CRF layer; Implicit Element Reasoning Layer: Based on knowledge graphs and rule-based reasoning, implicit element information is derived from explicit elements.

3. The method for summarizing and generating judicial documents based on image recognition according to claim 1, characterized in that: The specific content of image preprocessing in step 1.1 is as follows: Tilt correction: The Hough transform maps lines in image space to parameter space. ,in The perpendicular distance from the origin to the line is... For a straight line and The included angle of the axis is given by the formula: ; through the peak values ​​corresponding to the statistical parameter space The value determines the text tilt angle, and then a rotation matrix is ​​used. For image pixels Perform coordinate transformation to achieve tilt correction; Noise Removal: An adaptive median filtering algorithm is used to process speckle noise and paper texture noise in the scanned document; the specific definition of the filtering window is as follows. Window size range The processing steps are as follows: calculate the minimum value of the pixels within the window. Maximum value Median With center pixel value ;like Enter the second stage of judgment: if Output Otherwise output ;like or Expand the window to Repeat the above steps until the window reaches its maximum size. Final output ; Contrast Enhancement: The Retinex algorithm is used to enhance the contrast between text and background, specifically by adjusting the image... Decomposed into "reflection component" "and light component" ",Right now The illumination components are separated by logarithmic transformation and Gaussian filtering, as shown in the formula: ,in It is a Gaussian filter kernel. This represents the convolution operation.

4. The method for summarizing and generating judicial documents based on image recognition according to claim 1, characterized in that: In step 1.2, the specific optimization strategies adopted to improve recognition accuracy for different types of information are as follows: Handwritten annotation recognition: A handwritten text dataset was constructed, and the CNN-Transformer model was fine-tuned to optimize the recognition of cursive and variant characters; Stroke feature constraints help correct recognition errors by calculating the curvature and directional changes of the handwritten trajectory; Table data recognition: A two-step method of "table structure detection + cell content recognition" is adopted. First, the row and column boundaries of the table are detected by the Mask R-CNN model to generate cell coordinates; then, OCR recognition is performed on each cell individually, and the cell content is matched with the table header by coordinate association. Image annotation recognition: For text annotations in accident scene photos and injury photos, the text region is first detected by the YOLOv8 model, then cropped and input into the OCR model for recognition. At the same time, the image semantics are combined to help verify the rationality of the annotation content.

5. The method for summarizing and generating judicial documents based on image recognition according to claim 1, characterized in that: In step 2, after deriving and generating other related elements, there is also an element verification step: comparing the information of the same element in different case materials, if the information difference exceeds a preset threshold, it is marked as a suspicious element and prompts for manual review.

6. The method for summarizing and generating judicial documents based on image recognition according to claim 1, characterized in that: In step 3, when constructing various types of information for the judgment document, it is necessary to refer to the corresponding legal provisions database to ensure that the legal provisions are currently valid and match the facts of the case. During the persistent storage process, the element derivation process and rule reference log will be recorded to facilitate subsequent traceability and verification of the judgment document generation process.

7. The method for summarizing and generating judicial documents based on image recognition according to claim 1, characterized in that: In step 4, the judgment document has multiple format templates, and the appropriate template is automatically selected and the content is filled in according to the format requirements of the court with jurisdiction over the case.

8. The method for summarizing and generating judicial documents based on image recognition according to claim 1, characterized in that: Step 4, after generating the judgment document, also includes a document verification step: checking the document format, wording, and logical coherence. If there are any problems, it will prompt for modification until a judgment document that conforms to judicial norms is generated.

Citation Information

Cited By

  • UI anomaly detection method based on priori guidance semantic segmentation and related device

    CN122240512A

  • UI anomaly detection method based on priori guided semantic segmentation and related device

    CN122240512B