Traffic accident investigation document photo analysis method based on large model
By adopting a multi-layer collaborative processing architecture based on a large model, the problem of high-precision recognition and analysis of traffic accident investigation document photos was solved, realizing intelligent processing from document photos to case summaries, and improving the efficiency and accuracy of document information extraction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies are insufficient to meet the high-precision identification and analysis requirements of traffic accident investigation documents and photographs, especially in terms of document type identification, information extraction, and document structure preservation, which present complexities and challenges.
A multi-layer collaborative processing architecture based on a large model is adopted, including rapid classification and filtering, fine-grained document analysis, content understanding and document concatenation. Models such as MobileNet-V3, ResNet-152, RoBERTa-large, and GPT-3.5-turbo are used for feature extraction and semantic analysis to generate structured text information and perform document concatenation, ultimately generating a case summary.
It achieves a balance between high efficiency and high accuracy by efficiently filtering document photos from traffic accident investigation photos, accurately extracting information, and generating comprehensive accident analysis results.
Smart Images

Figure CN121789150A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and document processing, and in particular to a method for analyzing traffic accident investigation document photos based on a large model. Background Technology
[0002] Traffic accident investigations generate a large volume of highly complex and diverse photographs and documents. These materials primarily include photographs of case files, such as identification documents, driver's licenses, interrogation records, medical certificates, and insurance policies; document information from on-site photographs, such as license plate numbers, vehicle identification numbers, and transport documents; and multi-page continuous documents, such as accident reports and mediation agreements. This diverse range of documents presents significant technical challenges to automated processing, including the high complexity of document type identification, the extremely high accuracy requirements for information extraction, the significant technical challenges in maintaining document structure, and the high complexity of splicing multi-page documents.
[0003] A search of existing technical literature revealed a patent application (application number 202411051004.2) entitled "Multimodal Electronic Data Forensic Analysis Method, System, Medium, and Device." This patent acquires electronic data, including audio data, text data, and image data. It classifies the audio data using a language classifier to determine the language category; obtains the corresponding text content based on an automatic speech recognition model; obtains the text content of the image data using a visual language model; and inputs the text content into a multimodal large language model to obtain the analysis results of the electronic data. However, this patent has limitations in meeting the high-precision recognition and analysis requirements of document images in specific scenarios. Summary of the Invention
[0004] Therefore, it is necessary to provide a method for parsing traffic accident investigation document photos based on a large model to address the aforementioned technical problems, thereby achieving specialized and high-precision recognition and parsing of traffic accident investigation document photos.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a method for parsing photographs from traffic accident investigation documents based on a large model, the method comprising: S1: Based on the first multimodal large model, the traffic accident photo set is quickly classified and filtered to obtain document photos; S2: Perform fine document analysis on the document photo based on the second multimodal large model to obtain structured text information; S3: Based on the first text big model, perform content understanding on the structured text information to obtain an integrity assessment report; S4: Based on the second text big model, the document photo is stitched together according to the structured text information to obtain the stitched document photo; S5: Based on the third text big model, generate a case summary according to the document photo, the structured text information, the integrity assessment report, the spliced document photo, and the preset traffic accident case summary template.
[0006] Preferably, the first multimodal large model is composed of a MobileNet-V3 visual encoder and a BERT-Base text encoder. Based on the first multimodal large model, the traffic accident photo set is quickly classified and filtered, including: The traffic accident photo set is subjected to standardized preprocessing to obtain a standard traffic accident photo set; The MobileNet-V3 visual encoder was used to extract features from the standard traffic accident photo set to obtain visual feature vectors. The BERT-Base text encoder was used to extract features from the standard traffic accident photo set to obtain text feature vectors. An attention mechanism is used to fuse the visual feature vector and the text feature vector to obtain a first fused feature vector. A multilayer perceptron is used to classify the first fused feature vector to obtain classified document photos and corresponding confidence scores. The Softmax function is used to calculate the confidence score based on the confidence score corresponding to the classified document photos. Document photos are selected based on the confidence score being greater than a preset threshold. The document photos include ID document photos, document photos, and table document photos.
[0007] Preferably, the second multimodal large model is constructed using a ResNet-152 visual encoder and a RoBERTa-large text encoder. Based on the second multimodal large model, fine-grained document analysis is performed on the document photograph, including: The document image was analyzed for depth features using a ResNet-152 visual encoder to obtain visual depth features. The document image was subjected to deep feature analysis using the RoBERTa-large text encoder to obtain text depth features; An attention mechanism is used to fuse the visual depth features and the text depth features to obtain a second fused feature. The second fused feature is classified using a multilayer perceptron to obtain sub-document photos, which include ID card type sub-document photos, document type sub-document photos, and table type sub-document photos; Based on the sub-document photos, select the corresponding document photo processing strategy to extract information and obtain document information, which includes certificate document information, document document information and table document information; The document information is processed using a structured analysis algorithm of the corresponding type to obtain structured text information, which includes document-type structured text information, formal document-type structured text information, and table-type structured text information.
[0008] Preferably, information extraction is performed by selecting a document photo processing strategy of the corresponding type based on the sub-document photo, including: The document information includes: document border, document field positions, document information, and document photo. An improved Hough transform algorithm is used to detect rectangular borders in the document photo, and the Canny edge detection algorithm is used to extract the edges of the document photo to obtain the document border. A template matching algorithm is used to locate key fields in the document photo to obtain the document field positions. A CRNN algorithm is used to perform text recognition in the document photo to obtain the document information. An MTCNN algorithm is used to perform face detection in the document photo to obtain face bounding box information, and the document photo is segmented based on the face bounding box information. Document information includes: document structure, document table data, document seal, and document signature; the layout analysis algorithm is used to perform structural analysis on the document sub-document photos to obtain the document structure; the table recognition algorithm is used to extract the table data from the document sub-document photos to obtain the document table data; the YOLOv5 object detection algorithm is used to perform specific object detection on the document sub-document photos to obtain the document seal and document signature. The information for table-type documents includes: table lines, cell boundaries, and cell content; Hough transform is used to detect lines in the image of the table-type sub-document to obtain the table lines; cell segmentation algorithm is used to locate cell boundaries in the image of the table-type sub-document to obtain the cell boundaries; and DB text detection algorithm is used to identify cell content in the image of the table-type sub-document to obtain the cell content.
[0009] Preferably, the document information is processed using a structured analysis algorithm of the corresponding type, including: A pattern matching algorithm is used to match a preset field mapping table with the document information to obtain the field type; a coordinate transformation algorithm is used to convert the position of the document field into the absolute coordinates of the document border in a reference coordinate system to obtain the position information; an adaptive weighted fusion algorithm is used to calculate the confidence level based on the document field position, document information, and document photo; a multi-dimensional data fusion method is used to fuse the field type, position information, and confidence level to obtain document-type structured text information. The document structure is reconstructed using natural language processing technology to obtain the reconstructed document structure; named entity recognition technology combined with a deep learning model is used to identify the document's table data to obtain entity information; computer vision algorithms are used to separate the document's seal and signature to obtain image elements; layout analysis technology is used to extract formatting information from the document, including font type, font size, alignment, and line spacing; and a multi-dimensional data fusion method is used to fuse the reconstructed document structure, entity information, image elements, and formatting information to obtain the document's structured text information. Image processing algorithms are used to detect and reconstruct the table lines to obtain the table structure; a machine learning classifier is used to classify the cell content to obtain classified data, and a data format conversion algorithm is used to standardize the data to obtain standardized classified data; cells are merged according to their boundaries to obtain the merging relationship; a multi-dimensional data fusion method is used to fuse the table structure, standardized classified data, and merging relationship to obtain table-like structured text information.
[0010] Preferably, after obtaining the structured text information, the method further includes correcting the structured text information to obtain a quality verification report and the corrected structured text information.
[0011] Preferably, the first large text model adopts the GPT-3.5-turbo model, and content understanding of the structured text information is performed based on the first large text model, including: The GPT-3.5-turbo model is used to perform deep semantic understanding on the structured text information to obtain labeled structured text information; The improved TF-IDF algorithm combined with text embedding technology is used to extract keywords from the labeled structured text information to obtain a list of key information. The BERT model is used to calculate the text semantic similarity of the key information list to obtain a contradiction detection report; The integrity of the structured text information is verified based on the preset standard information checklist and the contradiction detection report to obtain an integrity assessment report.
[0012] Preferably, the second large-scale text model adopts the Sentence-BERT model, and based on the second large-scale text model, document stitching is performed on the document image according to the structured text information, including: The Sentence-BERT model was used to calculate the semantic similarity between the document photo and the structured text information to obtain the similarity analysis results. Based on the similarity analysis results, the document photos are logically ordered to obtain logically sorted document photos; Based on a preset standard document template, the missing content is inferred from the logically sorted document photos to obtain the missing content; The pHash image fingerprint algorithm is used to extract features from the logically sorted document photos to generate image fingerprints. The text hash fingerprint algorithm is used to process the structured text information to obtain text fingerprints. By comparing the image fingerprints and the text fingerprints, duplicate pages between the logically sorted document photos and the structured text information are detected to obtain duplicate pages. After filling in missing content and removing duplicate pages from the logically sorted document photos, the documents are then stitched together to obtain the stitched document photos.
[0013] Preferably, the third text model adopts the GPT-4 model.
[0014] Preferably, after step S5, the method further includes using RAG technology to retrieve the case summary based on the questions input by the user, and obtaining the retrieval results.
[0015] Compared with the prior art, the beneficial effects of the present invention are: This invention provides a method for parsing traffic accident investigation document photos based on a large model. It progresses from coarse-grained classification to fine-grained extraction, gradually delving into the document content. In the rapid classification and filtering stage, it accurately selects document photos from an unordered photo set. In the refined document analysis stage, it performs targeted processing for different document types. In the content understanding stage, it achieves accurate and comprehensive semantic analysis and content association. In the document stitching stage, it intelligently rearranges and restores content from multiple scattered pages. Finally, it generates comprehensive accident analysis results based on multi-source information. This hierarchical collaborative processing architecture is specifically designed for traffic accident investigation photo document processing, achieving an optimal balance between efficiency and accuracy through multi-layered collaboration. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of a traffic accident investigation document photo parsing method based on a large model in one embodiment; Figure 2 This is a schematic diagram of the complete architecture of a traffic accident investigation document photo parsing method based on a large model in one embodiment; Figure 3 This is a schematic diagram of the multimodal large model architecture of a traffic accident investigation document photo parsing method based on a large model in one embodiment; Figure 4 This is a schematic diagram of a rapid classification and filtering process for a traffic accident investigation document photo parsing method based on a large model in one embodiment; Figure 5 This is a schematic diagram comparing the technical performance of a traffic accident investigation document photo parsing method based on a large model in one embodiment. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0018] Example 1 like Figure 1 As shown in the figure, this embodiment proposes a method for parsing photos in traffic accident investigation documents based on a large model. The method includes: S1: Based on the first multimodal large model, the traffic accident photo set is quickly classified and filtered to obtain document photos; The specific implementation of this step is as follows: the rapid classification and screening layer plays a crucial role in the rapid classification and initial screening of documents. The design goal is to use a small-capacity, low-precision multimodal large model to rapidly process a large number of input traffic accident investigation photos, efficiently identifying and filtering photos containing documents and certificates. This layer employs a lightweight multimodal large model to maximize processing speed and throughput while ensuring basic classification accuracy. The classification process includes multiple sub-modules such as image preprocessing, rapid feature extraction, document type recognition, and relevance filtering, ensuring rapid and accurate filtering of the documents and certificates to be processed from a large number of photos. This layer also features a parallel processing mechanism, capable of processing multiple images simultaneously, providing high-quality input data for subsequent fine-tuning layers.
[0019] S2: Perform fine document analysis on the document photo based on the second multimodal large model to obtain structured text information; The specific implementation of this step is as follows: The refined document analysis layer, based on a large-scale multimodal model, achieves in-depth analysis and precise processing of documents and certificates. It can not only identify basic text and image elements in documents, but also deeply understand high-level information such as document layout, content format, and semantic relationships. This layer fully utilizes the advanced technical capabilities of the large-scale multimodal model, including high-precision OCR recognition, layout structure analysis, information type recognition, image element extraction, and structured processing, to achieve accurate recognition, high-precision information extraction, and standardized format output for various documents and certificates. This layer employs specialized processing strategies for different document types, including specialized algorithms for certificate processing, document processing, and table processing, ensuring optimal processing results for each document type.
[0020] S3: Based on the first text big model, perform content understanding on the structured text information to obtain an integrity assessment report; The specific implementation of this step is as follows: The content understanding layer is responsible for performing in-depth semantic analysis and understanding of the document content based on the extraction results of the fine-grained processing layer. This layer adopts a highly efficient large text model with powerful natural language understanding and analysis capabilities, enabling it to deeply understand the meaning, relevance, importance, and other multi-dimensional information of the document content. This layer employs a multi-level analysis strategy, including basic semantic understanding, content association analysis, important information identification, and contradiction detection, to construct a complete document content understanding result. Simultaneously, this layer also designs an intelligent context analysis mechanism, which can comprehensively analyze the content of multiple documents, providing an accurate content foundation for subsequent document splicing and summary generation.
[0021] S4: Based on the second text big model, the document photo is stitched together according to the structured text information to obtain the stitched document photo; The specific implementation of this step is as follows: The document stitching layer is responsible for intelligently stitching and restoring content from multi-page documents based on the analysis results of the content understanding layer. Leveraging the powerful language understanding capabilities of a large text model, this layer can accurately determine the logical relationships between document pages, identify missing and duplicate pages, and perform intelligent document reconstruction. This layer incorporates various innovative stitching algorithms and technologies, including content similarity analysis, logical relationship reasoning, page order judgment, and missing content inference, enabling it to handle various complex multi-page document stitching scenarios. This layer also implements quality monitoring and consistency checking mechanisms to ensure the integrity and accuracy of the stitching results.
[0022] S5: Based on the third text big model, generate a case summary according to the document photo, the structured text information, the integrity assessment report, the spliced document photo, and the preset traffic accident case summary template.
[0023] The specific implementation of this step is as follows. The core technical effect of this embodiment is to automatically generate a case summary document that meets the standardized requirements of the traffic police department through intelligent analysis and comprehensive processing of traffic accident scene photos. It transforms scattered, multi-format traffic accident investigation photos (including documents, forms, etc.) into a structured case analysis report, providing a complete intelligent solution for traffic accident handling. The final output case summary includes standardized content such as basic accident information, party information, vehicle information, accident details, loss situation, and liability analysis. Based on the processing results of the previous layers and the powerful content generation and question-answering capabilities of the large text model, a comprehensive case analysis is generated.
[0024] Example 2 This embodiment further explains the method for parsing traffic accident investigation documents and photos based on a large model proposed in Embodiment 1.
[0025] The first multimodal large model is composed of a MobileNet-V3 visual encoder and a BERT-Base text encoder. Based on the first multimodal large model, a fast classification and filtering of traffic accident photo sets is performed, including: The traffic accident photo set is subjected to standardized preprocessing to obtain a standard traffic accident photo set; The MobileNet-V3 visual encoder was used to extract features from the standard traffic accident photo set to obtain visual feature vectors. The BERT-Base text encoder was used to extract features from the standard traffic accident photo set to obtain text feature vectors. An attention mechanism is used to fuse the visual feature vector and the text feature vector to obtain a first fused feature vector. A multilayer perceptron is used to classify the first fused feature vector to obtain classified document photos and corresponding confidence scores. The Softmax function is used to calculate the confidence score based on the confidence score corresponding to the classified document photos. Document photos are selected based on the confidence score being greater than a preset threshold. The document photos include ID document photos, document photos, and table document photos.
[0026] The specific implementation of this step is as follows, taking a traffic accident photo set consisting of 10 document photos taken at the traffic accident scene (2 ID cards, 1 driver's license, 1 vehicle registration certificate, 3 pages of accident determination letter, 2 pages of medical diagnosis certificate, and 1 insurance policy) as an example. Figure 4 As shown, image normalization preprocessing uses bilinear interpolation to adjust all photos in the traffic accident photo set to 224×224 pixels, and adaptive histogram equalization is applied to enhance image quality, resulting in a standard traffic accident photo set, as shown. Figure 3 As shown, MobileNet-V3 was used to extract 188-dimensional visual feature vectors from a standard traffic accident photo set, and BERT-Base was used to extract 768-dimensional text feature vectors. An attention mechanism was used to fuse the visual and text feature vectors, and then a multilayer perceptron was used for fusion classification. The input layer was 956-dimensional (188+768), hidden layer 1 had 512 neurons (ReLU activation), hidden layer 2 had 128 neurons (ReLU activation), and the output layer had 4 neurons (documents, forms, and others). Each input image would receive a confidence score for one of four categories from the multilayer perceptron. The loss function used to train the multilayer perceptron was cross-entropy loss, the optimizer was Adam, and the learning rate was 0.001. The training strategy involved training with 10,000 labeled traffic accident document images, with a validation set ratio of 20%. The Softmax function is used to ensure that the sum of the four output values is 1. Each value represents the confidence level of the corresponding type. The confidence level threshold is set to 0.7. Only the classification with the highest confidence level exceeding 0.7 is accepted. For example, ID card 1: document type 0.95, document type 0.03, form type 0.01, other type 0.01, therefore classified as document type; Accident report 1: document type 0.12, document type 0.83, form type 0.03, other type 0.02, therefore classified as document type; Finally, there are 4 document type reports (average confidence level 0.92), 6 document type reports (average confidence level 0.89), 0 form type reports, and 0 other type reports.
[0027] The second multimodal large model is constructed using a ResNet-152 visual encoder and a RoBERTa-large text encoder. Based on this second multimodal large model, fine-grained document analysis is performed on the document photograph, including: The document image was analyzed for depth features using a ResNet-152 visual encoder to obtain visual depth features. The document image was subjected to deep feature analysis using the RoBERTa-large text encoder to obtain text depth features; An attention mechanism is used to fuse the visual depth features and the text depth features to obtain a second fused feature. The second fused feature is classified using a multilayer perceptron to obtain sub-document photos, which include ID card type sub-document photos, document type sub-document photos, and table type sub-document photos; Based on the sub-document photos, select the corresponding document photo processing strategy to extract information and obtain document information, which includes certificate document information, document document information and table document information; The document information is processed using a structured analysis algorithm of the corresponding type to obtain structured text information, which includes document-type structured text information, formal document-type structured text information, and table-type structured text information.
[0028] The specific implementation of this step is as follows: the filtered document photos (including image data and confidence scores) are subjected to deep feature analysis using a ResNet-152 visual encoder and a RoBERTa-large text encoder. The deep visual feature analysis specifically includes: layout structure features, analyzing visual features such as document layout, paragraph structure, table arrangement, and image position; text region features, identifying visual attributes such as the position, size, font, and orientation of text blocks; image element features, extracting feature representations of image elements such as photos, seals, signatures, and barcodes; and document quality features, evaluating image clarity, lighting conditions, tilt angle, and occlusion. Deep text feature analysis specifically includes: semantic content features, using the BERT model to extract deep semantic representations of document content; entity relationship features, identifying entities such as names of people, places, times, license plate numbers, and their interrelationships; document type features, determining the specific type and attributes of a document based on text content features; and language style features, analyzing the language style, terminology, and formatting specifications of a document. In the deep visual text feature analysis process, auxiliary algorithms were also employed, namely Faster R-CNN for document region detection and PaddleOCR 4.0 for text recognition. After mapping visual features and text features to the same semantic space, a cross-modal attention mechanism was used to integrate different types of features to obtain a second fusion feature. A multilayer perceptron was used to classify the second fusion feature to obtain sub-document photos, which include document type sub-document photos, formal document type sub-document photos, and table type sub-document photos. Specifically, from the document type, specific document types such as ID cards, driver's licenses, and vehicle registration certificates were accurately identified; from the formal document type, specific document types such as interrogation records, medical certificates, and insurance policies were accurately identified; and from the table type, the row and column structure and cell position of tables were accurately identified. Based on the sub-document photos, select the corresponding document photo processing strategy to extract information, obtain document information, and then use the corresponding type of structured analysis algorithm to process the document information into structured text information.
[0029] Based on the sub-document photos, select the corresponding document photo processing strategy for information extraction, including: The document information includes: document border, document field positions, document information, and document photo. An improved Hough transform algorithm is used to detect rectangular borders in the document photo, and the Canny edge detection algorithm is used to extract the edges of the document photo to obtain the document border. A template matching algorithm is used to locate key fields in the document photo to obtain the document field positions. A CRNN algorithm is used to perform text recognition in the document photo to obtain the document information. An MTCNN algorithm is used to perform face detection in the document photo to obtain face bounding box information, and the document photo is segmented based on the face bounding box information. Document information includes: document structure, document table data, document seal, and document signature; the layout analysis algorithm is used to perform structural analysis on the document sub-document photos to obtain the document structure; the table recognition algorithm is used to extract the table data from the document sub-document photos to obtain the document table data; the YOLOv5 object detection algorithm is used to perform specific object detection on the document sub-document photos to obtain the document seal and document signature. The information for table-type documents includes: table lines, cell boundaries, and cell content; Hough transform is used to detect lines in the image of the table-type sub-document to obtain the table lines; cell segmentation algorithm is used to locate cell boundaries in the image of the table-type sub-document to obtain the cell boundaries; and DB text detection algorithm is used to identify cell content in the image of the table-type sub-document to obtain the cell content.
[0030] The specific implementation of this step is as follows: For document types, the image can be preprocessed first, adjusted to a fixed resolution of 1920×1080 pixels, and contrast enhanced by the Adaptive Histogram Equalization (AHE) algorithm. Then, a nonlocal mean denoising algorithm (parameters h=10, search_window=21×21) is applied to suppress noise. The first step after preprocessing is document border detection. An improved Hough Transform algorithm is used to detect the rectangular borders of the document, and the Canny edge detection algorithm is used to extract image edges. A threshold of 50-150 is set. Contour detection and quadrilateral fitting are used to accurately locate the coordinates of the four corner points of the document. The detected borders are quality-assessed to ensure completeness and accuracy. The second step is key field localization. A template matching algorithm combined with deep learning is used to locate the key fields of ID cards, driver's licenses, and vehicle registration certificates. The U-Net semantic segmentation model is used to accurately locate the fields such as name, ID number, and address, outputting the precise coordinates and confidence score of each field. The third step is text information extraction. The CRNN algorithm is used for text recognition to extract information such as name, ID number, and address. The PaddleOCR engine is used for high-precision OCR recognition. Post-processing and error correction are performed on the recognition results to improve accuracy. The fourth step is document photo extraction. Face detection uses the MTCNN algorithm, achieving an accuracy of over 99% in document photo extraction. The photo area in the document is automatically located and cropped, outputting a standardized document photo.
[0031] For document-type documents, optional preprocessing steps include layout analysis, using a Hough transform-based line detection algorithm for document skew correction, an OTSU threshold segmentation algorithm to remove background interference, and a text block recognition algorithm based on connected component analysis for text region segmentation. The first step after preprocessing is layout structure analysis, using layout analysis algorithms to maintain the document structure of inquiry records and diagnostic certificates, and using deep learning models to identify structural elements such as paragraphs, headings, and tables. The second step is content extraction, using table recognition algorithms to extract table data from insurance documents, and using the YOLOv5 object detection algorithm for seal and signature recognition. The third step is format preservation, maintaining the original document's paragraph structure, hierarchical relationships, and formatting features.
[0032] For tables, optional preprocessing can be performed first. Morphological operations (kernel_size=3×3) can be applied to enhance table lines, Gaussian blur (σ=0.8) can be used to reduce image noise, and an adaptive thresholding algorithm (block_size=11, C=2) can be used for binarization. After preprocessing, the first step is to use Hough transform to detect table lines, the second step is to use a cell segmentation algorithm to determine cell boundaries, and the third step is to use the DB text detection algorithm for cell content recognition.
[0033] The document information is processed using a structured analysis algorithm of the corresponding type, including: A pattern matching algorithm is used to match a preset field mapping table with the document information to obtain the field type; a coordinate transformation algorithm is used to convert the position of the document field into the absolute coordinates of the document border in a reference coordinate system to obtain the position information; an adaptive weighted fusion algorithm is used to calculate the confidence level based on the document field position, document information, and document photo; a multi-dimensional data fusion method is used to fuse the field type, position information, and confidence level to obtain document-type structured text information. The document structure is reconstructed using natural language processing technology to obtain the reconstructed document structure; named entity recognition technology combined with a deep learning model is used to identify the document's table data to obtain entity information; computer vision algorithms are used to separate the document's seal and signature to obtain image elements; layout analysis technology is used to extract formatting information from the document, including font type, font size, alignment, and line spacing; and a multi-dimensional data fusion method is used to fuse the reconstructed document structure, entity information, image elements, and formatting information to obtain the document's structured text information. Image processing algorithms are used to detect and reconstruct the table lines to obtain the table structure; a machine learning classifier is used to classify the cell content to obtain classified data, and a data format conversion algorithm is used to standardize the data to obtain standardized classified data; cells are merged according to their boundaries to obtain the merging relationship; a multi-dimensional data fusion method is used to fuse the table structure, standardized classified data, and merging relationship to obtain table-like structured text information.
[0034] The specific implementation of this step is as follows: For document types, the first step is to map the OCR recognition results to standard fields using a predefined field mapping table, and to match the text content recognized by OCR with the preset field templates using a pattern matching algorithm, determining the field type based on keyword weights and positional features; the second step is to perform relative positioning calculations between the key field coordinates and the document border coordinates, and to convert the relative coordinates of each field to absolute coordinates in the document border reference coordinate system using a coordinate system transformation algorithm; the third step is to calculate the comprehensive confidence score based on OCR recognition confidence, field position confidence, and format conformity, and to use a multi-factor weighted fusion algorithm to comprehensively consider factors such as OCR recognition confidence, field position accuracy, and document photo format conformity; the fourth step is to use a recursive descent algorithm to construct a hierarchical JSON structure, and to use a hierarchical data organization method to construct a JSON object structure according to three dimensions: field type, position information, and confidence score, ultimately outputting a standardized document JSON object containing field information, coordinates, and confidence score.
[0035] For document-type documents, the first step is to analyze the document's hierarchical structure using natural language processing (NLP) techniques. NLP is used to perform semantic analysis of the text, reconstructing the document's hierarchical structure through title recognition, paragraph segmentation, and logical relationship analysis. The second step is to use the BERT-CRF model to identify key entities and mark their locations. Named entity recognition technology, combined with a deep learning model, is used to identify key entities such as names, places, times, and license plate numbers, and to record their precise locations within the document. The third step is to separate non-text elements such as signatures and seals into independent fields. Computer vision algorithms are used to separate non-text elements such as signatures, seals, and barcodes, calculating their spatial relationships and visual features. The fourth step is to extract formatting attributes such as font, size, and alignment. Page layout analysis techniques are used to extract formatting information such as font type, font size, alignment, and line spacing. Finally, a multi-dimensional data fusion method is used to integrate the text structure, entity information, image elements, and formatting attributes into a unified JSON format.
[0036] For tables, the first step is to reconstruct the table frame based on line detection and cell boundaries. Image processing algorithms are used to detect table border lines (Hough transform threshold: vote count ≥ 100, line angle tolerance: ±2). The first step involves reconstructing the table framework structure based on line intersection recognition and cell segmentation techniques (horizontal line detection: detecting horizontal table lines and calculating the y-coordinate array Y_lines; vertical line detection: detecting vertical table lines and calculating the x-coordinate array X_lines; grid construction: constructing a table grid matrix based on Y_lines and X_lines; cell verification: verifying cell integrity using connected component analysis). The second step involves using a machine learning classifier to identify the data type of each cell, distinguishing between numeric, text, date, and percentage types, and converting the identified text into a standard data format using a data format conversion algorithm. The third step involves analyzing the merging relationships and hierarchical structure between cells, determining the merging relationships between cells through spatial positional relationship analysis, and establishing a hierarchical structure model of the table. Finally, a JSON object is constructed based on a multi-dimensional data model, including table structure (two-dimensional array format, containing the number of rows and columns, cell coordinates, and merging information), cell content (structured data object, containing information such as row and column index: the cell's position in the table (row, col), extracted content: the recognized text, numbers, or date information, data type: text, numerical, or date classification, confidence score: recognition reliability score in the range of 0-1, boundary coordinates: the pixel coordinate position of the cell in the image, and merging status: a boolean value indicating whether it is a merged cell), table metadata (structured data of table title, unit information, and summary rows), and quality indicators (structure recognition accuracy, content extraction confidence, and integrity score).
[0037] After obtaining the structured text information, the process also includes correcting the structured text information to obtain a quality verification report and the corrected structured text information.
[0038] The specific implementation of this step is as follows: Based on the previously generated structured JSON data, a quality verification processing algorithm is used. The first step is to check the format of the verification fields. A pattern matching algorithm and data type verification rules are used to check the correctness of the extracted fields, including date formats, ID card number formats, and phone number formats. The second step is to perform cross-document information consistency verification and conflict detection. Semantic similarity calculation technology is used to verify the consistency of the same entity information across documents and identify potential information conflicts. The third step is to perform context-based error detection and automatic correction. Based on context analysis and knowledge base matching technology, errors found are automatically corrected, such as ID card verification code correction and date format standardization. The fourth step is to recalculate the field credibility by considering multiple factors. A multi-factor evaluation model is used to comprehensively consider the format verification results, consistency check results, error correction history, and other factors to recalculate the confidence score of each field. The fifth step is to calculate the comprehensive quality score based on completeness, accuracy, and consistency. A comprehensive quality assessment algorithm is used to calculate the overall data quality score based on completeness indicators, accuracy indicators, and consistency indicators. Finally, a templated report generation technology is used to automatically generate a quality verification report that includes verification results, error correction records, quality scores, and optimization suggestions.
[0039] Example 3 This embodiment further supplements the above embodiment's method for parsing traffic accident investigation documents and photos based on a large model.
[0040] The first large-scale text model employs the GPT-3.5-turbo model. Based on this model, content understanding is performed on the structured text information, including: The GPT-3.5-turbo model is used to perform deep semantic understanding on the structured text information to obtain labeled structured text information; The improved TF-IDF algorithm combined with text embedding technology is used to extract keywords from the labeled structured text information to obtain a list of key information. The BERT model is used to calculate the text semantic similarity of the key information list to obtain a contradiction detection report; The integrity of the structured text information is verified based on the preset standard information checklist and the contradiction detection report to obtain an integrity assessment report.
[0041] The specific implementation of this step is as follows: After using the BERT-CRF model to identify key entities such as names, place names, time, license plate numbers, and vehicle identification codes, the GPT-3.5-turbo large language model is used to analyze the semantic relationships and dependencies between entities, construct a knowledge graph, and then identify and understand the temporal sequence and causal relationships of accident-related events. Finally, labeled structured text information is output, with each entity labeled with the source document and location information. Entities include time entities (accident time, processing time, insurance period), location entities (accident location, processing location, hospital address), people entities (parties involved, witnesses, law enforcement officers, medical personnel), vehicle entities (license plate number, vehicle type, VIN code, engine number), and loss entities (property damage, personal injury, repair costs). This embodiment specifically designs an importance scoring algorithm (a fusion algorithm based on self-attention weights and TF-IDF values) after the attention mechanism output layer of the GPT-3.5-turbo model. It comprehensively considers word frequency, inverse document frequency, and position weight, and performs weighted calculation on the attention weight of each token. The parameters are set as word frequency weight 0.4, inverse document frequency weight 0.4, and position weight 0.2. Finally, it outputs a list of key information sorted by importance, classified into basic accident information, party information, vehicle information, and loss situation, etc. Compared with the traditional TF-IDF algorithm, the accuracy of key information recognition is improved by 15%. This embodiment specifically designs a contradiction identification logic. When the similarity is below 0.3 and contains the same entity, it is marked as a potential contradiction. Specifically, named entity recognition technology is used to identify different representations of the same entity. Cosine similarity is calculated based on the BERT model, with a threshold set to 0.3. The calculation expression is:
[0042] in Using the all-MiniLM-L6-v2 model from the sentence-transformers library, the types of contradictions can be classified as follows: temporal contradictions: inconsistent time descriptions of the same event in different documents; numerical contradictions: inconsistent numerical information such as amount and quantity; status contradictions: inconsistent information such as vehicle status and personnel status; and relational contradictions: inconsistent logical relationships such as the relationship between the parties and the determination of responsibility. The final output is a contradiction detection report, which includes contradictory items, similarity scores, and suggested verification content. The structured text information is verified for completeness based on a preset standard information checklist and the contradiction detection report. The preset standard information checklist is an information checklist established by the national traffic accident investigation standard. By calculating the coverage of necessary information, the system automatically identifies missing key information and finally outputs a completeness assessment report, which includes a list of missing information and supplementary suggestions.
[0043] The second large-scale text model employs the Sentence-BERT model. Based on this model, document stitching is performed on the document images according to the structured text information, including: The Sentence-BERT model was used to calculate the semantic similarity between the document photo and the structured text information to obtain the similarity analysis results. Based on the similarity analysis results, the document photos are logically ordered to obtain logically sorted document photos; Based on a preset standard document template, the missing content is inferred from the logically sorted document photos to obtain the missing content; The pHash image fingerprint algorithm is used to extract features from the logically sorted document photos to generate image fingerprints. The text hash fingerprint algorithm is used to process the structured text information to obtain text fingerprints. By comparing the image fingerprints and the text fingerprints, duplicate pages between the logically sorted document photos and the structured text information are detected to obtain duplicate pages. After filling in missing content and removing duplicate pages from the logically sorted document photos, the documents are then stitched together to obtain the stitched document photos.
[0044] The specific implementation of this step is as follows: Traffic accident investigations often generate multi-page continuous documents (such as accident determination reports, interrogation records, mediation agreements, etc.). These documents often suffer from problems such as disordered page order, repeated shooting, and missing pages when photographed at the scene. Intelligent stitching is necessary to restore the complete document content. First, the Sentence-BERT model is used to calculate the semantic similarity between pages of the document photos and structured text information. Cosine similarity is used to calculate the semantic relevance of page content. Simultaneously, structural features such as page format, layout, and number of paragraphs are compared. By analyzing the semantic and structural similarity between pages, the degree of association between pages is determined, and a page similarity matrix is constructed for subsequent logical order reasoning. This implementation uses a weighted comprehensive similarity calculation method, including text similarity (weight 0.5): calculating text similarity between pages based on semantic content; structural similarity (weight 0.3): comparing structural features such as page format and layout; and temporal order similarity (weight 0.2): inferring the page order based on time information. The comprehensive similarity calculation formula is as follows:
[0045] The final result is a similarity analysis including a similarity matrix between pages. Based on the similarity analysis results, a temporal relationship inference algorithm is used. When document photo pages contain time reference information, they are sorted by time sequence; when they contain number sequence, they are sorted by numerical order, supporting both Chinese and Arabic numerals; when they contain chapter titles, they are sorted by document structure logic. Furthermore, the page order is inferred based on the semantic relevance of the document photo content. Semantic relevance includes a semantic correlation between adjacent page content greater than or equal to 0.6, timestamp intervals within a reasonable range, and logical coherence of paragraphs and sentences. The temporal relationship inference algorithm first constructs a document photo page relationship graph: ,in As page nodes, each document node contains features such as text vectors, key entities, and timestamps. For relation edges, For weights, a graph algorithm is used for topological sorting to determine the logical order of pages, and edge weights are used. ,in =0.6, =0.2, =0.2; The PageRank algorithm is used to calculate the importance of nodes, and an important connection is filtered with a threshold of 0.1 to find the optimal page arrangement scheme; Finally, the logical order and continuity score of the pages are deduced, and then the logical order of the document photos is deduced to obtain the logically sorted document photos; The missing pages of logically sorted document images are inferred based on contextual content analysis and compared with standard document templates. Semantic completion technology is used to predict possible missing information based on the content of the preceding and following pages, resulting in the inference results and confidence scores for the missing page content. The missing content is then determined based on the confidence scores. For logically sorted document images and structured text information, the pHash image fingerprint algorithm and text hash fingerprint algorithm are used for duplicate detection, respectively. When the image similarity exceeds 95% and the text similarity exceeds 98%, it is determined to be a duplicate page. Finally, the duplicate page identification results and deduplication suggestions are output. After filling in missing content and removing duplicate pages from the logically sorted document images, the documents are then stitched together to obtain the stitched document images. This embodiment can also score the stitched document images from three dimensions: continuity, completeness, and logical consistency. The calculation expressions are as follows:
[0046] in Based on the importance of the fields, such as a weight of 0.3 for ID number and 0.2 for name, the final splicing quality assessment report and optimization suggestions are obtained.
[0047] The third large text model adopts the GPT-4 model.
[0048] The specific implementation of this step is as follows: Based on the GPT-4 model, a case summary is generated using document photos, structured text information, an integrity assessment report, stitched document photos, and a preset traffic accident case summary template. The preset traffic accident case summary template is a standardized template that conforms to the requirements of the "Road Traffic Accident Handling Work Specifications," containing a fixed structure including basic accident information, party information, vehicle information, accident details, loss assessment, and liability analysis. Specifically, it includes: basic accident information: time, location, weather, and road conditions; party information: detailed identity information and contact details of all parties involved; vehicle information: technical parameters, damage, and insurance information of the vehicles involved; accident details: a complete reconstruction of the accident process based on fragmented information; loss assessment: a detailed assessment of personal injury and property damage; and liability analysis: a preliminary determination of liability based on traffic regulations. Finally, a structured traffic accident case summary report is output based on the template, meeting the standardization requirements of the traffic police department.
[0049] After step S5, the method further includes using RAG technology to retrieve the case summary based on the questions input by the user, and obtaining the retrieval results.
[0050] The specific implementation of this step is as follows: Based on the user's natural language query, the system uses RAG (Retrieval-Augmented Generation) technology to retrieve relevant content from the case summary and generate an accurate and comprehensive answer. This question-and-answer system can accurately understand the professional terminology and query intent related to traffic accidents and generate a structured answer to the user's question, including information sources and credibility scores. Based on the case summary and the generated structured answer, a more comprehensive case summary is generated using various document generation technologies, supporting different output formats (Markdown format: structured document that is easy to read and edit; PDF format: standardized report that meets official requirements; JSON format: structured data that is easy to integrate into the system; visualization charts: graphical display of accident information).
[0051] like Figure 2As shown, the complete system architecture based on this embodiment and the solutions adopted in the above embodiments significantly improves information extraction accuracy for traffic accident investigation scenarios. The accuracy rate for extracting traffic accident investigation document information reaches over 97%, the accuracy rate for extracting document information fields reaches over 99%, and the document classification accuracy rate reaches over 90%. Processing efficiency is comprehensively optimized; through a layered architecture design, batch processing capabilities are achieved, with an average single document processing time of 1.23 seconds and a system response time controlled within 3 seconds. Functional completeness is greatly enhanced, supporting intelligent splicing of multi-page documents with an accuracy rate exceeding 95%, generating standardized case summary reports that meet the requirements of traffic police departments, achieving 100% completeness, and supporting natural language query and interaction functions. Its application value is outstanding, improving traffic accident investigation and processing efficiency by over 60%, reducing manual review workload by over 70%, and achieving end-to-end intelligent processing from scattered photos to structured case analysis. The technical performance of this embodiment is compared as follows: Figure 5 As shown.
Claims
1. A method for parsing photographs from traffic accident investigation documents based on a large model, characterized in that, include: S1: Based on the first multimodal large model, the traffic accident photo set is quickly classified and filtered to obtain document photos; S2: Perform fine document analysis on the document photo based on the second multimodal large model to obtain structured text information; S3: Based on the first text big model, perform content understanding on the structured text information to obtain an integrity assessment report; S4: Based on the second text big model, the document photo is stitched together according to the structured text information to obtain the stitched document photo; S5: Based on the third text big model, generate a case summary according to the document photo, the structured text information, the integrity assessment report, the spliced document photo, and the preset traffic accident case summary template.
2. The method for parsing traffic accident investigation documents and photos based on a large model according to claim 1, characterized in that, The first multimodal large model is composed of a MobileNet-V3 visual encoder and a BERT-Base text encoder. Based on the first multimodal large model, a fast classification and filtering of traffic accident photo sets is performed, including: The traffic accident photo set is subjected to standardized preprocessing to obtain a standard traffic accident photo set; The MobileNet-V3 visual encoder was used to extract features from the standard traffic accident photo set to obtain visual feature vectors. The BERT-Base text encoder was used to extract features from the standard traffic accident photo set to obtain text feature vectors. An attention mechanism is used to fuse the visual feature vector and the text feature vector to obtain a first fused feature vector. A multilayer perceptron is used to classify the first fused feature vector to obtain classified document photos and corresponding confidence scores. The Softmax function is used to calculate the confidence score based on the confidence score corresponding to the classified document photos. Document photos are selected based on the confidence score being greater than a preset threshold. The document photos include ID document photos, document photos, and table document photos.
3. The method for parsing traffic accident investigation documents and photos based on a large model according to claim 1, characterized in that, The second multimodal large model is constructed using a ResNet-152 visual encoder and a RoBERTa-large text encoder. Based on this second multimodal large model, fine-grained document analysis is performed on the document photograph, including: The document image was analyzed for depth features using a ResNet-152 visual encoder to obtain visual depth features. The document image was subjected to deep feature analysis using the RoBERTa-large text encoder to obtain text depth features; An attention mechanism is used to fuse the visual depth features and the text depth features to obtain a second fused feature. The second fused feature is classified using a multilayer perceptron to obtain sub-document photos, which include ID card type sub-document photos, document type sub-document photos, and table type sub-document photos; Based on the sub-document photos, select the corresponding document photo processing strategy to extract information and obtain document information, which includes certificate document information, document document information and table document information; The document information is processed using a structured analysis algorithm of the corresponding type to obtain structured text information, which includes document-type structured text information, formal document-type structured text information, and table-type structured text information.
4. The method for parsing traffic accident investigation documents and photos based on a large model according to claim 3, characterized in that, Based on the sub-document photos, select the corresponding document photo processing strategy for information extraction, including: The document information includes: document border, document field positions, document information, and document photo. An improved Hough transform algorithm is used to detect rectangular borders in the document photo, and the Canny edge detection algorithm is used to extract the edges of the document photo to obtain the document border. A template matching algorithm is used to locate key fields in the document photo to obtain the document field positions. A CRNN algorithm is used to perform text recognition in the document photo to obtain the document information. An MTCNN algorithm is used to perform face detection in the document photo to obtain face bounding box information, and the document photo is segmented based on the face bounding box information. Document information includes: document structure, document table data, document seal, and document signature; the layout analysis algorithm is used to perform structural analysis on the document sub-document photos to obtain the document structure; the table recognition algorithm is used to extract the table data from the document sub-document photos to obtain the document table data; the YOLOv5 object detection algorithm is used to perform specific object detection on the document sub-document photos to obtain the document seal and document signature. The information for table-type documents includes: table lines, cell boundaries, and cell content; Hough transform is used to detect lines in the image of the table-type sub-document to obtain the table lines; cell segmentation algorithm is used to locate cell boundaries in the image of the table-type sub-document to obtain the cell boundaries; and DB text detection algorithm is used to identify cell content in the image of the table-type sub-document to obtain the cell content.
5. The method for parsing traffic accident investigation documents and photos based on a large model according to claim 4, characterized in that, The document information is processed using a structured analysis algorithm of the corresponding type, including: A pattern matching algorithm is used to match a preset field mapping table with the document information to obtain the field type; a coordinate transformation algorithm is used to convert the position of the document field into the absolute coordinates of the document border in a reference coordinate system to obtain the position information; an adaptive weighted fusion algorithm is used to calculate the confidence level based on the document field position, document information, and document photo; a multi-dimensional data fusion method is used to fuse the field type, position information, and confidence level to obtain document-type structured text information. The document structure is reconstructed using natural language processing technology to obtain the reconstructed document structure; named entity recognition technology combined with a deep learning model is used to identify the document's table data to obtain entity information; computer vision algorithms are used to separate the document's seal and signature to obtain image elements; layout analysis technology is used to extract formatting information from the document, including font type, font size, alignment, and line spacing; and a multi-dimensional data fusion method is used to fuse the reconstructed document structure, entity information, image elements, and formatting information to obtain the document's structured text information. Image processing algorithms are used to detect and reconstruct the table lines to obtain the table structure; a machine learning classifier is used to classify the cell content to obtain classified data, and a data format conversion algorithm is used to standardize the data to obtain standardized classified data; cells are merged according to their boundaries to obtain the merging relationship; a multi-dimensional data fusion method is used to fuse the table structure, standardized classified data, and merging relationship to obtain table-like structured text information.
6. The method for parsing traffic accident investigation documents and photos based on a large model according to claim 5, characterized in that, After obtaining the structured text information, the process also includes correcting the structured text information to obtain a quality verification report and the corrected structured text information.
7. The method for parsing traffic accident investigation documents and photos based on a large model according to claim 1, characterized in that, The first large-scale text model employs the GPT-3.5-turbo model. Based on this model, content understanding is performed on the structured text information, including: The GPT-3.5-turbo model is used to perform deep semantic understanding on the structured text information to obtain labeled structured text information; The improved TF-IDF algorithm combined with text embedding technology is used to extract keywords from the labeled structured text information to obtain a list of key information. The BERT model is used to calculate the text semantic similarity of the key information list to obtain a contradiction detection report; The integrity of the structured text information is verified based on the preset standard information checklist and the contradiction detection report to obtain an integrity assessment report.
8. The method for parsing traffic accident investigation documents and photos based on a large model according to claim 1, characterized in that, The second large-scale text model employs the Sentence-BERT model. Based on this model, document stitching is performed on the document images according to the structured text information, including: The Sentence-BERT model was used to calculate the semantic similarity between the document photo and the structured text information to obtain the similarity analysis results. Based on the similarity analysis results, the document photos are logically ordered to obtain logically sorted document photos; Based on a preset standard document template, the missing content is inferred from the logically sorted document photos to obtain the missing content; The pHash image fingerprint algorithm is used to extract features from the logically sorted document photos to generate image fingerprints. The text hash fingerprint algorithm is used to process the structured text information to obtain text fingerprints. By comparing the image fingerprints and the text fingerprints, duplicate pages between the logically sorted document photos and the structured text information are detected to obtain duplicate pages. After filling in missing content and removing duplicate pages from the logically sorted document photos, the documents are then stitched together to obtain the stitched document photos.
9. The method for parsing traffic accident investigation documents and photos based on a large model according to claim 1, characterized in that, The third text model uses the GPT-4 model.
10. The method for parsing traffic accident investigation documents and photos based on a large model according to claim 1, characterized in that, After step S5, the method further includes using RAG technology to retrieve the case summary based on the questions input by the user, and obtaining the retrieval results.
Citation Information
Patent Citations
Electronic data forensic analysis method and system based on multiple modes, medium and equipment
CN119004024A