AI technology-based traffic transportation file evaluation method and system

By using AI technology for image recognition and cross-validation, the case file review process has been optimized, solving the problem of low efficiency in manual review in existing technologies and achieving efficient and accurate case file review.

CN121686485APending Publication Date: 2026-03-17HEBEI ALPHASTA TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-04
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

The current method of reviewing transportation case files relies on manual review of each file, which is inefficient and increases personnel costs and resource consumption.

Method used

An AI-based method for reviewing transportation case files was adopted. By cross-validating image recognition models and transcript information, the parameters of the image recognition model were optimized to improve recognition efficiency. Logical inference was used to confirm and evaluate the case file identification status.

Benefits of technology

It significantly improved the efficiency and accuracy of case file review, reduced the false alarm rate, reduced manual intervention, and improved the overall efficiency and accuracy of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121686485A_ABST
    Figure CN121686485A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic transportation file evaluation method and system based on an AI technology, and relates to the technical field of AI. The method comprises the following steps: acquiring a traffic image, record information and a first affirmation result in a target traffic file; importing the traffic image into an image recognition model to obtain a recognition result; performing cross validation on the identification result and the record information to obtain an available information set and a corresponding weight coefficient; importing the available information set and the corresponding weight coefficient into an evaluation model to obtain a second affirmation result; and judging the affirmation state of the target traffic file according to the first affirmation result and the second affirmation result, and then evaluating the affirmation state. According to the image recognition model, calculation of parameters is reduced, the recognition efficiency is improved, cross validation is carried out on the recognition result and the record information according to the recognition result, unavailable information in the file is removed, the quality of high-quality information is higher, the false alarm rate is reduced, and then the execution efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of AI technology, specifically relating to a method and system for reviewing transportation case files based on AI technology. Background Technology

[0002] Current methods for reviewing transportation case files often rely on manual, document-by-document review and analysis, which suffers from low efficiency and the risk of missing crucial information. By leveraging AI technology, through intelligent recognition, analysis, and mining of various data types, including text and images, case file reviews can be completed quickly and accurately. Utilizing AI's natural language processing and computer vision technologies, various information within transportation case files, such as vehicle information, personnel information, accident scene descriptions, and evidence, can be automatically analyzed to identify key elements, potential problems, and violations. This enables efficient and accurate case file reviews, providing strong support for decision-making by transportation management departments and effectively improving the efficiency and standardization of management in the transportation sector.

[0003] Existing technology (publication number: CN120599626A) provides a case file processing method, system, computer equipment, and storage medium based on AI, RPA, and OCR recognition. This addresses the shortcomings of existing technologies in case file material segmentation and standardized cataloging, as well as limitations in application scenarios. It extends the technical solution to complex business scenarios such as courts and hospitals, meeting the personalized needs of different fields, while also solving the problems of chaotic management of paper case files before archiving and inaccurate cataloging of electronic case files. The method performs OCR recognition on pre-processed electronic images, uses an AI engine to analyze and recognize case file content, extracts key information, and classifies it, segmenting case file materials according to different types. Based on the AI ​​recognition results and predefined cataloging rules, an RPA robot automatically generates standardized case file titles and catalog names, and categorizes the case file materials into corresponding directory levels, forming a complete electronic case file directory tree.

[0004] The aforementioned patents have solved the problems of chaotic management before case file archiving, inaccurate electronic case file cataloging, and case file classification. However, in practical applications, a large number of examiners are still required to review the case files, which increases personnel costs, thereby reducing the efficiency of the examination and increasing the consumption of transportation resources. Summary of the Invention

[0005] The purpose of this invention is to solve the problem that a large number of personnel are needed to review case files, which increases personnel costs, reduces review efficiency, and increases the consumption of transportation resources. Therefore, this invention proposes a transportation case file review method and system based on AI technology.

[0006] In a first aspect of this invention, a method for reviewing transportation case files based on AI technology is first proposed, the method comprising: Obtain traffic images, written records, and initial determination results from the target traffic case file; The traffic images are imported into an image recognition model to obtain recognition results; the recognition results include traffic behavior and subject information. The recognition results are cross-validated with the transcript information to obtain the usable information set and the corresponding weight coefficients; The available information set and corresponding weight coefficients are imported into the evaluation model to obtain the second determination result; The determination status of the target traffic file is determined based on the first determination result and the second determination result; the determination status includes qualified and unqualified. Target traffic files that are deemed qualified are stored in a preset database, while target traffic files that are deemed unqualified are marked as requiring rectification and pushed to the terminal.

[0007] Optionally, the image recognition model is an improvement upon the YOLOv8 model, specifically including: In the backbone network structure, the third, fifth, seventh and ninth layers of the original YOLOv8 model are replaced with C2f_PConv modules; The principle and process of the C2f_PConv module include: Replace the Bottleneck module in the C2f module of the original YOLOv8 model with the PaConvBottleneck module; The calculation process of the PaConvBottleneck module includes:

[0008] In this case, the number of channels in the Conv layer is replaced by C / 4 to obtain the PConv layer; The operators represent the PConv layer, where X represents the input features of the PaConvBottleneck module and Y represents the output features of the PaConvBottleneck module.

[0009] Optionally, the image recognition model is obtained by improving the YOLOv8 model, and the specific improvements include: A CS module is added between the C2f module and the three detection head modules in the neck network structure; The working principle of the CS module includes: The output features of the C2f module are used as the input feature tensor of the CS module; The horizontal and vertical feature tensors are obtained by performing horizontal average pooling and vertical average pooling on the input feature tensor, respectively. The first feature tensor is obtained by concatenating the horizontal feature tensor and the vertical feature tensor. The second feature tensor is obtained by performing a 3×3 convolution on the first feature tensor. The second feature tensor is subjected to a two-branch global average pooling to obtain the third and fourth feature tensors; The fifth feature tensor is obtained by performing a fully connected process on the third feature tensor and then ReLU activation. The sixth feature tensor is obtained by performing a fully connected process on the fourth feature tensor and then ReLU activation. The target feature tensor is obtained by fusing the fifth feature tensor and the sixth feature tensor.

[0010] Optionally, the step of cross-validating the recognition results with the transcript information to obtain the usable information set and the corresponding weight coefficients includes: Semantic analysis is performed on the transcript information to output a list of transcript statements and corresponding target tags; the target tags include: direct match, logical match, no match, and direct conflict; The direct matching refers to matching the written statement with the recognition result of the image; The logical matching refers to matching the written statement with the recognition result derived from logical reasoning; The inability to match means that the identification result does not contain information that can verify or refute the statement in the transcript. The direct conflict refers to a direct contradiction between the written statement and the identification result; Transcript statements that are directly or logically matched are marked as having high credibility and assigned a first proportional weighting coefficient, and are directly included in the available information set; Unmatched transcript statements are marked as medium confidence, included in the available information set, and assigned a second proportional weighting coefficient; the first proportional weighting coefficient is greater than the second proportional weighting coefficient. Conflicting statements in the transcripts are marked as low credibility, excluded from the available information set, and a record of evidence conflict is generated.

[0011] Optionally, the principle process of the logical deduction confirmation includes: All traffic safety rules that satisfy the recognition result are extracted from a preset rule base; each traffic safety rule has a preset confidence level; the preset rule base consists of multiple traffic safety rules; the confidence level includes high confidence, medium confidence, and low confidence. Generate inferred facts based on all traffic safety rules that satisfy the identification results and their corresponding confidence levels; The transcript statements are compared and verified with the deduced facts. If the semantics of the content are consistent, the logical inference is confirmed.

[0012] In another aspect of this invention, a transportation case file review system based on AI technology is first proposed, the system comprising: Data acquisition module: Acquires traffic images, record information, and the first determination result from the target traffic case file; Image recognition module: imports the traffic image into the image recognition model to obtain the recognition result; the recognition result includes traffic behavior and subject information; Verification module: Cross-validates the recognition results with the transcript information to obtain the usable information set and the corresponding weight coefficients; The identification and assessment module imports the available information set and corresponding weight coefficients into the assessment model to obtain the second identification result; Judgment module: Determines the determination status of the target traffic file based on the first determination result and the second determination result; the determination status includes qualified and unqualified; The assessment result module stores target traffic files with a qualified status in a preset database, marks target traffic files with a non-qualified status as needing rectification, and uploads them to the assessment terminal.

[0013] Optionally, the image recognition module is further configured to use an image recognition model that is an improvement upon the YOLOv8 model, specifically including: In the backbone network structure, the third, fifth, seventh and ninth layers of the original YOLOv8 model are replaced with C2f_PConv modules; The principle and process of the C2f_PConv module include: Replace the Bottleneck module in the C2f module of the original YOLOv8 model with the PaConvBottleneck module; The calculation process of the PaConvBottleneck module includes:

[0014] In this case, the number of channels in the Conv layer is replaced by C / 4 to obtain the PConv layer; The operators represent the PConv layer, where X represents the input features of the PaConvBottleneck module and Y represents the output features of the PaConvBottleneck module.

[0015] Optionally, the image recognition module is further configured to use an image recognition model that is an improvement upon the YOLOv8 model, specifically including: A CS module is added between the C2f module and the three detection head modules in the neck network structure; The working principle of the CS module includes: The output features of the C2f module are used as the input feature tensor of the CS module; The horizontal and vertical feature tensors are obtained by performing horizontal average pooling and vertical average pooling on the input feature tensor, respectively. The first feature tensor is obtained by concatenating the horizontal feature tensor and the vertical feature tensor. The second feature tensor is obtained by performing a 3×3 convolution on the first feature tensor. The second feature tensor is subjected to a two-branch global average pooling to obtain the third and fourth feature tensors; The fifth feature tensor is obtained by performing a fully connected process on the third feature tensor and then ReLU activation. The sixth feature tensor is obtained by performing a fully connected process on the fourth feature tensor and then ReLU activation. The target feature tensor is obtained by fusing the fifth feature tensor and the sixth feature tensor.

[0016] Optionally, the verification module is further configured to cross-validate the recognition results with the transcript information to obtain a usable information set and corresponding weight coefficients, including: Semantic analysis is performed on the transcript information to output a list of transcript statements and corresponding target tags; the target tags include: direct match, logical match, no match, and direct conflict; The direct matching refers to matching the written statement with the recognition result of the image; The logical matching refers to matching the written statement with the recognition result derived from logical reasoning; The inability to match means that the identification result does not contain information that can verify or refute the statement in the transcript. The direct conflict refers to a direct contradiction between the written statement and the identification result; Transcript statements that are directly or logically matched are marked as having high credibility and assigned a first proportional weighting coefficient, and are directly included in the available information set; Unmatched transcript statements are marked as medium confidence, included in the available information set, and assigned a second proportional weighting coefficient; the first proportional weighting coefficient is greater than the second proportional weighting coefficient. Conflicting statements in the transcripts are marked as low credibility and excluded from the usable information set.

[0017] Optionally, the principle process of the logical deduction confirmation includes: All traffic safety rules that satisfy the recognition result are extracted from a preset rule base; each traffic safety rule has a preset confidence level; the preset rule base consists of multiple traffic safety rules; the confidence level includes high confidence, medium confidence, and low confidence. Generate inferred facts based on all traffic safety rules that satisfy the identification results and their corresponding confidence levels; The transcript statements are compared and verified with the deduced facts. If the semantics of the content are consistent, the logical inference is confirmed.

[0018] The beneficial effects of this invention are: This invention proposes an AI-based method for reviewing transportation case files. By optimizing image recognition model parameters and simplifying the calculation process, this invention significantly improves recognition efficiency, thereby enhancing overall identification efficiency. Cross-validation of the recognition results with transcript information filters out invalid information in the case files, strengthens the quality of valid information, and reduces the false alarm rate, thus improving both the efficiency and accuracy of identification and ensuring the accuracy of case file review. Attached Figure Description

[0019] The invention will now be further described with reference to the accompanying drawings.

[0020] Figure 1 A flowchart illustrating an AI-based method for reviewing transportation case files, provided as an embodiment of the present invention; Figure 2 A network structure diagram of a YOLOv8 model provided in an embodiment of the present invention; Figure 3 This is a network structure diagram of an image recognition model provided in an embodiment of the present invention; Figure 4 A network structure diagram of a C2f_PConv module provided in an embodiment of the present invention; Figure 5 This is a framework diagram of a transportation case file review system based on AI technology, provided for an embodiment of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and B can represent: A alone, A and B simultaneously, and B alone. Furthermore, descriptions involving "first," "second," etc., in this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" can explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0022] Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] This invention provides a method for reviewing transportation case files based on AI technology. See also... Figure 1 , Figure 1 A flowchart illustrating an AI-based method for reviewing transportation case files, provided as an embodiment of the present invention. The method includes the following steps: Obtain traffic images, written records, and initial determination results from the target traffic case file; Traffic images are imported into an image recognition model to obtain recognition results; the recognition results include traffic behavior and subject information. The recognition results are cross-validated with the transcript information to obtain the usable information set and the corresponding weight coefficients; The available information set and corresponding weighting coefficients are imported into the evaluation model to obtain the second determination result; The determination status of the target traffic file is determined based on the first and second determination results; the determination status includes qualified and unqualified. Target traffic files that are deemed qualified are stored in a pre-set database, while target traffic files that are deemed unqualified are marked as requiring rectification and uploaded to the evaluation terminal.

[0024] The transportation case file review method based on AI technology provided by this invention improves recognition efficiency by reducing the number of parameters calculated in the image recognition model, thereby improving the overall identification efficiency. By cross-validating the recognition results and transcript information, unusable information in the case file is eliminated, resulting in higher quality information, reducing the false alarm rate, and thus improving execution efficiency and accuracy, thereby improving the accuracy of case file review.

[0025] Specifically, the traffic behaviors identified include running red lights, illegal parking, driving outside designated lanes, and overloading; key information includes vehicle license plate numbers and driver information; the preset database is set up based on the staff's historical experience; the first determination result is the percentage of responsibility among the three parties in the target traffic case file. For example, in a rear-end collision between two vehicles traveling in a lane, the following vehicle (Party B) and the preceding vehicle (Party A) failed to maintain a safe braking distance; in this accident, Party A's responsibility percentage is 99.9%, and Party B's responsibility percentage is 0.1%; therefore, it is inferred that Party A is fully responsible, and Party B is not responsible. Similarly, the second determination result is the same. In one implementation, the preset violation threshold is stored in a preset clause database and set based on the staff's historical experience; In one implementation, see [link to implementation details]. Figure 2 , Figure 2 A network structure diagram of a YOLOv8 model provided in an embodiment of the present invention; see also Figure 3 , Figure 3 This is a network structure diagram of an image recognition model provided in an embodiment of the present invention. The image recognition model is obtained by improving the YOLOv8 model, and the specific improvements include: In the backbone network structure, the third, fifth, seventh and ninth layers of the original YOLOv8 model are replaced with C2f_PConv modules; A CS module is added between the C2f module and the three detection head modules in the neck network structure; See Figure 4 , Figure 4 This invention provides a network structure diagram for a C2f_PConv module, and the principle process of the C2f_PConv module includes: Replace the Bottleneck module in the C2f module of the original YOLOv8 model with the PaConvBottleneck module; The calculation process of the PaConvBottleneck module includes:

[0026] In this case, the number of channels in the Conv layer is replaced by C / 4 to obtain the PConv layer; The operators represent the PConv layer; X represents the input features of the PaConvBottleneck module; Y represents the output features of the PaConvBottleneck module. In one implementation, it should be specifically noted that during the convolution operation in the Conv layer, the input features and output features map the same number of channels C, the width and height of the feature map are W and H respectively, and the convolution kernel is K; The formula for calculating the cost of the Conv layer is: The computational cost of the FW convolutional layer; During the convolution operation in the PConv layer, the number of input feature channels is C / 2, the width and height of the feature map are W and H respectively; the convolution kernel is K; and the output feature map is C / 4. The formula for calculating the cost of the PConv layer is: The computational cost of the PW convolutional layer; The PConv layer uses only regular convolutions with C / 4 channels, which reduces the computational cost of the PConv layer and thus improves the overall model processing speed.

[0027] In one implementation, there are N PaConvBottleneck modules; In one implementation, image recognition models can reduce computational costs and improve traffic images, thereby increasing image processing speed.

[0028] Utilizing this redundancy through PConv layers helps improve cost optimization. Unlike standard convolution, which uses filters on all input channels, increasing the number of parameters and computational load, PConv selectively processes certain channels, omitting others. This approach not only improves processing efficiency but also strikes a balance between computational requirements and accuracy; by improving processing efficiency, it further improves the recognition efficiency of traffic images.

[0029] In one implementation, the working principle of the CS module includes: The output features of the C2f module are used as the input feature tensor of the CS module; The horizontal and vertical feature tensors are obtained by performing horizontal average pooling and vertical average pooling on the input feature tensor, respectively. The first feature tensor is obtained by concatenating the horizontal and vertical feature tensors. The second feature tensor is obtained by performing a 3×3 convolution on the first feature tensor. The second feature tensor is subjected to a two-branch global average pooling to obtain the third and fourth feature tensors; The fifth feature tensor is obtained by performing a fully connected process on the third feature tensor and then ReLU activation. The sixth feature tensor is obtained by performing a fully connected process on the fourth feature tensor and then ReLU activation. The target feature tensor is obtained by fusing the fifth and sixth feature tensors.

[0030] In one implementation, it's worth noting that the CS module reduces computational costs and improves processing speed by simplifying feature dimensions and the computation process. The core logic involves first performing horizontal and vertical average pooling on the input features. This significantly compresses the dimensionality of the feature tensor while preserving key spatial information, reducing the amount of data required for subsequent convolution and fully connected operations. Simultaneously, through dual-branch global average pooling and targeted feature fusion, complex high-dimensional feature transformations and redundant computational steps are avoided, resulting in a simpler feature processing path and reduced computational complexity. Compared to existing modules, the CS module primarily optimizes feature extraction and computational costs. In feature extraction, it effectively reduces feature dimensions through pooling operations, avoiding the high computational costs of directly processing high-dimensional features in some existing modules. In the computational process, it adopts a lightweight architecture of "pooling-convolution-dual-branch fully connected-fusion," eliminating unnecessary feature transformations and repetitive computational steps. At the same time, precise feature fusion ensures information validity, improving computational efficiency without sacrificing core feature representation.

[0031] In one implementation, tensors are generated using horizontal pooling and vertical pooling. These two one-dimensional vectors are then concatenated, and intermediate features representing coordinate attention are generated using a shared 3×3 convolution and a non-linear activation function (such as ReLU). These intermediate features are divided into horizontal and vertical attention weight matrices. Each attention matrix is ​​applied weighted to the original feature map to enhance spatially sensitive feature responses. Finally, the weighted horizontal and vertical feature maps are fused to obtain the final attention-enhanced features, preserving accurate location information. The dual-branch global average pooling design enhances the joint learning ability of channel features and global representations, significantly improving the network's ability to represent features and effectively improving the model's generalization ability across different datasets. This results in a substantial improvement in traffic image recognition capabilities and a more comprehensive review of case files.

[0032] In one implementation, cross-validating the recognition results with the transcript information to obtain a usable information set and corresponding weight coefficients includes: Perform semantic analysis on the transcript information and output a list of transcript statements and corresponding target tags; the target tags include: direct match, logical match, no match, and direct conflict; Direct matching involves matching the written statement with the image recognition results; Logical matching involves matching the written statement with the recognition result derived from logical reasoning. The inability to match indicates that the identification result does not contain information that can verify or refute the statements in the transcript. A direct conflict occurs when the written statement contradicts the identification results. Transcript statements that are directly or logically matched are marked as having high credibility and assigned a first proportional weighting coefficient, and are directly included in the available information set; Unmatched transcript statements are marked as medium confidence, included in the available information set, and assigned a second proportional weighting coefficient; the first proportional weighting coefficient is greater than the second proportional weighting coefficient. Conflicting statements in the transcripts are marked as low credibility and excluded from the usable information set.

[0033] In one implementation, the first proportional weight coefficients corresponding to direct matching and logical matching, and the second proportional weight coefficients corresponding to non-matching are normalized. For example, the weight coefficient for direct matching is 0.9, the weight coefficient for logical matching is 0.95, and the weight coefficient for non-matching is 0.5. Normalization is performed so that the sum of the normalized weight coefficients is 1. Since logical matching is the result of integrating and analyzing multiple images, while direct matching is the result of analyzing a single image, the weight coefficient for logical matching is greater than the weight coefficient for direct matching, which is greater than the weight coefficient for non-matching.

[0034] In one implementation, the logical deduction and confirmation process includes: Extract all traffic safety rules that satisfy the recognition results from the preset rule base; each traffic safety rule has a preset confidence level; the preset rule base consists of multiple traffic safety rules; the confidence levels include high confidence, medium confidence, and low confidence. Generate inferred facts based on all traffic safety rules that satisfy the identification results and their corresponding confidence levels; The transcript statements are compared and verified with the deduced facts. If the semantics of the content are consistent, the logical inference is confirmed.

[0035] In one implementation, the traffic images (photos of violations), written statements (Zhang admitting to failing to yield at a certain time and place), and the first determination result of the target traffic case file are the processing results made by on-site or preliminary review staff based on the above evidence.

[0036] Traffic images (photos of traffic violations) are intelligently identified and analyzed to generate structured identification results (vehicle B, located at the zebra crossing). Cross-validate the written statement with the image recognition results. The specific validation steps are as follows: Once a statement (e.g., "(driver admits to failing to yield the right-of-way)") enters the system, the inference engine performs the following steps: Step 1: The engine checks all traffic safety rules (Rule R1, Rule R2, etc.) in the preset rule base; for specific traffic safety rules, such as Rule R1 (stop and yield rule): if there is (vehicle, located, within the stop line), (traffic light, status, red light), or (vehicle, behavior, movement); then it infers (vehicle, behavior, failure to stop and yield); and searches for traffic safety rules that can be satisfied by the recognition results of the current image.

[0037] Step 2: The identification results are F1: (Vehicle A, located within the stop line), F2: (Traffic light, status, red light), F3: (Vehicle A, behavior, movement). These three completely match the condition part of rule R1. The deduced fact is obtained by deriving the condition part (Vehicle A, behavior, failure to stop and yield). Step 3: If the written statement (driver, admits, failed to stop and yield) and the inferred facts (vehicle A, behavior, failed to stop and yield) are highly consistent in semantics, and the corresponding weight coefficient is 0.9~1.0, and is determined as the first weight coefficient, then the logical inference is confirmed.

[0038] Similarly, the label is directly determined by comparing (vehicle B, located at the zebra crossing) with (driver, admitted, failed to stop and yield), with a corresponding weight coefficient of 0.5~0.8, and is determined as the second weight coefficient.

[0039] The available information set and corresponding weighting coefficients are imported into the evaluation model to obtain the second determination result. By understanding the importance of different pieces of evidence (affected by weights), for example, a standardized penalty code "1032B" is output, corresponding to "driving a motor vehicle in violation of road traffic signals". The system compares the original first determination result with the second determination result and determines the determination status of the case file based on the comparison result: A qualified status corresponds to the condition that the first and second determination results are consistent in the determination of core facts and the application of law. Processing: The system automatically marks the case file as "AI Review Qualified" and archives it to the qualified case file database. This case file process ends. An unqualified status corresponds to the condition that there is a substantial difference between the first and second determination results (such as different determinations of violations, or excessively lenient or severe penalties).

[0040] Processing: The system automatically marked the file as "pending rectification" and uploaded it to the review terminal for processing.

[0041] In one implementation, the training process for evaluating the model includes: The process involves obtaining the available information set and corresponding weight coefficients from historical traffic files; the second assessment result is an integrated evaluation by experts based on the available information set and corresponding weight coefficients from each historical traffic file; the higher the weight coefficient corresponding to the available information set, the greater the role of the available information set, and the more accurate the second assessment result, thus bringing more convenience to the staff; the available information set, the weight coefficients corresponding to the available information set, and the second assessment result are integrated into several training data and test data. The training data is imported into the artificial intelligence model for training, and the tested data is used to test the trained artificial intelligence model. Specifically, the available information set and corresponding weight coefficients from the historical traffic files in the tested data are input into the trained artificial intelligence model, and a second determination result is output. The absolute value of the difference between the second determination result and the second determination result recorded in the tested data is checked to see if it is within an acceptable range. If yes, it means that the set of tested data has passed the test, and the next set of tested data is tested. If not, the relevant parameters of the artificial intelligence model need to be adjusted, and the test data continues to be used for testing. This continues until a set proportion of the tested data passes the test. Finally, an evaluation model is obtained, with the input being the available information set from the historical traffic files, the corresponding weight coefficients of the available information set, and the second determination result, and the output being the second determination result. The artificial intelligence model is an MLP model.

[0042] In one implementation, the MLP model includes an input layer, two fully connected hidden layers, and an output layer. The input dimension of the model is the same as the feature vector dimension of the available information set, and the output dimension is the same as the number of categories or regression values ​​of the identified results. The number of neurons in the hidden layer can be configured between 128 and 1024, and overfitting is prevented using techniques such as Dropout. The total number of trainable parameters of the MLP model ranges from hundreds of thousands to millions, depending on the complexity and accuracy requirements of the application.

[0043] Specifically, the assessment results are based on the overall evaluation indicators of the traffic case files; In one implementation, determining the identification status of the target traffic file based on the first identification result and the second identification result includes: When the first determination result and the second determination result are semantically consistent, the determination status of the target traffic file is qualified; When the first determination result and the second determination result are semantically inconsistent, the determination status of the target traffic file is unqualified.

[0044] Specifically, the first determination result is based on the data obtained by staff from the target file. The specific method of obtaining the data involves a sufficient number of samples in the target scenario, including both normal and abnormal data (normal data can be used if there is no abnormal data). Then, the data is preprocessed, and the core content is obtained through semantic extraction and judgment. Common methods include semantic feature extraction.

[0045] Based on the same inventive concept, this invention also provides an AI-based transportation case file review system. See also Figure 5 , Figure 5 A framework diagram of a transportation case file review system based on AI technology provided for embodiments of the present invention includes: Data acquisition module: Acquires traffic images, record information, and the first determination result from the target traffic case file; Image recognition module: Imports traffic images into the image recognition model to obtain recognition results; the recognition results include traffic behavior and subject information; Verification module: Cross-validates the recognition results with the transcript information to obtain the usable information set and the corresponding weight coefficients; The identification and assessment module imports the available information set and corresponding weight coefficients into the assessment model to obtain the second identification result; Judgment Module: Determines the determination status of the target traffic file based on the first determination result and the second determination result; the determination status includes qualified and unqualified. The assessment result module stores target traffic files with an assessment status of "qualified" in a preset database, marks target traffic files with an assessment status of "unqualified" as "needing rectification", and pushes them to the terminal.

[0046] The transportation case file review system based on AI technology provided in this invention improves recognition efficiency by reducing the number of parameters calculated in the image recognition model, thereby improving the overall identification efficiency. By cross-validating the recognition results and transcript information and eliminating unusable information in the case file, the system ensures higher quality information, reduces the false alarm rate, and improves execution efficiency and accuracy, thus enhancing the accuracy of case file review.

[0047] The foregoing has described one embodiment of the present invention in detail, but this content is merely a preferred embodiment and should not be considered as limiting the scope of the present invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the scope of the claims of this invention.

Claims

1. A method for reviewing transportation case files based on AI technology, characterized in that, The method includes: Obtain traffic images, written records, and initial determination results from the target traffic case file; The traffic images are imported into an image recognition model to obtain recognition results; the recognition results include traffic behavior and subject information. The recognition results are cross-validated with the transcript information to obtain the usable information set and the corresponding weight coefficients; The available information set and corresponding weight coefficients are imported into the evaluation model to obtain the second determination result; The determination status of the target traffic file is determined based on the first determination result and the second determination result; the determination status includes qualified and unqualified. Target traffic files that are deemed qualified are stored in a pre-set database, while target traffic files that are deemed unqualified are marked as requiring rectification and uploaded to the evaluation terminal.

2. The method for reviewing transportation case files based on AI technology according to claim 1, characterized in that, The image recognition model is an improvement upon the YOLOv8 model, specifically including: In the backbone network structure, the third, fifth, seventh and ninth layers of the original YOLOv8 model are replaced with C2f_PConv modules; The principle and process of the C2f_PConv module include: Replace the Bottleneck module in the C2f module of the original YOLOv8 model with the PaConvBottleneck module; The calculation process of the PaConvBottleneck module includes: ; In this case, the number of channels in the Conv layer is replaced by C / 4 to obtain the PConv layer; The operators represent the PConv layer, where X represents the input features of the PaConvBottleneck module and Y represents the output features of the PaConvBottleneck module.

3. The method for reviewing transportation case files based on AI technology according to claim 2, characterized in that, The image recognition model is an improvement upon the YOLOv8 model, and the specific improvements include: A CS module is added between the C2f module and the three detection head modules in the neck network structure; The working principle of the CS module includes: The output features of the C2f module are used as the input feature tensor of the CS module; The horizontal and vertical feature tensors are obtained by performing horizontal average pooling and vertical average pooling on the input feature tensor, respectively. The first feature tensor is obtained by concatenating the horizontal feature tensor and the vertical feature tensor. The second feature tensor is obtained by performing a 3×3 convolution on the first feature tensor. The second feature tensor is subjected to a two-branch global average pooling to obtain the third and fourth feature tensors; The fifth feature tensor is obtained by performing a fully connected process on the third feature tensor and then ReLU activation. The sixth feature tensor is obtained by performing a fully connected process on the fourth feature tensor and then ReLU activation. The target feature tensor is obtained by fusing the fifth feature tensor and the sixth feature tensor.

4. The method for reviewing transportation case files based on AI technology according to claim 1, characterized in that, The step of cross-validating the recognition results with the transcript information to obtain the usable information set and the corresponding weight coefficients includes: Semantic analysis is performed on the transcript information to output a list of transcript statements and corresponding target tags; the target tags include: direct match, logical match, no match, and direct conflict; The direct matching refers to matching the written statement with the recognition result of the image; The logical matching refers to matching the written statement with the recognition result derived from logical reasoning; The inability to match means that the identification result does not contain information that can verify or refute the statement in the transcript. The direct conflict refers to a direct contradiction between the written statement and the identification result; Transcript statements that are directly or logically matched are marked as having high credibility and assigned a first proportional weighting coefficient, and are directly included in the available information set; Unmatched transcript statements are marked as medium confidence, included in the available information set, and assigned a second proportional weighting coefficient; the first proportional weighting coefficient is greater than the second proportional weighting coefficient. Conflicting statements in the transcripts are marked as low credibility and excluded from the usable information set.

5. The method for reviewing transportation case files based on AI technology according to claim 4, characterized in that, The logical deduction and confirmation process includes: All traffic safety rules that satisfy the recognition result are extracted from a preset rule base; each traffic safety rule has a preset confidence level; the preset rule base consists of multiple traffic safety rules; the confidence level includes high confidence, medium confidence, and low confidence. Generate inferred facts based on all traffic safety rules that satisfy the identification results and their corresponding confidence levels; The transcript statements are compared and verified with the deduced facts. If the semantics of the content are consistent, the logical inference is confirmed.

6. A transportation case file review system based on AI technology, characterized in that, The system includes: Data acquisition module: Acquires traffic images, record information, and the first determination result from the target traffic case file; Image recognition module: imports the traffic image into the image recognition model to obtain the recognition result; the recognition result includes traffic behavior and subject information; Verification module: Cross-validates the recognition results with the transcript information to obtain the usable information set and the corresponding weight coefficients; The identification and assessment module imports the available information set and corresponding weight coefficients into the assessment model to obtain the second identification result; Judgment module: Determines the determination status of the target traffic file based on the first determination result and the second determination result; the determination status includes qualified and unqualified; The assessment result module stores target traffic files with a qualified status in a preset database, marks target traffic files with a non-qualified status as needing rectification, and uploads them to the assessment terminal.

7. The transportation case file review system based on AI technology according to claim 6, characterized in that, The image recognition module is further configured to use an image recognition model that is an improvement upon the YOLOv8 model. Specific improvements include: In the backbone network structure, the third, fifth, seventh and ninth layers of the original YOLOv8 model are replaced with C2f_PConv modules; The principle and process of the C2f_PConv module include: Replace the Bottleneck module in the C2f module of the original YOLOv8 model with the PaConvBottleneck module; The calculation process of the PaConvBottleneck module includes: ; In this case, the number of channels in the Conv layer is replaced by C / 4 to obtain the PConv layer; The operators represent the PConv layer, where X represents the input features of the PaConvBottleneck module and Y represents the output features of the PaConvBottleneck module.

8. The transportation case file review system based on AI technology according to claim 7, characterized in that, The image recognition module is further configured to use an image recognition model that is an improvement upon the YOLOv8 model. Specific improvements include: A CS module is added between the C2f module and the three detection head modules in the neck network structure; The working principle of the CS module includes: The output features of the C2f module are used as the input feature tensor of the CS module; The horizontal and vertical feature tensors are obtained by performing horizontal average pooling and vertical average pooling on the input feature tensor, respectively. The first feature tensor is obtained by concatenating the horizontal feature tensor and the vertical feature tensor. The second feature tensor is obtained by performing a 3×3 convolution on the first feature tensor. The second feature tensor is subjected to a two-branch global average pooling to obtain the third and fourth feature tensors; The fifth feature tensor is obtained by performing a fully connected process on the third feature tensor and then ReLU activation. The sixth feature tensor is obtained by performing a fully connected process on the fourth feature tensor and then ReLU activation. The target feature tensor is obtained by fusing the fifth feature tensor and the sixth feature tensor.

9. A transportation case file review system based on AI technology according to claim 6, characterized in that, The verification module is further configured to cross-validate the recognition results with the transcript information to obtain a usable information set and corresponding weight coefficients, including: Semantic analysis is performed on the transcript information to output a list of transcript statements and corresponding target tags; the target tags include: direct match, logical match, no match, and direct conflict; The direct matching refers to matching the written statement with the recognition result of the image; The logical matching refers to matching the written statement with the recognition result derived from logical reasoning; The inability to match means that the identification result does not contain information that can verify or refute the statement in the transcript. The direct conflict refers to a direct contradiction between the written statement and the identification result; Transcript statements that are directly or logically matched are marked as having high credibility and assigned a first proportional weighting coefficient, and are directly included in the available information set; Unmatched transcript statements are marked as medium confidence, included in the available information set, and assigned a second proportional weighting coefficient; the first proportional weighting coefficient is greater than the second proportional weighting coefficient. Conflicting statements in the transcripts are marked as low credibility and excluded from the usable information set.

10. A transportation case file review system based on AI technology according to claim 9, characterized in that, The logical deduction and confirmation process includes: All traffic safety rules that satisfy the recognition result are extracted from a preset rule base; each traffic safety rule has a preset confidence level; the preset rule base consists of multiple traffic safety rules; the confidence level includes high confidence, medium confidence, and low confidence. Based on all traffic safety rules that satisfy the identification results and their corresponding confidence levels, inferred facts are generated; the transcript statements are compared and verified with the inferred facts. If the semantics of the content are consistent, it is determined that the logical inference is confirmed.

Citation Information

Patent Citations

  • Document processing method based on AI, RPA and OCR

    CN120599626A