Risk identification method and device for vehicle loss assessment order
By using a multimodal visual language model to perform structured recognition and comparison of damage assessment reports, the problems of misjudgment and high cost in the traditional damage assessment model are solved, risk identification and cost control are achieved, and the accuracy and efficiency of the damage assessment process are improved.
Patent Information
- Application Number
- CN202610213082.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-13
- Publication Date
- 2026-05-19
AI Technical Summary
Traditional manual damage assessment methods are prone to misjudgment when dealing with complex vehicle models, new structures, or hidden damage, resulting in unreasonable repair plans, lack of bargaining power, long claims cycles and high costs, and difficulty in controlling claims risks and costs.
By acquiring the first damage assessment data generated based on logic and the damage assessment report image provided by a third party, a multimodal visual language model is used for structured recognition. Combined with preset prompt word constraints, the system identifies and compares parts information and price information to identify risk items, including parts code consistency, price deviation, and repair method matching rules.
It enables accurate identification of risk items in third-party damage assessment reports, reduces claims costs, avoids damage expansion and over-repair, and improves the accuracy and efficiency of the damage assessment process.
Smart Images

Figure CN122066530A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automotive information technology, and in particular to a method and apparatus for risk identification in vehicle damage assessment reports. Background Technology
[0002] With the continuous growth of car ownership, the volume of vehicle insurance claims is increasing rapidly. In the vehicle accident loss assessment (damage assessment) stage, the damage assessor needs to go to the scene in person, rely on personal experience to judge the scope of the loss, determine the repair plan, and manually take a large number of photos. Afterwards, the photos need to be manually linked with the damage assessment items and entered into the system.
[0003] However, traditional manual damage assessment methods have the following drawbacks: When faced with complex vehicle models, new structures, or hidden damage, inexperienced assessors are prone to misjudgments, leading to unreasonable repair plans or a lack of bargaining power when negotiating with repair shops, thus increasing insurance claims costs. Therefore, the skill level of the assessor directly affects the accuracy of the judgment. In addition, the lack of standardized operating guidelines for collecting on-site evidence (such as photos and fault codes) frequently results in problems such as incomplete photo angles, failure to photograph key damaged areas, and omissions of electronic diagnostic information, thereby prolonging the claims process and increasing operating costs. Furthermore, it requires manual matching of dozens or even hundreds of on-site photos with the damage assessment items one by one, which is time-consuming, labor-intensive, and prone to errors. Moreover, traditional damage assessment methods struggle to accurately identify risk points in damage estimates provided by third parties (such as repair shops), thus failing to effectively control the risks and costs of claims. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method and apparatus for risk identification of vehicle damage assessment reports, which can accurately identify risk items in third-party damage assessment reports.
[0005] To achieve the above objectives, according to one aspect of the present invention, a risk identification method for vehicle damage assessment reports is provided, comprising:
[0006] Obtain first damage assessment data generated based on the first logic, which includes parts price information;
[0007] Obtain an image of the damage assessment report provided by a third party;
[0008] Based on preset prompt word constraints, a multimodal visual language model is invoked to perform structured recognition on the image to obtain second damage assessment data, which includes accessory information and price information identified from the image.
[0009] By comparing the first loss assessment data with the second loss assessment data, the risk items in the second loss assessment data are identified and output.
[0010] Preferably, the structured recognition of images using a multimodal visual language model includes:
[0011] The image is input into the model along with predefined cue word constraints, where the cue word constraints are used to guide the model to extract structured information from predefined fields in the image.
[0012] Preferably, the prompt word constraint includes at least one of constraint rules for identifying the vehicle identification number and constraint rules for identifying the part name.
[0013] Preferably, the first damage assessment data further includes parts code and / or repair method information, and the second damage assessment data further includes parts code and / or repair method information identified from the image.
[0014] Preferably, by comparing the first loss assessment data with the second loss assessment data, risk items in the second loss assessment data are identified and output, including:
[0015] The first loss assessment data is compared with the second loss assessment data to identify discrepancies;
[0016] Based on preset business rules, at least some of the discrepancies will be identified as risk items.
[0017] Preferably, the business rules include at least one of the following:
[0018] Parts coding consistency verification rules;
[0019] Parts price deviation threshold rules;
[0020] Matching rules between repair methods and damage levels.
[0021] Preferably, the first damage assessment data is generated by parsing a natural language description of vehicle damage using a large language model.
[0022] According to another aspect of the present invention, a risk identification device for a vehicle damage assessment form is provided, comprising:
[0023] The first acquisition unit is used to acquire first damage assessment data generated based on the first logic, the first damage assessment data including parts price information;
[0024] The second acquisition unit is used to acquire images of the damage assessment report provided by a third party.
[0025] The model invocation unit is used to invoke a multimodal visual language model to perform structured recognition on images based on preset prompt word constraints, thereby obtaining second damage assessment data. This second damage assessment data includes parts information and price information identified from the images; and
[0026] The comparison unit is used to identify and output risk items in the second loss assessment data by comparing the first loss assessment data with the second loss assessment data.
[0027] According to another aspect of the present invention, a risk identification electronic device for vehicle damage assessment reports is provided, comprising:
[0028] One or more processors; and
[0029] Storage device for storing one or more programs.
[0030] When one or more programs are executed by one or more processors, the one or more processors implement the methods described above in the embodiments of the present invention.
[0031] According to another aspect of the present invention, a computer program is provided that, when executed by a processor, implements the methods described above in the embodiments of the present invention.
[0032] According to another aspect of the present invention, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the methods described above in the embodiments of the present invention.
[0033] One embodiment of the above invention has the following advantages or beneficial effects: it can accurately identify risk items in the third party's loss assessment report, thereby bringing risk control in the loss assessment process forward to the on-site loss assessment stage, avoiding situations such as expanded damage and excessive repairs, and effectively reducing the cost of claims.
[0034] The further effects of the aforementioned unconventional alternative methods will be explained below in conjunction with specific implementation methods. Attached Figure Description
[0035] The accompanying drawings are provided to better understand the invention and are not intended to unduly limit the scope of the invention. Wherein:
[0036] Figure 1 This is a schematic diagram of the main flow of the risk identification method for vehicle damage assessment forms according to an embodiment of the present invention;
[0037] Figure 2 This is an interactive schematic diagram of generating a damage assessment preview form using natural language according to an embodiment of the present invention;
[0038] Figure 3 This is a schematic diagram of the interface guiding the standard operation process of damage assessment according to an embodiment of the present invention;
[0039] Figure 4 This is a schematic diagram of the classification and association operation of damage assessment images according to an embodiment of the present invention;
[0040] Figure 5This is a schematic diagram of the damage assessment process according to an embodiment of the present invention;
[0041] Figure 6 This is a schematic diagram of the main modules of the risk identification device for a vehicle damage assessment form according to an embodiment of the present invention;
[0042] Figure 7 This is an exemplary system architecture diagram in which embodiments of the present invention can be applied;
[0043] Figure 8 This is a schematic diagram of the structure of a computer system suitable for implementing terminal devices or servers of the present invention. Detailed Implementation
[0044] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0045] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in this disclosed technical solution all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.
[0046] To collect data on the usage of our products / services, we will aggregate, analyze, and utilize technically processed user data, and share the processed statistical information with third parties. We will use secure encryption techniques and other methods to ensure that information recipients cannot re-identify specific individuals.
[0047] Figure 1 This is a schematic diagram illustrating the main flow of the risk identification method for vehicle damage assessment reports according to an embodiment of the present invention. Figure 1 As shown, the risk identification method for vehicle damage assessment forms includes steps S101 to S104.
[0048] like Figure 1As shown, in step S101, first damage assessment data generated based on the first logic is obtained. This first damage assessment data includes parts price information. Generating the first damage assessment data based on the first logic may include generating the first damage assessment data through parsing a natural language description of vehicle damage using a large language model. Furthermore, it is not limited to generating the first damage assessment data through a model based on the natural language description described above; the first damage assessment data can also be manually entered by the user or imported from other damage assessment systems.
[0049] The following is a detailed description of how the first damage assessment data is generated through parsing using a large language model. The system receives natural language information describing the vehicle damage. This natural language information can be, for example, the language content of a user (such as a damage assessor). For instance, the receiving operation of this natural language information can support three modes: (1) Natural speech generation damage assessment form mode, receiving the vehicle damage information verbally described by the damage assessor (once); (2) Natural chat log mode, continuously listening to the voice content during the communication between the damage assessor and the repair shop personnel; and (3) Speaking and photographing mode, receiving the voice information such as the names of parts and repair methods verbally described by the damage assessor during the photographing process. The system can parse the received natural language information in real time to generate a damage assessment preview form (structured data) as described below.
[0050] For example, based on natural language information and preset prompts, a large language model can be invoked for processing to generate structured data containing at least one damage assessment item. Then, based on the damage assessment items in the structured data, a pre-set parts database can be queried to generate first damage assessment data containing price information. The structured data could be, for example, a damage assessment preview form.
[0051] Specifically, the received natural language information is converted into text information. Based on preset prompts, a large language model is invoked to perform semantic understanding on the natural language (converted text information) of an assessor, such as identifying part names, damage types, and repair methods. The data is then output in a preset format, such as structured JSON, for subsequent damage assessment processing. The large language model (general-purpose large language model) possesses the ability to understand human natural language and can be pre-loaded with basic knowledge of the automotive repair field. Examples of general-purpose large language models include qwen3-max, doubao-1-5-pro-32k-250115, and deepseek-v3.2. Prompt engineering guides the model to output structured results that meet the needs of damage assessment. Preset prompts, without changing the parameters of the large language model, can guide the model's reasoning path, behavior, and output format through systematic design and constraints on the input prompts, thereby achieving controllable, stable, and reproducible output results in specific damage assessment scenarios. Furthermore, prompts can be used to replace minor tweaks to the model.
[0052] The core calculation of large language models is based on the formula: P(y∣x)=t∏P(yt∣x,y<t). The text generation of large language models is essentially a process of gradually generating an output token sequence by maximizing the conditional probability of the next token under the condition of a given input token sequence (i.e., the prompt and the user's text), based on the internal parameters obtained from its pre-training. Therefore, the essence of prompt engineering is the artificial intervention in the model's "probability space + attention distribution + semantic activation path". The key of this invention is not to change the internal parameters of the model, but to construct the input token sequence through domain-specific prompt engineering to guide the general model to stably output a structured token sequence (such as JSON format) that conforms to a predetermined format in the vehicle damage assessment scenario.
[0053] The text information after converting natural language information can be combined with a preset prompt, and the combined information can be used as input information to be processed by a large language model. For example, according to the format defined by the preset prompt, the text information is combined to form model input information; the model input information is input into the large language model, and the preset prompt is configured to guide the large language model to output structured data that conforms to a predetermined format based on the text information. The preset prompt may include, for example, at least one of the following for guiding the large language model: role definition information, task objective description, output format specification, and standardized mapping rules for colloquial expressions in the input text.
[0054] Specifically, the prompt may include the following content: role definition information, task objective description, standardized mapping rules, background and context, process guidance, constraints and rules, and output format (output format specification).
[0055] A design example (the defined format) of the prompt is as follows:
[0056] #Parts information extraction prompts (part name, repair method, damage type)
[0057] PART_INFO_EXTRACTION_PROMPT = """You are an expert in extracting automotive repair information. (i.e., role definition)
[0058] Task (i.e., task objective): Extract the following information from the text input by the user:
[0059] 1. Part name (part_name) - retain directional words (front / rear / left / right / inside / outside, etc.)
[0060] 2. Repair Method (repair_method) - Optional values: painting, disassembly and assembly, sheet metal work, external repair, replacement
[0061] Colloquial repair method expression mapping rules:
[0062] - "Change one", "Change to a new one", "Get a new one", "Change to a new one" → Replace
[0063] - "Spray paint", "Spray one", "Spray once" → Spray paint
[0064] - "Fix", "Repair", "Repair", "Restore" → Sheet metal
[0065] 3. Damage type - Optional values: crack, deformation, scratch, fracture, damage
[0066] Colloquialism Impairment Expression Mapping Rules:
[0067] - "shattered", "damaged", "broken", "broken" → cracked
[0068] - "bent", "flattened", "bent", "flattened", "dented" → deformed
[0069] - "scratched", "rubbed", "wiped" → scraped
[0070] Rules (i.e., constraints and rules):
[0071] 1. If there are multiple parts, extract only the last one. If it is not explicitly mentioned in the text, return null. Output only the part name or null, without any other content. Output only the part name, without any explanation.
[0072] 2. The repair method and damage type must be selected from the optional values. If not explicitly mentioned in the text, return null.
[0073] 3. Prioritize identifying colloquial expressions and mapping them to standard terminology.
[0074] 4. Return plain JSON format, without any explanation or additional content.
[0075] Example:
[0076] - Input: "Headlight paint repair and scratch repair" → Output: {{"part_name": "Headlight", "repair_method": "Paint", "damage_type": "Scratching"}} (i.e., output format)
[0077] Now please extract information from the following text:
[0078] {text}
[0079] Output only JSON, no other content: ""
[0080] The following is for reference. Figure 2 The generation of the damage assessment preview form (structured data) is further described through specific examples.
[0081] In the scenario of assessing vehicle damage, the damage adjuster clicks the voice input button in the mobile terminal application (mobile system) and describes the vehicle damage in natural voice, for example, "The front bumper is deformed and needs to be replaced, the left front headlight is broken and also needs to be replaced, and the right front fender is scratched and needs to be repainted."
[0082] After the claims adjuster finishes voice input, the system converts the claims adjuster's voice content into text, and calls the general large language model (LLM) based on the predetermined prompt words (prompt word engineering) to perform semantic parsing on the text and automatically identify the claims adjustment intention contained therein (such as part name, damage type, repair method, etc.).
[0083] After model analysis, the system extracts the corresponding damage assessment items from the text and generates the following structured results:
[0084] Front bumper - Replacement - Deformed
[0085] Left front headlight - Replacement - Broken
[0086] Right front fender - paint - scratch
[0087] The system returns the following JSON data from the program:
[0088] {
[0089] {"part_name": "Front bumper", "repair_method": "Replacement", "damage_type": "Deformation"}
[0090] {"part_name": "Left front headlight", "repair_method": "replace", "damage_type": "cracked"}
[0091] {"part_name": "Right front fender", "repair_method": "paint", "damage_type": "scratching"}
[0092] }
[0093] The structured results described above are displayed to the loss assessor as a one-way preview for use in the subsequent loss assessment process.
[0094] Furthermore, based on the structured data generated above, operational guidance information corresponding to the damage assessment items can be generated and output to the user. For example, after generating a damage assessment preview form, the damage assessor can verify the damage assessment items in the preview form. The operational guidance information includes recommended shooting instructions for specified parts. Specifically, based on the damage assessment items in the structured data, the corresponding parts information and / or vehicle information can be determined; based on the parts information and / or vehicle information, the corresponding operational guidance information can be retrieved from a predetermined standard operating procedure (SOP) knowledge base. The obtained operational guidance information can be displayed to the user (damage assessor) on, for example, a mobile terminal interface to guide the user in completing standardized damage assessment operations.
[0095] The SOP knowledge base can be pre-set to include, for example, recommended shooting guidelines for specific parts (such as standard shooting positions and shooting angle requirements), fault code collection prompts for electronic parts and corresponding damage judgment criteria, and judgment rules for overall replacement, partial replacement or repair of parts, etc.
[0096] refer to Figure 3 This describes a specific example of generating operation guidance information. Based on the generated damage assessment preview form, the SOP corresponding to the damage assessment item is retrieved from the backend database (SOP knowledge base) according to the vehicle model and part name (e.g., engine hood). The SOP content includes: the photo requirement is a 45-degree side view with a close-up; the fault code prompt is to check the collision sensor; the replacement logic is that if the deformation rate is greater than 30%, replacement is recommended. The above SOP content corresponding to the engine hood is displayed on the mobile terminal.
[0097] This step generates operational guidance information and guides users, which can effectively improve the rationality and accuracy of damage assessment, prevent unnecessary replacements or repairs, and thus reduce claims risks.
[0098] In addition, media files related to vehicle damage can be acquired and associated with damage assessment items in structured data to generate damage assessment results. Media files may include, for example, videos and images. Media files may be collected based on the operation guidance information generated through the above steps. Media files may also include files collected arbitrarily by the user without following the operation guidance.
[0099] The acquired media files can be deduplicated based on feature vector similarity. Image similarity algorithms can be used to analyze the similarity of images in the acquired media file set (e.g., damage assessment images) to identify and remove highly similar redundant images. Specifically, each image is vectorized through a feature extraction network to generate a corresponding image feature vector. Then, the cosine distance between the feature vectors of any two images is calculated to measure image similarity; the smaller the cosine distance, the higher the similarity between the two images. When the detected image similarity exceeds a preset threshold, the two images are determined to be a similar image set. Further, computer vision (CV) algorithms are used to evaluate image quality indicators such as image sharpness and exposure in this similar image set. The image with the highest sharpness is automatically retained, and the rest are removed, thus achieving image deduplication and quality optimization (sharpness optimization). The deduplicated media files can then undergo the classification process described later. Alternatively, the above deduplication process can be omitted, and the acquired media files can be directly classified.
[0100] Media files can be categorized, and images belonging to the vehicle photo category can be filtered out. The media file categorization process includes identification and classification using a trained deep learning model. For example, a visual Transformer-based image classification model can be used to identify the category of damage assessment-related images. The construction and training process of the deep learning model includes the following operations (1) to (4): (1) Collect a set of sample images related to damage assessment (e.g., images related to historical damage assessment cases, with a collection of approximately 6,000 images). The collected sample images can be divided into a training set (e.g., approximately 4,000 images) and a validation set (e.g., approximately 2,000 images); (2) Manually label and classify the collected sample images. For example, they can be classified as: document images (e.g., vehicle registration certificate, driver's license, ID card, etc.), appraisal report images, accident liability determination images, vehicle photos, etc.; (3) Construct a classification model. The backbone module of the classification model is based on the open-source DINO pre-trained visual large model. The classifier head module (Classification Header Module) built based on MLP (Multilayer Perceptron) is connected as an adapter. DINO is based on ViT (Vision in Transformer). The visual base model constructed by the architecture contains, for example, about 800 million parameters, which can effectively extract image features; (4) Using the labeled sample images, on a multi-card Nvidia A100-80G GPU server, the model is trained using SGD (stochastic gradient descent) based on the PyTorch deep learning framework. In addition, various image augmentation methods can be used during model training, including random flipping, random cropping, random rotation, random adjustment of image brightness and contrast, etc., to improve the generalization ability of the model. At the same time, the training parameters (batch size, learning rate, etc.) are appropriately adjusted in multiple rounds of training, so as to obtain the model with the highest accuracy on the validation set.
[0101] By classifying and processing media files as described above, images can be quickly retrieved and located in subsequent loss assessment processes (such as review, loss verification, or compensation verification), thereby improving overall processing efficiency.
[0102] For vehicle photos filtered through classification, a trained multimodal recognition model identifies vehicle parts within them. These identified parts are then matched with the damage assessment items to establish a correlation. In other words, after image classification, the system further automatically associates vehicle photos with the parts items listed on the damage assessment report.
[0103] While existing open-source multimodal large models based on the Transformer structure possess certain automotive parts recognition capabilities, they are insufficient to directly meet the needs of damage assessment operations and suffer from the following shortcomings: (1) the types of recognizable parts are limited, making it difficult to cover high-frequency (e.g., TOP50) parts in damage assessment scenarios, especially with weak recognition capabilities for internal components (non-exterior parts); (2) the spatial positioning accuracy of parts is insufficient, easily leading to problems such as indistinguishment between left and right, or front and back. In view of the above-mentioned shortcomings of the existing technology, this invention performs secondary training on the open-source multimodal large model to obtain a trained multimodal recognition model, thereby enabling effective and accurate identification and association of vehicle parts.
[0104] Specifically, the multimodal recognition model of the present invention is, for example, a model trained using supervised fine-tuning techniques. Training the multimodal recognition model may include: acquiring a sample image set of frequently occurring parts in vehicle damage assessments; for example, collecting images of the most frequently occurring parts (e.g., the top 50) in historical damage assessment reports; labeling the sample images in the sample image set to obtain training data; wherein, for example, labeling whether a specific part exists in the sample image (constructing positive and negative example labels); and labeling the location information of the specific part in the sample image (e.g., labeling in the form of bounding boxes); wherein at least 2000 sample images can be labeled for each type of part. Thus, training data containing (e.g., more than 120,000 images) of labeled records is obtained, and this data can be divided into a training set and a validation set.
[0105] The model training also includes supervised fine-tuning of the multimodal recognition model using labeled training data. Specifically, based on the labeled training data, the original model is trained using large-scale model SFT (supervised fine-tuning) optimization techniques. First, a multimodal visual-language model based on concept features is constructed. The backbone module uses the open-source SAM3 multimodal large-scale model's visual-language encoder to extract visual and linguistic features respectively. A multilayer perceptron (MLP) is used to build a projection-fusion module for visual and linguistic features after the backbone module. Finally, an MLP is used to construct the classification module. Then, using the training data, the model is trained on a multi-GPU Nvidia A100-80G server using the PyTorch deep learning framework and employing SGD (stochastic gradient descent). During training, various image augmentation methods can be used, including random cropping, random rotation, and random adjustment of image brightness and contrast, to improve the generalization ability of the model. At the same time, training parameters (batch size, learning rate, etc.) can be appropriately adjusted in multiple rounds of training to obtain the model with the highest accuracy on the validation set, thereby generating a multimodal recognition model suitable for loss assessment business scenarios.
[0106] After associating vehicle images with parts items on the damage assessment report, the associated images can be scored and sorted when multiple images correspond to the same part. Specifically, when calling the multimodal recognition model, the location information of the vehicle part in the image can be output, represented as a bounding box. After associating media files with damage assessment items, for multiple images associated with the same vehicle part, a score is calculated based on the area ratio of the bounding box of the part in each image, and the scores are then sorted and output.
[0107] Specifically, for example, images can be scaled proportionally to a preset size range. When the multimodal recognition model for accessories is invoked, the target bounding box information of the accessory is obtained, the area ratio of the target bounding box within the central region of the image is calculated, and the image is scored based on the area ratio; the larger the area ratio, the higher the image score. Based on the scoring results, multiple images associated with the same accessory are sorted and displayed. For example, the damage assessment results can include multiple sorted associated images.
[0108] Figure 4The diagram illustrates the automatic classification and association process for damage assessment images. First, images in the original image pool (acquired media images) undergo deduplication based on vector similarity. Then, a CNN convolutional network is used to classify the deduplicated images. After classification and filtering, a multimodal large model is used to identify accessories. Finally, through a scoring mechanism, the images are automatically associated with the damage assessment items, and the preview results are displayed in sorted order by score, serving as the damage assessment result.
[0109] return Figure 1 In step S102, an image of a damage assessment report provided by a third party is acquired. In step S103, based on preset prompt word constraints, a multimodal visual language model is invoked to perform structured recognition on the image to obtain second damage assessment data. The second damage assessment data includes parts information and price information identified from the image. In step S104, by comparing the first damage assessment data and the second damage assessment data, risk items in the second damage assessment data are identified and output. The first damage assessment data may also include parts codes and / or repair method information, and the second damage assessment data may also include parts codes and / or repair method information identified from the image. The repair method information in the first damage assessment data can be obtained through the aforementioned operation guidance information.
[0110] Specifically, after confirming the damage assessment items, the system automatically retrieves the corresponding part codes, prices, and repair hours from the backend parts database based on the part names and repair methods in the damage assessment preview form, generating an initial damage assessment form (first damage assessment data) containing price information. The damage assessor can then scan a damage assessment form (second damage assessment data, such as a computer-generated or handwritten form) provided by a third party (e.g., a repair shop) using a mobile terminal. The system uses a multimodal visual language model to perform structured recognition on the damage assessment form image, extracting information such as part names, part codes, repair methods, and prices. This information is then automatically compared with the initial damage assessment form to identify discrepancies and potential risks, thus assisting the damage assessor in accurately assessing vehicle damage (e.g., on-site negotiation and price negotiation).
[0111] Multimodal visual language models include qwen-vl-ocr (Qianwen OCR-VL multimodal model), qwen3-vl-plus (Qianwen 3-VL Plus multimodal model), gemini-3-pro-preview (Google Gemini3 multimodal large model), and gpt-5.2 (OpenAI gpt5.2 multimodal large model), etc.
[0112] Multimodal visual language models can perform structured image recognition by inputting images of third-party damage assessment reports along with pre-defined cue word constraints. These constraints guide the model to extract structured information from predefined fields in the image. Cue word constraints may include, for example, at least one of constraints for recognizing vehicle identification numbers (VINs) or for recognizing part names. Since damage assessment reports from different third parties (repair shops) vary significantly in format, especially handwritten reports which often lack fixed headers and are poorly written, pre-designing the cue words and providing clear feature constraints for various recognition elements before model invocation can improve the model's compatibility and recognition accuracy across different types of damage assessment reports.
[0113] Examples of prompt word constraints are as follows: For the Vehicle Identification Number (VIN) in the vehicle identification information, the following constraints are defined: the VIN consists of 17 letters and numbers, excluding the letters I, O, and Q, and is typically labeled in the damage assessment report image as "VIN", "Vehicle Identification Number", "Chassis Number", or "VIN" fields; For the part name, the following constraints are defined: the part name is characterized by a string composed of Chinese characters representing the vehicle part name, corresponding to fields such as "Materials Used", "Part Name", "Part Name", or "Description" in the table, but excluding the part code field.
[0114] Through constraints such as the aforementioned cue words, the multimodal visual language model outputs recognition results according to, for example, a pre-defined JSON data structure, without outputting other irrelevant content, thereby obtaining structured loss assessment data (second loss assessment data).
[0115] The prompt words and the image of the third-party damage assessment report are input into the multimodal visual language model. The model is invoked in the following way:
[0116] # Calling OpenAI compatible API
[0117] response = client.chat.completions.create(
[0118] model=config['model'],
[0119] messages=[
[0120] {
[0121] "role": "user",
[0122] "content": [
[0123] {
[0124] "type": "text",
[0125] "text": prompt
[0126] },
[0127] {
[0128] "type": "image_url",
[0129] "image_url": {
[0130] "url": f"data:{mime_type};base64,{base64_image}"
[0131] }
[0132] } ]
[0134] }
[0135] ],
[0136] max_tokens=8192 )
[0138] Here, prompt is a pre-designed prompt word, base64_image is the base64 encoding of the damage assessment form, and model is the name of the model, such as qwen-vl-ocr.
[0139] Here is an example of the model returning JSON:
[0140] {
[0141] {"name":"Front Bumper Removal and Installation", "price":"3276", "type":"labor", "work_type":"removal and installation", "part_code":""}
[0142] {"name":"Right front fender painting", "price":"8000", "type":"labor", "work_type":"painting", "part_code":""}
[0143] {"name":"Middle Frame", "price":"27500", "type":"part", "work_type":"Replacement","part_code":"52137308754"}
[0144] {"name":"Front Bumper Upper Body", "price":"964", "type":"part", "work_type":"Replacement", "part_code":"51117379434"}
[0145] {"name":"Left Front Fog Light", "price":"1200", "type":"part", "work_type":"Replacement", "part_code":"63177248911-A1"}
[0146] }
[0147] In addition, determining risk items by comparing the first loss assessment data with the second loss assessment data may include: comparing the first loss assessment data with the second loss assessment data to identify discrepancies; and, based on preset business rules, identifying at least a portion of the discrepancies as risk items.
[0148] Business rules may include at least one of the following: Parts code consistency verification rule, whether the parts codes in the first damage assessment data and the second damage assessment data are consistent; Parts price deviation threshold rule, whether the price deviation of parts in the first damage assessment data and the second damage assessment data exceeds a predetermined threshold; Repair method and damage degree matching rule, whether the repair method and damage degree of the first damage assessment data and the second damage assessment data match. Risk items may include, for example: Parts code risk, such as the parts code on the repair shop's damage estimate differing from the actual VIN code, potentially leading to under-insurance and over-equipment risk; Parts price risk, such as the parts price on the repair shop's damage estimate being higher than the official parts price collected from the price database in the damage assessment, or the price difference exceeding a predetermined value; Parts replacement standard risk, such as parts requiring replacement in the repair shop's damage estimate, while the damage assessment, based on the SOP process, determines they are repairable; Over-repair risk due to complete parts replacement, such as all parts requiring replacement in the repair shop's damage estimate, while the damage assessment, based on the SOP process and the composition of parts in the assembly, determines they can be partially replaced.
[0149] The risk items identified by comparing the first and second loss assessment data can be output and displayed to the user. After editing or adjusting the loss assessment information based on the risk items, a loss assessment report is generated. The aforementioned image association operation is then performed based on the loss assessment items in this report.
[0150] The risk identification method for vehicle damage assessment reports according to embodiments of the present invention can accurately identify risk items in third-party damage assessment reports, thereby bringing risk control in the damage assessment process forward to the on-site damage assessment stage, avoiding situations such as expanded damage and excessive repairs, and effectively reducing the cost of claims.
[0151] Figure 5A schematic flowchart of a vehicle damage assessment method according to an embodiment of the present invention is shown. Figure 5 As shown, the damage assessment process includes four stages. Stage 1: Generating a preview form using natural language. Natural language input includes three modes: natural voice input, natural chat mode, and simultaneous setup and shooting mode. The natural language information input in these modes is converted into text information. A large language model is used to identify the assessor's intentions, and natural language processing (NLP) is used to standardize the identified part names using word segmentation and retrieval techniques, thus generating a damage assessment preview form. Stage 2: Standard Operating Procedure (SOP)-guided generation of the damage assessment form. Information collection is guided by an SOP knowledge base. Stage 3: Valuation form identification and comparison with the damage assessment form. An initial damage assessment form is generated through querying. A multimodal large model is used to identify the valuation form image, obtaining the identified valuation form. The two are compared to highlight risks, and the final damage assessment form is edited based on these risks. After deduplication and classification of the collected images, a supervised optimization model is used to identify the images and associate them with the assessed parts, thus obtaining the damage assessment result.
[0152] The vehicle damage assessment method according to embodiments of the present invention can reduce the time spent on manual data entry, querying, and comparison by introducing intelligent assistance from large models, thereby accelerating the damage assessment process and improving the efficiency of vehicle damage assessment. Through standardized and structured processing, the completeness and consistency of damage assessment information collection are improved, reducing human error and thus increasing the accuracy of damage assessment. Through natural interaction, intelligent guidance, and automatic organization and association of damage assessment images, the workload of damage assessors is reduced, improving the efficiency of subsequent damage verification and providing reliable evidence. Furthermore, through SOP guidance and a difference comparison mechanism, risk control in the damage assessment process is brought forward to the on-site damage assessment stage, avoiding expanded damage and excessive repairs, effectively reducing compensation costs. Moreover, it provides complete, standardized, and traceable electronic damage assessment evidence for insurance claims.
[0153] Figure 6 This is a schematic diagram of the main modules of the risk identification device for a vehicle damage assessment form according to an embodiment of the present invention. Figure 6 As shown, the risk identification device 600 for vehicle damage assessment includes: a first acquisition unit 601, a second acquisition unit 602, a model calling unit 603, and a comparison unit 604.
[0154] The first acquisition unit 601 acquires first damage assessment data generated based on the first logic. This first damage assessment data includes parts price information. Specifically, the first damage assessment data is generated by parsing a natural language description of vehicle damage using a large language model.
[0155] The second acquisition unit 602 acquires the image of the damage assessment report provided by a third party.
[0156] The model calling unit 603 calls a multimodal visual language model to perform structured recognition on the image based on preset prompt word constraints, and obtains second damage assessment data. The second damage assessment data includes accessory information and price information identified from the image.
[0157] The model invocation unit 603 inputs the image and preset prompt word constraints into the model, where the prompt word constraints are used to guide the model to extract structured information of predefined fields from the image.
[0158] The prompt word constraint includes at least one of the constraint rules for identifying the vehicle identification number and the constraint rules for identifying the part name.
[0159] The first damage assessment data also includes part codes and / or repair method information, and the second damage assessment data also includes part codes and / or repair method information identified from the images.
[0160] The comparison unit 604 compares the first damage assessment data with the second damage assessment data to identify discrepancies; based on preset business rules, at least a portion of the discrepancies are identified as risk items. The business rules include at least one of the following: parts code consistency verification rules; parts price deviation threshold rules; and repair method and damage severity matching rules.
[0161] The risk identification device for vehicle damage assessment reports according to embodiments of the present invention can accurately identify risk items in third-party damage assessment reports, thereby bringing risk control in the damage assessment process forward to the on-site damage assessment stage, avoiding situations such as expanded damage and excessive repairs, and effectively reducing the cost of claims.
[0162] Figure 7 An exemplary system architecture 700 is shown, in which the risk identification method or device for vehicle damage assessment forms can be applied according to embodiments of the present invention.
[0163] like Figure 7 As shown, system architecture 700 may include terminal devices 701, 702, and 703, a network 704, and a server 705. Network 704 serves as the medium for providing communication links between terminal devices 701, 702, and 703 and server 705. Network 704 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0164] Users can use terminal devices 701, 702, and 703 to interact with server 705 via network 704 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 701, 702, and 703.
[0165] Terminal devices 701, 702, and 703 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0166] Server 705 can be a server that provides various services, such as a back-end management server (for example only) that supports interactive operations performed by users using terminal devices 701, 702, and 703. The back-end management server can analyze and process data such as received product information query requests, and feed back the processing results (such as operation guidance information - for example only) to the terminal devices.
[0167] It should be noted that the risk identification method for vehicle damage assessment forms provided in this embodiment of the invention is generally executed by server 705, and correspondingly, the risk identification device for vehicle damage assessment forms is generally set in server 705.
[0168] It should be understood that Figure 7 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0169] The following is for reference. Figure 8 It shows a schematic diagram of the structure of a computer system 800 suitable for implementing a terminal device of the present invention. Figure 8 The terminal device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0170] like Figure 8 As shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 802 or programs loaded from storage section 808 into random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the system 800. The CPU 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0171] The following components are connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 810 as needed so that computer programs read from it can be installed into storage section 808 as needed.
[0172] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 809, and / or installed from removable medium 811. When the computer program is executed by central processing unit (CPU) 801, it performs the functions defined above in the system of this invention.
[0173] It should be noted that the computer-readable medium shown in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0174] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0175] The units described in the embodiments of the present invention can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a first acquisition unit, a second acquisition unit, a model invocation unit, and a comparison unit. The names of these units do not necessarily limit the specific unit; for example, the second acquisition unit may also be described as "a unit for acquiring images of damage assessment reports provided by a third party."
[0176] In another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs, which, when executed by the device, cause the device to: acquire first damage assessment data generated based on first logic, the first damage assessment data including parts price information; acquire an image of a damage assessment report provided by a third party; based on preset prompt word constraints, invoke a multimodal visual language model to perform structured recognition on the image to obtain second damage assessment data, the second damage assessment data including parts information and price information identified from the image; and identify and output risk items in the second damage assessment data by comparing the first damage assessment data with the second damage assessment data.
[0177] According to the technical solution of the present invention, risk items in the third party's loss assessment report can be accurately identified, thereby bringing risk control in the loss assessment process forward to the on-site loss assessment stage, avoiding situations such as expanded damage and excessive repairs, and effectively reducing the cost of claims.
[0178] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for risk identification in vehicle damage assessment reports, characterized in that, include: Obtain first damage assessment data generated based on the first logic, which includes parts price information; Obtain an image of the damage assessment report provided by a third party; Based on preset prompt word constraints, a multimodal visual language model is invoked to perform structured recognition on the image to obtain second damage assessment data, which includes accessory information and price information identified from the image. By comparing the first loss assessment data with the second loss assessment data, the risk items in the second loss assessment data are identified and output.
2. The method according to claim 1, characterized in that, The structured recognition of the image using a multimodal visual language model includes: The image and the preset cue word constraints are input into the model, wherein the cue word constraints are used to guide the model to extract structured information of predefined fields from the image.
3. The method according to claim 2, characterized in that, The prompt word constraint includes at least one of the constraint rules for identifying the vehicle identification number and the constraint rules for identifying the part name.
4. The method according to claim 1, characterized in that, The first damage assessment data also includes parts codes and / or repair method information, and the second damage assessment data also includes parts codes and / or repair method information identified from the image.
5. The method according to claim 1, characterized in that, By comparing the first loss assessment data with the second loss assessment data, risk items in the second loss assessment data are identified and output, including: The first damage assessment data is compared with the second damage assessment data to identify discrepancies; Based on preset business rules, at least a portion of the discrepancies are identified as risk items.
6. The method according to claim 5, characterized in that, The business rules include at least one of the following: Parts coding consistency verification rules; Parts price deviation threshold rules; Matching rules between repair methods and damage levels.
7. The method according to claim 1, characterized in that, The first damage assessment data is generated by parsing a large language model based on a natural language description of vehicle damage.
8. A risk identification device for vehicle damage assessment reports, characterized in that, include: The first acquisition unit is used to acquire first damage assessment data generated based on the first logic, the first damage assessment data including parts price information; The second acquisition unit is used to acquire images of the damage assessment report provided by a third party. The model invocation unit is used to invoke a multimodal visual language model to perform structured recognition on the image based on preset prompt word constraints, and obtain second damage assessment data, which includes accessory information and price information identified from the image; as well as The comparison unit is used to identify and output the risk items in the second loss assessment data by comparing the first loss assessment data with the second loss assessment data.
9. An electronic device for risk identification of vehicle damage assessment reports, characterized in that, include: One or more processors; as well as Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.
10. A computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-7.
11. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.