Waybill identification method and computer readable storage medium

By combining multimodal large models and custom prompt words, the problems of low efficiency of manual data entry and difficulty of OCR in adapting to complex scenarios in waybill recognition are solved. This enables efficient and accurate recognition and structured data entry of waybill information, thereby improving the level of information management at construction sites.

CN120877294AActive Publication Date: 2025-10-31GLODON CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511004014.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-10-31
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

In the management of material entry and exit in the construction industry, existing technologies rely on manual input of waybill information, which is inefficient and prone to errors. General OCR technology is difficult to implement effectively in complex engineering scenarios and cannot adapt to the characteristics of waybill documents, such as non-fixed format and large field variability.

Method used

By employing a multimodal large model combined with custom or built-in prompts, and through image quality assessment, corner detection, and rotation correction, field names are automatically labeled. The multimodal large model is used to identify key field information in waybill images, achieving semantic understanding and extraction without relying on templates.

Benefits of technology

It significantly improves the robustness and generalization of waybill recognition, reduces manual data entry, provides accurate and timely material information, supports the recognition and structured entry of diverse form structures, and promotes the informatization and intelligent management of construction sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877294A_ABST
    Figure CN120877294A_ABST
Patent Text Reader

Abstract

The invention discloses a waybill identification method and a computer readable storage medium. The method comprises the following steps: acquiring a target waybill picture; determining a selection mode of a prompt word corresponding to the target waybill picture; when the selected mode is a user-defined mode, marking a plurality of field names in the target waybill picture; determining a to-be-identified key field name in the target waybill picture based on the marked field name; on the basis of the key field name, constructing a prompt word corresponding to the target waybill picture; and jointly inputting the target waybill picture and the cue word corresponding to the target waybill picture into a preset multi-modal large model, so that the multi-modal large model identifies a field corresponding to the key field name in the target waybill picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of waybill recognition, and particularly to a waybill recognition method and a computer-readable storage medium. Background Art

[0002] In the process of material in-and-out management in the construction industry, the accurate collection and efficient entry of waybill information are the key foundations for ensuring the timely supply of materials, controlling construction costs, and realizing the full-process traceability of materials. However, in current practical applications, the technical means in this link are still relatively weak and mainly face the following challenges: (1) Dependence on manual entry, with low efficiency and easy to make mistakes.

[0003] Currently, material clerks often need to manually enter multiple fields in the waybill (such as material name, specification model, quantity, etc.) item by item. Especially in complex list scenarios with "multiple materials per vehicle" or "multiple-field reuse", manual operations not only have high labor intensity and high repetition, but are also extremely prone to problems such as information omission and entry errors, seriously affecting the accuracy and integrity of data, and increasing the subsequent verification and error correction costs.

[0004] (2) General OCR (Optical Character Recognition) technology is difficult to be effectively implemented in complex engineering scenarios.

[0005] In existing waybill recognition methods based on OCR technology, they usually include template matching method and template-free recognition method. The template matching method is: pre-construct various waybill templates and extract fields through image matching; the template-free recognition method is: combine OCR text detection and NLP (Neuro-Linguistic Programming) technology to directly recognize and extract information in the field area. Although such methods have certain application value in general bill scenarios, they still face significant bottlenecks in the recognition of complex waybills at the construction site: Waybill documents in actual scenarios often have characteristics such as non-fixed layout and large field variability, and traditional template-based extraction methods are difficult to adapt to diverse form structures; most fields are optional items, and the content lacks structural constraints, making it difficult to establish a standard recognition process.

[0006] For the above problems, there is currently no effective solution. [[ID=#]]Summary of the Invention

[0007] The purpose of the present invention is to provide a waybill recognition method and a computer-readable storage medium to solve the above technical problems.

[0008] According to one aspect of the present invention, a waybill recognition method is provided, and the method includes: Obtain a target waybill picture; Determine the selection method for the prompt words corresponding to the target waybill image; When the selected method is a custom method, multiple field names are marked in the target waybill image; Based on the labeled field names, determine the key field names to be identified in the target waybill image; Based on the key field names, construct the prompt words corresponding to the target waybill image; The target waybill image and the corresponding prompt words are input into a preset multimodal big data model so that the multimodal big data model can identify the field corresponding to the key field name in the target waybill image.

[0009] Optionally, obtaining the target waybill image includes: Get the images of the waybills to be processed; Attempt to detect the four corner points of the waybill in the waybill image; When four corner points are detected and determined to be the four corner points of the waybill, the image quality of the waybill image is evaluated to see if it meets the preset quality requirements. When the image quality of the waybill image is evaluated to meet the preset quality requirements, the waybill image is rotated and corrected based on the four detected corner points so that the horizontal center axis of the waybill in the rotated and corrected waybill image is horizontal. The image size of the rotated and corrected waybill image is scaled to a preset image size to obtain the target waybill image.

[0010] Optionally, when the image quality of the waybill image is evaluated to meet the preset quality requirements, the waybill image is rotated and corrected based on the detected four corner points so that the horizontal center axis of the waybill in the rotated and corrected waybill image is horizontal, including: The four detected corner points are used as the four corner points of the rectangle, and the rectangle is drawn. Scale the rectangle and / or the waybill image so that the size of the waybill in the waybill image matches the size of the rectangle. The waybill within the rectangle is rotated and corrected based on affine transformation or perspective transformation so that the horizontal center axis of the rotated and corrected waybill is horizontal.

[0011] Optionally, when the selected method is a custom method, multiple field names are marked in the target waybill image, including: Based on the multimodal large model, all field names belonging to the detail category in the target waybill image are identified; Based on a preset OCR recognition model, all characters in the target waybill image and the position information of each character are identified; Match the field names identified by the multimodal large model with the characters identified by the OCR recognition model; Assign the position information of the successfully matched characters to the corresponding field name; Based on the location information of each field name, a label box for each field name is drawn in the target waybill image.

[0012] Optionally, determining the key field names to be identified in the target waybill image based on the labeled field names includes: Receive user selection instructions for annotation boxes, mark the selected annotation box in the target waybill image, and extract the field names within the marked annotation box; and The field names within the marked annotation boxes will be used as the key field names to be identified in the target waybill image; or, The system receives user input instructions for field names in a general category, parses out the field names carried in the input instructions, and uses the field names within the marked annotation boxes and the field names carried in the input instructions as the key field names to be identified in the target waybill image.

[0013] Optionally, when the selected method is a custom method, multiple field names are marked in the target waybill image, including: When the selected method is a custom method and there is no historical template associated with the supplier / shipper unit of the target waybill image, multiple field names are marked in the target waybill image; wherein, the historical template includes historical key field names; After determining the key field names to be identified in the target waybill image based on the labeled field names, the method further includes: The identified key field names are divided into several field name groups according to their respective field categories; wherein, the field category is a detailed category or a general category, and each field name group includes a field category and field names belonging to that field category; The segmented field names are grouped and stored as templates associated with the supplier / shipper unit of the target waybill image.

[0014] Optionally, when the selected method is a custom method and there is no historical template associated with the supplier / shipper unit of the target waybill image, multiple field names are marked in the target waybill image, including: When the selected method is a custom method and there is no historical template, multiple field names are marked in the target waybill image; When the selected method is a custom method and there are several historical templates, multiple optional buttons are displayed, and when the first button is selected, multiple field names are marked in the target waybill image; wherein, the multiple optional buttons include: a first button indicating that a historical template is not used and a second button corresponding to each historical template.

[0015] Optionally, the method further includes: When the second button is selected, retrieve the historical template corresponding to the selected second button; The historical key field names contained in the retrieved historical templates will be used as the key field names to be identified in the target waybill image.

[0016] Optionally, the method further includes: When the selected method is the built-in method, obtain the preset prompt words; The target waybill image and pre-set prompts are input into the multimodal big data model, so that the multimodal big data model can identify the required fields in the target waybill image based on the pre-set prompts.

[0017] To achieve the above objectives, the present invention also provides a waybill identification device, the device comprising: The acquisition module is used to acquire the target waybill image; The first determining module is used to determine the selection method of the prompt word corresponding to the target waybill image; The annotation module is used to annotate multiple field names in the target waybill image when the selected method is a custom method; The second determining module is used to determine the key field names to be identified in the target waybill image based on the marked field names; The module is used to construct prompt words corresponding to the target waybill image based on the key field names; The recognition module is used to input the target waybill image and the corresponding prompt words into a preset multimodal big data model, so that the multimodal big data model can identify the field corresponding to the key field name in the target waybill image.

[0018] To achieve the above objectives, the present invention also provides a computer device, the computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the waybill identification method described above.

[0019] To achieve the above objectives, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is used to implement the steps of the waybill identification method described above.

[0020] This invention provides a waybill recognition method and a computer-readable storage medium. Considering the characteristics of waybill documents in real-world scenarios, such as non-fixed formatting and high field variability, it proposes a field semantic structuring method based on a multimodal large model. Through customized prompt design, the large model is guided to recognize core field information in the image. This method does not rely on templates, possesses superior semantic understanding and extraction capabilities, significantly improves the robustness and generalization of recognition in non-template documents, and solves the problem that traditional template-based extraction methods are difficult to adapt to diverse form structures. Simultaneously, through automated and structured input of material information, it provides accurate and timely raw data support for project material flow, supply chain analysis, and construction management, promoting the informatization and intelligentization of construction sites. Attached Figure Description

[0021] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart of the waybill identification method provided in Example 1; Figure 2 This is a schematic diagram of the waybill identification scheme provided in Example 1; Figure 3 A block diagram of the waybill recognition device provided in Embodiment 2; Figure 4 A block diagram of a computer device suitable for implementing the waybill identification method, provided in Embodiment 3. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0023] Example 1 Embodiment 1 of the present invention provides a waybill identification method, such as Figure 1 As shown, the method includes steps S1 to S6, wherein: Step S1: Obtain the target waybill image.

[0024] The target waybill image is an image that can be directly recognized by the multimodal large model. The target waybill image contains a complete waybill and the image quality meets the recognition requirements.

[0025] Traditional OCR systems passively recognize images, lacking the ability to actively assess image quality (such as sharpness, crop integrity, tilt angle, and exposure). They cannot provide interactive feedback to users regarding image quality, leading to inconsistent image quality, frequent issues like blurriness, missing edges, and information loss, thus affecting the stability and accuracy of recognition results. To address this, this embodiment introduces an image quality assessment module to comprehensively analyze captured images, covering multiple dimensions including blur detection, crop integrity assessment, tilt angle estimation, and exposure analysis. Based on the analysis results, the system can provide real-time feedback to users on image quality issues (e.g., "Image is blurry, please retake the shot," "Edges are missing, please adjust the shooting angle"), enabling pre-emptive control of image quality and ensuring the accuracy of subsequent structured recognition from the source.

[0026] Optionally, step S1 includes: Get the images of the waybills to be processed; Attempt to detect the four corner points of the waybill in the waybill image; When four corner points are detected and determined to be the four corner points of the waybill, the image quality of the waybill image is evaluated to see if it meets the preset quality requirements. When the image quality of the waybill image is evaluated to meet the preset quality requirements, the waybill image is rotated and corrected based on the four detected corner points so that the horizontal center axis of the waybill in the rotated and corrected waybill image is horizontal. The image size of the rotated and corrected waybill image is scaled to a preset image size to obtain the target waybill image.

[0027] The waybill images to be processed can be waybill images taken by users through mobile terminals (such as mobile phones), or waybill images already taken and selected from the image library. The "photo-to-recognition" mode lowers the usage threshold for frontline personnel, simplifies the operation process, reduces training and maintenance costs, and helps the project to go live quickly and be promoted on a large scale.

[0028] After obtaining the image of the waybill to be processed, such as Figure 2As shown, the integrity of the waybill image is first checked: an attempt is made to detect the four corner points of the waybill in the image. When four corner points are detected and determined to be the four corner points of the waybill, the waybill image is considered to contain a complete waybill. The detected corner points follow a certain order, such as counter-clockwise or clockwise. The corner detection model used in this embodiment is based on the Gaussian heatmap mechanism, belonging to a single-stage object detection framework, with a typical representative algorithm being CenterNet. The model aims to accurately locate the four corner points of the waybill in the image (e.g., in the order of top left, top right, bottom right, bottom left) to achieve subsequent rotation correction.

[0029] Furthermore, such as Figure 2 As shown, if the waybill image is determined to be complete, the image quality of the waybill image is evaluated to see if it meets the preset quality requirements. This embodiment uses the BRISQUE (Blind / Referenceless Image Spatial Quality Evaluator) algorithm, a no-reference image quality assessment algorithm, to score the image's sharpness, noise, and other quality indicators. The system feeds back the quality assessment results to the user interface. If the image quality score is lower than a preset threshold, the user is prompted that the image quality is poor and is advised to retake the photo or reselect the image to ensure the accuracy of subsequent field extraction. BRISQUE is a no-reference image quality assessment algorithm, meaning it does not rely on an "ideal image" as a comparison benchmark, but only scores the quality based on the current image itself. Its core idea is as follows: Utilizing the natural scene statistics (NSS) of the image, the statistical distribution deviation of the image in the spatial domain is analyzed; the brightness normalization statistical features of the image patches (such as mean, variance, slope, etc.) are extracted; and a trained SVM (Support Vector Machine) regression model is used to output an image quality score from 0 to 100, with a lower score indicating a sharper image.

[0030] The image quality assessment process based on BRISQUE is as follows: The system preprocesses the rotated image (e.g., converting it to grayscale and normalizing it); it then uses the BRISQUE model to extract statistical features and calculate a quality score; the score is compared to a preset quality threshold; if the score is lower than the threshold, the system prompts: "Image quality is poor, we suggest retaking the image or selecting a new one." The interface simultaneously displays the rotated image and its corresponding quality score, allowing users to intuitively understand the image quality. If the image quality meets the standards, the user is allowed to proceed to the subsequent information structuring process.

[0031] Furthermore, such as Figure 2As shown, if the quality of the target waybill image meets the requirements, image rotation correction is performed based on the four detected corner points. The purpose of rotation correction is to straighten the waybill within the target waybill image, not to straighten the target waybill image itself.

[0032] Optionally, when the image quality of the waybill image is evaluated to meet the preset quality requirements, the waybill image is rotated and corrected based on the detected four corner points so that the horizontal center axis of the waybill in the rotated and corrected waybill image is horizontal, including: The four detected corner points are used as the four corner points of the rectangle, and the rectangle is drawn. Scale the rectangle and / or the waybill image so that the size of the waybill in the waybill image matches the size of the rectangle. The waybill within the rectangle is rotated and corrected based on affine transformation or perspective transformation so that the horizontal center axis of the rotated and corrected waybill is horizontal.

[0033] This embodiment first uses a key point detection model to locate the waybill in the image, detecting the positions of the four corner points of the waybill, thereby achieving automatic selection of the waybill area. After the initial selection is completed, the area will be displayed on the user interface. The user can manually adjust the selected area according to the actual situation, such as scaling the rectangle and / or the waybill image, so that the size of the waybill in the waybill image matches the size of the rectangle, ensuring the completeness and accuracy of the waybill content.

[0034] The process of rotating and correcting the waybill image is as follows: 1. Corner point acquisition Based on the key point detection module, the four corner points of the waybill are output in sequence to construct a perspective coordinate system.

[0035] 2. Calculation of affine or perspective transformations Construct a quadrilateral region in the source image based on the corner coordinates; Map it to a regular rectangular area (usually A4 scale or the standard size set by the model). Use perspective transformation (cv2.getPerspectiveTransform+cv2.warpPerspective) to rotate and straighten the image so that the list is horizontal.

[0036] 3. Rotational correction output Output the rotated image and pass it to the quality assessment module.

[0037] like Figure 2 As shown, after corner correction, the rotated waybill image is scaled to make its size consistent with the preset image size.

[0038] With a built-in image quality detection module, the system can determine in real time whether the waybill photo is blurry or incompletely cropped, and guide the user to retake the photo in a timely manner to ensure the quality of the input image. This improves the overall recognition accuracy and robustness from the source, ensuring the accuracy and stability of image recognition.

[0039] Step S2: Determine the selection method for the prompt word corresponding to the target waybill image.

[0040] The selection methods include: custom method and built-in method. The custom method involves setting prompts based on the user's personalized needs after interaction; the built-in method uses pre-set prompts. The prompts corresponding to the target waybill image are used to enable the multimodal large model to identify the corresponding fields in the target waybill image based on these prompts.

[0041] Step S3: When the selected method is the custom method, mark multiple field names in the target waybill image.

[0042] Step S4: Determine the key field names to be identified in the target waybill image based on the labeled field names.

[0043] The labeled field names can be field names belonging to the detail category and / or field names belonging to the general category in the waybill. Fields within the table of the waybill belong to the detail category, while fields outside the table belong to the general category. Each field includes a field name and a field value. Labeling multiple field names can be done by labeling some or all field names, such as labeling all field names in the detail category. Some or all field names can be selected from the labeled field names as the key field names to be identified in the target waybill image.

[0044] Optionally, step S3 includes: Based on the multimodal large model, all field names belonging to the detail category in the target waybill image are identified; Based on a preset OCR recognition model, all characters in the target waybill image and the position information of each character are identified; Match the field names identified by the multimodal large model with the characters identified by the OCR recognition model; Assign the position information of the successfully matched characters to the corresponding field name; Based on the location information of each field name, a label box for each field name is drawn in the target waybill image.

[0045] The annotation can be a label box drawn at the position of the field name so that the field name is located within the corresponding label box.

[0046] When recognizing field names based on a multimodal large model, pre-built prompt words are used for recognition. See the example below for the recognition process.

[0047] Please identify and extract the field names describing the product details from the sales order image, and return these fields in row order. The result should be a JSON array, with array elements arranged according to the actual order of the product details table.

[0048] [Task Description] 1. Extract the field names / column names corresponding to the product details section in the table. These are usually located in the first row or near the top of the table, but not necessarily in the first row; they may also be in the middle.

[0049] 2. Ensure that field names are arranged in natural order (e.g., the first row is “Material Name”, “Specification”, “Quantity”, etc.) and avoid duplicate fields.

[0050] 3. Do not extract the specific data from the table, only keep the field names.

[0051]

Processing Rules

[0052] 2. Each field is kept only once, and duplicate fields are removed.

[0053] 3. It does not return explanatory text or additional content, but only outputs the JSON array itself.

[0054] 4. Do not include items in the table that are not related to the product details.

[0055] [Output Example] [ "Material Name" "Specifications and Models" "quantity", "unit price", "Amount", "Remark" ] Based on the above output example, the fields of material name, specifications, quantity, unit price, amount, and remarks can be marked on the target waybill image.

[0056] This embodiment utilizes OCR technology combined with multimodal large model capabilities to automatically identify key field information in waybill images, replacing traditional manual entry, greatly reducing repetitive work for material handlers, and significantly improving the efficiency of waybill information entry.

[0057] Different enterprises or projects have varying requirements for field types and structures. Fixed field extraction methods lack flexibility and interactivity, limiting the system's adaptability. This embodiment uses a configurable and interactive field selection mechanism, combined with a lightweight OCR model, to extract image text and its location information. A visual interface binding images and fields is provided on the front end, allowing users to manually adjust field areas and add or delete field content, enabling flexible definition and dynamic configuration of the structured recognition process, as detailed below.

[0058] Optionally, step S4 includes: Receive user selection instructions for annotation boxes, mark the selected annotation box in the target waybill image, and extract the field names within the marked annotation box; and The field names within the marked annotation boxes will be used as the key field names to be identified in the target waybill image; or, The system receives user input instructions for field names in a general category, parses out the field names carried in the input instructions, and uses the field names within the marked annotation boxes and the field names carried in the input instructions as the key field names to be identified in the target waybill image.

[0059] In this embodiment, the user can choose to use the built-in prompt or a custom prompt depending on the specific usage scenario. If a custom prompt is selected, the system, based on image rotation correction, calls a multimodal large model to process the image and extract the possible list field keywords (Keys), i.e., field names belonging to the detail category, from the image. The general prompt used in this step guides the model to identify standard fields in the list, such as key field names like "material name," "quantity," and "unit price." To further enhance the user's control over the field extraction results, the system provides field interaction functionality. Specifically, based on the rotated image, the system calls a lightweight OCR model to perform full-area character recognition on the image, obtaining all character content and their corresponding bounding boxes. Combining the list field Key content identified by the aforementioned multimodal large model, the system performs positional and semantic matching between the two, presenting the recognition results on the interface in the form of "field name + location information." The user can select the field names to retain by clicking. For field names that are not automatically recognized but need to be extracted in actual business, i.e., general-type field names, the user can manually enter the field names.

[0060] Optionally, step S3 includes: When the selected method is a custom method and there is no historical template associated with the supplier / shipper unit of the target waybill image, multiple field names are marked in the target waybill image; wherein, the historical template includes historical key field names; After determining the key field names to be identified in the target waybill image based on the labeled field names, the method further includes: The identified key field names are divided into several field name groups according to their respective field categories; wherein, the field category is a detailed category or a general category, and each field name group includes a field category and field names belonging to that field category; The segmented field names are grouped and stored as templates associated with the supplier / shipper unit of the target waybill image.

[0061] The storage rules for historical templates are consistent with the rules for grouping templates by field name as described above. If a historical template exists in the system but it is not associated with the supplier / shipper unit of the target waybill image, the historical template cannot be used to determine the prompt words. Field names selected by the user within the annotation box and / or field names of general categories entered by the user are added to the current template. All user-confirmed field keys and their corresponding information will form a list of field templates. Users can name and save this template for quick extraction of similar document types in the future.

[0062] Example of template storage:

Detailed Categories

[0063] Optionally, the method further includes: When the second button is selected, retrieve the historical template corresponding to the selected second button; The historical key field names contained in the retrieved historical templates will be used as the key field names to be identified in the target waybill image.

[0064] When no historical templates exist, after the user selects the custom method, the backend calculates and directly displays the target waybill image labeled with multiple field names on the client. When historical templates exist but are not available, selecting the first button initiates the backend calculation process and displays the target waybill image labeled with multiple field names on the client. "No available historical templates" includes: no historical template associated with the supplier / shipper unit of the target waybill image; or, a historical template associated with the supplier / shipper unit of the target waybill image exists, but a new template needs to be created after the recognition requirements change.

[0065] Each second button uniquely corresponds to a historical template. Based on the supplier / shipper unit of the target waybill image, the available historical templates are determined from these templates, and the second button corresponding to the available historical template is selected.

[0066] This embodiment supports user-defined waybill field structures. By combining a lightweight OCR model with large model inference, it provides field location feedback to help users intuitively select and configure structured fields, meeting diverse business needs.

[0067] Step S5: Construct prompt words corresponding to the target waybill image based on the key field names.

[0068] The prompts generated in this embodiment can also be grouped and arranged according to the field category to which the key field name belongs. See the following example for a sample of the prompts generated in this embodiment: [Task Breakdown] 1. General category extraction (only extract if the category exists in its entirety): Extracted field: [key1] 2. Extraction of detailed categories: Parse each row according to the table / list structure, treating each row as an independent record. Extracted field: [key2]

Processing Rules

[0069] Based on the key field names confirmed by the user and combined with the built-in initialization template logic, the system constructs a custom prompt for subsequent multimodal model inference, ensuring the accuracy and controllability of the model extraction process. The initialization template can include information such as field names, field types, and format requirements, forming a model-oriented prompt instruction framework.

[0070] Step S6: Input the target waybill image and the corresponding prompt words into a preset multimodal big data model, so that the multimodal big data model can identify the field corresponding to the key field name in the target waybill image.

[0071] Based on the field names in the prompt words, the multimodal large model identifies the corresponding field values ​​in the target waybill image and outputs them in the format of "field name: field value".

[0072] By leveraging the multimodal understanding and reasoning capabilities of large language models, semantic analysis and structural extraction are performed on the content of images. Compared with traditional rule matching methods, this approach is more universal and adaptable, and can be adapted to waybill templates of different types and formats, thus achieving intelligent structuring capabilities for multimodal text and images.

[0073] Existing systems lack a unified standard for outputting field recognition results, which hinders subsequent data integration and processing. To address this issue, this embodiment designs a joint output interface for image information and structured data, employing a unified standard structure to express the results. The output includes information such as field name, field value, and field type. This output format can seamlessly integrate with third-party business systems, enabling automatic data entry, archiving, and verification, significantly improving the business usability of structured information and system integration efficiency.

[0074] Optionally, the method further includes: When the selected method is the built-in method, obtain the preset prompt words; The target waybill image and pre-set prompts are input into the multimodal big data model, so that the multimodal big data model can identify the required fields in the target waybill image based on the pre-set prompts.

[0075] The built-in method uses a pre-defined Prompt. An example of a built-in Prompt is as follows: Please carefully analyze the contents of the delivery note image and convert the information into structured JSON format according to the following requirements: [Task Breakdown] 1. Basic information extraction (only extract if the information exists completely): Supplier → Identify entries such as "Supplier" or "Consumer" in the header of documents. Numbering → Focus on finding identifiers such as "Number", "Document Number", "NO.", and "Order Number", paying attention to capitalization and variations of symbols. License plate number → Pay attention to keywords such as "license plate" and "vehicle number," and carefully distinguish between transport vehicle and trailer numbers. 2. Extraction of product details: The data is parsed line by line according to the table / list structure, with each line treated as an independent record. The field mapping table is as follows: Material Name → "Material Name" / "Name" / "Goods Name" / "Product Name" etc. Specifications and models → "specifications" / "model" / "specification parameters" / "model specifications", etc. Specifications and models differ from materials. Material → "Material" / "Materials" / "Material Properties" etc. Unit of measurement → "Unit" / "Unit of measurement" (Retrieved only if it exists, not inferred (e.g., without breaking down "200 tons"), only returns "". Waybill Quantity → Fields with numerical values ​​such as "Quantity" / "Volume" / "Weight" / "Shipping Volume" Unit Price → Clearly specify the numerical value of the "unit price", excluding monetary fields. Brand → Titles such as "Brand" / "Manufacturer" will not be extracted if no relevant field is found. Size → Combine and extract "Length" / "Width" / "Height" or "Dimension" combined value Note → Special explanatory text corresponding to this line

Processing Rules

[0076] The present invention has the following beneficial effects: (1) Intelligent photo-taking interaction mechanism based on image quality assessment An image quality assessment module is introduced to analyze waybill images uploaded or taken by users in real time, automatically determine whether there are problems such as blurriness, missing corners, tilt, or abnormal exposure, and provide users with intelligent prompts and interactive feedback to guide them to retake the photos, thus ensuring the image quality of structured recognition from the source.

[0077] (2) A semantic structuring method for waybill fields based on a multimodal large model A recognition framework combining OCR text detection and large language model inference capabilities is constructed. Through a carefully designed prompt, the large model is guided to perform semantic parsing on key information in waybill images, extracting field information including supplier, license plate number, material name, specifications, quantity, etc., thereby improving the structured accuracy and versatility in complex scenarios.

[0078] (3) Supports a structured recognition mechanism with configurable and interactive field selection. The large model, combined with a lightweight OCR model, returns the location information (bounding box) of all text in the image, allowing users to visually confirm, correct, or customize the selection of fields. This enables flexible configuration of field structures and interactive, controllable extraction, meeting the personalized needs of different projects and enterprises.

[0079] (4) Standardized interface design for joint output of image information and structural data It provides a unified return format for images and structured data, including field text and field type, which facilitates subsequent system integration and automatic database entry, thereby improving system integration and project implementation efficiency.

[0080] (5) Lightweight inference solution compatible with mobile devices By combining model pruning, inference optimization and other methods, key modules support running on mobile devices. In particular, the image quality detection and basic OCR modules have good resource adaptability and ensure a smooth user experience in weak network and low-performance terminal environments.

[0081] It should be noted that the training process of the corner detection model is as follows: I. Training Data Preparation 1. Input Image Size and Normalization All training images are uniformly scaled to a fixed resolution (e.g., 512×512) and normalized. Normalization methods include linear normalization to the [0, 1] interval, or standardization based on the mean and standard deviation of the training set.

[0082] 2. Corner point annotation format The four corner points of the waybill are manually marked in a fixed order, such as [x0, y0], [x1, y1], [x2, y2] and [x3, y3], which correspond to the top left corner, top right corner, bottom right corner and bottom left corner respectively; 3. Batch size settings Set a reasonable batch size based on the video memory resources of the training device, such as setting it to 16.

[0083] 4. Data Augmentation Strategies Geometric transformations: random rotation (±15°), scaling (0.8~1.2), translation (±10%), perspective transformation (simulating photographic viewpoint shift); Lighting transformation: Image brightness adjustment (±20%), contrast change (±30%), addition of Gaussian noise (σ=0.05); Occlusion enhancement: Simulates partial occlusion of the image and randomly erases local areas.

[0084] II. Gaussian Heatmap Generation This invention uses Gaussian heatmaps as a supervision signal to guide the model in learning precise corner locations. The specific process is as follows: 1. Downsampling and coordinate mapping Assuming the input image size is H×W (e.g., 512×512) and the output heatmap size is H / s×W / s (e.g., 128×128, downsampling factor s=4), map the coordinates (xi, yi) of each corner point in the original image to the coordinates (xi / s, yi / s) in the heatmap. 2. Gaussian kernel generation A two-dimensional Gaussian distribution is plotted at the mapped corner points. The standard deviation σ of the Gaussian kernel is set according to the size of the heatmap, and is usually taken as 1 / 50 to 1 / 30 of the heatmap width (e.g., for a 128×128 heatmap, σ=2~4).

[0085] Each corner point corresponds to an independent channel in the heatmap (4 channels in total). The Gaussian response is only plotted on its corresponding channel to ensure that there is no interference between corner points.

[0086] III. Subpixel offset prediction Because heatmap downsampling introduces quantization errors, the model simultaneously learns to predict sub-pixel-level offsets to improve corner prediction accuracy. The offset prediction branch outputs an 8-channel tensor, with each corner containing offset values ​​in both the x and y directions, corresponding to a total of 8 channels. During training, the offset loss is calculated only at locations where the true corners exist.

[0087] IV. Loss Function Design This invention employs a multi-task loss function to optimize model parameters, comprising two parts: heatmap loss and offset loss. 1. Heatmap Loss (L_heat) An improved form of Focal Loss is adopted to alleviate the problem of imbalance between positive and negative samples.

[0088] 2. Offset Loss (L_offset) Using the L1 loss function, calculations are only performed at locations where corners actually exist.

[0089] 3. Total Loss Function L = L_heat + λ × L_offset Where λ is the balance coefficient, which is empirically set to 0.1.

[0090] V. Optimization Strategies Optimizer: Adam; Initial learning rate: ; Learning rate decay: The learning rate decays to 0.7 of its original value every 20 training epochs; Total training rounds: approximately 200 rounds, depending on the specific convergence.

[0091] In this embodiment, the coordinates of the four corner points of the waybill image are accurately located by analyzing the Gaussian heatmap and offset branch output by the model during the inference phase. The entire prediction process mainly includes the following steps: I. Peak Extraction from Heat Map 1. Process each channel independently The heatmap output by the model contains four channels, corresponding to the top left, top right, bottom right, and bottom left corner points respectively. Each channel is processed separately to extract the corner point positions independently.

[0092] 2. Peak detection mechanism Perform a 3×3 max pooling operation on each heatmap channel to obtain a response map of the same size as the original heatmap.

[0093] Compare the values ​​of the original heatmap and the pooled heatmap. If the value at a certain position is the same and greater than a set threshold (such as 0.3), then that position is determined to be a possible peak point.

[0094] In each channel, the peak value with the highest response value is selected as the initial prediction position of that corner point, denoted as (u_c, v_c), where c∈{0,1,2,3} represents the channel number.

[0095] II. Subpixel offset correction To further improve positioning accuracy, the model also predicts the sub-pixel offset value of each corner point within the heatmap grid to compensate for quantization errors caused by downsampling.

[0096] 1. Offset Extraction At the predicted heatmap coordinates (u_c, v_c), extract the x / y offset values ​​from the corresponding channel of the offset branch.

[0097] 2. Restore coordinates to the original image space. Use a downsampling factor s (e.g., s=4) to map the heatmap coordinates and offset values ​​back to the original image coordinate system: x_c = (u_c + Δx_c) × s y_c = (v_c + Δy_c) × s III. Corner Point Sequence Output Since the four heatmap channels in the model output already correspond to the four corner points (channel 0: top left, channel 1: top right, channel 2: bottom right, channel 3: bottom left), no additional sorting or post-processing is required. The coordinates of the four corner points can be directly output in channel order and used for subsequent rotation correction or affine transformation operations.

[0098] Example 2 Embodiment 2 of the present invention provides a waybill identification device, such as Figure 3 As shown, the waybill identification device 30 includes: The acquisition module is used to acquire the target waybill image; The first determining module is used to determine the selection method of the prompt word corresponding to the target waybill image; The annotation module is used to annotate multiple field names in the target waybill image when the selected method is a custom method; The second determining module is used to determine the key field names to be identified in the target waybill image based on the marked field names; The module is used to construct prompt words corresponding to the target waybill image based on the key field names; The recognition module is used to input the target waybill image and the corresponding prompt words into a preset multimodal big data model, so that the multimodal big data model can identify the field corresponding to the key field name in the target waybill image.

[0099] Optionally, the acquisition module is specifically used for: Get the images of the waybills to be processed; Attempt to detect the four corner points of the waybill in the waybill image; When four corner points are detected and determined to be the four corner points of the waybill, the image quality of the waybill image is evaluated to see if it meets the preset quality requirements. When the image quality of the waybill image is evaluated to meet the preset quality requirements, the waybill image is rotated and corrected based on the four detected corner points so that the horizontal center axis of the waybill in the rotated and corrected waybill image is horizontal. The image size of the rotated and corrected waybill image is scaled to a preset image size to obtain the target waybill image.

[0100] Optionally, when the acquisition module performs rotation correction on the waybill image based on the detected four corner points after evaluating that the image quality of the waybill image meets the preset quality requirements, so that the horizontal center axis of the waybill in the rotated waybill image is horizontal, it is specifically used for: The four detected corner points are used as the four corner points of the rectangle, and the rectangle is drawn. Scale the rectangle and / or the waybill image so that the size of the waybill in the waybill image matches the size of the rectangle. The waybill within the rectangle is rotated and corrected based on affine transformation or perspective transformation so that the horizontal center axis of the rotated and corrected waybill is horizontal.

[0101] Optionally, the annotation module is specifically used for: Based on the multimodal large model, all field names belonging to the detail category in the target waybill image are identified; Based on a preset OCR recognition model, all characters in the target waybill image and the position information of each character are identified; Match the field names identified by the multimodal large model with the characters identified by the OCR recognition model; Assign the position information of the successfully matched characters to the corresponding field name; Based on the location information of each field name, a label box for each field name is drawn in the target waybill image.

[0102] Optionally, the second determining module is specifically used for: Receive user selection instructions for annotation boxes, mark the selected annotation box in the target waybill image, and extract the field names within the marked annotation box; and The field names within the marked annotation boxes will be used as the key field names to be identified in the target waybill image; or, The system receives user input instructions for field names in a general category, parses out the field names carried in the input instructions, and uses the field names within the marked annotation boxes and the field names carried in the input instructions as the key field names to be identified in the target waybill image.

[0103] Optionally, the annotation module is specifically used for: When the selected method is a custom method and there is no historical template associated with the supplier / shipper unit of the target waybill image, multiple field names are marked in the target waybill image; wherein, the historical template includes historical key field names; The device further includes: The segmentation module is used to divide the identified key field names in the target waybill image into several field name groups according to their respective field categories after the key field names to be identified are determined based on the labeled field names; wherein, the field category is a detailed category or a general category, and each field name group includes a field category and field names belonging to that field category; The storage module is used to group and store the divided field names as templates associated with the supplier / shipper of the target waybill image.

[0104] Optionally, when the annotation module annotates multiple field names in the target waybill image when the selected method is a custom method and there is no historical template associated with the supplier / shipper unit of the target waybill image, it is specifically used for: When the selected method is a custom method and there is no historical template, multiple field names are marked in the target waybill image; When the selected method is a custom method and there are several historical templates, multiple optional buttons are displayed, and when the first button is selected, multiple field names are marked in the target waybill image; wherein, the multiple optional buttons include: a first button indicating that a historical template is not used and a second button corresponding to each historical template.

[0105] Optionally, the device further includes: The retrieval module is used to retrieve the historical template corresponding to the selected second button when the second button is selected. The third determination module is used to use the historical key field names contained in the retrieved historical template as the key field names to be identified in the target waybill image.

[0106] Optionally, the device further includes: The extraction module is used to obtain pre-set prompt words when the selected method is the built-in method; The processing module is used to input the target waybill image and pre-set prompt words into the multimodal big data model, so that the multimodal big data model can identify the required fields in the target waybill image based on the pre-set prompt words.

[0107] Example 3 This embodiment also provides a computer device, such as a smartphone, tablet computer, laptop computer, desktop computer, rack server, blade server, tower server, or cabinet server (including a standalone server or a server cluster composed of multiple servers), etc., capable of executing programs. Figure 4 As shown, the computer device 40 in this embodiment includes, but is not limited to, a memory 401 and a processor 402 that are communicatively connected to each other via a system bus. It should be noted that... Figure 4 Only a computer device 40 with components 401-402 is shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0108] In this embodiment, the memory 401 (i.e., the readable storage medium) includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 401 may be an internal storage unit of the computer device 40, such as the hard disk or memory of the computer device 40. In other embodiments, the memory 401 may also be an external storage device of the computer device 40, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 40. Of course, the memory 401 may include both the internal storage unit and its external storage device of the computer device 40. In this embodiment, the memory 401 is typically used to store the operating system and various application software installed on the computer device 40. In addition, the memory 401 may also be used to temporarily store various types of data that have been output or will be output.

[0109] In some embodiments, processor 402 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. This processor 402 is typically used to control the overall operation of computer device 40.

[0110] Specifically, in this embodiment, the processor 402 is used to execute the program of the waybill identification method stored in the memory 401. The specific implementation process of the above method steps can be found in Embodiment 1, and will not be repeated here.

[0111] Example 4 This embodiment also provides a computer-readable storage medium, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, server, App application store, etc., which stores a computer program. When the computer program is executed by a processor, it implements the steps of the waybill identification method described in Embodiment 1. Specific implementation processes of the above method steps can be found in Embodiment 1, and will not be repeated here.

[0112] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0113] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0114] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0115] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A waybill identification method, characterized in that, The method includes: Obtain the target waybill image; Determine the selection method for the prompt words corresponding to the target waybill image; When the selected method is a custom method, multiple field names are marked in the target waybill image; Based on the labeled field names, determine the key field names to be identified in the target waybill image; Based on the key field names, construct the prompt words corresponding to the target waybill image; The target waybill image and the corresponding prompt words are input into a preset multimodal big data model so that the multimodal big data model can identify the field corresponding to the key field name in the target waybill image.

2. The method according to claim 1, characterized in that, The process of obtaining the target waybill image includes: Get the images of the waybills to be processed; Attempt to detect the four corner points of the waybill in the waybill image; When four corner points are detected and determined to be the four corner points of the waybill, the image quality of the waybill image is evaluated to see if it meets the preset quality requirements. When the image quality of the waybill image is evaluated to meet the preset quality requirements, the waybill image is rotated and corrected based on the four detected corner points so that the horizontal center axis of the waybill in the rotated and corrected waybill image is horizontal. The image size of the rotated and corrected waybill image is scaled to a preset image size to obtain the target waybill image.

3. The method according to claim 2, characterized in that, When the image quality of the waybill image is assessed to meet the preset quality requirements, the waybill image is rotated and corrected based on the four detected corner points so that the horizontal center axis of the waybill in the rotated and corrected waybill image is horizontal, including: The four detected corner points are used as the four corner points of the rectangle, and the rectangle is drawn. Scale the rectangle and / or the waybill image so that the size of the waybill in the waybill image matches the size of the rectangle. The waybill within the rectangle is rotated and corrected based on affine transformation or perspective transformation so that the horizontal center axis of the rotated and corrected waybill is horizontal.

4. The method according to claim 1, characterized in that, When the selected method is a custom method, multiple field names are marked on the target waybill image, including: Based on the multimodal large model, all field names belonging to the detail category in the target waybill image are identified; Based on a preset OCR recognition model, all characters in the target waybill image and the position information of each character are identified; Match the field names identified by the multimodal large model with the characters identified by the OCR recognition model; Assign the position information of the successfully matched characters to the corresponding field name; Based on the location information of each field name, a label box for each field name is drawn in the target waybill image.

5. The method according to claim 4, characterized in that, The process of determining the key field names to be identified in the target waybill image based on the labeled field names includes: Receive user selection instructions for annotation boxes, mark the selected annotation box in the target waybill image, and extract the field names within the marked annotation box; and The field names within the marked annotation boxes will be used as the key field names to be identified in the target waybill image; or, The system receives user input instructions for field names in a general category, parses out the field names carried in the input instructions, and uses the field names within the marked annotation boxes and the field names carried in the input instructions as the key field names to be identified in the target waybill image.

6. The method according to claim 1, characterized in that, When the selected method is a custom method, multiple field names are marked on the target waybill image, including: When the selected method is a custom method and there is no historical template associated with the supplier / shipper unit of the target waybill image, multiple field names are marked in the target waybill image; wherein, the historical template includes historical key field names; After determining the key field names to be identified in the target waybill image based on the labeled field names, the method further includes: The identified key field names are divided into several field name groups according to their respective field categories; wherein, the field category is a detailed category or a general category, and each field name group includes a field category and field names belonging to that field category; The segmented field names are grouped and stored as templates associated with the supplier / shipper unit of the target waybill image.

7. The method according to claim 6, characterized in that, When the selected method is a custom method and there is no historical template associated with the supplier / shipper unit of the target waybill image, multiple field names are marked in the target waybill image, including: When the selected method is a custom method and there is no historical template, multiple field names are marked in the target waybill image; When the selected method is a custom method and there are several historical templates, multiple optional buttons are displayed, and when the first button is selected, multiple field names are marked in the target waybill image; wherein, the multiple optional buttons include: a first button indicating that a historical template is not used and a second button corresponding to each historical template.

8. The method according to claim 7, characterized in that, The method further includes: When the second button is selected, retrieve the historical template corresponding to the selected second button; The historical key field names contained in the retrieved historical templates will be used as the key field names to be identified in the target waybill image.

9. The method according to claim 1, characterized in that, The method further includes: When the selected method is the built-in method, obtain the preset prompt words; The target waybill image and pre-set prompts are input into the multimodal big data model, so that the multimodal big data model can identify the required fields in the target waybill image based on the pre-set prompts.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it is used to implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Question recording method and system based on intelligent adaptation progressive intelligent learning

    CN119169630A

  • Form highlighting method and device based on large model, computer equipment and storage medium

    CN119578369A

  • Intelligent order recording method, device and equipment and storage medium

    CN119578379A

  • Document information extraction method, device and system and storage medium

    CN119942576A

  • Method for electronic data and signature collection, and system

    US20060288222A1