Port business confirmation function identification method, system and device and medium

Through the collaborative mechanism of primary extraction and secondary processing, the problem of missed recognition caused by blurred seals or abnormal forms in the port business confirmation letter was solved, the complete extraction of key information and structured data output were achieved, the recognition robustness and processing efficiency were improved, and the docking requirements of the business system were met.

CN120766288AActive Publication Date: 2025-10-10SHANDONG PORT LAND-SEA INT LOGISTICS GRP CO LTD +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510745885.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-10-10
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The existing port business confirmation letter recognition method lacks dynamic integrity judgment and secondary processing mechanisms. When the seal is blurred or the form is abnormal, the missed content cannot be automatically repaired, and the output data format is loose and difficult to apply directly.

Method used

A collaborative mechanism of primary extraction and secondary processing is adopted. The primary extraction is performed through the pre-trained business confirmation letter key information extraction model, seal position detection model and seal content recognition model, and the secondary processing is performed in combination with the key name library and seal template library to achieve automatic supplement of unsuccessfully recognized content and repair of seal content, and encapsulate the results into JSON structured data.

Benefits of technology

It achieves complete extraction of key information from port business confirmation letters, improves recognition robustness in complex scenarios, reduces the need for manual intervention, improves processing efficiency, and achieves seamless integration with business systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766288A_ABST
    Figure CN120766288A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of file identification, in particular to a port business confirmation letter identification method, system and device and a medium, and the method comprises the steps: obtaining a business confirmation letter file; converting the business confirmation letter file into business confirmation letter image data; performing primary extraction on the business confirmation letter image data to obtain primary extraction information; analyzing the primarily extracted information, judging whether the seal part content and the table part content are successfully identified, if the seal part content and the table part content are successfully identified, taking the primarily extracted information as a business confirmation letter identification result, and if a part of which the content is not successfully identified exists, carrying out secondary processing on the part to obtain a business confirmation letter identification result; and packaging the identification result of the business confirmation function into JSON structured data. According to the invention, through a primary extraction and secondary processing cooperation mechanism, dynamic identification result credibility evaluation and a structured data packaging technology, accurate extraction of total elements of the key information of the warranty, automatic closed-loop analysis and efficient interaction of a service system are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of document identification technology, and in particular to a method, system, device and medium for identifying a port business confirmation letter. Background Art

[0002] As the digital transformation of the port logistics industry accelerates, the need for automated identification of port business confirmation letters, core documents for port transportation operations, is becoming increasingly urgent. The port business confirmation letter process involves several steps: first, business personnel must understand the port terminal's rate information; second, prepare the port business confirmation letter based on business needs; and finally, send the port business confirmation letter to the port terminal via email, explaining the business situation and copying it to the billing center. Throughout this process, the identification and processing of port business confirmation letters still relies on manual operations, resulting in a low degree of digitalization, cumbersome operations, and low efficiency.

[0003] To improve the efficiency and automation of port business confirmation letter recognition, existing technologies employ a phased process: first, optical character recognition (OCR) is used to extract the letter's text content, then image segmentation algorithms are used to locate the seal area and perform text recognition, ultimately outputting the results as structured data. This approach, by processing text and seal information in a step-by-step manner, has initially achieved automated parsing of the letter's content.

[0004] However, the existing technology still has the following defects: a single recognition process cannot dynamically evaluate the extraction integrity of key information, resulting in missed recognition content directly entering the output link and requiring manual completion; there is a lack of automated compensation mechanism for unsuccessful recognition parts. When the seal area is blurred or the table structure is abnormal, the system cannot automatically repair the error; the output data format is loose and has not formed a standardized connection with the business system, which increases the subsequent data processing costs. Summary of the Invention

[0005] In response to the technical problems that the existing port business confirmation letter identification method lacks dynamic integrity judgment and secondary processing mechanism, missed identification content cannot be automatically repaired when the seal is blurred or the form is abnormal, and the output data format is loose and difficult to directly apply, this application provides a sea-rail intermodal transport business confirmation letter identification method, system, equipment and medium. Through the primary extraction and secondary processing collaborative mechanism, dynamic recognition result credibility assessment and structured data packaging technology, it realizes the full-factor accurate extraction of key information of the letter of guarantee, automated closed-loop analysis and efficient interaction with the business system.

[0006] In a first aspect, the present application provides a method for identifying a port business confirmation letter, comprising the following steps: S1. Obtain the business confirmation letter document, which is a single-page PDF file of the port business confirmation letter containing the title, form, seal, and other content; S2. Converting the business confirmation letter file into business confirmation letter image data; S3. Perform initial extraction of the business confirmation letter image data, including key information extraction, seal extraction, and seal content recognition, to obtain the initial extracted information; Among them, key information extraction is to input the business confirmation letter image data into the pre-trained business confirmation letter key information extraction model, and output the key information extraction results, including all text fragments in the business confirmation letter image data, the category label corresponding to each text fragment, and the key name-key value pair relationship label of the question-answer; Seal extraction includes locating the seal area, extracting the circular text image and the linear text image in the seal area, and preprocessing the circular text image; Seal content recognition includes recognizing the content of pre-processed circular text images and linear text images, and outputting the seal content recognition results; S4. Determine whether the initial information extraction successfully identifies the seal portion and the form portion; If all are successfully identified, the first extracted information will be used as the business confirmation letter identification result; If there is a part of the content that is not successfully identified, the part of the content that is not successfully identified is processed again to obtain the identification result of the business confirmation letter; Secondary processing includes secondary processing of the table part and secondary processing of the seal part. The secondary processing of the table part is to generate key name-key value pair relationship labels missing in the initial extraction information and supplement them into the key information extraction results. The secondary processing of the seal part is to match the seal area image with the seal template library and use the seal content information corresponding to the matched seal template image to replace the original seal content recognition result in the initial extraction information. S5. Encapsulate the business confirmation letter recognition result into JSON structured data.

[0007] It should be further explained that step S2 performs three-channel pixel conversion on the business confirmation letter file to generate business confirmation letter image data, and the business confirmation letter image data is in JPG, JPEG or PNG format.

[0008] It should be further explained that in step S3, the step of extracting key information using the pre-trained business confirmation letter key information extraction model includes: Extract all text fragments from the business confirmation letter image data; Categorize text snippets into categories such as questions, answers, titles, and other content; Perform key-value matching on questions and answers, with questions corresponding to key names and answers corresponding to key values; Output key information extraction results, including all text segments in the business confirmation letter image data, the category label corresponding to each text segment, and the key name-value pair relationship label of the question-answer.

[0009] It should be further explained that, in step S3, seal extraction includes: Use the pre-trained seal position detection model, a deep learning model, to locate the seal area and output the seal area bounding box coordinates and seal area image. A pre-trained text region segmentation model is used to extract circular text images and linear text images within the seal region image. The text region segmentation model is a deep learning model. Perform polar coordinate transformation on the annular text image to obtain a straightened annular text image.

[0010] It should be further explained that, in step S3, a pre-trained seal content recognition model is used to perform seal content recognition. The seal content recognition model is a deep learning model, including: The linear text image and the straightened circular text image are input into the seal content recognition model, and the seal content recognition results are output.

[0011] It should be further explained that the business confirmation letter key information extraction model is built based on the VI-LayoutXLM algorithm. The training steps include: Obtain a business confirmation letter document sample, and convert the business confirmation letter document sample into business confirmation letter sample image data; Manually classify the text snippets in the business confirmation letter sample image data and perform key-value matching on the questions and answers to obtain the actual key information of the business confirmation letter sample image data, including the category label corresponding to each text snippet and the key name-value pair relationship label of the question-answer relationship. Each set of business confirmation letter sample image data and its actual key information is used as a sample to construct a business confirmation letter sample key information dataset. The business confirmation letter key information dataset is used to train the business confirmation letter key information extraction model in two stages: In the first phase, the text segment category recognition capability of the business confirmation letter key information extraction model is trained based on the category label corresponding to each text segment, and the cross-entropy loss function is used to optimize the category label recognition results. In the second stage, based on the fixed category label corresponding to each text fragment, the key-value matching ability of the business confirmation letter key information extraction model is trained based on the key-name-key-value pair relationship labels of the question-answer, and the similarity comparison loss function is used to optimize the matching relationship of the key-name-key-value pairs.

[0012] It should be further explained that the seal position detection model is built based on the YOLO algorithm. The training steps include: Obtain a business confirmation letter document sample, and convert the business confirmation letter document sample into business confirmation letter sample image data; Manually use annotation tools to mark the seal area bounding box coordinates and category labels in the business confirmation letter sample image data. Each set of business confirmation letter sample image data and its seal area bounding box coordinates and category labels is used as a sample to construct the business confirmation letter sample seal position dataset; The seal position detection model is trained using a dataset of seal positions in business confirmation letters. Through a multi-scale feature map prediction mechanism, the intersection-over-union loss function is used to optimize the positioning accuracy of the seal region bounding box. The text region segmentation model is built based on the Mask R-CNN algorithm. The training steps include: Obtain seal sample images, manually use annotation tools to annotate circular text region segmentation masks and straight text region segmentation masks in the seal sample images, and use each set of seal sample images and the corresponding circular text region segmentation masks and straight text region segmentation masks as a sample to construct a text region segmentation dataset; The text region segmentation dataset is used to train the text region segmentation model, a pixel-level segmentation network is adopted, and the binary cross entropy loss function is used to optimize the text region segmentation results.

[0013] It should be further explained that the seal content recognition model is built based on the PPOCRv4 algorithm. The training steps include: Acquire a seal sample image, extract a sample circular text image and a sample linear text image from the seal sample image, perform polar coordinate transformation on the sample circular text image, and obtain a straightened sample circular text image; Manually identify the text content of the straightened sample circular text images and the sample linear text images, and use each straightened sample circular text image and its text content as a sample, and each linear text image and its text content as a sample to construct a seal content dataset; The seal content dataset is used to train the seal content recognition model, and the seal content recognition results are optimized using the sequence transcription loss function.

[0014] It should be further explained that step S6 is also included: adding the business confirmation letter image data and the corresponding business confirmation letter recognition results to the training data set, and using the training data set to iteratively update the business confirmation letter key information extraction model, seal position detection model, text area segmentation model and seal content recognition model through incremental learning.

[0015] It should be further explained that the rules for determining whether the seal content and the form content are successfully recognized in step S4 include: Extract the company name from the seal content recognition result and the company name from the title text segment from the key information extraction result, and then calculate the similarity between the two. If the similarity is lower than the preset threshold, it is determined that the seal content has not been successfully recognized; The key-value matching relationship of all questions and answers in the key information extraction results is verified. If the key name does not match the key value, it is determined that part of the table content was not successfully identified.

[0016] It should be further explained that in step S4, the secondary processing steps of the table portion include: S401 sets the key name library, the key name library stores the key name and its corresponding key value regular expression rules; S402. When the contents of the table portion are not successfully identified, the table portion is subjected to OCR recognition, all text fragments of the table portion are extracted, and the spatial coordinates of each key name that does not match the key value in the image are recorded; S403 matches each unmatched key name with the key name in the key name library. After a successful match, extract all text fragments that match the key name corresponding to the regular expression as candidate key values; S404 records the spatial position coordinates of the candidate key in the image, selects the candidate key with the closest spatial distance to the unmatched key according to the spatial position coordinates as the matching key of the key, forms the key name - key value pair relationship label of the key, and adds it to the key information extraction result; Secondary processing of seals includes: S411 set seal template library, seal template library stores a mapping relationship between the company name and the corresponding seal template image group, each seal template image group contains at least one seal template image corresponding to the company, each seal template image associated with pre-stored seal content information; S412 extracts the key information extraction result from the title text fragment of the company name, based on the extracted company name query seal template library, the company name corresponding to the seal template image group as a candidate seal template image group; S413. ORB feature point matching is performed on the seal area image and each seal template image in the candidate seal template image group to calculate the image matching similarity score. The seal template image with the highest image matching similarity score and above a preset threshold is used as the target seal template image. S414. The seal content information associated with the target seal template image is used as the actual seal content recognition result, replacing the original seal content recognition result in the initially extracted information.

[0017] It should be further explained that in step S403, if there is no key name with no matching key value in the key name library, the matching fails and the key value corresponding to the key name is defined as blank; After a successful match, if there is no text fragment that matches the regular expression corresponding to the key name in all text fragments, the key value corresponding to the key name is defined as blank.

[0018] In a second aspect, the present application provides a port business confirmation letter identification system for implementing the above-mentioned business confirmation letter identification method, comprising: A pre-processing module, used for acquiring a business confirmation letter file and converting the business confirmation letter file into business confirmation letter image data; The initial extraction module is used to perform initial extraction on the business confirmation letter image data to obtain initial extraction information; The analysis and judgment module is used to analyze the initially extracted information and determine whether the seal content and the form content are successfully recognized; A secondary processing module is used to perform secondary processing on the part of the content that was not successfully recognized to obtain the recognition result of the business confirmation letter; The structured processing module is used to encapsulate the business confirmation letter recognition results into JSON structured data.

[0019] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the steps of the above-mentioned business confirmation letter identification method when executing the computer program.

[0020] In a fourth aspect, the present application provides a storage medium having a computer program stored thereon, which implements the steps of the above-mentioned business confirmation letter identification method when executed by a processor.

[0021] It can be seen from the above technical solutions that this application has the following advantages: 1. This application solves the problem of missed recognition caused by blurred seals or complex forms in the existing technology through a collaborative mechanism of primary extraction and secondary processing, achieves complete extraction of key information and seal content in the business confirmation letter, and significantly improves the recognition robustness in complex scenarios.

[0022] 2. This application solves the defect of existing technologies that they cannot automatically evaluate the reliability of recognition by dynamically judging the completeness of the initial extraction results, realizes targeted secondary processing of unsuccessfully identified content, reduces the need for manual intervention, and improves processing efficiency.

[0023] 3. This application solves the problem of loose output format and difficulty in direct application of existing technologies through structured data encapsulation technology, and realizes standardized JSON output of business confirmation letter information, which facilitates seamless integration with business systems and improves data flow efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for the description. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0025] Figure 1 This is a flowchart of a method for identifying a business confirmation letter in one embodiment of the present application.

[0026] Figure 2 It is a schematic block diagram of a business confirmation letter identification system in one embodiment of the present application.

[0027] Figure 3 It is a schematic diagram of the hardware structure of an electronic device in one embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the application objectives, features, and advantages of this application more obvious and easy to understand, the technical solutions protected by this application will be clearly and completely described below using specific embodiments and drawings. Obviously, the embodiments described below are only part of the embodiments of this application, not all of them. Based on the embodiments in this patent, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this patent.

[0029] The business confirmation letter recognition method involved in this application is mainly aimed at the field of document recognition technology. Through the collaborative mechanism of primary extraction and secondary processing, it solves the problem of missed recognition caused by blurred seals or complex forms in the existing technology, realizes the complete extraction of key information and seal content of the business confirmation letter, and significantly improves the recognition robustness in complex scenarios; by dynamically judging the integrity of the primary extraction results, it solves the defect that the existing technology cannot automatically evaluate the recognition credibility, realizes targeted secondary processing of unsuccessfully recognized content, reduces the need for manual intervention, and improves processing efficiency; through structured data encapsulation technology, it solves the problem that the output result format of the existing technology is loose and difficult to directly apply, and realizes standardized JSON output of business confirmation letter information, which facilitates seamless connection with the business system and improves data flow efficiency.

[0030] The business confirmation letter recognition method involved in this application is mainly aimed at the technical problems that the existing port business confirmation letter recognition method lacks dynamic integrity judgment and secondary processing mechanism, missed recognition content cannot be automatically repaired when the seal is blurred or the form is abnormal, and the output data format is loose and difficult to directly apply.

[0031] The following describes in detail the business confirmation letter identification method involved in this application. Specific details such as specific system structures and technologies are provided for illustration rather than limitation to facilitate a thorough understanding of the embodiments of this application. However, it should be clear to those skilled in the art that this application may also be implemented in other embodiments without these specific details.

[0032] In the business confirmation letter identification method involved in this application, the term "including" is used to indicate the presence of the described features, entities, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, entities, steps, operations, elements, components and / or their collections. The terms "including," "comprising," "having" and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0033] To facilitate the clear description of the technical solutions of this application, the words "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or order of execution, and the words "first" and "second" do not necessarily mean different.

[0034] The phrases "one embodiment" or "some embodiments" described in this application mean that the specific features, structures, or characteristics described in the embodiment are included in one or more embodiments of the application. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in other embodiments," etc. that appear in different places in this application do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized.

[0035] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0036] The business confirmation letter identification method provided in the embodiment of the present application is executed by a computer device, and accordingly, the business confirmation letter identification system runs in the computer device.

[0037] The following are some explanations of terms in this plan to facilitate a better understanding of this plan: A port business confirmation letter refers to an electronic confirmation document issued by the business initiator to the port operator in a port logistics service scenario. This document records the core information of a specific business in a structured table format, including business elements such as the acceptance number of cargo transportation, container specifications / box numbers, operation time windows, fee details and settlement terms, and must be stamped with a registered corporate electronic seal or digital signature. Its characteristics are as follows: the document content is generated according to the standard template for port business, key fields (such as box number and fee item code) comply with port data specifications, and the electronic seal information is bound to the registration database. The port business confirmation letter is a key carrier for realizing automated billing, operational responsibility traceability and compliance audits in the digital process of port business.

[0038] Figure 1 This is a flowchart of a method for identifying a business confirmation letter according to an embodiment of the present application. Figure 1 The execution entity may be a port business confirmation letter identification system. According to different requirements, the order of the steps in the flowchart can be changed, and some steps can be omitted.

[0039] like Figure 1 As shown, the business confirmation letter identification method includes: Step S1, obtaining a business confirmation letter file, which is a single-page port business confirmation letter PDF file including a title, form, seal and other contents.

[0040] By limiting the business confirmation letter document to a single-page PDF format and including a title, form, and seal, the uniformity and integrity of the input data are ensured, providing a standardized data source for subsequent information extraction and structured processing.

[0041] Step S2: converting the business confirmation letter file into business confirmation letter image data.

[0042] By converting PDF files into image data, the complexity of PDF native format parsing is solved, providing a unified image input interface for subsequent computer vision-based text and seal extraction.

[0043] In some specific embodiments, three-channel pixel conversion is performed on the business confirmation letter file to generate business confirmation letter image data, and the business confirmation letter image data is in JPG, JPEG or PNG format.

[0044] By converting three-channel pixels into JPG / JPEG / PNG format images, the problem of missing color channels in PDF conversion is solved, ensuring the complete preservation of color features and format adaptation during subsequent model processing.

[0045] Step S3, performing initial extraction on the business confirmation letter image data, including key information extraction, seal extraction, and seal content recognition, to obtain initial extracted information; Among them, key information extraction is to input the business confirmation letter image data into the pre-trained business confirmation letter key information extraction model, and output the key information extraction results, including all text fragments in the business confirmation letter image data, the category label corresponding to each text fragment, and the key name-value pair relationship label of the question-answer; Seal extraction includes locating the seal area, extracting the circular text image and the linear text image in the seal area, and preprocessing the circular text image; Seal content recognition includes recognizing the contents of pre-processed circular text images and linear text images, and outputting the seal content recognition results.

[0046] By performing key information extraction, seal extraction and seal content recognition, parallel processing of text structured analysis and seal semantic understanding is achieved, improving overall recognition efficiency and reducing module coupling.

[0047] In some specific embodiments, the step of extracting key information using a pre-trained business confirmation letter key information extraction model includes: Extract all text fragments from the business confirmation letter image data; Categorize text snippets into categories such as questions, answers, titles, and other content; Perform key-value matching on questions and answers, with questions corresponding to key names and answers corresponding to key values; Output key information extraction results, including all text segments in the business confirmation letter image data, the category label corresponding to each text segment, and the key name-value pair relationship label of the question-answer.

[0048] By using pre-trained models to classify text fragments and match question-answer key values, semantic relationship analysis of non-fixed layout documents is achieved, avoiding the limitations of traditional rule-based methods that strongly rely on typesetting.

[0049] In some specific embodiments, stamp extraction includes: Use the pre-trained seal position detection model, a deep learning model, to locate the seal area and output the seal area bounding box coordinates and seal area image. A pre-trained text region segmentation model is used to extract circular text images and linear text images within the seal region image. The text region segmentation model is a deep learning model. Perform polar coordinate transformation on the annular text image to obtain a straightened annular text image.

[0050] Through the cascade processing of seal area positioning and text area segmentation models, the independent extraction of circular and linear arranged text within the seal is achieved, providing an adaptive input form for subsequent content recognition.

[0051] In some specific embodiments, a pre-trained seal content recognition model is used for seal content recognition. The seal content recognition model is a deep learning model including: The linear text image and the straightened circular text image are input into the seal content recognition model, and the seal content recognition results are output.

[0052] By unifying the recognition models for straightened circular text and linearly arranged text, the blind spots of traditional OCR in recognizing special typeset text are resolved, and the completeness and accuracy of seal content analysis are improved.

[0053] In some specific embodiments, the business confirmation letter key information extraction model is constructed based on the VI-LayoutXLM algorithm, and the training steps include: Obtain a business confirmation letter document sample, and convert the business confirmation letter document sample into business confirmation letter sample image data; Manually classify the text snippets in the business confirmation letter sample image data and perform key-value matching on the questions and answers to obtain the actual key information of the business confirmation letter sample image data, including the category label corresponding to each text snippet and the key name-value pair relationship label of the question-answer relationship. Each set of business confirmation letter sample image data and its actual key information is used as a sample to construct a business confirmation letter sample key information dataset. The business confirmation letter key information dataset is used to train the business confirmation letter key information extraction model in two stages: In the first phase, the text segment classification capability of the business confirmation letter key information extraction model is trained based on the category labels corresponding to each text segment. The cross-entropy loss function is used to optimize the classification results. In the second stage, based on the fixed category label corresponding to each text fragment, the key-value matching ability of the business confirmation letter key information extraction model is trained based on the key-name-key-value pair relationship labels of the question-answer, and the similarity comparison loss function is used to optimize the matching relationship of the key-name-key-value pairs.

[0054] By optimizing text classification and key-value matching tasks through a phased training strategy, the parameter conflict problem of multi-task models is solved, and the classification accuracy and relationship mapping capabilities of the key information extraction model are improved.

[0055] In some specific embodiments, the seal position detection model is constructed based on the YOLO algorithm, and the training steps include: Obtain a business confirmation letter document sample, and convert the business confirmation letter document sample into business confirmation letter sample image data; Manually use annotation tools to mark the seal area bounding box coordinates and category labels in the business confirmation letter sample image data. Each set of business confirmation letter sample image data and its seal area bounding box coordinates and category labels is used as a sample to construct the business confirmation letter sample seal position dataset; The seal position detection model is trained using a dataset of seal positions in business confirmation letters. Through a multi-scale feature map prediction mechanism, the intersection-over-union loss function is used to optimize the positioning accuracy of the seal region bounding box. The text region segmentation model is built based on the Mask R-CNN algorithm. The training steps include: Obtain seal sample images, manually use annotation tools to annotate circular text region segmentation masks and straight text region segmentation masks in the seal sample images, and use each set of seal sample images and the corresponding circular text region segmentation masks and straight text region segmentation masks as a sample to construct a text region segmentation dataset; The text region segmentation dataset is used to train the text region segmentation model, a pixel-level segmentation network is adopted, and the binary cross entropy loss function is used to optimize the text region segmentation results.

[0056] Through multi-scale feature prediction and pixel-level segmentation network, high-precision positioning of the seal area bounding box and refined segmentation of the text area are achieved, providing reliable spatial feature support for subsequent processing.

[0057] In some specific embodiments, the seal content recognition model is constructed based on the PPOCRv4 algorithm, and the training steps include: Acquire a seal sample image, extract a sample circular text image and a sample linear text image from the seal sample image, perform polar coordinate transformation on the sample circular text image, and obtain a straightened sample circular text image; Manually identify the text content of the straightened sample circular text images and the sample linear text images, and use each straightened sample circular text image and its text content as a sample, and each linear text image and its text content as a sample to construct a seal content dataset; The seal content dataset is used to train the seal content recognition model, and the seal content recognition results are optimized using the sequence transcription loss function.

[0058] Through polar coordinate transformation preprocessing and sequence transcription loss function, the recognition difficulty caused by the deformation of circular text is solved, and the end-to-end recognition efficiency of circularly arranged text in seals is significantly improved.

[0059] Step S4, analyzing the initially extracted information to determine whether the seal portion and the form portion are successfully recognized; If all are successfully identified, the first extracted information will be used as the business confirmation letter identification result; If there is a part of the content that is not successfully identified, the part of the content that is not successfully identified is processed again to obtain the identification result of the business confirmation letter; The secondary processing includes secondary processing of the table part and secondary processing of the seal part. The secondary processing of the table part is to generate the key name-key value pair relationship labels missing in the initial extracted information and add them to the key information extraction results. The secondary processing of the seal part is to match the seal area image with the seal template library, and use the seal content information corresponding to the matched seal template image to replace the original seal content recognition results in the initial extracted information.

[0060] By setting the branch judgment logic of the initial extraction results, targeted secondary processing of unrecognized content is achieved, avoiding the termination of the overall process due to local recognition failure and enhancing the system's fault tolerance.

[0061] In some specific embodiments, the rules for determining whether the seal content and the form content are successfully recognized include: Extract the company name from the seal content recognition result and the company name from the title text segment from the key information extraction result, and then calculate the similarity between the two. If the similarity is lower than the preset threshold, it is determined that the seal content has not been successfully recognized; The key-value matching relationship of all questions and answers in the key information extraction results is verified. If the key name does not match the key value, it is determined that part of the table content was not successfully identified.

[0062] Through similarity threshold determination and key value integrity verification rules, automatic quality assessment of recognition results is achieved to ensure that the output data meets business logic constraints and consistency requirements.

[0063] In some specific embodiments, the table portion secondary processing step includes: S401 sets the key name library, the key name library stores the key name and its corresponding key value regular expression rules; S402. When the contents of the table portion are not successfully identified, the table portion is subjected to OCR recognition, all text fragments of the table portion are extracted, and the spatial coordinates of each key name that does not match the key value in the image are recorded; S403 matches each unmatched key name with the key name in the key name library. After a successful match, extract all text fragments that match the key name corresponding to the regular expression as candidate key values; S404 records the spatial position coordinates of the candidate key in the image, selects the candidate key with the closest spatial distance to the unmatched key according to the spatial position coordinates as the matching key of the key, forms the key name - key value pair relationship label of the key, and adds it to the key information extraction result; Secondary processing of seals includes: S411 set seal template library, seal template library stores a mapping relationship between the company name and the corresponding seal template image group, each seal template image group contains at least one seal template image corresponding to the company, each seal template image associated with pre-stored seal content information; S412 extracts the key information extraction result from the title text fragment of the company name, based on the extracted company name query seal template library, the company name corresponding to the seal template image group as a candidate seal template image group; S413. ORB feature point matching is performed on the seal area image and each seal template image in the candidate seal template image group to calculate the image matching similarity score. The seal template image with the highest image matching similarity score and above a preset threshold is used as the target seal template image. S414. The seal content information associated with the target seal template image is used as the actual seal content recognition result, replacing the original seal content recognition result in the initially extracted information.

[0064] Through the regular matching of key name library and the spatial coordinate association strategy, the problem of dynamic supplement of unidentified key values ​​is solved, and combined with the seal template matching mechanism, the fault tolerance and repair capabilities in complex scenarios are improved.

[0065] In some specific embodiments, in step S403, if there is no key name with no matching key value in the key name library, the matching fails, and the key value corresponding to the key name is defined as blank; After a successful match, if there is no text fragment that matches the regular expression corresponding to the key name in all text fragments, the key value corresponding to the key name is defined as blank.

[0066] By defining blank key values ​​and handling regular expression matching failures, the output format of abnormal data is standardized to avoid downstream system parsing errors or process interruptions caused by missing content.

[0067] In some specific embodiments, in step S412, if the extracted company name does not exist in the seal template library, the current process is terminated and an error message is output; In step S413 , if there is no seal template image with an image matching similarity score higher than the preset threshold, the current process ends and an error message is output.

[0068] Step S5: Encapsulate the business confirmation letter recognition result into JSON structured data.

[0069] By encapsulating the recognition results in JSON format, standardized conversion of unstructured documents to machine-readable data is achieved, meeting the compatibility requirements of business systems for data interfaces.

[0070] In some specific embodiments, step S6 is also included: adding the business confirmation letter image data and the corresponding business confirmation letter recognition results to a training data set, and using the training data set to iteratively update the business confirmation letter key information extraction model, seal position detection model, text area segmentation model and seal content recognition model through incremental learning.

[0071] By dynamically updating model parameters through incremental learning, the problem of model performance degradation caused by changes in business data distribution is solved, ensuring the stability and adaptability of the recognition system in long-term operation.

[0072] In a specific embodiment, the steps of the business confirmation letter identification method include: Step S1, obtaining a business confirmation letter file, which is a single-page port business confirmation letter PDF file including a title, form, seal and other contents.

[0073] Step S2: Perform three-channel pixel conversion on the business confirmation letter file to generate business confirmation letter image data. The business confirmation letter image data is in JPG, JPEG or PNG format.

[0074] Step S3, performing initial extraction on the business confirmation letter image data, including key information extraction, seal extraction, and seal content recognition, to obtain initial extracted information; The key information extraction is performed using a pre-trained business confirmation letter key information extraction model. The steps include: Extract all text fragments from the business confirmation letter image data; Categorize text snippets into categories such as questions, answers, titles, and other content; Perform key-value matching on questions and answers, with questions corresponding to key names and answers corresponding to key values; Output key information extraction results, including all text segments in the business confirmation letter image data, the category label corresponding to each text segment, and the key-value pair relationship label of the question-answer relationship; Seal extraction includes: Use the pre-trained seal position detection model, a deep learning model, to locate the seal area and output the seal area bounding box coordinates and seal area image. A pre-trained text region segmentation model is used to extract circular text images and linear text images within the seal region image. The text region segmentation model is a deep learning model. Perform polar coordinate transformation on the annular text image to obtain a straightened annular text image; Use the pre-trained seal content recognition model for seal content recognition. The seal content recognition model is a deep learning model that includes: Inputting the linear text image and the straightened circular text image into the seal content recognition model and outputting the seal content recognition result; The business confirmation letter key information extraction model is built based on the VI-LayoutXLM algorithm. The training steps include: Obtain a business confirmation letter document sample, and convert the business confirmation letter document sample into business confirmation letter sample image data; Manually classify the text snippets in the business confirmation letter sample image data and perform key-value matching on the questions and answers to obtain the actual key information of the business confirmation letter sample image data, including the category label corresponding to each text snippet and the key name-value pair relationship label of the question-answer relationship. Each set of business confirmation letter sample image data and its actual key information is used as a sample to construct a business confirmation letter sample key information dataset. The business confirmation letter key information dataset is used to train the business confirmation letter key information extraction model in two stages: In the first phase, the text segment classification capability of the business confirmation letter key information extraction model is trained based on the category labels corresponding to each text segment. The cross-entropy loss function is used to optimize the classification results. In the second phase, based on the fixed category label corresponding to each text segment, the key-value matching capability of the business confirmation letter key information extraction model is trained based on the question-answer key-value pair relationship labels. The similarity comparison loss function is used to optimize the key-value pair matching relationship. The seal position detection model is built based on the YOLO algorithm. The training steps include: Obtain a business confirmation letter document sample, and convert the business confirmation letter document sample into business confirmation letter sample image data; Manually use annotation tools to mark the seal area bounding box coordinates and category labels in the business confirmation letter sample image data. Each set of business confirmation letter sample image data and its seal area bounding box coordinates and category labels is used as a sample to construct the business confirmation letter sample seal position dataset; The seal position detection model is trained using a dataset of seal positions in business confirmation letters. Through a multi-scale feature map prediction mechanism, the intersection-over-union loss function is used to optimize the positioning accuracy of the seal region bounding box. The text region segmentation model is built based on the Mask R-CNN algorithm. The training steps include: Obtain seal sample images, manually use annotation tools to annotate circular text region segmentation masks and straight text region segmentation masks in the seal sample images, and use each set of seal sample images and the corresponding circular text region segmentation masks and straight text region segmentation masks as a sample to construct a text region segmentation dataset; Use the text region segmentation dataset to train the text region segmentation model, adopt a pixel-level segmentation network, and optimize the text region segmentation results through the binary cross entropy loss function; The seal content recognition model is built based on the PPOCRv4 algorithm. The training steps include: Acquire a seal sample image, extract a sample circular text image and a sample linear text image from the seal sample image, perform polar coordinate transformation on the sample circular text image, and obtain a straightened sample circular text image; Manually identify the text content of the straightened sample circular text images and the sample linear text images, and use each straightened sample circular text image and its text content as a sample, and each linear text image and its text content as a sample to construct a seal content dataset; The seal content dataset is used to train the seal content recognition model, and the seal content recognition results are optimized using the sequence transcription loss function.

[0075] Step S4: Analyze the initially extracted information to determine whether the seal content and the form content are successfully recognized; If all are successfully identified, the first extracted information will be used as the business confirmation letter identification result; If there is a part of the content that is not successfully identified, the part of the content that is not successfully identified is processed again to obtain the identification result of the business confirmation letter; The rules for determining whether the seal content and the form content are successfully recognized include: Extract the company name from the seal content recognition result and the company name from the title text segment from the key information extraction result, and then calculate the similarity between the two. If the similarity is lower than the preset threshold, it is determined that the seal content has not been successfully recognized; Verify the key-value matching relationship of all questions and answers in the key information extraction results. If the key name does not match the key value, it is determined that part of the table content was not successfully recognized; The secondary processing steps for the table section include: S401 sets the key name library, the key name library stores the key name and its corresponding key value regular expression rules; S402. When the contents of the table portion are not successfully identified, the table portion is subjected to OCR recognition, all text fragments of the table portion are extracted, and the spatial coordinates of each key name that does not match the key value in the image are recorded; S403 matches each unmatched key name with the key name in the key name library. After a successful match, extract all text fragments that match the key name corresponding to the regular expression as candidate key values; S404 records the spatial position coordinates of the candidate key in the image, selects the candidate key with the closest spatial distance to the unmatched key according to the spatial position coordinates as the matching key of the key, forms the key name - key value pair relationship label of the key, and adds it to the key information extraction result; Secondary processing of seals includes: S411 set seal template library, seal template library stores a mapping relationship between the company name and the corresponding seal template image group, each seal template image group contains at least one seal template image corresponding to the company, each seal template image associated with pre-stored seal content information; S412 extracts the key information extraction result from the title text fragment of the company name, based on the extracted company name query seal template library, the company name corresponding to the seal template image group as a candidate seal template image group; S413. ORB feature point matching is performed on the seal area image and each seal template image in the candidate seal template image group to calculate the image matching similarity score. The seal template image with the highest image matching similarity score and above a preset threshold is used as the target seal template image. S414. The seal content information associated with the target seal template image is used as the actual seal content recognition result, replacing the original seal content recognition result in the initially extracted information.

[0076] Step S5: Encapsulate the business confirmation letter recognition result into JSON structured data.

[0077] Step S6: Add the business confirmation letter image data and the corresponding business confirmation letter recognition results to the training data set, and use the training data set to iteratively update the business confirmation letter key information extraction model, seal position detection model, text area segmentation model and seal content recognition model through incremental learning.

[0078] The following is an embodiment of the business confirmation letter identification system provided in the embodiments of the present application. The business confirmation letter identification system and the business confirmation letter identification methods of the above-mentioned embodiments belong to the same inventive concept. For details not fully described in the embodiments of the business confirmation letter identification system, please refer to the embodiments of the above-mentioned business confirmation letter identification method.

[0079] like Figure 2 As shown, the business confirmation letter identification system includes: A pre-processing module, used for acquiring a business confirmation letter file and converting the business confirmation letter file into business confirmation letter image data; The initial extraction module is used to perform initial extraction on the business confirmation letter image data to obtain initial extraction information; The analysis and judgment module is used to analyze the initially extracted information and determine whether the seal content and the form content are successfully recognized; A secondary processing module is used to perform secondary processing on the part of the content that was not successfully recognized to obtain the recognition result of the business confirmation letter; The structured processing module is used to encapsulate the business confirmation letter recognition results into JSON structured data.

[0080] The business confirmation letter recognition system of this embodiment is used to implement a business confirmation letter recognition method, and the steps include: S1. Obtain the business confirmation letter, which is a single-page PDF file containing the title, form, seal, and other content. S2. Converting the business confirmation letter file into business confirmation letter image data; S3. Perform initial extraction of the business confirmation letter image data, including key information extraction, seal extraction, and seal content recognition, to obtain the initial extracted information; Among them, key information extraction is to input the business confirmation letter image data into the pre-trained business confirmation letter key information extraction model, and output the key information extraction results, including all text fragments in the business confirmation letter image data, the category label corresponding to each text fragment, and the key name-value pair relationship label of the question-answer; Seal extraction includes locating the seal area, extracting the circular text image and the linear text image in the seal area, and preprocessing the circular text image; Seal content recognition includes recognizing the content of pre-processed circular text images and linear text images, and outputting the seal content recognition results; S4. Determine whether the initial information extraction successfully identifies the seal portion and the form portion; If all are successfully identified, the first extracted information will be used as the business confirmation letter identification result; If there is a part of the content that is not successfully identified, the part of the content that is not successfully identified is processed again to obtain the identification result of the business confirmation letter; Secondary processing includes secondary processing of the table part and secondary processing of the seal part. The secondary processing of the table part is to generate key name-key value pair relationship labels missing in the initial extraction information and supplement them into the key information extraction results. The secondary processing of the seal part is to match the seal area image with the seal template library and use the seal content information corresponding to the matched seal template image to replace the original seal content recognition result in the initial extraction information. S5. Encapsulate the business confirmation letter recognition result into JSON structured data.

[0081] This application also provides an electronic device for implementing each embodiment of this application. Figure 3A schematic diagram of a hardware structure of an electronic device for implementing various embodiments of the present application is shown in FIG. 1. Figure 3 As shown in FIG. 1, the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor.

[0082] Those skilled in the art can understand that the electronic device structure involved in the embodiments of the present application does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than shown, or combine certain components, or different component arrangements.

[0083] In the embodiments of the present application, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workstation, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the embodiments of the present application described and / or claimed herein.

[0084] In the embodiments of the present application, the processor can be implemented by using at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a processor, a controller, a microcontroller, a microprocessor, an electronic unit designed to perform the functions described herein, and in some cases, such an implementation can be implemented in a controller. For software implementation, the implementation of such as a process or a function can be implemented with a separate software module allowing at least one function or operation to be performed, and the software code can be implemented by a software application (or program) written in any appropriate programming language, which can be stored in a memory and executed by a controller.

[0085] In addition, the electronic device includes some functional modules that are not shown and will not be described here.

[0086] Those skilled in the art can understand that various aspects of the electronic device provided by the present application can be implemented as a system, a method or a program product. Therefore, various aspects of the present application can be embodied as a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, which can be collectively referred to as "circuitry", "module" or "system" herein.

[0087] The present application also provides a storage medium storing a program product capable of implementing the business confirmation letter identification method. In some possible implementations, various aspects of the present application may also be implemented in the form of a program product comprising program code. When the program product is executed on a terminal device, the program code is configured to cause the terminal device to execute the steps described in the "Exemplary Methods" section above according to various exemplary implementations of the present application.

[0088] The storage medium can be any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0089] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for identifying a port business confirmation letter, characterized in that: include: S1. Obtain the business confirmation letter, which is a single-page PDF file containing the title, form, seal, and other content. S2. Converting the business confirmation letter file into business confirmation letter image data; S3. Perform initial extraction of the business confirmation letter image data, including key information extraction, seal extraction, and seal content recognition, to obtain the initial extracted information; Among them, key information extraction is to input the business confirmation letter image data into the pre-trained business confirmation letter key information extraction model, and output the key information extraction results, including all text fragments in the business confirmation letter image data, the category label corresponding to each text fragment, and the key name-key value pair relationship label of the question-answer; Seal extraction includes locating the seal area, extracting the circular text image and the linear text image in the seal area, and preprocessing the circular text image; Seal content recognition includes recognizing the content of pre-processed circular text images and linear text images, and outputting the seal content recognition results; S4. Determine whether the initial information extraction successfully identifies the seal portion and the form portion; If all are successfully identified, the first extracted information will be used as the business confirmation letter identification result; If there is a part of the content that is not successfully identified, the part of the content that is not successfully identified is processed again to obtain the identification result of the business confirmation letter; Secondary processing includes secondary processing of the table part and secondary processing of the seal part. The secondary processing of the table part is to generate key name-key value pair relationship labels missing in the initial extraction information and supplement them into the key information extraction results. The secondary processing of the seal part is to match the seal area image with the seal template library and use the seal content information corresponding to the matched seal template image to replace the original seal content recognition result in the initial extraction information. S5. Encapsulate the business confirmation letter recognition result into JSON structured data.

2. The business confirmation letter identification method according to claim 1, characterized in that: In step S3, the steps of extracting key information using the pre-trained business confirmation letter key information extraction model include: Extract all text fragments from the business confirmation letter image data; Categorize text snippets into categories such as questions, answers, titles, and other content; Perform key-value matching on questions and answers, with questions corresponding to key names and answers corresponding to key values; Output key information extraction results.

3. The business confirmation letter identification method according to claim 1, characterized in that: In step S3, seal extraction includes: Use the pre-trained seal position detection model, a deep learning model, to locate the seal area and output the seal area bounding box coordinates and seal area image. A pre-trained text region segmentation model is used to extract circular text images and linear text images within the seal region image. The text region segmentation model is a deep learning model. Perform polar coordinate transformation on the annular text image to obtain a straightened annular text image.

4. The business confirmation letter identification method according to claim 3, characterized in that: In step S3, a pre-trained seal content recognition model is used to perform seal content recognition. The seal content recognition model is a deep learning model, including: The linear text image and the straightened circular text image are input into the seal content recognition model, and the seal content recognition results are output.

5. The business confirmation letter identification method according to any one of claims 2 to 4, characterized in that: It also includes step S6: adding the business confirmation letter image data and the corresponding business confirmation letter recognition results to the training data set, and using the training data set to iteratively update the business confirmation letter key information extraction model, seal position detection model, text area segmentation model and seal content recognition model through incremental learning.

6. The business confirmation letter identification method according to any one of claims 2 to 4, characterized in that: The rules for determining whether the seal content and the form content are successfully recognized in step S4 include: Extract the company name from the seal content recognition result and the company name from the title text segment from the key information extraction result, and then calculate the similarity between the two. If the similarity is lower than the preset threshold, it is determined that the seal content has not been successfully recognized; The key-value matching relationship of all questions and answers in the key information extraction results is verified. If the key name does not match the key value, it is determined that part of the table content was not successfully identified.

7. The business confirmation letter identification method according to any one of claims 2 to 4, characterized in that: In step S4, the secondary processing steps of the table part include: S401 sets the key name library, the key name library stores the key name and its corresponding key value regular expression rules; S402. When the contents of the table portion are not successfully identified, the table portion is subjected to OCR recognition, all text fragments of the table portion are extracted, and the spatial coordinates of each key name that does not match the key value in the image are recorded; S403 matches each unmatched key name with the key name in the key name library. After a successful match, extract all text fragments that match the key name corresponding to the regular expression as candidate key values; S404 records the spatial position coordinates of the candidate key in the image, selects the candidate key with the closest spatial distance to the unmatched key according to the spatial position coordinates as the matching key of the key, forms the key name - key value pair relationship label of the key, and adds it to the key information extraction result; Secondary processing of seals includes: S411 set seal template library, seal template library stores a mapping relationship between the company name and the corresponding seal template image group, each seal template image group contains at least one seal template image corresponding to the company, each seal template image associated with pre-stored seal content information; S412 extracts the key information extraction result from the title text fragment of the company name, based on the extracted company name query seal template library, the company name corresponding to the seal template image group as a candidate seal template image group; S413. ORB feature point matching is performed on the seal area image and each seal template image in the candidate seal template image group to calculate the image matching similarity score. The seal template image with the highest image matching similarity score and above a preset threshold is used as the target seal template image. S414. The seal content information associated with the target seal template image is used as the actual seal content recognition result, replacing the original seal content recognition result in the initially extracted information.

8. A port business confirmation letter identification system, characterized in that: The method for implementing the business confirmation letter identification method according to any one of claims 1 to 7 comprises: A pre-processing module, used for acquiring a business confirmation letter file and converting the business confirmation letter file into business confirmation letter image data; The initial extraction module is used to perform initial extraction on the business confirmation letter image data to obtain initial extraction information; The analysis and judgment module is used to analyze the initially extracted information and determine whether the seal content and the form content are successfully recognized; A secondary processing module is used to perform secondary processing on the part of the content that was not successfully recognized to obtain the recognition result of the business confirmation letter; The structured processing module is used to encapsulate the business confirmation letter recognition results into JSON structured data.

9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the steps of the business confirmation letter identification method according to any one of claims 1 to 7 when executing the computer program.

10. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the steps of the business confirmation letter identification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Document entry rechecking method, system, electronic equipment and medium

    CN113239893A

  • Electronic seal data configuration method and device, equipment and storage medium

    CN113435169A

  • Tax declaration form identification method and device

    CN114445841A

  • Certificate request analysis and verification method

    CN114706824A

  • Method, device and system for identifying container number of quay crane and computer equipment

    CN115527209A