A Deep Learning-Based Method for Recognizing and Processing Logistics Documents for Circular Packaging

By using deep learning technology, the physical damage of reusable packaging logistics documents is quantitatively assessed, enabling robust identification and automated processing in complex environments. This solves the problem of decreased identification accuracy caused by physical damage in existing technologies, ensuring the continuity of logistics document flow and the reliability of data.

CN121438335BActive Publication Date: 2026-04-03ANWOOD LOGISTICS SYSTEMS (SUZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies suffer from decreased recognition accuracy when faced with physical damage to reusable packaging logistics documents, leading to increased data recognition complexity, failure to output effective information, abnormal interruption of business processes, and increased manual intervention costs.

Method used

A deep learning-based approach is adopted to obtain the damage layer through an image segmentation model. Combined with predefined key information regions and business weight factors, the physical state mask and the physical entropy increase index of the key information regions are quantified. Multi-channel information extraction and nonlinear fusion are performed to adjudicate multimodal information conflicts and output decisions, generating credible data and business recommendations.

Benefits of technology

It significantly improves the robustness and automation of document recognition in complex environments, reduces the risk of abnormal interruptions, provides monitoring of the continuity of asset flow and long-term health status, and reduces the cost of manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121438335B_ABST
    Figure CN121438335B_ABST
Patent Text Reader

Abstract

This invention relates to the field of deep learning and logistics document recognition technology, specifically a deep learning-based method for recognizing and processing reusable packaging logistics documents. The method includes: acquiring original images of the documents captured on-site; acquiring business context; performing physical state analysis on the original document images to obtain a physical state mask and a physical entropy increase index for key information regions; obtaining original data fragments, optical character recognition confidence scores, and model uncertainties; combining the physical entropy increase index of key information regions, optical character recognition confidence scores, and model uncertainties to calculate a coupled data credibility score; obtaining credible data based on the coupled data credibility score, the original data fragments, and the business context; and generating structured data and actionable business suggestions from the credible data and the coupled data credibility score. This invention significantly improves the robustness and environmental adaptability of reusable packaging document recognition in complex on-site environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning and logistics document recognition technology, specifically a method for recognizing and processing reusable packaging logistics documents based on deep learning. Background Technology

[0002] With the development of modern logistics, reusable packaging logistics documents may encounter physical damage such as soiling, folding, and tearing during the circulation process, which significantly increases the complexity of data recognition.

[0003] Currently, automated processing mainly relies on optical character recognition technology; however, traditional OCR technology overemphasizes apparent accuracy and exhibits great vulnerability when faced with real-world physical damage. This vulnerability causes high-accuracy models to fail to output effective information when recognition fails, leading to abnormal interruptions in business processes and incurring high costs for manual intervention.

[0004] Therefore, how to overcome the vulnerability of existing technologies to physical damage and provide a robust and quantifiable data credibility assessment to ensure the continuity of the transfer of physical assets has become an urgent technical problem to be solved in this field. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention discloses a method for identifying and processing reusable packaging logistics documents based on deep learning. Specifically, the technical solution of this invention includes:

[0006] Acquire the original images of the documents captured on-site;

[0007] Obtain the business context, including obtaining the predefined key information area mask and obtaining the predefined business weight factor;

[0008] Based on the predefined key information region mask and the predefined business weight factor, the original image of the document is subjected to physical state analysis processing to obtain the physical state mask and the physical entropy increase index of the key information region.

[0009] Based on the original image of the document and the physical state mask, multi-channel information extraction processing is performed to obtain the original data fragment, optical character recognition confidence, and model uncertainty.

[0010] By combining the physical entropy increase index of the key information area, the confidence level of optical character recognition, and the uncertainty of the model, a nonlinear fusion process is performed to calculate the credibility score of the coupled data.

[0011] Based on the credibility score of the coupled data, the original data fragment, and the business context, multimodal information conflict adjudication processing is performed to obtain credible data;

[0012] Based on the credibility scores of the credible data and the coupled data, decision output processing is performed to generate structured data and actionable business recommendations.

[0013] Preferably, the step of performing physical state analysis processing on the original image of the document includes:

[0014] The original image of the document is processed using an image segmentation model, which divides it into multiple semantic layers to obtain the physical state mask;

[0015] The physical state mask includes a printed area layer, a handwritten area layer, a stamp area layer, and a damage layer.

[0016] Using the predefined key information region mask, the percentage of overlap between the damage layer and the key information region is calculated; using the predefined business weight factor, the percentage of overlap is weighted and summed to calculate the physical entropy increase index of the key information region.

[0017] Preferably, the step of performing multi-channel information extraction processing includes:

[0018] Run the printed optical character recognition model on the printed area layer of the physical state mask;

[0019] Run the handwritten optical character recognition model on the handwritten area layer of the physical state mask;

[0020] Run the stamp classification model on the stamp area layer of the physical state mask;

[0021] The outputs of the printed optical character recognition model, the handwritten optical character recognition model, and the seal classification model together constitute the original data segment.

[0022] The Monte Carlo Dropout technique is used to perform multiple inferences on each model run in the aforementioned steps, resulting in multiple inference results;

[0023] The confidence level of optical character recognition is determined based on the average softmax probability of the multiple inference results.

[0024] The uncertainty of the model is determined based on the variance of the multiple inference results or based on the information entropy of the multiple inference results.

[0025] Preferably, the step of performing nonlinear fusion processing includes:

[0026] Obtain the preset balance weights;

[0027] Based on the preset balance weights, the confidence level of optical character recognition and the uncertainty of the model are fused to calculate the cognitive credibility.

[0028] Obtain the preset degradation sensitivity coefficient;

[0029] Based on the physical entropy increase index of the key information region and the preset degradation sensitivity coefficient, the physical penalty factor is calculated.

[0030] The cognitive credibility score is obtained by multiplying the physical penalty factor by the cognitive credibility score.

[0031] Preferably, the step of calculating the physical penalty factor includes:

[0032] Subtract the physical entropy increase index of the key information region from 1 to obtain the difference;

[0033] The difference is compared with 0, and the maximum value is taken to obtain the base of the exponentiation operation;

[0034] The physical penalty factor is obtained by performing an exponential operation using the base and the preset degradation sensitivity coefficient.

[0035] Preferably, the step of performing multimodal information conflict resolution includes:

[0036] From the original data fragment, obtain the printed information and the handwritten information;

[0037] Obtain the credibility score of the coupled data corresponding to the printed information;

[0038] Obtain the credibility score of the coupled data corresponding to the handwritten information;

[0039] Based on the business context, it is determined that the printed information and the handwritten information are conflicting or supplementary information;

[0040] If it is determined to be supplementary information, the printed information and the handwritten information are merged to generate a merged fact, and the merged fact is used as the reliable data;

[0041] If it is determined to be supplementary information, the lower of the coupled data confidence score corresponding to the printed information and the coupled data confidence score corresponding to the handwritten information is used as the confidence level of the reliable data.

[0042] If the information is determined to be conflicting, the reliability score of the coupled data corresponding to the printed information is compared with the reliability score of the coupled data corresponding to the handwritten information.

[0043] If the information is determined to be conflicting, the information with a higher coupling data confidence score will be used as the trusted data, and its corresponding coupling data confidence score will be used as the confidence level of the trusted data.

[0044] Preferably, the step of performing decision output processing includes:

[0045] Obtain the preset automatic approval threshold;

[0046] Obtain the preset threshold for manual review;

[0047] Get a list of predefined key fields;

[0048] The trusted data and its corresponding coupled data credibility scores are output as the structured data;

[0049] The fields in the trusted data are compared with the list of key fields to determine whether there are any key fields whose credibility scores for the coupled data are less than or equal to the manual review threshold.

[0050] The credibility score of the coupled data for all fields in the trusted data is compared with the automatic approval threshold to determine whether there is a coupled data credibility score that is less than or equal to the automatic approval threshold.

[0051] If the credibility score of the coupled data with key fields is determined to be less than or equal to the manual review threshold, a rejection process suggestion is generated as the actionable business suggestion.

[0052] If it is determined that the credibility score of the coupled data without key fields is less than or equal to the manual review threshold, and it is determined that the credibility score of the coupled data is less than or equal to the automatic approval threshold, then a marked review suggestion is generated as the actionable business suggestion.

[0053] If it is determined that the credibility score of the coupled data for which no key field exists is less than or equal to the manual review threshold, and it is determined that the credibility score of the coupled data for which no field exists is less than or equal to the automatic approval threshold, then an automatic approval suggestion is generated as the actionable business suggestion.

[0054] Preferably, the physical state analysis and processing steps further include:

[0055] Calculate the global physical entropy increase index of the original image of the document;

[0056] The steps for processing the decision output also include:

[0057] Obtain the preset maintenance threshold;

[0058] The global physical entropy increase index is compared with the preset maintenance threshold;

[0059] If the global physical entropy increase index is greater than the preset maintenance threshold, a predictive maintenance tag warning will be generated in the actionable business suggestions.

[0060] If the global physical entropy increase index is less than or equal to the preset maintenance threshold, then a predictive maintenance tag warning will not be generated in the actionable business recommendations.

[0061] Compared with the prior art, the present invention has the following beneficial effects:

[0062] 1. This invention obtains the damaged layer through an image segmentation model and, combined with predefined key information regions and business weight factors, calculates the physical entropy increase index of the key information regions. This index will subsequently be used as a physical penalty factor, multiplied by the model's cognitive credibility. This design can quantitatively assess the negative impact of physical damage or dirt on the identification of key information, significantly improving the robustness and environmental adaptability of identifying reusable packaging documents in complex on-site environments.

[0063] 2. This invention employs a multi-channel information extraction strategy, running dedicated recognition models on different physical layers such as printed text, handwritten text, and seals. Simultaneously, it innovatively utilizes Monte Carlo Dropout technology to evaluate the confidence level and model uncertainty in optical character recognition, fusing both into a cognitive credibility score. Compared to a single confidence score, this approach provides a deeper understanding of the model's grasp of the recognition results, offering richer and more reliable evidence for subsequent credibility assessments.

[0064] 3. This invention utilizes coupled data credibility scores and business context to perform multimodal conflict resolution on printed and handwritten information, intelligently determining whether they are conflicting or complementary, and generating credible data according to rules. Combining a key field list, automatic approval thresholds, and manual review thresholds, the system can automatically output suggestions such as automatic approval, marking for review, or rejection. This achieves closed-loop processing from raw images to actionable business suggestions, greatly improving the automation level and decision-making efficiency of logistics document processing.

[0065] 4. This invention not only assesses the entropy increase in key information areas of documents but also calculates the global physical entropy increase index and compares it with a preset maintenance threshold. When the global entropy increase is too high, the system will proactively generate a predictive maintenance tag warning. This function goes beyond single data identification and extends to long-term health monitoring of cyclic packaging asset tags, providing decision support for preventative tag replacement and maintenance, and ensuring the long-term stability of data collection. Attached Figure Description

[0066] The present invention will be further explained below with reference to the accompanying drawings and embodiments:

[0067] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0068] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0069] Example 1

[0070] Please see Figure 1 A deep learning-based method for identifying and processing reusable packaging logistics documents includes:

[0071] Acquire the original images of the documents captured on-site;

[0072] Obtain the business context, including obtaining the predefined key information area mask and obtaining the predefined business weight factor;

[0073] Based on the predefined key information region mask and the predefined business weight factor, the original document image is subjected to physical state analysis processing to obtain the physical state mask and the physical entropy increase index of the key information region.

[0074] Based on the original image of the document and the physical state mask, multi-channel information extraction processing is performed to obtain the original data fragments, optical character recognition confidence, and model uncertainty.

[0075] By combining the physical entropy increase index of key information areas, the confidence level of optical character recognition, and model uncertainty, nonlinear fusion processing is performed to calculate the credibility score of coupled data.

[0076] Based on the coupled data credibility score, original data fragments, and business context, multimodal information conflict adjudication is performed to obtain credible data;

[0077] Based on the credibility scores of credible data and coupled data, decision output processing is performed to generate structured data and actionable business recommendations;

[0078] This embodiment provides a method for identifying and processing reusable packaging logistics documents based on deep learning;

[0079] This method is the core technical solution of the present invention, forming a complete and self-consistent technical closed loop. It aims to solve the vulnerability of existing OCR technology when facing real-world damaged and folded circular documents due to the excessive pursuit of appearance indicators.

[0080] The technical objective of this embodiment is to reconstruct the identification problem: instead of pursuing a binary and fragile accuracy metric, it provides a probabilistic and robust data credibility assessment to ensure the continuity of the flow of physical assets.

[0081] The complete process of this method includes:

[0082] Acquire raw images of documents captured on-site; this refers to real-time images of logistics documents collected during actual operations using devices such as fixed industrial cameras at warehouse entrances and handheld PDA terminals.

[0083] Obtain the business context; this refers to the business process information associated with the current image acquisition action, such as inbound, outbound, or inventory count; this information is crucial for subsequent determination of the logical relationship between multimodal information, such as printing and handwriting.

[0084] The original document image undergoes physical state analysis processing; in this embodiment, this step is performed by a physical state analysis module; this module receives the original document image, and its core purpose is to evaluate the degree of physical entropy increase or physical degradation of the image; it outputs two key results:

[0085] Physical state mask: Divides an image into different layers such as print, handwriting, stamp, stain, tear, etc.

[0086] Physical entropy increase index of key information area A quantitative, dimensionless index that accurately assesses the actual threat of physical damage to the readability of core data such as barcodes and serial numbers;

[0087] Based on the original document image and physical state mask, multi-channel information extraction processing is performed. This step is executed by a multi-channel information extraction module. This module receives the original document image and physical state mask, and its core purpose is to extract all possible information in parallel and assess the cognitive risk of the model itself. Guided by the physical state mask, it runs different AI models, such as printed text OCR and handwritten text OCR, on different layers and outputs three key results:

[0088] Raw data fragment: The collection of all identified raw text and symbols;

[0089] Optical character recognition confidence The confidence level of the model in the recognition results, such as the softmax probability;

[0090] Model uncertainty The degree of hesitation or stability of the model towards the recognition results, such as the variance of MCDropout;

[0091] The invention incorporates the physical entropy increase index of key information regions, the confidence level of optical character recognition, and model uncertainty, and performs nonlinear fusion processing. This step is the core innovation of the invention and is executed by a confidence quantification and fusion module. This module receives physical risks from the preceding steps. and perceived risk Its core purpose is to non-linearly couple these two completely different dimensions of risk; it calculates the final credibility score of the coupled data. ;

[0092] Based on the coupled data credibility score, original data fragments, and business context, multimodal information conflict adjudication is performed; this step is also executed by the credibility measurement and fusion module; its core purpose is to utilize the calculated... The score serves as a quantitative weight, and combined with the business context, it intelligently adjudicates potential conflicts in the original data fragments, such as a printed 10 and a handwritten damaged 2, and outputs a single, complete, and reliable data.

[0093] Based on the credibility scores of credible data and coupled data, decision output processing is performed; this step is executed by a decision output module; its core purpose is to base decisions on preset business rules such as automatic approval thresholds and manual review thresholds. The threshold transforms quantified credibility into actual business actions, enabling intelligent data degradation; it ultimately generates two types of outputs:

[0094] Structured data: contains reliable data and its corresponding data. The score is in JSON format.

[0095] Actionable business suggestions: such as automating approval, marking for review, or rejecting processes;

[0096] Through the complete closed loop of the above seven steps, this invention provides a robust and quantifiable data reliability assessment method; it fundamentally solves the vulnerability of high-accuracy models in the face of real physical damage in existing technologies; and it proactively quantifies physical risks. and perceived risk By deeply coupling it with business logic, this invention enables automated systems to perform hierarchical processing—that is, it can quantitatively assess the unreliability of images. Or determine the model uncertainty is high This replaces the traditional model where OCR failure means unusable, greatly ensuring the continuity of business processes and reducing the cost of manual intervention due to abnormal interruptions.

[0097] Example 2

[0098] The steps for physical state analysis processing of the original document image include:

[0099] An image segmentation model is used to process the original image of the document, dividing it into multiple semantic layers to obtain a physical state mask;

[0100] The physical state mask includes a printed area layer, a handwritten area layer, a stamp area layer, and a damage layer.

[0101] Using the predefined key information region mask, calculate the percentage of overlap between the damage layer and the key information region; use the predefined business weight factor to perform a weighted summation operation on the percentage of overlap to calculate the physical entropy increase index of the key information region.

[0102] This embodiment is based on Embodiment 1, and further specifies and optimizes the physical state analysis and processing steps.

[0103] It elaborates on how to accurately quantify the actual threat of physical damage to critical business data, that is, how to calculate the physical entropy increase index of critical information areas. ;

[0104] The steps for performing physical state analysis on the original image of the document, in this embodiment, specifically include:

[0105] An image segmentation model is used to process the original image of the document;

[0106] Image segmentation model refers to a deep learning model, such as U-Net or Mask R-CNN, which are well known to those skilled in the art; its purpose is to divide the input original document image into multiple layers with different semantic meanings.

[0107] Output: Physical state mask;

[0108] The physical state mask includes:

[0109] Print area layer : Indicates the position of all printed characters;

[0110] Handwritten area layer : Identify the location of all handwritten marks;

[0111] Stamp area layer : Identify the location of all seals;

[0112] Damage layer : Identify the location of various physical damages, such as stains Reflective Tearing, etc.;

[0113] Obtain the predefined key information region mask;

[0114] Key information area mask This refers to a predefined binary mask whose function is to identify the template location of areas on a document that are critical to the business process, such as all barcodes, QR codes, and serial number fields; its source is predefined according to business requirements, and it tells the system which areas' physical status should be given priority.

[0115] Obtain the predefined business weight factors for the damage layer;

[0116] Predefined business weight factors This refers to a dimensionless weighting coefficient, associated with each type of damage layer. Correspondingly, predefined business weight factors are pre-stored in the system's configuration database before the system performs physical state analysis and processing, and need to be directly called during analysis. Their function is to quantify business logic and inject it into the assessment of physical damage, reflecting the severity of different damage types and locations. These factors are derived from historical data analysis or expert experience, pre-calibrated based on business logic, or adaptively adjusted during system operation. For example, the weight factor corresponding to a crease passing through a barcode... For example, 1.5 should be much larger than the value corresponding to a slight stain in the blank area. For example, 0.8;

[0117] A calibration Exemplary methods include: collecting N image samples, each containing only one major damage type. Such as stains, tears, and covering of critical areas Business experts label the business unacceptability of each sample, for example... Where 0 represents no impact, 1 represents that verification is required, and 2 represents that the process must be interrupted; calculate the normalized damage area for each sample. ; The value can be calculated and The ratio ( Estimate by the statistical mean or median of ( );

[0118] A weighted mask calculation is performed by combining the damage layer, the key information area mask, and the predefined business weight factor;

[0119] This step involves calculating the physical entropy increase index of the key information region. The core of it is achieved through a weighted model based on area overlap rate;

[0120] The calculation formula is as follows:

[0121]

[0122] : Physical entropy increase index of key information area, dimensionless, the final output of this module;

[0123] : The i-th damage layer, mask, source: output of the image segmentation model;

[0124] Key information area mask; mask source: predefined business template.

[0125] : A calculation function used to calculate the area of ​​pixels covered by the mask;

[0126] : The business weight factor for the i-th type of damage, dimensionless, source: pre-calibrated;

[0127] This formula is passed Calculate the area of ​​the i-th type of damage within the critical region, and then divide it by the total area of ​​the critical region. Normalize the result and then multiply it by its corresponding business weight. In this calculation, the damage layer The source is the output of the image segmentation model, the key information region mask. The source is a predefined business template, business weight factor. The source is pre-calibrated;

[0128] This design precisely quantifies the actual threat of physical damage to the readability of core data; it no longer simply counts the area of ​​soiling across the entire image, but focuses on critical damage in key locations.

[0129] Through the above steps, this invention achieves accurate, quantitative, and tightly coupled assessment of physical damage with business logic; it solves the problem of the binary and coarse assessment method in existing technologies where damage equates to unreliability; by introducing... and The system can intelligently distinguish between minor stains that have no impact on business and critical barcode tears that cause process interruptions, providing accurate and robust physical risk input for subsequent credibility assessments.

[0130] Example 3

[0131] The steps for multi-channel information extraction and processing include:

[0132] Run the printed volume optical character recognition model on the printed area layer of the physical state mask;

[0133] Run the handwritten optical character recognition model on the handwritten region layer of the physical state mask;

[0134] Run the stamp classification model on the stamp region layer of the physical state mask;

[0135] The outputs of the printed optical character recognition model, the handwritten optical character recognition model, and the seal classification model together constitute the original data fragment.

[0136] The Monte Carlo Dropout technique is used to perform multiple inferences on each model run in the aforementioned steps, resulting in multiple inference results;

[0137] The confidence level for optical character recognition is determined based on the average softmax probability of multiple inference results.

[0138] The uncertainty of the model is determined based on the variance of the results of multiple inferences, or based on the information entropy of the results of multiple inferences.

[0139] This embodiment is based on Embodiment 1, and further specifies and optimizes the multi-channel information extraction and processing steps.

[0140] It elaborates on how to extract all information in parallel and quantifies the cognitive risk of the model by decomposing it into two dimensions: confidence and uncertainty.

[0141] The steps for multi-channel information extraction and processing, as described in this embodiment, specifically include:

[0142] Run the printed volume optical character recognition model on the printed area layer of the physical state mask; this step utilizes the physical state mask. The region only identifies printed text;

[0143] Run the handwritten optical character recognition model on the handwritten region layer of the physical state mask; this step utilizes A dedicated area for handling handwritten notes;

[0144] Run the stamp classification model on the stamp region layer of the physical state mask; this step utilizes The area identifies the type of seal, such as quality inspection qualified or invalid;

[0145] The outputs from the preceding steps collectively constitute the original data fragments;

[0146] The raw data fragment refers to a structured collection containing all extracted, unprocessed preliminary identification results; its role is to serve as the raw input for subsequent credible quantification and conflict resolution; for example: {"print":"SN12345","handwrite":"damaged 2","stamp":"inspected"};

[0147] The Monte Carlo Dropout technique is used to perform multiple inferences on each model that has been run in the aforementioned steps;

[0148] Monte Carlo refers to a technique well-known in the field that maintains the activation of the Dropout layer during the inference phase and performs N forward propagations, for example, N=20, on the same sample, i.e., the original image of the document; its core purpose is to obtain N not completely identical distributions of prediction results, thereby quantifying the cognitive uncertainty of the model itself.

[0149] The confidence level for optical character recognition is determined based on the average softmax probability of multiple inference results.

[0150] Optical character recognition confidence This refers to averaging the softmax probability values ​​obtained from N inferences; its function is to quantify the model's average confidence that the character is identified as A; it is a dimensionless value in the interval 0, 1.

[0151] The uncertainty of the model is determined based on the variance of the results of multiple inferences, or based on the information entropy of the results of multiple inferences.

[0152] Model uncertainty This refers to calculating the variance or entropy of the results of N inference iterations, such as N softmax outputs or N predicted categories. Its purpose is to quantify the model's hesitation or stability across N iterations. If the N results are highly consistent, such as when the variance is close to 0, then the model's performance is considered stable. It's very low; if the results of N trials show a large discrepancy, such as high variance, It is very high; this is a dimensionless value in the interval 0, 1.

[0153] By the confidence level of the model With model stability By performing separation and quantization, this solution addresses the problem that existing technologies typically rely solely on... However, this is problematic when the model makes a mistake with high confidence. It's very tall, but This solution also addresses the issue of complete failure at high speeds; by introducing... This provides a deeper understanding of the cognitive risks of the model, namely that the model can represent states it does not know, providing richer and more reliable decision-making basis for subsequent fusion modules, and significantly improving the system's awareness of the model's own limitations.

[0154] Example 4

[0155] The steps for performing nonlinear fusion processing include:

[0156] Obtain the preset balance weights;

[0157] Based on preset balanced weights, the confidence level of optical character recognition and model uncertainty are fused to calculate the cognitive confidence level.

[0158] Obtain the preset degradation sensitivity coefficient;

[0159] Based on the physical entropy increase index of the key information area and the preset degradation sensitivity coefficient, the physical penalty factor is calculated.

[0160] The coupled data credibility score is obtained by multiplying the cognitive credibility score with the physical penalty factor.

[0161] The steps to calculate the physics penalty factor include:

[0162] Subtract the physical entropy increase index of the key information area from 1 to obtain the difference;

[0163] Compare the difference with 0, take the maximum value, and obtain the base of the exponentiation operation;

[0164] The physical penalty factor is obtained by performing an exponential operation using the base and a preset degradation sensitivity coefficient;

[0165] This embodiment is based on Embodiment 1, and it further specifies and optimizes the nonlinear fusion processing steps. It also includes further limitations on the key steps therein.

[0166] This embodiment demonstrates the core algorithm implementation for achieving the robustness of this invention. It details how to nonlinearly couple physical risk and cognitive risk to calculate the final coupled data credibility score. ;

[0167] In this embodiment, the steps of nonlinear fusion processing are specifically implemented through a coupled data reliability factor. The evaluation model is used to achieve this; this fusion process targets each raw data segment output by the multi-channel information extraction module and its corresponding... , and The physical entropy increase index of the region where the fragment is located is calculated independently to determine the coupling data reliability score for each data fragment. Its core design philosophy is to increase the overall credibility. Considered as cognitive credibility and physical penalty factor The product of:

[0168]

[0169] Cognitive credibility Solution:

[0170] This step aims to integrate the model's own confidence and stability;

[0171] Obtain the preset balance weights;

[0172] Balance weight This refers to a calibration parameter in the range of 0 and 1; its function is to balance the confidence level of optical character recognition. and model uncertainty In the end The proportion of confidence and stability in the validation set; its source is the experimental calibration of the preference for confidence and stability in the application scenario through grid search on the validation set or based on historical data statistical analysis; for example, a calibration method based on grid search includes: preparing a validation set containing M samples, where each sample is manually labeled with a true confidence level. Define a loss function, such as mean squared error. ;exist Within the interval, for example, searching with a step size of 0.1, select the loss function that... smallest The value is used as the optimal parameter;

[0173] Based on preset balance weights, for and The data is then integrated to calculate the cognitive credibility.

[0174] Cognitive credibility The calculation formula is as follows:

[0175]

[0176] Optical character recognition confidence and The uncertainty of the model all originates from the output of the multi-channel information extraction module;

[0177] Physical penalty factor Solution:

[0178] This step aims to mitigate physical risks. Convert it into a penalty coefficient in the range of 0 and 1;

[0179] Obtain the preset degradation sensitivity coefficient:

[0180] Degradation sensitivity coefficient This refers to a key custom dimensionless parameter; its function is to control... Physical damage The degree of nonlinear sensitivity; its origin lies in the sensitivity requirements of different fields, such as barcodes and titles, to damage, determined through experiments or expert experience based on business logic. For example, barcodes are extremely sensitive to damage. A higher value, such as 3.0, can be used to make... A sharp decline due to minor damage; high title redundancy. A low value, such as 0.5, can be used during calibration. hour, The values ​​can be determined together through a grid search. For example, in the aforementioned grid search based on the validation set, simultaneously... Within the interval, for example, searching with a step size of 0.5, select a loss function that... ,like smallest The combination is used as the optimal parameter;

[0181] based on and The physical penalty factor was calculated.

[0182] Subtract the physical entropy increase index of the key information area from 1 to obtain the difference: that is, calculate... ;

[0183] Compare the difference with 0, take the maximum value, and obtain the base for the exponentiation: that is, calculate... ;

[0184] The physical penalty factor is obtained by performing an exponential operation using the base and a preset degradation sensitivity coefficient;

[0185] Combining the above steps, the physical penalty factor... The calculation formula is as follows:

[0186]

[0187] The source of the physical entropy increase index in the key information area is the output of the physical state analysis module;

[0188] The technical motivation behind this design lies in processing... In real-world scenarios, the value might be greater than 1, for example, when the damage weight... And when the damaged area is large;

[0189] pass The operation ensures that the base of exponentiation is always non-negative, which is especially important in mathematics. It is always defined when the integer is not an integer;

[0190] The technical consideration lies in its implementation of a physical veto mechanism: once weighted physical damage is assessed... If the value is greater than or equal to 1, the base is clamped to 0, resulting in... Forced reset to zero;

[0191] when When there is no damage, the base is 1. No penalty;

[0192] Final coupling:

[0193] Multiply the cognitive credibility by the physical penalty factor;

[0194] Final Coupling Data Credibility Score for: ;

[0195] The technical consideration behind using multiplication instead of addition for coupling lies in achieving a dual veto power:

[0196] Physical veto: If the physical damage is severe ,but ,lead to Forced to zero – regardless of the model's confidence level How tall;

[0197] Cognitive rejection: If the model is extremely uncertain or has low confidence. ,but No matter how clear the image is How tall;

[0198] A robust evaluation model that truly reflects fundamental value was constructed; it deeply integrates the risks of physical distortion and cognitive uncertainty of the model—two completely different dimensions—through nonlinear multiplicative coupling; this enables the system to perform hierarchical processing—allowing the system to determine its own high uncertainty. Or determine that the image is highly unreliable —This completely solves the problem of vulnerability and high confidence error of high-accuracy models in existing technologies when faced with real contaminated data.

[0199] Example 5

[0200] The steps for handling multimodal information conflict adjudication include:

[0201] Extract printed and handwritten information from the raw data fragments;

[0202] Obtain the credibility score of the coupled data corresponding to the printed information;

[0203] Obtain the credibility score of the coupled data corresponding to the handwritten information;

[0204] Based on the business context, determine whether the printed information and the handwritten information are conflicting or supplementary information;

[0205] If it is determined to be supplementary information, the printed information and the handwritten information will be merged to generate merged facts, and the merged facts will be used as reliable data.

[0206] If it is determined to be supplementary information, the lower of the coupled data confidence score corresponding to the printed information and the coupled data confidence score corresponding to the handwritten information shall be used as the confidence score of the credible data.

[0207] If the information is determined to be conflicting, the credibility score of the coupled data corresponding to the printed information is compared with the credibility score of the coupled data corresponding to the handwritten information.

[0208] If the information is determined to be conflicting, the information with a higher coupled data confidence score will be regarded as credible data, and its corresponding coupled data confidence score will be regarded as the confidence level of the credible data.

[0209] This embodiment is a refinement and optimization of the multimodal information conflict adjudication process based on Embodiment 1;

[0210] It explains in detail how to use the results calculated in the preceding steps. Scoring is used to intelligently and flexibly handle the relationships between information from different sources, such as printed, handwritten, and stamped information;

[0211] The steps for multimodal information conflict resolution in this embodiment specifically include:

[0212] Extract printed and handwritten information from the raw data fragments; for example, extract... =Quantity: 10 and = "Damaged 2";

[0213] Obtain the credibility score of the coupled data corresponding to the printed information; for example, calculate the... Because the print is clear, the model has high confidence.

[0214] Obtain the credibility score of the coupled data corresponding to the handwritten information; for example, calculate... Because the handwriting was illegible, the model had uncertainty. High;

[0215] Based on the business context, determine whether the printed information and the handwritten information are conflicting or supplementary information;

[0216] The business context plays a crucial role in semantic adjudication here;

[0217] For example, when the business context is "Inbound", the system rule base determines that the quantity and damage are valid supplementary information;

[0218] When the business context is "outbound", if =Quantity: 10 and If the value is ="Quantity: 8", the system determines that the two are conflicting messages.

[0219] If the information is determined to be supplementary information, the corresponding data entry scenario is as follows:

[0220] The printed information is merged with the handwritten information to generate a merged fact: for example, the merged fact is 10 items received, with the note: 2 items were damaged.

[0221] Integrating facts as credible data ;

[0222] The lower of the coupled data confidence scores is used as the confidence level of the credible data: that is, the final confidence level of the fused fact is... ;

[0223] If the information is determined to be conflicting, the corresponding outbound scenario is as follows:

[0224] Compare the printed information Corresponding to handwritten information : i.e., comparison and ;

[0225] Will have a high Information as credible data: rulings "Quantity: 10" wins and is awarded the prize. ;

[0226] and its corresponding Confidence level as reliable data: i.e. The credibility is 0.99;

[0227] This paper implements a flexible, interpretable, and quantified data-based multimodal information fusion mechanism. Existing technologies often rely on fragile hard-coded rules, such as handwriting priority, or directly discard handwritten / stamp information, resulting in a lack of context. This solution introduces... As a decision weight, it enables the system to dynamically determine whether something is supplementary or conflicting based on quantified credibility and business context; this improves the completeness of information extraction and ensures strong consistency between data in the digital world and physical state in the physical world, such as not losing important physical state information like two broken items.

[0228] Example 6

[0229] The steps for processing decision outputs include:

[0230] Obtain the preset automatic approval threshold;

[0231] Obtain the preset threshold for manual review;

[0232] Get a list of predefined key fields;

[0233] The credibility scores of the credible data and its corresponding coupled data are output as structured data.

[0234] Compare the fields in the trusted data with the list of key fields to determine whether there is any coupling of key fields. The data trust score is less than or equal to the manual review threshold.

[0235] The credibility score of coupled data in all fields of the trusted data is compared with the automatic approval threshold to determine whether there is any coupled data whose credibility score is less than or equal to the automatic approval threshold.

[0236] If the credibility score of the coupled data with key fields is determined to be less than or equal to the manual review threshold, a rejection process suggestion is generated as an actionable business suggestion.

[0237] If the credibility score of coupled data without key fields is determined to be less than or equal to the manual review threshold, and the credibility score of coupled data is determined to be less than or equal to the automatic approval threshold, then a review suggestion is generated as an actionable business suggestion.

[0238] If the credibility score of the coupled data without key fields is determined to be less than or equal to the manual review threshold, and the credibility score of the coupled data without fields is determined to be less than or equal to the automatic approval threshold, then an automatic approval suggestion is generated as an actionable business suggestion.

[0239] This embodiment is based on Embodiment 1, and further specifies and optimizes the decision output processing steps;

[0240] It explains in detail how to utilize The scoring system enables intelligent data degradation logic, transforming quantified credibility into actionable business recommendations.

[0241] The steps for processing the decision output, as described in this embodiment, specifically include:

[0242] Obtain the preset automatic approval threshold;

[0243] Automatic approval threshold This refers to a high credibility threshold; it is derived from statistical analysis of historical credible data and business results, and is determined based on the business's tolerance for risk, for example... ; The data is considered highly reliable; a specific calibration method is to analyze the historical validation set and plot... The curve showing the relationship between scores and actual data accuracy is used to select a curve that allows for an automatic approval accuracy rate greater than 99.9% or the highest level of accuracy acceptable to the business. Threshold, as ;

[0244] Obtain the preset threshold for manual review;

[0245] Manual review threshold This refers to a low credibility baseline; its source is also a business benchmark based on historical data, for example... One specific calibration method is to analyze the historical validation set and select a... A threshold is set to ensure that the proportion of data with incorrect key fields is below that threshold. This means the recall rate for errors is higher than 95%, or the highest recall rate required by the business, ensuring maximum interception of untrusted data. ; The data is considered unreliable; located in Data within a given range is considered of moderate reliability and requires manual review.

[0246] Get a list of predefined key fields;

[0247] A key field list refers to a predefined set of field names, such as ["serial_number", "barcode"]; its purpose is to identify those core fields that have veto power in business decisions.

[0248] The credibility scores of the credible data and its corresponding coupled data are output as structured data.

[0249] The following is an example of the structured data output:

[0250] {

[0251] "serial_number": {

[0252] "value": "SN12345",

[0253] "confidence": 0.99

[0254] },

[0255] "quantity": {

[0256] "value": "10",

[0257] "confidence": 0.95

[0258] },

[0259] "notes": {

[0260] "value": "Damaged",

[0261] "confidence": 0.70

[0262] }

[0263] }

[0264] Based on the above thresholds and field list, a three-stage decision-making logic is executed:

[0265] Logic 1, Rejection Process:

[0266] Determine if a key field exists. Less than or equal to the manual review threshold ;

[0267] If the condition is found to exist, a rejection process suggestion is generated;

[0268] If the serial_number in the example Only 0.5 less than or equal to This indicates that the core data is completely unreliable, the process must be interrupted, and a rejection process recommendation must be generated.

[0269] Logic 2, Mark for review / intelligent degradation:

[0270] If logic 1 is not true, and the condition exists... Less than or equal to the automatic approval threshold ;

[0271] Then, a review suggestion will be generated;

[0272] In the structured data example above, the confidence scores of the key fields `serial_number` (0.99) and `quantity` (0.95) are both greater than [a certain value]. However, the confidence level of the notes field, which represents notes information, is 0.70, which is at the lower end of the range. The system can continue its decision-making process, but it needs to generate a labeling and review recommendation, i.e., intelligent data downgrading: automatically approve trusted data and only label moderately trusted data.

[0273] Logic 3: Automatic Approval

[0274] If logic 1 is not true, then logic 2 is also not true;

[0275] Then an automatic approval suggestion will be generated;

[0276] If all fields All greater than This indicates that the data is completely reliable, and the system outputs an automatic approval recommendation;

[0277] This solution achieves intelligent data degradation, fundamentally ensuring the continuity of business processes. Traditional OCR, when encountering recognition failures such as the inability to recognize the "notes" field, renders its output invalid and may interrupt the entire process, requiring manual data entry. This solution addresses these issues by... The threshold and key field list transforms complete failures into partially reliable, on-demand reviews; it allows the system to automatically approve reliable core data and only mark low-reliability auxiliary data for human intervention, greatly reducing unnecessary human intervention and anomaly handling costs, and truly realizing the fundamental value of automated systems.

[0278] Example 7

[0279] The steps of physical state analysis and processing also include:

[0280] Calculate the global physical entropy increase index of the original image of the document;

[0281] The steps in decision output processing also include:

[0282] Obtain the preset maintenance threshold;

[0283] The global physical entropy increase index is compared with a preset maintenance threshold;

[0284] If the global physical entropy increase index is greater than the preset maintenance threshold, a predictive maintenance tag warning will be generated in the actionable business suggestions.

[0285] If the global physical entropy increase index is less than or equal to the preset maintenance threshold, a predictive maintenance tag warning will not be generated in the actionable business recommendations.

[0286] Based on Example 1, this embodiment achieves a novel predictive maintenance function through inter-module collaboration without increasing additional hardware costs.

[0287] It involves functional enhancements to two steps: physical state analysis and decision output processing.

[0288] In this embodiment, the physical state analysis and decision output processing steps further include:

[0289] In the physical state analysis and processing steps:

[0290] Calculate the global physical entropy increase index of the original image of the document;

[0291] Global physical entropy increase index This refers to an index used to assess the degree of physical degradation of the entire document;

[0292] Its calculation method is similar to the physical entropy increase index of the key information area. They are exactly the same, the only difference being the masking of the critical information area. Replace with the mask of the entire document Its function is no longer to assess the readability of the current data, but to quantify the physical lifespan loss of the entire label;

[0293] In the decision output processing steps:

[0294] Obtain the preset maintenance threshold;

[0295] The maintenance threshold refers to a preset global entropy increase alarm threshold, such as 0.8. This threshold can be determined by correlating the historical damage data of the tag with its actual service life, indicating that the physical degradation of the tag has approached the upper limit of its service life.

[0296] The global physical entropy increase index is compared with a preset maintenance threshold;

[0297] If the global physical entropy increase index is greater than the preset maintenance threshold, for example... :

[0298] Then, a predictive maintenance label alert will be generated in the actionable business recommendations; for example, along with the output of the automatic approval recommendation, an additional recommendation will be added: Predictive maintenance label alert: Replace the label on this cycle of packaging;

[0299] If the global physical entropy increase exponent is less than or equal to the preset maintenance threshold:

[0300] Then, predictive maintenance label alerts will not be generated;

[0301] The synergy between modules generates a predictive maintenance value of gain; a single physics analysis module can only output how much damage was caused; a single decision module can only determine whether it is credible; this solution integrates the capabilities of module one... The output, combined with the decision logic of Module 4, enables the system not only to identify current data but also to predict future failures; it can detect failures caused by excessive tag wear. Before it becomes completely ineffective, proactively issue an alert to replace the label; this changes the operating model from reactive failure identification and manual data entry to proactive maintenance, opening up new asset management and cost-saving value for enterprises.

[0302] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for identifying and processing reusable packaging logistics documents based on deep learning, characterized in that, include: Acquire the original images of the documents captured on-site; Obtain the business context, including obtaining the predefined key information area mask and obtaining the predefined business weight factor; Based on the predefined key information region mask and the predefined business weight factor, the original image of the document is subjected to physical state analysis processing to obtain the physical state mask and the physical entropy increase index of the key information region. Based on the original image of the document and the physical state mask, multi-channel information extraction processing is performed to obtain the original data fragment, optical character recognition confidence, and model uncertainty. By combining the physical entropy increase index of the key information area, the confidence level of optical character recognition, and the uncertainty of the model, a nonlinear fusion process is performed to calculate the credibility score of the coupled data. Based on the credibility score of the coupled data, the original data fragment, and the business context, multimodal information conflict adjudication processing is performed to obtain credible data; Based on the credibility scores of the credible data and the coupled data, decision output processing is performed to generate structured data and actionable business recommendations; The steps for performing multi-channel information extraction and processing include: Run the printed optical character recognition model on the printed area layer of the physical state mask; Run the handwritten optical character recognition model on the handwritten area layer of the physical state mask; Run the stamp classification model on the stamp area layer of the physical state mask; The outputs of the printed optical character recognition model, the handwritten optical character recognition model, and the seal classification model together constitute the original data segment. The Monte Carlo Dropout technique is used to perform multiple inferences on each model run in the aforementioned steps, resulting in multiple inference results; The confidence level of optical character recognition is determined based on the average softmax probability of the multiple inference results. The uncertainty of the model is determined based on the variance of the multiple inference results or based on the information entropy of the multiple inference results; The steps for multimodal information conflict resolution include: From the original data fragment, obtain the printed information and the handwritten information; Obtain the credibility score of the coupled data corresponding to the printed information; Obtain the credibility score of the coupled data corresponding to the handwritten information; Based on the business context, it is determined that the printed information and the handwritten information are conflicting or supplementary information; If it is determined to be supplementary information, the printed information and the handwritten information are merged to generate a merged fact, and the merged fact is used as the reliable data; If it is determined to be supplementary information, the lower of the coupled data confidence score corresponding to the printed information and the coupled data confidence score corresponding to the handwritten information is used as the confidence level of the reliable data. If the information is determined to be conflicting, the reliability score of the coupled data corresponding to the printed information is compared with the reliability score of the coupled data corresponding to the handwritten information. If the information is determined to be conflicting, the information with a higher coupling data confidence score will be used as the trusted data, and its corresponding coupling data confidence score will be used as the confidence level of the trusted data.

2. The method for identifying and processing reusable packaging logistics documents based on deep learning as described in claim 1, characterized in that, The step of performing physical state analysis processing on the original image of the document includes: The original image of the document is processed using an image segmentation model, which divides it into multiple semantic layers to obtain the physical state mask; The physical state mask includes a printed area layer, a handwritten area layer, a stamp area layer, and a damage layer. Using the predefined key information region mask, the percentage of overlap between the damage layer and the key information region is calculated; using the predefined business weight factor, the percentage of overlap is weighted and summed to calculate the physical entropy increase index of the key information region.

3. The method for identifying and processing reusable packaging logistics documents based on deep learning as described in claim 1, characterized in that, The steps for performing nonlinear fusion processing include: Obtain the preset balance weights; Based on the preset balance weights, the confidence level of optical character recognition and the uncertainty of the model are fused to calculate the cognitive credibility. Obtain the preset degradation sensitivity coefficient; Based on the physical entropy increase index of the key information region and the preset degradation sensitivity coefficient, the physical penalty factor is calculated. The cognitive credibility score is obtained by multiplying the physical penalty factor by the cognitive credibility score.

4. The method for identifying and processing reusable packaging logistics documents based on deep learning as described in claim 3, characterized in that, The step of calculating the physical penalty factor includes: Subtract the physical entropy increase index of the key information region from 1 to obtain the difference; The difference is compared with 0, and the maximum value is taken to obtain the base of the exponentiation operation; The physical penalty factor is obtained by performing an exponential operation using the base and the preset degradation sensitivity coefficient.

5. The method for identifying and processing reusable packaging logistics documents based on deep learning as described in claim 1, characterized in that, The steps for processing the decision output include: Obtain the preset automatic approval threshold; Obtain the preset threshold for manual review; Get a list of predefined key fields; The trusted data and its corresponding coupled data credibility scores are output as the structured data; The fields in the trusted data are compared with the list of key fields to determine whether there are any key fields whose credibility scores for the coupled data are less than or equal to the manual review threshold. The credibility score of the coupled data for all fields in the trusted data is compared with the automatic approval threshold to determine whether there is a coupled data credibility score that is less than or equal to the automatic approval threshold. If the credibility score of the coupled data with key fields is determined to be less than or equal to the manual review threshold, a rejection process suggestion is generated as the actionable business suggestion. If it is determined that the credibility score of the coupled data without key fields is less than or equal to the manual review threshold, and it is determined that the credibility score of the coupled data is less than or equal to the automatic approval threshold, then a marked review suggestion is generated as the actionable business suggestion. If it is determined that the credibility score of the coupled data for which no key field exists is less than or equal to the manual review threshold, and it is determined that the credibility score of the coupled data for which no field exists is less than or equal to the automatic approval threshold, then an automatic approval suggestion is generated as the actionable business suggestion.

6. The method for identifying and processing reusable packaging logistics documents based on deep learning as described in claim 1, characterized in that, The physical state analysis and processing steps also include: Calculate the global physical entropy increase index of the original image of the document; The steps for processing the decision output also include: Obtain the preset maintenance threshold; The global physical entropy increase index is compared with the preset maintenance threshold; If the global physical entropy increase index is greater than the preset maintenance threshold, a predictive maintenance tag warning will be generated in the actionable business suggestions. If the global physical entropy increase index is less than or equal to the preset maintenance threshold, then a predictive maintenance tag warning will not be generated in the actionable business recommendations.

Citation Information

Patent Citations

  • Logistics express method and logistics express bill

    CN109978119A

  • Intelligent waybill management system

    CN119884196A