Information processing method for port container number tallying
By deploying visual networks and heterogeneous OCR model clusters, combined with image fusion and knowledge graphs, the high error recognition rate problem of the port container number recognition system in harsh environments was solved, and efficient and reliable container number recognition and automated processing were achieved.
Patent Information
- Application Number
- CN202511296663.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-09-11
AI Technical Summary
The existing port container number recognition system has a high error rate in harsh environments and lacks image quality assessment and intervention mechanisms, resulting in an imbalance between resources and efficiency. It also lacks automatic verification and semantic reasoning capabilities, making it difficult to achieve efficient and fully automated operations.
Deploy visual networks for image quality assessment, use embedded lightweight convolutional neural network models to score image quality, combine multi-frame image fusion and deep neural network recognition, use heterogeneous OCR model clusters for parallel recognition and dynamic weighted fusion, and perform error correction and completion through knowledge graphs.
It improves the accuracy and reliability of data collection, reduces data loss, reduces resource waste, achieves high-precision recognition in complex environments, and improves the level of automation and recognition success rate.
Smart Images

Figure CN120782401A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing, and in particular to an information processing method for port container number tallying. Background Art
[0002] Information processing in port container tallying refers to the use of various technical means to distinguish and manage container numbers and related data. Existing technologies generally use cameras for data collection and image processing technology to identify the numbers on the containers. At the same time, data verification and updates are carried out in conjunction with the supporting operating system and other business systems at the terminal or port to ensure the accuracy and real-time nature of logistics information.
[0003] However, existing port container number recognition systems mostly rely on a single camera capturing images and then directly feeding them into OCR recognition, lacking a pre-emptive image quality assessment or intervention mechanism. When encountering rain, fog, oil stains, reflections, or obstructions caused by stacked containers, the captured images are often blurry, overexposed, or partially missing, yet the system still forcibly recognizes them, resulting in a high error rate. For example, if AABC is mistakenly identified as AABO, or 5 is identified as S, the currently designed recognition system cannot determine whether the result is reliable and can only output it as is. More seriously, these low-quality recognition results, regardless of the error type, are uniformly reviewed by humans or supporting systems, with no priority distinction or auxiliary judgment methods, resulting in an ineffective balance between resources and efficiency. Furthermore, traditional solutions lack the ability to automatically verify container number rules, let alone leverage historical data and knowledge graphs for semantic reasoning. Once characters are blurred or missing, they are helpless and difficult to effectively complete. Taken together, these problems make the existing design system more rigid and inefficient, making it difficult to support the high-efficiency and fully automated operations required by modern ports. Summary of the Invention
[0004] To achieve the above objectives, the present invention is implemented through the following technical solutions:
[0005] An information processing method for port container number tallying, the method comprising:
[0006] A visual network is deployed at key nodes within the target area to acquire a dataset related to container numbers. An embedded lightweight convolutional neural network model is used to assess the quality of each original image frame, outputting a comprehensive quality assessment value. This comprehensive quality assessment value is then compared with an assessment threshold, and the system determines whether to trigger a retake adjustment based on the comparison result. Multiple images of the same target container are captured at different times and angles, and a feature-matching-based image registration and fusion algorithm is used to output a reconstructed dataset related to container numbers. A two-level deep neural network is deployed to complete container number recognition.
[0007] After completing the box number recognition process, a pre-built heterogeneous OCR model cluster performs parallel recognition on each standard character, obtaining three sets of recognition results. Based on the recognition results, a revision and dynamic weighted fusion strategy is implemented, and the revised first, second, and third confidence levels are compared with the standard threshold range.
[0008] When the corrected first confidence level exceeds the upper limit of the standard threshold range, the recognition result is adopted;
[0009] When the corrected first confidence level is within the standard threshold interval, the recognition results of the heterogeneous OCR model cluster are weighted voted and output; when the corrected first confidence level is lower than the lower limit of the standard threshold interval, a secondary judgment mechanism is executed, and the corrected second confidence level and the corrected third confidence level are compared with the lower limit of the standard threshold interval. If both are lower than the lower limit of the standard threshold interval, they are marked as abnormal characters; otherwise, they are marked as ambiguous characters.
[0010] Furthermore, the container number-related data set includes at least: target container images taken from several angles; each frame of the original image is the corresponding target container image, and the process of quality scoring each frame of the original image is: weighted summation based on the scoring dimensions to obtain a comprehensive quality assessment value; wherein the scoring dimensions include at least: illumination uniformity, motion blur, key area coverage, and character clarity.
[0011] Furthermore, the process of running the feature matching-based image registration and fusion algorithm is as follows:
[0012] SIFT is used to extract stable feature points from each image, and the RANSAC algorithm is used to eliminate false matches and achieve sub-pixel alignment. ESRGAN-Port, a super-resolution model based on a generative adversarial network, is used to fuse and reconstruct multiple low-resolution images (LR) after sub-pixel alignment into a single high-resolution image (HR). The ESRGAN-Port model is trained on a dedicated dataset containing several real container images. During the reconstruction process, the ESRGAN-Port model prioritizes enhancing high-frequency details in the container number area on the target container through an attention mechanism.
[0013] Furthermore, the two-stage deep neural network includes: a box area detection model YOLOv7-Cargo and a character segmentation network SegNet-CR; a single high-resolution image HR is input to the box area detection model YOLOv7-Cargo, and the output is: the minimum enclosing rectangle of the box number area on the target container box; the character segmentation network SegNet-CR is used to segment the detected box number and obtain several standard characters.
[0014] Furthermore, the heterogeneous OCR model cluster includes at least: a main heterogeneous model, a first sub-heterogeneous model and a second sub-heterogeneous model; wherein, the main heterogeneous model adopts the deep convolutional neural network CRNN-CargoNet; the first sub-heterogeneous model adopts the Transformer-based visual recognition model ViT-CR; the second sub-heterogeneous model adopts the lightweight model MobileOCR-Edge; the confidence corresponding to the main heterogeneous model is marked as the first confidence, the confidence corresponding to the first sub-heterogeneous model is marked as the second confidence, and the confidence corresponding to the second sub-heterogeneous model is marked as the third confidence.
[0015] Furthermore, the correction and dynamic weighted fusion strategy process also includes: taking the comprehensive quality evaluation value as a basis, calculating the average of the comprehensive quality evaluation value corresponding to each frame within a preset period to obtain the comprehensive quality average Q, constructing a linear attenuation model based on the Q value, inputting the original confidence level and the comprehensive quality average Q, and outputting: a first-level corrected confidence level;
[0016] The constraints for the first-level corrected confidence level are set as follows:
[0017] Condition 1: The confidence level after the first level correction exceeds the set threshold;
[0018] Condition 2: The difference between the confidence levels after the first level correction does not exceed the error range. When the confidence level after the first level correction meets the constraint conditions, indicating that the effect is up to standard, the confidence level after the first level correction shall prevail. Otherwise, the second level correction mechanism shall be triggered.
[0019] Query the historical accuracy database to obtain the corresponding historical average accuracy; build a piecewise affine transformation calibration model, input the first-level corrected confidence and the corresponding historical average accuracy, and output: the second-level corrected confidence.
[0020] Furthermore, the function on which the piecewise affine transformation calibration model is based is: P2=f(P1, A_hist); where P2 represents the confidence after the second-level correction, P1 represents the confidence after the first-level correction, and A_hist represents the corresponding historical average accuracy; the specific expansion formula of the function on which the piecewise affine transformation calibration model is based is: when P1 is not lower than A_hist, then P2=A_hist+u*(P1-A_hist); where u represents the scaling factor, and u∈(0, 1); when P1 is lower than A_hist, then P2=P1.
[0021] Furthermore, the method also includes: when abnormal characters are detected, a dynamic error correction and completion mechanism based on context and business rules is triggered; when ambiguous characters are detected, the establishment of a port-level container number knowledge graph and association reasoning mechanism is triggered.
[0022] Further, the process of the dynamic error correction and completion mechanism based on context and business rules is: calling the container main code database, verifying the validity of the combination of the first several letters; cross-verification after obtaining the pre-provided information; applying the container number check code algorithm for mathematical verification, converting the first n characters into numerical values, calculating the modulus n+1 remainder after weighted summation, and detecting whether it matches the n+1th character; if not, calibrate the error bit by traversing all character combinations until the calibration passes; combined with the space-time context, the container numbers of adjacent target containers in the same target area are continuous, and if the front and rear container numbers have been confirmed, the intermediate container number is inferred.
[0023] Further, the process of establishing a port-level container number knowledge graph and associated reasoning mechanism is: constructing a container number knowledge graph under the target area, and the container number knowledge graph takes the container number as the entity node, and the associated attributes at least include: the company to which it belongs, the ship, the voyage, the origin port, the destination port, the cargo type, the historical port record and the RFID tag ID; calling the container number knowledge graph for associated reasoning, and outputting the associated reasoning result.
[0024] The present application provides an information processing method for port container number tallying, which has the following beneficial effects:
[0025] 1) The present application can obtain container numbers and their related information in complex environments by deploying a visual network, ensuring the quality and integrity of the data set; at the same time, an embedded lightweight convolutional neural network model is used to evaluate the quality of the image, and a retake adjustment instruction is triggered when necessary, further ensuring that the images entering the subsequent processing process meet the basic requirements for recognition. This scheme not only improves the accuracy and reliability of data acquisition, but also reduces the problem of data loss caused by harsh environments or technical limitations.
[0026] 2) The present application uses the comparison of the comprehensive quality evaluation value and the evaluation threshold to determine whether to trigger the retake adjustment instruction, and further obtains the comprehensive quality mean value on this basis, and constructs a linear decay model to make a first correction to the original confidence level. On the one hand, it ensures that only when the image quality is poor will additional measures be taken to ensure the quality and integrity of data acquisition, avoiding unnecessary waste of resources; on the other hand, through conservative calibration of the original confidence level based on image quality, the true reliability of the recognition result can be more accurately reflected, enhancing the robustness of the decision-making process, effectively solving the problem of high misjudgment rate in traditional schemes due to neglect of image quality, and achieving the goal of maintaining high-precision recognition in complex working conditions;
[0027] 3) This solution introduces a first-level correction based on Q-values and a second-level correction supported by historical accuracy. The combination of the two results in a revision and dynamic weighted fusion strategy, as well as error correction and information reasoning recognition methods based on decision-making results. This not only enables rapid response and calibration of initial recognition results, but also allows for deeper mining of historical data for more accurate secondary corrections when necessary. Through the application of associative reasoning and knowledge graphs, it achieves a leap from character recognition to semantic understanding, significantly improving the final recognition success rate for case numbers corresponding to abnormal or ambiguous characters.
[0028] 4) The differentiation mechanism given above effectively solves the problem of unbalanced efficiency and resources caused by the unified processing of all low-confidence results in traditional tallying schemes, realizes the layered processing effect, and embodies the optimal solution based on multi-source evidence fusion and setting of layers, making the overall solution both rigorous and flexible, greatly improving the overall recognition success rate and automation level in complex degradation scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 The figure is a schematic diagram of the overall flow of an information processing method for port container number tallying in the present invention. DETAILED DESCRIPTION
[0030] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0031] See also Figure 1 This embodiment provides an information processing method for port container number tallying. The method is for processing container numbers and related data information in the port area. The specific steps of the processing method are as follows:
[0032] S1. Collection and preprocessing of container number information based on multimodal perception:
[0033] S1.1. Deployment Information Collection:
[0034] Deploy a visual network at key nodes in the target area to obtain container number-related data sets;
[0035] The target area is the selected port area. For the key nodes in the target area, it includes: port gantry crane, shore crane, container transport channel and container stacking area, etc. The deployed visual network includes: array of several high-definition cameras, laser radar, ground coil and RFID, etc. The purpose is to ensure the overall and complete information or data. The container number related data set includes container images taken from several angles to ensure that the appearance of the target container is captured without dead angles. The specific content of the container image can be selected as: the identification area of the front, side, bottom and top of the complete target container.
[0036] Specifically, for each high-definition camera in the visual network, each high-definition camera should be equipped with automatic focusing, zoom and wide-angle lens to ensure that container images can be obtained under different lifting heights and different container stacking numbers. Each high-definition camera has a design of anti-salt mist and anti-vibration to adapt to the harsh environment of high humidity and strong wind in the port area, thereby ensuring the quality and integrity of the container number related data set from the physical layer.
[0037] The visual network will trigger a linkage mechanism when it works:
[0038] When the target container enters the target area, the ground coil activates the preliminary positioning, the RFID reads the equipment number carried by the target container and associates with the pre-provided information through networking. The pre-provided information includes: bill of lading number, ship name and voyage, loading and unloading plan, etc. At the same time, the laser radar scans the outline of the target container and triggers the corresponding high-definition camera under the key node to complete the continuous shooting of the target container. Among them, no less than 3 different angle images are collected at a time to ensure that the container number area on the target container is more complete and clear in at least 2 images. The use of this linkage mechanism solves the problem of easy shielding and limited angle of view of traditional single-point fixed camera to a certain extent. Through the joint operation of multiple devices, the complete acquisition of container number related data set under complex working conditions can be significantly improved.
[0039] S1.2, image information evaluation:
[0040] Based on the obtained container number related data set, an embedded lightweight convolutional neural network model is used to score the quality of each frame of original image. The scoring dimensions for quality scoring include: uniformity of illumination, motion blur, key area coverage and character clarity. The weighted sum of each scoring dimension is output to obtain a comprehensive quality evaluation value. The comprehensive quality evaluation value is compared with the evaluation threshold value, and the comparison result is used to select whether to trigger the retake adjustment instruction.
[0041] Among them, illumination uniformity is obtained by local standard deviation analysis to avoid overexposed / underexposed areas; motion blur is obtained based on gradient amplitude variance calculation; key area coverage is obtained by locating the target box number area through the pre-trained box ROI detector and calculating its pixel ratio; character clarity is obtained by calculating the sharpness index after edge extraction using the Sobel operator; the required comprehensive quality evaluation value can be obtained by weighted summation of illumination uniformity, motion blur, key area coverage and character clarity; the evaluation threshold is preset in advance according to actual needs; when the limit of the comprehensive quality evaluation value is within 1, the corresponding evaluation threshold is 1 / 7, that is, 0.7; Taking the evaluation threshold = 0.7 as an example, when the comprehensive quality evaluation value exceeds the evaluation threshold, no response action is taken; when the comprehensive quality evaluation value does not exceed the evaluation threshold, the reshoot adjustment instruction is triggered: it is automatically marked as a low-quality image, and the visual network parameters are synchronously adjusted before reshooting; among them, the visual network parameters correspond to the high-definition camera increasing exposure compensation and switching focal length, etc.; this step effectively filters out blurred and distorted images caused by rain, fog, oil pollution, reflection or rapid movement, ensuring that the images entering the subsequent processing flow meet the minimum requirements for recognition, reducing the consumption of invalid computing resources; compared with traditional fixed threshold filtering, this solution achieves more accurate quality judgment through model adaptive learning of image degradation patterns in the specific environment of the port.
[0042] Description: By deploying a visual network, this method can obtain container numbers and related information in complex environments, ensuring the quality and integrity of the dataset. Furthermore, an embedded lightweight convolutional neural network model is used to assess image quality and, when necessary, trigger retake and adjustment instructions, further ensuring that images entering subsequent processing meet basic recognition requirements. This solution not only improves the accuracy and reliability of data acquisition, but also reduces data loss caused by harsh environments or technical limitations.
[0043] S1.3 Multi-source image fusion and reconstruction:
[0044] Under the condition of collecting multiple frames of images of the same target container at different times and angles, an image registration and fusion algorithm based on feature matching is used to output the reconstructed container number related data set;
[0045] The process of running the image registration and fusion algorithm based on feature matching is as follows:
[0046] S1.3.1. Use SIFT to extract stable feature points in each image, and combine it with the RANSAC algorithm to eliminate false matches and achieve sub-pixel alignment; S1.3.2. Apply the super-resolution model ESRGAN-Port based on the generative adversarial network to fuse the multiple frames of low-resolution images LR after sub-pixel alignment and reconstruct them into a single high-resolution image HR, where the high resolution mentioned here is at least 4k; S1.3.3. The ESRGAN-Port model completes the training operation on the equipped dedicated dataset; the dedicated dataset contains several real container images; specifically, during the reconstruction process, the ESRGAN-Port model uses the attention mechanism to prioritize the enhancement of high-frequency details in the container number area on the target container and suppress background noise. The output single high-resolution image HR not only improves the clarity of the character edges, but also partially restores the stroke information that is slightly covered to ensure the input quality of subsequent data; the above steps combine multi-view redundant information with deep learning super-resolution technology to solve the problem of insufficient resolution of a single camera due to distance or occlusion.
[0047] S1.4. Container number area positioning and segmentation:
[0048] A two-level deep neural network is deployed on the reconstructed container number-related dataset to complete the container number recognition action;
[0049] Among them, the two-level deep neural network includes: the box area detection model YOLOv7-Cargo and the character segmentation network SegNet-CR; the 4k image is input to the box area detection model YOLOv7-Cargo, and the output is: the minimum enclosing rectangle of the box number area on the target container box; the box area detection model YOLOv7-Cargo is the existing set model, and its positioning accuracy can reach more than 99%; the character segmentation network SegNet-CR: The character segmentation network SegNet-CR performs pixel-level segmentation on the detected box number, and the segmentation obtains several standard characters, such as: 4-digit box master code + 6-digit Number + 1 check code; SegNet-CR adopts U-Net architecture, the encoder is ResNet-34, and the decoder combines with void convolution to expand the receptive field and output the binary mask of each character; in order to deal with situations such as character adhesion, breakage or oil coverage, conditional random field post-processing is introduced, and the fixed spacing and arrangement rules between characters are used to optimize the segmentation results by topological constraints, and the size of the single character image after segmentation is standardized; this step achieves robust positioning and accurate segmentation of the box number area under complex background through a dedicated network trained end-to-end, and reduces the error rate of the adhesion character processing by at least 60% compared with the traditional edge detection combined with projection method, which is a significant effect.
[0050] S2. Character recognition, confidence assessment and correction:
[0051] After completing the box number recognition process, a pre-built heterogeneous OCR model cluster performs parallel recognition of each standard character. The heterogeneous OCR model cluster includes a main heterogeneous model, a first-level heterogeneous model, and a second-level heterogeneous model. The main heterogeneous model is the deep convolutional neural network CRNN-CargoNet; the first-level heterogeneous model is the Transformer-based visual recognition model ViT-CR; and the second-level heterogeneous model is the lightweight model MobileOCR-Edge. Three sets of recognition results are obtained, and revision and dynamic weighted fusion strategies are implemented based on the recognition results.
[0052] Specifically, the deep convolutional neural network CRNN-CargoNet uses a CNN+BiLSTM+CTC architecture optimized for sequence recognition. The Transformer-based visual recognition model ViT-CR is used to divide character images into 16×16 patch sequence inputs and uses a self-attention mechanism to capture global context, making it more robust when handling slightly deformed and partially rotated characters. The partial rotation limit is ±15°. MobileOCR-Edge is compressed from CRNN-CargoNet using knowledge distillation technology, with 15% of the original model's parameters. It is deployed on edge computing nodes for real-time preliminary recognition.
[0053] The three sets of recognition results obtained are three sets of confidence levels;
[0054] The confidence level corresponding to the deep convolutional neural network CRNN-CargoNet is marked as the first confidence level, the confidence level corresponding to the Transformer-based visual recognition model ViT-CR is marked as the second confidence level, and the confidence level corresponding to the lightweight model MobileOCR-Edge is marked as the third confidence level;
[0055] The process of executing the correction and dynamic weighted fusion strategy is as follows:
[0056] S2.1. Level 1 correction: Using the comprehensive quality assessment value obtained in S1.2 as the basis, calculate the mean of the comprehensive quality assessment value corresponding to each frame within the preset period to obtain the comprehensive quality mean Q, construct a linear attenuation model based on the Q value, input the original confidence and the comprehensive quality mean Q, and output the confidence after level 1 correction; where the confidence after level 1 correction is: the product of the original confidence and the comprehensive quality mean Q; logical explanation: the credibility of the confidence of this model is proportional to the image quality, and can be executed on the edge node through conservative calibration of the product formula.
[0057] Illustration: In S1.2, it is decided whether to trigger the retake adjustment instruction by comparing the comprehensive quality evaluation value with the evaluation threshold, and on this basis, the comprehensive quality mean value is further obtained, and a linear decay model is constructed to make a first-level correction to the original confidence. On the one hand, it ensures that only when the image quality is poor will additional measures be taken to ensure the quality and integrity of data collection, avoiding unnecessary waste of resources. On the other hand, through conservative calibration of the original confidence based on image quality, the true reliability of the recognition result can be more accurately reflected, enhancing the robustness in the decision-making process.
[0058] The above method effectively solves the problem of high misjudgment rate caused by neglecting image quality in the traditional scheme, achieving the goal of maintaining high-precision recognition under complex working conditions. At the same time, the strategy adopted above embodies the design concept of combining intelligent perception and adaptive processing, not only improving the success rate of a single identification task, but also optimizing resource allocation to improve the efficiency and response speed of the overall scheme operation, ensuring that reliable tally decisions can be made to some extent even under low-quality image conditions, reflecting the automation level and accuracy of port container number tallying work.
[0059] S2.2, Effect Evaluation and Compliance Determination: The constraint condition of the first-level corrected confidence is: condition one, the first-level corrected confidence exceeds the set threshold; condition two, the difference between the first-level corrected confidence does not exceed the error range; wherein the first-level corrected confidence includes three confidences; when the first-level corrected confidence meets the constraint condition, it means that the effect meets the standard, and the first-level corrected confidence is used as the reference; when the first-level corrected confidence does not meet the constraint condition, the second-level correction mechanism is triggered; wherein the first-level corrected confidence is denoted as P1;
[0060] S2.3, Second-level Correction Mechanism: Query the historical accuracy rate database, which establishes a multi-dimensional lookup table according to the comprehensive quality evaluation value bin, model type, and character position (first / middle / last);
[0061] For example, the query entry: (Model = CRNN-CargoNet, Q_bin = [0.80, 0.85), Position = Middle) corresponds to a historical average accuracy rate of A_hist = 89.5%, which is only an example for reference;
[0062] Construct a piecewise affine transformation calibration model, input the first-level corrected confidence and the corresponding historical average accuracy, and output the second-level corrected confidence; wherein, the function based on the piecewise affine transformation calibration model is: P2=f(P1, A_hist); wherein P2 represents the second-level corrected confidence, and A_hist represents the corresponding historical average accuracy; the specific expansion formula of the function based on the piecewise affine transformation calibration model is: when P1 is not lower than A_hist, then P2=A_hist+u*(P1-A_hist); wherein u represents the scaling factor, and u∈(0,1), and in this embodiment, the value of u is usually 0.5 to prevent overestimation; when P1 is lower than A_hist, then P2=P1, conservatively maintained; the second-level corrected confidence shall prevail;
[0063] Regardless of whether the first-level revised confidence level or the second-level revised confidence level is used as the basis, the confidence levels ultimately obtained are the required revised first confidence level, revised second confidence level, and revised third confidence level.
[0064] The solution introduces a first-level correction based on Q value and a second-level correction supported by historical accuracy. The combination of the two results in a revision and dynamic weighted fusion strategy, as well as error correction and information reasoning and recognition methods based on decision results. It can not only quickly respond to and calibrate preliminary recognition results, but also deeply mine historical data when necessary to achieve more accurate secondary corrections. Through the application of associative reasoning and knowledge graphs, it realizes the leap from character recognition to semantic understanding, greatly improving the final recognition success rate of box numbers corresponding to abnormal or ambiguous characters. From the perspective of the overall solution, it constitutes an efficient and intelligent information processing closed loop, which effectively solves the problems of misjudgment and missed judgment that may occur in traditional information recognition solutions.
[0065] S2.4. Execute weighted fusion strategy:
[0066] The corrected first confidence level is compared with the standard threshold interval. If the corrected first confidence level exceeds the upper limit of the standard threshold interval, the recognition result is adopted. If the corrected first confidence level is within the standard threshold interval, the recognition results of the main heterogeneous model, the first sub-heterogeneous model, and the second sub-heterogeneous model are weighted voted and output. If the corrected first confidence level is lower than the lower limit of the standard threshold interval, the secondary judgment mechanism is executed, and the corrected second confidence level and the corrected third confidence level are compared with the lower limit of the standard threshold interval. If both are lower than the lower limit of the standard threshold interval, they are marked as abnormal characters and S3.1 is executed. Otherwise, they are marked as ambiguous characters and S3.2 is executed.
[0067] Note: When using the correction and dynamic weighted fusion strategy implemented above, Q value is preferred for rapid calibration, and historical data is called only when necessary, which enables rapid processing. In some complex scenarios, compared with a single correction method, overconfidence in low-quality images is avoided, thereby improving the accuracy of the final decision. Most recognition tasks meet the standards after the first level of correction, without the need to query the historical database, reducing system latency while achieving resource optimization. The constructed two-layer confidence calibration system enables the overall information processing solution to have both real-time response capabilities and deep learning capabilities, which is a key innovation in the field of intelligent port tallying.
[0068] S3. Error correction and completion and information reasoning and recognition:
[0069] S3.1. Dynamic error correction and completion mechanism based on context and business rules:
[0070] S3.1.1. Call the container master code database to verify the validity of the combination of the first few letters; S3.1.2. Obtain the pre-provided information from S1.1 and perform cross-validation based on it; S3.1.3. Apply the container number check code algorithm for mathematical verification, convert the first n characters into numerical values according to the rules, calculate the modulo n+1 remainder after weighted summation, and check whether it matches the n+1th bit; if it does not match, traverse all possible character combinations to calibrate the error bits until the calibration passes; in this embodiment, the value of n is 10; S3.1.4. Combined with the spatiotemporal context, the container numbers of adjacent target containers in the same area are continuous. If the previous and next container numbers have been confirmed, the middle container number is inferred.
[0071] Among them, the container master code database is a known corresponding database, for example: the first four-digit code set in most companies is mapped to its corresponding company name, which is used to verify the validity of the first four-digit combination; if the first three digits are determined and the fourth digit is ambiguous, the fourth digit is tried first rather than other letters; in S3.1.2, if a container number is known in advance, and the recognition result is the same as the predicted result, only the sixth digit is ambiguous, the sixth digit can be directly completed according to the predicted result; in S3.1.3, for example, if the recognition is: XXXX123456, and the verification fails, try to replace the 6 in the 10th digit with any number until the verification passes; the above-mentioned dynamic error correction and completion method based on context and business rules integrates the industry knowledge base, business data flow and mathematical rules, and increases the final recognition success rate of the container number corresponding to the abnormal characters from the original, for example: 70% by at least 20%.
[0072] S3.2. Establish a knowledge graph and association reasoning mechanism for port-level container numbers:
[0073] S3.2.1. Construct a container number knowledge graph for the target region. This knowledge graph uses container numbers as entity nodes and includes attributes such as company, vessel, voyage number, port of departure, port of destination, cargo type, historical entry and exit records, maintenance records, and RFID tag ID. S3.2.2. Call the knowledge graph to perform association reasoning and output the association reasoning results.
[0074] For example, if the first six digits of a container number are identified as: XXXX12, a query on the knowledge graph reveals that on the ocean vessels that have recently docked, the containers starting with XXXX12 are all from Port A, and are mostly loaded with electronic products. If the appearance characteristics of the current container are similar to those of such cargo boxes, XXXX12 will be used as the final result; conversely, if the box appears next to a bulk carrier loaded with ore, an abnormal warning will be triggered and S3.1 will continue to be executed; the knowledge graph also records historical recognition error patterns to achieve continuous learning and updating, and realize the leap from character recognition to semantic understanding, thereby achieving more complete and comprehensive information processing.
[0075] In this scheme, the corrected second confidence level and the corrected third confidence level are compared with the lower limit of the standard threshold interval. If both are lower than the lower limit of the standard threshold interval, they are marked as abnormal characters and S3.1 is executed; otherwise, they are marked as ambiguous characters and S3.2 is executed. This design reflects the refined classification and differentiated response strategies for the reasons for recognition failure. On the one hand, when the confidence levels of the heterogeneous OCR model clusters are all low, it indicates that the current image is severely degraded, which can be attributed to systematic recognition failure. Therefore, it is prioritized to rely on external strong constraint information for deterministic repair. The industry rules and mathematical verification provided by S3.1 can directly complete or correct the box number with a higher probability, so as to achieve the effect of quickly recovering key information. On the other hand, When only the main heterogeneous model has low confidence due to its sensitivity to quality, but other models still have a certain confidence, it means that although the characters are fuzzy, there are some recognizable features, which is a local uncertainty problem. At this time, the knowledge graph reasoning of S3.2 can combine the semantic context for intelligent inference, thereby achieving flexible decision-making in the absence of clear rules; the above-mentioned distinction mechanism effectively solves the problem of unbalanced efficiency and resources caused by the unified processing of all low-confidence results in the traditional tallying scheme, realizes the hierarchical processing effect, and embodies the optimal solution based on multi-source evidence fusion and setting stratification, so that the overall solution has both rigor and flexibility, greatly improving the overall recognition success rate and automation level in complex degradation scenarios.
[0076] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.
[0077] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.
[0078] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. An information processing method for port container number tallying, characterized by: The method includes: A visual network is deployed at key nodes within the target area to acquire a dataset related to container numbers. An embedded lightweight convolutional neural network model is used to assess the quality of each original image frame and output a comprehensive quality assessment value. This comprehensive quality assessment value is then compared with an assessment threshold, and the decision to trigger a retake adjustment is made based on the comparison result. Multiple images of the same target container are captured at different times and angles, and a feature-matching-based image registration and fusion algorithm is used to output a reconstructed dataset related to container numbers. A two-level deep neural network is deployed to complete container number recognition. After completing the box number recognition action and obtaining the standard characters, a pre-built heterogeneous OCR model cluster performs parallel recognition on each standard character to obtain three sets of recognition results. Based on the recognition results, a revision and dynamic weighted fusion strategy is implemented, and the revised first confidence level, second confidence level, and third confidence level are compared with the standard threshold range. When the corrected first confidence level exceeds the upper limit of the standard threshold range, the recognition result is adopted; When the corrected first confidence level is within the standard threshold interval, the recognition results of the heterogeneous OCR model cluster are weighted voted and output; when the corrected first confidence level is lower than the lower limit of the standard threshold interval, a secondary judgment mechanism is executed, and the corrected second confidence level and the corrected third confidence level are compared with the lower limit of the standard threshold interval. If both are lower than the lower limit of the standard threshold interval, they are marked as abnormal characters; otherwise, they are marked as ambiguous characters.
2. The information processing method for port container number tallying according to claim 1, characterized in that: The container number-related dataset includes at least: target container images taken from several angles; each frame of the original image is the corresponding target container image, and the process of scoring the quality of each frame of the original image is: weighted summation based on the scoring dimensions to obtain a comprehensive quality assessment value; among which, the scoring dimensions include at least: illumination uniformity, motion blur, key area coverage, and character clarity.
3. The information processing method for port container number tallying according to claim 1, characterized in that: The process of running the feature matching-based image registration and fusion algorithm is as follows: SIFT is used to extract stable feature points from each image, and the RANSAC algorithm is used to eliminate mismatches and achieve sub-pixel alignment. ESRGAN-Port, a super-resolution model based on a generative adversarial network, is used to fuse and reconstruct multiple low-resolution images (LR) after sub-pixel alignment into a single high-resolution image (HR). The ESRGAN-Port model is trained on a dedicated dataset containing several real container images. During reconstruction, the ESRGAN-Port model prioritizes enhancing high-frequency details in the container number area on the target container through an attention mechanism.
4. The information processing method for port container number tallying according to claim 3, characterized in that: The two-stage deep neural network includes: the box area detection model YOLOv7-Cargo and the character segmentation network SegNet-CR; a single high-resolution image HR is input to the box area detection model YOLOv7-Cargo, and the output is: the minimum enclosing rectangle of the box number area on the target container box; the character segmentation network SegNet-CR is used to segment the detected box number and obtain several standard characters.
5. The information processing method for port container number tallying according to claim 1, characterized in that: The heterogeneous OCR model cluster includes at least: a main heterogeneous model, a first sub-heterogeneous model and a second sub-heterogeneous model; among them, the main heterogeneous model adopts the deep convolutional neural network CRNN-CargoNet; the first sub-heterogeneous model adopts the Transformer-based visual recognition model ViT-CR; the second sub-heterogeneous model adopts the lightweight model MobileOCR-Edge; the confidence level corresponding to the main heterogeneous model is marked as the first confidence level, the confidence level corresponding to the first sub-heterogeneous model is marked as the second confidence level, and the confidence level corresponding to the second sub-heterogeneous model is marked as the third confidence level.
6. The information processing method for port container number tallying according to claim 1, characterized in that: The correction and dynamic weighted fusion strategy process also includes: taking the comprehensive quality evaluation value as a basis, calculating the average of the comprehensive quality evaluation value corresponding to each frame within a preset period to obtain the comprehensive quality average Q, constructing a linear attenuation model based on the Q value, inputting the original confidence level and the comprehensive quality average Q, and outputting: the first-level corrected confidence level; The constraints for the first-level corrected confidence level are set as follows: Condition 1: The confidence level after the first level correction exceeds the set threshold; Condition 2: The difference between the confidence levels after the first level correction does not exceed the error range. If the confidence level after the first level correction meets the constraint conditions, indicating that the effect is up to standard, the confidence level after the first level correction will prevail. Otherwise, the second level correction mechanism will be triggered. Query the historical accuracy database to obtain the corresponding historical average accuracy; build a piecewise affine transformation calibration model, input the first-level corrected confidence and the corresponding historical average accuracy, and output: the second-level corrected confidence.
7. The information processing method for port container number tallying according to claim 6, characterized in that: The function based on the piecewise affine transformation calibration model is: P2=f(P1, A_hist); where P2 represents the confidence after the second-level correction, P1 represents the confidence after the first-level correction, and A_hist represents the corresponding historical average accuracy; the specific expansion formula of the function based on the piecewise affine transformation calibration model is: when P1 is not lower than A_hist, then P2=A_hist+u*(P1-A_hist); where u represents the scaling factor, and u∈(0,1); when P1 is lower than A_hist, then P2=P1.
8. The information processing method for port container number tallying according to claim 1, characterized in that: The method also includes: when abnormal characters are detected, a dynamic error correction and completion mechanism based on context and business rules is triggered; when ambiguous characters are detected, the establishment of a port-level container number knowledge graph and an association reasoning mechanism is triggered.
9. The information processing method for port container number tallying according to claim 8, characterized in that: The process of the dynamic error correction and completion mechanism based on context and business rules is as follows: call the container master code database to verify the validity of the combination of the first few letters; perform cross-validation after obtaining pre-provided information; apply the container number check code algorithm for mathematical verification, convert the first n characters into a numerical value, calculate the modulo n+1 remainder after weighted summation, and check whether it matches the n+1th digit; if there is no match, traverse all character combinations to calibrate the error bit until the calibration is successful; combined with the spatiotemporal context, the container numbers of adjacent target containers in the same target area are continuous. If the previous and next container numbers have been confirmed, the intermediate container number is inferred.
10. The information processing method for port container number tallying according to claim 8, characterized in that: The process of establishing a port-level container number knowledge graph and association reasoning mechanism is as follows: constructing a container number knowledge graph under the target area, and the container number knowledge graph uses the container number as the entity node, and the associated attributes include at least: the company, ship, voyage, departure port, destination port, cargo type, historical entry and exit records, and RFID tag ID; calling the container number knowledge graph for association reasoning and outputting the association reasoning results.
Citation Information
Patent Citations
Foggy container number identification method based on lightweight deep convolutional neural network
CN116665216A
Multi-frame snapshot recognition method for port automation equipment
CN117475371A
Knotarization intelligent question and answer customer service method and system based on knowledge graph
CN119938816A
Container number tallying identification method
CN120411989A
Multi-scale lightweight brain tumor segmentation method based on improved YOLOv8n-Seg
CN120526155A