OCR (optical character recognition) method for quay crane and quay crane tallying system
Through multi-model fusion technology, the quay crane tallying system uses PP-OCR, Paddle OCR and CnOCR models to recognize images, aggregate prediction boxes and assign recognition weights, solving the problem of low OCR recognition accuracy and achieving efficient port operations.
Patent Information
- Application Number
- CN202510817085.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
AI Technical Summary
The OCR recognition accuracy in the quay crane tallying system is low and cannot meet the requirements of fully automated operation, mainly due to complex on-site working conditions, poor image quality, incomplete and damaged container numbers, and different position structures.
Using multi-model fusion technology, the image is recognized through PP-OCR, Paddle OCR and CnOCR models respectively, the prediction box is aggregated and the confidence is calculated, the recognition weight is assigned based on the model's proficiency index in the specified scenario, and finally the character recognition result is calculated.
The OCR recognition rate has been greatly improved, and the accuracy rate of container number recognition in port operations has reached 99%, ensuring the efficient operation of the intelligent tallying system and avoiding manual processing.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of automated ports, and in particular relates to a quay crane OCR recognition method and a quay crane tallying system. Background Art
[0002] With the acceleration of global integration and the increasing frequency of international economic activity, modern large-scale hub ports are becoming increasingly important as propellers and connectors of modern logistics services based on transportation, warehousing, and distribution. The quay crane container intelligent identification system, based on PLC data integration and video data acquisition, utilizes computer image processing, deep learning, and other video AI recognition technologies, combined with real-time data processing, to achieve the digital transformation of tally data collection services.
[0003] The quay crane intelligent tallying system needs to identify the dimensions of the containers in operation (including single 20-foot, double 20-foot, quad 20-foot, 40-foot, 45-foot, 48-foot, 53-foot, double long containers, etc.) and the various types of containers (including conventional containers, ultra-high containers, out-of-gauge containers, refrigerated containers, flat rack containers, tank containers, open top containers, etc.), including container number, ISO code, door direction, container damage, presence of lead seals, dangerous goods identification, special container types, loading and unloading container locations, and other information.
[0004] The optical character recognition (OCR) accuracy of the quay crane's intelligent tallying system is low due to the following reasons: 1. Complex on-site conditions lead to shadows, blur, and incomplete container numbers in captured images; 2. Corrosive coastal conditions cause container numbers to be damaged; and 3. The location and structure of container numbers vary. This low accuracy rate fails to meet the requirements of fully automated operation, limiting on-site operational efficiency. Summary of the Invention
[0005] In order to solve the problem of low OCR recognition accuracy in the quay crane tallying system, the present invention proposes a quay crane OCR recognition method and a quay crane tallying system based on multi-model fusion, which greatly improves the quay crane OCR recognition rate.
[0006] The present invention is achieved by adopting the following technical solutions:
[0007] A quay crane OCR recognition method is proposed, including:
[0008] S1: Use pre-trained multiple models to identify the image to be identified and obtain multiple text prediction boxes;
[0009] S2: Aggregate prediction box of multiple models;
[0010] S3: Perform multi-model recognition on the aggregated prediction boxes to obtain the recognition results and confidence of each model;
[0011] S4: Assign recognition weights to each model based on its proficiency index in a given task scenario;
[0012] S5: Calculate the final recognition result of each character by combining the recognition weight and confidence.
[0013] In some embodiments of the present invention, step S2 includes:
[0014] The confidence of the prediction boxes obtained by different models is brought into their coordinates to form a target trust coordinate group:
[0015] ;in, is the predicted box coordinate, n is the number of coordinates; The confidence of the model output prediction box, m is the number of models;
[0016] use = and = The prediction frames of each model are integrated to obtain the target prediction frame:
[0017] [ ];in, ; The confidence after fusion is adopted calculate.
[0018] In some embodiments of the present invention, step S4 includes:
[0019] Count the recognition accuracy of each model in a specified operating scenario;
[0020] The recognition accuracy of multiple models is ranked, and recognition weights are assigned to each model according to the ranking results.
[0021] In some embodiments of the present invention, step S5 includes:
[0022] Obtain the recognition results and confidence levels of each character output by different models in S3 after recognizing the prediction box;
[0023] based on Calculate the recognition result of each character, where is the recognition weight of the model, The confidence level of the character recognized by the model.
[0024] A quay crane tallying system is proposed, including:
[0025] A multi-model prediction unit is used to use pre-trained multiple models to identify the image to be identified and obtain text prediction boxes;
[0026] Aggregation unit, used to aggregate prediction boxes of multiple models;
[0027] The multi-model recognition unit is used to perform multi-model recognition on the aggregated prediction boxes to obtain the recognition results and confidence levels of each model. Based on the proficiency index of each model in a specified operating scenario, a recognition weight is assigned to each model. The final recognition result of each character is calculated by combining the recognition weight and confidence level.
[0028] In some embodiments of the present invention, the polymeric unit is specifically used to:
[0029] The confidence of the prediction boxes obtained by different models is brought into their coordinates to form a target trust coordinate group:
[0030] ;in, is the predicted box coordinate, n is the number of coordinates; The confidence of the model output prediction box, m is the number of models;
[0031] use = and = The prediction frames of each model are integrated to obtain the target prediction frame:
[0032] [ ];in, ; The confidence after fusion is adopted calculate.
[0033] In some embodiments of the present invention, the multi-model recognition unit is specifically configured to:
[0034] Count the recognition accuracy of each model in a specified operating scenario;
[0035] The recognition accuracy of multiple models is ranked, and recognition weights are assigned to each model according to the ranking results.
[0036] In some embodiments of the present invention, the multi-model recognition unit is specifically configured to:
[0037] Obtain the recognition results and confidence levels of each character output by different models after recognizing the prediction box;
[0038] based on Calculate the recognition result of each character, where is the recognition weight of the model, The confidence level of the character recognized by the model.
[0039] Compared with the prior art, the advantages and positive effects of the present invention are as follows: in the quay crane OCR recognition method and quay crane tallying system proposed in the present invention, multiple models are used to respectively recognize the image to be recognized to obtain multiple prediction frames, the multiple prediction frames are aggregated to obtain a final text prediction frame, and the aggregated prediction frames are respectively subjected to character recognition using multiple models to obtain the recognition result and confidence of each model, and a recognition weight is assigned to each model based on the proficiency index of each model in a specified operation scenario, and the final recognition result of each character is calculated in combination with the recognition weight and confidence; the present invention uses multi-model fusion technology in the two stages of prediction frame and character recognition, introduces the advantages of each model, effectively solves the influence of factors such as poor image capture effect, incomplete and damaged container numbers, and different container number position and structure states, and greatly improves the recognition rate of quay crane OCR. After actual port operation tests, the container number recognition accuracy under various working conditions is greater than 99%, thereby ensuring the efficient operation of the intelligent tallying system, avoiding manual processing caused by recognition errors, and achieving the goal of reducing costs and increasing efficiency in port operations.
[0040] Other features and advantages of the present invention will become more apparent after reading the detailed description of the embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings, as part of the present invention, are intended to provide a further understanding of the present invention. The exemplary embodiments and descriptions of the present invention are intended to explain the present invention but do not constitute undue limitations thereon. Obviously, the drawings described below are merely examples, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0042] Figure 1 This is a schematic diagram of the steps of the quay crane OCR recognition method proposed in the present invention;
[0043] Figure 2 It is a diagram of the prediction box obtained by multiple models in the method of the present invention;
[0044] Figure 3 This is a diagram of the prediction box after aggregating multiple model prediction boxes in the method of the present invention;
[0045] Figure 4 This is a schematic diagram of the system structure of the quay crane tallying system proposed in the present invention.
[0046] It should be noted that these drawings and textual descriptions are not intended to limit the conceptual scope of the present invention in any way, but rather to illustrate the concept of the present invention for those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. The following embodiments are used to illustrate the present invention but are not used to limit the scope of the present invention.
[0048] In the description of the present invention, it should be noted that the terms "upper", "lower", "front", "back", "left", "right", "vertical", "inside", "outside", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, they cannot be understood as limiting the present invention.
[0049] In the description of the present invention, it should be noted that, unless otherwise specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; and direct or through an intermediary. Those skilled in the art will understand the specific meanings of these terms in the present invention based on the specific circumstances.
[0050] The quay crane OCR recognition method proposed in this invention is as follows: Figure 1 As shown, the following steps are included:
[0051] S1: Use pre-trained multiple models to identify the image to be identified and obtain multiple text prediction boxes.
[0052] In this embodiment, the PP-OCR model, the Paddle OCR model, and the CnOCR model are used as examples to detect the items to be recognized and obtain three text prediction boxes.
[0053] For the PP-OCR and Paddle OCR models: The image is first normalized and resized before being fed into the backbone network for feature extraction. The backbone network progressively extracts image features through multiple convolutional and pooling layers, transforming the image from its original pixel space to a feature space. These features are abstracted into feature maps of varying scales and channels, containing information about the structure and strokes of the text within the image. The feature maps are then fed into the detection head, where convolution operations generate features relevant to object detection, including the location and category of the predicted box. This involves generating a set of coordinate values representing the predicted box location and a confidence score indicating whether the location contains text (the likelihood of the text being present). Finally, the detection head output undergoes post-processing operations such as non-maximum suppression (NMS). NMS removes highly overlapping prediction boxes based on their confidence scores, retaining only those with the highest confidence, resulting in the final prediction box.
[0054] For the CnOCR model: first, after a series of processing such as grayscale conversion, noise reduction, binarization, and tilt correction, the binarized image is subjected to connected domain analysis, and the interconnected white (or black) pixels in the image are divided into different connected domains. Each connected domain corresponds to a possible character or character combination, and the area, perimeter, circumscribed rectangle and other features of each connected domain are calculated; then, according to preset character feature conditions, such as the size range and aspect ratio of the character, the connected domains are screened to remove connected domains that do not meet the character characteristics, such as noise points, too small or too large areas, etc.; for some adhered characters, the adhered characters are segmented into individual characters based on the structural characteristics of the characters, stroke trends and other information. For each character or character combination that has been filtered and grouped, its bounding rectangle (preliminary prediction box) is calculated, and its coordinates are determined by the boundary coordinates of the connected domain. Finally, the generated preliminary prediction box is optimized and adjusted. For example, the size and position of the prediction box are fine-tuned according to the actual shape and distribution of the character to make it more tightly surround the character, thereby improving positioning accuracy. At the same time, prediction boxes with high overlap and obvious unreasonableness are also removed, and the prediction box with the highest confidence is retained.
[0055] The prediction frames obtained by the PP-OCR model, Paddle OCR model and CnOCR model are as follows Figure 2 shown.
[0056] S2: Aggregate prediction boxes from multiple models.
[0057] A(x 11 ,y 11 ), A(x 12 ,y 12 ),……,A(x 1n ,y 1n) represents the predicted box coordinates obtained by the PP-OCR model, and its confidence is recorded as S1.
[0058] Take B(x 21 ,y 21 ), B(x 22 ,y 22 ),……,B(x 2n ,y 2n ) represents the predicted box coordinates detected by the Paddle COR model, and its confidence is recorded as S2.
[0059] C(x 31 ,y 31 ), C(x 32 ,y 32 ),……,C(x 3n ,y 3n ) represents the predicted box coordinates obtained by the CnOCR model detection, and its confidence is recorded as S3.
[0060] (1) The confidence of the prediction boxes of different models is brought into their coordinates to form a target trust coordinate group.
[0061] The target trust coordinate group formed by the PP-OCR model is: A[x 11 ,y 11 , x 12 ,y 12 ,……, x 1n ,y 1n , S1]; the target trust coordinate group formed by the Paddle OCR model is: B[x 21 ,y 21 , x 22 ,y 22 ,……, x 2n ,y 2n , S2]; the target information coordinate group formed by the CnOCR model is: C[x 31 ,y 31 , x 32 ,y 32 ,……, x 3n ,y 3n , S3].
[0062] (2) Adoption , The prediction frames of each model are fused to obtain the target prediction frame, where the confidence after fusion is calculated using calculate.
[0063] in, is the number of models, is the number of coordinate points.
[0064] The calculated fused target prediction box is ,like Figure 3 shown.
[0065] S3: Perform multi-model recognition on the aggregated prediction box to obtain the recognition result and confidence of each model.
[0066] In this embodiment, the PP-OCR model, Paddle OCR model, and CnOCR model are still used as examples to identify the aggregated prediction boxes respectively:
[0067] The PP-OCR model performs character recognition on the aggregated prediction box: the image within the prediction box is preprocessed, including image resizing and normalization, to meet the input requirements of the recognition model. The preprocessed image is then input into a convolutional neural network (CNN). Through a series of operations such as convolution and pooling, the CNN extracts feature sequences from the image. These feature sequences contain control information and semantic information about the characters in the image. The feature sequences extracted by the CNN are input into an RNN (recurrent neural network). The RNN models the feature sequence and captures contextual information between characters. A bidirectional LSTM (Long Short-Term Memory) is typically used as the RNN unit to better capture contextual information. The RNN outputs a probability distribution sequence, with each time step corresponding to the probability distribution of a character. Due to the inconsistent lengths of the input and output sequences, the CTC (Connectionist Temporal Classification) loss function is used for training and decoding. The CTC decoding process seeks the most likely character sequence and obtains the final recognition result and confidence level by removing duplicate and blank characters.
[0068] The Paddle OCR model performs character recognition on the aggregated prediction box: the image within the prediction box is preprocessed, including image resizing and normalization. The preprocessed image is sent to the convolution layer, which performs convolution operations through multiple convolution kernels to extract local features of the image and form a feature sequence. These feature sequences contain information such as the shape and strokes of the characters. The feature sequence output by the convolution layer is input to the recurrent layer (using GRU (Gated Recurrent Unit) as the recurrent unit). The recurrent layer processes the sequence data and captures the contextual information between characters by considering the order of the characters. The output of the recurrent layer is a probability distribution sequence, and each time step corresponds to the probability distribution of a character. Since the length of the input feature sequence may be different from the actual character sequence length, and there may be blank spaces between characters, the CTC loss function is used for training and decoding. During the CTC decoding process, the most likely character sequence is found, and the final recognition result and the corresponding confidence level are obtained by removing repeated characters and blank characters.
[0069] The CnOCR model performs character recognition on the aggregated prediction box: the image within the prediction box is preprocessed, including cropping, grayscale conversion, normalization, and resizing. The preprocessed image is input into the CNN network, where it undergoes alternating processing through multiple convolutional and pooling layers to gradually extract features at different levels. Shallow convolutional layers extract simple local adjustments, such as edges and corners. As the network is further extended, the features extracted by the convolutional layers become increasingly abstract, containing more semantic information and better representing the overall characteristics of the character. The feature sequence extracted by the CNN is input into the RNN, which updates the current hidden state based on the current input and the previous hidden state. In this way, the network captures contextual information between characters, learns the sequential pattern of character sequences, and thus better understands the semantics of the characters. The RNN outputs a sequence of probability distributions, with each time step corresponding to the probability distribution of a character. Due to the mismatch in length between the input and output sequences, the CTC loss function is used for training and decoding. The CTC decoding process seeks the most likely character sequence and obtains the final recognition result and confidence level by removing duplicate characters and blank characters.
[0070] S4: Assign recognition weights to each model based on its proficiency index in a specified job scenario.
[0071] In this embodiment of the present invention, each model's recognition results are weighted based on their algorithm's proficiency in the dock operation scenario (represented by a proficiency index). The final recognition result is calculated by combining the weights with the confidence level of each character in the model's output recognition result. Specifically, the following factors are considered:
[0072] 1. Count the recognition accuracy of each model in a specified operation scenario.
[0073] For the three models of PP-OCR, Paddle OCR, and CnOCR, their respective recognition accuracy rates were statistically analyzed through experimental data, and the recognition accuracy rate was used as the proficiency index.
[0074] 2. Sort the recognition accuracy of multiple models and assign recognition weights to each model based on the sorting results.
[0075] According to practical statistics, in the container terminal container number recognition scenario, the accuracy of the three models of PP-OCR, Paddle OCR, and CnOCR for container number recognition results is as follows from high to low: Paddle OCR, PP-OCR, CnOCR, and the total weight assigned to each model is 1. , .
[0076] The above embodiment uses recognition accuracy as the proficiency index of the model in a specified operation scenario. In actual applications, other indicators can also be used as proficiency indexes based on needs, such as model stability, model calculation speed, etc.
[0077] S5: Calculate the final recognition result of each character by combining the recognition weight and confidence.
[0078] Get the recognition results and confidence of each character output by different models after recognizing the prediction box in step S3, based on Calculate the recognition result of each character, where is the recognition weight of the model, The confidence level of the character recognized by the model.
[0079] For example, the character recognition results and confidence levels output by the three models Paddle OCR, PP-OCR, and CnOCR after recognizing the prediction box are: "Z, 0.67", "2, 0.97", and "2, 0.86", respectively. 0. and For example, the recognition result of each character is calculated by weighted calculation:
[0080] "Z, 40%", ;
[0081] "2,57%", .
[0082] In combination with the above, the present invention also proposes a quay crane tally system, such as Figure 4 As shown, including:
[0083] The multi-model prediction unit is used to use pre-trained multi-models to recognize the image to be recognized and obtain the text prediction box.
[0084] Aggregation unit, used to aggregate prediction boxes of multiple models.
[0085] The multi-model recognition unit is used to perform multi-model recognition on the aggregated prediction boxes to obtain the recognition results and confidence levels of each model. Based on the proficiency index of each model in a specified operating scenario, a recognition weight is assigned to each model. The final recognition result of each character is calculated by combining the recognition weight and confidence level.
[0086] In some embodiments of the present invention, the polymeric unit is specifically used to:
[0087] The confidence of the prediction boxes obtained by different models is brought into their coordinates to form a target trust coordinate group:
[0088] ;in, is the predicted box coordinate, n is the number of coordinates; The confidence of the model output prediction box, m is the number of models;
[0089] use = and = The prediction frames of each model are integrated to obtain the target prediction frame:
[0090] ;in, ; The confidence after fusion is adopted calculate.
[0091] In some embodiments of the present invention, the multi-model recognition unit is specifically configured to:
[0092] Count the recognition accuracy of each model in a specified operating scenario;
[0093] The recognition accuracy of multiple models is ranked, and recognition weights are assigned to each model according to the ranking results.
[0094] In some embodiments of the present invention, the multi-model recognition unit is specifically configured to:
[0095] Obtain the recognition results and confidence levels of each character output by different models after recognizing the prediction box;
[0096] based on Calculate the recognition result of each character, where is the recognition weight of the model, The confidence level of the character recognized by the model.
[0097] The specific OCR recognition method of the quay crane tally system has been described in detail above and will not be repeated here.
[0098] It should be noted that, in the specific implementation process, the above-mentioned control part can be implemented by a hardware processor executing computer execution instructions in software form stored in the memory, which will not be elaborated here. The programs corresponding to the actions performed by the above-mentioned control circuit can be stored in the system's computer-readable storage medium in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0099] The computer-readable storage medium mentioned above may include volatile memory, such as random access memory; may also include non-volatile memory, such as read-only memory, flash memory, hard disk or solid-state drive; may also include a combination of the above types of memory.
[0100] The processor mentioned above can also be a collective term for multiple processing elements. For example, the processor can be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices (PLDs), discrete gate or transistor logic devices (LDDs), discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor, etc., and can also be a special-purpose processor.
[0101] It should be pointed out that the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by ordinary technicians in this technical field within the essential scope of the present invention should also fall within the scope of protection of the present invention.
Claims
1. A quay crane OCR recognition method, characterized in that: include: S1: Use pre-trained multiple models to identify the image to be identified and obtain multiple text prediction boxes; S2: Aggregate prediction box of multiple models; S3: Perform multi-model recognition on the aggregated prediction boxes to obtain the recognition results and confidence of each model; S4: Assign recognition weights to each model based on its proficiency index in a given task scenario; S5: Calculate the final recognition result of each character by combining the recognition weight and confidence.
2. The quay crane OCR recognition method according to claim 1, characterized in that: Step S2 includes: The confidence of the prediction boxes obtained by different models is brought into their coordinates to form a target trust coordinate group: ;in, is the predicted box coordinate, n is the number of coordinates; The confidence of the model output prediction box, m is the number of models; use = and = The prediction frames of each model are integrated to obtain the target prediction frame: [ ];in, ; The confidence after fusion is adopted calculate.
3. The quay crane OCR recognition method according to claim 1, characterized in that: Step S4 includes: Count the recognition accuracy of each model in a specified operating scenario; The recognition accuracy of multiple models is ranked, and recognition weights are assigned to each model according to the ranking results.
4. The quay crane OCR recognition method according to claim 1 or 3, characterized in that: Step S5 includes: Obtain the recognition results and confidence levels of each character output by different models in S3 after recognizing the prediction box; based on Calculate the recognition result of each character, where is the recognition weight of the model, The confidence level of the character recognized by the model.
5. A quay crane tallying system, characterized in that: include: A multi-model prediction unit is used to use pre-trained multiple models to identify the image to be identified and obtain text prediction boxes; Aggregation unit, used to aggregate prediction boxes of multiple models; The multi-model recognition unit is used to perform multi-model recognition on the aggregated prediction box to obtain the recognition result and confidence of each model; Assign recognition weights to each model based on its proficiency index in a given job scenario; The recognition weight and confidence are combined to calculate the final recognition result of each character.
6. The quay crane tallying system according to claim 5, characterized in that: The polymerization unit is specifically used for: The confidence of the prediction boxes obtained by different models is brought into their coordinates to form a target trust coordinate group: ;in, is the predicted box coordinate, n is the number of coordinates; The confidence of the model output prediction box, m is the number of models; use = and = The prediction frames of each model are integrated to obtain the target prediction frame: [ ];in, ; The confidence after fusion is adopted calculate.
7. The quay crane tallying system according to claim 5, characterized in that: The multi-model recognition unit is specifically used for: Count the recognition accuracy of each model in a specified operating scenario; The recognition accuracy of multiple models is ranked, and recognition weights are assigned to each model according to the ranking results.
8. The quay crane tallying system according to claim 5 or 7, characterized in that: The multi-model recognition unit is specifically used for: Obtain the recognition results and confidence levels of each character output by different models after recognizing the prediction box; based on Calculate the recognition result of each character, where is the recognition weight of the model, The confidence level of the character recognized by the model.