A ship name recognition method based on image recognition and font vocabulary similarity

By combining image recognition and character shape-vocabulary similarity algorithms with ship name structure preprocessing and vocabulary similarity calculation, the problem of inconsistent and mixed fonts in Chinese ship name recognition is solved, achieving high-accuracy ship name recognition.

CN115424281BActive Publication Date: 2025-12-09JIANGSU CENTURY INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211132922.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-16
Publication Date
2025-12-09
Estimated Expiration
2042-09-16

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as missing fonts, inconsistent sizes, and mixing of Chinese and English when identifying Chinese ship names, resulting in low recognition accuracy.

Method used

A method based on image recognition and glyph and word similarity is adopted. By preprocessing the ship name structure, using font similarity and word similarity algorithms, and combining artificial intelligence optical character recognition technology, ship name recognition is performed.

Benefits of technology

It effectively solves recognition errors in cases such as name line breaks, reverse order, characters that are too far apart, and Chinese characters mixed with pinyin, improving the recognition accuracy to over 98%, and adapting to complex environments such as poor lighting and unclear text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424281B_ABST
    Figure CN115424281B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of ship name identification method based on image recognition and character similarity, the method includes the following steps: step 1: ship picture is based on artificial intelligence optical character recognition, step 2: the possibility of judging the recognized text as ship name;Step 3: consider the error mode such as line break, reverse order, text is too far apart, Chinese mixed pinyin;Step 4: by ship name library, direct matching search is carried out;Step 5: for the text that cannot be matched as a whole, by ship name library, obtain the alternative ship name set;Step 6: according to the structural features of Chinese character, the difference part text is carried out character disassembly, and character similarity calculation is carried out etc..The technical scheme can effectively improve the ship name recognition rate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a calculation analysis method, in particular to a ship name recognition method based on image recognition and character font vocabulary similarity, and belongs to the technical field of image recognition. BACKGROUND

[0002] With the rapid development of information technology and the rapid development of optical character recognition technology based on artificial intelligence, related applications are gradually developing. However, because of the actual character identification, the font is missing, the font size is inconsistent, and the English and Chinese are mixed, which leads to the difficulty in font recognition, or the unreasonable combination of vocabulary. In view of the specific problem of Chinese ship name, on the basis of the relatively limited total number of about 300,000 ship names, by means of the Chinese specific structure and the relatively fixed characteristics of the ship name structure, the ship name structure preprocessing, font similarity and vocabulary similarity algorithms are constructed, and the optimization and enhancement of the ship name recognition are realized. SUMMARY

[0003] The application is just for the problems existing in the prior art, and provides a ship name recognition method based on image recognition and character font vocabulary similarity. The optical character recognition result of artificial intelligence is combined with the actual situation of ship name identification and the construction of vocabulary similarity based on Chinese characteristics. The technical scheme of the optical character recognition technology of artificial intelligence realizes the ship name structure preprocessing, font similarity and vocabulary similarity algorithms under the limited vocabulary set, effectively solves the problems of low ship name recognition accuracy caused by the actual situation such as name line change, name reverse order, character separation too far apart, Chinese mixed with pinyin, incomplete font structure, unclear font, and other image interference.

[0004] In order to realize the above purpose, the technical scheme of the application is as follows: a ship name recognition method based on image recognition and character font vocabulary similarity, the method comprising the following steps:

[0005] Step 1: performing optical character recognition based on artificial intelligence on the ship picture to obtain the position of the character in the picture and the possible first five character combinations;

[0006] Step 2: judging the possibility of the recognized character being a ship name;

[0007] Step 3: considering the error mode of line change, reverse order, character separation too far apart, and Chinese mixed with pinyin, and further preprocessing;

[0008] Step 4: directly matching and searching through the ship name library;

[0009] Step 5: for the characters that cannot be matched as a whole, obtaining a set of alternative ship names from the ship name library;

[0010] Step 6: For the set of alternative ship name, according to the structure characteristics of Chinese characters, the difference part of the characters is disassembled, the character similarity is calculated, and the overall similarity is calculated according to the way of high weight on both sides of the middle weight. The top 2 matching results and the overall similarity with the highest similarity are returned.

[0011] Step 2: Determine the possibility of identifying characters as ship names: According to the structure requirements and characteristics of Chinese ship names, the high probability of ship names enters step 4, and the high probability of non-ship name characters from step 1 enters step 3. For the high probability of non-ship name characters from step 3, return the non-ship name result and similarity 0%, and end the process. Step 3: Specifically as follows, including: combination of line break characters, reverse, combination of far apart characters, remove Chinese pinyin, combine the processed characters, and enter step 2 again for judgment;

[0012] Step 4: Directly match and search through the ship name library: 1) Directly return the matching result and similarity 100% for complete matching, and end the process; 2) Identify the middle part of the ship name as the matching result and similarity of the identified characters / matching ship length, and end the process; 3) For the above search, enter step 5 if no result is obtained;

[0013] Step 5: Obtain a set of alternative ship names from the ship name library: Case 1) For cases with numbers and more than 3 numbers, obtain a set of alternative ship names through number matching. If the set of alternative ship names is empty, obtain a set of alternative ship names through Chinese similarity matching based on the Chinese part, and the algorithm refers to case 2; Case 2) For other cases, remove and insert 1 character at any position to search and obtain a set of alternative ship names. If the set of alternative ship names is empty, add the number of removed and inserted characters to expand the search range until half of the original number of identified characters is reached. If the set of alternative ship names is still empty, return the identified characters and similarity 1%, end the process, and other cases enter step 6.

[0014] In step 3, the identification problem is preprocessed according to the ship name structure, and the method is as follows:

[0015] 1) For the case where the ship name character image has line breaks, if the font size is consistent, the combination method of two rows or two small rows plus one large row is used, and the words are combined from left to right and from top to bottom;

[0016] 2) For the case of considering the reverse order of ship name character image, the reverse order recombination method is used;

[0017] 3) For the case that the individual characters of the ship name character image are too far apart, 1) for the case that the tail font is too far apart, if the font size is consistent, the tail font is combined in the ship name, and the tail font set is collected, then the words are combined from left to right; 2) for the case that the overall font is too far apart, if the font size is consistent, the recognition is a single word, and the interval distance fluctuation is less than 10%, then the words are combined from left to right;

[0018] 4) For the case that Chinese characters are mixed with pinyin in the ship name character image, for example: su su wu xi hu 001, that is, the case that one Chinese character is combined with multiple letter combinations, the pinyin between the Chinese characters is removed to recombine.

[0019] In step 6, for the set of alternative ship names, the similarity of the single difference character is calculated, and the method is as follows:

[0020] Step 1, for the recognized character A and the matched character B, according to the structure of Chinese characters, including: 1) left and right structure, such as: struggle, great, rest, da, Ming, sand; 2) upper and lower structure, such as: Zhi, Miao, Zi, Wei, Sui, Army; 3) left, middle and right structure, such as: lake, foot, splash, Xie, do, porridge; 4) upper, middle and lower structure, such as: Xi, Jiao, Neng, Hen, Ying, Yan; 5) half-enclosed structure: such as: sentence, can, si, style, soldier, lice; 6) fully enclosed structure, such as: prisoner, team, cause, link, round, country; 7) inlaid structure, such as: sit, cool, clip, evil, wizard, and wipe out, to disassemble the characters;

[0021] Step 2, for the disassembled image structure, the similarity is counted, case 1) if the structure is exactly the same, then the part with the same structure is counted as 1; 2) if the structure is similar, such as: spoon, uniform, then count 0.5; case 3) complete dissimilarity, count 0;

[0022] Step 3, for the recognized character A and the matched character B, the similarity of the disassembled structure is accumulated and summarized, and the result is divided by the total number of disassembled structures, which is used as the character shape similarity.

[0023] In step 6, for the set of alternative ship names, the overall similarity is calculated on the basis of the similarity of the single difference text, and the method is as follows:

[0024] First, for the recognized character combination A1A2A3 and the matched character combination B1B1B2, the similarity of each character is calculated, which is 1 for complete consistency, and the difference is calculated according to the character shape similarity to obtain the similarity; then, the character shape similarity of each character is superimposed and summarized, and the length of the recognized character combination A1A2A3 and the matched character combination B1B1B2 is summed up to obtain the overall similarity.

[0025] Compared with the prior art, the present application has the following advantages:

[0026] 1) The actual situation of line break, reverse order, too far apart characters, Chinese mixed with pinyin in the ship name image is pre-processed, and the recognition rate of ship name recognition only relying on image recognition is improved from 0% to the same level of image recognition accuracy, generally more than 98%;

[0027] 2) Combined with the structure of Chinese characters, the recognition error caused by paint drop, shielding and other symbol interference is effectively solved, and the recognition rate of ship name recognition only relying on image recognition is improved from 0% to the same level of image recognition accuracy, generally more than 98%;

[0028] 3) According to the characteristics of limited ship name and fixed ship name structure, the accuracy of recognition is effectively improved through the whole word similarity algorithm;

[0029] 4) According to the characteristics of limited ship name and fixed ship name structure, the similarity is calculated by searching for a set of candidate ship names and calculating the similarity of Chinese characters, which has the advantages of simple algorithm structure, small calculation amount and high recognition accuracy compared with the word vector algorithm;

[0030] 5) The algorithm can well adapt to poor light, unclear characters and other image interference through preprocessing, Chinese character structure similarity calculation and word group structure similarity calculation, and has better pertinence, high efficiency and wide application range compared with other algorithms. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 The whole step schematic diagram of embodiment 1 of the application is shown in the figure;

[0032] Figure 2 The whole step schematic diagram of embodiment 2 of the application is shown in the figure. DETAILED DESCRIPTION

[0033] In order to deepen the understanding of the application, the following detailed description of the embodiment is made in combination with the drawings.

[0034] Embodiment 1: Referring to Figure 1 A ship name recognition method based on image recognition and word similarity, comprising the following steps:

[0035] Step 1: The ship image is recognized by the image recognition technology based on artificial intelligence to obtain the recognized character combination "100ad Da Gnem Alliance";

[0036] Step 2: Judge whether it is a Chinese ship name structure, that is, all Chinese, or Chinese plus numbers, or Chinese plus numbers plus "number", the result is not a Chinese ship name structure, enter step 3;

[0037] Step 3, preprocessing the non-Chinese ship name structure word group,

[0038] Step 3.1, first according to the first letter is a number, reverse processing, get "meng da 001";

[0039] Step 3.2, judging as a single Chinese + multiple letters structure, removing processing, get "meng da 001";

[0040] Step 4, "meng da 001" in Chinese ship name set for direct matching, return empty set, step 5;

[0041] Step 5, "meng da 001" through Chinese ship name set for alternative ship name set acquisition, get non-empty set ["ming da 001", "macro da 001"];

[0042] Step 6, calculate the lexical similarity, and return the result, as follows:

[0043] 1) "meng da 001" and "ming da 001" are calculated:

[0044] Step 6.1, difference font similarity calculation: meng da is day, month, and dish; Ming da is day, month, that is, the number of day + day + month + month is 4, divided by the total structure day, month, dish, day, month, which is 5, which is 80%;

[0045] Step 6.2, "meng da 001" and "ming da 001" lexical similarity calculation, 0.8 (meng) + 1 (da) + 1 (0) + 1 (0) + 1 (1) + 0.8 (ming) + 1 (da) + 1 (0) + 1 (0) + 1 (1) = 9.6, divided by 5 (meng da 001) + 5 (ming da 001) = 10, similarity is 96%;

[0046] 2) "meng da 001" and "macro da 001" are calculated:

[0047] Step 6.1, difference font similarity calculation: meng da is day, month, and dish; Macro da is nian, xiong, no similar structure, 0%;

[0048] Step 6.2, "meng da 001" and "macro da 001" lexical similarity calculation, 0 (meng) + 1 (da) + 1 (0) + 1 (0) + 1 (1) + 0 (macro) + 1 (da) + 1 (0) + 1 (0) + 1 (1) = 8, divided by 5 (meng da 001) + 5 (ming da 001) = 10, similarity is 80%.

[0049] Example 2: see Figure 2As shown in the figure, a ship name recognition method based on image recognition and glyph-lexical similarity includes the following steps:

[0050] Step 1: Through artificial intelligence-based image recognition technology, the ship image is recognized to obtain the recognized text combination "su Suquanhang 5311".

[0051] Step 2: Determine whether it is a Chinese ship name structure, that is, all Chinese, or Chinese plus numbers, or Chinese plus numbers plus "number". If the result is not a Chinese ship name structure, enter Step 3.

[0052] Step 3: Preprocess the non-Chinese ship name structure phrase.

[0053] Step 3.1: If it is determined to be a single Chinese + multiple letter structure, perform removal processing to obtain "Suquanhang 5311".

[0054] Step 4: Directly match "Suquanhang 5311" in the Chinese ship name set. If the returned set is empty, enter Step 5. <00001十一7>

[0055] Step 5: Obtain the alternative ship name set for "Suquanhang 5311" through the number "5311" to obtain a non-empty set ["Wanquan 5311", "Wanquanhang 5311", "Sugaohang 5311"].

[0056] Step 6: Calculate the lexical similarity and return the result, specifically as follows:

[0057] 1) Calculate "Suquanhang

[0058] Step 6.1: Calculate the similarity of different fonts: Su is disassembled into 艹 and 办; Wan is disassembled into 白 and 完, with no similar structure, which is 0%.

[0059] Step 6.2: Calculate the lexical similarity between "Suquanhang 5311" and "Wanquan 5311", which is 0(Su)+1(Quan)+0(Hang)+1(5)+1(3)+1(1)+1(1)+0(Wan)+1(Quan)+1(5)+1(3)+1(1)+1(1)=10, divided by 7(Suquanhang 5311)+7(Wanquan 5311)=13, and the similarity is 76.9%.

[0060] 2) Calculate "Suquanhang 5311" and "Wanquanhang 5311":

[0061] Step 6.1: Calculate the similarity of different fonts: Su is disassembled into 艹 and 办; Wan is disassembled into 白 and 完, with no similar structure, which is 0%.

[0062] Step 6.2, "Su Quan Hang 5311" and "Wan Quan Hang 5311" lexical similarity calculation, 0 (Su) + 1 (Quan) + 1 (Hang) + 1 (5) + 1 (3) + 1 (1) + 1 (1) + 0 (Wan) + 1 (Quan) + 1 (Hang) + 1 (5) + 1 (3) + 1 (1) + 1 (1) = 12.8, divided by 7 (Su Quan Hang 5311) + 7 (Wan Quan Hang 5311) = 14, similarity is 91.4%;

[0063] 3) "Su Quan Hang 5311" and "Su Jiao Hang 5311" calculation:

[0064] Step 6.1, difference font similarity calculation: Quan is broken into Bai, Ban; Jiao is broken into Bai, Da, Shi, that is, the number of Bai + Bai is 2, divided by the total structure Bai, Ban, Bai, Da, Shi, which is 5, is 40%;

[0065] Step 6.2, "Su Quan Hang 5311" and "Su Jiao Hang 5311" lexical similarity calculation, 1 (Su) + 0.4 (Quan) + 1 (Hang) + 1 (5) + 1 (3) + 1 (1) + 1 (1) + 1 (Su) + 0.4 (Jiao) + 1 (Hang) + 1 (5) + 1 (3) + 1 (1) + 1 (1) = 12.8, divided by 7 (Su Quan Hang 5311) + 7 (Su Jiao Hang 5311) = 14, similarity is 91.4%;.

Claims

1. A ship name recognition method based on image recognition and font lexicon similarity, characterized by, The method comprises the following steps: Step 1: AI-based optical character recognition is performed on the ship picture to obtain the position of the text in the picture and possible top 5 text combinations; Step 2: Determine the possibility of the recognized text being a ship name; Step 3: Further preprocessing is performed considering the error modes of line breaks, reverse order, text being too far apart, and Chinese mixed with pinyin; Step 4: Direct matching retrieval is performed through the ship name library; Step 5: For the text that cannot be matched as a whole, the ship name library is used to obtain a set of alternative ship names; Step 6: For the set of alternative ship names, the text is disassembled according to the characteristics of Chinese text structure, the similarity of the characters is calculated, and then the overall similarity is calculated according to the method of low weight in the middle and high weight on both sides, returning the top 2 matching results and the overall similarity with the highest similarity. In step 5, the ship name library is used to obtain a set of alternative ship names: case 1) for cases with numbers and more than 3 numbers, the alternative ship name set is obtained through number matching, if the alternative ship name set is empty, the Chinese part is used to obtain the alternative ship name set according to the Chinese similarity; case 2) for cases without numbers or with less than or equal to 3 numbers, the alternative ship name set is obtained by removing or inserting 1 character at any position, if the alternative ship name set is empty, the number of deletable and insertable characters is gradually increased from 1, the search range is expanded, and if the alternative ship name set is still empty, the recognized text and a similarity of 1% are returned, and the process ends. In step 6, the similarity of a single different character in the set of alternative ship names is calculated as follows: Step 1: Disassemble the recognized text A and the matching text B according to the structure of Chinese characters, including: 1) left and right structure; 2) top and bottom structure; 3) left, middle and right structure; 4) top, middle and bottom structure; 5) half-enclosed structure; 6) fully enclosed structure; 7) inlaid structure; Step 2: Count the similarity of the disassembled image structure, case 1) if the structure is exactly the same, count the structure as 1; 2) if the structure is similar, such as spoon and uniform, count as 0.5; case 3) completely different, count as 0; Step 3: Add up the structure similarity counts of the disassembled A and B, and divide by the total number of disassembled structures to obtain the character shape similarity; In step 6, the overall similarity calculation is based on the similarity of a single different character in the set of alternative ship names, as follows: First, calculate the similarity of each character in the recognized text combination A1A2A3 and the matching text combination B1B1B2, which is 1 for complete consistency and is obtained by calculating the character shape similarity for differences; then, add up the character shape similarity calculations of each character to obtain the overall similarity calculation of the recognized text combination A1A2A3 and the matching text combination B1B1B2. ​ 2. The ship name recognition method based on image recognition and font similarity according to claim 1, characterized in that, Step 2: Determine the possibility of recognizing the text as a ship name: According to the structure and characteristics of Chinese ship name, the high probability of ship name enters step 4, the high probability of non-ship name text from step 1 enters step 3, and the high probability of non-ship name text from step 3 returns the result of non-ship name and similarity 0%, and the process ends.

3. The ship name recognition method based on image recognition and font similarity according to claim 1, characterized in that, Step 3: Specifically as follows, including: combination of line break text, reverse, combination of far away text, remove Chinese inter-pinyin, combine the processed text, and enter step 2 again for judgment.

4. The ship name recognition method based on image recognition and font similarity according to claim 1, characterized in that, Step 4: Direct matching retrieval through the ship name library: 1) Directly return the matching result and similarity 100% for complete matching, and end the process; 2) The recognized text is the middle part of the ship name, and only the missing matching of the first and last is directly returned; 3) For the above retrieval, enter step 5 if no result is obtained.

5. The ship name recognition method based on image recognition and font similarity according to claim 1, characterized in that, In step 3, the recognition problem is preprocessed according to the structure of the ship name, and the method is as follows: 1) For the case of line break in ship name character image, if the font size is consistent, the combination mode of two rows or two small rows plus one large row is adopted, and the word combination is performed from left to right edge and from top to bottom; 2) For the case of considering the reverse order of ship name character image, the reverse order recombination method is adopted; 3) For the case of individual text being too far apart in ship name character image, 1) for the case of tail font being too far apart, if the font size is consistent, the tail font is combined with the ship name tail font set, then the word combination is performed from left to right edge, 2) for the case of overall font being too far apart, if the font size is consistent, the recognition is single character and the interval distance fluctuation is less than 10%, then the word combination is performed from left to right edge; 4) For the case of Chinese inter-pinyin in ship name character image, that is, the case of one Chinese plus multiple letter combinations, the method of removing Chinese inter-pinyin is adopted for recombination.

Citation Information

Patent Citations

  • Font pattern recognition method, electronic device and storage medium

    CN109857912A

  • A method and system for ship number recognition based on a combination of English and Chinese characters.

    CN114937269A