Text and Image Recognition Retrieval Method, Device, Storage Medium and Electronic Device

By evaluating and scoring text within images based on algorithm performance, semantic coherence, and visibility, the method enhances data cleaning and retrieval accuracy in grapheme-based systems.

CN113821666BActive Publication Date: 2025-07-15BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110126962.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-29
Publication Date
2025-07-15
Estimated Expiration
2041-01-29

AI Technical Summary

Technical Problem

In the prior art, the graphic and text recognition search system fails to make full use of the diverse text information in the image, resulting in poor cleaning effect of text recognition results, affecting the search accuracy, especially when short words are searched.

Method used

By receiving images and identifying multiple texts, multiple scoring values are calculated based on text information, including graphic and text recognition algorithm scores, semantic integrity scores and eye-catching scores, weighting operations are performed to obtain comprehensive scores, sort and store them in the database, establishing the index relationship between text and images, and optimizing the search effect.

Benefits of technology

It realizes effective cleaning of text recognition results, improves the accuracy of image retrieval, especially when short words retrieval, significantly improves the accuracy of retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113821666B_ABST
    Figure CN113821666B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for graphic and text recognition and retrieval, an electronic device, and a computer-readable storage medium, relating to the technical field of graphic and text recognition. The method for graphic and text recognition and retrieval includes: receiving an image and storing it in a database, and recognizing multiple pieces of text in the image; obtaining corresponding multiple score values according to the text information of each piece of text, and obtaining a comprehensive score corresponding to each piece of text based on the multiple score values; storing each piece of text in the database based on the comprehensive score, so as to search for the corresponding image based on the scored text in the database. After recognizing the text in the image, the present disclosure obtains multiple score values by combining the recognized text information, obtains a comprehensive score based on the multiple score values, and stores the result of the above text recognition based on the comprehensive score, so as to achieve the effect of optimizing the retrieval result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of graphic and text recognition, and in particular, to a graphic and text recognition and retrieval method, a graphic and text recognition and retrieval device, an electronic device, and a computer-readable storage medium. Background Art

[0002] Retrieval refers to the process of finding the information or data needed from information collections such as literature materials and network information. With the advent of the information age, the carriers of information have become more diverse. For example, information is carried in the form of images. Images have been widely used in many fields because they can convey information more vividly and intuitively. For example, in the advertising field, images can bring better visual stimulation and achieve better publicity effects.

[0003] In order to better adapt to the progress of informatization, the research on graphic and text recognition and retrieval technology has become increasingly important. In related technologies, graphic and text recognition and retrieval focuses on providing a graphic and text retrieval system or expanding the retrieval function, without considering issues such as data cleaning of the recognized text result set and how to use the diverse text information in the image to optimize the retrieval effect.

[0004] Therefore, in order to solve the above problems, a graphic and text recognition and retrieval method is needed, which can achieve data cleaning of the recognition result by integrating various aspects of information of the text in the image during the text recognition process, thereby optimizing the effect of retrieving images based on the text.

[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The purpose of the embodiments of the present disclosure is to provide a graphic and text recognition and retrieval method, a graphic and text recognition and retrieval device, an electronic device, and a computer-readable storage medium, so as to achieve data cleaning of the recognition result by integrating various aspects of information of the text in the image during the text recognition process, thereby optimizing the effect of retrieving images based on the text.

[0007] According to a first aspect of the present disclosure, a graphic and text recognition and retrieval method is provided, including:

[0008] Receiving an image and storing it in a database, and recognizing multiple texts in the image;

[0009] Obtaining corresponding multiple score values according to the text information of each text, and obtaining a comprehensive score corresponding to each text based on the multiple score values;

[0010] Store each of the texts in the database based on the comprehensive score, so as to search for the corresponding images based on the scored texts in the database.

[0011] In an exemplary embodiment of the present disclosure, the recognizing multiple texts in the image includes:

[0012] Extract the multiple texts in the image based on an optical character recognition algorithm, and obtain the text information of each of the texts.

[0013] In an exemplary embodiment of the present disclosure, the obtaining corresponding multiple score values according to the text information of each of the texts includes:

[0014] Obtain an algorithm score of an algorithm for recognizing the corresponding text in the image, a semantic integrity score of each of the texts, and a prominence score of each of the texts in the image based on the text information;

[0015] Wherein, the semantic integrity score is determined by the ratio of the number of words after word segmentation of each of the texts to the total number of words, and the prominence score is determined by the ratio of each of the texts to the total area of all texts in the image.

[0016] In an exemplary embodiment of the present disclosure, the obtaining the comprehensive score corresponding to each of the texts based on the multiple score values includes:

[0017] Perform a weighted operation on the algorithm score, the semantic integrity score, and the prominence score to obtain the comprehensive score corresponding to each of the texts.

[0018] In an exemplary embodiment of the present disclosure, the storing each of the texts in the database based on the comprehensive score includes:

[0019] Sort each of the texts based on the comprehensive score, divide each of the texts into multiple levels according to the sorting result, filter the texts according to the multiple levels, and store the filtered texts in the database according to a preset ratio.

[0020] In an exemplary embodiment of the present disclosure, the searching for the corresponding images based on the scored texts in the database includes:

[0021] Receive a retrieval keyword, match the scored texts based on the retrieval keyword to obtain multiple secondary scoring results, and obtain a secondary comprehensive score based on the multiple secondary scoring results;

[0022] Sort the secondary comprehensive score and return a result set based on the sorting result, where the result set includes multiple of the images.

[0023] In an exemplary embodiment of the present disclosure, searching for the corresponding image based on the scored text in the database includes:

[0024] Receiving a retrieval keyword to respectively match the retrieval keyword from texts of different levels in the database, obtaining a plurality of secondary scoring results, and obtaining a secondary comprehensive score based on the secondary scoring results;

[0025] Sorting the secondary comprehensive scores and returning a result set based on the sorting result, the result set including a plurality of the images.

[0026] According to a second aspect of the present disclosure, there is provided a graphic and text recognition and retrieval device, including:

[0027] A text recognition module, configured to receive an image and store it in a database, and recognize multiple texts in the image;

[0028] A data cleaning module, configured to obtain corresponding multiple scoring values according to the text information of each text, and obtain a comprehensive score corresponding to each text based on the multiple scoring values;

[0029] An identification and retrieval module, configured to store each text in the database based on the comprehensive score, so as to search for the corresponding image based on the scored text in the database.

[0030] According to a third aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the method described in any one of the above by executing the executable instructions.

[0031] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the above is implemented.

[0032] The exemplary embodiments of the present disclosure may have the following partial or all beneficial effects:

[0033] In the text and image recognition and retrieval method provided by the exemplary embodiment of the present disclosure, an image is received and stored in a database, and multiple texts in the image are recognized; corresponding multiple score values are obtained according to the text information of each text, and a comprehensive score corresponding to each text is obtained based on the multiple score values; each text is stored in the database based on the comprehensive score, so as to search for the corresponding image in the database based on the scored text. On the one hand, after recognizing multiple texts in the image, the text and image recognition and retrieval method provided by this exemplary embodiment will also obtain the comprehensive score corresponding to each text through the text information of each text, making full use of all aspects of the text information, thereby realizing effective cleaning of the text recognition results. On the other hand, after effectively cleaning the text recognition results, the text recognition results are put into the database, so that the image can also be retrieved in the database through the cleaned text, improving the accuracy of the retrieval.

[0034] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings according to these drawings without creative efforts.

[0036] Figure 1 FIG. shows a schematic diagram of an exemplary system architecture of a text and image recognition and retrieval method and apparatus to which the embodiments of the present disclosure can be applied;

[0037] Figure 2 FIG. shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure;

[0038] Figure 3 FIG. schematically shows a flowchart of a text and image recognition and retrieval method according to an embodiment of the present disclosure;

[0039] Figure 4 FIG. schematically shows a system architecture diagram of an application scenario of a text and image recognition and retrieval method according to an embodiment of the present disclosure;

[0040] Figure 5 FIG. schematically shows a flowchart of data cleaning according to a specific application scenario of the present disclosure;

[0041] Figure 6 FIG. schematically shows a flowchart of picture retrieval according to a specific application scenario of the present disclosure;

[0042] Figure 7 A block diagram of a graphic recognition and retrieval device according to an embodiment of the present disclosure is schematically shown. Detailed implementation manners

[0043] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.

[0044] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0045] Figure 1 A schematic diagram of a system architecture of an exemplary application environment of a graphic recognition and retrieval method and device to which the embodiments of the present disclosure can be applied is shown.

[0046] As Figure 1 shown, the system architecture 100 may include one or more of the terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The terminal devices 101, 102, 103 may be electronic devices having a photographing or image transmission function, including but not limited to desktop computers, portable computers, smart phones, and tablet computers, etc. It should be understood that Figure 1 the number of terminal devices, networks, and servers in

[0047] The text and image recognition and retrieval method provided by the embodiments of the present disclosure can be executed by terminal devices 101, 102, and 103. Correspondingly, the text and image recognition and retrieval device can also be disposed in terminal devices 101, 102, and 103. The text and image recognition and retrieval method provided by the embodiments of the present disclosure can also be jointly executed by terminal devices 101, 102, and 103 and server 105. Correspondingly, the text and image recognition and retrieval device can be disposed in terminal devices 101, 102, and 103 and server 105. In addition, the text and image recognition and retrieval method provided by the embodiments of the present disclosure can also be executed by server 105. Correspondingly, the text and image recognition and retrieval device can be disposed in server 105. No special limitation is made in this exemplary embodiment.

[0048] For example, in the present exemplary embodiment, the above-mentioned text and image recognition and retrieval method can be jointly executed by terminal devices 101, 102, and 103 and server 105. First, an image can be captured or received by the terminal device. After receiving or capturing the image, the terminal device transmits the image to the server. The server stores the image in the database and calls an optical character recognition (OCR) algorithm to recognize multiple pieces of text in the image, and obtains corresponding multiple score values based on the text information of each piece of text. A comprehensive score is obtained based on the obtained multiple score values. Finally, based on the comprehensive score, the server stores the above-mentioned multiple pieces of text in the database for subsequent searching for the corresponding image in the database based on the scored text.

[0049] Figure 2 FIG. shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure.

[0050] It should be noted that Figure 2 The computer system 200 of the electronic device shown is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of the present disclosure.

[0051] As Figure 2 shown, the computer system 200 includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 202 or a program loaded from a storage section 208 into a random access memory (RAM) 203. In the RAM 203, various programs and data required for system operation are also stored. The CPU 201, ROM 202, and RAM 203 are connected to each other through a bus 204. An input / output (I / O) interface 205 is also connected to the bus 204.

[0052] The following components are connected to the I / O interface 205: an input section 206 including a keyboard, a mouse, etc.; an output section 207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. as well as a speaker, etc.; a storage section 208 including a hard disk, etc.; and a communication section 209 including a network interface card such as a LAN card, a modem, etc. The communication section 209 performs communication processing via a network such as the Internet. A drive 210 is also connected to the I / O interface 205 as needed. A removable medium 211 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 210 as needed so that a computer program read from it can be installed into the storage section 208 as needed.

[0053] With the advent of the information age, the carriers of information have become more diverse. For example, images have been widely used in many fields because they can convey information more vividly and intuitively. In order to better adapt to the progress of informatization, the research on text and image recognition and retrieval technology has become increasingly important.

[0054] The text and image recognition and retrieval technology mainly includes two aspects: data cleaning in the process of text recognition and searching for images by text. In related technologies, the technical means for data cleaning mainly include two aspects. On the one hand, the recognition algorithm is optimized to improve the recognition accuracy. On the other hand, the recognition results are scored by an algorithm, and filtering is performed through a set score threshold. Although this method can achieve effective cleaning, it ignores information such as the position and size of the text in the image.

[0055] In addition, for searching for images by text, it mainly relies on the full-text indexing function of the database. By matching the sentence to be retrieved with the data set in the database, a result set is returned. Although this method can achieve retrieval, the returned result set only depends on the full-text retrieval matching result of the database, and does not consider issues such as how to use the diverse text information in the image to optimize the retrieval effect. The accuracy is relatively low when retrieving short words.

[0056] In order to solve the problems existing in the above methods, the present exemplary embodiment proposes a technical solution that can, in the process of text recognition, comprehensively use various aspects of information of the text in the image to perform data cleaning on the recognition results, thereby optimizing the effect of retrieving images based on text. The technical solution of the embodiments of the present disclosure will be elaborated in detail below:

[0057] The present exemplary embodiment first provides a text and image recognition and retrieval method. Refer to Figure 3 As shown, the text and image recognition and retrieval method specifically includes the following steps:

[0058] Step S310: Receive an image and store it in the database, and recognize multiple texts in the image;

[0059] Step S320: Obtain corresponding multiple scoring values according to the text information of each piece of text, and obtain the comprehensive score corresponding to each piece of text based on the above multiple scoring values;

[0060] Step S330: Store each piece of text into the above database based on the above comprehensive score, so as to search for the corresponding image based on the scored text in the above database.

[0061] In the text-image recognition and retrieval method provided by the exemplary embodiment of the present disclosure, on the one hand, after recognizing multiple pieces of text in an image, the text-image recognition and retrieval method provided by this exemplary embodiment will also obtain the comprehensive score corresponding to each piece of text through the text information of each piece of text, making full use of all aspects of the text information, so as to effectively clean the text recognition results. On the other hand, after effectively cleaning the text recognition results, the text recognition results are put into the database, so that the image can also be retrieved in the database through the cleaned text, improving the accuracy of the retrieval.

[0062] Next, in another embodiment, the above steps will be described in more detail.

[0063] In step S310, receive the image and store it in the database, and recognize multiple pieces of text in the image.

[0064] The text-image recognition and retrieval method provided by this exemplary embodiment is used to provide the function of recognizing text in an image and retrieving images through keywords. For example, the text-image recognition and retrieval method can be jointly executed by a terminal device and a server. Specifically, a database can be established in the server for storing text, images, and providing retrieval functions, etc. It should be noted that the above scenario is only an exemplary description, and the protection scope of this exemplary embodiment is not limited thereto.

[0065] In this exemplary embodiment, the above image is any image containing text content. For example, it can be an advertising image containing multiple advertising slogans, or other types of images containing text. This exemplary embodiment does not make special limitations on this. In addition, the above image can be taken by a terminal device and transmitted to the server, or input into the terminal device externally and transmitted to the server by the terminal device, or an image directly transmitted to the server externally. This exemplary embodiment does not make special limitations on this.

[0066] In this exemplary embodiment, after receiving the above-mentioned image, the received image is stored in the database, and multiple pieces of text included in the image are recognized through an optical character recognition algorithm. The above-mentioned optical character recognition algorithm is used to analyze and recognize an image file of a text document to obtain the text and layout information in the image. For example, the optical character recognition algorithm can be an OCR (Optical Character Recognition) algorithm. Through the OCR algorithm, the text in the image can be extracted for further analysis later, so as to give play to the value of the text. For example, retrieving images according to the text content, classifying images according to the text content, etc. It should be noted that the above scenario is only an exemplary illustration, and the protection scope of this exemplary embodiment is not limited thereto.

[0067] In step S320, corresponding multiple score values are obtained based on the text information of each piece of text, and a comprehensive score corresponding to each piece of text is obtained based on the above-mentioned multiple score values.

[0068] In this exemplary embodiment, after the above-mentioned multiple pieces of text in the image are recognized through the optical character recognition algorithm, a comprehensive score corresponding to each piece of text can also be obtained according to the text information of each piece of text. For example, the above-mentioned text information may include text content, position information of the text in the image, area occupied by the text, etc. It should be noted that the above scenario is only an exemplary illustration, and this exemplary embodiment is not limited thereto. For example, according to actual needs, the above-mentioned text information may include more or less information, which all fall within the protection scope of this exemplary embodiment.

[0069] In this exemplary embodiment, the above-mentioned score values can be obtained based on the above-mentioned text information. For example, the above-mentioned score values may include an optical character recognition algorithm score, a text prominence score, and a sentence meaning integrity score. Then the process of obtaining multiple score values based on the text information can be implemented as follows: obtaining the algorithm score of the optical character recognition algorithm, the semantic integrity score of each piece of text, and the prominence score of each piece of text in the image based on the above-mentioned text information. Among them, the above-mentioned semantic integrity score is determined by the ratio of the number of words after word segmentation of each piece of text to the total number of words, and the above-mentioned prominence score is determined by the ratio of each piece of text to the total area of all texts in the image.

[0070] Taking the above-mentioned image as an advertisement picture as an example, the advertisement picture includes multiple advertising slogans. Assuming any one of the advertising slogans is C i , then the above-mentioned semantic integrity score The calculation formula can be as follows:

[0071]

[0072] Among them, The number of words after word segmentation for each line of text, is the total number of words in this advertising slogan, as described above can represent the number of words after word segmentation for each line of text accounting for the total number of words The ratio represents the semantic integrity of the sentence. The higher this value, the higher the integrity of the corresponding sentence and the easier it is to understand.

[0073] In addition, the above-mentioned prominence score can be calculated as follows:

[0074]

[0075] Among them, a i is the area of the region where any line of text is located, is the area occupied by all text regions in the entire picture. Then is the ratio of the area of the region where this line of text is located to the area occupied by all text regions in the entire picture representing the prominence of this line of text in the picture. The larger this value, the more prominent it is in the picture and the higher its importance.

[0076] It should be noted that the above scenario is only an exemplary illustration, and the scope of protection of this exemplary embodiment is not limited thereto.

[0077] In this exemplary embodiment, after obtaining the above-mentioned multiple score values for each piece of text based on the recognized text information, a comprehensive score corresponding to each piece of text can also be obtained based on the above-mentioned multiple score values. Specifically, this process can be achieved as follows: perform a weighted operation on the above algorithm score, semantic integrity score, and prominence score to obtain the comprehensive score corresponding to each piece of text.

[0078] Taking the above advertising picture as an example, this comprehensive score can be obtained through the following formula:

[0079]

[0080] Among them, the above-mentioned is the weighted comprehensive score, is the graphic recognition algorithm score, is the above-mentioned semantic integrity score, For the above-mentioned prominence score. The above-mentioned α, β, and λ are the weighted ratios of the above-mentioned graphic and text recognition algorithm score, semantic integrity score, and prominence score respectively. This weighted ratio can be determined according to the actual situation and empirical values. For example, preferably, the ratio combination can be as follows: the value of α is 0.3, the value of β is 0.1, and the value of λ is 0.6. It should be noted that the above scenario is only an exemplary illustration, and the protection scope of the present exemplary embodiment is not limited thereto.

[0081] In step S330, each piece of text is stored in the above-mentioned database based on the above-mentioned comprehensive score, so as to search for the corresponding image based on the scored text in the above-mentioned database.

[0082] In the present exemplary embodiment, in order to achieve more effective data cleaning, after obtaining the above-mentioned comprehensive score through step S320, the above-mentioned multiple pieces of recognized text can also be correspondingly stored in the above-mentioned database based on the above-mentioned comprehensive score. For example, this process can be implemented as follows: sort each piece of text based on the above-mentioned comprehensive score, divide each piece of text into multiple levels according to the sorting result, filter the text according to the multiple levels, and store the filtered text in the database according to a preset ratio. In addition, after storing the text in the database, an index relationship between the text and the image can also be established, so as to search for the corresponding image based on the scored text in the above-mentioned database.

[0083] Specifically, taking the above-mentioned advertisement picture as an example, the above process can be: after calculating the weighted comprehensive score of all advertising slogans in the above-mentioned advertisement picture sort them from largest to smallest according to the score of the comprehensive score, and divide the multiple recognized advertising slogans into three categories: critical (very important), normal (generally important), and dirty (filterable) according to the sorting result. Among them, the classification rule can be: sort according to the weighted comprehensive score Sort all advertising slogans into two parts according to a certain ratio, and divide the advertising slogans in the previous part of the sorting into the critical category. This ratio can be determined according to actual needs and empirical values. Preferably, the value of this ratio can be 1:1; then, divide the advertising slogans in the latter part according to a preset fixed weighted score S'. The advertising slogans greater than S' are divided into the normal category, and the advertising slogans less than S' are divided into the dirty category. Among them, the value of S' can be determined according to actual needs and empirical values. Preferably, the value of this preset fixed weighted score can be 0.15; finally, filter out the advertising slogans in the above-mentioned dirty category, and store the critical category and the normal category in the above-mentioned database by category. In addition, an index relationship between the advertising slogan and the image can also be established in the database. It should be noted that the above scenario is only an exemplary illustration, and the protection scope of the present exemplary embodiment is not limited thereto.

[0084] In the present exemplary embodiment, the process of searching for corresponding images based on the scored text in the above database can be implemented as follows: receiving a retrieval keyword, matching the scored text based on the retrieval keyword to obtain multiple secondary scoring results, and obtaining a secondary comprehensive score based on the multiple secondary scoring results; sorting the secondary comprehensive scores, and returning a result set based on the sorting result, where the result set includes multiple pictures. Specifically, the matching of the scored text based on the retrieval keyword to obtain multiple secondary scoring results can be: respectively matching the retrieval keyword from texts of different levels in the above database, and obtaining multiple secondary scoring results.

[0085] Taking the above advertising picture as an example, the above process can be implemented as follows: performing secondary scoring based on the advertising slogans (the dirty category has been filtered during data cleaning) in the divided critical and normal categories and the scores of the full-text retrieval in the database, obtaining the above multiple secondary scoring results and sorting them, and returning a result set based on the sorting result. Specifically, assuming that for the retrieval keyword K, when the database performs full-text retrieval, the secondary scoring result for retrieving a certain image from the critical category is Sc, and the secondary scoring result for the retrieval result from the normal category is S n , then the above secondary comprehensive score can be calculated by the following formula:

[0086]

[0087] where α + β = 1, and the values of α and β can be determined according to the actual situation and experience. For example, preferably, the value of α can be 0.71, and the value of β can be 0.29. It should be noted that the above scenario is only an exemplary illustration, and the protection scope of the present exemplary embodiment is not limited thereto.

[0088] Since the above retrieval process re-sorts according to the scores of the secondary comprehensive scores of all results in the result set after calculating them, taking into account the semantic integrity of the text in the image and the prominence of the text in the image, the results obtained by the retrieval method are more accurate.

[0089] Next, taking the specific application scenario of graphic and text recognition retrieval of advertising pictures as an example, combined with Figures 4 to 6 , the above graphic and text recognition retrieval method will be fully elaborated. Among them, Figure 4 is the architecture diagram of the advertising picture text recognition retrieval system. As Figure 4 shown, the system architecture of the recognition retrieval system includes an external interaction layer 410, a recognition service module 420, a data cleaning module 430, a database 440, a result set optimization module 450, and a retrieval service module 460. Among them:

[0090] The above external interaction layer is used to submit pictures to the system directly or indirectly, and receive the result set retrieved by the system and send it to the caller.

[0091] The above recognition service module is used to recognize the text (multiple advertising slogans) in the pictures obtained through the above external interaction layer. Specifically, the advertising slogans in the pictures can be recognized through the optical character recognition (OCR) algorithm.

[0092] The above data cleaning module is used to clean the recognized advertising slogans. Specifically, the data cleaning process can be implemented by executing the Figure 5 process shown. As Figure 5 shown, this process includes the following steps:

[0093] Step S510: Obtain the scores recognized by the OCR algorithm.

[0094] In this step, obtain the score recognized by the OCR algorithm for any one of the above recognized advertising slogans C i Repeat this step to obtain the scores recognized by the OCR algorithm for all the recognized advertising slogans.

[0095] Step S520: Segment the sentences (advertising slogans) and calculate the proportion of the segmented words in the sentences.

[0096] In this step, segment each of the above advertising slogans and obtain the semantic integrity score by calculating the proportion of the segmented words in the sentence. Among them, the proportion of the segmented words in the sentence can be calculated by the following formula:

[0097]

[0098] Among them, is the number of words after segmenting each line of text, is the total number of words in this advertising slogan, and the above can represent the number of words after segmenting each line of text accounting for the proportion of the total number of words represents the semantic integrity of the sentence. The higher this value is, the higher the integrity of the corresponding sentence is and the easier it is to be understood.

[0099] Step S530: Calculate the proportion of the area of each advertising slogan in the picture to the area of all the text in the picture.

[0100] In this step, obtain the prominence score of each advertising slogan in the picture by calculating the proportion of the area of each advertising slogan in the picture to the area of all the text in the picture ​​It can be calculated by the following formula:

[0101]

[0102] where a i is the area of the region where any line of text is located, is the area occupied by all text regions in the entire image, then is the proportion of the area of the region where this line of text is located to the area of all text regions in the entire image which represents the prominence of this line of text in the image. The larger this value, the more prominent it is in the image and the higher its importance.

[0103] Step S540: Calculate the weighted score.

[0104] In this step, calculate the weighted score of the score recognized by the above OCR algorithm, the semantic integrity score, and the prominence score. This weighted score can be calculated by the following formula:

[0105]

[0106] where the above is the weighted comprehensive score, is the score of the graphic and text recognition algorithm, is the above semantic integrity score, is the above prominence score. The above α, β, and λ are the weighted proportions of the score of the graphic and text recognition algorithm, the semantic integrity score, and the prominence score respectively. This weighted proportion can be determined according to the actual situation and empirical values. For example, preferably, the proportion combination can be as follows: the value of α is 0.3, the value of β is 0.1, and the value of λ is 0.6.

[0107] Step S550: Divide all advertising slogans into three grades: critical, normal, and dirty according to the weighted score.

[0108] In this step, sort the scores of the above weighted scores in descending order, and divide all the recognized advertising slogans into three grades: critical (very important), normal (generally important), and dirty (filterable) according to the sorting result. Among them, the classification rule can be: according to the weighted comprehensive score Sort the advertising slogans and divide all of them into two parts according to a certain ratio. The advertising slogans in the first part are classified as critical. The above ratio can be determined according to actual needs and empirical values. Preferably, the value of this ratio can be 1:1. Then, split the advertising slogans in the second part according to a preset fixed weighted score S'. The advertising slogans greater than S' are classified as normal, and the advertising slogans less than S' are classified as dirty. Among them, the value of S' can be determined according to actual needs and empirical values. Preferably, the value of this preset fixed weighted score can be 0.15.

[0109] Step S560: Store the advertising slogans after the above data cleaning in the database.

[0110] In this step, filter out the advertising slogans of the above dirty category, and store the critical category and the normal category in the above database by category.

[0111] The above database is used to store the received pictures and the advertising slogans processed by the data cleaning module. In addition, an index relationship between the advertising slogans and the images can be established in this database.

[0112] The above result set optimization module is used to comprehensively score and sort the critical and normal category sentences divided in the data cleaning process and the scores given by the full-text retrieval of the database, and optimize the retrieved result set based on the sorting result. Most of the sentences of the dirty category obtained in the data cleaning process are advertising slogans with recognition errors or too small text to be noticed, so they are not considered during retrieval. Specifically, this result optimization module can implement this optimization process by executing the process as Figure 6 shown:

[0113] Step S610: Receive the retrieval keyword.

[0114] Step S620: Retrieve and match in the database from the advertising slogans of the critical and normal categories.

[0115] Step S630: Obtain the critical matching score and the normal matching score.

[0116] In this step, obtain the matching score Sc for retrieving a certain image from the critical category and the matching score S n .

[0117] Step S640: Calculate the weighted score.

[0118] In this step, calculate the weighted scores of the above critical matching scores and normal matching scores, and the weighted scores can be calculated by the following formula:

[0119]

[0120] Where α + β = 1, and the values of α and β can be determined according to the actual situation and experience. For example, preferably, the value of α can be 0.71 and the value of β can be 0.29.

[0121] Step S650: Sort according to the weighted scores and return the results.

[0122] In this step, sort the above weighted scores and return the retrieved result set based on the sorting result.

[0123] The above retrieval service module is used to provide retrieval services. For example, a result set can be retrieved through the key retrieval terms.

[0124] It should be noted that although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0125] Furthermore, in the present exemplary embodiment, an image and text recognition and retrieval device is further provided. Referring to Figure 7 as shown, the image and text recognition and retrieval device 700 may include a text recognition module 710, a data cleaning module 720, and an identification and retrieval module 730. Among them:

[0126] The text recognition module 710 can be used to receive an image and store it in the database, and recognize multiple pieces of text in the image;

[0127] The data cleaning module 720 can be used to obtain corresponding multiple score values based on the text information of each piece of text, and obtain the comprehensive score corresponding to each piece of text based on the multiple score values;

[0128] The identification and retrieval module 730 can be used to store each piece of text in the database based on the comprehensive score, so as to search for the corresponding image in the database based on the scored text.

[0129] In the present exemplary embodiment, the above-mentioned text recognition module includes a receiving unit and a text recognition unit. Among them, the above-mentioned receiving unit is used to receive an image. For example, an image can be obtained directly or indirectly through external interaction; the above-mentioned text recognition unit is used to recognize multiple texts in the image through a graphic and text recognition algorithm. For example, all texts in the image can be recognized through an optical character recognition algorithm.

[0130] In the present exemplary embodiment, the above-mentioned data cleaning module can clean the recognized text by performing the following method: obtaining the score of the graphic and text recognition algorithm, the semantic integrity score of each text, and the salience score of each text in the image based on the text information; wherein, the semantic integrity score is determined by the ratio of the number of words after word segmentation of each text to the total number of words, and the salience score is determined by the ratio of each text to the total area of all texts in the image; performing a weighted operation on the above algorithm score, semantic integrity score, and salience score to obtain a comprehensive score corresponding to each text.

[0131] In the present exemplary embodiment, the above-mentioned recognition and retrieval module may include a warehousing unit and a retrieval unit. The above-mentioned warehousing unit is used to sort each text based on the above comprehensive score, divide each text into multiple levels according to the sorting result, filter the texts according to multiple levels, and store the filtered texts in the database according to a preset ratio. The above-mentioned retrieval unit is used to receive a retrieval keyword, match the scored text based on the retrieval keyword to obtain multiple secondary scoring results, and obtain a secondary comprehensive score based on the multiple secondary scoring results; sort the secondary comprehensive score, and return a result set based on the sorting result, and the result set includes multiple retrieved pictures.

[0132] The specific details of each module or unit in the above-mentioned graphic and text recognition and retrieval device have been described in detail in the corresponding graphic and text recognition and retrieval method, so they will not be elaborated here.

[0133] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0134] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the methods described in the following embodiments. For example, the electronic device can implement the steps such as Figure 3 , Figure 5 or Figure 6 shown.

[0135] It should be noted that the computer-readable medium shown in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0136] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A method for graphic and text recognition and retrieval, characterized in that, including: Receiving an image and storing it in a database, and recognizing multiple pieces of text in the image; Obtaining corresponding multiple score values according to the text information of each piece of text, and obtaining a comprehensive score corresponding to each piece of text based on the multiple score values, where the multiple score values include an algorithm score for optical character recognition, a text prominence score, and a sentence integrity score; Storing each piece of text in the database based on the comprehensive score, so as to search for the corresponding image based on the scored text in the database.

2. The graphic and text recognition and retrieval method according to claim 1, wherein The recognizing multiple pieces of text in the image includes: Extracting the multiple pieces of text in the image based on an optical character recognition algorithm, and obtaining the text information of each piece of text.

3. The graphic recognition and retrieval method according to claim 1, wherein The obtaining corresponding multiple score values according to the text information of each piece of text includes: Obtaining an algorithm score of an algorithm for recognizing the corresponding text in the image, a semantic integrity score of each piece of text, and a prominence score of each piece of text in the image based on the text information; Among them, the semantic integrity score is determined by the ratio of the number of words after word segmentation of each piece of text to the total number of words, and the prominence score is determined by the ratio of each piece of text to the total area of all text in the image.

4. The graphic recognition and retrieval method according to claim 3, wherein The obtaining a comprehensive score corresponding to each piece of text based on the multiple score values includes: Performing a weighted operation on the algorithm score, the semantic integrity score, and the prominence score to obtain the comprehensive score corresponding to each piece of text.

5. The graphic recognition and retrieval method according to claim 1, wherein The storing each piece of text in the database based on the comprehensive score includes: Sorting each piece of text based on the comprehensive score, classifying each piece of text into multiple levels according to the sorting result, filtering the text according to the multiple levels, and storing the filtered text in the database according to a preset ratio.

6. The graphic recognition and retrieval method according to claim 1, characterized in that The searching for the corresponding image based on the scored text in the database includes: Receiving a retrieval keyword, matching the scored text based on the retrieval keyword to obtain multiple secondary score results, and obtaining a secondary comprehensive score based on the multiple secondary score results; Sorting the secondary comprehensive score, and returning a result set based on the sorting result, where the result set includes multiple images.

7. The graphic recognition and retrieval method according to claim 5, characterized in that The searching for the corresponding image based on the scored text in the database includes: Receiving a retrieval keyword, respectively matching the retrieval keyword from texts of different levels in the database to obtain multiple secondary score results, and obtaining a secondary comprehensive score based on the secondary score results; Sorting the secondary comprehensive score, and returning a result set based on the sorting result, where the result set includes multiple images.

8. An image and text recognition and retrieval device, characterized in that, including: A text recognition module, configured to receive an image and store it in a database, and recognize multiple pieces of text in the image; A data cleaning module, configured to obtain corresponding multiple score values according to the text information of each piece of text, and obtain a comprehensive score corresponding to each piece of text based on the multiple score values, where the multiple score values include an algorithm score for optical character recognition, a text prominence score, and a sentence integrity score; An identification and retrieval module, configured to store each of the texts into the database based on the comprehensive score, so as to search for the corresponding image based on the scored texts in the database.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the text and image identification and retrieval method according to any one of claims 1-7.

10. An electronic device, characterized in that, Comprising: A processor; A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the text and image identification and retrieval method according to any one of claims 1-7 by executing the executable instructions.

Citation Information

Patent Citations

  • Result re-ranking for object recognition

    US10769200B1