Image search method, device, equipment and computer program product
By combining low-dimensional and high-dimensional OCR recognition processing, the problems of time-consuming image search and difficult refined recognition in the existing technology are solved, and efficient and accurate image search is achieved.
Patent Information
- Application Number
- CN202011248141.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-10
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2040-11-10
AI Technical Summary
In the existing technology, OCR-based image search takes too long to process a large number of images, cannot accurately identify the text information in the images, and cloud-based recognition has the risk of upload and download failures and privacy leakage.
Combining low-dimensional OCR recognition processing and high-dimensional OCR recognition processing, low-dimensional OCR recognition processing quickly identifies easily recognizable content, and high-dimensional OCR recognition processing accurately identifies detailed content, achieving refined search.
It improves the efficiency and accuracy of image search, can more accurately identify and search text information in images, and avoids the risks of cloud-based recognition.
Smart Images

Figure CN112347948B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of Internet technology, and are related to, but not limited to, an image search method, apparatus, device, and computer program product. Background Art
[0002] In related technologies, image search based on optical character recognition (OCR) relies on OCR to identify text within an image before performing a search. When there are a large number of images, users may need to wait for OCR to complete all of the images before receiving search results, which is relatively time-consuming. Furthermore, if more detailed textual information within an image needs to be searched, such as notes or store names within the image, the existing search function cannot provide this refined search capability. Summary of the Invention
[0003] The present invention provides an image search method, apparatus, device, and computer program product, all of which relate to the field of artificial intelligence technology. By combining the recognition results of low-dimensional and high-dimensional OCR processes to perform image search, the method can more accurately search for text information in images, achieve refined search results, and improve search efficiency.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] This embodiment of the present application provides an image search method, including:
[0006] Obtaining an image search request, wherein the image search request includes keywords;
[0007] In response to the image search request, obtaining an OCR recognition result for each image in a preset image library; wherein the OCR recognition result includes at least one of the following: a low-dimensional OCR recognition result obtained by using a low-dimensional OCR recognition process based on an OCR recognition threshold, and a high-dimensional OCR recognition result obtained by using a high-dimensional OCR recognition process based on depth recognition, wherein the recognition accuracy of the low-dimensional OCR recognition process is less than the recognition accuracy of the high-dimensional OCR recognition process;
[0008] Traversing the pictures that have not completed the low-dimensional OCR recognition process and the high-dimensional OCR recognition process, and performing the low-dimensional OCR recognition process on each traversed picture to obtain a low-dimensional OCR recognition result for each corresponding picture;
[0009] Determining a target image matching the keyword in the preset image library according to the low-dimensional OCR recognition result or the high-dimensional OCR recognition result of each image;
[0010] The target image is determined as the search result of the image search request, and the search result is displayed.
[0011] In some embodiments, the method further comprises:
[0012] Determining the credibility of the low-dimensional OCR recognition result of each of the images;
[0013] Delete low-dimensional OCR recognition results whose confidence level is lower than the threshold.
[0014] The present invention provides an image search device, comprising:
[0015] An acquisition module, configured to acquire an image search request, wherein the image search request includes keywords;
[0016] a response module, configured to obtain, in response to the image search request, an OCR recognition result for each image in a preset image library; wherein the OCR recognition result includes at least one of the following: a low-dimensional OCR recognition result obtained by low-dimensional OCR recognition processing based on an OCR recognition threshold, and a high-dimensional OCR recognition result obtained by high-dimensional OCR recognition processing based on depth recognition, wherein the recognition accuracy of the low-dimensional OCR recognition processing is less than the recognition accuracy of the high-dimensional OCR recognition processing;
[0017] a processing module, configured to traverse images that have not completed the low-dimensional OCR recognition process and the high-dimensional OCR recognition process, and perform the low-dimensional OCR recognition process on each traversed image to obtain a low-dimensional OCR recognition result for each corresponding image;
[0018] A first determining module is configured to determine a target image matching the keyword in the preset image library based on the low-dimensional OCR recognition result or the high-dimensional OCR recognition result of each image;
[0019] The second determining module is configured to determine the target image as a search result of the image search request and display the search result.
[0020] An embodiment of the present application provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; wherein a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor is used to execute the computer instructions to implement the above-mentioned image search method.
[0021] The present invention provides an image search device, including:
[0022] The memory is used to store executable instructions; the processor is used to implement the above-mentioned image search method when executing the executable instructions stored in the memory.
[0023] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a processor to execute the executable instructions to implement the above-mentioned image search method.
[0024] The embodiments of the present application have the following beneficial effects: low-dimensional OCR recognition processing based on OCR recognition threshold and high-dimensional OCR recognition processing based on depth recognition are used to process images in a preset image library, and low-dimensional OCR recognition results and high-dimensional OCR recognition results are obtained accordingly. Based on the low-dimensional OCR recognition results or high-dimensional OCR recognition results of each image, the target image of the image search request is matched. In this way, since the image search is performed by combining the recognition results of the low-dimensional OCR recognition processing and the high-dimensional OCR recognition processing, the text information in the image can be searched more accurately, a refined search can be achieved, accurate search results can be obtained, and the search efficiency can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1A It is a schematic diagram of the image search process in the related art;
[0026] Figure 1B This is a schematic diagram of a search scenario in which the image to be searched is a note in the related art;
[0027] Figure 2 This is an optional architectural diagram of the image search system provided in an embodiment of the present application;
[0028] Figure 3 is a structural diagram of an electronic device provided in an embodiment of the present application;
[0029] Figure 4 This is an optional flowchart of the image search method provided in the embodiment of the present application;
[0030] Figure 5 This is an optional flowchart of the image search method provided in the embodiment of the present application;
[0031] Figure 6 This is an optional flowchart of the image search method provided in the embodiment of the present application;
[0032] Figure 7 This is an optional flowchart of the high-dimensional OCR recognition process provided by the embodiment of the present application;
[0033] Figure 8 This is an optional flowchart of the image search method provided in the embodiment of the present application;
[0034] Figure 9 This is a flowchart of the image search method provided in the embodiment of the present application;
[0035] Figure 10 This is a detailed flowchart of the image search method provided in the embodiment of the present application;
[0036] Figure 11 This is a schematic diagram of a simplified identification process provided in an embodiment of the present application;
[0037] Figure 12 This is a schematic diagram of a simplified identification process in an embodiment of the present application;
[0038] Figure 13 This is a flowchart of the depth recognition strategy provided by the embodiment of the present application;
[0039] Figure 14A This is a schematic diagram of an image before being divided into equal parts and enlarged, provided in an embodiment of the present application;
[0040] Figure 14B This is a schematic diagram of an image divided into equal parts and magnified according to an embodiment of the present application. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0042] In the following description, reference is made to "some embodiments," which describe a subset of all possible embodiments. However, it will be understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict. Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by those skilled in the art to which the embodiments of this application pertain. The terms used in the embodiments of this application are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0043] In order to better understand the image search method provided in the embodiments of the present application, the image search method in the related art is first described:
[0044] In related technologies, when searching for images (for example, searching for emoticon images), you can search by some keywords on the image, such as Figure 1AAs shown in the figure, it is a schematic diagram of the process of image search in the related art. When a user wants to search for an emoji image corresponding to "dejected", they can input "dejected", and the system will automatically match and output Image 11.
[0045] However, when a user wants to search for the detailed text information captured in some images, the search method in the related art cannot support such fine recognition. For example, when a user wants to search for the text information in a note, as Figure 1B As shown in the figure, it is a schematic diagram of a search scenario where the image to be searched is a note in the related art. The existing search function cannot support it because the threshold of OCR recognition during search is relatively low, and it is set to prioritize the recognition of easier-to-recognize content. For example, it prioritizes the recognition of large fonts 12, etc., so as to ensure that the recognition time during search will not be too high. Therefore, users cannot search for detailed text information 13 through the search function, so it is necessary to optimize the search function in this case. In addition, some solutions use the method of first uploading the image to the server for OCR recognition, and then synchronizing the results to the client after the recognition is completed. This solution can have a relatively high accuracy, but there is a certain risk of failure in requests such as server upload and download. Uploading large and long images will consume a lot of time and traffic, and at the same time, the privacy of user images is also difficult to guarantee.
[0046] To sum up, the solutions in the related art have at least the following problems: high-precision recognition takes too much time and cannot be applied to the search scenario; low-precision recognition results in the inability to find detailed content during search; cloud recognition faces risks of upload and download failure, upload and download time consumption, privacy leakage, and offline inavailability.
[0047] To solve at least one of the above problems existing in the image search method in the related art, the embodiment of the present application proposes an image search method, which combines fast recognition with precise recognition of secondary processing of images to optimize the efficiency and accuracy of OCR recognition for text information on images during the search process, making the results faster and more accurate.
[0048] The image search method provided in an embodiment of the present application first obtains an image search request, which includes keywords; then, in response to the image search request, obtains the OCR recognition result of each image in a preset image library; wherein the OCR recognition result includes at least one of the following: a low-dimensional OCR recognition result obtained by low-dimensional OCR recognition processing based on an OCR recognition threshold, or a high-dimensional OCR recognition result obtained by high-dimensional OCR recognition processing based on depth recognition, wherein the recognition accuracy of the low-dimensional OCR recognition processing is less than the recognition accuracy of the high-dimensional OCR recognition processing; then, traverses images that have not completed low-dimensional OCR recognition processing and high-dimensional OCR recognition processing, and performs low-dimensional OCR recognition processing on each traversed image to obtain a low-dimensional OCR recognition result of each corresponding image; based on the low-dimensional OCR recognition result or high-dimensional OCR recognition result of each image, determines a target image that matches the keyword in the preset image library; finally, determines the target image as the search result of the image search request, and displays the search result. In this way, since the recognition results of low-dimensional OCR recognition processing and high-dimensional OCR recognition processing are combined for image search, the text information in the image can be searched more accurately, a refined search can be achieved, accurate search results can be obtained, and search efficiency can be improved.
[0049] The following describes an exemplary application of the image search device of an embodiment of the present application. In one implementation, the image search device provided by the embodiment of the present application can be implemented as any terminal or electronic device that can play audio, such as a laptop computer, a tablet computer, a desktop computer, a mobile device (for example, a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device), an intelligent robot, etc. In another implementation, the image search device provided by the embodiment of the present application can also be implemented as a server. The following describes an exemplary application of the image search device when it is implemented as an electronic device, and a client on the electronic device can be used to perform image search.
[0050] See also Figure 2 , Figure 2: This is an optional architectural diagram of the image search system 10 provided in an embodiment of the present application. In order to accurately respond to image search requests and search for accurate target images, the image search system 10 provided in an embodiment of the present application includes a terminal 100 (i.e., an electronic device), a network 200, and a server 300. Among them, an image search application is running on the terminal 100, and the image search application corresponds to a preset image library 400. The preset image library 400 stores multiple images. The user can enter keywords through the client of the image search application to form an image search request. The client responds to the user's image search request to match at least one target image in the preset image library. In the embodiment of the present application, the client can also perform low-dimensional OCR recognition processing based on OCR recognition threshold on each image in the preset image library 400. The server 300 serves as a background server, which is used to perform high-dimensional OCR recognition processing based on depth recognition on each image in the preset image library 400 during idle time to obtain high-dimensional OCR recognition results and send the high-dimensional OCR recognition results to the terminal 100.
[0051] In an embodiment of the present application, when responding to an image search request, the terminal 100 obtains the OCR recognition result of each image in the preset image library in response to the image search request; wherein the OCR recognition result includes at least one of the following: a low-dimensional OCR recognition result obtained by low-dimensional OCR recognition processing based on an OCR recognition threshold, or a high-dimensional OCR recognition result obtained by high-dimensional OCR recognition processing based on depth recognition; traversing images that have not completed low-dimensional OCR recognition processing and high-dimensional OCR recognition processing, and performing low-dimensional OCR recognition processing on each traversed image to obtain a low-dimensional OCR recognition result for each corresponding image; determining a target image that matches the keyword in the preset image library based on the low-dimensional OCR recognition result or high-dimensional OCR recognition result of each image; determining the search result as the image search request, and displaying the search result on the current interface 100-1 of the terminal 100.
[0052] The image search method provided in the embodiments of the present application also relates to the field of artificial intelligence technology, and can be implemented at least through computer vision technology and machine learning technology in artificial intelligence technology. Among them, computer vision technology (CV) is a science that studies how to make machines "see". More specifically, it refers to machine vision such as using cameras and computers to replace human eyes to identify and measure targets, and further performing image processing so that the computer processing becomes an image more suitable for human observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to establish an artificial intelligence system that can obtain information from images or multidimensional data. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition. Machine learning (ML) is a multi-disciplinary interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning generally include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching. In the embodiments of the present application, machine learning technology is used to respond to network structure search requests to automatically search for the target network structure, as well as to train and optimize the controller and score model.
[0053] Figure 3 is a structural diagram of an electronic device 300 provided in an embodiment of the present application, Figure 3 The electronic device 300 shown includes: at least one processor 310, a memory 350, at least one network interface 320 and a user interface 330. The various components in the electronic device 300 are coupled together via a bus system 340. It is understood that the bus system 340 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 340 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 340 is not described in detail. Figure 3 Various buses are labeled as bus system 340 .
[0054] The processor 310 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0055] The user interface 330 includes one or more output devices 331 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 330 also includes one or more input devices 332, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0056] The memory 350 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, and the like. The memory 350 may optionally include one or more storage devices physically located away from the processor 310. The memory 350 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 350 described in the embodiments of the present application is intended to include any suitable type of memory. In some embodiments, the memory 350 is capable of storing data to support various operations, examples of which include programs, modules, and data structures, or subsets or supersets thereof, as exemplified below.
[0057] Operating system 351, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0058] A network communication module 352 for reaching other computing devices via one or more (wired or wireless) network interfaces 320 , exemplary network interfaces 320 including Bluetooth, WiFi, and USB;
[0059] The input processing module 353 is configured to detect one or more user inputs or interactions from one of the one or more input devices 332 and to translate the detected inputs or interactions.
[0060] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 3 The image search device 354 stored in the memory 350 is shown. The image search device 354 can be an image search device in the electronic device 300. The image search device 354 can be software in the form of a program or plug-in, and includes the following software modules: an acquisition module 3541, a response module 3542, a processing module 3543, a first determination module 3544, and a second determination module 3545. These modules are logical and can be arbitrarily combined or further separated according to the functions implemented. The functions of each module will be described below.
[0061] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the image search method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.
[0062] The following will describe the image search method provided by the embodiment of the present application in conjunction with the exemplary application and implementation of the electronic device 300 provided by the embodiment of the present application. Figure 4 , Figure 4 This is an optional flow chart of the image search method provided in the embodiment of the present application, which will be combined with Figure 4 The steps shown are explained.
[0063] Step S401: Obtain an image search request, where the image search request includes keywords.
[0064] Here, an image search application is running on the electronic device. A user can enter a keyword on the client of the image search application. The client then generates an image search request based on the user's input or the user's click on "search." This request requests the client to search for images corresponding to the keyword. The keyword can be the type of image, text within the image, or a summary of the text within the image.
[0065] In an embodiment of the present application, when the client performs an image search in response to an image search request, the search may be performed in an online state or an offline state.
[0066] Step S402: In response to the image search request, obtain the OCR recognition result of each image in the preset image library.
[0067] Here, the OCR recognition result includes at least one of the following: a low-dimensional OCR recognition result obtained by low-dimensional OCR recognition processing based on an OCR recognition threshold, or a high-dimensional OCR recognition result obtained by high-dimensional OCR recognition processing based on deep recognition. Among them, low-dimensional OCR recognition processing is a simple recognition strategy, and high-dimensional OCR recognition processing is a deep recognition strategy; low-dimensional OCR recognition processing performs simple recognition on images, while high-dimensional OCR recognition processing performs more detailed and accurate recognition on images, and the recognition accuracy of low-dimensional OCR recognition processing is lower than that of high-dimensional OCR recognition processing; low-dimensional OCR recognition processing is less difficult, has a higher recognition rate, and has lower resource consumption, while high-dimensional OCR recognition processing is more difficult, has a lower recognition rate, and has greater resource consumption.
[0068] The OCR recognition threshold is a value that is relatively balanced between recognition accuracy and recognition time. That is to say, when the OCR recognition threshold is met, not only is the recognition speed fast, but the recognition fault tolerance rate is also high. For example, the OCR recognition threshold may include a threshold corresponding to the font size or a threshold corresponding to the recognition credibility. That is, when recognizing text of a certain font size, not only the recognition accuracy but also the recognition efficiency can be guaranteed, then the font size value corresponding to the font size can be the OCR recognition threshold. Alternatively, when recognizing text in an image, when a certain credibility is reached, the recognition accuracy is high and the recognition efficiency is also high, then the credibility can be determined as the OCR recognition threshold.
[0069] In the embodiment of the present application, the low-dimensional OCR recognition processing is performed based on the OCR recognition threshold, that is, when the low-dimensional OCR recognition processing is performed on the image, the recognition parameters meet the OCR recognition threshold. For example, when the OCR recognition threshold includes a font size threshold, then during the low-dimensional OCR recognition processing, only the text in the image with a font size greater than the font size threshold is recognized, and the text with a font size smaller than the font size threshold is not recognized. That is, when the low-dimensional OCR recognition processing is performed on the image, if the image is a picture with more detailed text, such as a note, OCR recognition will not be performed on all the text in the image, and only OCR recognition will be performed on some of the text in the image that is easy to recognize, which can improve the recognition efficiency.
[0070] Deep recognition is a precise recognition method that also identifies detailed content within an image. During deep recognition, not only the overall content is recognized, but also the detailed text within the image is recognized. High-dimensional OCR processing based on deep recognition can recognize every word in the image, resulting in higher accuracy but also more time-consuming recognition.
[0071] In some embodiments, after performing low-dimensional OCR recognition processing on the image, a low-dimensional OCR recognition result is obtained, and after performing high-dimensional OCR recognition processing on the image, a high-dimensional OCR recognition result is obtained. After obtaining the low-dimensional OCR recognition result or the high-dimensional OCR recognition result, the corresponding low-dimensional OCR recognition result or the high-dimensional OCR recognition result, as well as the mapping relationship between the low-dimensional OCR recognition result and the image, and the mapping relationship between the high-dimensional OCR recognition result and the image, are stored in a preset storage unit.
[0072] Step S403 , traverse the pictures that have not completed the low-dimensional OCR recognition processing and the high-dimensional OCR recognition processing, and perform low-dimensional OCR recognition processing on each traversed picture to obtain a low-dimensional OCR recognition result for each corresponding picture.
[0073] Here, a check is performed on each image in the preset image library to determine whether each image has undergone low-dimensional OCR recognition processing and high-dimensional OCR recognition processing, thereby completing the traversal of the images in the preset image library. For example, the determination of whether each image has undergone low-dimensional OCR recognition processing or high-dimensional OCR recognition processing can be performed by searching a preset storage unit to determine whether a low-dimensional OCR recognition result or a high-dimensional OCR recognition result for each image is stored.
[0074] In an embodiment of the present application, for images that have not yet undergone low-dimensional OCR recognition processing and high-dimensional OCR recognition processing at the current moment, low-dimensional OCR recognition processing is performed on these images. Since the recognition efficiency of low-dimensional OCR recognition processing is relatively high, the efficiency of image recognition can be improved during this image search process, thereby improving the efficiency of image search.
[0075] It should be noted that, after performing low-dimensional OCR recognition processing on any image at the current moment, the low-dimensional OCR recognition result of the image may be correspondingly stored in a preset storage unit.
[0076] Step S404 : determining a target image matching the keyword in a preset image library according to the low-dimensional OCR recognition result or the high-dimensional OCR recognition result of each image.
[0077] In the embodiment of the present application, when matching a target image, matching can be performed not only based on the image's low-dimensional OCR recognition results, but also based on the image's high-dimensional OCR recognition results. If the image has a high-dimensional OCR recognition result, matching based on the high-dimensional OCR recognition result is preferred because the high-dimensional OCR recognition result has more recognition content and higher recognition accuracy than the low-dimensional OCR recognition result. If the image only has a low-dimensional OCR recognition result, matching is performed based on the low-dimensional OCR recognition result.
[0078] In some embodiments, when matching the target image, the keyword can be matched with the corresponding text content in the low-dimensional OCR recognition result or the high-dimensional OCR recognition result, the similarity between the text content corresponding to the low-dimensional OCR recognition result or the high-dimensional OCR recognition result and the keyword is determined, and the image with the highest similarity is determined as the target image, or, after determining the similarity between each image and the keyword, the images are sorted in descending order of similarity to form an image sequence, and then a specific number of images in the image sequence are selected as target images.
[0079] In other embodiments, when matching target images, the image keywords corresponding to each image can be determined based on the text content corresponding to the low-dimensional OCR recognition results or the high-dimensional OCR recognition results, and then the keywords in the image search request are matched with the image keywords of each image, and the images corresponding to the same image keywords or similar image keywords as the keywords in the image search request are determined as target images.
[0080] Step S405: determine the target image as the search result of the image search request, and display the search result.
[0081] In an embodiment of the present application, when the determined target picture is one, this picture is displayed on the current interface of the electronic device. When the determined target pictures are multiple, multiple pictures are displayed on the current interface of the electronic device at the same time, or multiple pictures are displayed in pages.
[0082] The image search method provided in the embodiment of the present application uses low-dimensional OCR recognition processing based on OCR recognition threshold and high-dimensional OCR recognition processing based on depth recognition to process images in a preset image library, and obtains low-dimensional OCR recognition results and high-dimensional OCR recognition results accordingly. Based on the low-dimensional OCR recognition results or high-dimensional OCR recognition results of each image, the target image of the image search request is matched. In this way, since the image search is performed by combining the recognition results of the low-dimensional OCR recognition processing and the high-dimensional OCR recognition processing, it is possible to more accurately search for text information in the image, realize refined search, obtain accurate search results, and improve search efficiency.
[0083] In some embodiments, low-dimensional OCR recognition processing can be performed in different ways. Figure 5 This is an optional flow chart of the image search method provided in the embodiment of the present application. Figure 5 As shown, the method includes the following steps:
[0084] Step S501: Obtain an image search request, where the image search request includes keywords.
[0085] Step S502: In response to the image search request, obtain the OCR recognition result of each image in the preset image library; wherein the OCR recognition result includes at least one of the following: a low-dimensional OCR recognition result obtained by low-dimensional OCR recognition processing based on an OCR recognition threshold, and a high-dimensional OCR recognition result obtained by high-dimensional OCR recognition processing based on depth recognition, and the recognition accuracy of the low-dimensional OCR recognition processing is less than the recognition accuracy of the high-dimensional OCR recognition processing.
[0086] Step S503 : When it is determined that there is at least one picture in the preset picture library that has not completed the low-dimensional OCR recognition process and the high-dimensional OCR recognition process, traverse the pictures that have not completed the low-dimensional OCR recognition process and the high-dimensional OCR recognition process.
[0087] It should be noted that steps S501 to S503 are the same as the above-mentioned steps S401 to S403, and will not be repeated in this embodiment of the application.
[0088] In some embodiments, the OCR recognition threshold includes a recognition speed threshold. Correspondingly, low-dimensional OCR recognition processing can be performed through the following steps:
[0089] Step S504: determining the recognition speed for each character in the image.
[0090] Here, the recognition speed for each character is the ratio of the time required to recognize a certain character to the number of characters in the character to be recognized. A higher recognition speed indicates that the corresponding character is easier to recognize, while a lower recognition speed indicates that the corresponding character is more difficult to recognize.
[0091] In the embodiment of the present application, the recognition speed for each type of text can be determined in advance based on the OCR recognition situation, thereby determining a suitable recognition speed threshold.
[0092] Step S505 , performing OCR recognition on the text whose recognition speed is greater than the recognition speed threshold, so as to complete the low-dimensional OCR recognition processing of the image.
[0093] Here, the characters whose recognition speed is greater than the recognition speed threshold are relatively easy to recognize characters. Only the relatively easy to recognize characters can be subjected to OCR recognition to complete the low-dimensional OCR recognition processing of the image.
[0094] In some embodiments, the OCR recognition threshold includes a font size threshold. Correspondingly, low-dimensional OCR recognition processing can be performed by the following steps:
[0095] Step S506: Determine the font size of each character in the image.
[0096] Here, text with a larger font size is relatively easier to recognize, and thus the recognition speed is higher, while text with a smaller font size is relatively more difficult to recognize, and thus the recognition speed is lower.
[0097] In the embodiment of the present application, the recognition speed for text of different font sizes can be determined in advance based on the OCR recognition situation, thereby determining a suitable font size threshold.
[0098] Step S507 , performing OCR recognition on the text whose font size is larger than the font size threshold, so as to complete the low-dimensional OCR recognition processing of the image.
[0099] In some embodiments, the image may also include variant forms of characters. Correspondingly, low-dimensional OCR recognition processing can be performed through the following steps:
[0100] Step S508: When the image includes variant characters, low-dimensional OCR recognition processing is not performed on the variant characters.
[0101] Here, since the computer cannot accurately recognize variant characters, variant characters are not recognized.
[0102] Step S509 : determining a target image matching the keyword in a preset image library according to the low-dimensional OCR recognition result or the high-dimensional OCR recognition result of each image.
[0103] Step S510: determining the target image as the search result of the image search request, and displaying the search result.
[0104] In an embodiment of the present application, when performing low-dimensional OCR recognition processing on an image, different OCR recognition thresholds can be set, and text recognition can be performed using different OCR recognition thresholds as reference conditions for recognition, thereby achieving a balance between recognition accuracy and recognition efficiency while ensuring recognition accuracy.
[0105] based on Figure 4 , Figure 6This is an optional flow chart of the image search method provided in the embodiment of the present application. In some embodiments, the recognition accuracy of the low-dimensional OCR recognition process is lower than the recognition accuracy of the high-dimensional OCR recognition process, such as Figure 6 As shown, step S404 can be implemented by the following steps:
[0106] Step S601: determine whether each image has a low-dimensional OCR recognition result.
[0107] If the judgment result is yes, step S602 is executed; if the judgment result is no, the process returns to step S403 to continue performing low-dimensional OCR recognition processing on the image.
[0108] Step S602: Determine whether each image has a high-dimensional OCR recognition result.
[0109] If the judgment result is yes, execute step S603; if the judgment result is no, execute step S604.
[0110] Step S603: When the picture has both a low-dimensional OCR recognition result and a high-dimensional OCR recognition result, the high-dimensional OCR recognition result is determined as the OCR recognition result of the picture.
[0111] Step S604: When the image only has a low-dimensional OCR recognition result, the low-dimensional OCR recognition result is determined as the OCR recognition result of the image.
[0112] Step S605: According to the OCR recognition result of the image, a target image matching the keyword is determined in a preset image library.
[0113] In the embodiment of the present application, since the accuracy of the high-dimensional OCR recognition result is higher than the accuracy of the low-dimensional OCR recognition result, when there are both low-dimensional OCR recognition results and high-dimensional OCR recognition results, the target image is matched based on the high-dimensional OCR recognition result with higher accuracy; and when there is only a low-dimensional OCR recognition result, in order to ensure the timeliness of this image search task and improve the search efficiency of this image search task, the target image is continued to be matched based on the low-dimensional OCR recognition result. At this time, since the low-dimensional OCR recognition result also has a certain degree of credibility and a certain degree of recognition accuracy, the accuracy of the final matching result can also be guaranteed to a certain extent.
[0114] Figure 7 This is an optional flow chart of the high-dimensional OCR recognition process provided by the embodiment of the present application, such as Figure 7 As shown, the method includes the following steps:
[0115] Step S701: Before obtaining an image search request, or after completing a response to the image search request, or when the search request response is interrupted, determine that images in a preset image library that have not completed high-dimensional OCR recognition processing are unprocessed images.
[0116] In an embodiment of the present application, high-dimensional OCR recognition processing can be implemented during idle time, that is, when the image search task is not being executed, high-dimensional OCR recognition processing can be executed in the background. Since the image search task is not executed before obtaining the image search request, or after completing the response to the image search request, or when the search request response is interrupted, high-dimensional OCR recognition processing can be performed during these time periods to complete high-dimensional OCR recognition processing for each image in the preset image library, so that in subsequent image search tasks, image searches can be performed based on the more accurate high-dimensional OCR recognition results.
[0117] It should be noted that unprocessed images are images that have not completed high-dimensional OCR recognition processing, regardless of whether the images have completed high-dimensional OCR recognition processing. That is, unprocessed images include not only images that have not completed low-dimensional OCR recognition processing and high-dimensional OCR recognition processing, but also images that have completed low-dimensional OCR recognition processing but not high-dimensional OCR recognition processing.
[0118] In step S702 , each unprocessed image is processed using high-dimensional OCR recognition processing to obtain a high-dimensional OCR recognition result for each unprocessed image.
[0119] In some embodiments, after performing low-dimensional OCR recognition processing on each traversed image and obtaining the low-dimensional OCR recognition result of each corresponding image, the low-dimensional OCR recognition result is stored in a preset storage unit; after using high-dimensional OCR recognition processing to process each unprocessed image and obtaining the high-dimensional OCR recognition result of each unprocessed image, the high-dimensional OCR recognition result is stored in a preset storage unit, and the low-dimensional OCR recognition result of the corresponding unprocessed image is deleted.
[0120] In an embodiment of the present application, after each low-dimensional OCR recognition processing or high-dimensional OCR recognition processing is completed, the obtained low-dimensional OCR recognition result or high-dimensional OCR recognition result is stored in a preset storage unit. In this way, it can be ensured that when the image search task is subsequently executed, the low-dimensional OCR recognition result or high-dimensional OCR recognition result can be quickly obtained directly from the preset storage unit, and fast keyword matching is performed based on the obtained low-dimensional OCR recognition result or high-dimensional OCR recognition result, without the need to perform low-dimensional OCR recognition processing or high-dimensional OCR recognition processing on the image, thereby improving the image search efficiency.
[0121] In some embodiments, step S702 may be implemented by the following steps:
[0122] Step S7021: Perform text sharpening processing on the unprocessed image to obtain a text sharpened image.
[0123] Here, the text sharpening process includes the following steps: first, segmenting the unprocessed image to form at least two sub-images; then, enlarging each sub-image to obtain an enlarged sub-image.
[0124] In an embodiment of the present application, the unprocessed image can be equally divided into at least two sub-images, or any segmentation method can be used, or based on certain segmentation rules, the unprocessed image can be segmented into at least two irregular or unequal sub-images.
[0125] When an unprocessed image is divided irregularly or unequally, for example, the left third of unprocessed image A is a pure image without any text, while the right two-thirds is a text image formed by text, then unprocessed image A can be divided into two parts, the first part being a sub-image formed by the left third of the pure image, and the second part being a sub-image formed by the right two-thirds of the text image. In this way, since the first part is a pure image, no OCR recognition is required, and since the second part is a text image, this division does not affect the continuity of the text in the second part, allowing for more accurate recognition of the text in the second part, and only requiring OCR recognition on the second part. This not only improves recognition accuracy, but also effectively improves recognition efficiency.
[0126] In an embodiment of the present application, since the high-dimensional OCR recognition result needs to also recognize the detailed content in the image, and the detailed content in the image, such as text, is usually relatively small, in order to improve the recognition accuracy, the segmented sub-image can be enlarged to reduce the difficulty of recognizing the detailed content.
[0127] Step S7022: Perform OCR recognition on the text in the image after the text sharpening process to obtain the high-dimensional OCR recognition result of each unprocessed image.
[0128] Here, performing OCR recognition on the text in the image after text sharpening processing in step S7022 can be achieved by performing OCR recognition on the text in the enlarged sub-image to obtain a sub-recognition result corresponding to each sub-image. Then, the sub-recognition results of each sub-image in at least two sub-images are fused to obtain a high-dimensional OCR recognition result of the unprocessed image.
[0129] Here, it is possible to determine whether there is overlapping content in the sub-recognition results corresponding to at least two sub-images; when there is overlapping content in the sub-recognition results, determine the non-overlapping content and overlapping content in the sub-recognition results of each sub-image; and fuse the non-overlapping content with the overlapping content to obtain the high-dimensional OCR recognition result of the unprocessed image.
[0130] It should be noted that fusing non-overlapping content with overlapping content here refers to deleting the duplicated portions of the overlapping content in the high-dimensional OCR recognition results. For example, when the sub-recognition results of the first sub-image include the four keywords A, B, C, and D, and the sub-recognition results of the second sub-image include the four keywords C, D, E, and F, the non-overlapping contents of the sub-recognition results of the first sub-image and the sub-recognition results of the second sub-image are: A, B, E, and F, while the overlapping contents are C and D. Therefore, after fusing the non-overlapping contents with the overlapping contents, the second OCR recognition result of the unprocessed image should be: A, B, C, D, E, F, and not: A, B, C, D, C, D, E, F. That is, the duplicated portions C and D in the high-dimensional OCR recognition results with the overlapping contents C and D need to be deleted.
[0131] When there is no overlapping content in the sub-recognition results, each sub-image is re-segmented, enlarged, OCR recognized and the sub-recognition results are fused to obtain the recognition result of each sub-image; based on the recognition result of each sub-image, the high-dimensional OCR recognition result of the unprocessed image is determined.
[0132] Here, when there is no overlapping content, in order to further improve the recognition accuracy, the sub-images may be segmented, enlarged, and recognized again, and the recognition results may be fused to obtain more accurate recognition results of the sub-images.
[0133] In some embodiments, the image search method can be implemented by a client in an image search system, a preset storage unit corresponding to the client, and a server. Figure 8 This is an optional flow chart of the image search method provided in the embodiment of the present application. Figure 8 As shown, the method includes the following steps:
[0134] In step S801 , before the client receives the image search request, the server processes each image in the preset image library using high-dimensional OCR recognition processing based on depth recognition to obtain a high-dimensional OCR recognition result for each image.
[0135] Here, the server performs high-dimensional OCR recognition processing on each image in the preset image library during idle time, which can effectively utilize resources and avoid the problem of reducing search efficiency due to high-dimensional OCR recognition processing during image search tasks.
[0136] In step S802 , the server stores the high-dimensional OCR recognition result in a preset storage unit.
[0137] In the embodiment of the present application, when a high-dimensional OCR recognition result of each image is obtained, the high-dimensional OCR recognition result is stored in a preset storage unit, so that the high-dimensional OCR recognition result can be used in a timely manner in the next image search task.
[0138] Step S803: The client obtains an image search request, which includes keywords.
[0139] In step S804, the client obtains an OCR recognition result of each image from a preset storage unit in response to the image search request; wherein the OCR recognition result includes at least one of the following: a low-dimensional OCR recognition result obtained by a low-dimensional OCR recognition process based on an OCR recognition threshold, and a high-dimensional OCR recognition result obtained by a high-dimensional OCR recognition process based on depth recognition, and the recognition accuracy of the low-dimensional OCR recognition process is less than the recognition accuracy of the high-dimensional OCR recognition process.
[0140] In step S805 , the client traverses the pictures that have not completed the low-dimensional OCR recognition processing and the high-dimensional OCR recognition processing, and performs low-dimensional OCR recognition processing on each traversed picture to obtain a low-dimensional OCR recognition result for each corresponding picture.
[0141] Step S806: When a new picture is added to the preset picture library, the client performs low-dimensional OCR recognition processing on the new picture.
[0142] In the embodiment of the present application, when a new image is added to the preset image library, low-dimensional OCR recognition processing is also required for the new image to ensure that each image in the preset image library has a low-dimensional OCR recognition result. Alternatively, in other embodiments, when a new image is added to the preset image library, low-dimensional OCR recognition processing can be performed on the new image in a timely manner during the next image search task.
[0143] In step S807 , the client determines the credibility of the low-dimensional OCR recognition result of each image.
[0144] In the embodiment of the present application, a specific OCR recognition model can be used for OCR recognition. When the OCR recognition model is used for OCR recognition, not only a low-dimensional OCR recognition result can be obtained, but also the credibility corresponding to the low-dimensional OCR recognition result can be obtained.
[0145] Factors influencing reliability include, but are not limited to, at least one of the following: image clarity, image type, and the number of recognized characters. For example, for blurry, low-definition images, the reliability of the recognition result will be relatively low. There are also differences in reliability between printed and handwritten text recognition, with handwritten text recognition being less reliable than printed text. When recognizing the same image, if the number of recognized characters is significantly smaller than the actual number of characters, the reliability of the recognition result will be low.
[0146] Step S808: The client deletes the low-dimensional OCR recognition results whose credibility is lower than the threshold.
[0147] In the embodiment of the present application, a low-dimensional OCR recognition result with high credibility is selected.
[0148] In step S809 , the client stores the low-dimensional OCR recognition result of each corresponding image in a preset storage unit.
[0149] In step S810 , the client determines a target image that matches the keyword in a preset image library based on the low-dimensional OCR recognition result or the high-dimensional OCR recognition result of each image.
[0150] In step S811 , the client determines the target image as the search result of the image search request and displays the search result.
[0151] In step S812 , the server continues to use the high-dimensional OCR recognition process based on depth recognition to process the pictures in the preset picture library that have not yet been subjected to the high-dimensional OCR recognition process, and obtains the high-dimensional OCR recognition results of the pictures.
[0152] Here, since the pictures in the preset picture library have not yet completed high-dimensional OCR recognition processing, during the idle time after completing an image search task, the background server can continue to process the pictures in the preset picture library that have not yet undergone high-dimensional OCR recognition processing based on deep recognition.
[0153] In step S813, the server stores the high-dimensional OCR recognition result in a preset storage unit.
[0154] Step S814: When a new picture is added to the preset picture library, the server performs high-dimensional OCR recognition processing on the new picture.
[0155] In an embodiment of the present application, when a new picture is added to the preset picture library, it is also necessary to perform high-dimensional OCR recognition processing on the new picture to ensure that each picture in the preset picture library has a high-dimensional OCR recognition result.
[0156] Step S815 , after each image is processed by high-dimensional OCR recognition processing to obtain a high-dimensional OCR recognition result of each image, the low-dimensional OCR recognition result of the corresponding image in the preset storage unit is deleted.
[0157] In the embodiment of the present application, since the recognition accuracy of the high-dimensional OCR recognition result is higher than that of the low-dimensional OCR recognition result, when any image has both a low-dimensional OCR recognition result and a high-dimensional OCR recognition result, only the high-dimensional OCR recognition result with higher recognition accuracy can be retained, and the low-dimensional OCR recognition result stored in the preset storage unit can be deleted. In this way, not only can the storage space in the preset storage unit be saved, but it can also ensure that when performing subsequent image search tasks, the high-dimensional OCR recognition result stored in the preset storage unit can be directly used for keyword matching, without having to determine the high-dimensional OCR recognition result with higher recognition accuracy from the low-dimensional OCR recognition result and the high-dimensional OCR recognition result, that is, one step of interpretation and selection is saved, further improving the search efficiency.
[0158] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0159] The embodiment of the present application provides an image search method. In actual product applications, the user only needs to enter the search keyword in the input interface of the image search application, and the image search application can automatically search for search results that match the search keyword, and the search results can accurately include the text in the image. In theory, the method of the embodiment of the present application can be applied to all image search scenarios.
[0160] like Figure 9 As shown in FIG, it is a flow chart of the image search method provided by the embodiment of the present application, as shown in FIG. Figure 9 As shown, the method includes the following steps:
[0161] Step S901: The user inputs a search keyword.
[0162] In step S902 , the system performs OCR recognition on the image and determines whether the OCR recognition result contains the search keyword to obtain search results.
[0163] Step S903: output search results.
[0164] Since OCR recognition takes a certain amount of time, the efficiency of OCR recognition needs to be fast when the user searches. Therefore, the overall search process will be based on the combination of background idle time recognition and fast recognition during search, ensuring that the search results can be quickly output while ensuring the accuracy and completeness of the image OCR recognition. In the embodiment of the present application, the simple recognition strategy and the deep recognition strategy can be combined to implement the image search method, wherein the simple recognition strategy corresponds to the low-dimensional OCR recognition processing in the above embodiment, and the deep recognition strategy corresponds to the high-dimensional OCR recognition processing in the above embodiment. The detailed process of the simple recognition strategy and the deep recognition strategy and the scheduling relationship between them will be elaborated in detail below.
[0165] Figure 10 This is a detailed flow chart of the image search method provided in the embodiment of the present application. Figure 10 As shown, the method includes the following steps:
[0166] Step S101: Obtain user search keywords and search for images.
[0167] Step S102: determining whether all pictures in the preset picture library have been scanned.
[0168] Here, if the preset image library has already undergone a simplified recognition of all images, the recognition result can be used for querying; if the preset image library has not yet completed the simplified recognition of all images, the simplified recognition process will be entered. In other words, if the judgment result is yes, the recognition result is used for querying and step S103 is executed; if the judgment result is no, step S104 is executed.
[0169] Step S103: output the search results.
[0170] In an embodiment of the present application, OCR is used to identify content for searching. If the preset image library has completed a full and brief recognition, the keyword can be used to search the recognition results, and the search results containing the OCR content can be output at the same time.
[0171] Step S104: determine to use a simplified recognition strategy for processing.
[0172] Step S105: traverse the pictures in the preset picture library.
[0173] Step S106: Perform OCR recognition on the text in the image.
[0174] Step S107 , determining whether to perform a deep scan (ie, whether to use a deep recognition strategy to perform background idle time deep recognition).
[0175] If the judgment result is yes, a deep scan is performed and step S110 is executed; if the judgment result is no, step S108 is executed.
[0176] Step S108, determine whether the traversal is completed.
[0177] If the judgment result is yes, the process returns to step S103 ; if the judgment result is no, the process returns to step S105 .
[0178] When it is determined to use the depth recognition strategy for processing, the depth recognition strategy includes the following steps:
[0179] Step S109: determine to use the depth recognition strategy for processing.
[0180] Step S110: Divide the image into four equal parts and enlarge the image.
[0181] Step S111: performing OCR recognition on the segmented image.
[0182] Step S112: determine whether the recognition result is the same as the existing result.
[0183] Here, the existing results refer to the results of recognizing other equal parts of any picture in the historical process when recognizing any equal part of any picture at present.
[0184] In this embodiment of the present application, a determination is made as to whether the current recognition result for any equal portion of any image overlaps with the recognition results for other equal portions of the same image in the past. If the determination result is yes, step S113 is executed; if the determination result is no, the process returns to step S110 to continue segmentation and recognition.
[0185] Step S113: record the recognition result.
[0186] The following is a detailed description of the simplified identification process:
[0187] Figure 11 This is a schematic diagram of a simplified identification process provided by an embodiment of the present application. Figure 11 As shown in the simplified recognition process, OCR is used to simply recognize the text on a picture, such as Figure 12 As shown, it is a schematic diagram of the simplified recognition process in the embodiment of the present application. In the simplified recognition process, for large characters and traditional characters, for example Figure 12 The characters 121 in will be the targets of active recognition, while small characters, variant characters, etc., such as Figure 12 The text 122 in the image requires more time and resources to recognize. Therefore, recognition of this type of font will be abandoned to ensure that the recognition time of a single image can be controlled within 10ms.
[0188] Please continue to refer to Figure 11 ,The simple recognition process includes the following steps:
[0189] Step S111, traverse the pictures in the preset picture library.
[0190] Step S112: Perform OCR recognition on the text in the image.
[0191] Step S113: determine whether the reliability of the recognition result is greater than 80%.
[0192] Here, the recognition results with high credibility are selected. During the simplified recognition process, the results with low credibility are excluded. This is because the simplified recognition is triggered when the user is searching and the deep recognition is not yet completed. Therefore, it is necessary to maintain a certain level of recognition accuracy to ensure that the user can search normally while avoiding excessive search interference caused by low credibility.
[0193] In the embodiment of the present application, the credibility is a value that can be obtained when performing OCR recognition, that is, when performing OCR recognition on an image, not only the recognition result is output, but also the credibility corresponding to the recognition result is output. Factors affecting the credibility include but are not limited to at least one of the following: the clarity of the image, the type of image, and the number of recognized words. For example, for images with relatively low clarity and relatively blurry photos, the credibility of the recognition result will be relatively low; for the recognition of printed and handwritten text, there is also a difference in credibility. Compared with printed text, the recognition result of handwritten text has a lower credibility; when recognizing the same image, if the number of recognized words is much smaller than the actual number of words, the credibility of the recognition result is low.
[0194] In step S113 , if the judgment result is yes, step S114 is executed; if the judgment result is no, step S116 is executed.
[0195] Step S114: save the recognition result.
[0196] In the embodiment of the present application, after all images are recognized, the OCR results of the corresponding images will be saved in the database.
[0197] Step S115, determine whether the traversal is completed.
[0198] If the judgment result is yes, the process ends; if the judgment result is no, the process returns to step S111.
[0199] Step S116: discard the recognition result.
[0200] Step S117: When a new picture is added, continue with step S112 to perform OCR recognition on the text in the new picture.
[0201] In the embodiment of the present application, when new photos are added, there is no need to perform full recognition. It is only necessary to recognize the newly added pictures once and save the recognition results to the database.
[0202] The following is a detailed description of the depth recognition strategy:
[0203] Figure 13 This is a flow chart of the depth recognition strategy provided by the embodiment of the present application. Figure 13 As shown, the depth recognition strategy includes the following steps:
[0204] Step S131, traverse the pictures in the preset picture library.
[0205] It should be noted that in the embodiment of the present application, deep recognition will be performed during idle time (for example, late at night, when charging and the application is not in use).
[0206] Step S132: divide the image into four equal parts and enlarge it.
[0207] In depth recognition, the image is segmented and enlarged to ensure that more information can be recognized. For example, during depth recognition, the image can be divided into four equal parts. The purpose of this operation is to recognize more text in the image. Figure 14A This is a schematic diagram of an image before being divided into equal parts and enlarged, provided in an embodiment of the present application. Figure 14B This is a schematic diagram of an image divided into equal parts and magnified according to an embodiment of the present application. Figure 14A and Figure 14B As shown, in the original image 141 before the image is divided into equal parts and enlarged, the text in the image is small and difficult to recognize, while in the partially enlarged image 142 after the image is divided into equal parts and enlarged, the text is enlarged and easy to recognize.
[0208] In the embodiment of the present application, the image will be recognized after segmentation. If a segmented image does not contain any text information, the segmented image will be discarded and the segmented image area will no longer be recognized.
[0209] Step S133: Perform OCR recognition on the segmented image.
[0210] Step S134: determine whether there is an OCR recognition result.
[0211] If the judgment result is yes, execute step S135; if the judgment result is no, execute step S136.
[0212] Step S135: determine whether the credibility of the recognition result is lower than a threshold.
[0213] In the embodiment of the present application, re-segmentation and recognition can be performed on the results with low credibility.
[0214] In some cases, after the image is divided into four equal parts, it may still contain too much text information (such as panoramic images, long screenshots, etc.). At this time, the credibility of the recognition result will be low. For this part of the image, the segmented image will be segmented again, and the second segmented image will be recognized in the same way. If the image has been recognized with high credibility (for example, credibility greater than 70%), or the image does not contain text information, there is no need to segment it again.
[0215] In step S135, if the judgment result is yes, the process returns to step S132 to continue segmenting and identifying the image; if the judgment result is no, the process proceeds to step S136.
[0216] Step S136: skip the segmented picture.
[0217] Step S137, determine whether all segmented pictures have been traversed.
[0218] If the judgment result is yes, step S138 is executed. If the judgment result is no, the segmented pictures are traversed continuously and the process returns to step S133.
[0219] Step S138: Determine whether the pictures in the preset picture library have been traversed.
[0220] If the judgment result is yes, the process ends; if the judgment result is no, the process returns to step S131 to continue traversing the pictures.
[0221] In this embodiment of the present application, the recognition results can be saved in a database. If the database already contains the results of the simplified recognition process for the image, the results of the simplified recognition process are replaced with the results of the deep recognition process. Similarly, if a new image is added, the new image can also be directly incrementally recognized, that is, deep recognition is performed on the new image.
[0222] The image search method provided in the embodiment of the present application can more accurately search for text information in photos when searching for photos, and provides more dimensional photo searches. It can utilize more search scenarios, such as: note search, screenshot search, etc., and has a high accuracy rate without the need for background cloud recognition, and can be used offline.
[0223] The following continues to describe the exemplary structure of the picture search device 354 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 3 As shown, the software module stored in the image search device 354 of the memory 350 may be the image search device in the electronic device 300, including:
[0224] An acquisition module 3541 is configured to acquire an image search request, wherein the image search request includes keywords;
[0225] a response module 3542 configured to obtain, in response to the image search request, an OCR recognition result for each image in a preset image library; wherein the OCR recognition result includes at least one of the following: a low-dimensional OCR recognition result obtained by a low-dimensional OCR recognition process based on an OCR recognition threshold, and a high-dimensional OCR recognition result obtained by a high-dimensional OCR recognition process based on depth recognition, wherein the recognition accuracy of the low-dimensional OCR recognition process is less than the recognition accuracy of the high-dimensional OCR recognition process;
[0226] The processing module 3543 is configured to traverse the images that have not completed the low-dimensional OCR recognition process and the high-dimensional OCR recognition process, and perform the low-dimensional OCR recognition process on each traversed image to obtain a low-dimensional OCR recognition result for each corresponding image;
[0227] A first determining module 3544 is configured to determine a target image matching the keyword in the preset image library based on the low-dimensional OCR recognition result or the high-dimensional OCR recognition result of each image;
[0228] The second determining module 3545 is configured to determine the target image as the search result of the image search request and display the search result.
[0229] In some embodiments, the processing module is also used to: when it is determined that there is at least one picture in the preset picture library that has not completed the low-dimensional OCR recognition processing and has not completed the high-dimensional OCR recognition processing, traverse the pictures that have not completed the low-dimensional OCR recognition processing and the high-dimensional OCR recognition processing, and perform the low-dimensional OCR recognition processing on each traversed picture.
[0230] In some embodiments, the OCR recognition threshold includes a recognition speed threshold, and the processing module is further used to: determine the recognition speed for each text in the image; perform OCR recognition on the text whose recognition speed is greater than the recognition speed threshold to complete the low-dimensional OCR recognition processing of the image.
[0231] In some embodiments, the OCR recognition threshold includes a font size threshold, and the processing module is further used to: determine the font size of each text in the image; perform OCR recognition on the text whose font size is greater than the font size threshold to complete the low-dimensional OCR recognition processing of the image.
[0232] In some embodiments, the processing module is further configured to: when the image includes variant characters, not perform the low-dimensional OCR recognition processing on the variant characters.
[0233] In some embodiments, the first determination module is further used to: when the picture has both the low-dimensional OCR recognition result and the high-dimensional OCR recognition result, determine the high-dimensional OCR recognition result as the OCR recognition result of the picture; when the picture only has the low-dimensional OCR recognition result, determine the low-dimensional OCR recognition result as the OCR recognition result of the picture; and determine a target picture matching the keyword in the preset picture library based on the OCR recognition result of the picture.
[0234] In some embodiments, the device also includes: a third determination module, which is used to determine that the pictures in the preset picture library that have not completed the high-dimensional OCR recognition processing are unprocessed pictures before obtaining the picture search request, or after completing the response to the picture search request, or when the search request response is interrupted; and a picture processing module, which is used to use the high-dimensional OCR recognition processing to process each of the unprocessed pictures to obtain the high-dimensional OCR recognition result of each of the unprocessed pictures.
[0235] In some embodiments, the image processing module is further used to: perform text sharpening processing on the unprocessed image to obtain a text sharpened image; perform OCR recognition on the text in the text sharpened image to obtain the high-dimensional OCR recognition result of each of the unprocessed images.
[0236] In some embodiments, the image processing module is also used to: segment the unprocessed image to form at least two sub-images; enlarge each of the sub-images to obtain an enlarged sub-image; perform OCR recognition on the text in the enlarged sub-image to obtain a sub-recognition result corresponding to each of the sub-images; and fuse the sub-recognition results of each sub-image in the at least two sub-images to obtain the high-dimensional OCR recognition result of the unprocessed image.
[0237] In some embodiments, the image processing module is also used to: when the sub-recognition result contains overlapping content, determine the non-overlapping content and the overlapping content in the sub-recognition result of each sub-image; fuse the non-overlapping content with the overlapping content to obtain the high-dimensional OCR recognition result of the unprocessed image; when the sub-recognition result does not contain overlapping content, re-segment, enlarge, OCR recognize and fuse the sub-recognition results of each sub-image to obtain the recognition result of each sub-image; and determine the high-dimensional OCR recognition result of the unprocessed image based on the recognition result of each sub-image.
[0238] In some embodiments, the device further includes: a storage module for performing the low-dimensional OCR recognition processing on each traversed image to obtain the low-dimensional OCR recognition result of each corresponding image, and then storing the low-dimensional OCR recognition result in a preset storage unit; and, after processing each of the unprocessed images using the high-dimensional OCR recognition processing to obtain the high-dimensional OCR recognition result of each unprocessed image, storing the high-dimensional OCR recognition result in the preset storage unit and deleting the low-dimensional OCR recognition result of the corresponding unprocessed image.
[0239] In some embodiments, the apparatus further includes: a fourth determining module for determining the credibility corresponding to the low-dimensional OCR recognition result of each of the images; and a deleting module for deleting the low-dimensional OCR recognition results whose credibility is lower than a threshold.
[0240] In some embodiments, the apparatus further includes: an OCR recognition processing module, configured to perform the low-dimensional OCR recognition processing or the high-dimensional OCR recognition processing on a new image when a new image is added to the preset image library.
[0241] It should be noted that the description of the device embodiment of the present application is similar to the description of the method embodiment described above, and has similar beneficial effects as the method embodiment, so it will not be repeated. For technical details not disclosed in the device embodiment, please refer to the description of the method embodiment of the present application for understanding.
[0242] The present invention provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method described above in the present invention.
[0243] The embodiment of the present application provides a storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will execute the method provided by the embodiment of the present application, for example, Figure 4 The method shown.
[0244] In some embodiments, the storage medium can be a computer-readable storage medium, such as a ferroelectric random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various devices including one or any combination of the above memories.
[0245] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0246] By way of example, executable instructions may, but need not necessarily, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as one or more scripts in a Hypertext Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions). By way of example, executable instructions may be deployed for execution on one computing device, or on multiple computing devices located at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.
[0247] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A picture search method, characterized in that: include: Obtaining an image search request, wherein the image search request includes keywords; In response to the image search request, obtaining an OCR recognition result for each image in a preset image library; wherein the OCR recognition result includes at least one of the following: a low-dimensional OCR recognition result obtained by using a low-dimensional OCR recognition process based on an OCR recognition threshold, and a high-dimensional OCR recognition result obtained by using a high-dimensional OCR recognition process based on depth recognition, wherein the recognition accuracy of the low-dimensional OCR recognition process is less than the recognition accuracy of the high-dimensional OCR recognition process; Traversing the pictures that have not completed the low-dimensional OCR recognition process and the high-dimensional OCR recognition process, and performing the low-dimensional OCR recognition process on each traversed picture to obtain a low-dimensional OCR recognition result for each corresponding picture; When the image search task is not performed, determining that the images in the preset image library that have not completed the high-dimensional OCR recognition processing are unprocessed images; Processing each of the unprocessed images using the high-dimensional OCR recognition process in the background to obtain the high-dimensional OCR recognition result of each of the unprocessed images; Determining a target image matching the keyword in the preset image library according to the low-dimensional OCR recognition result or the high-dimensional OCR recognition result of each image; The target image is determined as a search result of the image search request, and the search result is displayed.
2. The method according to claim 1, characterized in that The traversing of the pictures that have not completed the low-dimensional OCR recognition process and the high-dimensional OCR recognition process, and performing the low-dimensional OCR recognition process on each traversed picture, includes: When it is determined that there is at least one picture in the preset picture library that has not completed the low-dimensional OCR recognition processing and the high-dimensional OCR recognition processing, traverse the pictures that have not completed the low-dimensional OCR recognition processing and the high-dimensional OCR recognition processing, and perform the low-dimensional OCR recognition processing on each traversed picture.
3. The method according to claim 1, characterized in that The OCR recognition threshold includes a recognition speed threshold, and the low-dimensional OCR recognition processing is performed on each traversed image, including: Determining a recognition speed for each word in the image; Perform OCR recognition on the text whose recognition speed is greater than the recognition speed threshold to complete the low-dimensional OCR recognition processing of the image.
4. The method according to claim 1, wherein The OCR recognition threshold includes a font size threshold, and the low-dimensional OCR recognition processing is performed on each traversed image, including: Determine the font size of each text in the image; Perform OCR recognition on the text whose font size is larger than the font size threshold to complete the low-dimensional OCR recognition processing on the image.
5. The method according to claim 1, wherein The low-dimensional OCR recognition process is performed on each traversed image, including: When the picture includes variant characters, the low-dimensional OCR recognition process is not performed on the variant characters.
6. The method according to claim 1, characterized in that The determining, in the preset image library, a target image matching the keyword based on the low-dimensional OCR recognition result or the high-dimensional OCR recognition result of each image includes: When the picture has both the low-dimensional OCR recognition result and the high-dimensional OCR recognition result, determining the high-dimensional OCR recognition result as the OCR recognition result of the picture; When the picture only has the low-dimensional OCR recognition result, determining the low-dimensional OCR recognition result as the OCR recognition result of the picture; According to the OCR recognition result of the image, a target image matching the keyword is determined in the preset image library.
7. The method according to claim 6, characterized in that Determining a target image matching the keyword in the preset image library according to the OCR recognition result of the image includes: Determine the image keywords corresponding to each image based on the text content corresponding to the OCR recognition result of the image; The keyword in the image search request is matched with the image keyword of each image, and the image corresponding to the image keyword that is the same as the keyword is determined as the target image.
8. The method according to claim 1, characterized in that The displaying of the search results includes: When the determined target image is one, displaying the target image on the current interface; When the determined target pictures are multiple, the multiple target pictures are displayed on the current interface at the same time, or the multiple target pictures are displayed in pages.
9. The method according to claim 1, characterized in that The step of processing each of the unprocessed images using the high-dimensional OCR recognition process to obtain the high-dimensional OCR recognition result of each of the unprocessed images includes: Performing text-clearing processing on the unprocessed image to obtain a text-cleared image; Perform OCR recognition on the text in the image after the text sharpening process to obtain the high-dimensional OCR recognition result of each of the unprocessed images.
10. The method according to claim 9, characterized in that The performing text-clearing processing on the unprocessed image to obtain a text-cleared image includes: Segmenting the unprocessed image to form at least two sub-images; Performing an amplification process on each of the sub-images to obtain an amplified sub-image; The performing OCR recognition on the text in the image after the text sharpening process to obtain the high-dimensional OCR recognition result of each of the unprocessed images includes: Performing OCR recognition on the text in the enlarged sub-image to obtain a sub-recognition result corresponding to each sub-image; The sub-recognition results of each sub-image of the at least two sub-images are fused to obtain the high-dimensional OCR recognition result of the unprocessed image.
11. The method according to claim 10, characterized in that The fusing the sub-recognition results of each sub-image of the at least two sub-images to obtain the high-dimensional OCR recognition result of the unprocessed image includes: When the sub-recognition results have overlapping content, determining non-overlapping content and the overlapping content in the sub-recognition results of each of the sub-pictures; Fusing the non-overlapping content with the overlapping content to obtain the high-dimensional OCR recognition result of the unprocessed image; When the sub-recognition results do not contain the overlapping content, performing the re-segmentation, the amplification, the OCR recognition, and the fusion of the sub-recognition results on each of the sub-images to obtain a recognition result for each of the sub-images; The high-dimensional OCR recognition result of the unprocessed image is determined according to the recognition result of each sub-image.
12. The method according to claim 1, characterized in that The method further comprises: After performing the low-dimensional OCR recognition process on each traversed image and obtaining a low-dimensional OCR recognition result for each corresponding image, the low-dimensional OCR recognition result is stored in a preset storage unit; After each of the unprocessed images is processed using the high-dimensional OCR recognition process to obtain the high-dimensional OCR recognition result of each of the unprocessed images, the high-dimensional OCR recognition result is stored in the preset storage unit, and the low-dimensional OCR recognition result of the corresponding unprocessed image is deleted.
13. The method according to any one of claims 1 to 6, characterized in that The method further comprises: When a new picture is added to the preset picture library, the low-dimensional OCR recognition process or the high-dimensional OCR recognition process is performed on the new picture.
14. A picture search device, characterized in that: include: An acquisition module, configured to acquire an image search request, wherein the image search request includes keywords; a response module, configured to obtain, in response to the image search request, an OCR recognition result for each image in a preset image library; wherein the OCR recognition result includes at least one of the following: a low-dimensional OCR recognition result obtained by low-dimensional OCR recognition processing based on an OCR recognition threshold, and a high-dimensional OCR recognition result obtained by high-dimensional OCR recognition processing based on depth recognition, wherein the recognition accuracy of the low-dimensional OCR recognition processing is less than the recognition accuracy of the high-dimensional OCR recognition processing; a processing module, configured to traverse images that have not completed the low-dimensional OCR recognition processing and the high-dimensional OCR recognition processing, and perform the low-dimensional OCR recognition processing on each traversed image to obtain a low-dimensional OCR recognition result for each corresponding image; when not performing an image search task, determine that images in the preset image library that have not completed the high-dimensional OCR recognition processing are unprocessed images; and use the high-dimensional OCR recognition processing in the background to process each of the unprocessed images to obtain the high-dimensional OCR recognition result for each of the unprocessed images; A first determining module is configured to determine a target image matching the keyword in the preset image library based on the low-dimensional OCR recognition result or the high-dimensional OCR recognition result of each image; The second determining module is configured to determine the target image as a search result of the image search request and display the search result.
15. An image search device, characterized in that: include: a memory for storing executable instructions; The processor is configured to implement the image search method according to any one of claims 1 to 13 when executing the executable instructions stored in the memory.
16. A computer program product, characterized in that The computer program product includes computer instructions stored in a computer-readable storage medium; The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor is used to execute the computer instructions to implement the image search method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Image processing method and device, processing apparatus, and storage medium thereof
CN109033261A
Image forming apparatus
JP2015032017A
OCR multi-resolution method and apparatus
US20090169131A1