Image search methods and apparatus

By combining image and text search in image search, extracting text information from images and performing text search, the problem of low recall and precision in existing technologies is solved, achieving high recall and high precision image search and improving user experience.

CN116204671BActive Publication Date: 2026-08-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2023-01-31
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing image search technologies have low recall and accuracy rates when assisting users in identifying and purchasing medicines, failing to meet user needs and requiring additional user actions, thus reducing the user experience.

Method used

By combining image search and text search methods, text information is extracted from the image to be searched, and a text search is performed in a second database based on the text information. The image search and text search results are then merged to output the final search results.

Benefits of technology

It improves the recall and accuracy of image search, avoids additional user operations, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116204671B_ABST
    Figure CN116204671B_ABST
Patent Text Reader

Abstract

This disclosure provides an image search method and apparatus, relating to the field of artificial intelligence technology, specifically to the fields of computer vision, image processing, and deep learning. The image search method according to this disclosure includes: performing an image search on an image to be searched to obtain a first search result from a first database; extracting text information from the image to be searched; performing a text search based on the extracted text information to obtain a second search result from a second database; and outputting a search result for the image to be searched based on the first search result and the second search result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, specifically to the fields of computer vision, image processing, and deep learning, and more specifically to image search methods and apparatus applicable to scenarios such as optical character recognition (OCR) and smart healthcare. Background Technology

[0002] With the rapid development of artificial intelligence technology, image search technology has received increasing attention. When a user sees a product or item of interest, they can take a picture of it and perform an image search in a database based on the image. However, the recall and precision of image search largely depend on the richness of the database. Especially when image search technology is used to assist users in identifying and purchasing medicines, the recall and precision are low because the drug image database is often cluttered and incomplete. In such cases, users often need to manually perform further searches. Therefore, this search method cannot meet user needs and degrades the user experience due to the need for additional operations. Summary of the Invention

[0003] This disclosure provides an image search method and apparatus, an electronic device, a non-transitory computer-readable storage medium, and a computer program product.

[0004] According to one aspect of this disclosure, an image search method is provided, comprising: performing an image search on an image to be searched to obtain a first search result from a first database; extracting text information from the image to be searched; performing a text search based on the extracted text information to obtain a second search result from a second database; and outputting a search result for the image to be searched based on the first search result and the second search result.

[0005] According to another aspect of this disclosure, an image search apparatus is provided, comprising: an image search module configured to perform an image search on an image to be searched to obtain a first search result from a first database; a text extraction module configured to extract text information from the image to be searched; a text search module configured to perform a text search based on the extracted text information to obtain a second search result from a second database; and an output module configured to output search results for the image to be searched based on the first search result and the second search result.

[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method according to an example embodiment of this disclosure.

[0007] According to another aspect of this disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause the computer to perform the method described according to an example embodiment of this disclosure.

[0008] According to another aspect of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements a method according to an example embodiment of this disclosure.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0011] Figure 1 An exemplary system architecture of the image search method and apparatus according to embodiments of this disclosure is shown;

[0012] Figure 2 A flowchart of an image search method according to an exemplary embodiment of this disclosure is shown;

[0013] Figure 3 A flowchart illustrating the operation of performing an image search on an image to be searched according to an exemplary embodiment of this disclosure is shown;

[0014] Figure 4A A flowchart illustrating the operation of extracting text information from an image to be searched according to an exemplary embodiment of this disclosure is shown;

[0015] Figure 4B A schematic diagram is shown of a text box detected via scene text detection according to an exemplary embodiment of the present disclosure;

[0016] Figure 5 A flowchart illustrating an example of performing a text search in a second database based on extracted text information, according to an exemplary embodiment of this disclosure;

[0017] Figure 6 A schematic diagram illustrating a specific example of an image search method according to an exemplary embodiment of the present disclosure is shown;

[0018] Figure 7 An example block diagram of an image search apparatus according to an exemplary embodiment of the present disclosure is shown; and

[0019] Figure 8 A block diagram is shown of another example of an electronic device used to implement embodiments of the present disclosure. Detailed Implementation

[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0021] As mentioned earlier, existing image search methods that rely solely on images have low recall and precision, failing to meet user needs and degrading the user experience. Therefore, there is a need for an image search method and apparatus that can improve search recall and precision.

[0022] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0023] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0024] Figure 1 An exemplary system architecture of an image search method and apparatus according to embodiments of the present disclosure is illustrated.

[0025] It is important to note that Figure 1 The examples shown are merely examples of system architectures that can be applied to embodiments of this disclosure, intended to help those skilled in the art understand the technical content of this disclosure. However, they do not imply that embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For instance, in another embodiment, an exemplary system architecture to which image search methods and apparatus can be applied may include a terminal device. However, the terminal device can implement the image search methods and apparatus provided in embodiments of this disclosure without interacting with a server.

[0026] like Figure 1As shown, the system architecture 10 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as at least one of wired and wireless communication links.

[0027] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as at least one of the following: search applications, instant messaging tools, and e-commerce software.

[0028] Terminal devices 101, 102, and 103 can be various electronic devices capable of accessing a search engine. For example, electronic devices can include at least one of intelligent vehicles, smartphones, tablets, laptops, and desktop computers.

[0029] Server 105 can be a server providing various services. For example, server 105 can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the shortcomings of traditional physical hosts and virtual private servers (VPS, Virtual Private Server) services, such as high management difficulty and weak business scalability. It should be noted that the image search method provided in this embodiment can generally be executed by terminal devices 101, 102, or 103. Correspondingly, the image search method provided in this embodiment can also be set in terminal devices 101, 102, or 103.

[0030] Alternatively, the image search method provided in this embodiment can also generally be executed by server 105. Accordingly, the image search method provided in this embodiment can generally be set in server 105. The image search method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Accordingly, the image search device provided in this embodiment can also be set in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0031] It should be noted that the image search method provided in this embodiment can generally be executed by server 105. Correspondingly, the image search device provided in this embodiment can generally be located in server 105. The image search method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the image search device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0032] Alternatively, the image search method provided in the embodiments of this disclosure can also be executed by terminal devices 101, 102, or 103. Correspondingly, the image search device provided in the embodiments of this disclosure can also be disposed in terminal devices 101, 102, or 103.

[0033] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0034] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.

[0035] Figure 2 This is a flowchart of an image search method 200 according to an example embodiment of the present disclosure.

[0036] like Figure 2 As shown, the image search method 200 according to an example embodiment of this disclosure may include steps S210 to S240. In step S210, an image search is performed on the image to be searched to obtain a first search result from a first database. The image to be searched may be an image of a product or item that the user is interested in. For example, the image to be searched may be image data obtained by taking a picture of the product or item using a shooting device such as a mobile phone camera of any type, or by capturing video via a video recording device and subsequently segmenting the video frame by frame. The first database may be a database for image search and stores data on feature maps of multiple images.

[0037] In step S220, text information is extracted from the image to be searched. In one example, Optical Character Recognition (OCR) technology can be used to extract the text information contained in the image to be searched. The text information may be instructions for use of the product or item, brand information, product identification codes, etc.

[0038] Furthermore, in step S230, a text search can be performed based on the extracted text information to obtain a second search result from the second database. The second database can be a database used for text search and stores data about multiple text information. In one example, a text search can be performed directly in the second database based on the extracted text information. In another example, similar text with similar semantics can be generated based on the extracted text information; and a text search can be performed in the second database based on the similar text. By generating similar text based on the extracted text information, the search terms can be further enriched, thereby further improving the recall rate.

[0039] Step S240: Based on the first and second search results, output the search results for the image to be searched. In other words, the fusion result of text search and image search can be output as the search result for the image to be searched. In one example, the results of text search and image search can be sorted in reverse order of similarity or prediction score, and several results with the highest similarity (i.e., TopK results) can be selected as the output. In another example, the fusion result of text search and image search can be obtained by setting a threshold. For example, the first search result is compared with the threshold, and in response to the number of first search results being less than the threshold, the second search result is output as the search result for the image to be searched. By fusing the first and second search results, the problem of low recall and precision caused by relying solely on image search can be solved, thus providing a search method with both high recall and precision without additional user action.

[0040] It should be noted that the above description illustrates an image search method according to an exemplary embodiment of this disclosure by using image search from a first database and text retrieval from a second database as examples. However, those skilled in the art should understand that, depending on the data format stored in the databases, the first database and the second database may be the same or different databases.

[0041] Furthermore, those skilled in the art should understand that, for ease of description, the image search method according to the exemplary embodiments of this disclosure is described with specific step numbers. However, the image search method according to the exemplary embodiments of this disclosure is not limited to being implemented in the specific order shown by the step numbers above. The steps described in one or more orders may be parallel steps or steps executed in reverse order. For example, step S220 may be executed after step S210, simultaneously with step S210, or before step S210.

[0042] The above is for reference only. Figure 2This paper describes an image search method that improves recall and precision without requiring additional user intervention. It addresses the problem of low recall and precision caused by relying solely on image search by combining image search and text search.

[0043] Figure 3 A flowchart illustrating the operation of performing an image search on an image to be searched according to an exemplary embodiment of this disclosure is shown.

[0044] like Figure 3 As shown, the operation of performing an image search on the image to be searched may include sub-steps S311 to S313. In sub-step S311, a feature map of the image to be searched is extracted. For example, a residual network structure ResNet101 can be used to extract the feature map of the image to be searched, wherein the residual network structure ResNet101 may have a network model determined by using Neural Architecture Search (NAS) technology. In the image search method according to the example embodiment of this disclosure, the model of the feature extraction network determined by NAS technology is used, which enables the automatic design of the preferred network structure of the feature extraction network based on prior data, thereby saving a lot of resources. Therefore, when a user wants to search for a drug by a drug image, the drug image can be input into the feature extraction model ResNet101, and 512-dimensional image features can be obtained.

[0045] In sub-step S312, the extracted feature map can be compared with the image features stored in the first database to obtain one or more entries with similar features. As mentioned earlier, the first database can store information about multiple image features; for example, the first database can store data on image features of approximately 20 million drugs in approximately 150,000 categories. By comparing the 512-dimensional features extracted via sub-step S311 with the 20 million image features stored in the database, the most similar features can be obtained. In one example, similar image features can be sorted in reverse order of similarity, and the K most similar entries can be output.

[0046] In sub-step S313, one or more selected entries can be output as the first search result, i.e., the image search result. The accuracy of the first search result depends not only on the feature extraction capability of the feature extraction network but also on the richness of the image features stored in the first database. The multi-layer residual network structure ResNet101 can extract multi-dimensional image features to a great extent, enabling the extracted image features to more realistically reflect the characteristics of the image, thereby improving the accuracy of image search.

[0047] However, when the image features stored in the first database are missing or insufficient, it is necessary to combine text search to improve the accuracy and recall of the search. The following will combine... Figures 4A to 5 This describes the operation of extracting text information from an image to be searched and performing a text search based on that text information.

[0048] Figure 4A A flowchart illustrating the operation of extracting text information from an image to be searched according to an exemplary embodiment of this disclosure is shown.

[0049] like Figure 4A As shown, the operation of extracting text information from the image to be searched may include sub-steps S421 and S422. In sub-step S421, scene text detection is performed on the image to be searched to obtain text box information related to the text boxes containing text information, thereby locating the position of the text information. In one example, scene text detection can be performed on the image to be searched using an Efficient and Accurate Scene Text (EAST) detection model to obtain scene text box information. For example, the EAST detection model based on the 18-layer ResNet18_vd residual network structure can be used to perform scene text detection on the image to be searched.

[0050] Figure 4B This diagram illustrates text boxes detected via scene text detection. (Example) Figure 4B As shown, when using the EAST detection model based on the 18-layer residual network structure ResNet18_vd to perform scene text detection on the search image, it is possible to obtain text box information related to the text boxes containing text information in the search image, such as the position information of text boxes 401 to 407.

[0051] Continue to refer to Figure 4A The operation of extracting text information from the image to be searched may further include sub-step S422. In sub-step S422, text recognition is performed on the region within the text box based on the text box information to extract the text information. For example, a Connectionist Temporal Classification (CTC) recognition model can be used to perform text recognition on the results of scene text detection to extract the text information. In one example, a CTC text line recognition model based on a squeeze and excitation network structure (SeNet) can be used to recognize the text information. For example, a 34-layer squeeze and excitation network structure SeNet34 can be used as the backbone to recognize... Figure 4B The text information included in each text box shown is, for example, "clearing heat and detoxifying" or "used for viral colds".

[0052] It should be noted that Figure 4A and Figure 4B The example shown illustrates the location and recognition of text information at the text box level. However, those skilled in the art will recognize that this application is not limited to this, and the location and recognition of text information can also be performed at the keyword level.

[0053] After extracting the text information, a text search can be performed in the second database based on the extracted text information. Figure 5 A flowchart is shown as an example of an operation to perform a text search in a second database based on extracted text information, according to an exemplary embodiment of this disclosure.

[0054] exist Figure 5 In the example shown, in sub-step S531, similar text with similar semantics is generated based on the extracted text information. For example, the SimBERT model can be used to generate similar text based on the extracted text information. The SimBERT model uses an encoder and decoder structure to encode the input text into a fixed-size vector, and then decodes it by performing autoregression based on this vector to generate the corresponding text. In this process, the SimBERT model can complete both the similar text generation task and the text semantic vector acquisition task, thereby completing the similar text retrieval task. That is, the model has both similar text generation and similar text retrieval functions. By generating similar text based on the extracted text information, the search terms can be further enriched, thereby further improving the recall rate.

[0055] In sub-step S532, a text search is performed on the second database based on similar text. As mentioned earlier, since the SimBERT model has similar text generation and retrieval capabilities, it can perform similar text generation and retrieval. However, when the database information is extremely limited, a brute-force search can be used to perform a search based on the generated similar text, thus easily traversing all possible candidates.

[0056] The above combination Figure 5 An example of performing a text search from a second database based on extracted text information is described. However, those skilled in the art will recognize that methods for performing text searches in a second database are not limited to the methods described above. For example, text semantic vectors can be directly generated based on the extracted text information, and text retrieval tasks can be performed based on these text semantic vectors.

[0057] As mentioned earlier, in response to obtaining the first search result of the image search and the second search result of the text search, the first search result and the second search result can be merged to select the optimal search result output.

[0058] Figure 6 A schematic diagram illustrating a specific example of an image search method according to an exemplary embodiment of this disclosure is shown.

[0059] like Figure 6 As shown, a user can acquire an image 60 of interest by capturing it with a camera. The image 60 is input into an image feature extraction model, such as ResNet101, to generate 512-dimensional image features, as shown in label 611. These 512-dimensional image features are provided to a visual embedding layer to generate a low-dimensional, dense vector 1 representing the image features or the image itself, as shown in label 612. This vector 1 can be provided to a first database for image searching, as shown in label 613. The similarity between image features is determined by calculating the distance between vector 1 and vectors corresponding to image features stored in the first database, thereby determining the first search result.

[0060] also, Figure 6 The image search method shown also includes extracting text information based on the image to be searched 60, as indicated by label 620. For example, text information such as "clearing heat and detoxifying" and "used for viral colds" can be extracted by using the EAST detection model and CTC recognition model based on ResNet18_vd.

[0061] The extracted text information can be used for similar text generation, as shown by label 631. The similar text is provided to the text embedding layer to generate vector 2 representing the similar text, as shown by label 632. This vector 2 can then be provided to a second database for text searching, as shown by label 633. The similarity between the texts is determined by calculating the distance between vector 2 and the vector corresponding to the text stored in the second database, thereby determining the second search result.

[0062] The first and second search results are provided for output, as shown by label 640. For example, the first and second search results can be sorted in reverse order of similarity or prediction score, and several results with the highest similarity can be selected as output. Alternatively, a threshold can be set for image searches, and the first search results can be compared with the set threshold to determine the final search result. For example, if the number of first search results is less than the threshold, the second search result is output as the search result for the image to be searched. Those skilled in the art should understand that the fusion method for the first and second search results is not limited to this. Different weights can be applied to the first search result obtained through image search and the second search result obtained through text search, depending on the actual application needs, to fuse them into the final output result.

[0063] The above combination Figure 6 A specific example of an image search method according to an exemplary embodiment of the present disclosure is described. By merging a first search result and a second search result, the image search method according to an exemplary embodiment of the present disclosure can effectively solve the problem of low recall and precision in searches caused by relying solely on image search, without requiring additional user intervention.

[0064] Figure 7 An example block diagram of an image search apparatus 700 according to an exemplary embodiment of the present disclosure is shown.

[0065] like Figure 7 As shown, the image search device 700 may include an image search module 710, a text extraction module 720, a text search module 730, and an output module 740.

[0066] Image search module 710 can be configured to perform an image search on an image to be searched, in order to obtain a first search result from a first database. The image to be searched may be an image of a product or item that the user is interested in. The first database may be a database used for image search and stores information about feature maps of multiple images. In one example, image search module 710 can also be configured to: extract feature maps from the image to be searched; compare the extracted feature maps with image features stored in the first database to obtain one or more entries with similar features; and output one or more entries as the first search result.

[0067] The text extraction module 720 can be configured to extract text information from the image to be searched. In one example, OCR technology can be used to extract the text information contained in the image to be searched. The text extraction module 720 can be further configured to: perform scene text detection on the image to be searched to obtain text box information related to text boxes containing text information; and perform character recognition on the region within the text box based on the text box information to extract the text information.

[0068] The text search module 730 can be configured to perform a text search based on the extracted text information to obtain a second search result from a second database. The second database can be a database used for text search and stores data about multiple text information. In one example, a text search can be performed directly in the second database based on the extracted text information. In another example, the text search module 730 can be configured to: generate similar text with similar semantics based on the extracted text information; and perform a text search in the second database based on the similar text. By generating similar text based on the extracted text information, the search terms can be further enriched, thereby further improving the recall rate.

[0069] Output module 740 can be configured to output search results for the image to be searched based on the first search result and the second search result. In one example, output module 740 can sort the results of text search and image search in reverse order of similarity or prediction score, and select several results with the highest similarity as output. In another example, output module 740 can obtain a fusion result of text search and image search by setting a threshold. For example, the first search result is compared with the threshold, and in response to the number of first search results being less than the threshold, the second search result is output as the search result for the image to be searched.

[0070] Those skilled in the art will understand that, depending on the data format stored in the database, the first database and the second database can be the same or different databases.

[0071] The above is for reference only. Figure 7 An image search device that improves recall and precision without additional user intervention is described. It solves the problem of low recall and precision caused by relying solely on image search by combining image search and text search.

[0072] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0073] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0074] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0075] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0076] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and steps described above, for example, such as... Figures 2 to 6 The methods and steps are shown. For example, in some embodiments, Figures 2 to 6 The methods and steps illustrated can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by computing unit 801, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, computing unit 801 may be configured by any other suitable means (e.g., by means of firmware) to perform the methods and steps described above.

[0077] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0078] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0079] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0080] To provide interaction with a target object, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor for displaying information to the target object; and a keyboard and pointing device (e.g., a mouse or trackball) through which the target object provides input to the computer. Other types of devices can also be used to provide interaction with the target object; for example, feedback provided to the target object can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the target object can be received in any form (including sound input, voice input, or tactile input).

[0081] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a target computer with a graphical user interface or web browser through which a target object can interact with the implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0082] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0083] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0084] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An image search method, comprising: Perform an image search on the image to be searched to obtain a first search result from a first database; By using the EAST detection model to perform scene text detection on the image to be searched, and using the CTC recognition model to locate and recognize text information at the granular level of text boxes, text information is extracted from the image to be searched. The text information is extracted by obtaining text box information and locating and recognizing it at the granular level of text boxes, and can be directly used to perform text search in the second database. By using the SimBERT model, similar text with similar semantics can be generated based on the extracted text information; Based on the similar text, a text search is performed in the second database to obtain a second search result from the second database; and Based on the first search result and the second search result, search results for the image to be searched are output, wherein outputting search results for the image to be searched includes: comparing the first search result with a threshold; and in response to the number of first search results being less than the threshold, outputting the second search result as the search result for the image to be searched. Specifically, performing a text search in the second database based on the similar text to obtain a second search result from the second database includes: By providing the similar text to the text embedding layer, vectors representing the similar text are generated; By providing the vectors representing similar text to the second database, text search can be performed within the second database; and The similarity between texts is determined by calculating the distance between the vector representing similar text and the vector corresponding to the text stored in the second database, thereby determining the second search result.

2. The image search method according to claim 1, wherein Performing an image search on the image to be searched includes: Extract the feature map of the image to be searched; The extracted feature maps are compared with image features stored in the first database to obtain one or more entries with similar features; and The one or more entries are output as the first search result.

3. The image search method according to claim 1, wherein Extracting text information from the image to be searched includes: By using the efficient and accurate scene text detection model EAST, scene text detection is performed on the image to be searched to obtain text box information related to text boxes containing text information, wherein the text box information includes text box position information; and By using the CTC recognition model to locate and recognize text information at the granular level of text boxes, text recognition is performed on the area within the text box based on the text box information in order to extract the text information.

4. The image search method according to any one of claims 1 to 3, wherein, The first database and the second database are the same database.

5. An image search device, comprising: The image search module is configured to perform an image search on the image to be searched in order to obtain a first search result from a first database; The text extraction module is configured to perform scene text detection on the image to be searched using the EAST detection model, and to locate and recognize text information at the granular level of text boxes using the CTC recognition model, so as to extract text information from the image to be searched. The text information is extracted by obtaining text box information and locating and recognizing it at the granular level of text boxes, and can be directly used to perform text search in the second database. The text search module is configured to generate similar text with similar semantics based on the extracted text information using the SimBERT model; and to perform a text search in the second database based on the similar text to obtain a second search result from the second database; and An output module is configured to output search results for the image to be searched based on the first search result and the second search result, wherein the output module is further configured to: compare the first search result with a threshold; and, in response to the number of first search results being less than the threshold, output the second search result as the search result for the image to be searched. The text search module is further configured as follows: By providing the similar text to the text embedding layer, vectors representing the similar text are generated; By providing the vectors representing similar text to the second database, text search can be performed within the second database; and The similarity between texts is determined by calculating the distance between the vector representing similar text and the vector corresponding to the text stored in the second database, thereby determining the second search result.

6. The image search device according to claim 5, wherein, The image search module is configured as follows: Extract the feature map of the image to be searched; and The extracted feature maps are compared with image features stored in the first database to obtain one or more entries with similar features; and The one or more entries are output as the first search result.

7. The image search device according to claim 5, wherein, The text extraction module is also configured to: By using the efficient and accurate scene text detection model EAST, scene text detection is performed on the image to be searched to obtain text box information related to text boxes containing text information, wherein the text box information includes text box position information; and The text information is located and recognized at the text box level, and the text information is extracted by performing text recognition on the area within the text box based on the text box information.

8. The image search apparatus according to any one of claims 5 to 7, wherein, The first database and the second database are the same database.

9. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 4.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 4.

11. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 4.